# PROXIES.SX Compute Early-access managed inference on provider-owned Apple Silicon Macs. Product snapshot reviewed: 2026-09-11. Each guide has its own source review date. No authenticated rental, inference benchmark or payout test. Canonical overview: https://www.proxies.sx/compute JSON index: https://www.proxies.sx/compute/index.json Saved source responses: https://www.proxies.sx/compute/review.json Renter dashboard: https://compute.proxies.sx/login Required provider signup: https://farmer.proxies.sx/signup Provider compute dashboard: https://farmer.proxies.sx/compute Provider earnings: https://farmer.proxies.sx/earnings Provider payouts: https://farmer.proxies.sx/payouts Agent installation reference: https://agents.proxies.sx/compute/ Provider onboarding: a farmer account is required. Create the account, verify your email, sign the Partner Agreement and create a compute supply key on its Compute page before installing the compute agent. After model readiness, benchmark verification and supplier approval, list the node in the farmer compute dashboard. Manage earnings and payout requests in the same farmer account. Public signup and portal routes reviewed September 11, 2026. The review found no listed machines and zero online nodes. This is a dated snapshot, not live stock. The reviewed catalog contained three Qwen3.8 27B variants. Published prices and wider hardware examples do not establish availability. Current product data: https://api.proxies.sx/v1/peer/compute/catalog, https://api.proxies.sx/v1/peer/compute/tiers, https://api.proxies.sx/v1/peer/compute/marketplace, https://api.proxies.sx/v1/peer/compute/stats Compute uses an active rental and account credentials. Proxy MCP and x402 documentation does not establish compute purchasing support. Paid rentals accrue the admin-approved provider share for that node. The 45% default is a dated example, not a guaranteed share. Renewal is manual at the current quote; accepted purchase terms remain fixed for that intent. Accrual, settlement and completed payouts are separate states. Product facts: https://agents.proxies.sx/compute/facts.json # Rent an Apple Silicon Mac for dedicated inference Canonical: https://www.proxies.sx/compute/rent Audience: Renters Sources reviewed: 2026-09-11 A dedicated endpoint makes sense when you have recurring model traffic and want to reserve a machine for it. Start with the model your application needs, then check whether a listed Mac can serve it. ## What happens between finding a Mac and running a prompt? A balance top-up does not reserve compute. A rental exists after the quoted purchase successfully allocates a machine. Renew manually at a new quote for another 30 days. - Paid term: 30 days - Renewal: Manual - Request limits still apply: No token fee 1. Check ready stock: Find a suitable machine and review its effective price before adding funds. 2. Sign in and fund balance: Use Stripe card checkout or the CoinGate options offered at checkout. 3. Confirm the current quote: The portal checks price and pricing version. A pending operation is unresolved. 4. Use the allocated endpoint: A completed purchase supplies the rental. Confirm model readiness, then send text requests. Documented purchase sequence, not a completed checkout test. Failed allocation reverses its charge to account balance. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## What the rental includes PROXIES.SX Compute is an early-access marketplace for managed inference on provider-owned Apple Silicon Macs. The documented rental gives one customer a machine for 30 days through a chat-completions API. The provider runs the model through MLX. You receive API access; shell access and a remote macOS desktop are outside this product. [Provider documentation](https://agents.proxies.sx/compute/) There is no per-token charge. The accepted price covers a 30-day rental; request limits still apply and the machine has finite throughput. Context length, simultaneous requests and generation length affect how much useful work it can complete. A fixed bill does not establish a latency guarantee. ## Check stock before you plan a migration On September 11, 2026, the [public marketplace API](https://api.proxies.sx/v1/peer/compute/marketplace) returned an empty machine list. The [fleet statistics](https://api.proxies.sx/v1/peer/compute/stats) returned zero online nodes. These are dated observations. Open the [marketplace](https://compute.proxies.sx/market) for current stock. A published tier is a pricing category. A model catalog entry is a supported configuration. Neither means a machine is available to rent. Avoid scheduling a production cutover until a suitable listing exists and you can test the assigned endpoint. ## From an available machine to your first request Check the [marketplace](https://compute.proxies.sx/market) first. When a suitable ready machine exists, sign in, review its price and fund the shared USD account balance through Stripe card checkout or the cryptocurrency choices shown by CoinGate. Review the current machine quote again before confirming the rental. Currency and network choices come from hosted checkout; compute does not use x402. The customer portal submits the expected price and pricing version. A changed quote is rejected before charge. A pending purchase is still unresolved: let the portal reconcile that operation rather than creating another purchase. A successful balance top-up alone does not allocate a Mac. Once allocation completes and the model is ready, use the rental chat endpoint. [Current purchase contract](https://agents.proxies.sx/compute/skill.md) Completed purchases are non-refundable. If checkout cannot allocate the machine, its charge is reversed to account balance. Cancellation preserves access through the paid term. Renew manually for another 30 days at the displayed quote; there is no automatic renewal. Recorded outages pause remaining rental time, while a model download you request uses rental time. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [rental operation reference](https://agents.proxies.sx/compute/skill.md) ## Choose with your actual workload Take a small evaluation set from the work you intend to send. A document assistant should include long documents and questions whose answers are absent from the source. A coding assistant should include representative repository tasks and a way to check its output. Remove sensitive material from this first test. Record the exact model ID and quantization. Then measure answer quality, time to first token and completion time at your expected concurrency. A machine that performs well for one short prompt may queue under several long requests. Use the [benchmark worksheet](https://www.proxies.sx/compute/benchmarks) to make comparisons repeatable. ## Fit the application to the API The customer documentation shows a rental-specific base URL and account API key. Our [compute API guide](https://www.proxies.sx/compute/api) explains discovery and the documented request shape. Compatibility with chat completions does not establish support for every feature in a model SDK. The current documentation rejects tools and `response_format` request fields. Check those limits before choosing this endpoint for an agent that depends on native tool calling or constrained output. Run your agent orchestration and tools on infrastructure you control. The rented endpoint supplies model inference. If the agent also needs a carrier exit IP to collect web data, that is a separate [proxy integration](https://www.proxies.sx/for-ai-agents). ## Work through a rental acceptance sheet Before paying, write down the result your application must achieve. Use the following questions to decide whether a dedicated inference rental fits the job. | Your requirement | Evidence to collect | | --- | --- | | A particular language or extraction task | Answers from your saved evaluation set using the assigned model. | | Interactive responses | First-token and full-response timings at your busiest expected concurrency. | | Long documents | The complete prompt and output budget, plus the endpoint's request-size limits. | | Native tool calling | A supported request contract. The current compute docs reject tools fields. | | Predictable spending | The full 30-day price, renewal choice and any fallback service costs. | | Production continuity | Current stock, outage behavior and a tested application recovery path. | Keep the acceptance sheet with the rental ID and model revision. If you change the model, repeat the quality checks before moving all traffic. A result from a different precision or revision should not silently become the approval record for the replacement. The customer documentation says recorded offline time extends the rental term. That extension preserves paid rental time; it does not process a request while the machine is disconnected. Confirm your node's recorded extension in the dashboard, and keep an application deadline for time-sensitive jobs. [Rental operation documentation](https://compute.proxies.sx/docs) ## Decide what happens when a node is unavailable The customer documentation describes a `node_offline` error while a machine is disconnected. Decide whether your application can wait, return a clear failure or use an independently configured fallback. Set a maximum retry period so a background job cannot wait indefinitely. [Customer documentation](https://compute.proxies.sx/docs) Dedicated tenancy identifies who shares the rented capacity. It does not establish that a machine owner cannot observe data processed on their hardware. Use the [inference privacy guide](https://www.proxies.sx/compute/inference-privacy) to review prompt handling, retention and fallback destinations before using confidential inputs. The published retention policy allows temporary relay storage, and suppliers administer their own machines. The public materials do not establish confidential computing or a compute-specific uptime SLA. Review the [pricing calculation](https://www.proxies.sx/compute/pricing), then use the customer dashboard to inspect an available machine and its current terms. --- # Rent out your Mac for AI inference Canonical: https://www.proxies.sx/compute/providers Audience: Providers Sources reviewed: 2026-09-11 A farmer account is required to provide compute. Create it first, then connect a Mac you already own and can keep available. Your farmer portal is where you manage nodes, listings, earnings and payout requests. ## An installed agent is one step toward an approved listing. You need a verified farmer account, a countersigned agreement, a ready model, a passing benchmark and admin approval. Review the approved price and share in the Compute dashboard. - Reviewed agent version: 0.3.1 - Minimum installed RAM: 24 GB - Free disk for smallest setup: 26 GB 1. Account and agreement: Create your farmer account, verify email and complete the signed and countersigned agreement. 2. Supply key and runtime: Create the key in Farmer → Compute. Install the native Apple Silicon agent on macOS 14+. 3. Readiness and review: The pinned model must serve successfully. Benchmark verification and admin admission are separate checks. 4. Approved terms and listing: Check locked pricing and your share. A proposed price has no effect until admin approves it. Keep the Mac plugged in, awake and online. Starting at login does not prevent lid-close sleep. Listing and free tests do not earn rental income. [Provider documentation](https://agents.proxies.sx/compute/), [Dated product facts](https://agents.proxies.sx/compute/facts.json), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## Create the required farmer account [Create a farmer account](https://farmer.proxies.sx/signup) first. A provider account is mandatory, even when you already own a suitable Mac. It links the machine to you and gives you the portal for node management, rental earnings and payout requests. The renter account serves a different purpose. Verify your email, then complete the [Partner Agreement](https://farmer.proxies.sx/agreement). In the farmer [Compute page](https://farmer.proxies.sx/compute), use **Create compute agent key**. The reviewed portal requests `account:read` and `compute:supply` scopes for this key. Copy it into the installer's hidden prompt; keep it private. The signup client gates protected pages on email verification. [Provider documentation](https://agents.proxies.sx/compute/) An installed node also needs supplier approval before listing. The documentation says the team reviews the measured tier and a signed, countersigned Partner Agreement. Account creation, model readiness and approval are separate steps. ## Check the Mac before downloading The current agent requires Apple Silicon, at least 24 GB of memory, macOS 14 or later, native arm64 Node.js 18 or later and native arm64 Python 3.10 or later. Use a supported Node.js release that meets the minimum. A Node or Python installation running through Rosetta can fail the architecture checks. [Agent source](https://agents.proxies.sx/compute/compute-agent.js), [Node.js releases](https://nodejs.org/en/about/previous-releases) Allow at least 26 GB of free disk space for the smallest setup. Its catalog model is approximately 16 GB; the installer reserves working space as well. Larger configurations need more disk and memory. Check the [model catalog](https://www.proxies.sx/compute/models) and [Apple Silicon hardware guide](https://www.proxies.sx/compute/apple-silicon) for the assigned model. Keep the Mac plugged in, awake and connected during a rental. The service starts at user login; launchd does not prevent lid-close sleep. Restarting macOS or signing out deserves an availability check before you assume that the node is serving again. ## Install the agent and its managed runtime The September 11 review covers agent version 0.3.1. Its installer creates a private Python environment at `~/.proxies-compute/runtime-0.31.3`; you do not need the older manual MLX installation sequence. The source pins MLX LM 0.31.3, MLX 0.31.2 and Transformers 5.4.0. These are reviewed installer versions, not a claim that they are the latest upstream packages. [Installer source](https://agents.proxies.sx/compute/compute-agent.js) ```sh node --version node -p 'process.arch' python3 --version python3 -c 'import platform; print(platform.machine())' curl --fail --show-error --location --output compute-agent.js https://agents.proxies.sx/compute/compute-agent.js # Read the downloaded source before installing it. node compute-agent.js --install ``` The installer asks for your compute supply key without echoing it, saves the key in `~/.proxies-compute/config.json` with mode 600 and copies the agent into `~/.proxies-compute`. Moving the original download should therefore leave the installed copy in place. It creates the `sx.proxies.compute` launchd service. Python's [virtual environment documentation](https://docs.python.org/3/library/venv.html) explains how a private environment separates packages from system Python. These commands document the current installer. No provider installation was performed for this guide. Use [troubleshooting](https://www.proxies.sx/compute/provider-troubleshooting) if a prerequisite or download fails. ## Complete verification and supplier approval Run the installed copy's status command: ```sh node ~/.proxies-compute/compute-agent.js --status ``` Then open [Compute in the farmer portal](https://farmer.proxies.sx/compute). The agent selects a catalog model that fits the machine, downloads its pinned revision and serves it locally. Check each stage before proceeding. | Stage | What to confirm | | --- | --- | | Account | Verified email, Partner Agreement and compute supply key. | | Agent | Installed service and a recent heartbeat in the portal. | | Model | The assigned catalog model reaches ready state. | | Verification | The platform has measured the node with its benchmark. | | Supplier approval | The platform has approved the node and countersigned agreement. | | Pricing | Review the locked rental price, platform percentage and your remaining share in Compute. | | Listing | Enable List for rent in the Compute dashboard. | The provider documentation describes another benchmark after seven days when the node is free. More installed RAM alone does not guarantee a higher rental tier. [Verification and admission](https://agents.proxies.sx/compute/#hardware) ## Manage income and costs in the farmer portal Admin locks the rental price and platform percentage for your node. Your share is the remainder, visible in [Compute](https://farmer.proxies.sx/compute). You can submit one pending price proposal per node with a reason. Approval changes the effective price for future purchases and renewals; rejection leaves it unchanged and shows the review note. Accepted rental terms keep their agreed price and split. [Pricing reference](https://agents.proxies.sx/compute/skill.md) Successful paid rentals accrue your approved share. Listing alone and free test allocations produce no rental earnings. Accrued earnings, settled ledger entries and completed payouts are separate states. The September 11 facts snapshot reports settlement in shadow mode: admin review and settlement are required before payout. Check current status in the dashboard. [Earnings and payment facts](https://agents.proxies.sx/compute/facts.json) Use [Earnings](https://farmer.proxies.sx/earnings) to review income and [Payouts](https://farmer.proxies.sx/payouts) to manage payout requests. The [provider calculator](https://www.proxies.sx/compute/provider-earnings) accepts your approved price and share alongside unrented time, electricity and other costs. Compute nodes serve inference and are excluded from proxy routing. The [Peer Network](https://www.proxies.sx/peers) is the separate bandwidth product. Do not add a second bandwidth revenue stream to a compute forecast without a separately supported arrangement. [Provider product scope](https://agents.proxies.sx/compute/) Before updating or removing an actively rented node, arrange the interruption through the portal or support. Reinstallation restarts the service. The documented uninstall command is `node ~/.proxies-compute/compute-agent.js --uninstall`; model caches can remain afterward. Inspect individual snapshots before deleting files another local application may share. --- # Apple Silicon memory and performance for LLM inference Canonical: https://www.proxies.sx/compute/apple-silicon Audience: Providers Sources reviewed: 2026-09-11 The chip name alone cannot tell you which model a Mac will serve. Start with usable memory, account for the context your workload needs, then measure performance on that configuration. ## Memory requirements are not download sizes. The smallest catalog configuration requires 24 GB of RAM and approximately 16 GB on disk. Model weights, runtime allocations, context state and macOS use different parts of the capacity budget. - Catalog memory requirements: 24 / 48 / 96 GB - Approximate model disk sizes: 16 / 30 / 55 GB - Verified tier and admission: Separate review Switch the chart metric to inspect RAM, disk or context separately. Installed RAM alone does not establish a tier or guarantee rental demand. [Live model catalog](https://api.proxies.sx/v1/peer/compute/catalog), [Provider documentation](https://agents.proxies.sx/compute/) ## What unified memory changes MLX uses Apple Silicon's shared memory architecture. The CPU and GPU can work with arrays in the same memory without copying them between separate CPU and GPU memory pools. This is useful for large model weights, but it does not turn installed RAM into a promise of generation speed. [MLX unified memory documentation](https://ml-explore.github.io/mlx/build/html/usage/unified_memory.html) macOS and other applications also need memory. Model execution adds temporary allocations, and a conversation's stored attention state can grow with context. Read a catalog's memory requirement as a service configuration requirement, not as the size of the model download. ## A first estimate for model weights For a dense model with 27 billion parameters, four bits per parameter gives a raw weight estimate of 13.5 billion bytes, about 12.6 GiB. Eight bits gives 27 billion bytes. Sixteen bits gives 54 billion bytes. These are arithmetic estimates before quantization metadata, non-quantized tensors, working memory and context state. That explains why a 16 GB download and a 24 GB memory requirement can both be correct. It does not prove that every 27B model will run in 24 GB, or that two quantizations will produce equivalent answers. Use the requirements for the exact model repository and the [4-bit versus 8-bit evaluation guide](https://www.proxies.sx/compute/quantization) to test the quality your application needs. ## The reviewed PROXIES.SX requirements The September 11 catalog lists Qwen3.8 27B in three configurations: 4-bit at 24 GB required memory and 16,384 maximum context; 8-bit at 48 GB and 32,768 context; bf16 at 96 GB and 32,768 context. These are API fields, not independent benchmark results. The [model guide](https://www.proxies.sx/compute/models) includes the exact IDs and repositories. The reviewed tiers now correspond to the same three 27B configurations. Each machine still needs benchmark verification and supplier approval. Verify a catalog model and an approved listing together; installed memory alone does not determine the admitted tier. ## Mac mini, Mac Studio or MacBook Pro For an existing Mac mini or Mac Studio, check its configured memory rather than assuming that every machine in the product line has the same capacity. A permanently connected desktop avoids a laptop's lid, battery and travel interruptions. Leave room for airflow and measure power under the workload you will actually serve. A MacBook Pro can be useful if it can stay powered and available for the rental. If it is also your daily work machine, account for competing memory and GPU use. The provider guide recommends desktop Macs and does not recommend the fanless MacBook Air for sustained serving. [Hardware guidance](https://agents.proxies.sx/compute/#hardware) ## Read memory pressure while the workload runs On a Mac you operate, open Activity Monitor's Memory tab while serving a representative workload. Apple describes memory pressure as a combination of free memory, swap rate, wired memory and cached files. Record pressure and swap alongside the response timings. One free-memory number before loading the model misses the conditions under load. [Apple memory monitoring](https://support.apple.com/guide/activity-monitor/view-memory-usage-actmntr1004/mac) Repeat the observation with a longer prompt and the simultaneous requests you intend to serve. If pressure rises while responses slow, reduce the workload and inspect the model's configuration before committing that machine to a rental. This is a diagnostic comparison, not a universal threshold for every Mac. Context state is commonly called the KV cache, short for key-value cache. It stores attention information used during generation. Its size depends on the architecture, runtime, sequence lengths and active requests. Quantizing model weights does not automatically apply the same precision to that cache. The [quantization guide](https://www.proxies.sx/compute/quantization) separates those choices. For renters, these host observations may be unavailable through the managed API. Mark them unknown and evaluate the endpoint's behavior directly. Providers can add host measurements to the [benchmark worksheet](https://www.proxies.sx/compute/benchmarks). ## Separate fitting a model from serving a customer Loading a model proves that it fits that runtime configuration. A useful service also needs acceptable response time, answer quality and recovery from interruption. Test long prompts and several simultaneous requests instead of relying on the fastest short completion. Record the machine's configured RAM, runtime version, model revision and power conditions with the result. The [benchmark guide](https://www.proxies.sx/compute/benchmarks) supplies a worksheet. If your aim is rental income, test existing hardware first and put measured operating costs into the [earnings calculation](https://www.proxies.sx/compute/provider-earnings). --- # Compute model catalog and memory requirements Canonical: https://www.proxies.sx/compute/models Audience: Renters Sources reviewed: 2026-09-11 The public catalog identifies models and their memory and context requirements. Match a pinned model revision to an approved machine listing before planning a rental. ## Compare the supported catalog by the limit you need. All three entries belong to Qwen3.8 27B. Precision, required memory, disk and context differ. The exact model IDs and immutable revisions appear in the tables below. - In the reviewed catalog: 3 entries - Model family size: 27B - Repository revisions: Pinned Data from the September 11 catalog response. Model support is separate from rentable stock; confirm both before choosing a machine. [Live model catalog](https://api.proxies.sx/v1/peer/compute/catalog) ## Catalog checked on September 11, 2026 The [catalog API](https://api.proxies.sx/v1/peer/compute/catalog) returned these three entries. Memory, disk and context figures below reproduce its fields. This table records the API response on the review date. The [saved source responses](https://www.proxies.sx/compute/review.json) preserve the catalog, tier and inventory evidence; check the live API for changes. | Model ID | Quantization | Required memory | Approximate disk | Maximum context | | --- | --- | --- | --- | --- | | qwen3.8-27b-4bit | 4-bit | 24 GB | 16 GB | 16,384 | | qwen3.8-27b-8bit | 8-bit | 48 GB | 30 GB | 32,768 | | qwen3.8-27b-bf16 | BF16 (16-bit) | 96 GB | 55 GB | 32,768 | The corresponding model repositories are [Qwen3.8-27B-4bit](https://huggingface.co/mlx-community/Qwen3.8-27B-4bit), [Qwen3.8-27B-8bit](https://huggingface.co/mlx-community/Qwen3.8-27B-8bit) and [Qwen3.8-27B-bf16](https://huggingface.co/mlx-community/Qwen3.8-27B-bf16). Their public repository metadata was reachable during review. Repository availability does not establish inference quality or rental stock. ## Use the catalog ID in the rental API The short ID, such as `qwen3.8-27b-4bit`, identifies the platform model. The longer `mlx-community/...` name identifies the repository used by the serving runtime. Keep those fields separate in your integration. Copy the assigned model ID from the customer dashboard and use the [documented rental request](https://www.proxies.sx/compute/api). For a reproducible evaluation, also record the repository revision. A repository name can remain unchanged while its files are updated. The September 11 response includes a 40-character revision hash for each entry. The reviewed agent downloads that revision and serves its local snapshot. Preserve the full hash with evaluation results. The hash identifies the intended model files. This review did not verify a running rental against those files. [Pinned download behavior](https://agents.proxies.sx/compute/compute-agent.js) ## Preserve the exact model revision The September 11 catalog contains these revision hashes. Public Hugging Face revision metadata resolved each hash during the review. This checks repository identity and reachability; it does not test generated answers. | Catalog model | Pinned revision | | --- | --- | | qwen3.8-27b-4bit | `3e6447f082e89cc7f0bc6e5441afd38dfce760ff` | | qwen3.8-27b-8bit | `815b83c0df8ffd1d1b5244cf75fd6ef14fca9ef9` | | qwen3.8-27b-bf16 | `6f265714824f3c38d4452baa1628aef3d9b9aae9` | Hugging Face supports selecting a repository snapshot by its full commit hash. Saving the hash helps separate a runtime change from a change in the model files. [Download revision documentation](https://huggingface.co/docs/huggingface_hub/guides/download) Read the model card associated with the revision for intended uses, training context and evaluation limits. Hugging Face's model-card format includes these fields, but their presence and completeness vary by repository. A conversion repository can require checking the original model card as well. [Model card documentation](https://huggingface.co/docs/hub/model-cards) Keep tokenizer and prompt-format changes in the evaluation record too. The same weight revision can produce different behavior when the application changes its system instructions or generation settings. ## Quantization is a choice to evaluate Lower-bit weights reduce the space used to store model parameters. Whether the resulting model is good enough depends on the task. The [quantization comparison guide](https://www.proxies.sx/compute/quantization) explains how to evaluate the variants before paying for more memory or accepting a lower-memory configuration. For extraction, check whether the output contains the requested fields and preserves exact numbers. For code, run the generated code against meaningful checks. For an assistant, include questions where an honest uncertainty response is preferable to a plausible answer. Keep the prompts and scoring rules identical between variants. ## Context is also a capacity limit An advertised maximum context is not an instruction to fill every request to that limit. Your prompt, conversation history and requested answer need an appropriate budget. Longer prompts can take more time before the first token and consume more runtime memory. Build an input-length check into your application. When a document is too large, choose a deliberate strategy such as selecting relevant passages, shortening history or rejecting the request with a clear explanation. Do not let accidental truncation decide what the model reads. ## A catalog entry is not a machine listing The same review found no machines in the [marketplace API](https://api.proxies.sx/v1/peer/compute/marketplace). The current tier descriptions also refer to these 27B configurations. The September 11 response no longer includes the older Ultra tier. Arbitrary repositories, private fine-tunes and clustered Macs are outside the supported release described in the customer documentation. Check both endpoints again when you are ready to rent. Providers can use this table to check memory and disk before following the [setup guide](https://www.proxies.sx/compute/providers). Renters should compare the model evaluation with the [30-day cost](https://www.proxies.sx/compute/pricing). --- # Dedicated inference pricing for a 30-day Mac rental Canonical: https://www.proxies.sx/compute/pricing Audience: Renters Sources reviewed: 2026-09-11 Consider a dedicated endpoint when its model passes your evaluation and you have recurring work for it. Compare the full 30-day rental price with the token bill for that workload. ## A tier default is a reference, not your machine quote. Admin may approve a different price for an individual Mac. Checkout uses the effective machine price and pricing version. Accepted terms remain fixed for that intent; a new renewal needs a current quote. - Compute purchase funding: USD balance - Term, not a calendar month: 30 days - Approved price overrides: Per node Admin defaults returned September 11, 2026, in USD per 30 days. These bars are neither current offers nor an availability claim. [Live tier defaults](https://api.proxies.sx/v1/peer/compute/tiers), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## Dated admin tier defaults The [tiers API](https://api.proxies.sx/v1/peer/compute/tiers) returned these admin defaults on September 11, 2026. A node can have an approved price override. The effective machine price and pricing version in the [marketplace API](https://api.proxies.sx/v1/peer/compute/marketplace) take precedence for a purchase. This dated table is not a checkout quote or evidence of stock. Billing is per 30 days, with no per-token charge; request limits apply. | Tier | Minimum RAM field | Default per 30 days, Sep 11 | Default divided by 720 hours | | --- | --- | --- | --- | | Starter | 24 GB | $249 | $0.346 | | Pro | 48 GB | $449 | $0.624 | | Max | 96 GB | $799 | $1.110 | The September 11 response contains three tiers. The older Ultra tier is absent and is excluded from this current price table. The hourly figures are arithmetic comparisons. PROXIES.SX does not thereby offer hourly billing. The [marketplace](https://compute.proxies.sx/market) determines which specific machines can be rented. At review time, the public machine list was empty. ## Calculate the alternative token bill Use input and output tokens separately because their prices can differ. Include any cache discounts or extra charges from the provider you are comparing. ```text token bill = (input tokens / 1,000,000 × input price) + (output tokens / 1,000,000 × output price) dedicated cost per million output tokens = rental price / delivered output tokens × 1,000,000 ``` For a purely hypothetical API charging $1 per million input tokens and $3 per million output tokens, 50 million input and 10 million output tokens cost $80. An assumed $249 rental costs more under those assumptions; use the approved machine price for your own comparison. At 200 million input and 50 million output tokens, the hypothetical bill is $350. These are worked examples, not a competing provider's quote or a claim that the Mac can deliver those volumes. The second formula allocates the whole rental bill to output for comparison. It is not a price PROXIES.SX charges per token. Input processing still consumes capacity even when the rental bill stays fixed. ## Check whether the workload fits into 30 days Suppose your own test measures 10 output tokens per second. Producing 10 million output tokens would take about 278 hours of decoding alone. This excludes time spent reading prompts, loading models, waiting for work and recovering from errors. Ten tokens per second is an illustrative input, not a measured PROXIES.SX result. Do this calculation using your own [benchmark](https://www.proxies.sx/compute/benchmarks). A theoretical price saving has little value if work misses its deadline or needs a second provider to handle peaks. Include that fallback bill in the comparison. ## Compare equivalent work Two APIs with the same chat-completions request shape can serve different models. Compare answer quality before cost. A cheaper response that fails your extraction checks or breaks the agent's tool selection is not equivalent output. Also compare what you manage. A bare-metal Mac rental can include operating-system access and require you to install the model service. PROXIES.SX's documented product supplies a managed inference endpoint. The [cloud comparison](https://www.proxies.sx/compute/vs-gpu-cloud) explains how to separate those buying decisions. ## Calculate cost per accepted result Use successful work as the denominator when the model's output must pass a check. For example, a $249 rental that delivers 10,000 usable extractions allocates $0.0249 to each accepted extraction. If only 8,000 pass, the figure becomes about $0.0311. These are arithmetic scenarios; neither volume has been measured on this service. Include retries and rejected outputs in the work the node must process. Keep the acceptance rule fixed when comparing services. Otherwise, an API can appear cheaper simply because you allowed it to return less accurate work. For a batch job, compare the expected number of attempts with the successful request rate measured in your [benchmark](https://www.proxies.sx/compute/benchmarks). For interactive work, retain a latency target as well. A 30-day period with plenty of spare capacity can still have a peak hour that fails the application. Completed rentals are non-refundable. A checkout that cannot allocate a machine is automatically reversed to account balance. Cancellation preserves access through the paid term. Read the current checkout terms before paying; this guide has not tested either transaction path. [Published rental terms](https://compute.proxies.sx/docs) ## Confirm the term in the dashboard Renewal is manual. Top up the shared USD balance through Stripe card checkout or the CoinGate currencies and networks offered at checkout, then explicitly purchase another 30 days at the current quote. Check suitable stock before funding a new compute purchase. A balance top-up is not a rental and does not reserve a machine. The portal submits `expectedPriceUsd` and `pricingVersion`. If terms changed, it must obtain a fresh quote before charge. Accepted purchase and renewal intents retain their agreed price and split; future renewals need a new quote. Recorded outages pause the remaining rental time. A customer-requested model download uses rental time. These are published terms, not independently tested payment results. [Current pricing and renewal contract](https://agents.proxies.sx/compute/skill.md) If you are supplying the Mac, the rental price is customer revenue before the provider share and operating costs. Use the separate [provider calculation](https://www.proxies.sx/compute/provider-earnings). --- # Compute API integration for AI agents Canonical: https://www.proxies.sx/compute/api Audience: Developers Sources reviewed: 2026-09-11 An agent can read the public catalog before it has an account. Sending inference requests requires an active rental and its credentials. Keep discovery, purchasing and model execution as separate application steps. ## Frequency, queue capacity and generation time are different limits. The published limit is 60 requests per minute per rental, with one active job per node. A long answer can occupy that node while other requests wait. Measure completed work as well as admitted requests. - Maximum output per request: 4,096 tokens - Request rate per rental: 60 / minute - Active on a node: 1 job 1. Validate the request: Use the assigned model, text messages and an output budget. Tools and response_format are unsupported. 2. Wait for the node: Bound your application queue. The request-rate allowance does not promise immediate execution. 3. Generate and stream: The documented generation deadline is 180 seconds; idle generation fails after 60 seconds. 4. Handle the result: Check completion and errors. Use a deadline and a bounded retry policy for inference. A schematic of documented limits, not a measured latency timeline. Use a read/inference key for model requests and keep purchase permissions separate. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## Discover the product without credentials These read-only endpoints returned public JSON during the September 11 review. None of the following requests creates a rental or sends a payment. ```sh curl --fail --show-error https://api.proxies.sx/v1/peer/compute/catalog curl --fail --show-error https://api.proxies.sx/v1/peer/compute/tiers curl --fail --show-error https://api.proxies.sx/v1/peer/compute/marketplace curl --fail --show-error https://api.proxies.sx/v1/peer/compute/stats ``` The responses contain `models`, `tiers` and `machines` arrays, with fleet state in the stats response. Treat an empty machine list as no visible rental supply. Do not substitute the tier list for inventory. Validate the response shape and handle network errors before an automated workflow makes a decision. This site's [compute index](https://www.proxies.sx/compute/index.json) supplies page links, source endpoints and dated review context as JSON. The [plain-text guide](https://www.proxies.sx/compute/llms.txt) contains the same editorial material as these HTML pages. They are convenience formats; the live product API is the authority for stock and catalog changes. ## Use the rental-specific endpoint The [customer documentation](https://compute.proxies.sx/docs) and its published client example use this base URL: ```text https://api.proxies.sx/v1/peer/compute/rentals/ ``` The customer client example supplies a key with `compute:read` and `compute:infer` scopes in `X-API-Key`. Create it from the renter API page. Setting a generic SDK's bearer-token option alone should not be assumed sufficient. Use the rental ID, key and assigned model shown in your dashboard. The following is a documented request shape, not an authenticated test result: ```sh # Set COMPUTE_RENTAL_ID and COMPUTE_API_KEY from your account. # Replace the example model with the model assigned to your rental. curl --fail-with-body --show-error --max-time 190 \ "https://api.proxies.sx/v1/peer/compute/rentals/${COMPUTE_RENTAL_ID}/chat/completions" \ -H "X-API-Key: ${COMPUTE_API_KEY}" \ -H 'Content-Type: application/json' \ --data '{"model":"qwen3.8-27b-4bit","messages":[{"role":"user","content":"Reply with hello."}],"max_tokens":32,"stream":false}' ``` Keep credentials on your application server or in its secret store. A public browser bundle or a pasted conversation is not an appropriate place for an account key. Consult the customer dashboard for key management. ## Test the features your agent needs Start with a short response, then test streaming and your real prompt lengths. The current customer documentation explicitly rejects tools, stop sequences and `response_format` with HTTP 400. An agent requiring those request fields needs a supported alternative or a different application design. Chat completions do not imply an embeddings endpoint. [Documented request contract](https://compute.proxies.sx/docs) Run the orchestration loop and external tools on your own system. The rented model endpoint returns model output. For document applications, the [RAG hosting guide](https://www.proxies.sx/compute/rag-inference) maps retrieval, source checks and generation to their respective components. A host's larger memory does not automatically make its model reliable at choosing tools or following a schema. ## Bound retries and queue growth The customer client identifies errors including `node_offline`, `model_not_ready` and `timeout`. Use a bounded retry policy with increasing delay, a request deadline and a queue limit. Show the failure to the calling application when those limits are reached. Record latency and the assigned model with each evaluation, while keeping confidential prompt content out of routine logs. ## Enforce the published request limits The September 11 customer client documents the following contract. These are published limits, not results from authenticated API tests. | Field or condition | Documented behavior | | --- | --- | | `messages` | 1–64 messages, each up to 64 KB. | | `max_tokens` | 1–4096; default 1024. | | Other supported fields | `model`, `temperature`, `stream`. | | Streaming | Server-sent events. | | Non-streaming deadline | Up to three minutes for the full answer. | | Rate limit | 60 requests per minute per rental. | | `tools`, stop sequences, `response_format` | Rejected with HTTP 400. | | Expired, offline or unready rental | HTTP 409 with a specific error code. | A separate conservative input-byte budget limits the complete request against the model context. The maximum message count and per-message limit are not permission to send their product as one prompt. A byte is not a token; use the actual model tokenizer when budgeting text and handle the API's rejection explicitly. [Customer request documentation](https://compute.proxies.sx/docs) Bound concurrent work in your application. A rate limit describes admitted request frequency; it does not guarantee that sixty long requests can finish every minute. Use the [benchmark method](https://www.proxies.sx/compute/benchmarks) to establish useful throughput. ## Keep purchase permissions separate from inference An inference key needs `compute:read` and `compute:infer` according to the customer docs. Programmatic checkout and renewal additionally require `compute:purchase` and an `Idempotency-Key` header containing 8–128 letters, digits, dashes or underscores. Reuse the same operation key and request after a timeout or `operation_pending`; a new purchase intent gets a new operation key. [Customer checkout documentation](https://compute.proxies.sx/docs) Keep purchase credentials outside the model's prompt and grant spending permission in application code. A retry of a model request and a retry of a purchase have different consequences. The public JSON discovery calls above remain read-only. This review did not create a rental or exercise the documented purchase API. ## Handle a changed quote differently from a pending purchase The current public contract requires `expectedPriceUsd` and `pricingVersion` from the machine quote. Existing rentals expose their current `renewalQuote` through `GET /v1/peer/compute/rentals/mine`. Renewal is an explicit purchase for another 30 days. This website hands checkout to the customer portal. [Purchase and renewal contract](https://agents.proxies.sx/compute/skill.md) | State | What an integration should do | | --- | --- | | Request timed out or `operation_pending` | Retry the existing intent with its original body and idempotency key. | | Terminal `409 price_changed` | No charge occurs for the stale quote. Refresh the quote, obtain confirmation and create a new intent key. | | Allocation fails | Check the operation result and balance reversal; a funded balance alone does not mean a rental exists. | | Rental committed | Use the returned rental and assigned model. The accepted price and split stay fixed for that intent. | Never replace the body of an unresolved intent with a newer quote. Resolve it first. Treat a confirmed rental operation as the purchase result; a CTA click or top-up redirect is not proof of a purchase. ## Proxy payments are a separate integration The existing [x402 guide](https://www.proxies.sx/x402) documents purchasing proxy resources. It does not establish an x402 compute checkout or compute tools in the [proxy MCP server](https://www.proxies.sx/mcp). Use the compute customer flow for rentals and review its current terms. The [agent proxy guide](https://www.proxies.sx/for-ai-agents) remains relevant when an agent needs network access as well as model inference. --- # Mac compute provider earnings after operating costs Canonical: https://www.proxies.sx/compute/provider-earnings Audience: Providers Sources reviewed: 2026-09-11 The useful number is the amount left after costs and unrented time. Use the calculator with a machine you own, a measured power figure and an occupancy assumption you are prepared to question. ## Earned share, settled balance and paid cash are different amounts. For each successful paid term, accrual follows its accepted price and approved provider share. Free test allocations earn nothing. Use the rental ledger to check earned amounts and the payout record to check received funds. - Approved share and rental price: Editable - Basis for actual accrual: Paid terms - Occupancy assumptions: No guarantee 1. Paid rental accrues your share: Use the accepted price × paid terms × agreed share. Sum terms separately if their pricing differs. 2. Review and settle earnings: The dated snapshot reports shadow settlement mode. Admin review and settlement are required. 3. Track the completed payout: A payout request is not paid cash. Check its final state in the farmer portal. Existing peer ledger entries are already net of the platform share. Do not subtract that percentage a second time. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## Start with the published share Use the approved rental price and provider percentage shown for your node in [Compute](https://farmer.proxies.sx/compute). The September 11 default snapshot is 55% platform and 45% provider, but admin can approve node-specific terms. The calculator starts with $249 and 45% as an editable example. One paid 30-day term at those example terms accrues $112.05 before operating costs, adjustments and taxes. Neither input is a guaranteed offer. [Dated pricing facts](https://agents.proxies.sx/compute/facts.json) On September 11, 2026, the public fleet API reported zero nodes and the marketplace returned no listings. We have no observed paid occupancy or payout history from those endpoints. A forecast that quietly assumes continuous rentals would therefore be unsupported. [Fleet statistics](https://api.proxies.sx/v1/peer/compute/stats), [machine listings](https://api.proxies.sx/v1/peer/compute/marketplace) ## What the calculator means by occupancy Occupancy represents the share of time rented across several 30-day periods. For example, 50% could mean one full period rented and one period without a renter. It is not a statement that customers can buy half a month, or that the platform prorates a rental by the day. At zero occupancy, rental revenue is zero. A powered machine can still consume electricity. Enter an average wall-power measurement that includes the conditions you expect, such as loaded and idle hours. If you power down between rentals, adjust the average accordingly. ```text actual rental accrual = accepted price × paid terms × approved share fraction estimated share per 30 days = assumed price × approved share fraction × occupancy fraction electricity = average watts / 1,000 × 720 hours × price per kWh estimated operating result = estimated share − electricity − other costs ``` The 720 hours represent exactly 30 days. The result is before taxes and before any costs you have not entered. Occupancy prorating is a planning estimate, not billing or ledger logic. If accepted terms differ between rentals, calculate each term separately and sum its accrual. Confirm your approved terms in the dashboard. ## Include costs that are easy to miss If you already own the Mac, distinguish cash spending from the value of keeping it available for other work. A machine you need during the day may be a poor fit for a dedicated rental, even if its electricity bill is low. If you are considering buying hardware, include depreciation or financing and the risk that demand does not arrive. Also consider incremental connectivity, maintenance and time spent handling interruptions. Put a 30-day allowance for those costs in the calculator instead of treating the electricity figure as the entire expense. ## Use scenarios before expanding The calculator shows zero, 25%, 50% and full occupancy so the dependence on demand remains visible. These scenarios are assumptions, not a probability forecast. A break-even occupancy above 100% means the entered price and costs cannot break even under this model. Measure one existing machine first. Record when it is listed, when it is rented, actual power use, credits received and any downtime. Several completed rental periods are more useful for expansion decisions than a single successful install. ## Measure loaded and idle electricity separately If the Mac stays online between requests, measure both serving and idle power at the wall. Multiply each measurement by the hours spent in that state. A power-adapter rating is not the machine's average consumption. For a hypothetical 30-day period with 240 hours at 70 watts and 480 hours at 20 watts, energy use is 26.4 kWh. At an assumed $0.20 per kWh, that is $5.28. The equivalent average is about 36.7 watts for the calculator. These inputs are examples, not measurements of a Mac model or an electricity tariff. This gives a more useful input than assuming the machine draws full-load power for every hour. Add network equipment only when its cost belongs to this activity, and explain the allocation you use. ## Reconcile earned revenue with available payout balance A farmer account is required to supply the machine. Use its [Compute dashboard](https://farmer.proxies.sx/compute) for listing state, [Earnings](https://farmer.proxies.sx/earnings) for income and [Payouts](https://farmer.proxies.sx/payouts) for payout requests. An estimated share is not cleared cash in your bank or wallet. Successful paid rentals accrue the agreed share; free test allocations earn nothing. Accrual, settlement and completed payout are separate states. The September 11 snapshot reports automatic peer settlement in shadow mode, with admin settlement and payout review required. This operational setting can change; check the current dashboard. Existing peer earnings ledger rows already represent the net farmer share, so do not deduct the platform percentage again. [Settlement facts](https://agents.proxies.sx/compute/facts.json), [earnings reference](https://agents.proxies.sx/compute/skill.md) Keep a record of each rental period, adjustments, amounts credited and completed payouts before estimating future returns. Confirm current payout conditions in the portal. This review has no completed payout record and does not assume a payout speed, minimum or fee. The published revenue-share basis is in the [provider documentation](https://agents.proxies.sx/compute/). ## Keep compute and bandwidth earnings separate Compute rents inference capacity. The [Peer Network](https://www.proxies.sx/peers) pays for routed bandwidth under its own terms. The provider documentation says compute nodes are excluded from proxy routing, so do not add a second bandwidth income assumption to a compute forecast without a separately supported arrangement. When the costs look acceptable, follow the [provider setup guide](https://www.proxies.sx/compute/providers). Confirm the machine can serve a [catalog model](https://www.proxies.sx/compute/models) before listing it. --- # Benchmark Apple Silicon inference on your workload Canonical: https://www.proxies.sx/compute/benchmarks Audience: Developers Sources reviewed: 2026-09-11 A tokens-per-second number is incomplete without the model, prompt length and measurement method. Keep those details with the result so another person can repeat the test and understand its limits. ## One local observation is the current performance evidence. The product snapshot reports approximately 19–20 output tokens per second on one M4 Max with 36 GB RAM and Qwen3.8 27B at 4-bit, tested September 10. It does not establish fleet performance. - Reported local hardware: M4 Max · 36 GB - Reported model configuration: 4-bit · 27B - Evidence scope: 1 local test Reported by the product team, not independently measured for this guide. The snapshot omits the prompt set, trial count, timing logs and token-counting method. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [Provider documentation](https://agents.proxies.sx/compute/) ## Begin with a task and a pass condition Choose the workload before choosing the benchmark prompt. For document extraction, define the fields and compare them with known answers. For code, run an appropriate test or compiler. For an agent that selects tools, check the tool name and arguments against the intended action. Use a saved set of prompts and record how you selected them. Include difficult cases and failures you care about. Keep the same set when comparing models. A fast completion that fails the task should remain visible as a failure. ## Record model identity and environment Save the platform model ID, repository name and revision if available, quantization, runtime version, machine memory and chip. State the date and whether the model was already loaded. If the service does not expose a field, mark it unknown rather than guessing from a tier name. The [PROXIES.SX catalog](https://www.proxies.sx/compute/models) exposes model requirements. Its September 11 response includes immutable repository revisions, but no public performance results. The provider guide describes automatic verification, but does not supply enough benchmark details to reproduce its tier assignment independently. [Provider verification description](https://agents.proxies.sx/compute/#hardware) ## What the preliminary local result establishes The [published facts snapshot](https://agents.proxies.sx/compute/facts.json) reports one isolated local test on September 10, 2026: an M4 Max with 36 GB RAM running `qwen3.8-27b-4bit` at approximately 19–20 output tokens per second. That is a reported result for one setup. It is not a fleet comparison, a customer rental test or an availability promise. The snapshot does not include the prompt set, token-counting method, trial count, detailed runtime settings or raw timing logs. Those omissions prevent an independent reproduction from the snapshot alone. Use the methodology below before comparing another result. The same snapshot marks a live paid rental and a real supplier 24-hour soak as unproven. We did not run this local test or validate its speed independently. ## Separate waiting from generation Measure time to first token from sending the request to receiving the first generated content token. State whether that includes network travel, queueing and prompt processing. A timestamp taken only after response headers arrive misses some of the user's waiting time. For output speed, record output tokens and the generation interval you used as the denominator. A client receiving streamed text chunks should not count chunks as tokens. Use returned token usage when available, or document the tokenizer and counting method. Measure whole-request duration as well so a long prompt cannot disappear from the comparison. ## Repeat under the concurrency you need Test one request first, then run the simultaneous workload the application actually produces. Keep a separate record for each concurrency level. Report the number of attempts, successful responses, timeouts and other errors. For latency, include a median and a high percentile only with enough observations to make it meaningful. A p95 calculated from a handful of requests is unstable. Preserve failed attempts in the dataset rather than quietly removing them to improve the average. A separate 2025 comparative study used a 192 GB M2 Ultra Mac Studio and Qwen-2.5 models to examine five local runtimes. It reported different leaders for sustained generation and first-token latency under its settings. Those results explain why both measurements matter; they do not rank the Qwen3.8 rental configurations here. [Apple Silicon runtime study](https://arxiv.org/abs/2511.05502) ## Use consistent latency and throughput definitions NVIDIA's GenAI-Perf documents distinct measurements for first response, request duration and token throughput. That separation is useful when comparing any inference endpoints; its example results are not Apple Silicon measurements. [GenAI-Perf metric definitions](https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/perf_analyzer/genai-perf/README.html) | Measurement | Record in your worksheet | | --- | --- | | Time to first token (TTFT) | Request start to first generated content. State if your tool instead uses the first server event. | | Completion latency | Request start to final response. | | Decode rate | Output tokens divided by the declared generation interval. State whether the first token is excluded. | | Aggregate throughput | Completed useful work across all concurrent requests per wall-clock interval. | | Task pass rate | Accepted results divided by all scored attempts. Keep errors visible separately. | If a stream starts with metadata, timing that first event can understate the wait for readable output. Match definitions before comparing numbers from different tools. Use several prompt lengths, with the same output cap and generation settings for each candidate. Run a separate cold-load case, then a warm sequence. Finish with expected peak concurrency and report how many cases you ran. A single fast request cannot establish sustained capacity or a stable high-percentile latency. ## Download the recording worksheet The [CSV worksheet](https://www.proxies.sx/compute/benchmark-template.csv) is an empty recording format. It contains no measured PROXIES.SX performance. Store a prompt-set identifier instead of confidential prompt text when sharing results. Record at least request start, first content token and completion times; input and output token counts; concurrency; status; and the task's quality result. Add power readings if you are evaluating a provider machine. Mark requests made while loading the model separately from requests made after it is ready. ## Turn the result into a capacity decision Estimate whether the successful throughput can finish the expected work inside its deadline. Allow for bursts and interruption. An overnight classification job may tolerate slower responses that would make an interactive assistant frustrating. Then use the [pricing calculation](https://www.proxies.sx/compute/pricing) with measured useful output. For providers, use wall power and observed rental periods in the [earnings calculator](https://www.proxies.sx/compute/provider-earnings). Publish the test conditions with any future speed claim, including the cases that did not pass. No paid rental or hardware benchmark was run for this guide. It provides a test method and a recording format, not a performance ranking. --- # Mac inference compared with GPU cloud and bare-metal Macs Canonical: https://www.proxies.sx/compute/vs-gpu-cloud Audience: Renters Sources reviewed: 2026-09-11 A managed model endpoint, a remote Mac and a GPU server give you different kinds of access. Decide who needs to install software and what the application must run before you compare the bill. ## Choose access and workload before comparing prices. A managed Mac endpoint supplies catalog text inference. A bare-metal Mac provides operating-system access under its provider terms. A GPU host must match your runtime, memory and software requirements. - PROXIES.SX access: Model API - On a managed rental: No SSH - PROXIES.SX term: 30 days A workload comparison, not a speed or price ranking. Features of a cloud category vary by provider; verify the individual offer. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [AWS EC2 Mac access and allocation](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-mac-instances.html), [MacStadium access and pricing](https://macstadium.com/pricing) ## Match the access model to the job PROXIES.SX documents a managed inference endpoint on a dedicated provider Mac. The customer sends model requests through an API. It is suitable to evaluate when a catalog model can do the job and you want the serving process managed for you. [Compute documentation](https://agents.proxies.sx/compute/) MacStadium describes bare-metal Macs with root access, where the customer can configure the server. AWS documents EC2 Mac instances on Dedicated Hosts with a minimum host allocation period. Those products can support operating-system work that a managed chat endpoint does not expose. [MacStadium](https://macstadium.com/pricing), [AWS EC2 Mac](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-mac-instances.html) | Decision | Managed PROXIES.SX inference | Bare-metal Mac | GPU cloud | | --- | --- | --- | --- | | Access needed | Model API | macOS access under provider terms | VM or container access under provider terms | | Software choice | Current catalog and service API | Customer-managed Mac software | Hardware, image and runtime dependent | | Main evaluation | Model quality and endpoint capacity | Mac configuration and administration | GPU memory, runtime and workload fit | | Cost comparison | Full 30-day rental and fallback costs | Rental plus serving and administration | Instance, storage, networking and serving costs | This table is a buying framework. GPU clouds differ, and a category label does not establish an individual provider's features or price. ## When a GPU server is the closer fit If your application requires a CUDA-dependent package, custom container, fine-tuning job or a model outside the managed catalog, establish that requirement first. A chat-compatible API on a Mac does not grant the ability to run arbitrary training code. Confirm the GPU provider supports your image, memory requirement and storage needs before renting. If the only requirement is inference, compare equivalent models with the same prompts and quality checks. A large-memory Mac and a datacenter GPU can have different bottlenecks. Neither the memory number nor the purchase price proves which service will meet your latency target. ## When a remote Mac is the closer fit For Xcode builds, macOS application tests or remote desktop work, you need operating-system capabilities. Compare dedicated Mac services on that basis. Installing your own model server also gives you more control, alongside responsibility for updates, access control and recovery. MacStadium's reviewed pricing page lists a 24 GB M4 Mac mini at $249 per month. PROXIES.SX's September 11 Starter default is $249 per 30 days, with a 24 GB minimum RAM field; approved node prices can differ. Equal headline prices do not make the offers identical: they provide different access and operational responsibilities. The PROXIES.SX marketplace was empty at review time. [MacStadium pricing](https://macstadium.com/pricing), [compute tiers](https://api.proxies.sx/v1/peer/compute/tiers), [compute stock](https://api.proxies.sx/v1/peer/compute/marketplace) ## Compare the reservation period and recovery work AWS documents a minimum 24-hour allocation for EC2 Mac Dedicated Hosts and one Mac instance per host. That minimum applies even when the task you wanted to run is much shorter. Its operating-system access and storage model also differ from a managed inference endpoint. [EC2 Mac considerations](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-mac-instances.html) PROXIES.SX publishes a 30-day inference rental. Compare the bill over the time you must reserve, then include application hosting, persistent storage and recovery work where they apply. A short benchmark run and a continuously available endpoint can lead to different buying decisions. Write down who restores service after a failure. On a server you administer, that may include rebuilding the runtime, restoring weights and restarting the endpoint. With managed inference, your application still needs to handle failed requests and any fallback destination. See the [rental behavior](https://www.proxies.sx/compute/rent) and [privacy review](https://www.proxies.sx/compute/inference-privacy). For CUDA-dependent work, first establish that the software can run on the candidate GPU host. For Xcode or macOS testing, establish the required Mac operating-system access. Then compare prices for the resulting shortlist. This avoids using a model-only API price as a substitute for a server that your workload actually needs. ## If you want to supply hardware Provider programs also need separate comparisons. Vast.ai's hosting documentation addresses a GPU rental host's operational setup. PROXIES.SX's provider flow serves catalog models on Apple Silicon through its agent. Being able to run one program does not demonstrate eligibility for the other. [Vast.ai hosting](https://docs.vast.ai/host/hosting-overview) Compare the supported hardware, workload access, payment basis, demand and operating obligations. A revenue-share percentage by itself says little about expected income. For PROXIES.SX, begin with the [provider requirements](https://www.proxies.sx/compute/providers) and the [cost scenarios](https://www.proxies.sx/compute/provider-earnings). For a rental decision, put the candidate services through the same [benchmark method](https://www.proxies.sx/compute/benchmarks), then compare the [total bill](https://www.proxies.sx/compute/pricing). --- # Troubleshoot the MLX compute agent on your Mac Canonical: https://www.proxies.sx/compute/provider-troubleshooting Audience: Providers Sources reviewed: 2026-09-11 A loaded background service can still have no model ready to serve. Check the agent, Python environment and local model endpoint separately to find where setup stops. ## Find the first failed stage before reinstalling. An online agent, a serving model, a passed benchmark and an approved listing are separate states. Check runtime logs and the farmer dashboard together before changing a working installation. - Documented heartbeat interval: 60 seconds - Missing-heartbeat offline threshold: 3 minutes - Reviewed agent release: 0.3.1 1. Agent cannot start: Check native arm64 Node 18+, Python 3.10+, macOS 14+ and runtime logs. 2. Model cannot become ready: Check disk space, download progress, the pinned revision and startup generation errors. 3. Node is paused or offline: Check power, thermal state, sleep and heartbeat. launchd does not keep a closed laptop awake. 4. Ready but not listed: Review benchmark status, the countersigned agreement and separate admin approval. A diagnostic path based on the current agent reference, not a report of faults in your node. Redact keys from logs before requesting support. [Provider documentation](https://agents.proxies.sx/compute/), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## Check the installed agent version first Use the installed copy when it exists, so you inspect the version launchd actually starts. The September 11 source review covers agent 0.3.1. Its installer copies the script to `~/.proxies-compute` and uses a pinned Python runtime. [Compute agent source](https://agents.proxies.sx/compute/compute-agent.js) ```sh node ~/.proxies-compute/compute-agent.js --status ``` If that file does not exist, inspect the original download and follow the [current provider setup](https://www.proxies.sx/compute/providers). A farmer account, verified email and compute supply key are required. A renter inference key is a different credential. A status report distinguishes the service from the model endpoint. A loaded launchd service can still be downloading a model or waiting for approval. Read the node's state in the [farmer Compute dashboard](https://farmer.proxies.sx/compute) alongside the local result. ## Diagnose architecture and Python runtime errors The agent checks for native arm64 Node and Python. These read-only commands report the executables you are using: ```sh node -p 'process.arch' python3 -c 'import platform, sys; print(platform.machine()); print(sys.version)' sw_vers -productVersion ``` The documented minimums are macOS 14, Node.js 18, Python 3.10 and 24 GB of memory on Apple Silicon. Use a currently supported Node release that meets those checks. [Release schedule](https://nodejs.org/en/about/previous-releases) Agent 0.3.1 installs packages under `~/.proxies-compute/runtime-0.31.3`. A terminal reporting `mlx_lm.server: command not found` does not establish that this private runtime is broken. Inspect the installer's Python directly: ```sh ~/.proxies-compute/runtime-0.31.3/bin/python3 -m pip show mlx-lm mlx transformers ``` That path is version-specific. For a later agent, read its runtime path before copying the command. Avoid replacing the pinned packages with a global `pip install` as a first repair. Python environments have their own packages and executable paths. [Python environment behavior](https://docs.python.org/3/library/venv.html) ## Check free disk and the pinned model download The smallest installation requires at least 26 GB free. The agent checks the volume that holds the Hugging Face cache, which can differ from your home volume when `HF_HOME` or `HF_HUB_CACHE` is configured. Larger models need more headroom. [Agent disk checks](https://agents.proxies.sx/compute/compute-agent.js) A repository directory can exist while its requested revision is incomplete. The source checks the revision snapshot for configuration and weight files before serving it. Hugging Face supports downloads by full commit hash; compare the expected hash with the [catalog revision](https://www.proxies.sx/compute/models). [Hub download documentation](https://huggingface.co/docs/huggingface_hub/guides/download) Preserve the first download error and check connectivity to that repository. Do not delete the entire shared cache as a routine fix. Other applications may use its files. ## Check the local model endpoint The default server listens on loopback port 8734. Substitute your configured `MLX_PORT` if different: ```sh curl --fail --show-error --max-time 5 http://127.0.0.1:8734/v1/models ``` A successful response means a local server answered. Compare its model with the assigned snapshot; another process on that port does not establish readiness. Connection refusal means nothing accepted that connection at that moment. It does not diagnose an account problem by itself. Avoid starting a second model server on the agent's port. The agent manages its serving process. If you reinstall to repair the runtime, allow for a restart and coordinate an actively rented node's interruption first. ## Match the dashboard state to the next check | Observation | Next check | | --- | --- | | No registered node | Verify the farmer account and compute supply key, then read registration errors. | | Downloading or loading | Check the first error, cache volume space, exact model revision and memory. | | Ready locally, offline in the portal | Check internet access and heartbeat errors. The documented offline threshold is three minutes without a heartbeat. | | Paused for battery or thermal conditions | Restore power or cooling and inspect the reported pause reason. | | Verified but cannot list | Check supplier approval and the signed, countersigned Partner Agreement. | | Repeated server restarts | Record the initial runtime error before restarting again. The agent limits restart attempts. | These are diagnostic steps derived from the [published source](https://agents.proxies.sx/compute/compute-agent.js) and [provider admission documentation](https://agents.proxies.sx/compute/). They are not results from a completed provider installation. ## Collect useful evidence for support The default log is `~/Library/Logs/proxies-compute.log`. Share a short redacted excerpt around the first failure with the time, agent version, macOS version, memory and model ID. Keep the API key, private device identifiers and customer prompts out of public reports. The current executable constant sets a 175-second generation deadline and a 60-second gap-without-output deadline. An older comment still mentions 300 seconds; the executable constant is the basis for this guide. The customer documentation allows up to three minutes for a non-streaming response. These are recovery limits, not measured latency or capacity. [Agent source](https://agents.proxies.sx/compute/compute-agent.js), [customer API documentation](https://compute.proxies.sx/docs) Use the [benchmark worksheet](https://www.proxies.sx/compute/benchmarks) to record prompt length, output length and concurrency when a specific workload times out. When the node is ready and approved, list it in the farmer portal and track income through your account's [earnings and cost calculation](https://www.proxies.sx/compute/provider-earnings). --- # Check data privacy before renting an LLM inference endpoint Canonical: https://www.proxies.sx/compute/inference-privacy Audience: Renters Sources reviewed: 2026-09-11 A dedicated machine reserves capacity for a customer. To assess privacy, also establish where prompts go, who operates the hardware and what happens to request data after generation. ## Trace the prompt beyond your application. Prompts pass through the platform relay and a supplier-administered Mac. Dedicated tenancy limits who rents that capacity; it does not make the supplier host confidential computing. - Normal completed-content relay scrub: ~1 hour - Job-record expiry, with cleanup lag: ~24 hours - Machine administration: Supplier-owned 1. Your application: Select permitted text. Check your own logs, stored conversations and fallback destinations. 2. Platform relay: Temporary relay storage supports request delivery. Published cleanup windows are not zero retention. 3. Supplier Mac: The supplier administers the hardware. The agent does not intentionally log prompts. 4. Response and saved copies: The answer returns to your application. Its retention policy is separate from relay cleanup. Documented data path. Backend deletion and supplier-host behavior were not independently audited. No confidential-computing or compliance certification is established here. [Dated product facts](https://agents.proxies.sx/compute/facts.json), [Compute operation reference](https://agents.proxies.sx/compute/skill.md) ## Trace the request to the provider's Mac The published PROXIES.SX flow sends a customer's request through the compute API, delivers inference jobs to a provider agent and runs the model on that provider's Apple Silicon Mac. The agent posts generated results back to the platform. This describes the application flow visible in public materials; it is not a complete audit of backend storage or infrastructure. [Provider documentation](https://agents.proxies.sx/compute/), [agent source](https://agents.proxies.sx/compute/compute-agent.js) ```text Your application -> PROXIES.SX compute API and job delivery -> provider Mac and local MLX model server -> result returned through the platform -> your application ``` A document excerpt included in a prompt becomes part of that request. Keeping the full document database on your own server does not keep the selected excerpts there. The same applies to conversation history, tool results and identifiers added by your application. ## Separate dedicated capacity from data protection Single-customer rental describes allocation. The reviewed public materials do not establish that the hardware owner is technically unable to inspect prompts or outputs. They do not establish encrypted processing or independently attested execution. The current customer documentation explicitly describes temporary relay retention. HTTPS protects the network connection to the endpoint. Binding the local model server to loopback limits its network exposure. Neither establishes who can inspect data inside the machines processing it. | Property | Evidence to request before relying on it | | --- | --- | | Request retention | How the published relay windows apply to failed jobs, operational logs and backups. | | Operator access | Which platform and node operators can access requests or stored outputs. | | Training use | Whether prompts and completions can be used for model training or evaluation. | | Location | Where the assigned machine and any additional data processing take place. | | Deletion | How deletion is requested and what it covers, including backups. | | Incident handling | A contact and process for suspected data exposure. | These questions concern the coverage and enforcement of the published policy. A stated relay cleanup window does not answer every question about application logs, backups or access inside the supplier machine. ## Read the published relay retention windows The published operation reference says the agent does not intentionally log prompts. The supplier still administers the Mac, so this statement alone cannot establish what the host retains. The customer documentation describes relay storage used to deliver prompts and responses, with completed-job content normally scrubbed after one hour and job records expiring after 24 hours, subject to database cleanup timing. [Provider privacy description](https://agents.proxies.sx/compute/), [customer retention description](https://compute.proxies.sx/docs) Treat those as published operating terms. This review has not inspected the backend database, a supplier machine or deletion logs. The one-hour statement describes normal cleanup of completed jobs; it should not be rewritten as a guarantee that all copies disappear exactly sixty minutes after a request. The documented service permits temporary retention. Trace your own application copies too. A saved conversation, an error report or a tracing service can retain text after the inference relay has removed it. Decide what each system needs and set its retention accordingly. A confidential workflow needs evidence covering the whole request path. ## Keep credentials and unnecessary data out of prompts Call the rental API from your application backend and keep the account key there. A browser-delivered bundle can expose embedded credentials to its users. The [API guide](https://www.proxies.sx/compute/api) explains the documented header and rental-specific endpoint. Prepare a sample request and inspect everything it contains before allowing real application traffic. Remove account secrets, unrelated conversation history and document sections that the task does not need. If a task only needs an order status, consider supplying the status and an internal reference instead of the customer's full record. Record operational metrics such as duration and status separately from prompt text. Debug logging is an application choice: check your SDK, proxy, error tracker and tracing configuration as well as the inference service. ## Check retrieval and fallback paths For a [RAG application](https://www.proxies.sx/compute/rag-inference), enforce document access before selecting excerpts for the model. A user must not receive another user's document through retrieval, even if the final answer is well written. Keeping access checks in application code makes that boundary independent of the model's response. A fallback can introduce another processor. If an offline node causes your code to send the same prompt to a second model service, review that destination under the same data-handling requirements. Do not silently broaden where confidential requests are sent. External documents can also contain instructions aimed at the model. OWASP describes this as indirect prompt injection and notes that RAG does not remove the vulnerability. Treat retrieved content as input data, and enforce tool permissions in application code. This reduces potential impact; a prompt instruction alone does not guarantee containment. [OWASP prompt injection guidance](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) ## Choose inputs that match the evidence Public product descriptions, synthetic test cases and approved non-sensitive documents can help evaluate a model while data-handling questions are resolved. If your workload depends on a specific privacy guarantee, obtain supporting terms and technical evidence before sending it. The [rental guide](https://www.proxies.sx/compute/rent) covers availability and workload fit. For the service's current setup references, use the [agent documentation hub](https://agents.proxies.sx/) and [compute provider documentation](https://agents.proxies.sx/compute/). This review did not inspect a rented machine or backend logs. --- # Choose between 4-bit, 8-bit and bf16 LLM inference Canonical: https://www.proxies.sx/compute/quantization Audience: Renters Sources reviewed: 2026-09-11 Fewer bits can reduce model weight storage, but the useful choice depends on the answers your application needs. Compare quantizations with the same tasks before paying for a larger machine. ## Compare capacity without assuming a quality ranking. BF16 is a 16-bit format. The catalog pairs its three weight formats with specific memory and context limits; those configurations do not prove a universal speed or quality hierarchy. - Reviewed weight formats: 4 / 8 / 16 bits - For meaningful quality comparisons: Same prompts - KV cache and working memory: Separate state Published catalog requirements, not benchmark results. The 4-bit context limit belongs to this configuration; it is not a general law of quantization. [Live model catalog](https://api.proxies.sx/v1/peer/compute/catalog), [MLX LM quantized inference](https://github.com/ml-explore/mlx-lm) ## What changes when model weights use fewer bits Quantization represents model weights with fewer bits. That reduces the space used by those weights and changes their numerical representation. The total runtime memory still includes other allocations, such as context state and temporary work. MLX LM supports model conversion and quantized inference on Apple Silicon. [MLX LM documentation](https://github.com/ml-explore/mlx-lm) A 4-bit label does not specify total serving memory. Check the quantization method, model architecture, runtime and context as well. The [Apple Silicon memory guide](https://www.proxies.sx/compute/apple-silicon) explains the difference between a raw weight estimate, download size and usable serving capacity. ## What research establishes Dettmers and Zettlemoyer's ICML 2023 study compared the relationship between model size, weight precision and zero-shot accuracy across several model families. Its results support evaluating 4-bit models when memory is constrained. The experiments covered BLOOM, OPT, NeoX/Pythia and GPT-2; they do not establish the quality or speed of the Qwen3.8 variants in the PROXIES.SX catalog. [The case for 4-bit precision](https://proceedings.mlr.press/v202/dettmers23a.html) Do not turn a general research finding into a fixed percentage of quality loss for your application. A model can pass a broad test and still fail a particular extraction format, language or coding task. A 2025 profiling study by Benazir and Lin tested several Apple Silicon and NVIDIA configurations and examined dequantization overhead alongside memory and compute limits. It reports that lower weight precision does not consistently produce faster inference across hardware. Its configurations do not establish performance for this rental catalog. [Apple Silicon quantization profiling](https://arxiv.org/abs/2508.08531) ## What the compute catalog establishes The [saved September 11 catalog](https://www.proxies.sx/compute/review.json) lists three Qwen3.8 27B configurations. Its 4-bit entry requires 24 GB and publishes a 16,384 context limit; 8-bit requires 48 GB and bf16 requires 96 GB, both with 32,768 context. These are service fields, not measured comparisons. Exact repository names and download sizes are in the [model catalog guide](https://www.proxies.sx/compute/models). The context limits belong to these published configurations; 4-bit quantization does not inherently halve context capacity. BF16 stores values in a 16-bit floating-point format. Its catalog label describes the weights, not the precision of every runtime calculation. ## Separate weight precision from context settings MLX LM documents prompt caching, a rotating KV cache and configurable prefill steps. These affect how the runtime handles context and memory. They are separate from converting model weights to a lower precision. [MLX LM context features](https://github.com/ml-explore/mlx-lm#long-prompts-and-generations) A managed rental does not give you direct control over every runtime option. Compare the configuration the endpoint actually exposes. Keep model revision, input text, output cap and generation settings constant when evaluating precision; record an unavailable setting as unknown. Suppose an extraction test fails an exact number. Save that case and check whether the same failure occurs with each precision before attributing it to quantization. A prompt-format error, missing evidence or tokenizer difference can affect the result as well. For [RAG](https://www.proxies.sx/compute/rag-inference), use the same retrieved passages in each run. Choose the lowest-cost configuration that passes your task and capacity requirements. That is an application decision from your measurements, not a universal ranking of 4-bit, 8-bit and bf16 models. Use [cost per accepted result](https://www.proxies.sx/compute/pricing#calculate-cost-per-accepted-result) when failed outputs require retries or review. ## Compare quality before throughput Build a small saved evaluation set from your real tasks. Include exact-answer cases, long inputs and cases where the correct behavior is to say the available information is insufficient. Establish the pass condition before seeing model outputs. | Task | Check beyond whether the answer sounds plausible | | --- | --- | | Structured extraction | Required fields, source-backed values and handling of missing data. | | Code generation | Compilation, relevant tests and behavior on invalid inputs. | | Document answers | Correct source references, unsupported statements and refusal when evidence is absent. | | Classification | Per-category errors, especially the mistakes that cost the most to correct. | Keep the base model, prompt set, chat template and sampling settings as comparable as the service allows. Record any field you cannot inspect. Compare the same context length first; otherwise a change in retrieved evidence can be mistaken for a quantization effect. For a [RAG workload](https://www.proxies.sx/compute/rag-inference), test retrieval separately. More precise model weights cannot recover a document your retriever never supplied. ## Measure latency on the actual configuration Run timing tests only after defining acceptable task quality. Record first-token latency, completion time, errors and concurrency with the [benchmark worksheet](https://www.proxies.sx/compute/benchmarks). A shorter weight representation does not prove a particular speedup on a given Mac and runtime. If the 4-bit and 8-bit endpoints run on different chips, describe the result as a comparison of those complete configurations. It cannot isolate the effect of quantization. Keep output lengths comparable so a shorter answer does not appear faster simply because it contains less work. ## Choose the smallest configuration that passes Start with a configuration whose documented memory and context limits fit the task. If it passes your quality and latency requirements, a larger allocation needs a concrete benefit to justify its [rental cost](https://www.proxies.sx/compute/pricing). If it fails, inspect the failed examples before assuming precision is the cause. A different prompt, better retrieval or a different base model may address the failure. Those options require their own test; the current managed service is limited to its published catalog and assigned model. Check [current machine listings](https://compute.proxies.sx/market) before scheduling any comparison. No rented endpoint or quantization benchmark was tested for this guide. --- # Plan RAG inference hosting on a dedicated Mac endpoint Canonical: https://www.proxies.sx/compute/rag-inference Audience: Developers Sources reviewed: 2026-09-11 A RAG application retrieves relevant source material and supplies it to a language model. A managed chat endpoint can provide the generation step while your application retains the document index and retrieval logic. ## Reserve context for the answer as well as the evidence. Your application handles ingestion, retrieval and document permissions. The rented endpoint generates text from the excerpts you send. Budget instructions, history, retrieved text and output together. - Catalog context limits: 16,384 / 32,768 - Maximum requested output: 4,096 tokens - Retrieval and access checks: Application-owned An editable planning example, not a tokenizer or an acceptance guarantee. The API also enforces a conservative input-byte budget. Retrieved text remains untrusted input. [Live model catalog](https://api.proxies.sx/v1/peer/compute/catalog), [Dated product facts](https://agents.proxies.sx/compute/facts.json) ## Separate retrieval from generation The original retrieval-augmented generation research combined a retrieval component with a language generator for knowledge-intensive tasks. In an application design, it is useful to examine those stages separately: finding suitable evidence and generating an answer from it can fail for different reasons. [Lewis et al., Retrieval-Augmented Generation](https://arxiv.org/abs/2005.11401) For PROXIES.SX, the documented rental exposes model inference through a chat-completions API. It does not give the renter a shell for installing a vector database. The reviewed customer materials do not establish a managed embeddings, indexing or reranking endpoint. Plan those components separately. [Compute rental documentation](https://compute.proxies.sx/docs) ```text Documents -> ingestion and index you operate Question + user permissions -> retrieval you operate Question + permitted excerpts -> rented chat inference endpoint Answer + source checks -> your application response ``` This is a proposed application architecture, not a deployed PROXIES.SX RAG service. Start with the [compute API integration](https://www.proxies.sx/compute/api) to establish the supported request shape. ## Decide what stays in your application Keep document ingestion, access permissions, retrieval and source identifiers in your backend. If you use an embedding model, confirm its license and hosting requirements separately. A chat-compatible endpoint does not imply an embeddings-compatible endpoint. Give each retrieved excerpt a stable document identifier and a location such as a section or page. Preserve the mapping in your application so a response can link to the source the user is allowed to read. Avoid asking the language model to invent a source URL. For changing documents, decide how updates and deletions reach the index. An accurate generator can still answer from an outdated excerpt. Store the document version or retrieval timestamp with an evaluation case so you can distinguish stale evidence from a model mistake. ## Budget the complete context Count the question, instructions, conversation history, excerpts and planned output against the assigned configuration's context limit. The [reviewed catalog](https://www.proxies.sx/compute/models) publishes 16,384 for the 4-bit entry and 32,768 for the other two. Confirm the current endpoint's behavior before relying on the full limit. A planning example for a 16,384-token limit could allocate 12,000 input tokens and 2,000 output tokens, leaving 2,384 for additional formatting and margin. This is arithmetic for planning, not a measured capacity recommendation. Use the model's tokenizer or documented usage fields to check actual counts. Select excerpts by relevance and remove duplicates before expanding the prompt. Extra material has a cost in prompt processing and can make it harder to diagnose why an answer failed. Compare [quantizations](https://www.proxies.sx/compute/quantization) on the same evidence set. ## Check citations against the retrieved evidence Ask for source identifiers beside factual claims, then validate that each identifier came from the retrieved set. A valid identifier is only the first check: the cited passage must also support the claim. Create evaluation questions with known answers and include questions whose answers are absent from the indexed material. Record retrieval success, answer correctness, source support and unsupported claims separately. If the right passage was never retrieved, increasing inference capacity does not fix that failure. The [benchmark worksheet](https://www.proxies.sx/compute/benchmarks) can record model timing. Add your retrieval duration and source-check result alongside it to measure what the reader actually experiences. ## Test evidence position and missing answers Liu and colleagues' 2023 study found that performance on its retrieval and question-answering tasks could change when relevant evidence moved within a long context. The tested models do not establish a failure rate for this catalog. The finding supports adding position changes to your own evaluation. [Lost in the Middle](https://arxiv.org/abs/2307.03172) Take a question with one known supporting passage. Test the passage near the beginning, middle and end of the supplied excerpts, then remove it entirely. Keep the question and answer check fixed. This helps distinguish a model that uses supplied evidence from one that produces a plausible answer without it. | Evaluation case | Useful outcome to inspect | | --- | --- | | Correct passage retrieved | Answer matches the passage and cites its real identifier. | | Passage missing | The application reports that the indexed evidence is insufficient. | | Two document versions disagree | The answer uses the intended version or explains the conflict. | | Passage exists but user lacks access | Retrieval excludes it before inference. | For changing material, preserve the source version with the test case. Rerun retrieval checks after indexing changes and generation checks after model or prompt changes. This keeps a faster endpoint from masking an outdated or incorrectly permissioned index. The [model catalog](https://www.proxies.sx/compute/models) now exposes revision hashes for repeatable generation comparisons. ## Enforce permissions before model access Filter retrieval by the authenticated user's permissions before sending excerpts to the model. Do not rely on a system prompt to hide documents that the user should never have received. Treat retrieved text as external data. Instructions embedded in a document can try to redirect the model, and RAG does not eliminate that risk. Keep tool authorization in application code and test with adversarial documents as well as normal questions. [OWASP indirect prompt injection guidance](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) Selected excerpts leave your application when sent to the inference service. Review the [inference privacy questions](https://www.proxies.sx/compute/inference-privacy), including logging and any fallback provider, before using confidential documents. ## Test demand before a 30-day commitment Measure retrieval time, first-token latency and completion time under the simultaneous traffic you expect. If a busy period queues requests, a fast single-user demonstration is insufficient evidence for launch. Compare the [full rental cost](https://www.proxies.sx/compute/pricing) with the amount of successful work completed, and include the index, embedding service and application hosting in your own budget. Check [current stock](https://compute.proxies.sx/market) before depending on a rental: the saved September 11 snapshot contained no listed machines. Use the [agent documentation hub](https://agents.proxies.sx/) for the wider service references and [compute provider reference](https://agents.proxies.sx/compute/) for how the model is served. This guide specifies an architecture and evaluation method; it does not report a tested RAG deployment.