Choose where the work happens.
Run inference in your own infrastructure, a private cloud or an approved hosted service. Define the location, access and logging controls that your data requires.
AiAx offers locally hosted and cloud deployments of open-weight AI, including EU-hosted options. Start with straightforward cloud access, or choose more control over location and retention when your work requires it.
The model’s trained parameters are available to download and run under its licence. That gives you options beyond a single hosted API. It does not necessarily include the training data or all the freedoms of open-source software.
Run inference in your own infrastructure, a private cloud or an approved hosted service. Define the location, access and logging controls that your data requires.
Separate the model from the company hosting it. Where licensing and technical compatibility allow, move the same weights between providers or retain a tested version.
Evaluate and fine-tune a suitable model for your terminology, document types or recurring tasks. Compare it against a baseline before putting it to work.
Explore fine-tuning →Steady workloads can make dedicated infrastructure economical. Compare the full cost of hardware, hosting, operations and evaluation against a managed API. Self-hosting is not automatically cheaper.
Fine-tuning adapts a pretrained model using carefully selected examples of the work and the responses you want. It updates model parameters or a smaller adapter, rather than training a new model from scratch.
It is most useful for repeated tasks with a clear standard of quality.
Train on approved examples of classifications, extracted fields and report formats. The aim is fewer corrections on the tasks your team handles every day.
Teach task-specific terminology, tone and conventions. For example, turn service observations into the categories and action descriptions your team uses.
Specialisation can make a smaller model suitable for a narrow workload. Measure quality, latency and total cost before deciding whether it beats a general-purpose model.
Curate approved examples → Tune a compatible model → Test on separate, unseen cases.
Start with a prompting baseline. Methods such as LoRA can reduce training requirements. Deploy only after the tuned version meets your quality checks.Consider an agent that classifies incoming service reports and extracts the work to be done. The vocabulary is familiar, the task repeats and the output has a clear structure.
Some teams need a defined data boundary. Others want a capable model with a simple cloud setup. Your requirements set the level of control.
Run models in your own environment. Keep inference close to internal systems and apply your organisation’s access and network controls.
Agree the hardware, operations and support responsibilities for your deployment. Local hosting still needs explicit policies for logs, backups and retention.Choose straightforward cloud access when you have no special requirements for data location or retention. Focus on model quality, speed and cost. Dedicated hosting can be selected for workloads that need it. Zero retention is configured only where it is a customer requirement.
Standard cloud services use the provider’s data-handling terms. If you need zero retention or a specific region, select an endpoint and configuration that meet those requirements.Choose local hosting or an EU-hosted deployment, with access and retention requirements defined for your workload. Zero retention is provided where the customer requires it and the selected setup supports it.
GDPR compliance also depends on lawful processing, data-processing agreements, security, retention and handling people’s rights. EU hosting alone is not a compliance guarantee.
Confirm the full data path, including model endpoints, backups, support access and subprocessors. Any transfers outside the EEA need an appropriate legal basis for the transfer and applicable safeguards.
Explore security & data handling →We provide zero retention where the customer requires it. It is not the default for every deployment.
Zero-retention inference is an endpoint-specific commitment for handling prompts and outputs. Its scope, exceptions and eligible features must be checked against the provider agreement or the local logging configuration. It is not a blanket promise that no data is ever retained. Workspace history, files and AiAxBrain memory have their own retention rules, so your team can keep the work it needs.
Open weights alone do not guarantee zero retention. Hosting configuration and service terms determine it.
It gives you another deployment choice. An open-weight model can itself have frontier-level capabilities. In our routing diagram, “frontier API” describes the separately approved provider route. Compare quality, cost and data handling for each actual endpoint.
Check the exact model licence. Permissions for commercial use, modification, redistribution and hosting vary between releases.
The Router page explains how task requirements, context and endpoint policy inform model selection. OpenWeight explains why you may want this deployment option in the first place.
Explore AiAx Router →Bring a recurring task and your data requirements. Use them to compare the deployment options.