Which AI Model Nearshore Developers Actually Trust

Transformative software solutions aren't about rebuilding your team. First Factory's nearshore talent delivers long-term value and a true partnership.

August 25, 2026

Table of contents

Key Takeaways

  • 87% of First Factory engineers use AI tools daily or multiple times a day, and 85% require a thorough review before that code reaches production. Only 6% ship on a light review, and the standard does not loosen with heavier use or greater seniority (First Factory internal survey, August 2026).
  • 70% of our engineers name Claude as the model they trust most for production code, and the same share would recommend it to a client standardizing on one model. Claude leads in all six task categories and all six industry categories we measured.
  • That diverges sharply from global usage charts, where ChatGPT leads work coding at 28% against Claude’s chatbot at 7% and Gemini at 8%. It aligns with the loyalty data in the same study: Claude Code hit 18% workplace adoption, a sixfold jump from roughly 3% in mid-2025, with 91% customer satisfaction and a +54 Net Promoter Score (JetBrains, April 2026).
  • Cost ranked sixth out of eight selection criteria among our engineers, behind accuracy at 84% and reasoning on complex problems at 76%. Speed ranked seventh. Engineers choosing a model for client production work optimize being right, not cheap.
  • Global surveys report 29% developer trust in AI output and 46% active distrust, down from 40% in 2024 (Stack Overflow, Feb 2026). Roughly 60% of enterprise leaders across 24 countries say they are balancing AI rollout against risk effectively, with worker sentiment splitting near 13% enthusiastic and 55% open but unconvinced (Deloitte, "State of AI in the Enterprise," 2026 edition).

What "trust" means when a developer rates an AI coding model

Adoption and trust measure different things, and conflating them is where most "best AI model" content goes wrong. Adoption asks whether a developer opened the tool this week. Trust asks whether they shipped its output without verifying it first. A developer can use a tool every day precisely because they do not trust it unsupervised. It is the same reason usage charts for AI assistants and rankings of which model to trust never line up.

Our survey separated those questions. Of our engineers surveyed, spanning full stack, backend, front end, mobile, DevOps, QA, and the ML engineers and AI engineers working on generative-AI projects, 87% reported daily or multiple-daily use. When asked what they do with the output, 85% require a thorough review before production, 4% restrict AI-generated code to non-critical paths, 4% keep it out of production entirely, and 6% ship after a light review.

How they reach those models matters as much as which one they pick. Our engineers work through in-IDE AI coding assistants like GitHub Copilot inside VS Code, chat interfaces, direct API calls inside scripts, and command line tools handling terminal-based coding tasks. An inline code completion suggestion and an AI coding agent running in agent mode across a repository are not the same review problem, and the 85% figure covers both.

The interesting part is what does not move that number. Among engineers using AI multiple times a day, 82% still require a thorough review. Among those with 11 or more years of experience, 87% do. Nobody in the senior cohort ships on a light review. A standard that holds steady across usage intensity and experience level is not a personal preference. It is a practice, and a practice is something a client can be shown. 

Stack Overflow’s February 2026 analysis of its 2025 Developer Survey found trust in AI output at 29%, down from 40% the year before, with 46% actively distrusting what the tools produce. Read alongside our data, that 29% looks less like collapsing confidence than a measurement artifact. Ask a working engineer whether they trust AI-generated code, and the honest answer of "not unreviewed" registers as a no. Ask which model they trust most for production, and 70% of our people name one without hesitating.

Which AI model wins on trust versus usage

The common assumption is that the model most developers use is the model most developers trust. Our data says those are two different populations answering two different questions.

Among our engineers, 74% use Claude, 38% use Gemini, and 26% use Chat GPT. And 60% use two or more large language models regularly, so this is a comparison-informed preference rather than whatever was installed first. When asked which model they trust most for production code, 70% said Claude, 19% said it depends on the task, and 6% said they do not trust AI for production code at all.

That preference holds across every kind of work we measured, and where it weakens is instructive.

Task Share using AI for it Trust Claude most No clear preference
Code review and refactoring 81% 74% 16%
Writing new feature code (code generation) 94% 73% 18%
Debugging and troubleshooting 83% 72% 15%
Architecture and design decisions 85% 72% 15%
Writing tests 83% 72% 26%
Documentation and explaining code 94% 64% 23%

Source: First Factory internal engineering survey, August 2026. Trust and preference shares are calculated among engineers who use AI for that task.

Consolidation sits between 72% and 74% on every task requiring judgment and drops to 64% on documentation, which also draws the highest share reporting no clear preference. Where the work is hard, the team converges. Where any competent model will do, preference stops mattering. That pattern is a better argument for model selection than any benchmark chart, because it shows engineers discriminating only where discrimination pays off. A SWE-bench Pro score tells you how a model performs on curated tasks. It does not tell you how it behaves in your repository. Note also that 85% bring AI into architecture and design decisions. This is not confined to autocomplete and code completion.

Now set that against the global picture.

Interface Category Work coding usage (Jan 2026) Loyalty signal
ChatGPT General chatbot 28% Not reported
Claude Code Specialized coding agent 18%, up 6x from ~3% (Apr-Jun 2025) 91% CSAT, +54 NPS
Gemini General chatbot 8% Not reported
Claude (chatbot) General chatbot 7% Not reported

Source: JetBrains, "Which AI Coding Tools Do Developers Actually Use at Work?" (April 2026), based on its January 2026 AI Pulse survey of more than 10,000 professional developers.

A general chatbot wins on raw reach because nearly every developer has one open in a browser tab regardless of whether they use it for engineering work. Narrow the population to engineers shipping client production code under contract, most with six to fifteen years of experience, and the distribution changes completely. Our 70% is not a rebuttal of JetBrains’s 28%. It is what that study’s loyalty numbers look like from inside a single team that adopted the specialized tool.

Where trust varies by industry, and what enterprise leaders see

Trust is not evenly distributed across the work. We asked engineers which model they trust most in each industry they have actually shipped in, and the spread says more than the leader in any one of them.

Industry Trust Claude most No clear preference
Financial Services / FinTech 74% 26%
Education 61% 26%
Professional Services 55% 31%
Healthcare 50% 38%
Lifestyle Brands 50% 27%

Source: First Factory internal survey, August 2026. Shares calculated among engineers with production experience in that industry.

Financial services shows the tightest consolidation and the lowest indifference on the board. Healthcare shows the widest spread, with more than a third reporting no clear preference. Where correctness is under audit, our engineers converge. Where compliance pressure is real but the failure modes are less legible in code, model selection is still an open engineering question worth advising clients on rather than defaulting. Regulated work also shifts what an engineer weighs: where the data is processed, whether enterprise deployment options and security controls are available, and how well the model handles domain-specific logic.

Deloitte’s 2026 "State of AI in the Enterprise" survey of leaders across 24 countries found roughly 60% believe they are balancing rapid AI implementation against risk effectively, while naming trust in generative AI as one of the biggest remaining barriers to putting their AI investments to work at enterprise scale. Worker sentiment inside those organizations splits into about 13% highly enthusiastic, roughly 55% open but not yet convinced, and a smaller share actively resistant.

That 55% is the group most vendors misread. A vendor promising full AI automation is selling to the 13%. What our survey describes, a team at full daily adoption that has not relaxed a single review standard to get there, is what the 55% is actually looking for. That is what mature AI-assisted coding looks like from the buyer’s side.

There are two caveats we will not gloss over. Of our engineers, 26% selected "other" for the models they use without any write-in data offered, so part of that distribution is uncharacterized, and the review practices here are self-reported standards rather than observed behavior. Our team of engineers is a credible practitioner sample but a thinner statistical one. We also asked about model families rather than specific versions, so a preference for Claude covers whatever tier an engineer had on hand that week, whether that was Claude Opus or something lighter. Directionally, the findings held across every cut we ran. Precisely, they are one team’s answer in the current month.

How to evaluate an AI coding model for a nearshore engineering team

Picking a model by popularity treats a coding tool like a consumer app, and that is entirely the wrong approach. It writes code a real client’s system has to run in production. Model selection is a software engineering decision, not a procurement one. We asked our engineers to name their top three criteria for choosing a model, and the ranking is a usable evaluation framework on its own: code accuracy at 84%, reasoning on complex problems at 76%, large context window at 42%, follows instructions precisely at 38%, fewer hallucinations at 36%, cost at 29%, speed at 22%, and tooling integration at 18%.

Cost sixth. Speed seventh. Four criteria follow from that ranking.

  1. Correctness over price: Any client evaluating AI tooling on cost per token is measuring the variable our engineers rank sixth. Token efficiency and prompt caching move a platform bill, not the odds that the code is right, and the rework cost of a wrong answer in a production system dwarfs the difference between models.
  2. Context retention across a real session: A large context window ranked third, ahead of hallucination rate, because a model that holds constraints established twenty minutes into a pair programming session produces fewer invented methods and fewer regressions than one treating every prompt as a fresh start. This is where context engineering and repository understanding earn their keep, and it is why multi-file changes and long-horizon work separate models that look identical on a single prompt.
  3. A verification step that exists on paper and in practice: 76% of our engineers watch for incorrect or hallucinated logic, 53% for security vulnerabilities, and 44% for data privacy or leakage. The right question is not whether a team uses AI; it is what the reviewer is checking for, whether code quality is enforced at the gate or measured after the fact, and whether the security practices behind that check are written down. Licensing and IP concerns sit on the same list, which is why IP assignment language in a client contract should cover AI-assisted output explicitly.
  4. Fit with the business logic already in the codebase, not just the language: A model can write syntactically correct code in any popular language and still misunderstand domain rules a legacy system has encoded for a decade. Legacy codebase discovery is engineering work, not a prompt, and code suggestions that skip it add technical debt faster than they remove it.

The distinction between a model and the harness around it matters more as agentic systems take on more of the work. Vendors now ship that harness as a product, from the Claude Agent SDK to Amazon Q Developer and GitHub Copilot’s agent mode, which is why comparing bare model names is a shrinking part of the decision.

Johnn Castro, First Factory’s Director of Engineering, put it this way: "We don't hand every engineer the same model and hope it works out. What earns Claude Code space in a client's codebase isn't the model alone. Anthropic builds the model and the harness together and tunes them against each other, so the tool-calling, the context handling, and the pause before a destructive action are all shaped by the harness the model was trained to work inside. The other models hold their own where the work is easy. Our team splits on documentation and converges on architecture, and that's the whole answer: The gap doesn't show up in a single prompt, it shows up twenty minutes into a session when the constraints you established at minute three still have to hold."

That is also why AI tooling decisions inside a staff augmentation engagement are not made once and left alone. An engineer embedded in a client’s team adapts to whatever review process and risk tolerance that client already runs, which means the best model is whichever clears that client’s actual bar. First Factory’s AI Solution Services work the same way: Model selection, and increasingly model routing across more than one model, is a per-engagement decision made against the client’s business logic, compliance requirements, AI requirements, and existing systems and tech stack, not a single default handed to every project.

When chasing the most popular model is the wrong call

A general chatbot’s 28% work-coding share is a real number and the wrong statistic to build a tooling decision around. Popularity measures how many developers have a tab open. It says nothing about whether the code that tab produced shipped without a second look. A team that standardizes on whatever leads the adoption chart this quarter will standardize on something different next quarter, and a codebase does not benefit from switching its primary AI dependency every two quarters to chase a leaderboard. The AI-powered coding market moves fast enough that this quarter’s leader is a snapshot, not a strategy.

There is also a quieter risk in the data worth naming. Thirty-one percent of our engineers flagged over-reliance and skill erosion as something they actively watch for. Nobody markets that concern, and it signals a team thinking past the current quarter. An engineer worried about skill erosion is protecting the long-term value of the team a client is paying for. A tooling decision that optimizes purely for throughput buys velocity this year against capability next year.

The same logic applies to switching away from a tool a team already trusts. A newer model topping this quarter’s chart is not by itself a reason to retrain engineers and reset the review habits they have built. Migration costs weeks of reduced velocity while a team relearns where a new tool is reliable and where it needs a second pass, and the feedback loops that tell them which is which take a sprint or two to re-form. That cost is worth paying when the evidence says the tool reduces defects or rework. It is not worth paying because it has more downloads this month. A short pilot project inside a real sprint, measured against defect rate and rework rather than developer enthusiasm, tells a team more in two weeks than any survey published this year, including ours.

FAQ

Which AI model do nearshore developers actually trust for production code? 

In First Factory’s August 2026 engineering survey, 70% named Claude as the model they trust most, and the same share would recommend it to a client standardizing on one. Claude led all six task categories and all six industry categories measured.

Do developers actually trust AI-generated code? 

Not unreviewed. Stack Overflow found only 29% of developers trust AI-generated output, down from 40% in 2024, while 46% actively distrust it. Our data suggests that framing undercounts conditional trust: 87% of our engineers use AI daily, 85% require a thorough review, and 70% can still name the specific model they trust most.

Why does First Factory’s data differ from global usage charts? 

Global surveys measure reach across the entire developer population, including casual and non-production use. Our sample is of engineers shipping client production code, most with six to fifteen years of experience, 60% of whom use two or more models regularly. Narrow the population to people accountable for production output and the distribution shifts toward the specialized tools JetBrains’ own loyalty scores already favor.

Does the review standard slip when developers use AI more heavily? 

Not in our data. Among engineers using AI multiple times a day, 82% still require a thorough review, and among those with eleven or more years of experience, 87% do. No senior engineer in the sample ships AI-generated code on a light review.

How does First Factory decide which AI tools its engineers use on a project? 

Per engagement, against the client’s existing systems and tech stack, compliance requirements, project scope, and review process, rather than defaulted company-wide. Our industry data supports that: consolidation runs at 74% in financial services and 50% in healthcare, so a single default would be wrong somewhere.

Will AI coding tools replace the need for nearshore developers? 

The data argues against it. Trust in unsupervised output is falling globally, enterprise leaders name that same gap as a barrier to scaling AI, and the engineers using these tools most intensively have kept a human review gate on every line that ships. The gap between usage and trust is exactly the space a verified human engineer occupies. Nearshore outsourcing has not been made redundant by AI-assisted development. It has been re-priced around judgment.

Do you need nearshore AI engineers to build AI features, or can a generalist team handle it? 

It depends on the AI requirements. Most artificial intelligence work inside a product today is integration work: wiring an existing model into an application, which is ordinary software development that any strong engineer can do. Nearshore AI development that involves evaluation, model deployment, and the monitoring systems around it is AI engineering, and that is a different skill set. When we scope AI projects we staff for the actual work, with dedicated teams that pair application engineers with ML engineers where the problem calls for it, sized to a pilot project first. It is also the honest test of any nearshore partner selling AI development services: ask who on the proposed team has shipped an AI feature into production, not how many people are available.

Ready to see how AI actually fits your nearshore team

First Factory has delivered nearshore software development out of Costa Rica for more than two decades, with an average client tenure of over three years and a 30-day satisfaction guarantee on every resource for the life of the engagement. Costa Rica puts our engineers in US time zones, which is the practical difference between real-time collaboration and the overnight handoff that offshore delivery from Eastern Europe or Asia depends on: shared standups, same-day answers, and Agile delivery inside your sprint cadence, with the cultural alignment that comes from working the way US engineering teams already work. If the work is AI work, the same logic that applies to models applies to people. Latin America’s tech talent market has matured, and the AI talent pools are deep enough now that you can select nearshore AI engineers on production evidence rather than on availability.

Compare how AI fits into an actual engagement model and see the evidence behind it: Compare engagement models. Ready to talk specifics: Book a call.

Don Gregori is the Chief Operating Officer of First Factory, a multinational software solutions provider based in New York with nearshore operations in Costa Rica. A certified AI Business Leader, Don brings over 25 years of experience helping businesses from startups to Fortune 500 companies navigate product development, digital transformation, and AI adoption. He is a contributing author to The AI Journal and the author of The Emergent Leader, releasing June 16, 2026.