A mortgage AI benchmark is back in focus after new industry reporting highlighted accuracy and bias problems in leading general-purpose models, just as lenders accelerate their use of automation. An Aug. 17 mortgage-industry survey found 82% of mortgage professionals expect AI and automation to be the industry’s biggest drivers of change.
The MortarBench study, developed by Columbia University researchers and mortgage technology company Tidalwave, tested AI models on tasks such as matching payroll deposits, identifying large deposits, reviewing joint accounts, and reconciling bank statements with application data. On its exact-match measure, Gemini 3.1 Pro produced fully correct answers in 77.1% of cases. GPT-5.5 scored 76.8%, while Claude Sonnet 4.6 scored 51.4%.
The results do not mean AI is getting one in four mortgage decisions wrong. MortarBench tested general-purpose models on synthetic loan files, not deployed lender systems making final credit decisions.
What the mortgage AI test measured
MortarBench includes 188 test cases built from 47 unique questions. More than half required information from the Uniform Loan Application Dataset, while 57.4% required models to return lists of transaction IDs.
Transaction extraction proved especially difficult. Researchers found models could miss relevant transactions or return the wrong ones, including treating a personal loan as buy-now-pay-later debt, assuming wire transfers were international, and labeling a one-time housing payment as recurring.
Applied to Gemini 3.1 Pro, the researchers’ CRIT confidence framework improved exact-match accuracy from 77.1% to 80.5%. It also reduced, but did not eliminate, the measured name-related bias.
Mortgage lenders are accelerating AI adoption
AI is already moving deeper into buyer qualification and mortgage workflows, including platforms that assess financial readiness before connecting consumers with agents or lenders. The August survey found mortgage professionals already use AI for tasks such as research, marketing, and guideline interpretation, while respondents expect document collection and other administrative work to become increasingly automated.
The regulatory framework is evolving alongside that adoption. Fannie Mae’s AI governance requirements require seller-servicers using AI or machine learning in covered origination or servicing activities to maintain risk-management policies and oversee vendor use. They must also disclose their AI use and safeguards to Fannie Mae when requested.
Questions to ask before the referral
Agents do not need to audit a lender’s technology, but they can ask practical questions before referring clients to a lender using AI-assisted processes:
- Is borrower information sent to an outside AI provider?
- Who reviews AI-generated flags before they affect the loan file?
- Is AI used only to organize or flag documents, or can its output influence credit decisions?
- Has the lender independently evaluated the system it uses?
Names changed how models flagged deposits
In a separate experiment, researchers asked models which deposits “could be of foreign origin.” Across the three baseline models, transactions associated with English-language names were classified that way 13.3% of the time, compared with 77% for non-English names.
The experiment did not measure loan denials, pricing, or other real-world borrower outcomes. The results showed a strong difference in how the models classified deposits associated with English and non-English names.
Human review should be part of lender vetting
Agents referring buyers to AI-enabled lenders should know where human review enters the process. If an automated review incorrectly flags a document or transaction, it could trigger additional questions or manual review before the file moves forward. When vetting lender partners, agents should understand whether AI is assisting a human reviewer or shaping decisions before a person sees the file.