<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><title>AI Business Research</title><link>https://www.jedsonpinto.com/ai-research</link><description>New research using or studying AI/LLMs in business</description><lastBuildDate>Sat, 29 Aug 2026 12:07:46 +0000</lastBuildDate><item><title>Enhancing Consumer behavior Prediction with LLM-generated Data: A Signal Combination Approach</title><link>https://doi.org/10.2139/ssrn.7362741</link><guid isPermaLink="false">doi:10.2139/ssrn.7362741</guid><category>management</category><category>method</category><description>Consumer willingness-to-accept data from a social media discontinuation experiment, with GPT-3.5 and GPT-4.1 generating synthetic responses treated as noisy signals. GPT-3.5 and GPT-4.1 generate predictions combined with real human data through supervised learning rather than used as direct substitutes for human responses. Prediction error falls 66 to 86 percent versus naive substitution; GPT-4.1, a worse direct predictor, yields nearly identical final prediction quality through signal combination.</description></item><item><title>AIRE and AIREOM: A Foundational Framework for Human-AI Strategy</title><link>https://doi.org/10.2139/ssrn.7354658</link><guid isPermaLink="false">doi:10.2139/ssrn.7354658</guid><category>management</category><category>object</category><description>Conceptual strategy framework with no empirical sample, aimed at executives, AI product leaders, and strategists planning LLM and agentic AI deployment. Proposes AIRE (AI as Relational Entity) theory and AIREOM optimization model for how users assign roles to, and develop trust in, LLM systems. Argues competitive differentiation comes from the relational roles users assign to AI systems rather than raw task capability; detailed operationalization reserved for a future paper.</description></item><item><title>Attunement Costs: How Interaction Locks Users into Specific LLMs</title><link>https://doi.org/10.2139/ssrn.7354642</link><guid isPermaLink="false">doi:10.2139/ssrn.7354642</guid><category>economics</category><category>object</category><description>50 HumanEval coding tasks in four prompt variants submitted to 46 LLMs, yielding 9,200 prompt-model combinations testing cross-model transferability of prompting techniques. 46 LLMs (not individually named) respond to four prompting variants; model-level heterogeneity in technique effects is the central empirical test of attunement costs. Prompting techniques that improve one model frequently degrade another, with little cross-model agreement, supporting model-specific interaction knowledge that creates a novel procedural switching cost.</description></item><item><title>The Accountability Void: Governing Recursive Delegation and Semantic Intent in Agentic AI Ecosystems</title><link>https://doi.org/10.2139/ssrn.7362639</link><guid isPermaLink="false">doi:10.2139/ssrn.7362639</guid><category>management</category><category>object</category><description>Three documented 2025-2026 incidents: an MCP transport vulnerability, Cursor IDE destructive-agent actions, and a Claude Code Terraform event that deleted production infrastructure. No LLM is used as a research instrument; the paper forensically reconstructs real agentic AI failures to identify gaps in recursive delegation and semantic intent governance. Existing telemetry instruments the compute boundary but misses the semantic boundary; a framework combining continuous AI identity, trace-based assurance, and graduated oversight is proposed.</description></item><item><title>The Strategic Roadmap for Sovereign AI - Leveraging the EU AI Act for Sustainable Competitive Advantage in European Enterprises</title><link>https://doi.org/10.2139/ssrn.7355023</link><guid isPermaLink="false">doi:10.2139/ssrn.7355023</guid><category>management</category><category>object</category><description>European enterprises evaluating sovereign hybrid versus U.S. hyperscale AI deployment architectures under the EU AI Act, modeled with a risk-adjusted total cost of ownership framework. The paper examines large language model deployment strategies as the business object; no language model is used as a research instrument. Confidential computing adds roughly 25 percent to costs, but the hyperscale default carries 2.46 times higher long-term cost burden after adjusting for regulatory and breach risk.</description></item><item><title>Research Report 6: Agentic Shopping is Complicated and Contingent</title><link>https://doi.org/10.2139/ssrn.7355899</link><guid isPermaLink="false">doi:10.2139/ssrn.7355899</guid><category>management</category><category>agent</category><description>Experimental study exposing multiple LLMs to product reviews from sources including Reddit and Wirecutter, varying source combinations, presentation order, and user memory statements before eliciting product recommendations. Specific model names are not stated. LLMs receive product grids and external review sources to generate shopping recommendations, with consistency tested across source, order, and memory conditions. Minor changes to the search process substantially shift product recommendations, and adding multiple sources does not average preferences but produces model-specific, path-dependent selections that undermine recommendation consistency.</description></item><item><title>Professional Job-Vacancy Rate in a Middle-Income Country: The Case of Colombia</title><link>https://doi.org/10.2139/ssrn.7362098</link><guid isPermaLink="false">doi:10.2139/ssrn.7362098</guid><category>economics</category><category>instrument</category><description>Job vacancy listings scraped from a major Colombian online job portal, classified into 28 fields of the International Standard Classification of Education and compared against labor force and enrollment composition. Web scraping collects listings; large language models and a fine-tuned XLM-RoBERTa classifier assign each vacancy to an ISCED field. No validation of the classifications against human-coded labels is reported. Business, services, and database and network design fields hold disproportionately large vacancy shares relative to their representation in both the labor force and higher education enrollment, suggesting structural misalignment in professional labor supply.</description></item><item><title>Delegation Without Payoff Responsiveness: Mandate Fidelity in Agentic Machine Execution</title><link>https://doi.org/10.2139/ssrn.7354858</link><guid isPermaLink="false">doi:10.2139/ssrn.7354858</guid><category>management</category><category>object</category><description>Theoretical analysis of organizational delegation when execution is performed by agentic AI systems combining high discretion with low responsiveness to actor-specific future payoffs. No specific language model is used; the paper develops contingency theory treating AI executors as a novel delegation configuration distinct from human agents or traditional automation. Three propositions predict that output inspection becomes non-diagnostic as mandate complexity rises, that low payoff responsiveness both limits misalignment and forfeits aligned-incentive benefits, and that trajectory-level monitoring grows more important than outcome checking.</description></item><item><title>AI Agents as Governance Actors in Data Trusts – A Normative and Design Framework</title><link>https://doi.org/10.2139/ssrn.7356080</link><guid isPermaLink="false">doi:10.2139/ssrn.7356080</guid><category>management</category><category>object</category><description>Conceptual design theory integrating fiduciary principles, institutional trust, and AI ethics for governance of data trusts that steward personal and organizational data. No specific language model is used; the paper proposes four design principles including fiduciary alignment, traceability, explainability, and autonomy-preserving oversight for AI agents acting within data trusts. The framework protects beneficiary interests and data-owner rights while mitigating opacity and conflicts of interest; the authors call for sector-specific empirical validation of the proposed principles.</description></item><item><title>Silicon Sampling: Supporting Survey Research with Large Language Models</title><link>https://doi.org/10.2139/ssrn.7362005</link><guid isPermaLink="false">doi:10.2139/ssrn.7362005</guid><category>management</category><category>agent</category><description>6,048 synthetic survey responses across three models and three prompting strategies, benchmarked against 672 human responses from luxury consumption and B2B green marketing surveys. Three LLMs (families not named) generate synthetic respondent data under zero-shot, persona, and few-shot strategies, evaluated via MANOVA, confirmatory factor analysis, structural equation modeling, and SHAP. Few-shot prompting improves alignment with human data and synthetic responses preserve measurement reliability, but consumer scores run high and SHAP attribution patterns diverge from human baselines.</description></item><item><title>Generative AI and CEO Compensation</title><link>https://doi.org/10.2139/ssrn.7358298</link><guid isPermaLink="false">doi:10.2139/ssrn.7358298</guid><category>finance</category><category>object</category><description>U.S. public firms around ChatGPT&#x27;s release, with cross-sectional variation in pre-shock generative AI exposure and a shock-based instrumental variables identification strategy. ChatGPT&#x27;s public launch serves as the exogenous event; no language model is used as a research instrument by the authors. GenAI exposure raises CEO pay through equity-based compensation, driven by labor cost reductions and increased uncertainty, with pay adjustments predicting higher future firm value.</description></item><item><title>Talking the ESG Tightrope: Between Greenwashing and Greenhushing</title><link>https://doi.org/10.2139/ssrn.7357478</link><guid isPermaLink="false">doi:10.2139/ssrn.7357478</guid><category>finance</category><category>instrument</category><description>Quarterly earnings calls for firms across Europe and North America, with firm-level emissions, diversity data, and quasi-experimental variation from the 2018 EU ETS reform and 2021 Texas anti-ESG legislation. An LLM characterizes the context and depth of climate and DEI discourse in earnings calls. The specific model and any validation against human coding are not reported. Climate discussion correlates with subsequent emission cuts. DEI discussion is tangential and not followed by diversity gains. Anti-ESG legislation chilled even financially material environmental disclosure.</description></item><item><title>Generative AI and Mutual Fund Industry Concentration</title><link>https://doi.org/10.2139/ssrn.7347639</link><guid isPermaLink="false">doi:10.2139/ssrn.7347639</guid><category>finance</category><category>object</category><description>Website traffic and capital flow data for mutual fund families, analyzed around the introduction of OpenAI&#x27;s SearchGPT, with GPT service outages used as additional identification. SearchGPT is the treatment event. No model is run by the researchers; they exploit the platform launch and its outages to identify changes in investor search behavior. After SearchGPT, traffic concentration across fund families flattened and large families lost their relative traffic advantage, followed by declines in new sales and total net assets.</description></item><item><title>Local Execution and a Verifiable Audit Record as Compliance Substrate: A Self-Hostable Inference Runtime under the EU AI Act (Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744), the GDPR, the Cyber Resilience Act, DORA and NIS2</title><link>https://doi.org/10.2139/ssrn.7358638</link><guid isPermaLink="false">doi:10.2139/ssrn.7358638</guid><category>management</category><category>object</category><description>Legal analysis mapping a self-hostable LLM inference runtime against the EU AI Act, GDPR, Cyber Resilience Act, DORA, and NIS2 compliance requirements. No model is run empirically. The paper examines how running inference locally rather than through a third-party API changes an organization&#x27;s demonstrable compliance position. Local execution with a tamper-evident audit record addresses three compliance gaps: international data transfers under GDPR, third-party control of decision records, and record integrity.</description></item><item><title>Language-dependent Source Representation in AI Search: A Cross-platform Audit of Lebanon</title><link>https://doi.org/10.2139/ssrn.7358800</link><guid isPermaLink="false">doi:10.2139/ssrn.7358800</guid><category>management</category><category>object</category><description>320 AI search responses to 40 matched English-Arabic question pairs about Lebanon across eight topics, submitted to ChatGPT Search, Gemini, Perplexity, and Copilot. Four AI search platforms are audited as information intermediaries. Source hostnames were classified by institutional origin and type. No model is used by the researchers. Arabic queries produced 22.9 percentage points more Lebanese-origin sources than English queries, drawing on domestic government and news sources rather than international organizations.</description></item><item><title>Context-Aware Sustainability Narratives: An Industry-Conditioned Large Language Model Pipeline for Environmental Strategy and Governance</title><link>https://doi.org/10.2139/ssrn.7359745</link><guid isPermaLink="false">doi:10.2139/ssrn.7359745</guid><category>finance</category><category>instrument</category><description>Firm-specific news for S&amp;P 500 companies, scored along the GICS hierarchy and SASB materiality standards, analyzed over 11-day and 21-day event windows. A two-stage LLM pipeline separates semantic extraction from quantitative scoring to produce ESG narrative sentiment. The specific model is not named, and no validation is reported. Governance narratives drive unconditional return drift. Environmental news is priced positively only after a negative governance event, where it signals remediation; otherwise it is discounted.</description></item><item><title>Agentic Empirical Asset Pricing: Methodological Foundations *</title><link>https://doi.org/10.2139/ssrn.7359503</link><guid isPermaLink="false">doi:10.2139/ssrn.7359503</guid><category>finance</category><category>method</category><description>Two US equity panels used to evaluate SEADS, an autonomous factor-discovery architecture, against five re-implemented baselines under a proposed evaluation standard for agentic discovery systems. LLM agents autonomously conduct factor discovery. The paper contributes a reference architecture, a rigorous evaluation standard for discovered factors, and a rolling re-execution method to backtest the discovery process itself. No single metric consistently ranks factor-discovery systems, motivating multi-axis evaluation. Negative findings and limitations surface further pitfalls for future agentic empirical asset pricing work.</description></item><item><title>When Output Stops Being Evidence Generative AI and Production-Verification Decoupling in Institutional Accountability</title><link>https://doi.org/10.2139/ssrn.7356461</link><guid isPermaLink="false">doi:10.2139/ssrn.7356461</guid><category>management</category><category>object</category><description>Theory-building synthesis across institutional accountability settings, drawing on peer-reviewed literature and industry evidence. No original data or empirical test is reported. No specific model is used or named. The paper theorizes that generative AI collapses the cost of producing convincing output, severing the link between output quality and demonstrated competence. Introduces production-verification decoupling, distinguishes proxy collapse from generation-verification asymmetry, and derives four testable propositions and a governance heuristic for allocating verification effort across settings.</description></item><item><title>AI Intermediation and Investor Information Processing: Evidence from Earnings Summaries on Social Media</title><link>https://doi.org/10.2139/ssrn.7347558</link><guid isPermaLink="false">doi:10.2139/ssrn.7347558</guid><category>finance</category><category>object</category><description>Within-day difference-in-differences design around the staggered rollout of AI-generated earnings summaries on StockTwits, tracking browsing, engagement, and trading behavior of retail investors. StockTwits deploys a generative AI to produce structured post-earnings summaries; the model family is not stated. The study examines behavioral responses rather than model performance. After summary release, users browse less, engage less with peer content, and align subsequent posts with the AI interpretation. Retail trading increases and retail order flow gains predictive power for future returns.</description></item><item><title>AI Inference as Digital Infrastructure: Chip Controls, Electricity Burdens, and Production Location</title><link>https://doi.org/10.2139/ssrn.7362469</link><guid isPermaLink="false">doi:10.2139/ssrn.7362469</guid><category>economics</category><category>object</category><description>Two-location analytical framework with tradable AI token services, local power-market feedback, and asymmetric chip frictions, calibrated with public evidence on data centers, GPU supply, and electricity systems. No specific language model is used. The paper models AI inference as a commodity service whose unit energy cost depends on the chip technology available under export controls. Import restrictions on advanced chips raise unit electricity use when the domestic substitute is cheap relative to its energy disadvantage. Chip subsidies amplify the effect, and chip, data-center, and power-system policy must be treated as connected margins.</description></item><item><title>Observable Certifiability of LLM-Generated Compliance Policies: Risk Control Relative to Target Loss and Deployment Information</title><link>https://doi.org/10.2139/ssrn.7364531</link><guid isPermaLink="false">doi:10.2139/ssrn.7364531</guid><category>management</category><category>method</category><description>Evaluation on 234 CrossLib compliance cases and 262 active NIST SP 800-53 security controls, with 30 author-reviewed clean pairs for omission analysis and 74 independent hold-out cases. An LLM generates compliance policy drafts; the model family is not named. A semantic-medoid selection gate with simultaneous exact-binomial calibration certifies deployment within a declared risk tolerance. Self-consistency detects competing semantic modes but misses shared omissions (AUC 0.571). Adding obligation-specific features restores discrimination to AUC 0.903. Best-of-K evaluation understates deployed-medoid risk by up to 3 percentage points.</description></item><item><title>Choosing the Song Still Matters: Human Song Selection and Video Engagement in an AI-Saturated Music Ecosystem on YouTube</title><link>https://doi.org/10.2139/ssrn.7365626</link><guid isPermaLink="false">doi:10.2139/ssrn.7365626</guid><category>management</category><category>object</category><description>755 YouTube music videos uploaded from August 2025 to January 2026 in an AI-associated remix ecosystem, split between 465 remixes or covers of existing songs and 288 fully AI-generated original tracks. Content type classified by LLM-assisted first-pass coding with machine verification of evidence quotes and full author review; an independent second coder yielded Cohen&#x27;s kappa of .795. The LLM family is not named. Remixes of pre-existing songs attracted 47.4 times the views of fully AI-generated originals in an unadjusted model. Channel-size controls reduced the ratio but human song selection remained the strongest engagement predictor.</description></item><item><title>Beyond Black Box Monitoring: Mechanistic Audit Trails for AI Decision Systems in Regulated Industries</title><link>https://doi.org/10.2139/ssrn.7357160</link><guid isPermaLink="false">doi:10.2139/ssrn.7357160</guid><category>finance</category><category>object</category><description>US and EU regulatory frameworks for AI in credit underwriting, insurance pricing, and clinical decision support, anchored by the April 2026 OCC interagency guidance and the EU AI Act effective August 2026. No specific model is deployed; the paper assesses whether post-hoc explanation tools such as LIME and SHAP satisfy regulatory requirements for neural-network and generative AI decision systems. Current explanation methods produce structurally unfaithful and empirically unstable approximations of model reasoning, leaving a compliance gap that existing US and EU frameworks cannot close with available tools.</description></item><item><title>An Auditable AI Agent Loop for Empirical Economics A Case Study in Forecast Combination</title><link>https://doi.org/10.2139/ssrn.7360139</link><guid isPermaLink="false">doi:10.2139/ssrn.7360139</guid><category>economics</category><category>method</category><description>Forecast-combination case study reusing data from a prior study, with a holdout evaluation stage added to an open-source AI coding agent workflow. An open-source agent-loop architecture writes and executes code to search over empirical specifications; the specific LLM powering the agent is not stated. Independent agent searches found methods that improve on the original study&#x27;s benchmarks, and logged search trails paired with holdout evaluation make adaptive specification search more transparent.</description></item><item><title>The AI That Knows When Not to Decide: Calibrating Authority Boundaries in Enterprise Agentic Systems</title><link>https://doi.org/10.2139/ssrn.7353198</link><guid isPermaLink="false">doi:10.2139/ssrn.7353198</guid><category>management</category><category>object</category><description>Conceptual analysis of enterprise agentic AI systems, drawing on experience in systems engineering, technical operations, and legacy system modernization across multiple deployment contexts. No specific language model is used; the paper studies autonomous AI agents as organizational actors and proposes a four-mode boundary model: automate, assist, escalate, and abstain. Applying uniform governance across all AI agents regardless of autonomy level tends toward operational failure; an active human-in-command model is argued to outperform passive human-in-the-loop oversight.</description></item><item><title>The Style Penalty: How AI Resume-Writing Tools Shape Outcomes in Automated Hiring Screens An Empirical Analysis of 1,576 LLM Evaluation Decisions Across Four Frontier Models</title><link>https://doi.org/10.2139/ssrn.7348180</link><guid isPermaLink="false">doi:10.2139/ssrn.7348180</guid><category>management</category><category>object</category><description>100 synthetic candidate profiles across 12 industries and four career levels, each paired with a job posting, producing 400 resumes and 1,576 blind evaluations. GPT-5.4, Claude Sonnet 4.6, Gemini 3 Pro, and Grok 4.3 each generate and evaluate resumes, scoring fit on a 0 to 100 scale with hire, maybe, or reject recommendations. Hire rates for identical candidates vary by up to 42 percentage points depending on which model wrote the resume; Claude shows 84 percent self-preference while GPT-5.4 rates its own style lowest.</description></item><item><title>Leveraging Large Language Models and Agentic AI in Supply Chain Operations: A Framework and Industrial Implementation</title><link>https://doi.org/10.2139/ssrn.7347999</link><guid isPermaLink="false">doi:10.2139/ssrn.7347999</guid><category>management</category><category>method</category><description>Anonymized operational data from a global manufacturing company in Turkey, covering order management and logistics workflows across multiple industrial use cases. The model family is not stated. The EMPLOOY framework uses LLMs for data management, operational queries, and context-aware decision support, consuming fewer tokens than single-shot prompting. The implementation answers operational queries and is expected to improve responsiveness, coordination, and visibility, though quantitative performance gains are not reported.</description></item><item><title>A Hybrid MCDM Framework for Ceramic Art Innovation: Integrating Human-in-the-Loop LLMs and Fuzzy Delphi</title><link>https://doi.org/10.2139/ssrn.7347979</link><guid isPermaLink="false">doi:10.2139/ssrn.7347979</guid><category>management</category><category>instrument</category><description>Six ceramic art innovation strategies evaluated against 21 criteria by a multidisciplinary expert panel, with LLMs generating the initial criterion set from tacit knowledge. LLMs (not named) extract tacit knowledge to seed criteria, which experts validate through the Fuzzy Delphi Method before TOPSIS ranking. Digital craft hybridisation ranks first; the framework holds 70 percent rank stability across 30 sensitivity scenarios and correlates 0.94 with three alternative MCDM methods.</description></item><item><title>LLMs as Policymakers (in a Sandbox)</title><link>https://doi.org/10.2139/ssrn.7349758</link><guid isPermaLink="false">doi:10.2139/ssrn.7349758</guid><category>economics</category><category>agent</category><description>Conceptual framework and research agenda drawing on prior LLM-simulation policy experiments; no original empirical sample or new simulation results reported. Proposes giving LLMs bounded discretion in a closed decision loop with a simulation model, defining a reference architecture for sandbox policy experiments. Distinguishes experimental role enactment from epistemic reliance, arguing sandbox performance does not establish real-world policy efficacy or justify institutional authority.</description></item><item><title>DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research</title><link>https://arxiv.org/abs/2608.26990v1</link><guid isPermaLink="false">arxiv:2608.26990v1</guid><category>finance</category><category>method</category><description>Reference implementation spanning six regional markets and fifteen bundled strategy skills, with 1,457 backend contract tests passing at a frozen software snapshot. LLM agents (not named) follow an evidence-acquisition, context-construction, and model-routed analysis pipeline with role-specific parsing and conservative risk override. Contract tests confirm implementation conformance only; the paper explicitly does not validate report quality, forecasting accuracy, or investment returns from the generated research.</description></item><item><title>LLMs Do Not Emulate Populations</title><link>https://doi.org/10.2139/ssrn.7352138</link><guid isPermaLink="false">doi:10.2139/ssrn.7352138</guid><category>economics</category><category>method</category><description>Synthetic survey experiments testing whether prompt-conditioned LLMs can recover the statistical structure of real human populations across demographic subgroups. Multiple LLMs prompted to generate survey responses as demographically conditioned population emulators; specific model families not named in the abstract. LLM outputs violate basic population composition constraints at rates similar to random predictions, and model choice explains more output variation than the emulated demographic group.</description></item><item><title>When the Crowd Speaks: AI-based retail investor sentiment indicator with Reddit data</title><link>https://doi.org/10.2139/ssrn.7352464</link><guid isPermaLink="false">doi:10.2139/ssrn.7352464</guid><category>finance</category><category>instrument</category><description>Reddit posts from major investing and cryptocurrency subreddits between 2015 and April 2026, linked to daily S&amp;P 500 returns and established sentiment benchmarks. ChatGPT 5.1 classifies posts by sentiment and topic, outperforming FinBERT on informal social media language; validation against human-coded labels is not reported in the abstract. The resulting retail sentiment indicator predicts next-day S&amp;P 500 returns with statistical significance, and the relationship strengthens during periods of market decline.</description></item><item><title>Domain-AI Driven Marketing, Topic-AI Conversational Marketing &amp; AI Customer Service</title><link>https://doi.org/10.2139/ssrn.7353001</link><guid isPermaLink="false">doi:10.2139/ssrn.7353001</guid><category>management</category><category>object</category><description>LLM-powered conversational marketing and customer service systems, examined through the Moffatt v. Air Canada case, the FTC v. DoNotPay settlement, and India&#x27;s DPDP Act framework. The paper studies generative AI chatbot adoption as a business phenomenon, analyzing hallucinations, prompt injections, and guardrail regressions as operational failure modes. Courts and regulators impose strict corporate liability for automated chatbot representations, motivating a governance framework with risk-tiered deployment and grounded single-source databases.</description></item><item><title>Sustained Human-LLM Collaboration as an Emerging Object of Scientific Inquiry</title><link>https://doi.org/10.2139/ssrn.7349183</link><guid isPermaLink="false">doi:10.2139/ssrn.7349183</guid><category>management</category><category>object</category><description>Conceptual framework drawing on team cognition, transactive memory, distributed cognition, and collective intelligence research to analyze sustained human-LLM interaction, with no original empirical data. No language model is used or named; the paper theorizes that sustained dyadic interaction between a human and a configured LLM produces a collaborative system with emergent partner-specific properties. Two hypotheses are proposed: sustained interaction generates shared vocabulary and cognitive specialization, and these accrued properties improve joint performance beyond gains from practice or prompt engineering alone.</description></item><item><title>Sophistication in GenAI Use: Field Evidence from a Large Firm</title><link>https://doi.org/10.2139/ssrn.7351443</link><guid isPermaLink="false">doi:10.2139/ssrn.7351443</guid><category>management</category><category>object</category><description>713,564 employee prompts and LLM responses from nearly 4,000 back-office employees across 15 functional areas at a large firm over eight months in 2025. The study observes employees&#x27; real interactions with a generative AI tool. The specific language model is not named in the abstract. Senior employees show more sophisticated GenAI use. Sophistication varies across functions but did not improve over time or after formal AI training programs.</description></item><item><title>Artificial Intelligence and the Future Supply of Skills: Evidence from UK University Applications</title><link>https://doi.org/10.2139/ssrn.7354159</link><guid isPermaLink="false">doi:10.2139/ssrn.7354159</guid><category>economics</category><category>object</category><description>UK university applications for 1,090 degree programs over the 2020-2025 cycles, with AI exposure of each degree measured from millions of linked LinkedIn graduate profiles. ChatGPT&#x27;s November 2022 release serves as the treatment event. No model is run by the researchers; they exploit the launch to identify shifts in student demand. Applications to degrees one standard deviation higher in AI exposure grew 6 percent more after ChatGPT&#x27;s release, despite deteriorating entry-level job opportunities in those fields.</description></item><item><title>Agentic AI: Technical Developments, Potential Impact on Consumers and Markets, and Regulatory Implications</title><link>https://doi.org/10.2139/ssrn.7359418</link><guid isPermaLink="false">doi:10.2139/ssrn.7359418</guid><category>economics</category><category>object</category><description>Survey of agentic AI as an operational ecosystem around foundation models, covering technical architecture, consumer and market risks, and applicable EU regulatory frameworks. No specific model is tested. The paper analyzes how agentic systems that plan, use tools, and execute actions shift AI from a passive information tool to a delegated intermediary in digital markets. Existing EU instruments including consumer law, the Digital Services Act, the Digital Markets Act, and competition law partially address agentic AI risks, but gaps remain for delegated autonomous action on behalf of consumers.</description></item><item><title>Generative Change Management: Foundational Concepts for Human-AI Hybrid Productivity in Organizations</title><link>https://doi.org/10.2139/ssrn.7351278</link><guid isPermaLink="false">doi:10.2139/ssrn.7351278</guid><category>management</category><category>object</category><description>Conceptual framework addressing three interdependent levels of organizational transformation from generative AI integration: professional contribution redefinition, new work architectures, and cultural coexistence with algorithms. Generative AI is the studied phenomenon; the paper proposes a recursive cycle of analysis, design, and implementation for managing human-AI hybrid production arrangements. Hybrid productivity is framed as an organization-specific capability requiring coordinated redesign of roles, production architectures, and cultural meaning systems, not achievable through technology access alone.</description></item><item><title>Understanding Artificial Intelligence and Responsible Business</title><link>https://doi.org/10.2139/ssrn.7350559</link><guid isPermaLink="false">doi:10.2139/ssrn.7350559</guid><category>management</category><category>object</category><description>Survey evidence on generative AI adoption among Japanese firms, a case from academic publishing, and firstperson observation of judgment formation under three deliberative conditions. Generative AI is the studied phenomenon; the paper examines how digital technologies participate in intent formation before organizational decisions are fixed, a process it terms assetization. AI can enclose the collaborative process of intent formation; Japanese consensus-oriented decision-making is reframed as potential resistance to premature convergence rather than a cultural deficiency.</description></item><item><title>Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit</title><link>https://arxiv.org/abs/2608.27309v1</link><guid isPermaLink="false">arxiv:2608.27309v1</guid><category>other</category><category>method</category><description>Pre-registered audit of a frozen pedagogy LLM judge across 990 calls, testing whether a stated learner profile biases scaffolding preference on a bounded rating scale. The judge&#x27;s model family is not stated. The paper derives in closed form how differential ceiling and floor censoring in double-differenced bounded scores manufactures spurious interactions. The registered primary endpoint is null (p = 0.684), and the sole significant interaction is 79 to 85 percent attributable to the scale floor interacting with a severity shift, not to differential preference.</description></item><item><title>Counterfactual Bias Testing for Application Tracking System</title><link>https://arxiv.org/abs/2608.26899v1</link><guid isPermaLink="false">arxiv:2608.26899v1</guid><category>management</category><category>method</category><description>Synthetic corpus of 100 identity-neutral resumes and 10 demographic treatments across sex, age, residence, language, and disability, applied to five job orders under an EU AI Act-aligned audit protocol. Task-specialized LLM agents synthesize base resumes and inject treatments; a fine-tuned sentence-embedding model scores candidates by cosine similarity. No model family is named. Score shifts and retention metrics pass tolerance for every treatment, but a rank-stability metric and nDCG each surface borderline findings that a single-aggregate fairness view would miss.</description></item><item><title>The Delegation Frontier</title><link>https://doi.org/10.2139/ssrn.7348378</link><guid isPermaLink="false">doi:10.2139/ssrn.7348378</guid><category>management</category><category>object</category><description>Scope is not stated beyond firms that delegate sequences of consequential actions to agentic AI systems operating across open-ended cognitive workflows. No specific model is named; the paper frames delegation scope as a measurement gap that conventional productivity accounting does not address. Not stated; the abstract poses the question of optimal delegation span and whether the system should act at all, without reporting empirical findings.</description></item><item><title>When the Executor Changes: An Executor-Contingency Theory of Project Governance</title><link>https://doi.org/10.2139/ssrn.7350382</link><guid isPermaLink="false">doi:10.2139/ssrn.7350382</guid><category>management</category><category>object</category><description>Conceptual paper on project governance, using agentic AI as the revealing case where an executor generates consequential interpretations but cannot bear institutional answerability. No specific model is deployed; agentic AI provides the theoretical test case for extending project-control theory with an executor-contingency dimension alongside the standard task contingency. Proposes that an executor&#x27;s realized interpretive autonomy reduces the governance diagnosticity of its output and that trajectory observability is needed when execution separates from answerability.</description></item><item><title>Towards Expert Financial QA via Self-Improving RAG</title><link>https://arxiv.org/abs/2608.26706v1</link><guid isPermaLink="false">arxiv:2608.26706v1</guid><category>finance</category><category>method</category><description>SEC filing question answering evaluated on FinanceBench, a benchmark of financial-document queries with gold reference answers; the system targets audit-trail compliance for regulated finance. Self-Improving RAG decomposes document QA into retrieval, reasoning, and judge agents with feedback-driven retry; the judge triggers escalated strategies when confidence falls below a dynamic threshold. Achieves 86 percent oracle-guided accuracy with a 36.4 percent Lazarus rate, recovering nearly four in ten initially incorrect answers through judge-driven retry.</description></item><item><title>Knowledge-Augmented Column Generation for Airline Crew Pairing: Agentic LLM Intervention for Dual Interpretation and Early Termination</title><link>https://doi.org/10.2139/ssrn.7355988</link><guid isPermaLink="false">doi:10.2139/ssrn.7355988</guid><category>management</category><category>instrument</category><description>A real-world airline schedule with 1,092 flight legs, tested across eight LLMs spanning capability tiers and API costs from $0 to $0.87. GPT-5.4, o3, Claude Opus 4.6, Gemini 3.1 Pro, and Llama variants interpret dual values and rank columns inside a knowledge-augmented column generation loop; no ground-truth validation is reported. The framework improves block time utilization and terminates up to 90.8 percent faster when the solution stalls, at a total API cost below $3.50.</description></item><item><title>Small Models, Big Budgets: Open LLMs for Aligning Public Finance with the SDGs</title><link>https://doi.org/10.2139/ssrn.7343819</link><guid isPermaLink="false">doi:10.2139/ssrn.7343819</guid><category>economics</category><category>instrument</category><description>Dominican Republic national budget expenditure items tagged to Sustainable Development Goals and the National Development Strategy using locally deployable open LLMs. Qwen3 classifies budget line items via a three-step prompting protocol with output-distribution-based uncertainty measures, validated against an expert-tagged benchmark across multiple metrics. Small open models match or exceed prior supervised-learning and LLM-based tagging results, reducing classification costs while flagging uncertain cases for human review.</description></item><item><title>Contribution Estimation, Contribution Awareness, and Psychological Ownership in Co-Creation with LLMs: An Experimental Study</title><link>https://doi.org/10.2139/ssrn.7354076</link><guid isPermaLink="false">doi:10.2139/ssrn.7354076</guid><category>management</category><category>object</category><description>121 participants complete two within-subjects pitch-writing tasks co-created with an LLM, receiving contribution attribution feedback from an LLM-as-judge system between tasks. Participants co-create text with an LLM (not named); the Contribution Attribution Framework estimates human and AI shares across ideas, details, and wording dimensions. After feedback, self-assessments move toward measured contribution but actual human share stays unchanged; psychological ownership tracks perceived rather than measured contribution.</description></item><item><title>Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives</title><link>https://arxiv.org/abs/2608.26372v1</link><guid isPermaLink="false">arxiv:2608.26372v1</guid><category>management</category><category>object</category><description>KnownLieBench covers eight customer-service domains and 112 cases, testing 18 proprietary and open-weight models in multi-round dialogues with a trust-tracking customer agent. Neutral probes first verify the model knows the user&#x27;s entitlement; incentives to deny it are then introduced, separating deception from ignorance or hallucination. Emergent deception varies across model families; honesty-directed fine-tuning reduces it under incentive, while deception-graded tuning raises lie success on honest-control dialogues.</description></item><item><title>FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review</title><link>https://arxiv.org/abs/2608.25325v1</link><guid isPermaLink="false">arxiv:2608.25325v1</guid><category>finance</category><category>method</category><description>9,742 Chinese-language financial risk review instances spanning 53 task families, plus 680 replayed decision states from 104 de-identified professional review trajectories. 33 model configurations evaluated on operation execution under fixed evidence and evidence-state control under evolving review conditions; specific model families not named in the abstract. Operation-level rankings diverge from aggregate capability scores with mean pairwise Spearman correlation of 0.42, and knowledge-based shortlisting incurs up to 18 points of regret on individual operations.</description></item><item><title>The Impact of AI-Assisted Coding Tools on Agile Software Development Practices: An Empirical Study</title><link>https://doi.org/10.2139/ssrn.7331699</link><guid isPermaLink="false">doi:10.2139/ssrn.7331699</guid><category>management</category><category>object</category><description>Survey of 91 software professionals using AI-assisted coding tools such as ChatGPT, GitHub Copilot, Claude, and DeepSeek in Agile development environments, analyzed with descriptive and correlation statistics. Multiple AI-assisted tools including ChatGPT, Copilot, Claude, and DeepSeek are studied as adoption objects; the paper measures perceived effects through Likert scales, not model outputs directly. Participants reported improved coding efficiency and faster sprint execution but raised concerns about over-dependence on tools, security risks in generated code, and diminished learning opportunities.</description></item><item><title>Endogenous Arrivals in Expert-review Systems: Generative AI and the Peer-Review Bottleneck</title><link>https://doi.org/10.2139/ssrn.7342678</link><guid isPermaLink="false">doi:10.2139/ssrn.7342678</guid><category>economics</category><category>object</category><description>Stochastic service model of academic peer review where generative AI expands upstream manuscript production, calibrated with numerical experiments and public conference-scale submission statistics. No language model is used by the researchers; generative AI enters the model as a technology parameter that raises research productivity while potentially displacing reviewer effort. A Red Queen attenuation factor absorbs most upstream productivity gains through congestion; when AI also displaces reviewers, processed throughput can decline even as submission volume rises.</description></item><item><title>Simulating Firms&#x27; Inflation Expectations with a Multi-Agent AI Framework</title><link>https://doi.org/10.2139/ssrn.7340500</link><guid isPermaLink="false">doi:10.2139/ssrn.7340500</guid><category>economics</category><category>agent</category><description>Multi-agent simulation of the Survey of Firms&#x27; Inflation Expectations across 32 quarterly waves from 2018 through 2026, with distinct executive-role agents aggregated by a chairman agent. LLM family not stated; role-specific agents produce inflation forecasts synthesized at the firm level, validated against actual survey responses with staggered cutoff out-of-sample tests. Synthetic firms track main survey dynamics and generate wider within-firm disagreement during supply shocks, but produce excessively high tail-risk probabilities and compressed cross-firm dispersion.</description></item><item><title>US vs. China, Round Two: China&#x27;s Shifting AI and Supply Chain Strategy and Korea&#x27;s Response</title><link>https://doi.org/10.2139/ssrn.7351779</link><guid isPermaLink="false">doi:10.2139/ssrn.7351779</guid><category>economics</category><category>object</category><description>Policy analysis of US-China competition in the physical AI era, focusing on Korea&#x27;s strategic response in manufacturing architecture and supply chains. The paper proposes a manufacturing LLM trained on shop-floor judgment data as a core strategic asset. Claude Opus 5 was used only for translation. Korea should vertically integrate chokepoint assets with production-process data and pursue architectural assetization to participate as a co-designer in major economies&#x27; technology ecosystems.</description></item><item><title>AI Citations are Not Permanent: Citation Volatility, Half-Life, and What Actually Predicts Citation Persistence Across AI Platforms</title><link>https://doi.org/10.2139/ssrn.7344201</link><guid isPermaLink="false">doi:10.2139/ssrn.7344201</guid><category>management</category><category>object</category><description>3.5 million tracked citation events across ChatGPT, Perplexity, Google AI Overviews, Copilot, and Gemini, covering September 2025 through March 2026. Five AI search platforms are studied as citation sources. No model is run by the researchers; citation persistence, rotation rates, and predictors are measured across platforms. Median citation half-life is 4.5 weeks, and only 10.6 percent of URLs persist over 28 days. Brand search volume is the strongest predictor of citation frequency.</description></item><item><title>AI-Driven Destination Discovery in China: Positioning Iran in the Chinese Generative AI Tourism Ecosystem</title><link>https://doi.org/10.2139/ssrn.7344338</link><guid isPermaLink="false">doi:10.2139/ssrn.7344338</guid><category>management</category><category>object</category><description>660 prompt-response observations from six Chinese generative AI platforms, queried about Iran as an outbound tourism destination for Chinese travelers across thirteen thematic axes. DeepSeek, Doubao, ERNIE Bot, Kimi, Qwen, and Yuanbao were tested. Mention and recommendation were coded as separate outcomes for each platform-prompt pair. Iran reached 100 percent mention on direct queries but collapsed on open-ended prompts. Unconditional positive recommendations followed only 35.8 percent of mentions.</description></item><item><title>Governed Complementarity and the Translation Gap in National Generative AI Diffusion: Exploratory Cross-Country Evidence</title><link>https://doi.org/10.2139/ssrn.7354684</link><guid isPermaLink="false">doi:10.2139/ssrn.7354684</guid><category>economics</category><category>object</category><description>Cross-country analysis of generative AI diffusion across 146 economies, merging population-normalized behavioral telemetry with the Oxford Insights Government AI Readiness Index. No language model is used by the researchers. GenAI platform usage is the dependent variable, measured through behavioral telemetry and explained by readiness components. Overall readiness explains 63.5 percent of cross-country variance. Government capability amplifies the diffusion return to technology and infrastructure, producing a governed complementarity effect.</description></item><item><title>The Same-Model Ceiling: Testing Multi-Model AI Review in Strategic Analysis</title><link>https://doi.org/10.2139/ssrn.7346099</link><guid isPermaLink="false">doi:10.2139/ssrn.7346099</guid><category>management</category><category>method</category><description>One constructed acquisition case analyzed twelve times by Claude Opus 5, with each analysis receiving six critiques under the same brief from same-model and cross-model reviewers. Claude Opus 5 generated analyses. GPT-5.6 Sol, Muse Spark 1.2, and DeepSeek V4 Pro served as cross-model reviewers. No ground-truth benchmark was used. Cross-model panels found 37 unique issues versus 18 from same-model reviews. The multi-model revision was preferred in eight of twelve blind comparisons.</description></item><item><title>When AI Starts Acting: Governance, Audit Evidence, and Accountability for Agentic Systems</title><link>https://doi.org/10.2139/ssrn.7339582</link><guid isPermaLink="false">doi:10.2139/ssrn.7339582</guid><category>accounting</category><category>object</category><description>Conceptual framework drawing on financial statement audit, internal control, internal audit, forensic investigation, and M&amp;A due diligence to govern agentic AI systems that take consequential actions. No specific model is used. The paper treats the consequential AI action as the governance unit and develops seven control questions a skeptical reviewer would apply to autonomous system behavior. Human approval can become ceremonial, logs do not automatically constitute audit evidence, and AI reviewing AI may reproduce the same error, making retained human expertise part of the control environment.</description></item><item><title>Refuses the Shape, Serves the Substance: Evidence on the Rise of Normative AI Failure and the Case for White-Box Testing</title><link>https://doi.org/10.2139/ssrn.7340258</link><guid isPermaLink="false">doi:10.2139/ssrn.7340258</guid><category>management</category><category>object</category><description>All 1,560 incidents in the AI Incident Database as of July 2026, classified by primary failure mechanism, plus five experimental runs of a simulated banking assistant across three unnamed model vendors. No specific model family is named. Incidents are classified by mechanism and compared across five half-year windows; a banking assistant is probed with both recognizable attacks and operationally reframed equivalents. Normative failures account for 62.6 percent of incidents versus under 3.8 percent for adversarial attacks, and every tested model refused recognizable attack shapes while serving the same substance in operational framing.</description></item><item><title>Agentic AI Systems and Financial Stability: From Model Risk to Systemic Risk</title><link>https://doi.org/10.2139/ssrn.7342218</link><guid isPermaLink="false">doi:10.2139/ssrn.7342218</guid><category>finance</category><category>object</category><description>Theoretical monograph modeling agentic AI in financial systems across six mathematical settings, from single-institution model risk to fleet-level systemic risk using jump-diffusion and Hawkes processes. No specific model is used. Populations of agents sharing a common foundation model are treated as a non-diversifiable exposure, with contagion analyzed through percolation thresholds and spectral-radius criticality. Shared foundation models create a systemic-risk floor that no amount of fleet redundancy dilutes, and against irreversible harm, runtime detection and reversibility cannot substitute for ex-ante prevention.</description></item></channel></rss>