XobinLooking for talent assessment and AI hiring software?Check out Xobin
← Articles

Xobin Research · 2026 Mid-Year Edition

Human vs AI
Skills Report

What should hiring teams assess when candidates can use AI? This report brings together Xobin’s AI benchmark and hiring-related records from the first half of 2026, with matched 2024 comparisons for assessment requests and selected emerging skill groups.

Author
Guruprakash Sivabalan
Published
9 September 2026
Study cutoff
30 June 2026
Main records
1 Jan to 30 Jun 2026
Author on LinkedIn →

72%

of 683 tested skill groups met Xobin’s criteria for full or partial delegation to AI, with at least one tested setup qualifying for each group; the qualifying setup could differ between groups.[1] Capability under test, not a statement that these skills stop mattering.

Quick highlights

Five findings from Xobin’s first-half 2026 research

  1. 72%of 683 tested skill groups

    72% of 683 tested skill groups met Xobin’s benchmark criteria, with at least one tested AI setup qualifying for each group; the qualifying setup could differ between groups.[1]

  2. 52%average template weight

    EQ-related criteria averaged 52% of total weight across 113 templates from 92 employers; all other criteria averaged 48%.[2]

  3. Top 5across 35 technical roles

    Turning business requirements into technical solutions was among the five most requested skills across 35 technical roles.[3]

  4. 2024 → 2026request share

    Between the first halves of 2024 and 2026, traditional coding tests lost request share while AI code-generation, testing/debugging and explanation tasks gained it.[4]

  5. 24selected skill groups

    Each capability became more common within the same 24 skill groups selected using employer demand signals and tracked in 2024 and 2026.[5]

Xobin data through June 2026. Bracketed numbers refer to the data notes and references at the end of the report. Explore each finding for the chart, definitions and methodology.

The hiring question is expanding: what can a candidate produce, and how well can they direct, verify and improve AI-assisted work?

In the first half of 2026, Xobin tested AI models against the 683 skill groups in its 2024 assessment framework.[1] It also examined how employers weighted leadership scorecard templates,[2] which skills technical job descriptions and assessment requests asked for,[3] how the mix of requested technical assessments compared with the same months of 2024,[4] and how often two capabilities appeared in a fixed set of emerging skill groups across both periods.[5] Those are the five measured results in this report.

Our interpretation, stated as a question rather than a finding: if AI can meet the criteria for many tested tasks, should assessments evaluate the reasoning, checking and judgement behind an answer as well as the answer itself? Xobin did not measure prices, relative difficulty or economy-wide trends, and this report does not claim them.

Each finding counts something different: skill groups, scorecard templates, job roles, assessment requests and a fixed set of skill groups. They are separate sets of records and cannot be added together. Where a figure was supplied as a range, it is shown as a range and never turned into a single number. This is one benchmark plus several sets of Xobin’s own records, not a national hiring survey.

How to read the labels used throughout. "What this means" explains the measured result. "Our interpretation" is Xobin Research’s editorial reasoning about what a result could imply; it is not an additional finding, and a citation next to it supports only the factual premise, not the reasoning. "Illustrative example" marks a written scenario, never something observed in the data. "External context" summarises a separate publication dated 2026, several of which analyse activity from 2025; these add background and sometimes a counterpoint, and none of them confirm a Xobin number. "The takeaway" pairs Xobin Research’s interpretation with the figures behind it and a citation readers can reuse. Only the five numbered findings are Xobin results.

Finding 01 · AI delegability

72% of tested skill groups met criteria for full or partial AI delegation

Xobin tested all 683 skill groups classified as relevant in its 2024 assessment framework. Each group was set several tasks, and the outputs were produced by OpenAI and Anthropic models using tool and agent workflows. One expert then judged each group against written criteria. Across the benchmark, 72% of groups met the criteria for full or partial delegation to AI, with full and partial combined.[1]

Terms used in this report

Skill group:
a set of related capabilities in Xobin’s assessment framework.
AI-delegable:
in this benchmark, at least one tested AI model or workflow met the written criteria for doing all or part of the work. Partial delegation can still involve a person.
Reported whole-percent result:
a percentage supplied by the report author as a rounded whole number. The exact numerator and rounding rule were not supplied.

These are plain-language reading definitions, not a reproduced scoring rubric. The operational classification behind this finding is described in Xobin data note[1].

Figure 01Reported delegation rates by skill family, against the 72% overall rate

Swipe to explore chart

0%25%50%75%100%72% overallAnalytical reasoning and problem-solving90–100%Technical and digital skills72–89%Communication and language50–71%Leadership and interpersonal skills< 50%Operations and execution28%

Reported delegation rates and ranges. Ranges are reporting bands, not confidence intervals. Family denominators were not supplied, so the family figures cannot be weighted to reconstruct the 72% overall rate, and operations (28%) cannot be ranked against leadership (below 50%) because that band contains 28%. A dashed end marks an open bound.

Source: Xobin data note[1]

View chart data
Skill familyReported delegation rate
Analytical reasoning and problem-solving90% to 100%
Technical and digital skills72% to 89%
Communication and language50% to 71%
Leadership and interpersonal skillsbelow 50%
Operations and execution28% (reported single value)
All tested skill groups72% (reported whole-percent result; 683 groups)

How to read this chart

The horizontal scale shows the percentage of skill groups within each family that met the criteria for full or partial delegation. A bar shows the reported range; the dot marks the reported 28% operations result. The dashed end at 50% means the leadership result is below 50%. The 72% line marks the reported result across all 683 groups.[1]

What this means

AI’s tested capability varied by type of work. Analytical reasoning and problem-solving had the highest reported range, 90–100%. Technical and digital skills followed at 72–89%, and communication and language at 50–71%. Operations and execution was reported at 28%, and leadership and interpersonal skills was below 50%. Because the leadership band contains 28%, the data does not show whether operations sat above or below leadership. Across all 683 groups the reported rate was 72%.[1]

Three questions this result does and does not answer

It helps to keep three separate questions apart. First: can AI perform a task under test conditions? Second: can an employer use it reliably inside a real workflow, with real data, deadlines and consequences? Third: what must a person still understand and be accountable for once the tool is in place? This benchmark answers the first question only. It does not measure how often employers deploy AI, how much work it saves, or which jobs change.[1]

The families differed in their reported rates, from 90–100% for analytical reasoning and problem-solving to 28% for operations and execution. The benchmark judged outputs against written criteria and did not test why any family scored as it did, so this report offers no cause for the differences.[1]

Our interpretation: from producing an answer to judging one

Where an assessment permits AI and scores only a finished answer, that answer alone may reveal less about what the candidate can explain or verify. This benchmark did not test assessment validity or changes in candidate performance. The question we raise is conditional: how did the candidate choose the inputs, and how would they know if the answer were wrong? A task can become easier to produce while remaining important to understand.

If assessors wanted to see that, they would need to look at the choices behind an answer, the moment an error was caught, and the reasoning used to defend a decision. That is a design suggestion, not an observation from the data.

Illustrative example

An analyst uses AI to draft a forecast. Someone still has to ask whether the underlying data fits the question, how uncertain the result is, and what decision it should inform. The drafting got faster; the responsibility did not move. This scenario illustrates the distinction and is not a task from the benchmark.

External context

Anthropic’s January 2026 report distinguishes automated use from people iterating with Claude. Its analysis of November 2025 Claude.ai and first-party API activity uses model-based classifications and reports lower success on more complex tasks. These observations help explain why task capability, user involvement and dependable operation should be considered separately. They do not independently validate Xobin’s benchmark.[6]

Anthropic, 15 January 2026. anthropic.com. Context only; it does not validate the Xobin figures.

Methodology and scope

  • 72% describes full or partial delegation under the tested setups, with full and partial combined. It does not mean 72% of jobs disappear, that 72% of skills are unimportant, or that any work was actually replaced.
  • All 683 groups classified as relevant in the 2024 framework were tested. Each group contained multiple tasks, with the number varying by group. A group qualified if at least one tested model or workflow met the written criteria.
  • One expert per group judged the overall task evidence against written criteria, with no fixed task-pass-percentage cutoff and no independent second review.
  • Outputs were generated by 30 June 2026 using OpenAI and Anthropic tool and agent setups. Model versions were not supplied.
  • There was no comparison with human performance, and no claim that every tested model succeeded.
  • No exact numerator, separate full and partial breakdown, or family denominators were supplied. 72% is a reported whole-percent result and was not independently recomputed.

Source: Xobin data note[1]. Full method: method note 1.

Finding 02 · Leadership priorities

EQ averaged 52% of leadership scorecard weight

Xobin examined 113 distinct leadership scorecard templates from 92 employers, analysed for the January–June 2026 period. A template sets out what an employer plans to score in a leadership hire, and how heavily each part counts. Averaged across templates, EQ-related criteria received 52% of total weight and all other criteria combined received 48%.[2]

Terms used in this report

Emotional intelligence, or EQ:
in this report, the employer scorecard criteria Xobin grouped as EQ-related, broadly criteria about understanding and working with other people. The detailed coding key was not supplied, and this is not a validated EQ inventory.
Scorecard weight:
the share of the total evaluation assigned to a set of criteria. More weight means those criteria are intended to count more.
All other criteria:
every criterion not grouped as EQ-related, of any kind. The breakdown of this category was not provided.

These are plain-language reading definitions, not a reproduced scoring rubric. The operational classification behind this finding is described in Xobin data note[2].

Figure 02Average share of leadership scorecard weight, EQ-related criteria against all other criteria

Swipe to explore chart

52%
EQ-related criteria, combined average weight
48%
All other criteria, combined average weight

One cell is one percentage point of average total template weight across 113 templates. A cell is not an employer, a template or a candidate. The 48% covers all other criteria of any kind; it is not a technical-only figure and not a measure of intelligence.

Source: Xobin data note[2]

View chart data
Criteria groupAverage share of total template weight
EQ-related criteria (combined)52%
All other criteria (combined)48%
Templates113 templates from 92 employers

How to read this chart

Each square represents one percentage point of average scorecard weight. 52 of the 100 squares represent EQ-related criteria; 48 represent all other criteria. Squares are not candidates, employers or templates.[2]

What this means

Across 113 leadership scorecard templates from 92 employers, EQ-related criteria received an average 52% of total weight, against 48% for everything else combined. That is a small majority of the intended evaluation on average. It does not show that every template weighted EQ above half, and it is not proof that EQ matters more than any other single capability: the 48% combines all other criteria, and the breakdown of that category was not provided. It records what employers planned to score, not who they hired, who performed well, or how EQ compares with intelligence.[2]

Our interpretation: people skills could help AI adoption

In this report, EQ refers to the employer criteria grouped as EQ-related; the detailed coding key was not supplied. As an editorial illustration of leadership behaviour, and not a list of the traits Xobin coded, that could include surfacing concerns before they become failures, explaining a decision so that people can act on it, handling open disagreement without punishing it, and hearing a concern about a tool’s output without dismissing the person who raised it.

Those behaviours could matter when organisations put AI into daily work. A technically sound system can go underused if staff are afraid to admit a mistake it caused, or over-trusted if nobody feels able to challenge its output in front of a manager. That is a plausible mechanism for why people-related criteria could carry substantial weight in leadership hiring. Xobin measured the weight, not the mechanism.

Illustrative example

A manager announces a new AI-assisted workflow. An employee points out a case the workflow handles badly. What happens next is observable: does the manager ask for the evidence, listen without penalising the person who raised it, and decide how the exception will be reviewed? Those are the kinds of moments a people-weighted scorecard is trying to reach. They are our illustration of leadership behaviour, not traits Xobin scored.

External context

Microsoft’s 2026 Work Trend Index surveyed 20,000 AI-using knowledge workers across ten markets, alongside separate product-usage analysis. Survey respondents prioritised checking AI output and critical thinking, while supportive organisational conditions were associated with self-reported AI value. These associations concern AI users and their workplaces; they do not show that EQ causes successful AI adoption or confirm Xobin’s leadership scorecard weights.[7]

Microsoft, 5 May 2026. microsoft.com. Context only; it does not validate the Xobin figures.

Making leadership criteria observable

An employer can assign high weight to emotional intelligence and still measure it badly. A percentage on a template says nothing about the evidence behind the score. Employers could make these assessments more transparent by writing criteria as observable behaviours, using comparable scenarios and recording more than one example before assigning a rating. This is a practical proposal, not an observed practice among the employers in this study.

Methodology and scope

  • 113 distinct templates from 92 employers, analysed for January–June 2026. Template creation dates were not supplied for this report.
  • The 52% and 48% are template-level averages. Employers with multiple templates contribute more than once.
  • A combined category may receive more weight because it contains more criteria.
  • Templates state intentions. They are not completed candidate evaluations, hiring outcomes or evidence of a predictive relationship.
  • The detailed trait-to-category coding key was not supplied, and no IQ comparison was made.

Source: Xobin data note[2]. Full method: method note 2.

Finding 03 · Technical hiring asks

Business-to-technical translation ranked among the top five requested skills

Xobin reviewed 35 distinct technical roles represented in January–June 2026 job descriptions and assessment requests, with duplicate roles removed, across software, data, infrastructure and other technical work. Turning business requirements into technical solutions appeared among the five most frequently requested skills.[3]

Terms used in this report

Forward deployed engineer, or FDE:
an engineer who works closely with customers or business teams to turn a practical problem into a working technical solution. Business-to-technical translation is one capability associated with this job model.
Business-to-technical translation:
understanding the business need and turning it into clear technical requirements and a workable solution.
Requested skill:
a skill named in a role’s job description or assessment request. One role can request several skills.

These are plain-language reading definitions, not a reproduced scoring rubric. The operational classification behind this finding is described in Xobin data note[3].

Figure 03Five most frequently requested skills, count of roles out of 35

Swipe to explore chart

010203035AI integration and automation3035Analytical problem-solving3035Technical communication and stakeholders2029Software development and coding1019Translating business needs into solutions1019COUNT OF ROLES (OF 35)

Counts are reported as bands. Non-overlapping bands establish broad frequency differences. Exact counts and the order within the shared 30–35 and 10–19 bands are not known. Several skills may be requested for one role; counts must not be summed. This is a January–June 2026 snapshot with no earlier period to compare against.

Source: Xobin data note[3]

View chart data
Requested skillRoles requesting it (of 35)
AI integration and automation30 to 35 roles
Analytical problem-solving30 to 35 roles
Technical communication and stakeholder management20 to 29 roles
Software development and coding10 to 19 roles
Translating business requirements into technical solutions10 to 19 roles

How to read this chart

Each bar shows the reported band of roles requesting a skill, out of 35 distinct technical roles. A role may request several skills. Bands that do not overlap establish a broad difference in frequency; the order of skills that share a band is not known.[3]

What this means

Business-to-technical translation was one of the five most frequently requested skills in the roles reviewed. It appeared in 10–19 of 35 roles. AI integration and analytical problem-solving each appeared in 30–35 roles, technical communication and stakeholder management in 20–29, and software development and coding in 10–19. Translation and coding share the 10–19 band, so the data does not establish which was requested more often. AI integration and analytical problem-solving were more common, each in 30–35 roles. Exact ordering within each shared band is unknown.[3]

Our interpretation: defining the problem is part of the build

Business-to-technical translation sounds like a soft skill and behaves like a technical one. It covers three concrete jobs: deciding what outcome the work is actually for, writing down the constraints that the solution must respect, and knowing what evidence would show that the solution works. Get those wrong and a well-written system solves a problem nobody had.

The forward deployed engineer is one job model that puts this work at the centre, pairing an engineer directly with customers and business context. What Xobin observed is one associated capability being requested inside ordinary technical roles, across software, data and infrastructure work.[3] One skill ranking highly does not show that FDE jobs are growing or replacing software roles: there is no earlier sample for comparison. The data does not identify which roles requested which combinations of skills.

Illustrative example

A business team asks for a support chatbot. Before writing anything, the engineer asks whether the goal is shorter waiting times, more correct resolutions or better escalation to a human, because those three goals produce different systems. They check which customer information the system is allowed to use, and they agree in advance what result would count as success. Building straight from the original request could deliver a fast chatbot that resolves nothing. This scenario illustrates the capability; it is not a case Xobin observed.

External context

Anthropic analysed roughly 400,000 Claude Code sessions from October 2025 to April 2026. People generally made planning decisions while Claude handled execution; greater apparent task expertise was associated with higher session success. Both expertise and success were inferred from transcripts using classifiers, rather than independent tests of workers or real-world outcomes. This is relevant context for problem definition, not evidence of growth in FDE hiring.[8]

Anthropic, 16 June 2026. anthropic.com. Context only; it does not validate the Xobin figures.

What this could mean for engineers at either end of a career

For graduates, a portfolio project that only shows screenshots can leave a candidate’s reasoning unexplained. A project that explains the user problem, the constraints, the tests and the result gives hiring teams more evidence of the candidate’s reasoning. For senior engineers, domain knowledge can remain useful when evaluating AI output: knowing how a business actually operates is what lets someone judge whether an AI-generated answer is plausible. Neither point is a promise about job security, and neither was tested by this study.

Methodology and scope

  • A snapshot of 35 deduplicated roles represented in January–June 2026 job descriptions and assessment requests, across a broad technical mix.
  • There is no earlier sample for comparison, so this finding does not show rising demand over time or FDE job titles replacing software engineers.
  • No skill co-occurrence table was supplied, so exact combinations of requested skills and the roles containing them are unknown.
  • Non-overlapping bands establish broad frequency differences. Exact counts and the order within shared bands are not known, so the top five has a broad order but no exact internal rank.

Source: Xobin data note[3]. Full method: method note 3.

Finding 04 · The coding-assessment shift

LeetCode-style tests lost request share as AI-assisted coding tasks gained it

Xobin compared the technical assessments employers requested in January–June 2024 with the same months in 2026, using the same assessment definitions in both periods. Traditional coding tests fell from 75–100% of technical assessment requests to 25% to under 50%, while AI generation, testing and explanation tasks rose from below 25% to 50% to under 75%.[4]

Terms used in this report

Traditional coding assessment:
an algorithm or data-structure problem solved by writing code without AI. This report calls these LeetCode-style tests. The label describes the task type; the data is not from LeetCode and this is not a finding about that platform.
AI-assisted coding assessment:
a task in which the candidate generates code with AI, tests or debugs it, and explains the solution.
Request share:
the percentage of all technical assessment requests in a period that included that assessment type. A falling share does not mean a falling count.

These are plain-language reading definitions, not a reproduced scoring rubric. The operational classification behind this finding is described in Xobin data note[4].

Figure 04Share of technical assessment requests by type, 2024 against 2026

Swipe to explore chart

Jan to Jun 2024Jan to Jun 20260%25%50%75%100%Traditional coding (no AI)202475–100%202625% to < 50%AI generation, testing and explanation2024< 25%202650% to < 75%

Shares are reported as intervals; a dashed end marks an open upper bound. The denominator is all technical assessment requests in each period. The categories were analysed separately; overlap and coverage of all request types were not established, so the percentages must not be added and are not stacked. 2024 base: 500 to 999 technical assessment requests. 2026 base: 1,000 or more.

Source: Xobin data note[4]

View chart data
Assessment typePeriodShare of technical assessment requests
Traditional coding (no AI)Jan to Jun 202475% to 100%
Traditional coding (no AI)Jan to Jun 202625% to under 50%
AI generation, testing and explanationJan to Jun 2024below 25%
AI generation, testing and explanationJan to Jun 202650% to under 75%
Total technical assessment requestsJan to Jun 2024500 to 999
Total technical assessment requestsJan to Jun 20261,000 or more

How to read this chart

Compare each assessment type across January–June 2024 and January–June 2026. Bars show reported ranges, not exact percentages. A dashed end means that endpoint is excluded. The categories are shown separately and should not be added together.[4]

What this means

Traditional coding appeared in 75–100% of technical assessment requests in 2024, compared with 25% to under 50% in 2026. AI-assisted coding moved in the opposite direction: from below 25% to 50% to under 75%. These are shares of requests, not counts. The 2026 base was larger than the 2024 base, so a smaller share can still mean a similar or larger number of traditional tests being ordered. The employers behind the two periods may differ, so this is not a measured change within the same employers. Nothing here shows that LeetCode-style tests stopped working or that AI-assisted formats predict job performance better; it shows that employers asked for them less often relative to everything else.[4]

Our interpretation: working code does not show the whole picture

For roles where engineers are allowed to use AI, permitting it in an assessment can make the task closer to that working setup. That is an assessment-design proposal, not a measured validation. The weakness arrives at the same moment: if the deliverable is working code, and AI can produce working code, the finished artefact alone may not show how much the candidate understood. A candidate could pass a straightforward task and miss an edge case they never thought to look for.

The useful signal, in our reading, sits around the code rather than inside it: can the person explain what the program does, design a test that would catch a failure, and find the fault when one appears? Those are behaviours an assessor can watch. This is our reasoning about assessment design; Xobin measured what employers requested, not which format predicts job performance.

Illustrative example

A candidate asks an AI assistant for a booking function and gets working code in a minute. The assessment then asks what happens when two bookings land in different time zones, what the function does at the boundary of an open slot, and why the candidate chose the tests they wrote. A practical format built on that idea might run in five steps: restate the requirement, use AI freely, inspect and test the result, respond to one changed requirement, then explain the decisions. This is a proposal for discussion, not a format Xobin has validated.

External context

In a randomised study, 52 participants with Python experience learned the unfamiliar Trio library, with or without AI assistance. The AI-assisted group scored lower on an immediate skills quiz, while the overall difference in completion time was not statistically significant. This narrow learning experiment does not establish permanent skill loss or the predictive validity of a hiring test. It raises a practical question: how can an assessment distinguish producing code from understanding it?[9]

Shen, Judy Hanwen, and Tamkin, Alex, First submitted 28 January 2026; revised 1 February 2026. arxiv.org. Context only; it does not validate the Xobin figures.

The pipeline question this raises

Within Xobin’s requests, tasks that involve testing, debugging and explaining AI-generated code took a larger share in 2026 than in 2024.[4] Reviewing code can draw on understanding built through practice, debugging and study. If junior engineers delegate every difficult step from their first week, where does that practice come from? We raise this as a question prompted by one small learning study, not as evidence that workforce skills are collapsing.

Methodology and scope

  • The 2024 base was 500 to 999 technical assessment requests; the 2026 base was 1,000 or more. Exact counts by type and exact denominators were not supplied.
  • The same assessment definitions were used in both periods, but the employers behind the requests could differ. No within-employer change or cause is inferred.
  • These percentages describe request share, not total request volume, test validity or candidate performance.
  • The categories were analysed separately; overlap and coverage of all request types were not established, so the percentages must not be added.

Source: Xobin data note[4]. Full method: method note 4.

Finding 05 · Emerging human capabilities

Collaboration and non-linear thinking appeared in more tracked skill groups

Xobin selected 24 emerging skill groups using employer demand signals. The same groups were compared across January–June 2024 and January–June 2026 using consistent classification criteria. Collaboration and non-linear thinking were counted separately. Collaboration appeared in 4 groups in 2024 and 12–17 in 2026; non-linear thinking appeared in 0–5 groups in 2024 and 6–11 in 2026.[5]

Terms used in this report

Collaboration:
working with other people to combine knowledge, coordinate decisions and complete a shared task.
Non-linear thinking:
as read in this report, exploring more than one approach and changing direction when new information or constraints make the first approach unsuitable. This is an operational reading definition, not a claim that the ability is unique to people.
Emerging skill groups:
the 24 skill groups selected by Xobin’s research team using employer demand signals and tracked in both comparison periods.

These are plain-language reading definitions, not a reproduced scoring rubric. The operational classification behind this finding is described in Xobin data note[5].

Figure 05Emerging skill groups featuring each capability, out of 24 tracked groups

Swipe to explore chart

Jan to Jun 2024Jan to Jun 202606121824Collaboration · 20244 (exact)Collaboration · 202612–17Non-linear thinking · 20240–5Non-linear thinking · 20266–11GROUPS OF 24 TRACKED

The same 24 groups were tracked in both periods with identical classification criteria. The 2024 collaboration figure is an exact count; the other three values are reported bands, drawn as intervals with no midpoint. The groups themselves are not identified.

Source: Xobin data note[5]

View chart data
CapabilityPeriodGroups featuring it (of 24)
CollaborationJan to Jun 20244 (exact)
CollaborationJan to Jun 202612 to 17
Non-linear thinkingJan to Jun 20240 to 5
Non-linear thinkingJan to Jun 20266 to 11

How to read this chart

The scale runs from 0 to 24 groups. The dot shows that collaboration appeared in exactly 4 groups in 2024. Other bars show reported bands. Each capability was counted separately.[5]

What this means

Collaboration appeared in 4 of the 24 groups in 2024 and 12–17 in 2026. Non-linear thinking appeared in 0–5 groups in 2024 and 6–11 in 2026. Both became more common inside the same tracked set. The matched comparison shows how the emphasis on these two capabilities changed within the same demand-informed set of 24 skill groups.[5]

Our interpretation: behaviours, not slogans

Collaboration and non-linear thinking are easy words to put in a job advertisement and hard words to assess. Both describe things a person does that another person can observe. Someone collaborating asks a colleague for knowledge they do not have, states a shared goal out loud and changes a plan after hearing an objection. Someone thinking non-linearly holds a second option open, says why the first approach no longer fits and explains what changed their mind. Non-linear thinking is not erratic thinking, and it does not automatically mean more creative work; it means the person can revisit an assumption when conditions move.

When AI helps a team produce several plausible plans, choosing among those plans remains a separate task. In an illustrative workflow where people choose between AI-generated plans, that choice needs an agreed goal, a way of settling disagreement and attention to the consequences each option creates for other people. That is a reading of why these two capabilities could be gaining emphasis.

Illustrative example

A designer, an engineer and a salesperson each use AI to draft a launch plan. The three drafts are all defensible and they disagree. The team still has to settle which customer problem the launch is for, which deadline is real and which trade-off they accept. Two weeks later the budget changes and they revise the plan. The observable behaviours are asking for knowledge the others hold, explaining the revision and taking criticism into the next version. These behaviours are our illustration of what the finding could look like in practice, not a scored Xobin rubric.

Assessing how the pieces fit together

An assessment of one person working alone can miss how they coordinate with a team. A team can still coordinate badly while every member produces more text than before. The practical question for assessors is whether their process ever looks at how separate pieces of work fit together, or only at who produced the most. That is practical reasoning about assessment design, not a statistical result.

External context

The OECD’s June 2026 policy brief identifies continuing roles for problem-solving, creativity, managerial skills and social skills alongside AI. It also notes mixed evidence: some earlier employer research suggests reduced demand for certain social and emotional skills, so a universal upward trend cannot be assumed. The brief synthesises earlier OECD research; it is not a new 2026 survey and does not validate Xobin’s selected-group comparison.[10]

OECD, 5 June 2026. oecd.org. Context only; it does not validate the Xobin figures.

Methodology and scope

  • Xobin’s research team selected 24 emerging skill groups using employer demand signals.
  • The same groups and classification criteria were used for January–June 2024 and January–June 2026.
  • Collaboration and non-linear thinking were counted separately; results describe the frequency of each capability within these 24 groups.
  • The 2024 collaboration count is 4. The other three counts are shown at their reported ranges.

Source: Xobin data note[5]. Full method: method note 5.

Supporting material

Methodology, provenance and references

The Xobin findings use aggregate figures and methodology descriptions supplied by the report author. Underlying records were not provided for independent verification. External publications provide context and do not validate the Xobin results.

The report synthesises one completed benchmark and several observational Xobin records. It is not a representative national hiring survey. The units of analysis differ between findings: skill groups, scorecard templates, distinct roles, assessment requests and a fixed set of skill groups are separate populations and cannot be summed into a single sample size. Reported whole-percent figures and author-reported ranges are printed at the precision supplied. Bands shown in the figures are reporting ranges, not confidence intervals, and no midpoints, denominators or interval estimates have been inferred inside them.

No row-level dataset accompanies this edition, so there is no raw-data download. Every figure carries a "View chart data" table exposing the same aggregate values shown in the graphic, including open bounds. Where a value is an exact count it is labelled exact; where it is a reported single value or band it is labelled as such.

Evidence base by finding

FindingUnit of analysisPeriodData note
01 Delegability683 skill groupsOutputs generated by 30 Jun 2026[1]
02 Leadership113 templates, 92 employersJan to Jun 2026[2]
03 Technical hiring35 distinct technical rolesJan to Jun 2026[3]
04 Assessment mixTechnical assessment requests (500 to 999 in 2024; 1,000 or more in 2026)Jan to Jun 2024 and 2026[4]
05 Emerging skills24 selected emerging skill groupsJan to Jun 2024 and 2026[5]

Method notes

These five subsections are the method descriptions referred to by data notes [1] to [5]. They summarise what the report author supplied; they are not a reproduced scoring rubric.

Method note 1: AI-delegability benchmark (683 skill groups)

  • All 683 skill groups classified as relevant in Xobin’s 2024 framework were tested. Each group contained multiple tasks, with the number varying by group.
  • Outputs were produced with OpenAI and Anthropic tool and agent setups and generated by 30 June 2026. Model versions were not supplied.
  • A group qualified when at least one tested setup met the written criteria. One expert per group judged the overall evidence against those criteria, with no fixed task-pass-percentage threshold, no independent second review and no comparison with human performance.
  • Full and partial delegation are combined in the 72%. No exact numerator, separate full and partial breakdown or family denominators were supplied. 72% is a reported whole-percent result and was not independently recomputed.

Cited as Xobin data note [1].

Method note 2: Leadership scorecard weights (113 templates, 92 employers)

  • 113 distinct leadership scorecard templates from 92 employers, analysed for January–June 2026.
  • The template-level average gives EQ-related criteria 52% and all other criteria 48% of total weight. An employer with multiple templates contributes more than once, and the number of criteria in a category can affect its combined weight.
  • No completed candidate outcomes and no IQ comparison were included. The detailed trait-to-category coding key was not supplied.

Cited as Xobin data note [2].

Method note 3: Technical role requirements (35 distinct roles)

  • 35 deduplicated roles drawn from combined job descriptions and assessment requests for January–June 2026, across a broad technical mix.
  • Request frequencies are reported as bands, as supplied. Non-overlapping bands establish broad differences; exact counts inside a band are not known.
  • No historical sample, exact within-band counts or skill co-occurrence table were supplied.

Cited as Xobin data note [3].

Method note 4: Technical assessment request mix (Jan–June 2024 vs Jan–June 2026)

  • The same assessment definitions were applied to January–June 2024 and January–June 2026. Each figure is a share of all technical assessment requests in that period.
  • Base sizes were supplied as ranges: 500 to 999 requests in 2024 and 1,000 or more in 2026. Exact counts by type and exact denominators were not supplied.
  • The employer mix could differ between periods. No within-employer change or cause is inferred. The categories were analysed separately; overlap and coverage of all request types were not established, so the percentages must not be added.

Cited as Xobin data note [4].

Method note 5: Matched emerging skill groups comparison (24 groups)

  • Xobin’s research team selected 24 emerging skill groups using employer demand signals.
  • The same groups and classification criteria were used for January–June 2024 and January–June 2026.
  • Collaboration and non-linear thinking were counted separately; results describe the frequency of each capability within these 24 groups.
  • The 2024 collaboration count is 4. The other three counts are shown at their reported ranges.

Cited as Xobin data note [5].

References and data notes

Numbered in order of first appearance. Entries [1] to [5] are data notes within this report, not separately published studies or publicly verified datasets. Entries [6] to [10] are external publications dated 2026; several analyse activity recorded during 2025, as noted. They provide context and do not validate the Xobin results.

Xobin data notes

  1. [1]Xobin data note. Xobin-supplied aggregate data and methodology for the AI-delegability benchmark (683 skill groups). Supplied by the report author for this report. Method note 1 · Finding 01
  2. [2]Xobin data note. Xobin-supplied aggregate data and methodology for leadership scorecard weights (113 templates, 92 employers). Supplied by the report author for this report. Method note 2 · Finding 02
  3. [3]Xobin data note. Xobin-supplied aggregate data and methodology for technical role requirements (35 distinct roles). Supplied by the report author for this report. Method note 3 · Finding 03
  4. [4]Xobin data note. Xobin-supplied aggregate data and methodology for technical assessment request mix (Jan–June 2024 vs Jan–June 2026). Supplied by the report author for this report. Method note 4 · Finding 04
  5. [5]Xobin data note. Xobin-supplied aggregate data and methodology for the matched emerging skill groups comparison (24 groups). Supplied by the report author for this report. Method note 5 · Finding 05

External publications

  1. [6]Anthropic. Anthropic Economic Index report: Economic primitives. 15 January 2026. Company research report. Analyses November 2025 Claude.ai and first-party API activity, classified by model. anthropic.com · context in Finding 01
  2. [7]Microsoft. Agents, human agency, and the opportunity for every organization. 2026 Work Trend Index. 5 May 2026. Survey of 20,000 AI-using knowledge workers across ten markets, with separate product-usage analysis. microsoft.com · context in Finding 02
  3. [8]Anthropic. Agentic coding and persistent returns to expertise. 16 June 2026. Observational analysis of session transcripts. Roughly 400,000 Claude Code sessions, October 2025 to April 2026. anthropic.com · context in Finding 03
  4. [9]Shen, Judy Hanwen, and Tamkin, Alex. How AI Impacts Skill Formation. arXiv:2601.20245v2. First submitted 28 January 2026; revised 1 February 2026. Preprint, randomised study of 52 participants with Python experience learning the Trio library. arxiv.org · DOI · context in Finding 04
  5. [10]OECD. AI and skills: What we know so far. OECD Publishing. 5 June 2026. Policy brief synthesising earlier OECD research, not a new 2026 survey. oecd.org · DOI · context in Finding 05

How to cite

Sivabalan, G. (9 September 2026). Human vs AI Skills Report, 2026 Mid-Year Edition. Xobin Research. https://xobin.com/research/articles/human-vs-ai-skills-2026/

For methodology questions or press enquiries, get in touch.

§ FAQ

Frequently asked questions

What does the Human vs AI Skills Report 2026 measure?

It reports five findings from Xobin's own records for the first half of 2026: how many of 683 tested skill groups met the criteria for full or partial AI delegation, how much leadership scorecard weight went to EQ-related criteria, which skills were most requested across 35 technical roles, how the request share of technical assessment types changed between the first halves of 2024 and 2026, and how often collaboration and non-linear thinking appeared in 24 tracked emerging skill groups.Source: Xobin data notes [1] to [5] · Methodology

What does the 72% delegability figure mean?

Across the 683 skill groups classified as relevant in Xobin's 2024 framework, 72% were judged fully or partly delegable to AI when at least one tested model or workflow met the written criteria for that group. It is a reported whole-percent result describing capability under the tested setups. It is not a comparison with human performance and does not mean those skills stop mattering.Source: Xobin data note [1] · Method note 1

Why are several figures shown as ranges rather than exact numbers?

Some results were supplied as reporting bands rather than point estimates, so they are drawn and tabulated as bands. These are reporting ranges, not confidence intervals, and no midpoints or denominators have been inferred inside them.Source: Methodology and data notes
§ From the lab

Get in touch with the research team.

The methodology described here is what runs inside every Xobin assessment and AI interview, so hiring teams can make talent decisions on evidence, not instinct. See Xobin.