Confidence alone is a weak guide to whether an emotional strength matches your behavior. Research found limited accuracy when professional students estimated their MSCEIT performance, but it did not test how employees predict everyday conduct. Before feedback, turn one claimed strength into a prediction about a visible action in a specific situation; afterward, compare it with what happened and an observer’s concrete account, if available.
Before the review, what does a strength claim predict?
Before a performance conversation, a person may feel sure they stay curious when a proposal is challenged. The confidence matters because it implies a prediction: when disagreement comes, I will ask what I have missed rather than defend the first idea. To find out whether that belief is informative, make the claim specific enough that an actual exchange could confirm or complicate it. “I am good at disagreement” is too broad to check on its own.
The answer depends on what “good at disagreement” refers to. It could mean feeling capable, answering emotion problems well on the MSCEIT, behaving constructively in a meeting, or receiving a favorable overall rating from a supervisor. Those are different observations. The direct study discussed here asks people to estimate their emotional-intelligence standing and their likely MSCEIT performance. It does not ask them to predict what they will do when a colleague challenges a proposal. So the research can tell us how well confidence tracked that test result; a workplace prediction needs to be checked against conduct and feedback from the relevant situation.
A useful prediction names the moment and the action: “When my proposal is challenged in Thursday’s review, I will ask one clarifying question before replying.” That sentence can be revisited after the conversation. It does not settle whether the person has a broad emotional strength, but it gives confidence something observable to answer to. The article’s evidence begins one step earlier, with estimates about a test, and the difference shapes what its findings can support.
What did participants have to judge?
The central evidence comes from “Emotionally Unskilled, Unaware, and Uninterested in Learning More: Reactions to Feedback About Deficits in Emotional Intelligence.” Across three studies, participants completed the MSCEIT before estimating their own general emotional-intelligence standing and their expected percentile on that test. Their estimates therefore had a defined comparison: how they thought they would rank against other test takers, rather than how a colleague would describe their behavior in a meeting.
The participants were professional students with work experience. Study 1 included 157 master's students; Studies 2 and 3 drew on MBA students. The paper describes several years of post-undergraduate work experience in these samples, a reason the question may resonate with people thinking about management and development. They were still course-based convenience samples from US universities, not a cross-section of employees or workplaces. The study task also matters: estimating test standing is narrower than forecasting how one will listen, regulate frustration, or repair a strained exchange under particular work conditions.
The later studies added feedback procedures after testing, so the researchers could examine reactions to disclosed results. There is an unresolved count discrepancy for Study 2: its methods section reports 66 MBA participants, while Table 2 labels the sample N=166. The paper does not explain the difference, so neither number can be silently treated as settled. That inconsistency is relevant when reading the study's later feedback findings, but it does not change what participants were asked to estimate before feedback: their general EI and expected MSCEIT percentile.
How closely did confidence track performance?
Across the three studies, participants placed their general emotional intelligence, on average, at the 77th percentile among U.S. adults, while their average MSCEIT performance was around the 41st percentile. When asked specifically to estimate their MSCEIT standing, they predicted the 74th percentile. These averages describe two different judgments: a broad view of one’s EI and a forecast of one’s place on the test just completed. Both estimates ran high relative to the test criterion. Yet the estimates were not wholly unrelated to performance. The association between people’s self-ratings and actual performance was weak, averaging r = .09 across studies; the relationship between predicted MSCEIT percentile and actual test performance was also about .09. The latter association was marginal in the paper’s analysis. In practical terms, knowing a participant’s estimate gave little help in ordering participants by how they performed, even though higher estimates tended, slightly, to accompany higher performance. The gap between the mean estimate and mean result therefore should not be read as proof that every person was equally mistaken, or that no one had any signal about their relative standing.
The group averages conceal the sharpest difference in the paper. People at the low end of the performance distribution made the largest errors: participants around the 10th percentile overestimated their general EI by roughly 63 to 69 percentile points, depending on the study, and overestimated MSCEIT performance by about 62 to 63 points. At the 90th percentile, participants were much closer, underestimating by approximately 5 to 20 points on the same estimates. This pattern describes how estimation error varied with performance in these samples; it does not identify which colleague will misread a tense meeting, or how accurately an individual predicts a particular workplace response. It does show why a single average can mislead: modest association across the whole group coexisted with substantial overestimation among lower scorers and comparatively smaller errors higher in the distribution.
There is a statistical objection to that shape. When a measure is imperfect and people tend to rate themselves generously, those with low measured results may appear to overestimate more, while those with high results may appear to underestimate, simply through regression toward the average. The authors of Emotionally Unskilled, Unaware, and Uninterested in Learning More: Reactions to Feedback About Deficits in Emotional Intelligence tested the objection by examining absolute error: the distance between each estimate and actual performance, regardless of whether the estimate was too high or too low. If the pattern were only a consequence of regression plus general overestimation, accuracy should stop improving and then deteriorate at the top, where underestimation was most likely to create a gap. Instead, the absolute-error analyses showed accuracy generally improving as performance rose, though the improvement slowed at high levels. At the 99th percentile the slope still indicated better accuracy, significantly so when studies were combined, for both general EI estimates and MSCEIT estimates. The authors therefore argued that a simple regression account did not explain the full pattern. That analysis addresses one alternative explanation for the observed shape; it does not turn the estimates into precise forecasts for an individual.
The result is best read at the level the task measured. Participants estimated percentile standing against a performance-based emotional-intelligence test, after completing it. They were not asked to estimate whether they would recognize frustration in a project handoff, hold a boundary in a one-to-one, or repair a conversation after speaking sharply. Percentile calibration on the MSCEIT and forecasting behavior in a particular team are different prediction problems. The first has an explicit score criterion; the second depends on what behavior is meant, the situation, and what counts as a useful observation. The findings cannot be transferred automatically from the former to the latter.
For these professional-student samples, confidence alone offered weak evidence about test performance, with the largest errors concentrated among lower performers and a small positive group-level relationship remaining. That combination matters: a confident estimate was not meaningless, but it was too imprecise to substitute for the measured result. It also leaves unanswered the question a worker may actually care about: whether a claimed strength will show up in a specific exchange at work. The study’s test estimates cannot answer that workplace question; they establish how far self-assessed standing and MSCEIT standing came apart in these samples.
The percentile language can obscure what “accuracy” means here. A prediction of the 74th percentile is a prediction of rank among the comparison group, not a percentage chance that a person will handle emotions well, nor a rating of each skill in a real interaction. The paper compares those expected ranks with actual MSCEIT outcomes and examines how far estimates missed the performance criterion. It does not report a simple pass-or-fail threshold separating people who know themselves from people who do not. A small positive correlation can coexist with a large average gap because the two statistics answer separate questions: the gap compares typical predicted and observed levels, while the correlation asks whether relative differences in estimates line up with relative differences in performance. Reading both avoids turning the result into the claim that every estimate was useless.
What happened after participants saw their scores?
In Study 2, participants received their MSCEIT results after first estimating their emotional-intelligence standing. They then rated how accurate they thought the test was. The study asked whether a concrete score changed people’s initial judgments, but the reported relationship was uneven: better performance went with higher ratings of the MSCEIT’s accuracy, while lower performance went with more doubt about it. The association between performance and perceived accuracy was r = .37 in the authors’ analysis. Participants who had performed less well, as a group, were more likely to question the measure after seeing their feedback. This is a relationship between test performance and a post-feedback rating in one study setting; it does not reveal why any particular person doubts a result.
The setting gave participants a tangible way to respond as well. After disclosing the results, the researchers offered them a discounted copy of The Emotionally Intelligent Manager. A participant who said they wanted the book was contacted a week later to pay for it and receive the copy. The decision was therefore more concrete than a general statement that development sounds worthwhile, although it remained a single purchase opportunity attached to a research exercise. It was not a measure of sustained practice, later improvement, or willingness to seek feedback from colleagues.
Performance was associated with interest in buying the book: higher MSCEIT performers were more likely to indicate interest and follow through than lower performers. In the paper’s table, book interest correlated .34 with actual EI performance; the authors also report that participants in the higher performance quartiles were more likely to buy. The direction runs against the simple expectation that people who might benefit most from a development resource would be most eager to purchase it. But one optional book purchase cannot establish a general motivational profile. Price, timing, interest in that particular book, or other unmeasured considerations could shape an individual choice, and the study does not show what purchasers did with the material afterward.
The combined pattern is narrower than a claim that negative feedback inevitably makes people defensive. In this group, lower scores were associated both with lower perceived test accuracy and with a lower likelihood of the offered purchase; higher scores went with more favorable accuracy judgments and a greater likelihood of buying. Those associations fit the authors’ account that feedback can be discounted, but Study 2 alone does not settle the motive behind a response. An employee may have sound grounds for questioning a measure—for example, if its task does not match the skill they need to develop. The study did not adjudicate individual objections or test employees responding to a manager’s review.
Study 2 also has a sample-count inconsistency already identified in its methods and results table, so the reported pattern should be read with that unresolved detail in view rather than attached to a confidently reconstructed denominator. Its clearest contribution is the sequence it documents: people saw their test results, rated the test’s accuracy, and faced an actual, discounted purchase choice with a later payment follow-up. The performance-linked differences across those responses make the feedback reaction worth examining. They do not show that questioning an EQ result is always evasive, or that purchasing a book demonstrates lasting development.
The researchers also analyzed the two Study 2 responses together in their account of feedback reactions: actual performance was related to the accuracy rating and to buying the book, while the accuracy rating itself predicted the purchase decision after performance was controlled. Their mediation analysis treated perceived accuracy as one possible link between performance and buying. That model is consistent with the idea that dismissing feedback can accompany less interest in development, but statistical mediation in this setting does not establish the private motive behind one employee’s objection. This detail sharpens the group-level association; it does not convert a book sale into evidence of changed conduct.
Can a person be steered toward one interpretation of feedback?
Study 3 of “Emotionally Unskilled, Unaware, and Uninterested in Learning More: Reactions to Feedback About Deficits in Emotional Intelligence” tested a narrower question than whether people simply reject bad news: does their response shift when a questionnaire has already committed them to one view of the test or of emotional intelligence? Before taking the MSCEIT, MBA participants were randomly assigned to rate either how accurate they expected the test to be or how relevant they thought emotional intelligence was. After they saw their results, the researchers asked them to evaluate both. The first rating constrained one interpretive route; participants had less room to dismiss that attribute without contradicting a position they had already recorded.
The reported pattern followed that difference in room to maneuver. Participants with lower MSCEIT performance rated the attribute they had not rated beforehand less favorably: those who had first judged the test’s expected accuracy were more negative about EI’s relevance after feedback, while those who had first rated EI’s relevance were more negative about the test’s accuracy. The paper interprets this condition-specific response as evidence that people can discount an unfavorable result through the attribute left open to judgment. Random assignment gives the comparison causal leverage inside this particular survey sequence: the prior rating condition differed by design, rather than being inferred from a naturally occurring difference between workers.
The manipulation also sharpens what Study 2 could not isolate. In that earlier study, participants received their result and could rate the test’s accuracy; a lower rating might reflect a considered objection, a reaction to the score, or both. Study 3 placed a prior commitment on one of the two possible judgments and then observed which judgment remained more available after feedback. The authors’ account is therefore about constrained interpretation after a specific ability-test result, not a general claim that disagreement with feedback proves defensiveness. The assignment does not reveal every participant’s private reasoning, and a person can have a sound reason to question a test’s relevance or accuracy.
The two conditions are useful because they make the alternative routes visible. If someone first says the test should be accurate, then receives an unfavorable result, criticizing EI’s relevance leaves the test judgment intact. If someone first says EI matters, criticizing the test’s accuracy leaves the relevance judgment intact. The design asks whether the more flexible evaluation becomes more negative among lower performers. That pattern is consistent with motivated discounting, but it does not measure a conscious strategy or establish that participants deliberately protected their self-image. Prior commitment may shape which explanation is easiest to endorse; the observed ratings show the shift in evaluation, while the psychological explanation remains the authors’ interpretation of that shift. This separation matters when applying the finding: a questionnaire can reveal a response pattern under controlled choices without giving an observer direct access to why a person chose it.
The outcomes after the ratings need their own labels. Study 3 asked about intentions to pursue development and willingness to pay for development resources. The authors report that these indicators, like the feedback judgments, were related to test performance; they were not observations of participants practicing a skill over time, changing conduct, or responding to later feedback. Nor should willingness-to-pay responses here be blended with Study 2’s actual discounted-book offer and follow-up purchase procedure. One is a stated response within Study 3; the other was a concrete choice in a separate study.
The study began with 157 MBA participants and analyzed 141 after exclusions. That is the study-specific base for the manipulation; it does not repair the separate unresolved Study 2 count discrepancy. For someone receiving workplace feedback, the useful implication is modest: a conclusion about a result may depend partly on which part of the feedback is still open for debate. The experiment shows that this can happen when researchers deliberately constrain one judgment before an ability test. It does not establish that an ordinary review conversation creates the same constraint, or that the same response will follow a manager’s comment about behavior. In a review, the test’s accuracy and the relevance of a skill may not even be the live questions; the person may instead disagree about a specific event, expectation, or observation. Those conversations require evidence about the behavior being discussed.
What does an ability-test score actually represent?
The MSCEIT score in these studies came from performance on emotion-related problems, which is why it can serve as a criterion for participants’ estimates of their test standing. The measure presents tasks involving emotional information and scores responses against a keyed answer approach. It is an ability-oriented test: the result summarizes how a person answered those problems under test conditions. It is not a tally of what that person usually says when a colleague challenges a plan. That distinction identifies what the self-estimate study compared without requiring the score to stand for every meaning of emotional intelligence. A performance task asks what the respondent can do with the presented material at that moment; a report of typical behavior asks what they say they tend to do. The formats answer different questions even when both use the label EI.
“Measuring emotional intelligence with the MSCEIT V2.0” reports how the instrument’s answer basis was examined. The validation study compared responses with judgments from 21 emotion experts and a standardization sample of 2,112 participants. The abstract reports overlap between those judgments, which supports the scoring approach as more than an arbitrary key: responses had agreement with answers endorsed by people identified as experts in emotion. The result does not mean experts can observe a test taker’s emotional life or settle what behavior will be effective in every workplace situation. It tells the reader something specific about how answers were evaluated on this instrument.
The same validation report describes reasonable reliability and confirmatory support for theoretical factor models. Reliability matters here because a score used as a comparison point should show some consistency; the paper’s abstract gives a qualitative summary, not coefficients to reproduce. Factor-structure support concerns whether patterns among the test’s tasks fit the proposed organization of emotional abilities. Those findings give the MSCEIT a documented measurement basis, while the abstract-only record available for this account does not supply enough detail to add numerical estimates or claim that every proposed interpretation has been settled.
That basis makes the original study’s question legible. Participants estimated their MSCEIT percentile after taking it; the researchers could compare their expected standing with performance on the instrument’s emotion problems. This is a defined criterion, and it is why the paper can report calibration against test performance rather than rely on participants’ own descriptions as both predictor and outcome. The estimate-performance relationship remains a claim about that comparison. Scoring against emotion problems does not directly record the habits involved in a tense handoff, receiving criticism, or repairing an exchange.
A separate comparison, “Convergent, Discriminant, and Incremental Validity of Competing Measures of Emotional Intelligence,” helps prevent the MSCEIT result from being mistaken for a universal EQ quantity. The study compared MSCEIT performance with two self-report measures, the EQ-i and SREIT. Its abstract reports minimal relations between the ability measure and the self-reports, while the two self-reports were moderately interrelated and had different external associations. The measures therefore do not collapse into interchangeable versions of one score. A person can answer emotion problems in a way that earns a stronger MSCEIT result and describe their typical tendencies differently on a self-report; the discrepancy reflects different measurement approaches, not automatically a contradiction to resolve by choosing one as the person’s whole truth. This comparison is a construct-level warning about substituting one instrument’s output for another’s. It does not validate every self-report, and it cannot tell a reader which measure best captures a particular employee’s day-to-day conduct.
For the study of self-estimates, MSCEIT performance is a useful comparison because its answers are scored under an ability-test framework with published validation work behind it. That makes the authors’ calibration question more concrete than asking whether participants felt emotionally skilled. The interpretation should stay attached to that evidence: how well people anticipated performance on the MSCEIT, not a transcript of everyday conduct. When the question shifts to recurring behavior at work, the evidence must shift too.
Sources: Measuring emotional intelligence with the MSCEIT V2.0; Convergent, Discriminant, and Incremental Validity of Competing Measures of Emotional Intelligence

Why can self-description and test performance part company?
A self-report records what people say about their perceived characteristics or typical behavior. An ability measure such as the MSCEIT scores performance on emotion problems. They concern emotional intelligence, but they ask for different kinds of evidence: one asks what a respondent believes tends to describe them; the other examines how they handle presented tasks. A gap between the results is therefore not automatically a scoring mistake or a failure of honesty. It may be the expected consequence of asking different questions.
The abstract of “Convergent, Discriminant, and Incremental Validity of Competing Measures of Emotional Intelligence” compares MSCEIT performance with two self-reports, the EQ-i and SREIT. It reports minimal relations between the ability test and either self-report, moderate relations between the two self-reports, and different associations with external measures. The pattern matters more than any one pairwise comparison: the two self-reports shared some variance with each other, while neither simply reproduced the performance result. The authors’ comparison treats these approaches as distinct measures, not interchangeable ways to generate one common EQ value. The abstract does not provide the sample, detailed procedures, or coefficients, so the relation should remain qualitative here.
That modest convergence makes sense when the response task changes. A person might describe themselves as usually noticing when others are uncomfortable, yet miss a cue in a structured emotion problem. Someone else might perform well on such a task while being less confident about how they usually respond in conversation. Neither example proves what that person does at work; they show why a report about typical behavior and a score on an ability task need not line up neatly. The two methods also have different sources of error: a self-description depends on what the respondent notices and chooses to report, while a performance score depends on the particular tasks and scoring approach.
The external associations reported in the study add another reason not to average the outputs. The abstract says the measures had different associations with external variables, which suggests that their differences were not merely two noisy readings of an identical quantity. But without the full details, it would overstate the paper to name particular outcomes, rank the instruments, or claim that one predicts real workplace conduct better. Its supported contribution is narrower: the choice of measure changes what evidence is being collected and what outside criteria it relates to.
Averaging also hides the distinction the comparison was designed to examine. Suppose two instruments return results that point in different directions: a combined figure would make it impossible to tell whether the respondent reported a habit, solved an emotion problem, or both. It could appear precise while answering neither question clearly. Keeping separate labels preserves useful follow-up: ask for an example of the reported habit, or inspect what kind of task produced the performance result. Those are not interchangeable checks, and the abstract gives no common scale on which to combine them.
For a reader comparing results, keep the instrument’s question attached to its result. Say “this is how I described my typical behavior” for a self-report, and “this is how I performed on these emotion tasks” for an ability measure. If the accounts diverge, look first at what each asked and what evidence each could capture. Do not treat the difference as an arithmetic error to erase by averaging scores onto a single scale; that would create a number whose meaning the comparison study did not establish. The discrepancy can instead identify two separate questions for reflection: what do I think I usually do, and what did I demonstrate on this task?
These measures may each be useful for different reflective questions. Their low convergence does not establish that either one is the true account of everyday behavior, and the comparison does not tell us how accurately a particular worker judges a specific strength before feedback. It does establish a practical discipline for reading test evidence: name the kind of response behind a result before drawing a conclusion from it.
Sources: Convergent, Discriminant, and Incremental Validity of Competing Measures of Emotional Intelligence
What can colleagues see that a test cannot?
A colleague’s impression starts from interaction: what they noticed while working with someone over time. An ability-test result starts from responses to tasks. Those sources can speak to different parts of the question, even when both are described as emotional intelligence. The useful detail in observer research is not simply whether peers agree with one another. It is what their shared view aligns with, how much it aligns with the person’s own report, and what the observers had a chance to see.
In Study 1 of “The Social Perception of Emotional Abilities: Expanding What We Know About Observer Ratings of Emotional Intelligence,” 931 full-time MBA students worked in 186 teams for about two months before rating one another. The design gave team members a period of shared activity from which to form impressions. Agreement among observers was modest, and self-observer alignment was limited. The paper also found modest links between observer ratings and some team or work criteria, with the pattern depending on the ability branch being considered. These are connected but separate results: teammates did not produce fully interchangeable views, and some observer-rated aspects related to criteria in that particular setting.
The same paper reports weak convergence between observer ratings and ability-test scores. That finding fits the distinction between a task result and a social impression. A colleague can see whether someone invites a quieter teammate into a discussion, notices tension during a handoff, or returns to a disagreement constructively. The test records responses to its own emotion tasks. A test cannot witness the actual exchange; a teammate’s account cannot show performance on tasks they never saw. The paper’s MBA team setting matters here: it demonstrates what ratings looked like after a bounded period of collaboration, not how every supervisor or peer would rate emotional skill in any job.
Modest agreement does not mean that teammates had no useful information. It means their judgments were not identical enough to treat a group impression as a single, self-evident fact. Each rater encountered the same person through a particular role and set of exchanges; one might see the person facilitating a meeting, another negotiating a deadline. The observer paper’s branch-dependent criterion links likewise caution against treating “observer-rated EI” as one undifferentiated quality. Some aspects related to some team or work criteria in the study, while other aspects did not show the same pattern. The study supports attention to what was rated and what it related to; it does not support a universal claim that peers can reliably identify overall emotional ability.
A separate empathy study offers a narrow parallel. Its accessible abstract reports analyses involving 152 primary participants and 357 informants. Self-other agreement was relatively stronger for a multi-item empathy questionnaire than for a single direct question, while informant ratings were largely independent of participants’ performance on an emotion-recognition task. The study concerns empathy, not overall EI or workplace performance. It illustrates why a colleague’s impression and task performance can carry different information, but does not establish which gives the more accurate account of someone’s emotional strengths.
Together, the papers make the colleague’s field of view central to interpreting a rating. A teammate may have repeated opportunities to see how someone handles disagreement in project meetings, but no basis for judging how that person regulates emotion in a private client call or under a different manager. Ask what the colleague actually observed, over what period, and in which situations. A concrete account of a recurring action can illuminate behavior in those conditions; a broad label such as “empathetic” or “good with emotions” gives less information about what the observer means.
That distinction changes how feedback can be used in practice. “You are not empathetic” gives the recipient little to check: it compresses an impression into a trait-like label and leaves the relevant scene unstated. “In yesterday’s handoff, you moved to solutions before I had finished describing the impact” identifies an observable moment and a possible response to discuss. The second phrasing does not make the observer’s account automatically correct; it makes its basis available for examination. The receiver can ask what else was happening, whether this has occurred in other shared situations, and what response would have helped. This is an application of the studies’ field-of-view distinction, not an outcome tested by either paper.
This is why an observed episode and a test response should be read side by side only with their origins intact. If a colleague says someone tends to cut off objections, the useful follow-up is the interaction they witnessed and what happened next. If an ability measure indicates difficulty with a particular kind of emotion problem, that result remains about task performance. One does not certify or cancel the other. The disagreement may reflect different situations, opportunities to observe, or aspects of emotional skill; these studies do not determine which explanation applies to any one person.
Observer ratings can add a view of behavior that a test cannot directly see, while the observer-rating study’s limited self-observer alignment and branch-specific criteria counsel against treating a peer judgment as a correction key. A colleague’s account becomes informative when its scope is clear: the behavior, setting, and shared interactions behind it. That bounded description is more useful for a conversation about development than asking whether the colleague or the test has the final word.
Sources: The Social Perception of Emotional Abilities: Expanding What We Know About Observer Ratings of Emotional Intelligence; The Self-Other Agreement of Multiple Informants on Empathy Measures and Its Relation to Empathic Accuracy
When are two workplace ratings even comparable?
Before asking which rating is right, check whether the self-rating and supervisor rating describe the same work. A self-assessment of “communication” may mean keeping colleagues informed across a project; a supervisor may be thinking of how the employee handles objections in meetings. Those judgments can differ without directly contradicting each other. The comparison becomes meaningful only when both people have a reasonably shared target, period, and basis for observation.
“Self-Other Agreement in Job Performance Ratings: A Meta-Analytic Test of a Process Model” pooled 128 independent samples and reported a self–supervisor correlation of .22; its corrected estimate was .34 across 115 samples and 37,752 people. These are associations across general job-performance ratings. They are not EQ findings, accuracy rates, or evidence that supervisors know a worker’s emotional strengths better. A correlation describes how ratings varied together across samples; it does not identify the more accurate rater in a particular conversation.
The same meta-analysis found that correspondence varied with job position and with the rating indicator, format, and content. That matters because the label on a rating form can hide unlike targets. “Works well with others” might invite one person to recall routine cooperation and another to recall a single tense exchange. If the form asks about overall performance while the conversation focuses on listening during disagreement, the two answers have not yet met on common ground. The abstract reports these moderators but not their direction, so it cannot tell us which format or content reliably produces closer agreement.
The distinction between correlation and accuracy is especially important when a worker hears that their self-view “doesn’t match reality.” The pooled correlation summarizes co-variation: higher self-ratings tended, on average, to accompany higher supervisor ratings to a limited degree. It does not tell us whether a given person’s rating was too high or too low, how far apart two scores were in a particular workplace, or whether a supervisor’s judgment matched some independent record of conduct. The corrected estimate adjusts the association under the meta-analysis’s assumptions; it is still an association, not a percentage of people who judged themselves accurately. The underlying studies also used performance-rating indicators, not a common direct measure of emotional skill.
A second distinction concerns who is agreeing with whom. “Similarity and Agreement in Self-and Other Perception: A Meta-Analysis” separates self–other agreement—the match between a person’s view and another person’s view—from consensus among observers, or similarity in how observers judge the same target. Its 24-study synthesis reports that group size moderated self–other agreement and assumed similarity more strongly than it moderated observer consensus and assimilation. A group can therefore converge on an impression without that consensus showing that the person shares it, or that the impression is accurate.
For example, imagine a worker and supervisor both using the broad label “doesn’t listen,” while a second colleague disagrees. That split alone does not show whether the worker has a blind spot: the observers may have seen different exchanges, or may use the label for different conduct. In an illustrative feedback conversation, they could instead ask whether the concern is about interrupting during planning meetings, which period is in view, and which occasions each person actually observed. If they mean one behavior in one setting, the employee can examine that account against their own recollection. This example illustrates a way to clarify the disagreement; it is not a finding from either meta-analysis.
Observer access can also be uneven within the same job. A supervisor may see scheduled reviews and final deliverables, while a peer may share the fast-moving conversations where a person reacts to disagreement. If one rating concerns calm responses in formal meetings and another reflects rushed messages during a deadline, averaging them would blur the condition that may matter. Asking each rater for the situations behind the judgment preserves that difference long enough to find out whether the behavior recurs across settings or belongs to a narrower circumstance.
A practical comparison rule follows: name the target behavior, specify the rating period, and establish what each observer had an opportunity to see. Then ask for an occasion that makes the rating interpretable. Someone who attended the relevant meetings can speak to those exchanges; a supervisor who saw only project outcomes may have a different evidence base. The worker’s account also has a scope: it reflects what they noticed and remember. These differences may explain a mismatch, but they do not decide it automatically. A concrete account can still be mistaken, and a person may genuinely miss a behavior others repeatedly see.
Adding opinions does not settle which interpretation is accurate. More observers can reveal that an impression is shared, or that views vary across roles and situations; neither result by itself establishes the truth of a claim about emotional skill. The meta-analyses describe patterns in ratings, not an adjudication procedure for an individual case. When ratings diverge, the most useful next step is to make their subject and evidence comparable enough to discuss: one action, during a named period, from people who had a chance to witness it.
Sources: Self-Other Agreement in Job Performance Ratings: A Meta-Analytic Test of a Process Model; Similarity and Agreement in Self-and Other Perception: A Meta-Analysis
What should you do before your next feedback conversation?
Before a relevant conversation, write one claimed emotional strength as a prediction about something visible: “When a colleague challenges my recommendation, I expect to ask what concern is behind the objection before I defend the plan.” Afterward, compare that prediction with what happened and, if available, one observer’s concrete account of the exchange. One episode is a prompt for reflection, not a stable estimate of your ability; the point is to give a broad self-description a specific behavior you can examine next time. The private 32-item Emotional Skills Profile can also help you reflect on recent emotional habits and judgments about authored scenarios. Explore the profile at [/assessment](/assessment), with optional report guidance at [/report](/report). Keep the note narrow: write the setting, action, and what happened next, without converting a single success or lapse into a verdict about your character. If the account differs from your prediction, ask what detail would help you see the moment more clearly or what response you want to try in a similar exchange. The profile is a reflection aid for those questions, not a hiring score or diagnosis, and it does not guarantee a particular workplace outcome. Bring that note to the conversation.
Sources
- Emotionally Unskilled, Unaware, and Uninterested in Learning More: Reactions to Feedback About Deficits in Emotional Intelligence
Primary evidence for advance estimates of general EI and MSCEIT performance, reactions to disclosed results, experimentally constrained feedback interpretations, and subsequent development intentions or choices among professional students. Preserve Study 2's methods N=66 versus results-table N=166 discrepancy; Study 3 analyzed 141 after exclusions.
- Measuring emotional intelligence with the MSCEIT V2.0
The validation report describes an ability-oriented MSCEIT and examines agreement on answers, reliability, and factor structure; the abstract reports agreement among 21 emotion experts and a 2,112-person standardization sample, reasonable reliability, and support for theoretical factor models.
- Convergent, Discriminant, and Incremental Validity of Competing Measures of Emotional Intelligence
The comparison of MSCEIT ability performance with EQ-i and SREIT self-reports reports minimal relations between MSCEIT and the self-reports, moderate interrelation of the self-reports, and different external associations. Measures do not offer interchangeable accounts.
- The Social Perception of Emotional Abilities: Expanding What We Know About Observer Ratings of Emotional Intelligence
Three studies totaling 2,521 participants examine observer-rated EI. Study 1 had 931 full-time MBA students in 186 teams after about two months of interaction; observer agreement was modest, self-observer alignment limited, associations with some team/work criteria modest and branch-dependent, and observer scores weakly converged with ability tests.
- Self-Other Agreement in Job Performance Ratings: A Meta-Analytic Test of a Process Model
A meta-analysis synthesizes 128 independent samples; abstract-reported self-supervisor job-performance correlation is .22 and corrected rho is .34 across 115 samples and 37,752 people. Agreement varies with job position and rating indicator, format, and content.
- Similarity and Agreement in Self-and Other Perception: A Meta-Analysis
The meta-analysis distinguishes self-other agreement from consensus among observers and reports group size as a stronger moderator of self-other agreement and assumed similarity than of observer consensus and assimilation.
- The Self-Other Agreement of Multiple Informants on Empathy Measures and Its Relation to Empathic Accuracy
The publisher abstract reports analyses involving 152 primary participants and 357 informants. Self-other agreement was relatively stronger for the multi-item Toronto Empathy Questionnaire than for a single direct empathy judgment; informant ratings were largely independent of participants’ emotion-recognition performance. This is a narrow empathy finding, not evidence about overall EI or workplace performance.
Apply it to the real situation
Turn a broad EQ impression into one behavior to examine
From this guide: If you are unsure whether a reported strength appears in your work conversations, start with a situation where feedback or disagreement could change your response.
The Emotional Skills Profile offers a private way to reflect on recent emotional habits and judgments about authored situations. Its result can help you choose a question to examine, such as how you recognize frustration or respond when a colleague disagrees. Pair that reflection with a specific example from your own experience to decide what behavior to notice next.
