Short answer

Yes. An EQ score can change because the situation or testing occasion changed, but a difference alone cannot show why. First identify what the instrument asked you to report or do, then compare the sessions and the measure’s precision evidence. A repeated shift under comparable conditions, alongside a related behavior pattern, gives a stronger basis for interpreting change.

A changed score is a clue, not a change record

You retake an EQ test after a difficult stretch at work and the result is lower. Two explanations are possible: your emotional skills changed, or the occasion changed what the test captured. The second result is a numerical difference; it is not a record of personal change. To interpret it, first ask what the instrument asked you to report or do, whether the two sessions were comparable, and finally whether relevant behavior shifted beyond the test.

That first question matters because “EQ score” can refer to different kinds of evidence. In a community sample of 223, the study “The Assessment of Emotional Intelligence: A Comparison of Performance-Based and Self-Report Methodologies” found that the named performance test and self-report measure were not related to one another; their relationships with other variables also differed. The result concerns those two instruments there, not every EQ test. It still shows why a change between unlike measures cannot be read as a change on one shared scale.

Even with the same measure, a score can vary across occasions. ETS’s “Test Reliability—Basic Concepts” describes reliability in terms of how consistent test scores are across occasions, test editions, or raters, and explains measurement error: the unavoidable imprecision in a score. That general guide supplies no change threshold for an EQ instrument. So a higher or lower number may invite a question, but it cannot settle what changed. Before drawing that conclusion, identify the kind of result you compared.

Sources: The Assessment of Emotional Intelligence: A Comparison of Performance-Based and Self-Report Methodologies; Test Reliability—Basic Concepts

What do the two numbers measure?

Before asking why your EQ score changed, check whether the two results are answers to the same kind of question. There is no single EQ scale shared by every test. A self-report asks you to describe yourself: Its result reflects your account of typical behavior, An ability-based test instead presents emotion problems and scores responses according to the test’s scoring method. A situational judgment test gives described situations and asks you to choose or evaluate a response. The distinction matters when a person compares two reports and treats the numerical gap as a personal change. If the first test asked about recent habits and the second asked what someone should do in a written conflict scenario, the task itself changed. A difference may say something about the methods or the responses they invite; by itself it cannot show that emotional skill increased or declined. Before considering the occasion, ask what was measured, how the answer was produced, and what the score represents. A self-report can be useful for examining perceived habits, but it is still a report about oneself. A performance task samples answers under its own rules. Neither should be silently translated into the other’s meaning. A comparative study reported in “The assessment of emotional intelligence: a comparison of performance-based and self-report methodologies” makes this more than a wording concern. The researchers compared the Mayer-Salovey-Caruso Emotional Intelligence Test, a performance measure with problems thought to have correct responses, with a self-report EI measure. In a diverse community sample of 223 people, the two named measures were not related to one another. Their patterns of association also differed: self-reported EI related consistently to self-reported coping styles and depressive affect, while the performance measure showed stronger relations with age, education, and receiving psychotherapy. It does show why a change between different methods cannot be read simply as a change in one common quantity. A scenario test adds another distinction. “Development and Validation of a Situational Judgment Test of Emotional Intelligence” describes a 46-item measure built from emotional situations. Its reported factors concerned using one’s own emotion, sensing another person’s emotion, and understanding emotional context. That is different from reporting how often they acted in a certain way in ordinary life. A scenario answer may be relevant to how someone interprets a described emotional problem; it does not directly record what that person did in a real meeting or relationship. For a fair comparison, match the test name and version, language, instructions, response task, and scoring approach as closely as possible. Compare a self-report with another self-report about typical behavior, an ability task with the same kind of ability task, and scenario judgments with comparable scenarios. If any of those changed, write that down before interpreting the difference. Also check whether the report combines several parts into a total, or presents separate components; the same total can conceal different response patterns. The point is not to audit every technical detail before learning anything, but to establish whether the comparison is meaningful at all. This check does not decide whether you changed; it prevents a method change from being mistaken for personal development. Once the two results are genuinely comparable, the next questions concern the testing occasion and whether the shift appears again in behavior.

Sources: The Assessment of Emotional Intelligence: A Comparison of Performance-Based and Self-Report Methodologies; Development and Validation of a Situational Judgment Test of Emotional Intelligence

Stable over time does not mean identical in every situation

A trait-oriented EQ measure can be stable across months or years while a person’s momentary state, self-view, or behavior varies by situation. Stability describes how consistently a measure represents a broader pattern across occasions; it does not promise identical answers in every session. The distinction matters because a stable tendency and a changing experience can both be real. A lower result during a difficult period is not automatically evidence of lasting decline, but context alone cannot explain every difference either.

The 2024 study “Assessing the temporal stability of a measure of trait emotional intelligence: Systematic review and empirical analysis” examined the Trait Emotional Intelligence Questionnaire (TEIQue), a self-report measure of trait emotional intelligence. Its accessible abstract reports data collected over intervals from 30 to 1,444 days, ranging from about a month to four years. The authors report strong temporal stability at the global, factor, and facet levels. That finding pushes back against the easy claim that an EQ result merely records whatever happened that week: for this measure, the broad pattern showed substantial continuity across the studied intervals.

The result has a defined scope. It concerns the TEIQue and the intervals studied, and the abstract says that comparisons between different trait EI measures remain a future research need. It cannot tell a reader how much a result from another test should move, or whether a particular personal difference reflects growth. Nor does stability mean that every answer stays fixed. A measure can preserve broad differences among people over time while a person’s experiences, judgments, or behavior vary from one occasion to another. A retest result is therefore evidence about a pattern only to the extent that the instrument and its evidence support that interpretation.

Workplace conditions offer one plausible route for variation in current experience. In “An experience sampling investigation of workplace interactions, affective states, and employee well-being,” 60 full-time employees reported interaction characteristics and affect across ten workdays. The researchers analyzed 380 day-level observations and found that interaction characteristics were associated with affect and daily job satisfaction. The study measured work interactions, feelings, and satisfaction; it did not measure EQ scores or emotional-skill change. Its findings support the narrower point that a person’s affect at work can vary alongside daily interactions.

Applying that finding to a self-report EQ result is an inference, not a result tested in the study. If a questionnaire asks someone to summarize typical behavior or recent experience, current feelings may shape which episodes come to mind or how the person judges them. A tense stretch of work could change a response without showing that an enduring skill has changed. A different situation could also make a skill easier to use, while a repeated shift across comparable sessions and observable behavior might support a development explanation. These are possibilities, not conclusions from one score. The two studies make room for both continuity and variation: trait-level stability in one named measure, and day-to-day affective variation in a workplace sample.

Sources: Assessing the temporal stability of a measure of trait emotional intelligence: Systematic review and empirical analysis; An experience sampling investigation of workplace interactions, affective states, and employee well-being

How work conditions can change what you notice or report

A self-report about emotional habits depends partly on what comes to mind when you answer. A tense week can leave different episodes available for recall than a quieter one; it can also change how you judge your own response. That is a plausible route from situation to score when questions ask about recent behavior. It is an interpretation of the measurement task, not a finding that workplace stress directly lowers EQ scores.

The experience-sampling study An Experience Sampling Investigation of Workplace Interactions, Affective States, and Employee Well-Being followed 60 full-time employees across ten workdays. Participants reported interpersonal interaction characteristics and affect during each day, then job satisfaction at day’s end. Across 380 day-level observations, interaction characteristics were associated with affect and job satisfaction; affect also mediated the relationship between interactions and satisfaction. The study documents that work experience and affect can vary together within people over days. It did not administer an EQ test, compare EQ retests, or show that a demanding day changes emotional skill. Applying its finding to a recent-behavior questionnaire is therefore an inference: the work events that shape a person’s current emotional experience may also shape which examples they remember while describing their habits.

Several pathways can sit behind that inference. A state is a person’s current feeling, such as irritation after a sharp exchange. A recall window is the period or set of episodes a question asks the person to summarize. Opportunity means whether the situation gave them a chance to use the behavior at all. Self-appraisal is the judgment they make about how well they handled it. These can move separately. A person may feel more strained, remember more difficult conversations, have fewer chances to pause and clarify, or apply a harsher standard to the same response. A score that asks about typical behavior may blend these influences differently from one asking about a specific recent interval.

Consider this illustrative contrast. During a week with clear priorities and room to recover between meetings, a project coordinator notices frustration early and asks a colleague what a terse update means before replying. During a week of shifting deadlines and unclear decision authority, the same person may have little time to pause, may recall the rushed replies more readily, and may judge the week as evidence of poor regulation. The contrast does not establish that their underlying skill rose or fell. It identifies a question for interpretation: did the test capture a changed pattern of behavior, a different set of opportunities and remembered episodes, or both? The Dimotakis study makes the first part of that question credible by showing day-to-day links between workplace interactions and affect; it cannot settle the score attribution.

When a result is lower after a strained stretch, look at the questionnaire’s wording and recall period before treating it as a durable deficit. Note which kinds of work episodes were represented and whether you had a real chance to use the behavior being rated. If the score is intended to summarize recent habits, a changed work period may be relevant information about where a skill is harder to access. If it claims to measure something broader, the situation still gives context, but the test’s own evidence must show how much that context should affect interpretation. In either case, the score alone cannot tell you which pathway produced the difference.

Sources: An experience sampling investigation of workplace interactions, affective states, and employee well-being; The power of affect: Predicting intention to repeat workplace behavior from momentary affective states

Sometimes the test includes the situation by design

In a situational judgment test, the situation is part of what is measured. The person reads a described emotional situation and chooses or evaluates a response. A changed result can therefore reflect a different task, scenario, or interpretation of the prompt. That result does not directly say how the person usually behaves in everyday work. Before treating two scores as evidence of development, check whether the test asked the same kind of question both times. The 2013 study “Development and Validation of a Situational Judgment Test of Emotional Intelligence” illustrates this distinction. Its authors created 80 situations, each with three response alternatives, drawing on existing theoretical models. In data from 213 participants, their factor analysis produced a 46-item test with three factors: using one’s own emotion, sensing another person’s emotion, and understanding emotional context. They do not establish that every EQ test uses scenarios, or that a score from this test can be compared with a self-report profile as if both measured the same thing. A self-report habit question asks for a judgment about typical behavior or recent experience: for example, how often someone notices frustration before replying. A scenario item instead presents a constructed moment and asks what response fits, or how a person would respond. The first relies on self-description; the second samples judgment about a described situation. Neither automatically records what the person did in a real meeting. A person may recognize a constructive response in a vignette yet struggle to use it while receiving sharp feedback. That gap is a reflection question, not proof of hidden ability or inconsistency. Scenario content matters because different prompts can call for different judgments. A question about noticing one’s own emotion is not the same task as one about reading another person or understanding why a reaction changed. The 2013 study’s three factors make that point concrete: even within one scenario-based instrument, the reported dimensions were distinct. The abstract does not show that participants’ answers changed because they encountered different situations, or identify familiarity, wording, and interpretation as causes of score differences. Those are possibilities to check when comparing administrations, not findings to attribute to this study. There is still a practical value in noticing where a response seems easier to endorse than to carry out. Treat that contrast as a development hypothesis: perhaps the skill is available in a calm, clearly described example but harder to access under time pressure, uncertainty, or interpersonal strain. A scenario result cannot establish that explanation by itself. Look for the behavior in relevant episodes: Did the person name their own reaction, ask what the other person meant, pause before responding, or repair a misunderstanding? One episode should not stand in for a stable pattern. When two situational scores differ, compare the instrument version, instructions, and scenario demands before attributing the gap to personal change. If the prompts are not comparable, the scores may be answering different questions. If the same kind of task produces a repeated shift, it can be a useful signal to investigate alongside behavior over time. The test’s situation is part of the evidence, and a reason to ask what it says about life beyond the page.

Sources: Development and Validation of a Situational Judgment Test of Emotional Intelligence; Situational judgment tests: A review of recent research

Measurement error can mimic a personal shift

Even when the same skill and test are involved, every score carries some uncertainty. A small difference may reflect ordinary measurement variation rather than a meaningful change in the person. Whether the gap is larger than expected depends on evidence about that particular instrument’s precision. Without it, the two numbers show that the recorded results differ; they do not establish why.

The ETS memorandum Test Reliability—Basic Concepts describes reliability as score consistency across occasions, test editions, or raters. Practically, does the measure tend to give similar results when the underlying target has not substantially changed? Measurement error is the part of an observed result that makes it an imperfect estimate, rather than an exact reading. It can occur even when a test is administered as intended. The memorandum discusses standard error of measurement, a way to describe expected score imprecision, but gives no value for an EQ test. That information must come from evidence for the specific instrument and score.

This uncertainty is different from a situation effect. A person may answer differently because current strain changes what they notice or recall; that is a possible influence on the response, not simply random imprecision. Measurement error concerns the limits of the score as an estimate. Both can contribute to a difference, and one does not tell us how much of the other is present. Reliability evidence characterizes consistency; an account of the occasion may explain what was happening when the person answered.

A fair retest comparison starts with practical details: Was it the same instrument and version, language, response instructions, and scoring method? Were access conditions, timing, and broad circumstances reasonably comparable? A revised form or scoring rule can alter the comparison before personal development enters the picture. Even when details match, consistency evidence is needed to judge whether a gap exceeds expected measurement variation. A difference that looks substantial on a screen is not automatically substantial in measurement terms.

The Standards for Educational and Psychological Testing, a joint work of the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, guides test development and interpretation. Its relevance here is that evidence must fit the interpretation being made. General standards cannot supply an EQ instrument’s reliability estimate or threshold for meaningful change; those claims require information about the particular measure and score.

If a publisher provides no precision information or defensible method for deciding whether a retest difference exceeds expected uncertainty, do not turn the raw gap into a confident claim that you have changed. Treat it as a prompt to ask what differed between occasions and which observable behavior you want to examine. The numbers show that recorded results differ; interpreting that as personal change requires support they cannot provide alone.

Sources: Test Reliability—Basic Concepts; Standards for Educational and Psychological Testing; Standards for Educational and Psychological Testing

A split illustration shows one person working at a desk on the left and several people seated around a table with speech bubbles on the right, alongside documents and vertical bars.
A split illustration shows one person working at a desk on the left and several people seated around a table with speech bubbles on the right, alongside documents and vertical bars.

What would justify saying you have changed?

A change claim needs comparison, not just two numbers. Check whether both results came from the same measure and version, with the same instructions, language, scoring, and comparable conditions. If one session asked about recent habits and the other asked which response is best in a written scenario, the difference may describe two different tasks. It cannot yet show that an emotional skill changed. A different recall window or pressured work period changes it. Next ask whether the shift appears again. A single retest can reflect occasion-specific influences and ordinary measurement uncertainty. The ETS memorandum “Test Reliability—Basic Concepts” defines reliability through consistency across testing occasions, test editions, or raters, and explains error of measurement. Read a raw difference against evidence about how much this measure’s scores vary when the target is relatively stable. A general reliability statement cannot show whether this gap is meaningful. You need evidence for the instrument, version, scoring method, and comparison at hand. Without it, do not invent a cutoff or call a numerical gap reliable change. A repeated shift is more informative when a related behavior moves in the same direction over time. Choose a narrow behavior, such as pausing to clarify a colleague’s concern before defending a decision. Look across several relevant exchanges and record what happened, not what you assume another person felt. This is not another EQ test. It asks where the response appears and under what conditions. If the score changes but the behavior does not, your self-description may have shifted, the measure may not capture that behavior closely, or the difference may sit within its uncertainty. The mismatch calls for interpretation, not a verdict. There is a credible case for treating a repeated difference as evidence of development when the instrument supports the intended interpretation, administrations are comparable, its precision evidence makes the shift larger than expected variation, and relevant behavior also changes. Even then, the conclusion should stay proportionate: evidence of a developing pattern in observed situations, not proof the skill will appear everywhere. The UCL record for “Assessing the temporal stability of a measure of trait emotional intelligence: Systematic review and empirical analysis” reports strong temporal stability for the TEIQue at global, factor, and facet levels across intervals from 30 to 1,444 days. This weighs against attributing every change to context. It concerns one trait EI measure and cannot settle what a different EQ instrument’s retest means. The conclusion would become firmer with instrument-specific precision information and longitudinal evidence about the same measure, alongside relevant behavior. Until then, describe the result precisely: “My score differed on this retest, and I’m checking whether the pattern repeats.” If a work behavior has changed too, name what you observed and where it appears. A useful conversation might be: “I’ve been pausing to clarify before responding to difficult feedback. Have you noticed that in more than one project discussion?” The answer can add an observation; it cannot turn one score into a complete account of personal change.

Sources: Assessing the temporal stability of a measure of trait emotional intelligence: Systematic review and empirical analysis; Test Reliability—Basic Concepts; Standards for Educational and Psychological Testing

Use a short behavior record to locate the difference

An episode record can answer a narrower question than a retest: in which situations does a particular response show up, and what tends to happen just before it? Choose one behavior, such as pausing to clarify a terse message before replying. For a short period, note occasions when that behavior was relevant. The aim is to notice a pattern, not build a case for or against yourself. Keep each note brief and concrete. Record the setting and cue; what you felt or noticed; what you did next; and any observable response or later repair. “The project lead wrote ‘Need this today’; I felt tense, replied sharply, then asked to reset the conversation” separates an event, an internal cue, an action, and consequence. It is more useful than “I have poor regulation,” which turns one exchange into a verdict. Include different conditions where the behavior could be used: a routine clarification, a deadline, or a disagreement with someone who has different authority. If you record only the hardest moments, the log describes your response under strain, not your usual response. Record successes too, and note when the behavior was not an option. A chance to pause and ask a question differs from a conversation that has already moved on. This is a reflection aid, not a validated assessment procedure or experiment. In “An experience sampling investigation of workplace interactions, affective states, and employee well-being,” 60 full-time employees reported interactions and affect across ten workdays, producing 380 day-level observations. Analyses found associations between interaction characteristics and affect. The study measured daily affect and workplace interactions, not EQ scores or skill change. It supports attending to the conditions around an episode; applying that point to a personal record is a reasoned suggestion, not a result tested by the study. “Development and Validation of a Situational Judgment Test of Emotional Intelligence” describes an instrument built around emotional situations. It shows that scenarios can form part of an assessment task, but does not validate a personal log. After several entries, look for a bounded pattern. Does clarifying before replying happen in routine exchanges but become harder when feedback feels abrupt? Has the response appeared in more than one setting, or only with one person or kind of pressure? The pattern may suggest where a behavior is easier or harder to use. It cannot establish that an underlying EQ construct or score changed; notes reflect selected episodes, memory, and your account. Keep them private and time-limited. Choose one small practice, then see whether a later, comparable situation gives you a different opportunity to respond.

Sources: An experience sampling investigation of workplace interactions, affective states, and employee well-being; Development and Validation of a Situational Judgment Test of Emotional Intelligence

Use a profile as a behavior prompt, not a verdict

The Emotional Skills Profile is a prompt for reflection: it can help you examine recent emotional habits and judgments about described situations, then choose one behavior to practice. EQ Test describes it as a private, 32-item educational profile. Its recent-behavior summaries and authored-scenario judgments are separate report components, so read them as different kinds of information rather than one claim about your ability. The profile has no validated norms and is not a standardized ability test. A result cannot show how you compare with a population or establish that your emotional skill has changed. Suppose a result draws your attention to how you respond when someone is upset. Treat that as a question to investigate, not a verdict about how you always behave. You might choose one practice: before replying to tense feedback, pause to check what you heard and ask a clarifying question. Then notice whether you try this in ordinary interactions, what makes it easier or harder, and what happens next. Those observations can help you judge whether the practice is useful. They do not validate the score or prove that the profile caused a change. A profile can also give you language for a work conversation, if you choose to share it. You could describe the behavior you want to practice and ask a colleague for feedback on that behavior, rather than asking them to confirm a score. Keep the discussion focused on an observable exchange and your next step. The report is a private development aid: its value is helping you notice a pattern and select a practical experiment. Conclusions about change still depend on comparable evidence beyond the result. To begin that reflection, explore the Emotional Skills Profile at /assessment.

Sources: Emotional Skills Profile

Retest only when the next comparison can teach you something

Before retesting, decide what you want the comparison to answer. To ask whether a habit has shifted, use the same instrument, version, language, instructions, and scoring method where possible, under reasonably comparable conditions. A different questionnaire or task may produce a different signal without showing personal change. The useful interval depends on the measure and its stability evidence.

Pair a repeated result with a few episodes involving the behavior you care about. Note what was happening, what you did, and what followed. This record can show whether a response is becoming more available in different settings, but cannot calculate an EQ score or prove why a score moved. The ETS research memorandum explains why score consistency and expected measurement uncertainty matter. A change claim needs instrument-specific evidence that a repeated difference exceeds expected uncertainty, alongside a matching behavior pattern.

If the tool does not support score comparisons over time, treat another result as a reflection prompt and focus on behavior. The 32-item Emotional Skills Profile supports reflection on recent habits and scenario judgments. Explore it at /assessment, then initiate one conversation: ask a colleague what makes feedback easier to act on, and listen for one behavior to try.

Sources: Test Reliability—Basic Concepts; Emotional Skills Profile

Sources

  1. The Assessment of Emotional Intelligence: A Comparison of Performance-Based and Self-Report Methodologies

    Supports the distinction between two named performance and self-report measures in a community sample of 223.

  2. Assessing the temporal stability of a measure of trait emotional intelligence: Systematic review and empirical analysis

    Reports strong temporal stability for the TEIQue across the studied intervals.

  3. An experience sampling investigation of workplace interactions, affective states, and employee well-being

    Supports within-person links between workplace interaction characteristics and daily affect, not EQ score change.

  4. Development and Validation of a Situational Judgment Test of Emotional Intelligence

    Describes an emotional-intelligence judgment test built from situations and its reported factors.

  5. Test Reliability—Basic Concepts

    Explains score consistency, measurement error, and standard error of measurement as general testing concepts.

  6. The power of affect: Predicting intention to repeat workplace behavior from momentary affective states

    Assigned in the plan as further context for situational pathways; it does not establish EQ score change.

  7. Situational judgment tests: A review of recent research

    Assigned in the plan as context for scenario-based assessment; it does not establish why a particular score changed.

  8. Standards for Educational and Psychological Testing

    Describes the Standards as guidance for test development and interpretation.

  9. Standards for Educational and Psychological Testing

    Provides the Standards source assigned in the plan for test interpretation and score-use context.

  10. Emotional Skills Profile

    Supports the assigned product description as a private educational reflection profile with separate report components.

Apply it to the real situation

Turn an EQ result into one behavior to practice

From this guide: Use a result to choose a specific emotional habit to reflect on, then notice when that behavior is easier or harder to use.

A score difference cannot explain what changed in your work or what to practice next. The Emotional Skills Profile offers a private reflection on recent emotional habits and judgments about described situations. Use it to identify one behavior you want to notice, such as pausing to clarify before responding to difficult feedback. Then reflect on how that behavior appears across real situations. The profile is an educational development aid, not a normed ability score.

Explore the Emotional Skills ProfileView the report