Imas, intelligent multi-modal working memory assessment system, terminal and computerized adaptive testing system

CN122549565APending Publication Date: 2026-08-11SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-23
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]鉴于以上所述现有技术的缺点,本发明的目的在于提供一种iMAS智能多模态工作记忆测量系统、终端及计算机化自适应测试系统,用于解决传统方法依赖人工编制静态题库,无法动态生成理论驱动、难度可控、内容新颖的多模态刺激材料导致测量构念与理论模型严重不匹配、题目同质化导致公平性缺失以及文化局限性与生态效度低下等技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549565A_ABST
    Figure CN122549565A_ABST
Patent Text Reader

Abstract

The application provides an iMAS intelligent multi-modal working memory measurement system, a terminal and a computerized adaptive testing system, adopts a multi-modal stimulation generation engine to drive a cloud large language model with a structured generation instruction, generates multi-modal stimulation content in real time, and presents the question by fusing language, vision and spatial relationship description. The dynamic measurement equivalent control module extracts control parameters from the preset equivalence parameter matrix and binds them with the target psychometric properties, and then feeds them back to the engine. The adaptive difficulty adjustment module estimates the posterior distribution based on the Bayesian IRT adaptive algorithm, calculates the subsequent stimulation difficulty parameters according to the real-time response data of the subjects, and feeds them back. The real-time validity monitoring module calculates the standardized residual of each question to control the quality of the question. The application dynamically generates ternary stimulation with the help of LLM and forces multi-modal information binding to improve the theoretical fit; through the IRT parameter locking mechanism, the fairness and measurement equivalence of the question are guaranteed; the application can also generate localized stimulation according to the language background of the subjects to avoid cultural bias, and the adaptive algorithm shortens the test time, optimizes the cost, realizes 'testing as calibration', and makes the question parameter library evolve continuously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to an iMAS intelligent multimodal working memory measurement system, terminal, and computerized adaptive testing system. Background Technology

[0002] Working memory (WM), as the core of higher cognitive functions in humans, is clearly defined by the classic theoretical model—the Baddeley-Hitch multicomponent model (Baddeley & Hitch, 1974; Baddeley, 2000). This model clarifies that working memory is not a single system, but a dynamic processing system composed of four subsystems: the phonological loop, the visuospatial canvas, the central executive system, and the episodic buffer. These subsystems are responsible for processing auditory information, visuospatial information, attentional control and cognitive coordination, and multimodal information integration, respectively. Subsequent research has further confirmed that in real cognitive tasks, working memory must simultaneously process the encoding, retention, and manipulation of multimodal information flows, including linguistic, visual, spatial, and auditory information (Cowan, 2001). However, current mainstream working memory measurement tools, such as the Digit Span Test, Reading Span Test, Manipulative Span Test, and Spatial Span Test, all have design limitations and cannot fully reflect the true function of working memory.

[0003] Existing working memory measurement tools suffer from a severe mismatch between measurement constructs and theoretical models. These tools often assess only a single subsystem, such as verbal or spatial processing, in isolation. For example, digit span tests only activate the phonological loop, and spatial span tests only measure the visuospatial canvas, completely ignoring the essential characteristic of working memory: parallel processing and dynamic integration of multimodal information (Alloway & Alloway, 2010). This mismatch leads to low criterion validity of the measurement results, making it unable to effectively predict the working memory demands of complex cognitive performances requiring cross-modal coordination in the real world, such as art learning and clinical diagnosis.

[0004] Existing measurement tools suffer from issues such as item homogenization, cultural limitations, and low ecological validity. Traditional tests use fixed item banks, resulting in significant fluctuations in difficulty and high content similarity among items (e.g., letter sequences only change in order). This leads to increased measurement error, and fixed item banks are prone to practice effects and leakage risks, making it impossible to ensure the reliability of repeated measurements (Hicks et al., 2016). Furthermore, many existing tools are designed based on Western alphabetic scripts and Arabic numeral systems, leading to cultural biases for non-Latin-speaking participants and making it difficult to guarantee measurement fairness (Engel de Abreu et al., 2013). In addition, abstract symbolic tasks are disconnected from real-life scenarios, failing to reflect the actual working memory needs in applications such as education and healthcare, resulting in low ecological validity. The root cause of these shortcomings lies in the fact that traditional methods rely on manually compiled static item banks, failing to dynamically generate theoretically driven, controllable-difficulty, and novel multimodal stimulus materials. Summary of the Invention

[0005] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide an iMAS intelligent multimodal working memory measurement system, terminal and computerized adaptive testing system to solve the technical problems of traditional methods relying on manually compiled static question banks, which cannot dynamically generate theoretically driven, difficulty-controllable and novel multimodal stimulus materials, resulting in serious mismatch between measurement constructs and theoretical models, homogenization of questions leading to lack of fairness, and cultural limitations and low ecological validity.

[0006] To achieve the above and other related objectives, this invention provides an iMAS intelligent multimodal working memory measurement system. The system includes: a multimodal stimulus generation engine, used to drive a cloud-based large language model to generate multimodal stimulus content containing linguistic, visual, and spatial relational descriptions in real time, based on structured generation instructions, to display questions; a dynamic measurement equivalence control module, used to extract control parameters bound to target psychometric attributes from a preset equivalence parameter matrix and feed them back to the multimodal stimulus generation engine; an adaptive difficulty adjustment module, used to estimate the posterior distribution of the subject's ability based on the subject's real-time response data using a Bayesian IRT adaptive algorithm and calculate the stimulus difficulty parameters for subsequent tests, feeding them back to the multimodal stimulus generation engine; and a real-time validity monitoring module, used to calculate the standardized residuals of each question and perform question quality control.

[0007] In one embodiment of the present invention, the multimodal stimulus generation engine is configured to receive stimulus difficulty parameters and ability posterior distribution from the adaptive difficulty adjustment module and generate structured generation instructions; send the structured generation instructions to a cloud-based large language model to drive it to generate original text containing triples of language, visual description, and spatial relationship description; receive control parameters from the dynamic measurement equivalence control module, and render the original text of triples into standardized visual image stimuli, screen text stimuli, and interactive spatial stimuli according to the control parameters, so as to present the questions to the subjects.

[0008] In one embodiment of the present invention, the equivalence parameter matrix stored in the dynamic measurement equivalence control module includes a discrete parameter set corresponding to a specific psychometric attribute, calibrated through large-scale pre-experiments; the control parameters in the discrete parameter set include: stimulus presentation time, number of inter-stimulus interference terms, and response window duration.

[0009] In one embodiment of the present invention, the adaptive difficulty adjustment module is used to obtain the posterior distribution of the subject's ability parameters based on the subject's real-time response data using the Bayesian IRT adaptive algorithm, and to calculate the expected gain of Fisher information at the next difficulty level. The difficulty level with the largest expected gain is selected as the stimulus difficulty parameter for subsequent test rounds and output together with the posterior distribution of the subject's ability parameters to the multimodal stimulus generation engine.

[0010] In one embodiment of the present invention, the real-time validity monitoring module is used to calculate the standardized residual between the observed response pattern and the expected response pattern of each question; when the absolute value of the standardized residual is greater than a preset threshold, the corresponding question is marked as an abnormal question, triggering the multimodal stimulus generation engine to regenerate alternative questions, and feeding back the abnormal parameters to the equivalence parameter matrix for online learning updates.

[0011] In one embodiment of the present invention, the method of obtaining the posterior distribution of the subject's ability parameters based on the Bayesian IRT adaptive algorithm according to the subject's real-time response data includes: after each preset number of trials, the response data is weighted by using the entropy value of the current ability posterior distribution of each trial as the weight of the response data of each trial, and the posterior distribution of the subject's ability is calculated.

[0012] In one embodiment of the present invention, the subject's response data includes: response data of the subject's synchronously executed n-back matching task and response data of the cross-modal information binding judgment task.

[0013] In one embodiment of the present invention, during the first round of testing, the corresponding norm is first called from the parameter database based on the input demographic variables to generate initial ability values, which are then sent to the multimodal stimulus generation engine.

[0014] To achieve the above and other related objectives, the present invention provides an electronic terminal, comprising: one or more memories and one or more processors; the one or more memories are used to store a computer program; the one or more processors are connected to the memories and are used to run the computer program to perform the functions of the iMAS intelligent multimodal working memory measurement system.

[0015] To achieve the above and other related objectives, the present invention provides a computerized adaptive testing system, comprising: a testing terminal, a display, a network communication module, a cloud-based large language model inference server, and a parameter database server; the testing terminal is used to run the iMAS intelligent multimodal working memory measurement system as described; the display is used to show questions to the test subject; the network communication module is used to transmit data between the testing terminal and the cloud-based large language model inference server; the cloud-based large language model inference server is used to run a cloud-based large language model and provide generation services; and the parameter database server is used to store and provide test parameters, norms, and calibration data.

[0016] As described above, this invention is an iMAS intelligent multimodal working memory measurement system, terminal, and computerized adaptive testing system, which has the following beneficial effects: This invention employs a multimodal stimulus generation engine that uses structured generation instructions to drive a cloud-based large language model, producing multimodal stimulus content that integrates language, visual, and spatial relationship descriptions to present questions in real time. A dynamic measurement equivalence control module extracts control parameters from a preset equivalence parameter matrix, binds them to the target psychometric attributes, and feeds them back to the engine. An adaptive difficulty adjustment module, based on the Bayesian IRT adaptive algorithm, estimates the posterior distribution of ability based on the subject's real-time response data, calculates subsequent stimulus difficulty parameters, and feeds them back. A real-time validity monitoring module calculates the standardized residuals of each item to control item quality. This invention utilizes LLM to dynamically generate ternary stimuli and force multimodal information binding, improving theoretical fit; it ensures item fairness and measurement equivalence through an IRT parameter locking mechanism; it can also generate localized stimuli according to the subject's language background, avoiding cultural bias; and the adaptive algorithm shortens testing time and optimizes costs, achieving "testing as calibration," allowing the item parameter library to continuously evolve. Attached Figure Description

[0017] Figure 1 The diagram shown is a structural schematic of the iMAS intelligent multimodal working memory measurement system according to an embodiment of the present invention.

[0018] Figure 2 The diagram shown is a schematic representation of the multimodal stimulus content display in one embodiment of the present invention.

[0019] Figure 3The diagram shown is a structural schematic of an electronic terminal according to an embodiment of the present invention.

[0020] Figure 4 The diagram shown is a structural schematic of a computerized adaptive testing system according to an embodiment of the present invention.

[0021] Figure 5 The diagram shown is a structural schematic of a computerized adaptive testing system according to an embodiment of the present invention. Detailed Implementation

[0022] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0023] It should be noted that in the following description, reference is made to the accompanying drawings, which illustrate several embodiments of the invention. It should be understood that other embodiments may also be used, and changes in mechanical composition, structure, electrical system, and operation may be made without departing from the spirit and scope of the invention. The following detailed description should not be considered limiting, and the scope of the embodiments of the invention is defined only by the claims of the published patents. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. Spatially related terms, such as “upper,” “lower,” “left,” “right,” “below,” “below,” “lower part,” “above,” “upper part,” etc., may be used herein to illustrate the relationship between one element or feature shown in the figures and another element or feature.

[0024] Throughout this specification, when it is said that a part is "connected" to another part, this includes not only "direct connection" but also "indirect connection" by placing other elements in between. Furthermore, when it is said that a part "includes" a certain constituent element, unless otherwise stated otherwise, this does not exclude other constituent elements, but rather means that other constituent elements may also be included.

[0025] The terms "first," "second," and "third," etc., used herein are for the purpose of describing various parts, components, regions, layers, and / or segments, but are not limiting. These terms are used only to distinguish one part, component, region, layer, or segment from others. Therefore, the "first part," "component," "region," "layer," or "segment" described below may refer to a "second part," "component," "region," "layer," or "segment" without departing from the scope of this invention.

[0026] Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” indicate the presence of the stated feature, operation, element, component, item, kind, and / or group, but do not preclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms “or” and “and / or” as used herein are interpreted as inclusive, or mean any one or any combination thereof. Thus, “A, B, or C” or “A, B, and / or C” means “any one of: A; B; C; A and B; A and C; B and C; A, B, and C.” Exceptions to this definition arise only when combinations of elements, functions, or operations are inherently mutually exclusive in some manner.

[0027] This invention provides an iMAS intelligent multimodal working memory measurement system. It employs a multimodal stimulus generation engine that uses structured generation instructions to drive a cloud-based large language model, producing multimodal stimulus content that integrates linguistic, visual, and spatial relationship descriptions to present questions in real time. A dynamic measurement equivalence control module extracts control parameters from a preset equivalence parameter matrix, binds them to target psychometric attributes, and feeds them back to the engine. An adaptive difficulty adjustment module, based on a Bayesian IRT adaptive algorithm, estimates the posterior distribution of ability based on real-time response data, calculates subsequent stimulus difficulty parameters, and feeds them back. A real-time validity monitoring module calculates the standardized residuals for each item to control item quality. This invention utilizes LLM to dynamically generate ternary stimuli and enforces multimodal information binding, improving theoretical fit; it ensures item fairness and measurement equivalence through an IRT parameter locking mechanism; it can also generate localized stimuli according to the subject's language background, avoiding cultural bias; and the adaptive algorithm shortens testing time and optimizes costs, achieving "testing as calibration," allowing the item parameter library to continuously evolve.

[0028] The present invention will now be described in detail with reference to the accompanying drawings, so that those skilled in the art can readily implement it. The present invention can be embodied in many different forms and is not limited to the embodiments described herein.

[0029] like Figure 1 This diagram illustrates the structure of an iMAS intelligent multimodal working memory measurement system according to an embodiment of the present invention.

[0030] iMAS stands for Intelligent Multimodal Adaptive System, and the system includes:

[0031] The multimodal stimulus generation engine 1 is used to drive a large cloud-based language model to generate multimodal stimulus content, including linguistic descriptions, visual descriptions, and spatial relationship descriptions, in real time based on structured generation instructions to display the questions.

[0032] The dynamic measurement equivalence control module 2 is used to extract control parameters that are bound to the target psychometric attributes from the preset equivalence parameter matrix and feed them back to the multimodal stimulus generation engine 1.

[0033] The adaptive difficulty adjustment module 3 is used to estimate the posterior distribution of the subject's ability based on the subject's real-time response data and calculate the stimulus difficulty parameters for subsequent tests based on the Bayesian IRT adaptive algorithm, and then feed them back to the multimodal stimulus generation engine 1.

[0034] The real-time validity monitoring module 4 is used to calculate the standardized residuals of each question and to perform question quality control.

[0035] In one embodiment, the multimodal stimulus generation engine 1 first receives stimulus difficulty parameters and capability posterior distribution from the adaptive difficulty adjustment module 3. Based on this key information, it generates structured generation instructions. Subsequently, the engine precisely sends these structured generation instructions to a cloud-based large language model (LLM) to drive the LLM to perform its work. The LLM generates original text containing triplets of linguistic, visual, and spatial relational descriptions according to the instructions. This process is not arbitrary but has clear goals and specifications. To ensure that this process is standardized and orderly and that the output results are traceable and verifiable, a parameterized prompt template is used, which explicitly includes four essential fields: {modality type, difficulty level, stimulus features, and equivalence ID}, thereby strictly constraining the output of the LLM.

[0036] The generated triplet original text is sent to the dynamic measurement equivalent control module 2, and then the multimodal stimulus generation engine 1 receives the control parameters fed back from the dynamic measurement equivalent control module 2.

[0037] Based on the received control parameters, the multimodal stimulus generation engine 1 performs detailed and precise rendering processing on the original triplet text. It transforms the text into standardized visual image stimuli, allowing participants to visually perceive relevant images; into screen text stimuli, presenting information in clear text form; and into interactive spatial stimuli, providing participants with opportunities to interact with spatial elements. Through these three different forms of stimulation, the current test items are presented to participants comprehensively and vividly, ensuring that participants receive information from multiple dimensions, thereby completing the test more accurately.

[0038] In practice, the generation of multimodal stimulus content has specific requirements. For example, it requires generating a ternary combination containing screen text stimuli (such as 2-4 syllable words), visual image stimuli (such as detailed descriptions of the colors and shapes of everyday objects), and spatial stimuli (such as changes in the position of a 3×3 grid). The large language model generates a qualified original text description. This description is not only rich and accurate in content but also cognitively aligned with the participants' comprehension abilities, ensuring the validity and reliability of the test.

[0039] This innovative approach utilizes a large language model as a "real-time question factory" for cognitive testing, a significant breakthrough that overcomes the limitations of traditional static question banks. Unlike simple text replacement in traditional methods, this solution employs a meticulously designed structured prompt process to rigorously control the output of the large language model. This ensures that the output accurately matches pre-defined Item Response Theory (IRT) parameters across three key dimensions: semantic complexity, visual features, and spatial topology. Ultimately, this achieves the technical effect of infinitely random content generation while maintaining constant measurement attributes, bringing greater flexibility, accuracy, and scientific rigor to cognitive testing.

[0040] In one embodiment, rendering the original triplet text into standardized visual image stimuli, screen text stimuli, and interactive spatial stimuli includes:

[0041] For on-screen text stimuli, the relevant text is presented directly in the center of the screen, ensuring that the participant's attention is quickly focused there. Simultaneously, considering the visual characteristics and reading needs of participants of different ages, the font size is adaptively adjusted according to the participant's age. For example, the font size is appropriately enlarged for children to ensure they can read clearly and easily; while for adult participants, the font size is adjusted to a more moderate and comfortable size. Through this adaptive adjustment, participants of different ages can have a clear and comfortable reading experience, thereby more accurately understanding the test content.

[0042] When processing visual image stimuli, a professional text-to-image generation model, such as StableDiffusion, is invoked. This model possesses powerful image generation capabilities, accurately converting item descriptions generated by Large Language Models (LLM) into standardized images. Specifically, it transforms the descriptions into 128×128 pixel icons, a standardized size that facilitates unified display and management within the testing interface. Simultaneously, to ensure consistency and stability in visual presentation, the key parameters of the icons—brightness and contrast—are strictly locked using control parameters extracted from the dynamic measurement isometry control module 2. This ensures that the generated visual image stimuli maintain the same visual effect regardless of the testing environment, preventing different visual perceptions from occurring due to parameter variations, thus avoiding any impact on the accuracy of the test results.

[0043] For interactive spatial stimuli, a 2D / 3D graphics engine is used to render the spatial descriptions generated by the large language model into interactive virtual grids or stereoscopic views. For example, if the description is a change in position within a 3×3 grid, the system will generate a 3×3 virtual grid, which participants can interact with using interactive devices such as a mouse or touchscreen. To eliminate the interference of device differences on test results and ensure that all participants have the same experience during interaction, the frame rate of the motion trajectory is fixed at 60fps. This setting creates a stable and uniform interactive environment for participants, allowing their operational responses to more realistically reflect their cognitive abilities without being affected by differences in device performance.

[0044] As shown in Figure 2, the multimodal presentation interface of the testing system has a reasonable layout, divided into four main areas: a language stimulus area, a visual stimulus area, a spatial stimulus area, and a response button area. The specific layout is as follows:

[0045] Screen Partitioning: The entire screen is divided into three areas: top, middle, and bottom. The top area is the language stimulus area: This area presents language descriptions relevant to the test. For example, the example text "blue square" clearly conveys key information to the participant. The middle area is the visual stimulus area: This area displays 128×128 pixel icons generated by a text-to-image generation model. Using the example text above, a blue square icon is presented here, corresponding to the language description above, providing the participant with intuitive visual information. The bottom area is the spatial stimulus area: This area presents interactive spatial stimuli, such as a 3×3 grid in the example, highlighting specific cells, such as the cell in the 2nd row and 3rd column, to guide the participant's attention to specific spatial locations and encourage interaction. The response button area: Located at the bottom center of the screen, this area features "match / non-match" response buttons. Participants need to make a judgment by clicking the corresponding button based on their comprehensive understanding of the language, visual, and spatial stimuli. On either side of the response button are displayed a countdown timer (1500ms remaining) and a trial progress indicator (8 / 20). The countdown timer reminds the subject to complete the operation within the specified time, while the trial progress indicator lets the subject know the current progress of the test.

[0046] In one embodiment, during the first round of testing, the system first retrieves the corresponding norms from the parameter database based on the input demographic variables to generate an initial ability value, which is then sent to the multimodal stimulus generation engine 1. Inputting demographic variables such as the subject's age, education level, and language background, the system retrieves the corresponding norms from the parameter database to generate an initial ability value θ0 (standard score M=0, SD=1).

[0047] In the first round of testing, the system followed a specific data processing workflow. First, based on the input demographic variables, it precisely retrieved corresponding norm data from a pre-built parameter database. These demographic variables covered key information such as the participants' age, education level, and language background. The parameter database stored a large amount of norm data based on different combinations of demographic characteristics; this data underwent rigorous statistical analysis and validation, ensuring high reliability and validity. After obtaining the input demographic variables, an efficient retrieval algorithm quickly located and extracted the corresponding norms. Based on the retrieved norms, an initial ability value was generated using a specific computational model and sent to the multimodal stimulus generation engine 1. This initial ability value was presented as a standard score θ0, with a mean M set to 0 and a standard deviation SD set to 1. This standard score setting ensured the comparability of initial ability values ​​among different participants, providing a unified standard for subsequent test analysis and result interpretation.

[0048] In one embodiment, the dynamic measurement equivalence control module 2 stores an equivalence parameter matrix. This matrix contains a set of discrete parameters carefully calibrated through large-scale pre-experiments, which closely correspond to specific psychometric properties.

[0049] The control parameters in the discrete parameter set cover multiple key dimensions. The stimulus presentation time is set according to different modalities: 2000ms ± 100ms for the language modality and 3000ms ± 150ms for the visual modality. This fine-tuning aims to adapt to the information processing characteristics of different sensory channels. The number of inter-stimulus interference items is represented by the lure ratio in the n-back task, specifically set at 30%. This ratio has been repeatedly verified to effectively balance task difficulty and the cognitive load of the participants. The response window duration is set within a dynamic range of 2500ms to 5000ms, providing participants with reasonable reaction time while avoiding interference with measurement results due to excessively long or short durations.

[0050] These control parameters were not set arbitrarily, but rather derived from large-scale preliminary experiments (sample size N > 5000). During the experiment, Item Response Theory (IRT) was used to rigorously calibrate the parameters. Through the IRT model, it is possible to deeply analyze the impact of each parameter on the psychometric characteristics of the items (including difficulty b, discrimination a, and guessing parameter c), ensuring that any generated items maintain a high degree of consistency in these key characteristics.

[0051] Based on the above parameter framework, the Large Language Model (LLM) must strictly follow established rules when generating content. This means that although the generated content is random, ensuring the diversity and novelty of the test, the measurement attributes remain constant, thus providing a solid guarantee for dynamic measurement equivalence and ensuring the comparability and reliability of measurement results across different test rounds and items.

[0052] In one embodiment, the participants' response data encompassed two key components: response data from the simultaneous n-back matching task and response data from the cross-modal information binding judgment task. The experiment employed a dual-task paradigm to comprehensively and accurately assess participants' working memory integration capabilities. The primary task was an n-back variation task, requiring participants to judge whether a currently presented stimulus matched stimuli presented n trials prior. This task design effectively examined participants' short-term storage and retrieval abilities for sequential information, as well as their ability to distinguish and compare different stimulus features. The secondary task was a modal integration judgment task, such as asking participants questions like, "Did you see the red ball just now in the grid to the left of the voice prompt?" This task required participants to integrate and judge information from different modalities (such as visual and auditory), necessitating the construction of a unified contextual representation in their minds to accurately answer the question.

[0053] This dual-task paradigm design has unique scientific value. Through this design, it is possible to effectively measure the binding of multimodal information, filling the long-standing technical gap of lacking direct measurement tools for working memory integration function, and providing new research methods and perspectives for a deeper understanding of the integration mechanism of working memory.

[0054] In one embodiment, the adaptive difficulty adjustment module 3, based on the Bayesian IRT (Item Response Theory) adaptive algorithm, estimates the posterior distribution of the subject's ability parameters by analyzing the subject's real-time response data. Subsequently, it calculates the expected gain of Fisher information at the next difficulty level and selects the difficulty level with the largest expected gain as the stimulus difficulty parameter for subsequent test rounds. This parameter, along with the subject's posterior distribution of ability, is output to the multimodal stimulus generation engine.

[0055] Preferably, after a preset number of trials, the module employs a Bayesian IRT adaptive algorithm, where the response data from each trial is weighted by the entropy value of the current ability posterior distribution, thereby updating the posterior distribution of the subject's ability. Based on this distribution, the expected gain of Fisher information at the next difficulty level is calculated, and the difficulty level with the largest expected gain is selected as the stimulus difficulty parameter for subsequent test rounds.

[0056] For example, after every 3 trials, the system updates the posterior distribution of the subjects' abilities (denoted as θ'), then calculates the expected gain, and selects the stimulus difficulty parameter that maximizes the improvement in measurement accuracy (e.g., n ranges from 1 to 4, corresponding to a stimulus load of 2 to 6 items). The selected parameter is then sent to the Large Language Model (LLM) for the next round of stimulus generation.

[0057] In one embodiment, the real-time validity monitoring module 4 is used to monitor the validity of test items. Its core function is to calculate the standardized residual between the actual observed response pattern and the theoretically expected response pattern based on the IRT model for each item in the current test round. When the absolute value of the standardized residual of an item exceeds a preset threshold, the item is marked as anomalous, and the multimodal stimulus generation engine is triggered to regenerate a replacement item. Simultaneously, the parameters of the anomalous item are fed back into the system's equivalence parameter matrix for online learning and updating. Specifically, the validity filter built into the module executes the following process: First, it calculates the residual matrix for each item (comparing the observed response pattern with the IRT expected pattern). If the absolute value of the standardized residual of an item |Z| > 2.5, it is marked as an anomalous item. Subsequently, the system triggers the Large Language Model (LLM) to regenerate a replacement item and feeds back the anomalous parameters to the equivalence parameter matrix for online updating.

[0058] In addition, the module also records the reaction time distribution of the participants to identify possible random responses or fatigue effects (e.g., reaction times shorter than 200 milliseconds or longer than 10,000 milliseconds). When such patterns are detected, the system will automatically issue a rest prompt or terminate the test according to the strategy.

[0059] The iMAS intelligent multimodal working memory measurement system provided in this embodiment of the invention can be implemented on the terminal side or the server side. For the hardware structure of the electronic terminal, please refer to [link to relevant documentation]. Figure 3 This is a schematic diagram of an optional hardware structure of an electronic terminal 1000 provided in an embodiment of the present invention. The terminal 1000 can be a mobile phone, computer device, tablet device, personal digital processing device, factory back-end processing device, etc. The terminal 1000 includes: at least one processor 1001, a memory 1002, at least one network interface 10010, and a user interface 1009. The various components in the device are coupled together through a bus system 1005. It is understood that the bus system 1005 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 1005 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 All buses are labeled as bus systems.

[0060] The user interface 1009 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.

[0061] It is understood that memory 1002 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.

[0062] In this embodiment of the invention, the memory 1002 is used to store various types of data to support the operation of the terminal 1000. Examples of this data include: any executable program for operation on the terminal 1000, such as the operating system 10021 and application program 10022; the operating system 10021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 10022 may contain various applications, such as a media player, browser, etc., for implementing various application services. The implementation of the iMAS intelligent multimodal working memory measurement system provided in this embodiment of the invention can be included in the application program 10022.

[0063] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1001 or by instructions in the form of software. The processor 1001 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1001 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 1001 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in a memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.

[0064] In an exemplary embodiment, the terminal 1000 may be used to execute the aforementioned method by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).

[0065] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented using computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0066] In the embodiments provided in this application, the computer-readable and writable storage medium may include read-only memory, random access memory, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, flash memory, USB flash drive, portable hard drive, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Additionally, any connection may be appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. However, it should be understood that computer-readable and writable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are intended for non-transient, tangible storage media. The disks and optical discs used in the application include compact optical discs (CDs), laser optical discs, optical discs, digital multifunction optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs use lasers to copy data optically.

[0067] like Figure 4 This illustration shows a schematic diagram of a computerized adaptive testing system according to an embodiment of the present invention. The computerized adaptive testing system consists of a test terminal, a display, a network communication module, a cloud-based large language model inference server, and a parameter database server. Through the collaborative work of the various hardware components and real-time interaction with the large language model (LLM), a closed-loop feedback intelligent evaluation system is formed, which can efficiently and accurately measure the working memory of the test subjects.

[0068] The testing terminal is the core execution device of the system, undertaking the crucial task of running the iMAS Intelligent Multimodal Working Memory Measurement System. As the initiator and interaction center of the entire testing process, it coordinates the work of various parts of the system to ensure smooth testing. The testing terminal can be various common electronic devices, such as computers and tablets. It possesses powerful computing and processing capabilities, enabling stable operation of the iMAS Intelligent Multimodal Working Memory Measurement System. The system precisely controls the testing process and pace based on test parameters obtained from the parameter database server and question information generated by the cloud-based large language model inference server. For example, it adjusts the difficulty and type of subsequent questions in real time based on the subject's answers, achieving adaptive testing. Simultaneously, the testing terminal collects the subject's answer data and transmits it to other servers via network communication modules for further analysis and processing. Furthermore, the testing terminal receives real-time responses from the subject using input devices such as keyboards and mice, including clicking to select answers and dragging elements, and promptly feeds this data back to other parts of the system for processing.

[0069] A monitor is a visual output device that interacts directly with the test subject. Its main function is to clearly and accurately display the test questions to the test subject and provide them with an intuitive test interface.

[0070] The network communication module acts as a bridge connecting the test terminal and the cloud-based large language model inference server, responsible for the efficient and stable transmission of data between the two. It utilizes wired network (such as Ethernet) or wireless network (such as Wi-Fi, 4G / 5G) technologies to ensure fast and accurate data transfer between the test terminal and the cloud-based large language model inference server. During testing, the network communication module promptly uploads information such as question requests and participant response data sent by the test terminal to the cloud-based large language model inference server; simultaneously, it accurately downloads the results generated by the cloud-based large language model inference server and the feedback control parameters to the test terminal, ensuring the smooth operation of the entire testing process.

[0071] The cloud-based large language model inference server runs a large cloud-based language model and provides generation services. This server possesses powerful computing capabilities, enabling it to rapidly process massive amounts of text data. It receives question generation requests from test terminals via a network communication module, combines them with relevant test parameters obtained from a parameter database server, and uses a large language model (such as a Transformer architecture model with hundreds of billions of parameters) to generate original multimodal question text that meets the testing requirements. For example, based on parameters such as the test taker's ability level and the required test difficulty, it generates information text containing specific linguistic descriptions, visual features, and spatial relationships. After generating the original text, the server returns the results to the test terminal and may also store relevant data in the parameter database server as needed for subsequent analysis and calibration.

[0072] The parameter database server is the system's data storage center, responsible for storing and providing test parameters, norms, and calibration data.

[0073] Since the implementation principle of the iMAS intelligent multimodal working memory measurement system has been described in the foregoing embodiments, it will not be repeated here.

[0074] In one embodiment, such as Figure 5 The testing terminal connects to a monitor via HDMI / DisplayPort to present multimodal stimuli including text, images, and videos. Participants input responses using a keyboard / touchscreen, and the data is encrypted and transmitted to the cloud via 5G / WiFi. The cloud-based LLM inference server generates stimulus descriptions and equivalence parameter packages based on the parameter database information and sends them back, while the local rendering engine executes the rendering. The system also features an optional eye tracker / EEG interface for collecting physiological data to verify the physiological validity of the test, achieving an efficient, scientific, and comprehensive testing process.

[0075] Compared with the prior art, the present invention has the following advantages:

[0076] 1. A revolutionary improvement in the matching degree of theoretical concepts.

[0077] This protocol utilizes LLM to dynamically generate a three-dimensional stimuli of language, vision, and space, compelling participants to bind multimodal information within a task and directly measuring the integrative function of the context buffer. Preliminary experiments with N = 800 validated this protocol, demonstrating a significant improvement in construct validity compared to traditional unimodal tests, achieving Δr = 0.31 - 0.42. Furthermore, the protocol showed a high correlation (r = 0.68) with fMRI brain functional connectivity indicators (prefrontal-parietal network activation intensity), while traditional tools showed a correlation (r < 0.35). This clearly demonstrates a qualitative leap in theoretical construct matching, enabling a more accurate reflection of participants' cognitive abilities and providing a more reliable theoretical basis for related research.

[0078] 2. Fairness and measurement equivalence are absolutely guaranteed.

[0079] The dynamic measurement equivalence control module of this scheme employs an IRT parameter locking mechanism, which ensures the interchangeability of randomly generated items in terms of psychometric properties. Experimental data strongly demonstrates this: when the same subject used this system in different test sessions (with a 2-week interval), the standard error of the score difference (SEM) was only 2.1 points, while the SEM of the traditional reading span test was 5.7 points, reducing the measurement error by 63%. This means that this system can completely eliminate item similarity bias and the practice effect, fully meeting the stringent fairness requirements of the Educational and Psychological Measurement Standards (AERA, 2014), and providing a fair and impartial testing environment for all subjects.

[0080] 3. Maximize cultural adaptability and ecological effectiveness.

[0081] LLM can generate localized stimuli in real time based on the participant's language background, such as Chinese idioms, Japanese kana, and Arabic symbols, effectively avoiding the influence of cultural bias on test results. At the same time, the stimulus content is closely based on real-life scenarios, such as remembering shopping lists and planning navigation routes, which significantly improves the ecological validity coefficient (related to daily forgetfulness questionnaires) from r = 0.28 of traditional tools to r = 0.51. This improvement allows the system to be directly applied not only to educational diagnosis but also to predicting math performance. It can also play an important role in clinical screening, with a sensitivity of up to 92% in identifying mild cognitive impairment, providing a more practical testing tool for the education and clinical fields.

[0082] 4. Measurement efficiency and economic costs are significantly optimized.

[0083] By employing an adaptive algorithm, the test length can be shortened by 50% to achieve the same reliability as traditional methods. The time per test is significantly reduced from 45 minutes to 18 minutes, greatly improving testing efficiency and saving participants' time and effort. Furthermore, the cloud-based LLM can serve an unlimited number of concurrent users, with marginal costs approaching zero. Compared to the traditional method's ¥200 printing and labor costs per paper-and-pencil test, this system's cost per test is less than ¥0.5, a 99.75% reduction. This cost advantage makes this system highly economically feasible for large-scale testing applications, enabling it to provide affordable testing services to more organizations and individuals.

[0084] 5. Possesses real-time quality control and continuous evolution capabilities.

[0085] This solution implements the advanced function of "testing as calibration," and the question parameter library can be continuously optimized through online learning. The automatic elimination mechanism for abnormal questions further ensures the quality of the question bank; after 6 months of operation, the question bank quality index (QualityIndex = average discrimination / residual variance) improved by 40%. This demonstrates that the system has self-iterative capabilities, continuously adapting to new testing requirements and environmental changes, completely avoiding the problem of lagging version updates in traditional tools, consistently maintaining the accuracy and effectiveness of testing, and providing users with a consistently high-quality testing experience.

[0086] In summary, the iMAS intelligent multimodal working memory measurement system, terminal, and computerized adaptive testing system of this invention employ a multimodal stimulus generation engine that uses structured generation instructions to drive a cloud-based large language model, producing multimodal stimulus content that integrates language, visual, and spatial relationship descriptions to present questions in real time. The dynamic measurement equivalence control module extracts control parameters from a preset equivalence parameter matrix, binds them to the target psychometric attributes, and feeds them back to the engine. The adaptive difficulty adjustment module, based on the Bayesian IRT adaptive algorithm, estimates the posterior distribution of ability based on the subject's real-time response data, calculates subsequent stimulus difficulty parameters, and feeds them back. The real-time validity monitoring module calculates the standardized residuals for each item to control item quality. This invention utilizes LLM to dynamically generate ternary stimuli and enforces multimodal information binding, improving theoretical fit; it ensures item fairness and measurement equivalence through the IRT parameter locking mechanism; it can also generate localized stimuli according to the subject's language background, avoiding cultural bias; and the adaptive algorithm shortens testing time and optimizes costs, achieving "testing as calibration," enabling continuous evolution of the item parameter library. Therefore, this invention effectively overcomes the various shortcomings of the prior art and has high industrial application value.

[0087] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. An iMAS intelligent multi-modal working memory measurement system, characterized in that, The system includes: The multimodal stimulus generation engine is used to drive a large cloud-based language model to generate multimodal stimulus content, including linguistic descriptions, visual descriptions, and spatial relationship descriptions, in real time based on structured generation instructions to display the questions. The dynamic measurement equivalence control module is used to extract control parameters that are bound to the target psychometric attributes from the preset equivalence parameter matrix and feed them back to the multimodal stimulus generation engine. The adaptive difficulty adjustment module is used to estimate the posterior distribution of the subject's ability based on the subject's real-time response data and calculate the stimulus difficulty parameters for subsequent tests, based on the Bayesian IRT adaptive algorithm, and feed them back to the multimodal stimulus generation engine. The real-time validity monitoring module is used to calculate the standardized residuals of each question and to perform question quality control.

2. The iMAS intelligent multimodal working memory measurement system according to claim 1, characterized in that, The multimodal stimulus generation engine is used to receive stimulus difficulty parameters and ability posterior distribution from the adaptive difficulty adjustment module and generate structured generation instructions; send the structured generation instructions to the cloud-based large language model to drive it to generate original text containing triples of language, visual description, and spatial relationship description; receive control parameters from the dynamic measurement equivalence control module, and render the original text of triples into standardized visual image stimuli, screen text stimuli, and interactive spatial stimuli according to the control parameters, so as to present the questions to the subjects.

3. The iMAS intelligent multi-modal working memory measurement system as claimed in claim 1, wherein, The equivalence parameter matrix stored in the dynamic measurement equivalence control module includes a discrete parameter set corresponding to specific psychometric attributes, calibrated through large-scale pre-experiments; the control parameters in the discrete parameter set include: stimulus presentation time, number of inter-stimulus interference terms, and response window duration.

4. The iMAS intelligent multi-modal working memory measurement system as claimed in claim 1, wherein, The adaptive difficulty adjustment module is used to obtain the posterior distribution of the subject's ability parameters based on the subject's real-time response data using the Bayesian IRT adaptive algorithm, and to calculate the expected gain of Fisher information at the next difficulty level. The difficulty level with the largest expected gain is selected as the stimulus difficulty parameter for subsequent test rounds and output together with the posterior distribution of the subject's ability parameters to the multimodal stimulus generation engine.

5. The iMAS intelligent multi-modal working memory measurement system as claimed in claim 1, wherein, The real-time validity monitoring module is used to calculate the standardized residual between the observed response pattern and the expected response pattern for each question. When the absolute value of the standardized residual is greater than a preset threshold, the corresponding question is marked as an abnormal question, triggering the multimodal stimulus generation engine to regenerate an alternative question, and feeding the abnormal parameters back to the equivalence parameter matrix for online learning and updating.

6. The iMAS intelligent multi-modal working memory measurement system as claimed in claim 4, wherein, The Bayesian IRT-based adaptive algorithm for obtaining the posterior distribution of the subject's ability parameters based on the subject's real-time response data includes: after each preset number of trials, a Bayesian update algorithm is used, and the entropy value of the current ability posterior distribution of each trial's response is used as the weight of the response data for each trial to perform weighted processing and calculate the subject's ability posterior distribution.

7. The iMAS intelligent multimodal working memory measurement system according to claim 1, characterized in that, The participants' response data included: response data from the n-back matching task performed synchronously by the participants and response data from the cross-modal information binding judgment task.

8. The iMAS intelligent multi-modal working memory measurement system as claimed in claim 1, wherein, In the first round of testing, the corresponding norms are retrieved from the parameter database based on the input demographic variables to generate initial ability values, which are then sent to the multimodal stimulus generation engine.

9. An electronic terminal, characterized in that include: One or more memories and one or more processors; The one or more memories are used to store computer programs; The one or more processors are connected to the memory for running the computer program to perform the functions of the iMAS intelligent multimodal working memory measurement system as described in any one of claims 1 to 8.

10. A computerized adaptive testing system, characterized by The system includes: a test terminal, a display, a network communication module, a cloud-based large language model inference server, and a parameter database server; The test terminal is used to run the iMAS intelligent multimodal working memory measurement system as described in any one of claims 1 to 8; A display screen is used to show the questions to the test subjects. The network communication module is used to transmit data between the test terminal and the cloud-based large language model inference server. The cloud-based large language model inference server is used to run the cloud-based large language model and provide generation services; The parameter database server is used to store and provide test parameters, norms, and calibration data.