Multi-modal cognitive assessment question paper construction system and method

By building a multimodal cognitive assessment question system, collecting users' first modal response behaviors, generating a modal preference distribution model, and optimizing the question structure, the problem of the inability to track the modal switching rhythm in the existing system is solved, and the accuracy and personalized adaptability of the assessment results are improved.

CN120634813AInactive Publication Date: 2025-09-12SHANGHAI JILIXUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511142638.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing multimodal cognitive assessment question systems lack in-depth modeling of the user's information reception process and are unable to track the information modality that users pay primary attention to and the rhythm of switching between modalities, resulting in inaccurate assessment results.

Method used

By constructing a multimodal cognitive assessment question system, the user's first modal response behavior during the answering process is collected, the modal response time vector and weight vector are extracted, and the user modal preference distribution model is generated. Based on this, the question structure is optimized to achieve personalized adjustment.

Benefits of technology

It improves the accuracy of assessment results and the smoothness of interactive experience, adapts to individuals with different cognitive characteristics, dynamically matches the assessment question structure with individual cognitive processing paths, reduces cognitive jumps, and improves the efficiency of information reception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634813A_ABST
    Figure CN120634813A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal cognitive evaluation question paper construction system and method, and relates to the technical field of question paper evaluation.A first modal response behavior of a user in the actual answering process is collected, a modal response time vector is extracted, a modal response sequence vector and a modal response weight vector are further deduced, and a multi-modal cognitive evaluation question paper is constructed. Performing statistics in a multi-question range to construct a modal response weight set, forming a user modal preference distribution model according to the modal response weight set, performing personalized adjustment on the modal structure of each question in the original evaluation question paper on the basis of the user modal preference distribution model, generating an optimized evaluation question paper set, and collecting user interaction data again to obtain an evaluation question paper set; an optimized user modal preference distribution model is formed, the defects that the information processing process is invisible and cannot be quantified are overcome, a modal structure self-adaption mechanism is further introduced, and finally dynamic matching between an evaluation question paper structure and an individual cognition processing path is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of question paper evaluation technology, and in particular to a multimodal cognitive assessment question paper construction system and method. Background Art

[0002] With the ongoing convergence of contemporary cognitive science and intelligent assessment systems, AI-assisted cognitive assessment has become a key research direction in diverse fields, including intelligent education, professional competency assessment, and psychological diagnosis and treatment. This has spawned more specific subfields, such as cognitive state modeling and the construction of personalized questionnaire systems. Within this specific technology landscape, traditional single-modality questionnaire design is increasingly unable to meet the demands for high-precision and adaptable assessments, particularly in assessments of diverse dimensions such as attention, comprehension, and logical reasoning. Consequently, a new approach, centered around multimodality, has emerged.

[0003] In assessment systems across various scenarios, questions often combine text descriptions, images, and voice prompts to create a realistic and interactive assessment environment. However, most current systems only record and score test takers' responses from a static perspective, lacking in-depth modeling of their information-receiving processes. In such cases, even if the system can capture basic information about answer times or accuracy, it struggles to capture the user's strategic preferences, attentional distribution, and cognitive sequencing tendencies when processing information.

[0004] Specifically, existing technologies often overlook the order in which users process information: whether they read text first, then view images, or listen to audio first. These differences in order vary significantly across tasks and among people with different cognitive levels, and these variations may be linked to information extraction efficiency, cognitive resource allocation, and even psychological well-being. However, traditional test-taking systems often treat "answering behavior" as a black box operation, lacking mechanisms to track users' primary information modalities, the rhythm of switching between modalities, and even the cognitive preferences and ability structures underlying them. Summary of the Invention

[0005] In response to the deficiencies of the prior art, the present invention provides a multimodal cognitive assessment question paper construction system and method, which solves the problems mentioned in the background technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for constructing a multimodal cognitive assessment paper, comprising the following steps: S1. Using the content loading module in the question paper system, a question content combination structure including audio modality, image modality, and text modality is constructed to form an original modal response data set; S2. Extracting the user's first response time for each question in the audio modality, image modality, and text modality from the original modal response data set to form a modal response time vector, sorting the vectors to form a modal response order vector, and then calculating a modal response weight vector based on the modal response time vectors; S3. Statistically summarize the modal response weight vectors generated in all questions, construct a modal response weight set, average the response weights of each modal type, and obtain a user modal preference distribution model; S4. Based on the user modality preference distribution model, perform personalized structural optimization on the content of the assessment questions. After performing modal structure optimization on all questions in the original question paper, generate an optimized assessment question paper set. S5. Re-execute the answering process in the optimized assessment question set, record the user's modal response behavior data in the optimized question set, recalculate the modal response weight vector, and generate an optimized user modal preference distribution model based on this.

[0007] Preferably, said S1 includes S11 and S12; S11. Using the content loading module in the question paper system, read the multimodal question content preset in the question bank, combine and configure the audio mode, image mode, and text mode according to the question design requirements, and construct a multimodal question content combination structure; The constructed multimodal question content combination structure is loaded into the answering interface of the assessment paper, and the audio mode, image mode and text mode are presented in a synchronous and parallel manner according to the standard display order. Through this process, the user can simultaneously receive information input from different modalities at the beginning of each question, forming the initial input environment for multi-channel cognitive processing.

[0008] Preferably, S12, after the user enters the question answering interface, dynamically record the user's interactive behaviors in the audio mode, image mode, and text mode during the operation process, wherein the interactive behaviors include: clicking the play audio button, sliding in the image area, and focusing on the text area for reading; The built-in behavior monitoring mechanism records the specific time when each modality is first activated, forming paired data of modality type and response time; The first response behavior data of each modality is structured and packaged to generate an original modal response data set, which includes the current question number, modality type label and corresponding response time field.

[0009] Preferably, said S2 includes S21 and S22; S21. For each question in the original modal response data set, extract the user's first response time information for the audio modality, image modality, and text modality, and mark the first response time of the audio modality as the audio modal response time, the first response time of the image modality as the image modal response time, and the first response time of the text modality as the text modal response time. The audio modal response time, image modal response time, and text modal response time are combined in a fixed order to form a modal response time vector, which is used to describe the user's initial processing response rhythm to multimodal information in the current question; the modal response time vector is a direct input for subsequent calculations of modal processing order and response efficiency.

[0010] Preferably, S22, based on the time sequence relationship in the modal response time vector, sorting the response times of the audio modal response time, the image modal response time, and the text modal response time in ascending order to generate a modal response sequence vector; The modal response order vector is used to represent the user's modal processing order preference under the current question; Calculating a modal response weight vector based on the modal response time vector, wherein the modal response weight vector is used to quantify the priority of each mode in the user processing path; The modal response weight vector is obtained through steps S221, S222 and S223; S221. Extract the minimum response time value in the modal response time vector and mark it as the earliest response time; S222, calculating the time difference between the modal response time and the earliest response time for the audio modal response time, the image modal response time, and the text modal response time, and adding a constant smoothing factor; Among them, the constant smoothing factor is a positive number used to prevent the denominator from being zero, and its value is 0.1; S223 . Using the reciprocal of the time difference as a response weight corresponding to the modal response time, a modal response weight vector is formed, including an audio modal response weight, an image modal response weight, and a text modal response weight.

[0011] Preferably, said S3 includes S31 and S32; S31, summarizing and arranging the modal response weight vectors in order of question numbers and storing them as a modal response weight set; Each item in the modal response weight set corresponds to a modal response weight vector of a question, forming a modal processing response information matrix at the question paper level; S32. Aggregate and statistically analyze all modal response weight vectors in the modal response weight set by modality type, average the corresponding audio modal response weights in all questions to obtain an audio modality preference value; average the corresponding image modal response weights in all questions to obtain an image modality preference value; and average the corresponding text modal response weights in all questions to obtain a text modality preference value. The audio modality preference value, the image modality preference value, and the text modality preference value are integrated to form a user modality preference distribution model.

[0012] Preferably, said S4 includes S41 and S42; S41: Based on the order of the audio modality preference value, the image modality preference value, and the text modality preference value in the user modality preference distribution model, determine the current user's processing response order tendency for the three modalities and generate a modality presentation priority strategy for each question. The modal presentation priority strategy includes: Adjust the modal display position so that the modal information with the highest modal preference value among the three modal preferences is located in the user's sight priority area; The order of modal loading is rearranged so that the modal information with the highest modal preference value among the three modal information is activated first at the beginning of page loading; S42, applying the modality presentation priority strategy to all questions in the original assessment paper, dynamically reconstructing the modal structure of each question, performing operations including adjusting the modality loading logic, optimizing the information hierarchy, and fine-tuning the interface interaction mode, to generate an optimized assessment paper set; The presentation order of the audio modality, image modality and text modality of each question in the optimized assessment question set is consistent with the modality processing order preference reflected by the user modality preference distribution model.

[0013] Preferably, said S5 includes S51 and S52; S51. Based on the optimized assessment question set, reorganize the user's answering process and execute the question answering operation. During the answering process, record the user's first response behavior to the audio modality, image modality, and text modality for each question, extract the audio modality response time, image modality response time, and text modality response time, and construct a new modality response time vector. Based on the new modal response time vector, according to the same modal response weight vector acquisition method in step S2, the optimized modal response weight vector of each question is generated, and the optimized modal response weight set is constructed based on the optimized modal response weight vector of each question; The optimized modal response weight set is then subjected to the same aggregation statistics as in step S3 to obtain optimized audio modal preference values, image modal preference values, and text modal preference values, forming an optimized user modal preference distribution model.

[0014] Preferably, in step S52, the optimized user modality preference distribution model is compared one by one with the user modality preference distribution model generated in step S3, specifically calculating the modality preference offset value by Euclidean distance; then, the modality preference offset value is compared with a preset deviation threshold, and the validity of the question paper structure is determined based on the comparison result; When the modality preference offset value is less than the deviation threshold, the optimization result is determined to be valid and the current question structure is maintained; When the modal preference offset value is greater than or equal to the deviation threshold, the optimization result is determined to be invalid and the question paper structure continues to be optimized.

[0015] A multimodal cognitive assessment question paper construction system includes a content loading module, a modal vector extraction module, a user modality preference construction module, a question paper structure optimization module, and an execution feedback module; The content loading module constructs a question content combination structure including audio modality, image modality and text modality through the content loading module in the question paper system to form an original modality response data set; The modal vector extraction module extracts the user's first response time for the audio modality, image modality, and text modality in each question from the original modal response data set to form a modal response time vector, sorts the vectors to form a modal response order vector, and then calculates a modal response weight vector based on the modal response time vector. The user modality preference construction module statistically summarizes the modal response weight vectors generated in all questions, constructs a modal response weight set, and averages the response weights of each modality type to obtain a user modality preference distribution model; The question paper structure optimization module performs personalized structural optimization on the question content in the assessment question paper based on the user modal preference distribution model, and generates an optimized assessment question paper set after performing modal structure optimization on all questions in the original question paper; The execution feedback module re-executes the answering process in the optimized evaluation question paper set, records the user's modal response behavior data in the optimized question paper, recalculates the modal response weight vector, and generates an optimized user modal preference distribution model based on this.

[0016] The present invention provides a multimodal cognitive assessment question paper construction system and method, which has the following beneficial effects: (1) By collecting the user's first modal response behavior in the actual question-answering process, extracting the modal response time vector, and further deriving the modal response order vector and modal response weight vector, the modal response weight set is statistically constructed within the scope of multiple questions, and based on this, a user modal preference distribution model is formed. Subsequently, based on the user modal preference distribution model, the modal structure of each question in the original assessment paper is personalized adjusted to generate an optimized assessment paper set, and user interaction data is collected again to form an optimized user modal preference distribution model, which makes up for the invisible and non-quantifiable defects of the information processing process. On the basis of the optimized user modal preference distribution model, a modal structure adaptive mechanism is further introduced, and finally a dynamic match between the assessment paper structure and the individual cognitive processing path is achieved, effectively improving the accuracy of the assessment results, the smoothness of the interactive experience, and the adaptability of the system to individuals with different cognitive characteristics.

[0017] (2) Calculate the modal type dimension of all modal response weight vectors, and obtain the audio modal preference value, image modal preference value and text modal preference value respectively. These three are used as core indicators to construct a user modal preference distribution model. This user modal preference distribution model not only accurately expresses the user's processing response tendency for multiple modal information in the entire set of assessment questions, but also provides a data basis for the construction of personalized assessment paths as a stable expression of individual cognitive preference characteristics. Under the guidance of the user modal preference distribution model, it can truly realize the orderly regulation of the multimodal question structure, so that the information presentation rhythm of each question is accurately matched with the user's processing rhythm. Different from the problem of fixed modal content display structure and no consideration of individual reception efficiency differences in known technologies, this method drives the dynamic reconstruction of the assessment content structure through the modal presentation priority strategy, so that users can reduce cognitive jumps in the process of answering questions, improve reception efficiency, and maintain the fluency and coherence of thought processing, thereby improving the overall answering experience and the ability to control cognitive load.

[0018] (3) Based on the generated optimized assessment question set, the user's answering process is reorganized, and the user's first response behavior to the audio mode, image mode and text mode during the answering process is collected again to construct a new modal response time vector. Based on the modal response time vector, according to the established modal response weight vector calculation mechanism, the optimized modal response weight vector corresponding to each question is obtained, and then the optimized modal response weight set is summarized. The optimized audio modal preference value, image modal preference value and text modal preference value are obtained through aggregate statistical calculation, thereby constructing an optimized user modal preference distribution model, realizing secondary modeling of individual cognitive processing characteristics, and realizing automatic judgment and self-feedback control of the optimization effect of the question paper structure. Different from the existing technology that relies on expert experience or manual intervention to evaluate the optimization effect, the modal preference offset value is used to achieve objective judgment on whether the structure needs further optimization, and an automated closed-loop mechanism of evaluation optimization execution-behavior re-collection-effect evaluation-iterative adjustment is established, which improves the intelligence level and dynamic response capability of the question paper content adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic diagram of the steps of a method for constructing a multimodal cognitive assessment paper according to the present invention; Figure 2 This is a schematic diagram of a multimodal cognitive assessment question paper construction system of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0021] Example 1 The present invention provides a method for constructing a multimodal cognitive assessment paper. Figure 1 , including the following steps: S1. Using the content loading module in the question paper system, a question content combination structure including audio modality, image modality, and text modality is constructed to form an original modal response data set; S2. Extracting the user's first response time for each question in the audio modality, image modality, and text modality from the original modal response data set to form a modal response time vector, sorting the vectors to form a modal response order vector, and then calculating a modal response weight vector based on the modal response time vectors; S3. Statistically summarize the modal response weight vectors generated in all questions, construct a modal response weight set, average the response weights of each modal type, and obtain a user modal preference distribution model; S4. Based on the user modality preference distribution model, perform personalized structural optimization on the content of the assessment questions. After performing modal structure optimization on all questions in the original question paper, generate an optimized assessment question paper set. S5. Re-execute the answering process in the optimized assessment question set, record the user's modal response behavior data in the optimized question set, recalculate the modal response weight vector, and generate an optimized user modal preference distribution model based on this.

[0022] In this embodiment, by collecting the user's first modal response behavior during the actual question-answering process, extracting the modal response time vector, and further deriving the modal response order vector and modal response weight vector, a modal response weight set is statistically constructed within the scope of multiple questions, and a user modal preference distribution model is formed accordingly. Subsequently, based on the user modal preference distribution model, the modal structure of each question in the original assessment paper is personalized and adjusted to generate an optimized assessment paper set. User interaction data is then collected again to form an optimized user modal preference distribution model. The advantage of a question paper generation path that can achieve dynamic personalized optimization based on user cognitive behavior characteristics is that it can extract the user's actual information processing order from the user's modal response time vector, and reflect the individual's cognitive preference characteristics through the modal response sequence vector and the modal response weight vector, thereby establishing a real and quantifiable user modal preference distribution model, and reversely drive the question paper structure adjustment based on this model, so that the final optimized evaluation question paper set is more in line with the subject's modal processing habits and perceptual pathways.

[0023] It makes up for the invisible and unquantifiable defects of the information processing process, and further introduces the modal structure adaptation mechanism based on the optimized user modal preference distribution model, and finally achieves a dynamic match between the assessment question structure and the individual cognitive processing path, effectively improving the accuracy of the assessment results, the smoothness of the interactive experience and the adaptability of the system to individuals with different cognitive characteristics.

[0024] Example 2 specifically: S1 includes S11 and S12; S11. Using the content loading module in the question paper system, read the multimodal question content preset in the question bank, combine and configure the audio mode, image mode, and text mode according to the question design requirements, and construct a multimodal question content combination structure; The constructed multimodal question content combination structure is loaded into the answering interface of the assessment paper, and the audio mode, image mode and text mode are presented in a synchronous and parallel manner according to the standard display order. Through this process, the user can simultaneously receive information input from different modalities at the beginning of each question, forming the initial input environment for multi-channel cognitive processing.

[0025] S12. After the user enters the question answering interface, dynamically record the user's interactive behaviors in the audio mode, image mode, and text mode during the operation process, wherein the interactive behaviors include: clicking the play audio button, sliding in the image area, and focusing on the text area for reading; The built-in behavior monitoring mechanism records the specific time when each modality is first activated, forming paired data of modality type and response time; The first response behavior data of each modality is structured and encapsulated to generate an original modal response data set. The original modal response data set includes the current question number, modal type label and corresponding response time field, which is an important input basis for calculating the modal response time vector, modal response sequence vector and modal response weight vector in subsequent steps.

[0026] In this embodiment, by uniformly loading the question contents of audio modality, image modality and text modality and displaying them in a structured manner, the user can simultaneously receive cognitive stimulation from different information channels in the initial stage of answering questions, thereby constructing a question paper interactive environment that truly restores multi-sensory input. At the same time, through the behavior monitoring mechanism embedded in the answering interface, the user's first response behavior during the operation process is captured in real time, such as clicking the play audio button, sliding the image area, focusing on the text content, etc., and then the response time of each modality is structured and encapsulated to generate the original modal response data set. This set not only retains the pairing information of the modality type and the response time, but also realizes the precise positioning and sub-question tracking of the interactive data through the question number.

[0027] Not only does it achieve the synchronous loading and unified management of audio modalities, image modalities, and text modalities, but it also establishes a refined perception mechanism for multimodal response behaviors. Compared to traditional test paper systems that only focus on the final answer result or total time, this step makes the user's first response time to each modal information a quantifiable behavioral parameter, thereby making the original modal response data set complete, structured, and traceable. This data structure provides a unique and accurate original input for the subsequent calculation of the modal response time vector, modal response sequence vector, and modal response weight vector, giving the cognitive processing path analysis a real behavioral basis.

[0028] Example 3 specifically: S2 includes S21 and S22; S21. For each question in the original modal response data set, extract the user's first response time information for the audio modality, image modality, and text modality, and mark the first response time of the audio modality as the audio modal response time, the first response time of the image modality as the image modal response time, and the first response time of the text modality as the text modal response time. The audio modal response time, image modal response time, and text modal response time are combined in a fixed order to form a modal response time vector, which is used to describe the user's initial processing response rhythm to multimodal information in the current question; the modal response time vector is a direct input for subsequent calculations of modal processing order and response efficiency.

[0029] S22. Sort the audio modal response time, the image modal response time, and the text modal response time in ascending order based on the time sequence relationship in the modal response time vector to generate a modal response sequence vector; The modal response order vector is used to represent the user's modal processing order preference for the current question, clarifying which modality is the primary source of information input and which modality is the focus of subsequent attention; Calculating a modal response weight vector based on the modal response time vector, wherein the modal response weight vector is used to quantify the priority of each mode in the user processing path; The modal response weight vector is obtained through steps S221, S222 and S223; S221. Extract the minimum response time value in the modal response time vector and mark it as the earliest response time; S222, calculating the time difference between the modal response time and the earliest response time for the audio modal response time, the image modal response time, and the text modal response time, and adding a constant smoothing factor; Among them, the constant smoothing factor is a positive number used to prevent the denominator from being zero, and its value is 0.1; S223 . Using the reciprocal of the time difference as a response weight corresponding to the modal response time, a modal response weight vector is formed, including an audio modal response weight, an image modal response weight, and a text modal response weight.

[0030] In this implementation, by extracting the user's first response time information for each question across audio, image, and text modalities and constructing a modal response time vector in a unified order, this system, for the first time, achieves a quantifiable representation of the user's onset rhythm of multimodal information processing at the assessment paper level. The system then sorts the response times of each modality in the modal response time vector to generate a modal response sequence vector, accurately depicting the user's priority logic for receiving and processing different modal information in the current question.

[0031] Furthermore, a modal response weight vector calculation mechanism is introduced. By extracting the earliest response time from the modal response time vector and setting a smoothing constant, a stable and universal modal response priority assessment model is established. This calculation mechanism allows audio modal response time, image modal response time, and text modal response time to exist not only as raw time series data but also to be converted into response strength parameters that represent user processing efficiency, thereby achieving dual modeling of modal order preference and modal response efficiency.

[0032] The advantage lies in not only being able to identify the order in which users are exposed to information modalities, but also reflecting the relative emphasis of their information processing rhythm through modal response weight vectors. Unlike existing technologies that can only observe users' final answers but cannot understand their information processing paths, this method integrates "sequential preference modeling" and "priority intensity quantification" in cognitive processing, significantly improving the analytical granularity and computational accuracy of users' multimodal reception behavior, and providing more stable and interpretable basic feature parameter support for the subsequent construction of user modal preference distribution models.

[0033] Example 4 specifically: S3 includes S31 and S32; S31, summarizing and arranging the modal response weight vectors in order of question numbers and storing them as a modal response weight set, which respectively reflects the processing priority of each modal content by the tested user in each question; Each item in the modal response weight set corresponds to a modal response weight vector for a question, forming a modal processing response information matrix at the question paper level. The generation process of the modal response weight set completes the integration of response behavior structured data from the single question dimension to the full paper dimension, and is the statistical basis for modeling user modal preference characteristics. S32. Aggregate and statistically analyze all modal response weight vectors in the modal response weight set by modality type, average the corresponding audio modal response weights in all questions to obtain an audio modality preference value; average the corresponding image modal response weights in all questions to obtain an image modality preference value; and average the corresponding text modal response weights in all questions to obtain a text modality preference value. The audio modality preference value, the image modality preference value and the text modality preference value are integrated to form a user modality preference distribution model, which is used to reflect the user's overall processing response tendency and priority processing strategy for different modal information in the entire set of test papers. The user modality preference distribution model is used as an individual feature expression result to guide the modal structure reconstruction in the subsequent test paper optimization process.

[0034] Said S4 includes S41 and S42; S41: Based on the order of the audio modality preference value, the image modality preference value, and the text modality preference value in the user modality preference distribution model, determine the current user's processing response order tendency for the three modalities and generate a modality presentation priority strategy for each question. The modal presentation priority strategy includes: Adjust the modal display position so that the modal information with the highest modal preference value among the three modal preferences is located in the user's sight priority area; The order of modal loading is rearranged so that the modal information with the highest modal preference value among the three modal information is activated first at the beginning of page loading; S42, applying the modality presentation priority strategy to all questions in the original assessment paper, dynamically reconstructing the modal structure of each question, performing operations including adjusting the modality loading logic, optimizing the information hierarchy, and fine-tuning the interface interaction mode, to generate an optimized assessment paper set; The display order of the audio modality, image modality and text modality of each question in the optimized assessment question set is consistent with the modality processing order preference reflected by the user modality preference distribution model, thereby improving the user's information reception efficiency during the answering process, reducing cognitive jump costs and improving answer coherence.

[0035] In this example, by calculating the modal type dimension of all modal response weight vectors, we derive audio, image, and text modality preference values, respectively. These three values ​​serve as core indicators to construct a user modality preference distribution model. This user modality preference distribution model not only accurately reflects the user's processing and response tendencies for various modal information across the entire assessment test set, but also serves as a stable expression of individual cognitive preference characteristics, providing a data foundation for the construction of personalized assessment paths.

[0036] By determining the order of the three preference values ​​in the user modality preference distribution model, the user's processing response ranking tendency for audio modality, image modality, and text modality is determined, thereby generating a modal presentation priority strategy for each question. The presentation priority strategy completes differentiated processing based on the original question structure by adjusting the modal display position and rearranging the modal loading order. The optimization strategy is applied to all questions, and finally an optimized assessment paper set is generated. The multimodal display order of each question in the optimized assessment paper set is consistent with the processing order in the user modality preference distribution model, forming a complete closed loop from modeling to content reconstruction.

[0037] The advantage lies in the ability to systematically regulate the structure of multimodal test papers, guided by a user modality preference distribution model, ensuring that the information presentation rhythm of each question precisely matches the user's processing rhythm. Unlike existing techniques that employ a fixed modal content presentation structure and disregard individual differences in reception efficiency, this method uses a modality presentation priority strategy to drive the dynamic reconstruction of assessment content structure. This reduces cognitive jumps during the user's answering process, improves reception efficiency, and maintains the fluency and coherence of thought processing, thereby enhancing the overall answering experience and the ability to control cognitive load.

[0038] Embodiment 5 specifically: said S5 includes S51 and S52; S51. Based on the optimized assessment question set, reorganize the user's answering process and execute the question answering operation. During the answering process, record the user's first response behavior to the audio modality, image modality, and text modality for each question, extract the audio modality response time, image modality response time, and text modality response time, and construct a new modality response time vector. Based on the new modal response time vector, according to the same modal response weight vector acquisition method in step S2, the optimized modal response weight vector of each question is generated, and the optimized modal response weight set is constructed based on the optimized modal response weight vector of each question; The optimized modal response weight set is then subjected to the same aggregation statistics as in step S3 to obtain optimized audio modal preference values, image modal preference values, and text modal preference values, forming an optimized user modal preference distribution model.

[0039] S52: Compare the optimized user modality preference distribution model with the user modality preference distribution model generated in step S3 one by one. Specifically, calculate the modality preference offset value through Euclidean distance to measure whether the user exhibits a more stable and consistent modality processing path in the optimized test paper. Then, compare the modality preference offset value with a preset deviation threshold, and determine the validity of the test paper structure based on the comparison result. When the modality preference offset value is less than the deviation threshold, the optimization result is determined to be valid and the current question structure is maintained; When the modal preference offset value is greater than or equal to the deviation threshold, the optimization result is determined to be invalid and the question paper structure continues to be optimized.

[0040] In this embodiment, based on the generation of an optimized set of assessment questions, the user's answering process is reorganized, and the user's first response behavior for the audio modality, image modality, and text modality during the answering process is again collected to construct a new modal response time vector. Based on this modal response time vector, according to the established modal response weight vector calculation mechanism, the optimized modal response weight vector corresponding to each question is obtained. This is then aggregated to form an optimized modal response weight set. Optimized audio modality preference values, image modality preference values, and text modality preference values ​​are then obtained through aggregate statistical calculation. This constructs an optimized user modality preference distribution model, achieving secondary modeling of individual cognitive processing characteristics.

[0041] On this basis, the optimized user modal preference distribution model was compared one-by-one with the original user modal preference distribution model. The modal preference offset value was calculated using the Euclidean distance algorithm. This modal preference offset value served as the core criterion for evaluating the effectiveness of the optimization and was compared with a set deviation threshold. When the modal preference offset value was less than the deviation threshold, it indicated that the user exhibited a more stable and consistent modal processing path in the optimized test paper. The current structural optimization was deemed effective and retained. Conversely, when the modal preference offset value was greater than or equal to the deviation threshold, the optimization was deemed invalid and the next round of test paper structure adjustment was automatically initiated.

[0042] The advantage lies in the introduction of modal preference offsets as a quantitative evaluation parameter for the degree of change in user processing behavior after optimization, enabling automatic identification and self-feedback control of the optimization effect of the test paper structure. Unlike existing technologies that rely on expert experience or manual intervention to evaluate optimization results, this method uses modal preference offsets to objectively determine whether the structure needs further optimization. This establishes an automated closed-loop mechanism: evaluating optimization execution—behavior recollection—effect evaluation—and iterative adjustment, enhancing the intelligent level of adaptability and dynamic responsiveness of test paper content.

[0043] Example 6 A multimodal cognitive assessment question paper construction system, please refer to Figure 2 ,Specifically: including content loading module, modal vector extraction module, user modal preference construction module, question paper structure optimization module and execution feedback module; The content loading module constructs a question content combination structure including audio modality, image modality and text modality through the content loading module in the question paper system to form an original modality response data set; The modal vector extraction module extracts the user's first response time for the audio modality, image modality, and text modality in each question from the original modal response data set to form a modal response time vector, sorts the vectors to form a modal response order vector, and then calculates a modal response weight vector based on the modal response time vector. The user modality preference construction module statistically summarizes the modal response weight vectors generated in all questions, constructs a modal response weight set, and averages the response weights of each modality type to obtain a user modality preference distribution model; The question paper structure optimization module performs personalized structural optimization on the question content in the assessment question paper based on the user modal preference distribution model, and generates an optimized assessment question paper set after performing modal structure optimization on all questions in the original question paper; The execution feedback module re-executes the answering process in the optimized evaluation question paper set, records the user's modal response behavior data in the optimized question paper, recalculates the modal response weight vector, and generates an optimized user modal preference distribution model based on this.

[0044] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a multimodal cognitive assessment paper, characterized by: The following steps are involved: S1. Using the content loading module in the question paper system, a question content combination structure including audio modality, image modality, and text modality is constructed to form an original modal response data set; S2. Extracting the user's first response time for each question in the audio modality, image modality, and text modality from the original modal response data set to form a modal response time vector, sorting the vectors to form a modal response order vector, and then calculating a modal response weight vector based on the modal response time vectors; S3. Statistically summarize the modal response weight vectors generated in all questions, construct a modal response weight set, average the response weights of each modal type, and obtain a user modal preference distribution model; S4. Based on the user modality preference distribution model, perform personalized structural optimization on the content of the assessment questions. After performing modal structure optimization on all questions in the original question paper, generate an optimized assessment question paper set. S5. Re-execute the answering process in the optimized assessment question set, record the user's modal response behavior data in the optimized question set, recalculate the modal response weight vector, and generate an optimized user modal preference distribution model based on this.

2. A method for constructing a multimodal cognitive assessment paper according to claim 1, characterized in that: Said S1 includes S11 and S12; S11. Using the content loading module in the question paper system, read the multimodal question content preset in the question bank, combine and configure the audio mode, image mode, and text mode according to the question design requirements, and construct a multimodal question content combination structure; The constructed multimodal question content combination structure is loaded into the answering interface of the assessment paper, and the audio mode, image mode and text mode are presented in a synchronous and parallel manner according to the standard display order. Through this process, the user can simultaneously receive information input from different modalities at the beginning of each question, forming the initial input environment for multi-channel cognitive processing.

3. A method for constructing a multimodal cognitive assessment paper according to claim 2, characterized in that: S12. After the user enters the question answering interface, dynamically record the user's interactive behaviors in the audio mode, image mode, and text mode during the operation process, wherein the interactive behaviors include: clicking the play audio button, sliding in the image area, and focusing on the text area for reading; The built-in behavior monitoring mechanism records the specific time when each modality is first activated, forming paired data of modality type and response time; The first response behavior data of each modality is structured and packaged to generate an original modal response data set, which includes the current question number, modality type label and corresponding response time field.

4. A method for constructing a multimodal cognitive assessment paper according to claim 3, characterized in that: Said S2 includes S21 and S22; S21. For each question in the original modal response data set, extract the user's first response time information for the audio modality, image modality, and text modality, and mark the first response time of the audio modality as the audio modal response time, the first response time of the image modality as the image modal response time, and the first response time of the text modality as the text modal response time. The audio modal response time, image modal response time, and text modal response time are combined in a fixed order to form a modal response time vector, which is used to describe the user's initial processing response rhythm to multimodal information in the current question; the modal response time vector is a direct input for subsequent calculations of modal processing order and response efficiency.

5. A method for constructing a multimodal cognitive assessment paper according to claim 4, characterized in that: S22. Sort the audio modal response time, the image modal response time, and the text modal response time in ascending order based on the time sequence relationship in the modal response time vector to generate a modal response sequence vector; The modal response order vector is used to represent the user's modal processing order preference under the current question; Calculating a modal response weight vector based on the modal response time vector, wherein the modal response weight vector is used to quantify the priority of each mode in the user processing path; The modal response weight vector is obtained through steps S221, S222 and S223; S221. Extract the minimum response time value in the modal response time vector and mark it as the earliest response time; S222, calculating the time difference between the modal response time and the earliest response time for the audio modal response time, the image modal response time, and the text modal response time, and adding a constant smoothing factor; Among them, the constant smoothing factor is a positive number used to prevent the denominator from being zero, and its value is 0.1; S223 . Using the reciprocal of the time difference as a response weight corresponding to the modal response time, a modal response weight vector is formed, including an audio modal response weight, an image modal response weight, and a text modal response weight.

6. A method for constructing a multimodal cognitive assessment paper according to claim 5, characterized in that: Said S3 includes S31 and S32; S31, summarizing and arranging the modal response weight vectors in order of question numbers and storing them as a modal response weight set; Each item in the modal response weight set corresponds to a modal response weight vector of a question, forming a modal processing response information matrix at the question paper level; S32. Aggregate and count all modal response weight vectors in the modal response weight set according to modality type, average the corresponding audio modal response weights in all questions, and obtain an audio modality preference value; The image modality response weights corresponding to all questions are averaged to obtain the image modality preference value; The text modality response weights corresponding to all questions are averaged to obtain the text modality preference value; The audio modality preference value, the image modality preference value, and the text modality preference value are integrated to form a user modality preference distribution model.

7. A method for constructing a multimodal cognitive assessment paper according to claim 6, characterized in that: Said S4 includes S41 and S42; S41: Based on the order of the audio modality preference value, the image modality preference value, and the text modality preference value in the user modality preference distribution model, determine the current user's processing response order tendency for the three modalities and generate a modality presentation priority strategy for each question. The modal presentation priority strategy includes: Adjust the modal display position so that the modal information with the highest modal preference value among the three modal preferences is located in the user's sight priority area; The order of modal loading is rearranged so that the modal information with the highest modal preference value among the three modal information is activated first at the beginning of page loading; S42, applying the modality presentation priority strategy to all questions in the original assessment paper, dynamically reconstructing the modal structure of each question, performing operations including adjusting the modality loading logic, optimizing the information hierarchy, and fine-tuning the interface interaction mode, to generate an optimized assessment paper set; The presentation order of the audio modality, image modality and text modality of each question in the optimized assessment question set is consistent with the modality processing order preference reflected by the user modality preference distribution model.

8. A method for constructing a multimodal cognitive assessment paper according to claim 7, characterized in that: Said S5 includes S51 and S52; S51. Based on the optimized assessment question set, reorganize the user's answering process and execute the question answering operation. During the answering process, record the user's first response behavior to the audio modality, image modality, and text modality for each question, extract the audio modality response time, image modality response time, and text modality response time, and construct a new modality response time vector. Based on the new modal response time vector, according to the same modal response weight vector acquisition method in step S2, the optimized modal response weight vector of each question is generated, and the optimized modal response weight set is constructed based on the optimized modal response weight vector of each question; The optimized modal response weight set is then subjected to the same aggregation statistics as in step S3 to obtain optimized audio modal preference values, image modal preference values, and text modal preference values, forming an optimized user modal preference distribution model.

9. A method for constructing a multimodal cognitive assessment paper according to claim 8, characterized in that: S52: performing a one-to-one comparison between the optimized user modality preference distribution model and the user modality preference distribution model generated in step S3, specifically calculating a modality preference offset value by using Euclidean distance; Then, the modality preference offset value is compared with the preset deviation threshold, and the validity of the question paper structure is determined based on the comparison results; When the modality preference offset value is less than the deviation threshold, the optimization result is determined to be valid and the current question structure is maintained; When the modal preference offset value is greater than or equal to the deviation threshold, the optimization result is determined to be invalid and the question paper structure continues to be optimized.

10. A multimodal cognitive assessment question paper construction system, applied to a multimodal cognitive assessment question paper construction method according to any one of claims 1 to 9, characterized in that: It includes content loading module, modal vector extraction module, user modal preference construction module, question paper structure optimization module and execution feedback module; The content loading module constructs a question content combination structure including audio modality, image modality and text modality through the content loading module in the question paper system to form an original modality response data set; The modal vector extraction module extracts the user's first response time for the audio modality, image modality, and text modality in each question from the original modal response data set to form a modal response time vector, sorts the vectors to form a modal response order vector, and then calculates a modal response weight vector based on the modal response time vector. The user modality preference construction module statistically summarizes the modal response weight vectors generated in all questions, constructs a modal response weight set, and averages the response weights of each modality type to obtain a user modality preference distribution model; The question paper structure optimization module performs personalized structural optimization on the question content in the assessment question paper based on the user modal preference distribution model, and generates an optimized assessment question paper set after performing modal structure optimization on all questions in the original question paper; The execution feedback module re-executes the answering process in the optimized evaluation question set, records the user's modal response behavior data in the optimized question set, recalculates the modal response weight vector, and generates an optimized user modal preference distribution model based on this.