Method for automatically generating visual cognitive evaluation of images based on multi-modal language model

A visual cognition assessment method that automatically generates images using a multimodal language model expands the image library and conducts voice prompt tests, achieving efficient and accurate visual cognition assessment suitable for users of different ages.

CN120203519BActive Publication Date: 2025-12-16SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510369172.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-12-16
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

Existing visual cognition assessment methods suffer from high labor costs, susceptibility to subjective influence on assessment results, coarse classification of visual complexity levels, difficulty in achieving differentiated assessments across age groups, and failure to fully utilize visual language models, resulting in low test accuracy.

Method used

The system automatically generates images using a multimodal language model, expands the image recognition database using a visual language model, conducts visual cognition tests based on a speech generation module, and generates evaluation reports using a visual language model, thereby achieving dynamic expansion and accurate evaluation of the image library.

Benefits of technology

It significantly reduces labor costs, improves the quality and richness of knowledge base content, enhances the accuracy of assessments and their applicability across age groups, and solves the problem of insufficient recognition accuracy in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120203519B_ABST
    Figure CN120203519B_ABST
Patent Text Reader

Abstract

The application relates to a visual cognitive evaluation method for automatically generating images based on a multimodal language model. The method comprises the following steps: first, an initial image group is used to form an image recognition database, a visual language model is used to expand the image recognition database, then, corresponding images are selected from the image recognition database to form test questions based on a current evaluation task, then, a voice prompt is generated by using a voice generation module based on the test questions, a subject is subjected to a visual cognitive test, and finally, a visual cognitive test result is input into a visual language model to generate a visual cognitive evaluation report. Through the VLM visual language model, accurate image descriptions are automatically generated, and high-precision similar semantic images with great visual differences but small semantic differences are generated by using a semantic fine-tuning technology, dynamic expansion of an image knowledge base and accurate evaluation of subtle differences in visual semantics are realized, and the quality and richness of the content of the knowledge base are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cognitive function assessment technology, and in particular to a visual cognitive assessment method based on automatically generating images using a multimodal language model. Background Technology

[0002] With the iterative upgrade of artificial intelligence technology, the assessment of visual cognitive abilities across different age groups has gradually become a research hotspot. Traditional assessment systems mostly employ manual evaluation methods, i.e., testing through manually designed questions and standardized questionnaires. This traditional method has significant limitations. On the one hand, the development and implementation of standardized questionnaires incur significant human resource costs, with testing cycles typically lasting 2-3 weeks, and assessment results are easily influenced by the subjective judgment of the test takers. On the other hand, existing test materials have a coarse classification of visual complexity levels and lack a systematic design based on cognitive development theories, making it difficult to achieve differentiated assessments across age groups.

[0003] While existing technologies, particularly computer vision and artificial intelligence, offer methods for aiding cognitive assessment, they typically focus solely on the surface visual features of images, neglecting the inherent semantic differences. For instance, significant visual differences between images may lead to the misinterpretation of different events; conversely, images with minimal semantic differences may be difficult to distinguish due to visual discrepancies. Therefore, effectively designing test content and accurately assessing the semantic understanding of images across different age groups is a pressing issue. Furthermore, most current visual cognitive assessment systems fail to fully utilize the advanced technology of Visual Language Models (VLMs) and rely excessively on text input at the human-computer interaction level, reducing test accuracy.

[0004] Therefore, there is an urgent need in related technologies for a way to automatically generate test content and accurately assess and analyze visual cognitive abilities. Summary of the Invention

[0005] Therefore, it is necessary to provide a visual cognitive assessment method based on a multimodal language model that can automatically generate images and accurately evaluate and analyze visual cognitive abilities, addressing the aforementioned technical problems.

[0006] Firstly, this application provides a visual cognitive assessment method based on automatically generating images using a multimodal language model. The method includes:

[0007] An initial image database is obtained and composed of images. The image recognition database is then expanded using a visual language model.

[0008] Based on the current assessment task, corresponding images are selected from the image recognition database to form test questions;

[0009] Based on the test questions, a speech generation module is used to generate speech prompts to conduct a visual cognition test on the subject.

[0010] The visual cognition test results are input into the visual language model to generate a visual cognition assessment report.

[0011] Optionally, in one embodiment of this application, obtaining initial images to form an image recognition database and expanding the image recognition database using a visual language model includes:

[0012] The image is classified and labeled according to the features of the subject being measured.

[0013] Optionally, in one embodiment of this application, the step of selecting corresponding images from the image recognition database to form test questions based on the current evaluation task includes:

[0014] Select images with similar text descriptions within the same category based on similar text descriptions;

[0015] Select images with similar text descriptions from different categories based on similar text descriptions;

[0016] Select the corresponding image based on the semantically different text description.

[0017] Optionally, in one embodiment of this application, the similar text descriptions are obtained by similarity matching based on a large language model.

[0018] Optionally, in one embodiment of this application, the visual cognition test on the subject includes:

[0019] Calculate the accuracy rates for similar categories, dissimilar categories, and semantic differences to assess the subject's visual cognitive ability.

[0020] Secondly, this application also provides a visual cognitive assessment device for automatically generating images based on a multimodal language model. The device includes:

[0021] An image recognition database construction module is used to acquire initial images to form an image recognition database, and to expand the image recognition database using a visual language model;

[0022] The image test question selection module is used to select corresponding images from the image recognition database to form test questions based on the current evaluation task.

[0023] The visual cognition test and assessment module is used to generate voice prompts based on the test questions using the voice generation module, and to conduct visual cognition tests on the subject.

[0024] The visual cognition assessment results output module is used to input the visual cognition test results into the visual language model to generate a visual cognition assessment report.

[0025] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the steps of the methods described in the various embodiments above.

[0026] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the methods described in the various embodiments above.

[0027] The aforementioned visual cognition assessment method based on a multimodal language model (VLM) for automatically generating images first acquires an initial image recognition database, which is then expanded using a visual language model. Next, based on the current assessment task, corresponding images are selected from the database to form test questions. Then, a speech generation module generates voice prompts based on these questions to conduct a visual cognition test on the subject. Finally, the visual cognition test results are input into the visual language model to generate a visual cognition assessment report. In other words, by automatically generating accurate image descriptions using the VLM visual language model and utilizing semantic fine-tuning technology to generate high-precision similar semantic images with significant visual differences but minor semantic differences, the method achieves dynamic expansion of the image knowledge base and accurate assessment of subtle visual semantic differences. This significantly reduces manual labor costs and improves the quality and richness of the knowledge base content. Simultaneously, RAG technology enables image retrieval and matching based on semantic similarity, solving the problem of insufficient recognition accuracy in traditional visual difference assessment methods. Attached Figure Description

[0028] Figure 1 This is an application environment diagram of a visual cognitive assessment method based on a multimodal language model that automatically generates images, as shown in one embodiment.

[0029] Figure 2 This is a flowchart illustrating a visual cognition assessment method for automatically generating images based on a multimodal language model, as shown in one embodiment.

[0030] Figure 3 This is a structural block diagram of a visual cognition assessment device that automatically generates images based on a multimodal language model, as shown in one embodiment.

[0031] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0033] The visual cognitive assessment method based on a multimodal language model for automatically generating images provided in this application can be applied to, for example... Figure 1 The application environment is illustrated. The terminal communicates with the server via a network. The data storage system stores the data the server needs to process. This system can be integrated onto the server, or it can be hosted in the cloud or on other network servers. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0034] In one embodiment, such as Figure 2 As shown, a visual cognitive assessment method based on a multimodal language model for automatically generating images is provided, and this method is applied to... Figure 1 Taking the server in the example, the following steps are included:

[0035] S201: Obtain initial images to form an image recognition database, and expand the image recognition database using a visual language model.

[0036] In this embodiment, images are first collected from multiple channels (domestic and international picture books and related image datasets, which meet adult cognitive standards and can accommodate a wider user group) to form an image recognition database, ensuring rich and diverse content while meeting the cognitive and aesthetic needs of people of all ages. The image recognition database is a binary knowledge base composed of images and text. Each image is accompanied by a corresponding text description, which is automatically generated by a Visual Language Model (VLM) based on the image, accurately capturing and expressing key information in the image.

[0037] Different images often exhibit significant visual differences. Due to variations in environment, object placement, and personnel distribution, pixel differences between images are substantial, allowing for datasets with significant visual variations to be obtained without special processing. However, images with significant visual differences may not have significant semantic differences. For example, the same ball game may appear vastly different from different perspectives. In such cases, the test subject might perceive the events depicted in the images as entirely different due to the large visual differences, placing significant demands on the subject's cognitive abilities. Finding image datasets with approximate semantic information is quite difficult. Visual Language Models (VLMs) can generate corresponding images from text descriptions. Furthermore, by modifying the text descriptions, similar images with different semantic information can be generated. Examining the test subject's ability to distinguish these images provides a deeper assessment of whether the subject truly understands the semantic information of the images. Therefore, this study employs VLM to generate corresponding images based on the same text description and similar images based on similar text descriptions to expand the image recognition database.

[0038] In one embodiment of this application, the step of acquiring initial images to form an image recognition database and expanding the image recognition database using a visual language model includes:

[0039] The image is classified and labeled according to the features of the subject being measured.

[0040] In one embodiment of this application, the ability to understand and appreciate images may vary depending on the age of the tested subject. For example, children's cognitive abilities improve rapidly in early childhood, normal adults have high cognitive abilities, while the cognitive abilities of older adults tend to decline. Therefore, it is also important to rationally classify the images in the image recognition database. On the one hand, based on the age characteristics of children, images are divided into three levels—simple, medium, and complex—according to cognitive difficulty and content complexity. For example, children in kindergarten (lower, middle, and upper grades) are suitable for images that are brightly colored, simple in form, and easy to understand (simple and medium level), while students in grades 1-3 of primary school can gradually be exposed to images with richer content and more details (complex level). On the other hand, the process of human cognition of the world is divided into three different levels based on image information density (i.e., the range of scenes and field of vision shown in an image). The first level is the most complex landscape category, the second level is interpersonal interaction, and the third level is object category (further subdivided into images of moving robots and recognition of static single objects; single objects, considering that the test subjects include children, are divided into sports equipment, stationery, musical instruments, and fruits, etc.). Specifically, for adults with normal cognitive function, the image content is designed to be more complex and abstract, aiming to challenge their higher cognitive abilities. This level includes multi-layered scenes (such as city panoramas) and abstract art (such as modern paintings), requiring users to have strong analytical and reasoning abilities. For older adults with cognitive decline, the image content is simpler, more intuitive, and closer to daily life. This level includes scenes of family life (such as kitchen scenes) and social activities (such as walks in the park), aiming to help older adults maintain and exercise basic cognitive functions.

[0041] In this embodiment, an innovative classification method for image complexity and semantic information density is proposed based on the cognitive characteristics of subjects at different age stages, making the assessment process more in line with the cognitive development characteristics of the subjects.

[0042] S203: Select corresponding images from the image recognition database to form test questions based on the current evaluation task.

[0043] In this embodiment, different subjects have different abilities and levels of understanding of images. It is difficult for young children or elderly people with severe cognitive decline to provide a complete description of an image. Therefore, selecting images that match the description from different images is a simple and effective method for visual cognitive assessment. Based on different assessment tasks, Retrieval Enhancement Generative Algorithm (RAG) technology is used to select different images from the image recognition database to form test questions. The assessment tasks include comparison within the same category, comparison between different categories, and semantic difference comparison.

[0044] Specifically, in one embodiment of this application, the step of selecting corresponding images from the image recognition database to form test questions based on the current evaluation task includes:

[0045] S301: Select images with similar text descriptions within the same category based on similar text descriptions.

[0046] S303: Select images with similar text descriptions from different categories based on similar text descriptions.

[0047] S305: Select the corresponding image based on the text description with semantic differences.

[0048] In one embodiment of this application, for same-category comparisons, for text descriptions of the same type, the image with the most similar text description is selected to form the question based on the degree of similarity between the texts. For example, based on an image and its corresponding text description, three other images and their corresponding text descriptions with a high degree of similarity in the same category dataset are found. The test subject needs to find the image that matches the correct description from the four images. For dissimilar comparisons, the classification is to assess the test subject's cognitive ability under different categories of images. At the same time, when data for both same-category and dissimilar comparisons are available, it is possible to assess whether the test subject's cognitive problems are caused by the difference in categories. Similarly, text similarity is used as the standard to select the most similar image. The difference is that the image with the most similar text description is selected from different categories to form the question. For example, one image that is most similar to the correct image description is selected from each of the three categories that are different from the correct image to form the question.

[0049] For semantic difference comparison, the ability of the test subject to recognize images with minor semantic differences can be assessed. Specifically, one image is selected as the only correct one, and the corresponding text description is taken as the correct semantic information. The selected text description is then subtly adjusted to achieve a similar but slightly different effect. For example, for landscape images, the scenery and weather can be modified to change the semantic information; for images of interpersonal interaction, the number of people in the image can be modified to add or remove some semantic information. Simultaneously, the multimodal capabilities of VLM (Visual Metaphor) are utilized to describe the changed semantic information, achieving an effect with only subtle semantic differences from the original text. Based on the modified semantics, corresponding images are generated to form the questions. Thus, even if the visual differences between the images are significant, the semantic differences are only subtle. This requires the test subject to carefully compare the images to examine the detailed information in the descriptions.

[0050] In one embodiment of this application, the similar text descriptions are obtained by similarity matching based on a large language model.

[0051] In one embodiment of this application, for selecting similar texts, an advanced Large Language Model (LLM) is used for automatic similarity matching. The LLM scores the similarity between different texts to select the constituent objects of the question. The scoring criteria are as follows: 1. Content described; 2. Time described (occurrence time, duration); 3. Location described (spatial size of the event scene and its semantic information); 4. Composition of the described objects: whether there are people and their composition. Specifically, taking similar comparisons as an example, the Prompt is shown in Table 1 below:

[0052]

[0053] Table 1

[0054] S205: Based on the test questions, a voice generation module is used to generate voice prompts to conduct a visual cognition test on the subject.

[0055] In this embodiment, due to differences in literacy abilities among users of different ages—for example, children and some elderly people may have difficulty understanding text and therefore cannot accurately match image content during testing—the traditional text recognition task is eliminated, and only image recognition ability is assessed. A Funasr speech generation module is introduced to convey the text description of the test questions to the test subjects in speech form. With the help of speech prompts, test subjects can obtain key information without reading, thus allowing them to focus more on the image recognition task and ensuring that the test results more accurately reflect their visual cognitive abilities. Test subjects perform differently under different image categories, and also vary under questions with semantic and visual differences. Therefore, users have individual differences in their image recognition abilities. These differences can be analyzed using LLM to understand their underlying causes. Furthermore, dynamically adjusting the type ratio and the proportion of semantic and visual differences in subsequent training questions allows test subjects to continuously optimize their recognition level.

[0056] In one embodiment of this application, the visual cognition test on the subject includes:

[0057] Calculate the accuracy rates for similar categories, dissimilar categories, and semantic differences to assess the subject's visual cognitive ability.

[0058] In one embodiment of this application, to evaluate the subject's ability to recognize images, different accuracy rates were calculated for three evaluation methods: same-category accuracy, different-category accuracy, and semantic difference accuracy. A Comprehensive Evaluation Index (CEI) was calculated based on the same-category accuracy, different-category accuracy, and semantic difference accuracy, combined with different levels of influence, to assess the overall recognition ability. Specifically, the same-category accuracy (SCA) measures the ability to correctly recognize images within the same category, calculated by dividing the number of correctly recognized same-category images by the total number of images tested in that category. The different-category accuracy (CCA) measures the ability to distinguish between images of different categories, calculated by dividing the number of correctly classified cross-category images by the total number of cross-category images tested. The semantic difference accuracy (SDA) measures the ability to correctly recognize images with significant semantic differences, calculated by dividing the number of correctly recognized semantically different images by the total number of semantically different images tested. In addition to CEI, each index includes an overall accuracy rate and accuracy rates for four subcategories: Landscape, Interpersonal Interaction, Sports, and Static Objects. The specific calculation formulas are as follows:

[0059]

[0060]

[0061] CEI=ω1·SCA+ω2·CCA+ω3·SDA

[0062] Where ω1, ω2, and ω3 are weighting coefficients.

[0063] S207: Input the visual cognition test results into the visual language model to generate a visual cognition assessment report.

[0064] In this embodiment, during the evaluation phase, various indicators for similar, dissimilar, and semantically different categories were obtained. These indicators construct the recognition performance of the tested subject. When the category performance in the similar category accuracy rate is consistent with that in the dissimilar category accuracy rate, the category factor plays a crucial role in the entire image recognition task. Conversely, when the semantically different accuracy rate differs significantly from the former two, it indicates a problem with the tested subject's understanding of details. The evaluation results are then used as pre-defined information input into the visual language model. Furthermore, different categories of image datasets contain different semantic information. Summarizing this information is also used as pre-defined information input into the visual language model. The textual descriptions and category names of all similar images are input into the visual language model, thereby summarizing the overall characteristics of the categories for analysis. Additionally, due to differences in age and education level, image recognition abilities vary considerably. This information can also be used as background information for the tested subject input into the visual language model during analysis. Based on all received pre-defined information, the visual language model can analyze the tested subject's performance in specific categories, generating a capability evaluation report, clearly presented in PDF image format.

[0065] The aforementioned visual cognition assessment method based on multimodal language models (VLM) for automatically generating images first obtains an initial image recognition database, which is then expanded using a visual language model. Next, based on the current assessment task, corresponding images are selected from the database to form test questions. Then, a speech generation module generates voice prompts based on the test questions to conduct a visual cognition test on the subject. Finally, the visual cognition test results are input into the visual language model to generate a visual cognition assessment report. In other words, by automatically generating accurate image descriptions using the VLM visual language model and utilizing semantic fine-tuning technology to generate high-precision similar semantic images with significant visual differences but minor semantic differences, the method achieves dynamic expansion of the image knowledge base and accurate assessment of subtle visual semantic differences. This significantly reduces manual labor costs and improves the quality and richness of the knowledge base content. Simultaneously, RAG technology enables image retrieval and matching based on semantic similarity, solving the problem of insufficient recognition accuracy in traditional visual difference assessment methods.

[0066] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0067] Based on the same inventive concept, this application also provides a visual cognitive assessment device for automatically generating images based on a multimodal language model, used to implement the aforementioned visual cognitive assessment method for automatically generating images based on a multimodal language model. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the visual cognitive assessment device for automatically generating images based on a multimodal language model provided below can be found in the limitations of the visual cognitive assessment method for automatically generating images based on a multimodal language model described above, and will not be repeated here.

[0068] In one embodiment, such as Figure 3 As shown, a visual cognition assessment device 300 based on a multimodal language model for automatically generating images is provided, including: an image recognition database construction module 301, an image test question selection module 303, a visual cognition test assessment module 305, and a visual cognition assessment result output module 307, wherein:

[0069] The image recognition database construction module 301 is used to acquire initial images to form an image recognition database and to expand the image recognition database using a visual language model.

[0070] The image test question selection module 303 is used to select corresponding images from the image recognition database to form test questions based on the current evaluation task.

[0071] The visual cognition test assessment module 305 is used to generate voice prompts based on the test questions using the voice generation module to conduct a visual cognition test on the subject.

[0072] The visual cognition assessment result output module 307 is used to input the visual cognition test results into the visual language model to generate a visual cognition assessment report.

[0073] In one embodiment of this application, the image recognition database construction module is further configured to:

[0074] The image is classified and labeled according to the features of the subject being measured.

[0075] In one embodiment of this application, the image test question selection module is further configured to:

[0076] Select images with similar text descriptions within the same category based on similar text descriptions;

[0077] Select images with similar text descriptions from different categories based on similar text descriptions;

[0078] Select the corresponding image based on the semantically different text description.

[0079] In one embodiment of this application, the similar text descriptions are obtained by similarity matching based on a large language model.

[0080] In one embodiment of this application, the visual cognition test evaluation module is further configured to:

[0081] Calculate the accuracy rates for similar categories, dissimilar categories, and semantic differences to assess the subject's visual cognitive ability.

[0082] The modules in the aforementioned visual cognition assessment device that automatically generates images based on a multimodal language model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0083] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a visual cognitive assessment method that automatically generates images based on a multimodal language model. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0084] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0085] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0086] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0087] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0089] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0090] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0091] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A visual cognitive assessment method based on automatically generated images using a multimodal language model, characterized in that, The method includes: An initial image database is obtained and composed of images. The image recognition database is then expanded using a visual language model. Based on the current assessment task, corresponding images are selected from the image recognition database to form test questions; Based on the test questions, a speech generation module is used to generate speech prompts to conduct a visual cognition test on the subject. Input the visual cognition test results, image categories, text descriptions and category names of similar images, age, and education level into the visual language model to generate a visual cognition assessment report; The process of acquiring initial images to form an image recognition database and expanding the image recognition database using a visual language model includes: The image is classified and labeled according to the features of the subject being measured. A visual language model is used to generate corresponding images based on the same text description and similar images based on similar text descriptions, in order to expand the image recognition database; The step of selecting corresponding images from the image recognition database to form test questions based on the current evaluation task includes: Select images with similar text descriptions within the same category based on similar text descriptions; Select images with similar text descriptions from different categories based on similar text descriptions; Select the corresponding image based on the semantically different text description; The visual perception test of the subject includes: Calculate the accuracy rates for similar categories, dissimilar categories, and semantic differences to assess the subject's visual cognitive ability.

2. The visual cognitive assessment method based on multimodal language model for automatically generating images according to claim 1, characterized in that, The similar text descriptions are obtained by similarity matching based on a large language model.

3. A visual cognitive assessment device for automatically generating images based on a multimodal language model, characterized in that, The device includes: An image recognition database construction module is used to acquire initial images to form an image recognition database, and to expand the image recognition database using a visual language model; The image test question selection module is used to select corresponding images from the image recognition database to form test questions based on the current evaluation task. The visual cognition test and assessment module is used to generate voice prompts based on the test questions using the voice generation module, and to conduct visual cognition tests on the subject. The visual cognition assessment results output module is used to input visual cognition test results, image categories, text descriptions and category names of similar images, age, and education level information into the visual language model to generate a visual cognition assessment report; The process of acquiring initial images to form an image recognition database and expanding the image recognition database using a visual language model includes: The image is classified and labeled according to the features of the subject being measured. A visual language model is used to generate corresponding images based on the same text description and similar images based on similar text descriptions, in order to expand the image recognition database; The step of selecting corresponding images from the image recognition database to form test questions based on the current evaluation task includes: Select images with similar text descriptions within the same category based on similar text descriptions; Select images with similar text descriptions from different categories based on similar text descriptions; Select the corresponding image based on the semantically different text description; The visual perception test of the subject includes: Calculate the accuracy rates for similar categories, dissimilar categories, and semantic differences to assess the subject's visual cognitive ability.

4. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 2.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Cognitive competence evaluation system and method

    CN109589122A

  • Fine-grained multi-mode prompt learning method based on visual language pre-training model

    CN119538179A