Visual cognition evaluation method for automatically generating image based on multi-modal language model
Through the visual cognitive evaluation method based on multimodal language model, images and test questions are automatically generated, and databases are expanded and evaluation reports are generated using the visual language model, which solves the problems of high manual evaluation costs and neglected semantic information differences in the existing technology, and achieves efficient and accurate visual cognitive evaluation.
Patent Information
- Application Number
- CN202510369172.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing visual cognitive evaluation methods have problems such as high manual evaluation costs, long test cycles, and easy to be affected by subjective judgments. They have failed to effectively utilize the visual language model and ignore the differences in the internal semantic information of the image.
The multimodal language model is used to automatically generate images, and the image recognition database is formed by obtaining the initial image, and the database is expanded by using the visual language model, the corresponding images are selected to form test questions, and a voice prompt is generated for visual cognitive testing. Finally, the test results are input into the visual language model to generate an evaluation report.
It realizes automatic generation of test content, accurately evaluates visual cognitive ability, significantly reduces labor costs, improves the accuracy of evaluation and the quality and richness of the knowledge base content.
Smart Images

Figure CN120203519A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of cognitive function assessment, and particularly to a visual cognitive assessment method for automatically generating images based on a multimodal language model. Background Art
[0002] With the iterative upgrade of artificial intelligence technology, the assessment of visual cognitive abilities of people of different ages has gradually become a research hotspot. Traditional assessment systems mostly adopt manual evaluation methods, that is, tests are conducted by manually setting questions and standardized questionnaires. Such traditional methods have obvious limitations. On the one hand, there are significant human cost problems in the compilation and implementation of standardized questionnaires. The test cycle is usually as long as 2 - 3 weeks, and the assessment results are vulnerable to the subjective judgment of the testers. On the other hand, the hierarchical division of the visual complexity of existing test materials is rough, lacking systematic design based on cognitive development theory, and it is difficult to achieve differential assessment across different age groups.
[0003] In the prior art, although there have been some methods for assisting in cognitive assessment by means of computer vision and artificial intelligence technology, these methods usually only focus on the surface visual features of images and ignore the differences in the intrinsic semantic information of images. For example, when the visual differences between different pictures are large, it may be misjudged that different events are described; while for combinations of images with small semantic differences, it may be difficult to effectively distinguish them due to visual differences. In this case, how to effectively design test content and accurately assess the ability of people of different ages to understand image semantics has become an urgent problem to be solved. In addition, most existing visual cognitive assessment systems have not fully utilized the advanced technology of visual language models (VLM), and rely too much on text input at the human - machine interaction level, reducing the accuracy of the test.
[0004] Therefore, in the related art, there is an urgent need for a way to automatically generate test content and accurately assess and analyze visual cognitive abilities. Summary of the Invention
[0005] Based on this, in view of the above - mentioned technical problems, it is necessary to provide a visual cognitive assessment method for automatically generating images based on a multimodal language model, which can automatically generate test content and accurately assess and analyze visual cognitive abilities.
[0006] In a first aspect, this application provides a visual cognitive assessment method for automatically generating images based on a multimodal language model. The method includes:
[0007] Obtain initial images to form an image recognition database, and expand the image recognition database using a visual language model;
[0008] Select corresponding images from the image recognition database based on the current assessment task to form test questions;
[0009] Generate voice prompts using the voice generation module based on the test questions, and conduct visual cognitive tests on the subject;
[0010] Input the visual cognitive test results into the visual language model to generate a visual cognitive assessment report.
[0011] Optionally, in an embodiment of the present application, the obtaining the initial images to form an image recognition database and expanding the image recognition database using the visual language model includes:
[0012] Classify and divide the images based on the characteristics of the subject to be tested, and mark them by category.
[0013] Optionally, in an embodiment of the present application, the selecting corresponding images from the image recognition database to form test questions based on the current evaluation task includes:
[0014] Select similar text description pictures of the same category based on similar text descriptions;
[0015] Select similar text description pictures of different categories based on similar text descriptions;
[0016] Select corresponding pictures based on text descriptions with semantic differences.
[0017] Optionally, in an embodiment of the present application, the similar text descriptions are obtained by performing similarity matching based on a large language model.
[0018] Optionally, in an embodiment of the present application, the conducting visual cognitive tests on the subject includes:
[0019] Calculate the correct rate of the same category, the correct rate of different categories, and the correct rate of semantic differences to evaluate the visual cognitive ability of the subject.
[0020] In a second aspect, the present application also provides a visual cognitive assessment device for automatically generating images based on a multimodal language model. The device includes:
[0021] An image recognition database construction module, configured to obtain initial images to form an image recognition database, and expand the image recognition database using a visual language model;
[0022] An image test question selection module, configured to select corresponding images from the image recognition database to form test questions based on the current evaluation task;
[0023] A visual cognitive test evaluation module, configured to generate voice prompts using the voice generation module based on the test questions, and conduct visual cognitive tests on the subject;
[0024] A visual cognitive assessment result output module is used to input the visual cognitive test results into a visual language model to generate a visual cognitive assessment report.
[0025] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the steps of the methods described in the above respective embodiments.
[0026] In a fourth aspect, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the steps of the methods described in the above respective embodiments are implemented.
[0027] For the above visual cognitive assessment method for automatically generating images based on a multi-modal language model, first, an initial image is obtained to form an image recognition database, and the image recognition database is expanded using a visual language model. Then, corresponding images are selected from the image recognition database based on the current assessment task to form test questions. After that, a voice generation module is used to generate voice prompts based on the test questions to conduct a visual cognitive test on the subject. Finally, the visual cognitive test results are input into the visual language model to generate a visual cognitive assessment report. That is to say, by automatically generating accurate image descriptions through the VLM visual language model and using semantic fine-tuning technology to generate high-precision similar semantic images with large visual differences but small semantic differences, the dynamic expansion of the image knowledge base and the accurate assessment of subtle visual semantic differences are achieved, significantly reducing the labor cost and improving the quality and richness of the knowledge base content. At the same time, through the RAG technology, image retrieval and matching based on semantic similarity are realized, solving the problem of insufficient recognition accuracy of traditional visual difference assessment methods. Description of the Drawings
[0028] Figure 1 It is an application environment diagram of the visual cognitive assessment method for automatically generating images based on a multi-modal language model in an embodiment;
[0029] Figure 2 It is a flowchart of the visual cognitive assessment method for automatically generating images based on a multi-modal language model in an embodiment;
[0030] Figure 3 It is a structural block diagram of the visual cognitive assessment device for automatically generating images based on a multi-modal language model in an embodiment;
[0031] Figure 4 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0032] In order to make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.
[0033] The visual cognitive evaluation method for automatically generating images based on a multimodal language model provided by an embodiment of this application can be applied to an application environment such as Figure 1 shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be integrated on the server, or placed in the cloud or on other network servers. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0034] In one embodiment, as Figure 2 shown, a visual cognitive evaluation method for automatically generating images based on a multimodal language model is provided. Taking the case where this method is applied to the Figure 1 server as an example, the method includes the following steps:
[0035] S201: Obtain an initial image to form an image recognition database, and expand the image recognition database using a vision-language model.
[0036] In an embodiment of this application, first, images are collected through multiple channels (such as domestic and foreign picture books and related picture data sets, which have reached the cognitive standards of adults and can adapt to a wider user group) to form an image recognition database, ensuring that the content is rich and diverse, and at the same time meeting the cognitive and aesthetic needs of people of all ages. The image recognition database is a binary knowledge base composed of images and texts. Each image is accompanied by a corresponding text description, and the text description is automatically generated by a vision-language model (VLM) according to the picture, and can accurately capture and express the key information in the image.
[0037] There are often significant visual differences between different images. Due to differences in the environment, the placement of objects, and the distribution of people, the pixel differences between different images are significant. Without special processing of the dataset, a dataset with large visual differences can be obtained. However, images with large visual differences may not have significant semantic differences. For example, the differences in different perspectives of the same football game are extremely large. At this time, the subject being tested may think that the events described by the images are completely different due to the large visual differences, which places a high demand on the cognitive ability of the subject being tested. However, it is relatively difficult to find an image dataset with approximate semantic information. The Visual Language Model (VLM) can generate corresponding images through text descriptions. At the same time, by simply modifying the text description, images with different but similar semantic information can be generated. Examining the ability of the subject being tested to distinguish such images can more deeply determine whether the subject being tested truly understands the semantic information of the images. Therefore, the VLM is used to generate corresponding images based on the same text description and similar images based on similar text descriptions to expand the image recognition database.
[0038] In one embodiment of the present application, the step of obtaining the initial images to form the image recognition database and using the visual language model to expand the image recognition database includes:
[0039] Classify and divide the images based on the characteristics of the subject being tested and mark them by category.
[0040] In one embodiment of the present application, for the subject to be measured, due to their different ages, their ability to understand and appreciate images may also be different. For example, children's cognitive abilities improve rapidly in the early childhood stage, normal adults have relatively high cognitive abilities, while the cognitive abilities of the relatively older elderly group have degenerated. Therefore, it is also very important to reasonably divide the images in the image recognition database. On the one hand, according to the age characteristics of children, the images are divided into three levels: simple, medium, and complex, according to the cognitive difficulty and content complexity. For example, children in small, middle, and large classes in kindergarten are suitable for images with bright colors, simple images, and easy to understand (simple and medium levels), while students in grades 1-3 of primary school can gradually be exposed to images with richer content and more details (complex level). On the other hand, the process of humans understanding the world is divided into 3 different levels according to the picture information density (that is, the scene range and field of view width shown in a picture). Among them, the first level is the most complex landscape category, the second level is human interaction, and the third level is the object category (further subdivided into robot images that can move and the recognition of static single objects. Considering that the test subjects include children, single objects are divided into sports equipment, stationery, musical instruments, fruits, etc.). Specifically, for adults with normal cognition, the image content is designed to be more complex and abstract, aiming to challenge their advanced cognitive abilities. This level includes multi-level scenes (such as city panoramas), abstract art (such as modern paintings), etc., and requires users to have strong analysis and reasoning abilities. For the elderly with cognitive degeneration, the image content is simpler, more intuitive, and closer to daily life. This level includes family life (such as kitchen scenes), social activities (such as taking a walk in the park), etc., aiming to help the elderly maintain and exercise basic cognitive functions.
[0041] In this embodiment, by targeting the cognitive characteristics of the subjects to be measured at different age stages, an innovative classification method for image complexity and semantic information density is proposed, making the evaluation process more in line with the cognitive development characteristics of the subjects to be measured.
[0042] S203: Select corresponding images from the image recognition database to form test questions based on the current evaluation task.
[0043] In the embodiment of the present application, different subjects have different abilities and degrees of understanding of images. It is relatively difficult to directly ask young children or the elderly with severe cognitive degeneration to give a complete description of the images. Therefore, using the method of selecting pictures that match the description from different images is a simple and effective visual cognitive evaluation method. Different images are selected from the image recognition database to form test questions based on different evaluation tasks using the retrieval-augmented generation (RAG) technology for testing. The evaluation tasks include comparison of the same category, comparison of different categories, and semantic difference comparison.
[0044] Specifically, in an embodiment of the present application, the selection of corresponding images from the image recognition database based on the current evaluation task includes:
[0045] S301: Select similar text description pictures of the same category based on similar text descriptions.
[0046] S303: Select similar text description pictures of different categories based on similar text descriptions.
[0047] S305: Select corresponding pictures based on text descriptions with semantic differences.
[0048] In an embodiment of the present application, for the same-category comparison, for the same type of text description, according to the similarity degree between texts, select the pictures of the most similar text description to form the question. For example, based on a picture and its corresponding text description, find three other pictures and corresponding text descriptions with a relatively high similarity degree in the same-category dataset. The tested subject needs to find the picture that matches the correct description from the four pictures. For the different-category comparison, classifying categories is to evaluate the cognitive ability of the tested subject under different-category pictures. At the same time, when there are data for both same-category and different-category comparisons, it can be evaluated whether the cognitive problems of the tested subject are caused by different categories. Similarly, using the text similarity degree as the criterion, select the most similar pictures. The difference is that select the pictures of the most similar text description in different categories to form the question. For example, select one picture with the most similar description to the correct image from each of the three categories different from the correct image to form the question.
[0049] For the semantic difference comparison, it can evaluate the recognition ability of the tested subject for pictures with small semantic differences. Specifically, select one picture as the only correct picture, use the corresponding text description as the correct semantic information, make subtle adjustments to the selected text description to achieve a similar but slightly different effect. For example: for landscape-category images, we can change the landscape weather to change the semantic information of the landscape. For pictures of interpersonal interaction, we can change the number of people in the whole picture to increase or delete some semantic information. At the same time, with the help of the multi-modal ability of VLM, describe the changed semantic information to achieve the effect of only having subtle semantic differences from the original text. Generate corresponding pictures according to the modified semantics to form the question. In this way, even if the visual differences of the pictures may be large, the semantics only have subtle differences, which requires the tested subject to make a detailed comparison of the pictures to examine the detailed information in the description.
[0050] In an embodiment of the present application, the similar text descriptions are obtained by performing similarity matching based on a large language model.
[0051] In one embodiment of the present application, for selecting similar texts, an advanced large language model (LLM) is used for automatic similarity matching, and the constituent objects of the title are selected by having the large language model score the similarity degree between different texts. The scoring basis is as follows: 1. Described content; 2. Described time (occurrence time, duration); 3. Described location (the size of the event occurrence scene in space and the semantic information of the event occurrence scene); 4. Described object composition: whether there are people and the composition of the personnel. Specifically, taking the same-kind comparison as an example, the Prompt is shown in Table 1 below:
[0052]
[0053] Table 1
[0054] S205: Based on the test questions, a voice generation module is used to generate voice prompts for visual cognitive testing of the subject.
[0055] In the embodiment of the present application, due to the differences in literacy abilities of users of different age groups, for example, children and some elderly people may have difficulty accurately matching image content during the test because they have difficulty understanding the text, and there are often situations where they have difficulty understanding the text and cannot match the image. Therefore, during the test, traditional text recognition tasks are eliminated, and only the image recognition ability is examined. The Funasr voice generation module is introduced to convey the text description of the test questions to the subject in the form of voice. With the help of voice prompts, the subject can obtain key information without reading, so that they can concentrate more on the image recognition task, ensuring that the test results more accurately reflect their visual cognitive ability. The subject has different performances under different category images, and also shows different performances under test questions with semantic differences and visual differences. Therefore, users have individual differences in image recognition ability. These differences can be analyzed by the LLM to find out the internal reasons, and at the same time, the type ratio of subsequent training questions and the ratio of semantic and visual differences can be dynamically adjusted, enabling the subjects to continuously optimize their recognition level.
[0056] In one embodiment of the present application, the visual cognitive testing of the subject includes:
[0057] Calculating the correct rate of the same kind, the correct rate of different kinds, and the correct rate of semantic differences to evaluate the visual cognitive ability of the subject.
[0058] In one embodiment of the present application, to evaluate the recognition ability of the subject under test for pictures, different correct rates were statistically analyzed for three evaluation forms, namely the same-class correct rate, the different-class correct rate, and the semantic difference correct rate. And a comprehensive evaluation index (CEI) was calculated based on the same-class correct rate, the different-class correct rate, and the semantic difference correct rate combined with different influence degrees to evaluate the overall recognition ability. Specifically, the same-class correct rate (SCA) measures the ability to correctly recognize images within the same class, and is calculated as the number of correctly recognized same-class images divided by the total number of images tested in that class. The different-class correct rate (CCA) measures the ability to distinguish different-class images, and is calculated as the number of correctly classified cross-class images divided by the total number of cross-class images tested. The semantic difference correct rate (SDA) measures the ability to correctly recognize images with significant semantic differences, and is calculated as the number of correctly recognized semantic difference images divided by the total number of semantic difference images tested. Except for CEI, each index includes the overall correct rate and the correct rates for four sub-categories: Landscape, Interpersonal Interaction, Sports, and Static Objects. The specific calculation formulas are as follows:
[0059]
[0060]
[0061] CEI = ω1·SCA + ω2·CCA + ω3·SDA
[0062] Among them, ω1, ω2, and ω3 are weight coefficients.
[0063] S207: Input the visual cognitive test results into the visual language model to generate a visual cognitive evaluation report.
[0064] In the embodiments of the present application, various indicators of the same category, different categories, and semantic difference categories are obtained respectively in the evaluation stage. These indicators construct the recognition performance of the subject to be measured. When the category performance in the same-category accuracy rate is consistent with the category performance in the different-category accuracy rate, the category factor plays a very important role in the entire image recognition task. At the same time, when there are obvious differences between the semantic difference accuracy rate and the former two, it can be found that the subject to be measured has problems in understanding details. The evaluation results are input into the vision-language model as pre-information. At the same time, in different categories of image datasets, there are different semantic information. Summarizing the information is also input into the vision-language model as pre-information. The text descriptions and category names of all the same-category pictures will be input into the vision-language model, and thus the overall characteristics of the category are summarized, which is convenient as a factor for analysis. In addition, due to the differences in age and education level, there are also significant differences in the image recognition ability. This part of information can also be input into the vision-language model as the background information of the subject to be measured. Based on all the received pre-information, the vision-language model can analyze the performance of the subject to be measured under specific categories and generate an ability evaluation report, which is clearly presented in the form of a pdf picture.
[0065] In the above visual cognitive evaluation method for automatically generating images based on a multi-modal language model, first, an initial image is obtained to form an image recognition database, and the vision-language model is used to expand the image recognition database. Then, based on the current evaluation task, corresponding images are selected from the image recognition database to form test questions. Then, based on the test questions, a voice generation module is used to generate voice prompts to conduct a visual cognitive test on the subject. Finally, the visual cognitive test results are input into the vision-language model to generate a visual cognitive evaluation report. That is to say, by using the VLM vision-language model to automatically generate accurate image descriptions and using semantic fine-tuning technology to generate high-precision similar semantic images with large visual differences but small semantic differences, the dynamic expansion of the image knowledge base and the accurate evaluation of subtle visual semantic differences are realized, significantly reducing the labor cost and improving the quality and richness of the knowledge base content. At the same time, through the RAG technology, image retrieval and matching based on semantic similarity are realized, solving the problem of insufficient recognition accuracy of traditional visual difference evaluation methods.
[0066] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0067] Based on the same inventive concept, an embodiment of the present application further provides a visual cognitive evaluation device for automatically generating images based on a multi-modal language model for implementing the above-described visual cognitive evaluation method for automatically generating images based on a multi-modal language model. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the visual cognitive evaluation device for automatically generating images based on a multi-modal language model provided below can refer to the limitations on the visual cognitive evaluation method for automatically generating images based on a multi-modal language model in the above text, and will not be elaborated herein.
[0068] In one embodiment, as Figure 3 shown, a visual cognitive evaluation device 300 for automatically generating images based on a multi-modal language model is provided, including: an image recognition database construction module 301, an image test question selection module 303, a visual cognitive test evaluation module 305, and a visual cognitive evaluation result output module 307, where:
[0069] The image recognition database construction module 301 is configured to obtain initial images to form an image recognition database, and expand the image recognition database using a visual language model.
[0070] The image test question selection module 303 is configured to select corresponding images from the image recognition database based on the current evaluation task to form test questions.
[0071] The visual cognitive test evaluation module 305 is configured to generate voice prompts using a voice generation module based on the test questions to conduct a visual cognitive test on the subject.
[0072] The visual cognitive evaluation result output module 307 is configured to input the visual cognitive test result into a visual language model to generate a visual cognitive evaluation report.
[0073] In one embodiment of the present application, the image recognition database construction module is further configured to:
[0074] Classify and divide the images based on the characteristics of the subject to be measured, and mark them by category.
[0075] In one embodiment of the present application, the image test question selection module is further configured to:
[0076] Select pictures with similar text descriptions of the same category based on similar text descriptions;
[0077] Select pictures with similar text descriptions of different categories based on similar text descriptions;
[0078] Select corresponding pictures based on text descriptions with semantic differences.
[0079] In one embodiment of the present application, the similar text descriptions are obtained by performing similarity matching based on a large language model.
[0080] In one embodiment of the present application, the visual cognitive test evaluation module is further configured to:
[0081] Calculate the correct rate of the same category, the correct rate of different categories, and the correct rate of semantic differences to evaluate the visual cognitive ability of the subject.
[0082] Each module in the above visual cognitive evaluation device for automatically generating images based on a multi-modal language model can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0083] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a visual cognitive assessment method for automatically generating images based on a multi-modal language model. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0084] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0085] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0086] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0087] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0088] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0089] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0090] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0091] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A visual cognitive assessment method for automatically generating images based on a multimodal language model, characterized in that: The method comprises: Obtaining initial images to form an image recognition database, and expanding the image recognition database using a visual language model; Selecting corresponding images from the image recognition database to form test questions based on the current assessment task; Based on the test questions, a speech generation module is used to generate speech prompts to conduct a visual cognition test on the subject; The visual cognition test results are input into the visual language model to generate a visual cognition assessment report.
2. The visual cognition assessment method for automatically generating images based on a multimodal language model according to claim 1, characterized in that: The obtaining of the initial image to form an image recognition database and the use of the visual language model to expand the image recognition database include: The images are classified and divided based on the characteristics of the tested subjects and marked by category.
3. The visual cognition assessment method for automatically generating images based on a multimodal language model according to claim 1, characterized in that: The selecting corresponding images from the image recognition database to form test questions based on the current assessment task includes: Select similar text description images of the same category based on similar text descriptions; Selecting pictures with similar text descriptions of different categories based on similar text descriptions; Select corresponding images based on semantically different text descriptions.
4. The visual cognition assessment method for automatically generating images based on a multimodal language model according to claim 3, characterized in that: The similar text descriptions are obtained by performing similarity matching based on a large language model.
5. The visual cognition assessment method for automatically generating images based on a multimodal language model according to claim 1, characterized in that: The visual cognition test performed on the subject comprises: The same category accuracy, different category accuracy, and semantic difference accuracy were calculated to evaluate the subject’s visual cognitive ability.
6. A visual cognition assessment device for automatically generating images based on a multimodal language model, characterized in that: The device comprises: An image recognition database building module is used to obtain an initial image to form an image recognition database, and expand the image recognition database using a visual language model; An image test question selection module, used to select corresponding images from the image recognition database to form test questions based on the current assessment task; A visual cognition test assessment module, for generating voice prompts based on the test questions using a voice generation module to conduct a visual cognition test on the subject; The visual cognition assessment result output module is used to input the visual cognition test results into the visual language model to generate a visual cognition assessment report.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Cognitive competence evaluation system and method
CN109589122A
Incremental learning method and device for visual joint features and storage medium
CN116958778A
Cognitive competence testing method and device, equipment and storage medium
CN117838047A
Performance evaluation method and device of visual language model in positioning task
CN118736355A
Fine-grained multi-mode prompt learning method based on visual language pre-training model
CN119538179A