Visual language model training method, reasoning method, system, electronic equipment, storage medium and program product

The visual language model addresses the challenge of emotion and state perception in multi-modal models by training on facial attributes and providing personalized feedback, enhancing user experience in human-machine interaction.

CN120316554AActive Publication Date: 2025-07-15INST OF AUTOMATION CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510786859.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-15
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

The existing multimodal model lacks ability to perceive user emotions and states, and it is difficult to provide adaptive feedback, limiting its application in human-computer interaction scenarios.

Method used

By building a training data set for facial attribute perception, visual language models are trained, including visual encoder, projection layer, text word segmenter and large language model, the perception of wearable, appearance and expression attributes is realized, and adaptive feedback is generated through multimodal fusion.

Benefits of technology

It realizes accurate perception of user emotions and state, generates adaptive and emotional intelligent feedback, and improves user experience and interactive effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316554A_ABST
    Figure CN120316554A_ABST
Patent Text Reader

Abstract

The invention provides a visual language model training method and system, a reasoning method and system, electronic equipment, a storage medium and a program product. The visual language model comprises a visual encoder, a projection layer, a text word segmentation device and a large language model, and the training method comprises the following steps: constructing a first training data set for face attribute perception; performing first-stage training on the visual language model based on the first training data set to update parameters of a visual encoder and a projection layer, the first-stage training being used for enabling the visual language model to have a face attribute perception function; constructing a second training data set comprising a wearing attribute-oriented question and answer data sample, an appearance attribute-oriented question and answer data sample and an expression attribute-oriented question and answer data sample; and performing second-stage training on the visual language model based on the second training data set to update parameters of the projection layer and the large language model, the second-stage training being used for enabling the visual language model to have a function of providing targeted feedback for face attributes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of multimodal models, and more specifically, to a training method, an inference method, a system, an electronic device, a storage medium, and a program product for a vision-language model. Background Art

[0002] In recent years, with the rapid development of deep learning and large-scale pre-trained models, multimodal models have made remarkable progress in the fields of image understanding and natural language processing. These models can handle the associations between images, text, and other data types.

[0003] During the process of human-computer interaction, if the user's emotions and states can be accurately perceived and adaptive feedback can be provided, the user experience can be significantly improved.

[0004] However, existing multimodal models cannot well perceive the user's emotions and states and provide adaptive feedback based on the perception results, which limits the wide application of multimodal models in actual human-computer interaction scenarios. Summary of the Invention

[0005] Exemplary embodiments of the present disclosure are directed to a training method, an inference method, a system, an electronic device, a storage medium, and a program product for a vision-language model, which can solve at least one of the above problems existing in the prior art.

[0006] According to a first aspect of an embodiment of the present disclosure, there is provided a training method for a vision-language model, where the vision-language model includes: a vision encoder, a projection layer, a text tokenizer, and a large language model. The training method includes: constructing a first training dataset for face attribute perception, where the face attributes include: wearing attributes, appearance attributes, and expression attributes; performing a first-stage training on the vision-language model based on the first training dataset to update the parameters of the vision encoder and the projection layer, where the first-stage training is used to enable the vision-language model to have a face attribute perception function; constructing a second training dataset, where the second training dataset includes: question-and-answer data samples for the wearing attributes, question-and-answer data samples for the appearance attributes, and question-and-answer data samples for the expression attributes; after completing the first-stage training, performing a second-stage training on the vision-language model based on the second training dataset to update the parameters of the projection layer and the large language model, where the second-stage training is used to enable the vision-language model to have a function of providing targeted feedback for face attributes.

[0007] Optionally, the training method further includes: constructing a third training dataset for character style transfer; after completing the second-stage training, performing third-stage training on the vision-language model based on the third training dataset to update the parameters of the projection layer and the large language model; wherein, the third-stage training is used to enable the vision-language model to have the function of providing personalized feedback reflecting a specific character style.

[0008] Optionally, the wearable attributes include at least one of the following items: whether wearing a mask, whether wearing glasses, whether wearing a hat, whether wearing earrings; the appearance attributes include at least one of the following items: gender, age, ethnicity, mouth, eyes, beard, makeup, hair length, hair color, hairstyle; the expression attributes include at least one of the following items: happy, surprised, neutral, sad, afraid, angry, disgusted.

[0009] Optionally, the function of providing targeted feedback for face attributes includes: the function of providing life and health suggestions for the wearable attributes; the function of providing image management suggestions for the appearance attributes; the function of providing emotional responses for the expression attributes.

[0010] Optionally, constructing the third training dataset for character style transfer includes: generating character portraits of multiple specific characters that appear in the multiple literary works based on the structured dialogues extracted from the multiple literary works, and generating intrinsic knowledge Q&A data samples for each of the multiple specific characters; for each of the multiple specific characters, adjusting the Q&A data samples in the second training dataset based on the character portrait of the specific character to obtain Q&A data samples reflecting the style of the specific character; for each of the multiple specific characters, generating general Q&A data samples reflecting the style of the specific character based on the character portrait of the specific character.

[0011] According to a second aspect of the embodiments of the present disclosure, there is provided a reasoning method for a vision-language model, where the vision-language model includes: a vision encoder, a projection layer, a text tokenizer, and a large language model, and the reasoning method includes: obtaining text content corresponding to a user input; obtaining a face image of the user captured by an image acquisition device; inputting the face image into the vision encoder, and inputting the vector output by the vision encoder into the projection layer to obtain visual features output by the projection layer; inputting the text content into the text tokenizer to obtain text features output by the text tokenizer; inputting the visual features and the text features into the large language model to obtain feedback content output by the large language model for providing feedback on the face attributes of the user input and the face image; wherein, the vision-language model is trained by executing the training method as described above.

[0012] According to a third aspect of the embodiments of the present disclosure, there is provided a training system for a vision-language model, where the vision-language model includes: a vision encoder, a projection layer, a text tokenizer, and a large language model. Among them, the training system includes: a first data construction unit configured to construct a first training dataset for face attribute perception, where the face attributes include: wearing attributes, appearance attributes, and expression attributes; a first training unit configured to perform a first-stage training on the vision-language model based on the first training dataset to update the parameters of the vision encoder and the projection layer, where the first-stage training is used to enable the vision-language model to have the function of face attribute perception; a second data construction unit configured to construct a second training dataset, where the second training dataset includes: question-and-answer data samples for the wearing attributes, question-and-answer data samples for the appearance attributes, and question-and-answer data samples for the expression attributes; a second training unit configured to, after completing the first-stage training, perform a second-stage training on the vision-language model based on the second training dataset to update the parameters of the projection layer and the large language model, where the second-stage training is used to enable the vision-language model to have the function of providing targeted feedback for face attributes.

[0013] According to a fourth aspect of the embodiments of the present disclosure, there is provided an inference system for a vision-language model, where the vision-language model includes: a vision encoder, a projection layer, a text tokenizer, and a large language model. Among them, the inference system includes: a text acquisition unit configured to acquire the text content corresponding to the user input; an image acquisition unit configured to acquire the face image of the user captured by an image acquisition device; a vision feature acquisition unit configured to input the face image into the vision encoder and input the vector output by the vision encoder into the projection layer to obtain the vision feature output by the projection layer; a text feature acquisition unit configured to input the text content into the text tokenizer to obtain the text feature output by the text tokenizer; a feedback content acquisition unit configured to input the vision feature and the text feature into the large language model to obtain the feedback content output by the large language model for providing feedback on the face attributes of the user input and the face image; where the vision-language model is trained by executing the training method as described above.

[0014] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing instructions, which when executed by a processor of an electronic device, enables the electronic device to execute the training method and / or the inference method as described above.

[0015] According to a sixth aspect of the embodiments of the present disclosure, an electronic device is provided, which includes: at least one processor; and at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method and / or the inference method as described above.

[0016] According to a seventh aspect of the embodiments of the present disclosure, a computer program product is provided, including computer-executable instructions, which implement the training method and / or the inference method as described above when executed by at least one processor.

[0017] According to the training method, inference method, system, electronic device, storage medium and program product of the vision-language model according to the exemplary embodiments of the present disclosure, in the process of human-computer interaction, by recognizing and analyzing the user's face image, the user's emotion and state can be accurately perceived, and behavior feedback adapted to the user's emotion and state can be generated, thereby significantly improving the user experience.

[0018] In addition, the style of a specific role (for example, a pre-specified system role) can be transferred to the behavior feedback to form personalized behavior feedback after specific role-playing.

[0019] In the following description, some aspects and / or advantages of the general concept of the present disclosure will be elaborated, and some aspects and / or advantages will be learned from the following description or the implementation of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] These and / or other aspects and advantages of the present application will become clearer and easier to understand from the following detailed description of the embodiments of the present application in conjunction with the accompanying drawings, where: Figure 1 An example of a vision-language model according to an exemplary embodiment of the present disclosure is shown; Figure 2 A flowchart of a training method of a vision-language model according to an exemplary embodiment of the present disclosure is shown; Figure 3 An example of a training method of a vision-language model according to an exemplary embodiment of the present disclosure is shown; Figure 4 A flowchart of a method for constructing a third training dataset for role style transfer according to an exemplary embodiment of the present disclosure is shown; Figure 5 An example of a method for constructing a third training dataset for role style transfer according to an exemplary embodiment of the present disclosure is shown; Figure 6 A flowchart of an inference method of a vision-language model according to an exemplary embodiment of the present disclosure is shown; Figure 7 An example of an inference method of a vision - language model according to an exemplary embodiment of the present disclosure is shown; Figure 8 A structural block diagram of a training system of a vision - language model according to an exemplary embodiment of the present disclosure is shown; Figure 9 A structural block diagram of an inference system of a vision - language model according to an exemplary embodiment of the present disclosure is shown; Figure 10 A structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown. Detailed implementation manners

[0021] Reference will now be made in detail to the embodiments of the present disclosure. Examples of the embodiments are shown in the accompanying drawings, in which the same reference numerals always refer to the same components. The following embodiments will be described with reference to the drawings in order to explain the present disclosure.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above - mentioned drawings are used to distinguish similar objects and do not necessarily have to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0023] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel cases including "any one of the several items", "any combination of multiple items of the several items", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0024] Currently, most multimodal models are still in the initial stage of their ability to perceive human face attributes (such as emotions, gender, age, etc.), and have not been able to effectively provide immediate behavioral feedback based on human face attribute perception. This limits the wide application of multimodal models in actual human-computer interaction scenarios, especially in scenarios that require personalized interaction. In addition, most existing human face attribute recognition systems focus on object recognition in static images, and have insufficient understanding of the changes in human face attributes and their meanings in dynamic environments, making it difficult to achieve efficient human-computer interaction. Therefore, it is necessary to develop a system that can perceive human face attributes and provide immediate behavioral feedback to improve the user experience and interaction effect.

[0025] The present disclosure aims to solve the problem of insufficient visual information perception ability of existing dialogue systems. By combining human face attribute perception and text analysis, more adaptable and emotional intelligent feedback is generated. The following will be described in detail in conjunction with Figures 1 to 10 for detailed description.

[0026] Figure 1 An example of a vision-language model according to an exemplary embodiment of the present disclosure is shown.

[0027] The vision-language model according to an exemplary embodiment of the present disclosure includes: a visual encoder, a projection layer, a text tokenizer, and a large language model.

[0028] The visual encoder is used to encode a human face image to obtain a corresponding vector.

[0029] The projection layer is used to extract features from the vector output by the visual encoder to obtain visual features.

[0030] The text tokenizer is used to process text content to obtain text features. As an example, the text tokenizer can capture the semantic information of the text content and generate context-related embedding vectors for each word in the text content.

[0031] The large language model is used to perform multimodal fusion on the visual features and text features, and output feedback content for providing feedback on the text content and the human face attributes of the human face image. In addition, the large language model can also transfer the style of a specific role (for example, a pre-specified system role) to the feedback content to form personalized feedback content after specific role-playing.

[0032] Specifically, the large language model can perceive the user's emotions and state based on the visual features, and generate appropriate behavioral feedback in real time according to the perception results and text analysis results to respond to the user's state changes. In addition, it can also inject the tone and expression style of a specific role (such as language style, knowledge background, and catchphrases), etc., to reflect the role personality in the behavioral feedback and enhance the interaction fun and user experience.

[0033] Figure 2A flowchart showing a training method of a vision - language model according to an exemplary embodiment of the present disclosure.

[0034] Referring to Figure 2 and Figure 3 In step S201, a first training dataset for face attribute perception is constructed.

[0035] Face attributes include: wearing attributes, appearance attributes, and expression attributes.

[0036] As an exemplary embodiment, the wearing attributes may include but are not limited to at least one of the following items: whether wearing a mask, whether wearing glasses, whether wearing a hat, whether wearing earrings.

[0037] As an exemplary embodiment, the appearance attributes may include but are not limited to at least one of the following items: gender, age, ethnicity, mouth, eyes, beard, makeup, hair length, hair color, hairstyle.

[0038] As an exemplary embodiment, the expression attributes may include but are not limited to at least one of the following items: happy, surprised, neutral, sad, afraid, angry, disgusted.

[0039] As an exemplary embodiment, through prompt engineering (Prompt) and a specified large - language model (e.g., a large - language model other than the above vision - language model), each piece of data (face image - attribute label) in the publicly available face attribute dataset and / or private face attribute dataset can be converted into a question - and - answer data format (face image - question - answer about face attribute description). Then, the converted data can be data - cleaned to remove ambiguous and fuzzy answers, and the cleaned data can be used as the question - and - answer data samples of the first training dataset.

[0040] In step S202, the vision - language model is trained in the first stage based on the first training dataset to update the parameters of the vision encoder and the projection layer. The first - stage training is used to enable the vision - language model to have the face attribute perception function.

[0041] As an example, in the first - stage training, the parameters of the text tokenizer and the large - language model can be frozen, and only the parameters of the vision encoder and the projection layer are trained.

[0042] As an example, in the first-stage training, the face images in the Q&A data samples in the first training dataset can be input into the visual encoder, and the vectors output by the visual encoder are input into the projection layer to obtain the visual features output by the projection layer; the questions in the Q&A data samples are input into the text tokenizer to obtain the text features output by the text tokenizer; the visual features and text features are input into the large language model to obtain the inference results output by the large language model; based on the difference between the answers and the inference results in the Q&A data samples, the loss is calculated; based on the calculated loss, the parameters of the visual encoder and the projection layer are updated.

[0043] In step S203, a second training dataset is constructed.

[0044] The second training dataset includes: Q&A data samples for wearable attributes, Q&A data samples for appearance attributes, and Q&A data samples for expression attributes.

[0045] The Q&A data samples for wearable attributes may include: face images, questions, and answers for wearable attributes of the face images. As an example, the answers for wearable attributes of the face images can be: life and health suggestions provided for whether the user wears a mask, glasses, hat, earrings, etc. For example, if the user does not wear a mask, suggestions for wearing a mask to maintain health in an environment with poor air quality can be provided for the user.

[0046] The Q&A data samples for appearance attributes may include: face images, questions, and answers for appearance attributes of the face images. As an example, the answers for appearance attributes of the face images can be: image management suggestions provided for the user's appearance attributes such as gender, age, ethnicity, eyes, mouth, hair, etc. For example, suitable hairstyle changes can be recommended according to the user's hairstyle and hair color, or suitable makeup styles and products can be recommended according to the user's skin color and facial features.

[0047] The Q&A data samples for expression attributes may include: face images, questions, and answers for expression attributes of the face images. As an example, the answers for expression attributes of the face images can be: emotional responses provided for the emotions and states reflected by the user's expression attributes. For example, targeted feedback (such as encouragement, comfort), suggestions (such as guiding the user to carry out emotional regulation activities). For example, comforting statements or recommendations for relaxation activities can be provided for the user's negative emotions.

[0048] As an exemplary embodiment, through prompt engineering and a specified large language model (e.g., a large language model other than the above-mentioned vision-language model), question-and-answer data samples for wearable attributes, question-and-answer data samples for appearance attributes, and question-and-answer data samples for expression attributes can be generated. Then, the generated question-and-answer data samples can be data-cleaned to remove ambiguous and vague answers, and the cleaned data can be used as the question-and-answer data samples of the second training dataset.

[0049] It should be understood that step S203 may be executed before step S204, and the present disclosure does not limit the execution order of step S203, step S201, and step S202. For example, step S203 may be executed after step S201.

[0050] In step S204, after completing the first-stage training, the vision-language model is trained in a second stage based on the second training dataset to update the parameters of the projection layer and the large language model.

[0051] The second-stage training is used to enable the vision-language model to have the function of providing targeted feedback for face attributes. As an exemplary embodiment, the function of providing targeted feedback for face attributes may include but is not limited to: the function of providing life and health suggestions for wearable attributes; the function of providing image management suggestions for appearance attributes; the function of providing emotional responses for expression attributes.

[0052] As an example, in the second-stage training, the parameters of the text tokenizer and the vision encoder can be frozen, and only the parameters of the projection layer and the large language model are trained.

[0053] As an example, in the second-stage training, the face images in the question-and-answer data samples of the second training dataset can be input into the vision encoder, and the vectors output by the vision encoder are input into the projection layer to obtain the visual features output by the projection layer; the questions in the question-and-answer data samples are input into the text tokenizer to obtain the text features output by the text tokenizer; the visual features and the text features are input into the large language model to obtain the inference results output by the large language model; based on the difference between the answers and the inference results in the question-and-answer data samples, the loss is calculated; based on the calculated loss, the parameters of the projection layer and the large language model are updated. For example, the loss of image-text matching and language generation can be used for optimization, and an adaptive learning rate and a mixed-precision strategy can be used to improve the training efficiency.

[0054] The vision-language model trained according to the exemplary embodiment of the present disclosure can, through multi-modal fusion technology, real-time perceive the user's emotions and states and generate feedback strategies for wearable matching, image management, emotional comfort, etc.

[0055] In addition, with reference to Figure 3, the training method of the vision - language model according to the exemplary embodiments of the present disclosure may further include: step S205 and step S206.

[0056] In step S205, a third training dataset for character style transfer is constructed.

[0057] Next, the exemplary embodiments of step S205 will be described in conjunction with Figure 4 and Figure 5 and will not be elaborated here for the time being.

[0058] In step S206, after the second - stage training is completed, the vision - language model is trained in a third stage based on the third training dataset to update the parameters of the projection layer and the large - language model.

[0059] The third - stage training is used to enable the vision - language model to have the function of providing personalized feedback reflecting a specific character style. That is, in the third - stage training, personalized behavior feedback learning is carried out, enabling the model to have the ability to provide feedback according to the speaking style of a character (such as Sun Wukong), so as to be able to reflect the character personality in the feedback, enhancing the interactive interest and user experience.

[0060] As an example, in the third - stage training, the parameters of the text tokenizer and the vision encoder can be frozen, and only the parameters of the projection layer and the large - language model are trained.

[0061] Figure 4 The flowchart showing the method for constructing a third training dataset for character style transfer according to the exemplary embodiments of the present disclosure is shown.

[0062] Referring to Figure 4 and Figure 5 , in step S401, based on the structured conversations extracted from multiple literary works, character portraits of multiple specific characters appearing in the multiple literary works are generated, and inherent knowledge Q&A data samples for each of the multiple specific characters are generated.

[0063] As an example, the types of literary works may include but are not limited to: novels, film and television scripts, dramas, and fairy tales.

[0064] It should be understood that the multiple literary works may include diverse literary works with high quality and wide coverage, ensuring the richness, representativeness, and integrity of the data sources.

[0065] As an example, GPT can be used to segment the conversations in literary works to extract structured conversations.

[0066] As an example, based on the structured conversations extracted from literary works, character portraits can be generated using GPT, that is, character definitions can be made. A character portrait is a description of a character, which may specifically include, but is not limited to: personality, language style (e.g., catchphrases, tone, expression style), behavior tendencies, and emotional characteristics.

[0067] As an example, based on the structured conversations extracted from literary works, through Prompt engineering and GPT, Q&A data for character-specific knowledge can be generated to cover common questions regarding character-specific knowledge. For example, the Q&A data for character-specific knowledge can be triples consisting of questions - answers - confidence levels, and high-quality non-duplicate data can be selected from them based on the confidence levels as the Q&A data samples in the third training dataset.

[0068] In step S402, for each of the multiple specific characters, based on the character portrait of that specific character, the Q&A data samples in the second training dataset are adjusted to obtain Q&A data samples that embody the style of that specific character.

[0069] That is, a character style transfer is performed on the Q&A data samples in the second training dataset. As an example, character prompts, system instructions, and retrieval enhancements based on dialogue engineering can be designed, and GPT can be used to modify the Q&A data samples in the second training dataset. For example, the language style, knowledge background, catchphrases, etc. of the character can be injected into the answers.

[0070] In step S403, for each of the multiple specific characters, based on the character portrait of that specific character, general Q&A data samples that embody the style of that specific character are generated.

[0071] That is, a character style transfer is performed on the general Q&A data samples other than those for face attributes. As an example, character prompts, system instructions, and retrieval enhancements based on dialogue engineering can be designed, and GPT can be used to modify the Q&A data samples in the general Q&A dataset. For example, the language style, knowledge background, catchphrases, etc. of the character can be injected into the answers.

[0072] It should be understood that the Q&A data samples generated in steps S401 - S403 together constitute the third training dataset.

[0073] Figure 6 A flowchart showing an inference method of a vision - language model according to an exemplary embodiment of the present disclosure is presented. The vision - language model includes: a vision encoder, a projection layer, a text tokenizer, and a large - language model. The vision - language model is trained by performing the training method as described in the above - mentioned exemplary embodiment.

[0074] Refer to Figure 6, in step S601, obtain the text content corresponding to the user input.

[0075] As an example, the user input can be voice input or text input. If the user input is voice input, convert the voice input into text.

[0076] In step S602, obtain the face image of the user captured by the image acquisition device.

[0077] In step S603, input the face image into the visual encoder, and input the vector output by the visual encoder into the projection layer to obtain the visual features output by the projection layer.

[0078] In step S604, input the text content into the text tokenizer to obtain the text features output by the text tokenizer.

[0079] In step S605, input the visual features and text features into the large language model to obtain the feedback content output by the large language model for providing feedback on the user input and the face attributes of the face image, so as to respond to the user's state changes in real time.

[0080] In addition, the large language model can also inject the language style, knowledge background, and catchphrases of a specific role into the feedback content, that is, generate feedback content that can reflect the personality of a specific role, enhancing the interaction fun and user experience.

[0081] In addition, as an exemplary embodiment, refer to Figure 7 , the historical interaction data with the user can be combined, and the deep fusion of multi-modal information can be achieved through an adaptive learning mechanism. As an example, the feedback strategy can be optimized using reinforcement learning based on the historical interaction data with the user to improve the interaction effect. For example, the historical interaction data with the user can be recorded, and more suitable feedback content can be generated according to the user's historical emotional responses, preferences, and feedback through reinforcement learning. That is, the feedback content can be dynamically adjusted through an adaptive learning mechanism to optimize the interaction experience.

[0082] The present disclosure innovatively applies multi-modal data fusion to the field of behavior feedback, not only improving the accuracy of behavior feedback, but also enhancing the personalization and interactivity of the user experience. It is applicable to multiple industries such as education, medical care, and entertainment, and has important application potential in fields such as intelligent assistants, virtual customer service, and affective computing, with broad application prospects and social value.

[0083] Figure 8 Show a structural block diagram of a training system of a vision-language model according to an exemplary embodiment of the present disclosure. The vision-language model includes: a visual encoder, a projection layer, a text tokenizer, and a large language model.

[0084] Refer to Figure 8, the training system 800 of the vision - language model according to an exemplary embodiment of the present disclosure includes: a first data construction unit 801, a first training unit 802, a second data construction unit 803, and a second training unit 804.

[0085] Specifically, the first data construction unit 801 is configured to construct a first training data set for face - attribute perception, where the face attributes include: wearing attributes, appearance attributes, and expression attributes.

[0086] The first training unit 802 is configured to perform a first - stage training on the vision - language model based on the first training data set to update the parameters of the vision encoder and the projection layer, where the first - stage training is used to enable the vision - language model to have the function of face - attribute perception.

[0087] The second data construction unit 803 is configured to construct a second training data set, where the second training data set includes: question - and - answer data samples for the wearing attributes, question - and - answer data samples for the appearance attributes, and question - and - answer data samples for the expression attributes.

[0088] The second training unit 804 is configured to perform a second - stage training on the vision - language model based on the second training data set after completing the first - stage training to update the parameters of the projection layer and the large - language model, where the second - stage training is used to enable the vision - language model to have the function of providing targeted feedback for face attributes.

[0089] Figure 9 The structural block diagram of the inference system of the vision - language model according to an exemplary embodiment of the present disclosure is shown. The vision - language model includes: a vision encoder, a projection layer, a text tokenizer, and a large - language model. The vision - language model is trained by executing the training method as described in the above - mentioned exemplary embodiment.

[0090] Refer to Figure 9 , the inference system 900 of the vision - language model according to an exemplary embodiment of the present disclosure includes: a text acquisition unit 901, an image acquisition unit 902, a visual feature acquisition unit 903, a text feature acquisition unit 904, and a feedback content acquisition unit 905.

[0091] Specifically, the text acquisition unit 901 is configured to acquire the text content corresponding to the user input.

[0092] The image acquisition unit 902 is configured to acquire the face image of the user captured by an image acquisition device.

[0093] The visual feature acquisition unit 903 is configured to input the face image into the visual encoder, and input the vector output by the visual encoder into the projection layer to obtain the visual features output by the projection layer.

[0094] The text feature acquisition unit 904 is configured to input the text content into the text tokenizer to obtain the text features output by the text tokenizer.

[0095] The feedback content acquisition unit 905 is configured to input the visual features and the text features into the large language model to obtain the feedback content output by the large language model for providing feedback on the user input and the face attributes of the face image.

[0096] It should be understood that the specific processes executed by the above system have been described in detail with reference to Figures 1 to 7 and the relevant details will not be elaborated here.

[0097] It should be understood that each unit in the above system can be implemented as a hardware component and / or a software component.

[0098] Figure 10 A structural block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.

[0099] Referring to Figure 10 , the electronic device 1000 includes: at least one memory 1001 and at least one processor 1002. A set of computer-executable instructions is stored in the at least one memory 1001. When the set of computer-executable instructions is executed by the at least one processor 1002, the training method and / or the inference method described in the above exemplary embodiment is executed.

[0100] As an example, the electronic device can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device does not have to be a single electronic device, and can also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device that is interconnected with a local or remote device (for example, via wireless transmission).

[0101] In the electronic device, the processor 1002 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example and not a limitation, the processor 1002 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0102] The processor 1002 can run instructions or code stored in the memory 1001, where the memory 1001 can also store data. The instructions and data can also be sent and received over a network via a network interface device, where the network interface device can employ any known transmission protocol.

[0103] The memory 1001 can be integrated with the processor 1002. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 1001 can include separate devices such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 1001 and the processor 1002 can be operatively coupled or can communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 1002 can read files stored in the memory.

[0104] In addition, the electronic device can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device can be connected to each other via a bus and / or a network.

[0105] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are run by at least one processor, the at least one processor is caused to execute the training method and / or the inference method as described in the above exemplary embodiment. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), card memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0106] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, wherein the instructions in the computer program product can be executed by at least one processor to perform the training method and / or the inference method as described in the above exemplary embodiment.

[0107] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0108] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for a visual language model, characterized in that The visual language model includes: a visual encoder, a projection layer, a text tokenizer, and a large language model. Among them, the training method includes: Construct a first training dataset for face attribute perception, where the face attributes include: wearing attributes, appearance attributes, and expression attributes; Based on the first training dataset, conduct the first-stage training on the visual language model to update the parameters of the visual encoder and the projection layer. Among them, the first-stage training is used to enable the visual language model to have the function of face attribute perception; Construct a second training dataset, where the second training dataset includes: question-and-answer data samples for the wearing attributes, question-and-answer data samples for the appearance attributes, and question-and-answer data samples for the expression attributes; After completing the first-stage training, based on the second training dataset, conduct the second-stage training on the visual language model to update the parameters of the projection layer and the large language model. Among them, the second-stage training is used to enable the visual language model to have the function of providing targeted feedback for face attributes.

2. The training method according to claim 1, wherein The training method further includes: Construct a third training dataset for character style transfer; After completing the second-stage training, based on the third training dataset, conduct the third-stage training on the visual language model to update the parameters of the projection layer and the large language model; Among them, the third-stage training is used to enable the visual language model to have the function of providing personalized feedback reflecting a specific character style.

3. The training method according to claim 1, wherein The wearing attributes include at least one of the following items: whether wearing a mask, whether wearing glasses, whether wearing a hat, whether wearing earrings; The appearance attributes include at least one of the following items: gender, age, ethnic group, mouth, eyes, beard, makeup, hair length, hair color, hairstyle; The expression attributes include at least one of the following items: happy, surprised, neutral, sad, afraid, angry, disgusted.

4. The training method according to claim 1, wherein The function of providing targeted feedback for face attributes includes: The function of providing life and health suggestions for the wearing attributes; The function of providing image management suggestions for the appearance attributes; The function of providing emotional responses for the expression attributes.

5. The training method according to claim 2, characterized in that The construction of the third training dataset for character style transfer includes: Based on the structured conversations extracted from multiple literary works, generate character portraits of multiple specific characters that appear in the multiple literary works, and generate inherent knowledge question-and-answer data samples for each specific character among the multiple specific characters; For each specific character among the multiple specific characters, adjust the question-and-answer data samples in the second training dataset based on the character portrait of the specific character to obtain question-and-answer data samples reflecting the style of the specific character; For each specific character among the multiple specific characters, based on the character portrait of the specific character, generate general question-and-answer data samples reflecting the style of the specific character.

6. A reasoning method for a vision language model, characterized in that, The visual language model includes: a visual encoder, a projection layer, a text tokenizer, and a large language model. Among them, the inference method includes: Obtain the text content corresponding to the user input; Obtain the face image of the user captured by the image acquisition device; Input the face image into the vision encoder, and input the vector output by the vision encoder into the projection layer to obtain the visual features output by the projection layer; Input the text content into the text tokenizer to obtain the text features output by the text tokenizer; Input the visual features and the text features into the large language model to obtain the feedback content output by the large language model for providing feedback on the face attributes of the user input and the face image; Wherein, the vision language model is trained by executing the training method described in any one of claims 1 to 5.

7. A training system for a visual language model, characterized in that, The vision language model includes: a vision encoder, a projection layer, a text tokenizer, and a large language model. Among them, the training system includes: A first data construction unit configured to construct a first training dataset for face attribute perception, where the face attributes include: wearing attributes, appearance attributes, and expression attributes; A first training unit configured to perform a first-stage training on the vision language model based on the first training dataset to update the parameters of the vision encoder and the projection layer, where the first-stage training is used to enable the vision language model to have the face attribute perception function; A second data construction unit configured to construct a second training dataset, where the second training dataset includes: question-and-answer data samples for the wearing attributes, question-and-answer data samples for the appearance attributes, and question-and-answer data samples for the expression attributes; A second training unit configured to, after completing the first-stage training, perform a second-stage training on the vision language model based on the second training dataset to update the parameters of the projection layer and the large language model, where the second-stage training is used to enable the vision language model to have the function of providing targeted feedback for face attributes.

8. An inference system for a visual language model, characterized in that, The vision language model includes: a vision encoder, a projection layer, a text tokenizer, and a large language model. Among them, the inference system includes: A text acquisition unit configured to obtain the text content corresponding to the user input; An image acquisition unit configured to obtain the face image of the user captured by the image acquisition device; A visual feature acquisition unit configured to input the face image into the vision encoder, and input the vector output by the vision encoder into the projection layer to obtain the visual features output by the projection layer; A text feature acquisition unit configured to input the text content into the text tokenizer to obtain the text features output by the text tokenizer; A feedback content acquisition unit configured to input the visual features and the text features into the large language model to obtain the feedback content output by the large language model for providing feedback on the face attributes of the user input and the face image; Wherein, the vision language model is trained by executing the training method described in any one of claims 1 to 5.

9. A computer-readable storage medium for storing instructions, characterized in that, When the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the training method according to any one of claims 1 to 5 and / or the inference method according to claim 6.

10. An electronic device, characterized in that, The electronic device includes: at least one processor; at least one memory storing computer-executable instructions, wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method according to any one of claims 1 to 5 and / or the inference method according to claim 6.

11. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by at least one processor, the training method according to any one of claims 1 to 5 and / or the inference method according to claim 6 is implemented.

Citation Information

Patent Citations

  • A method and apparatus for controlling a vehicle

    CN109131167A

  • Multi-modal pre-training model training method and device, equipment and storage medium

    CN118133241A

  • Adapter-based large language model multi-modal lightweight fusion method and system

    CN118364066A

  • Visual question-answering model training method, visual question-answering method and visual question-answering system

    CN118798373A

  • Facial expression recognition method, device and equipment

    CN118918624A

Cited By

  • Multi-modal large model training method and device, equipment and storage medium

    CN120632470A

  • Post-training method and device for multi-modal model

    CN121145976A