Method for predicting target object attribute information, electronic device, and computer storage medium

By combining the acquisition of image and problem description information in a multimodal large model, and using visual encoder and text decoder to extract features, the problem of low credibility of object attribute information prediction is solved, and more accurate and reliable attribute information prediction is achieved.

CN119785042BActive Publication Date: 2025-07-18ZHEJIANG DAHUA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510263945.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-18
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

In the prior art, the reliability of object attribute information prediction is not high, resulting in a single and inaccurate attribute numerical estimate.

Method used

By obtaining the acquired image and problem description information, determining the target prompt information, and inputting it into the multimodal large model, combining the visual encoder and text decoder for feature extraction and decoding, the multimodal large model is used to predict the attribute information of the target object.

Benefits of technology

It improves the accuracy and credibility of attribute information prediction, alleviates the learning bias caused by pure attribute numerical prediction, and enhances the reliability of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785042B_ABST
    Figure CN119785042B_ABST
Patent Text Reader

Abstract

The present application discloses a method for predicting target object attribute information, an electronic device, and a computer storage medium. The method includes: obtaining a collected image and prediction instruction data; determining target prompt information of a target object in the collected image according to the collected image and the prediction instruction data; and inputting the target prompt information, the collected image, and the prediction instruction data into a multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model. Thereby, the credibility of attribute information prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and particularly relates to a method for predicting attribute information of a target object, an electronic device, and a computer storage medium. Background Art

[0002] The technology of object attribute information estimation has shown broad application prospects in various fields, including market analysis and marketing, personalized recommendation, entertainment and interaction, medical and health, academic research, intelligent transportation, financial services, and education. Currently, usually only the attribute numerical estimation of object samples is carried out, resulting in the problems of single attribute numerical estimation and low credibility.

[0003] Based on this, how to predict attribute information with high credibility has become a technical problem that urgently needs to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a method for predicting attribute information of a target object, an electronic device, and a computer-readable storage medium, which can improve the credibility of attribute information prediction.

[0005] To solve the above technical problem, a technical solution adopted by this application is: providing a method for predicting attribute information of a target object, the method for predicting attribute information of the target object includes: obtaining an acquired image and problem description information of an attribute to be predicted of a target object in the acquired image; determining target prompt information of the target object in the acquired image according to the acquired image and the problem description information; inputting the target prompt information, the acquired image, and the problem description information into a multimodal large model to obtain the attribute information of the target object predicted by the multimodal large model.

[0006] In an embodiment, the multimodal large model includes a first sub-model and a second sub-model. The step of inputting the target prompt information, the acquired image, and the problem description information into the multimodal large model to obtain the attribute information of the target object predicted by the multimodal large model includes: inputting the acquired image and the problem description information into the first sub-model to obtain word element information predicted by the first sub-model; inputting the word element information, the problem description information, and the target prompt information into the second sub-model to obtain the attribute information of the target object predicted by the second sub-model.

[0007] In one embodiment, the first sub-model includes a visual encoder and a text decoder. The step of inputting the captured image and the problem description information into the first sub-model to obtain the token information predicted by the first sub-model includes: encoding the captured image through the visual encoder to obtain a visual feature vector; performing text decoding processing on the problem description information and the visual feature vector through the text decoder to obtain a text feature vector; and determining the token information according to the visual feature vector and the text feature vector.

[0008] In one embodiment, the token information includes context tokens and content tokens. The step of determining the token information according to the visual feature vector and the text feature vector includes: performing compression processing on the text feature vector and the visual feature vector to obtain a compressed feature vector; performing conversion processing on the compressed feature vector to obtain the context tokens; and performing conversion processing on the visual feature vector to obtain the content tokens.

[0009] In one embodiment, the step of inputting the target prompt information, the captured image, and the problem description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model includes: obtaining the face pose angle information and / or the face occlusion information in the target prompt information and the age problem in the problem description information; and inputting the face pose angle information and / or the face occlusion information, the captured image, and the age problem into the multi-modal large model to obtain the age information of the target object predicted by the multi-modal large model.

[0010] In one embodiment, the problem description information includes problem text. The step of determining the target prompt information of the target object in the captured image according to the captured image and the problem description information includes: extracting features from the captured image to obtain the prior attribute information of the target object in the captured image; and determining the target prompt information according to the prior attribute information and the problem text.

[0011] In one embodiment, the step of determining the target prompt information according to the prior attribute information and the problem text includes: determining a target prompt template from a preset prompt template according to the problem text; obtaining each attribute value in the prior attribute information; and filling each attribute value into the target prompt template correspondingly to obtain the target prompt information.

[0012] In one embodiment, the method further includes: obtaining a sample image, prediction instruction data of a sample object in the sample image, and prompt information about the sample object, where the prediction instruction data includes problem description information of a to-be-predicted attribute of the sample object and answer information of the problem description information; inputting the sample image, the prediction instruction data of the sample object, and the prompt information of the sample object into a multi-modal large model to be trained, and obtaining a predicted result of attribute information output by the multi-modal large model to be trained; adjusting the multi-modal large model according to the predicted result of attribute information, the prediction instruction data of the sample object, and the prompt information of the sample object until a multi-modal large model that meets the requirements is obtained.

[0013] To solve the above technical problems, another technical solution adopted by this application is: providing an electronic device, including a memory and a processor, where the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the above-mentioned target object attribute information prediction method.

[0014] To solve the above technical problems, another technical solution adopted by this application is: providing a computer-readable storage medium including stored program data, where the program data is used to implement the above-mentioned target object attribute information prediction method when executed by a processor.

[0015] In the above solution, the target prompt information of the target object in the acquisition image is determined according to the acquired acquisition image and the attribute description information of the to-be-predicted attribute of the target object in the acquisition image; the target prompt information, the acquisition image, and the attribute description information are input into a multi-modal large model, and the attribute information of the target object predicted by the multi-modal large model is obtained. When the multi-modal large model predicts the attribute information of the target object, by combining the acquisition image, the attribute description information of the to-be-predicted attribute of the target object in the acquisition image, and the target prompt information, the accuracy of attribute information prediction can be improved while the credibility of attribute information prediction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0017] Figure 1 is a schematic flowchart of an exemplary embodiment of the target object attribute information prediction method shown in the present application;

[0018] Figure 2 is Figure 1Flow diagram of an exemplary embodiment of step S130 in the shown method for predicting target object attribute information;

[0019] Figure 3 Yes Figure 1 Flow diagram of an exemplary embodiment of step S120 in the shown method for predicting target object attribute information;

[0020] Figure 4 Structural diagram of an exemplary embodiment of the multimodal large model shown in this application;

[0021] Figure 5 Structural diagram of an exemplary embodiment of the attribute information prediction device shown in this application;

[0022] Figure 6 Structural diagram of an embodiment of the electronic device provided in this application;

[0023] Figure 7 Structural diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed implementation

[0024] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. It can be understood that the specific embodiments described here are only used to explain this application and are not intended to limit this application. Additionally, it should be noted that for the sake of description, only parts related to this application rather than all structures are shown in the accompanying drawings. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.

[0025] For details, please refer to Figure 1 , Figure 1 Flow diagram of an exemplary embodiment of the method for predicting target object attribute information shown in this application.

[0026] The execution subject of the method for predicting target object attribute information can be a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The execution subject of the method for predicting target object attribute information can also be an attribute information prediction device. In some possible implementation manners, the method for predicting target object attribute information can be implemented by a processor calling computer-readable instructions stored in a memory.

[0027] In the embodiments of the present application, the attribute information prediction device is used as the execution subject for description. Specifically, the method for predicting the target object attribute information in this embodiment includes the following steps:

[0028] S110, obtain the captured image and the problem description information of the attribute to be predicted of the target object in the captured image.

[0029] The target object can be an animal, a plant, a person, etc.

[0030] The attribute to be predicted can be age, gender, posture, etc.

[0031] The problem description information can be determined by the attribute information prediction device in response to the user's input. As an example, the attribute information prediction device receives the voice information input by the user, and the attribute information prediction device analyzes the voice information to obtain the problem description information. As another example, the attribute information prediction device receives the text information input by the user and uses the text information as the problem description information. It should be noted that the problem description information is also the user's question about the attribute of the target object in the captured image or the user's problem about the attribute of the target object in the captured image. For example, if the attribute to be predicted is posture, the problem description information can be what is the posture of the target object in the captured image? If the attribute to be predicted is age, the problem description information can be how old is the target object in the captured image? If the attribute to be predicted is gender, the problem description information can be what is the gender of the target object in the captured image?

[0032] S120, determine the target hint information of the target object in the captured image according to the captured image and the problem description information.

[0033] The target hint information is used to guide the multi-modal large model to use or partially use the target hint information during prediction to improve the prediction accuracy and credibility.

[0034] As an example, the attribute information prediction device can input the captured image and the problem description information into the trained hint model to obtain the target hint information of the target object in the captured image output by the hint model. For example, if the attribute to be predicted is age, the problem description information can be age-related information, specifically how old is the target object in the captured image? The target hint information can be age-related information, specifically the facial posture angle information or the facial occlusion information of the target object. It should be noted that the facial posture angle information can be the specific degree of the facial upward viewing posture angle. Another example, if the attribute to be predicted is gender, the problem description information can be gender-related information, specifically what is the gender of the target object in the captured image? The target hint information can be gender-related information, specifically the body feature information such as the hair length or Adam's apple of the target object.

[0035] S130. Input the target prompt information, the captured image, and the problem description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model.

[0036] The attribute information may be the gender, age, possible age range, pose information, information prediction credibility, etc. of the target object. It should be noted that the attribute information includes not only the attribute information of the attribute to be predicted, but also other attribute information related to the attribute to be predicted. Exemplarily, if the attribute to be predicted is age, the predicted attribute information may be the possible range of the target object's age, the impossible range of the age, the degree of partial occlusion of the target object, the pose angle of the target object, etc.

[0037] The attribute information prediction device inputs the target prompt information, the captured image, and the problem description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model. As an example, the attribute information prediction device obtains the face pose angle information and / or face occlusion information in the target prompt information and the age problem in the problem description information, and inputs the face pose angle information and / or face occlusion information, the captured image, and the age problem into the multi-modal large model to obtain the age information of the target object predicted by the multi-modal large model.

[0038] It can be seen that the target object attribute information prediction method of the embodiment of the present application determines the target prompt information of the target object in the captured image according to the captured image and the attribute description information of the attribute to be predicted of the target object in the captured image; inputs the target prompt information, the captured image, and the attribute description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model. When the multi-modal large model predicts the attribute information of the target object, by combining the captured image, the attribute description information of the attribute to be predicted of the target object in the captured image, and the target prompt information, it can improve the accuracy of attribute information prediction while improving the credibility of attribute information prediction.

[0039] Based on the above embodiments, the embodiments of the present application use Figure 2 a flowchart to elaborate in detail how the multi-modal large model predicts the attribute information of the target object in the captured image according to the target prompt information, the captured image, and the problem description information. Please refer to Figure 2 , Figure 2 which Figure 1 is a schematic flowchart of an exemplary embodiment of step S130 in the target object attribute information prediction method shown. Specifically, the process of step S130 inputting the target prompt information, the captured image, and the problem description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model includes the following steps:

[0040] S210. Input the acquired image and the problem description information into the first sub-model to obtain the token information predicted by the first sub-model.

[0041] The multi-modal large model includes a first sub-model and a second sub-model. Among them, the first sub-model is used to extract token information from the input acquired image and problem description information.

[0042] The attribute information prediction device processes the acquired image and the problem description information through the first sub-model to obtain the token information predicted by the first sub-model. Exemplarily, the first sub-model includes a visual encoder and a text decoder. The attribute information prediction device encodes the acquired image through the visual encoder to obtain a visual feature vector; decodes the problem description information and the visual feature vector through the text decoder to obtain a text feature vector; and then determines the token information according to the visual feature vector and the text feature vector. It should be noted that the visual feature vector can be represented as X, , N is the number of image patches in the acquired image, C is the number of feature channels, and the text feature vector can be represented as Q, , M represents the number of text feature vectors. Further, it should be noted that the text decoder is mainly used for cross-modal interaction between images and texts, and the text decoder in the embodiments of the present application can be a Q-Former structure.

[0043] For the method of encoding the acquired image through the visual encoder to obtain a visual feature vector, the attribute information prediction device performs feature extraction processing on the region where the target object is located in the acquired image through the visual encoder to obtain a visual feature vector. Among them, the attribute information prediction device can use the YOLOv5 model to obtain the region where the target object is located in the acquired image. It should be noted that the acquired image can be represented as , H represents the height of the acquired image, W represents the width of the acquired image, and the region where the target object is located in the acquired image can be represented as , and the visual feature vector can be represented as , C represents the number of feature channels. Among them, , N represents the number of image patches, and P represents the size of the image patch.

[0044] For the method of determining token information based on visual feature vectors and text feature vectors, as an example, the attribute information prediction device can perform fusion processing on the visual feature vectors and text feature vectors to obtain token information. As another example, the attribute information prediction device can construct a context attention module to process the visual feature vectors and text feature vectors through the context attention module to aggregate visual feature vectors related to the text and obtain token information. Specifically, the attribute information prediction device compresses the text feature vectors and visual feature vectors to obtain compressed feature vectors; performs transformation processing on the compressed feature vectors to obtain context tokens; and performs transformation processing on the visual feature vectors to obtain content tokens. It should be noted that for the method of performing transformation processing on the visual feature vectors to obtain content tokens, the attribute information prediction device can perform transformation processing on the visual feature vectors through adaptive pooling and a multi-layer perceptron (MLP) projector to obtain content tokens. For the method of performing transformation processing on the compressed feature vectors to obtain context tokens, the attribute information prediction device can perform transformation processing on the compressed feature vectors through an MLP projector to obtain context tokens, so as to achieve alignment between the content tokens and the context tokens.

[0045] Specifically, the compressed feature vectors satisfy the following formula:

[0046]

[0047] where E represents the compressed feature vectors, Mean() represents the mean function, Soft max() represents the normalization exponential function, Q represents the text feature vectors, represents the transpose of the visual feature vectors, and X represents the visual feature vectors.

[0048] S220, input the token information, problem description information, and target prompt information into the second sub-model to obtain the attribute information of the target object predicted by the second sub-model.

[0049] The second sub-model can be an LLM (Large Language Model) model.

[0050] It can be seen that the target object attribute information prediction method according to the embodiments of the present application inputs the collected image and the problem description information into the first sub-model to obtain the token information predicted by the first sub-model; inputs the token information, the problem description information, and the target prompt information into the second sub-model to obtain the attribute information of the target object predicted by the second sub-model. Thus, predicting the attribute information based on the problem description information and the target prompt information can alleviate the learning bias caused by network prediction with simple attribute values and improve the credibility of the attribute information prediction compared to only predicting the attribute values.

[0051] Based on the above embodiments, the embodiments of the present application use Figure 3 The flowchart details how to determine the target prompt information. Please refer to Figure 3 , Figure 3 is Figure 1 A schematic flowchart of an exemplary embodiment of step S120 in the target object attribute information prediction method shown. Specifically, the process of step S120 for determining the target prompt information of the target object in the collected image according to the collected image and the problem description information includes the following steps:

[0052] S310, perform feature extraction on the collected image to obtain the prior attribute information of the target object in the collected image.

[0053] The prior attribute information GPO (Gender, Pose, Occlusion) includes the gender, head pose, occlusion state, etc. of the target object in the collected image.

[0054] As an example, the attribute information prediction device can input the collected image into the FaceXFormer expert model to obtain the prior attribute information of the target object in the collected image. As another example, the attribute information prediction device can compare the features in the collected image with the preset features to determine the prior attribute information. Specifically, if the similarity between the features in the collected image and the preset features is greater than the preset similarity threshold, the attribute information prediction device determines the prior attribute information of the target object according to the attribute corresponding to the preset features and the feature values in the collected image.

[0055] S320, determine the target prompt information according to the prior attribute information and the problem text.

[0056] As an example, the attribute information prediction device determines a target prompt template from a preset prompt template according to the question text; obtains each attribute value in the attribute prior information; fills in each attribute value into the target prompt template correspondingly to obtain the target prompt information. As another example, the attribute information prediction device determines the target prompt information according to the matching degree between the question text and the attribute prior information. For example, the question text is "What is the age of the target object in the captured image?" The attribute prior information includes an estimated age of 15 years, an estimated height of 150, and an estimated gender of male. Then, it can be determined that the attribute prior information of the estimated age of 15 years matches the question text, and the corresponding estimated age of 15 years is determined as the target prompt information.

[0057] It should be noted that after the attribute information prediction device determines the target prompt information of the target object in the captured image, it can determine the credibility of the estimated attribute information according to the target prompt information. Exemplarily, the attribute information prediction device can compare the target prompt information with the preset information to determine the credibility of the estimated attribute information. For example, the target prompt information is that the pose angle of the target object is 60 degrees and the occlusion of the target part exceeds two-thirds. Then, it is judged whether the pose angle of the target object is greater than 45 degrees and whether the occlusion of the target part exceeds one-third. If the pose angle is greater than 45 degrees and the occlusion state of the target part exceeds one-third, it is determined that the prediction credibility of the corresponding attribute information is not high and there is a risk of inaccuracy; if the pose angle is less than or equal to 45 degrees and the occlusion state of the target part exceeds one-third, it is determined that the prediction credibility of the corresponding attribute information is relatively high; if the pose angle is greater than 45 degrees and the occlusion state of the target part does not exceed one-third, it is determined that the prediction credibility of the corresponding attribute information is relatively high; if the pose angle is less than or equal to 45 degrees and the occlusion state of the target part does not exceed one-third, it is determined that the prediction credibility of the corresponding attribute information is high.

[0058] It can be seen that the method for predicting the attribute information of the target object in the embodiment of the present application extracts features from the captured image to obtain the attribute prior information of the target object in the captured image; determines the target prompt information according to the attribute prior information and the question text. It can determine the target prompt information based on the attribute prior information, and then guide the model to selectively use the target prompt information based on the determined target prompt information, improving the accuracy and prediction credibility of predicting attribute information using the multi-modal large model.

[0059] Based on the above embodiments, in order to obtain a multi-modal large model with accurate prediction, the attribute information prediction device trains the multi-modal large model before predicting the attribute information of the target object. Specifically, the attribute information prediction device obtains a sample image, prediction instruction data of the sample object in the sample image, and prompt information about the sample object. The prediction instruction data includes problem description information of the attribute to be predicted of the sample object and answer information of the problem description information; inputs the sample image, the prediction instruction data of the sample object, and the prompt information of the sample object into the multi-modal large model to be trained, and obtains the attribute information prediction result output by the multi-modal large model to be trained; adjusts the multi-modal large model according to the attribute information prediction result, the prediction instruction data of the sample object, and the prompt information of the sample object until a multi-modal large model that meets the requirements is obtained. It should be noted that in the embodiments of the present application, LoRA (Low-Rank Adaptation) can be used for fine-tuning the model training. Specifically, the attribute information prediction device obtains the information difference between the attribute information prediction result and the answer information of the sample object and the prompt information of the sample object; adjusts the weight parameters in the first sub-model according to the information difference until the attribute information prediction result output by the second sub-model meets the requirements, and a multi-modal large model that meets the requirements is obtained.

[0060] For the acquisition method of the prediction instruction data of the sample object in the sample image, the attribute information prediction device can use the Qwen2-VL multi-modal model to generate prediction instruction data related to the attribute, that is, question-answer dialogue instruction data. Specifically, the attribute information prediction device inputs the collected image and the problem description information about the target object in the collected image into the Qwen2-VL multi-modal model, and obtains the answer information output by the Qwen2-VL multi-modal model. For example, the problem description information is how old is the target object in the collected image, and the answer information is the predicted age and the possible age range and impossible age range of the target object, so that the prediction of the attribute information can be extended by using the Qwen2-VL multi-modal model.

[0061] For the acquisition method of the prompt information of the sample object, the attribute information prediction device can perform detection processing on the sample image through YOLOv5 to obtain the image area where the target object is located; then input the image area where the target object is located into the FaceXFormer expert model to obtain prior attribute information; then use the problem description information and the prior attribute information to guide the multi-modal large model to selectively use the prior attribute information during training to obtain the prompt information of the sample object.

[0062] To elaborate on the above embodiments in detail, take the Figure 4 example in Figure 4 as an illustration. Please continue to refer to Figure 4The framework example diagram of the multi-modal large model is shown. As shown in the figure, the attribute information prediction device inputs the collected image and the problem description information about the target object in the collected image into the multi-modal large model. The collected image is encoded by the frozen visual encoder in the multi-modal large model to obtain the visual feature vector X. The problem description information "Question: What is the age of the face in the picture?" and the visual feature vector X are text-decoded by the text decoder in the multi-modal large model to obtain the text feature vector Q, that is, the text query. The visual feature vector X and the text feature vector Q are compressed by the context attention module in the multi-modal large model to obtain the compressed feature vector E. The compressed feature vector E is processed by the projector to obtain the context token, that is, the context word element. The visual feature vector X is pooled to obtain the pooled visual feature vector, and the pooled visual feature vector X is processed by the projector to obtain the content token, that is, the content word element. Then, the content token, the context token, the problem description information "Question: What is the age of the face in the picture?" and the target prompt information determined based on the problem description information, that is, the manual prompt words in the figure, are input into the LLM model fine-tuned based on LoRA to obtain the attribute information about the target object. For example, "answer: The age in the picture is 40-50 years old, the gender is male, there is no occlusion, there is no pose angle exceeding the threshold, and the confidence of the age prediction value is high" in the figure.

[0063] Please refer to Figure 5 , Figure 5 is the structural schematic diagram of an exemplary embodiment of the attribute information prediction device shown in the present application. The attribute information prediction device 500 includes an acquisition module 510, a determination module 520, and a prediction module 530. The acquisition module 510 is used to acquire the collected image and the problem description information of the attribute to be predicted of the target object in the collected image. The determination module 520 is used to determine the target prompt information of the target object in the collected image according to the collected image and the problem description information. The prediction module 530 is used to input the target prompt information, the collected image, and the problem description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model.

[0064] In the above solution, the attribute information prediction device determines the target hint information of the target object in the acquired image based on the acquired image and the attribute description information of the attribute to be predicted of the target object in the acquired image; inputs the target hint information, the acquired image, and the attribute description information into the multi-modal large model to obtain the attribute information of the target object predicted by the multi-modal large model. When the multi-modal large model predicts the attribute information of the target object, by combining the acquired image, the attribute description information of the attribute to be predicted of the target object in the acquired image, and the target hint information, it can improve the accuracy of attribute information prediction while improving the credibility of attribute information prediction.

[0065] Among them, the functions of each module can be referred to in the embodiments of the target object attribute information prediction method, which will not be elaborated here.

[0066] To implement the target object attribute information prediction method in the above embodiments, the present application proposes another electronic device. For details, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an embodiment of the electronic device provided by the present application.

[0067] The electronic device 600 includes a memory 601 and a processor 602, where the memory 601 and the processor 602 are coupled.

[0068] The memory 601 is used to store program data, and the processor 602 is used to execute the program data to implement the target object attribute information prediction method in the above embodiments.

[0069] In this embodiment, the processor 602 can also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 602 may be an integrated circuit chip with signal processing capabilities. The processor 602 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 602 may also be any conventional processor, etc.

[0070] The present application also provides a computer-readable storage medium, as Figure 7 shown, the computer-readable storage medium 700 is used to store program data 701, and when the program data 701 is executed by the processor, it is used to implement the target object attribute information prediction method in the method embodiments of the present application.

[0071] In the embodiments of the method for predicting the attribute information of the target object of the present application, when the method involved exists in the form of a software functional unit and is sold or used as an independent product, it can be stored in a device, for example, a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0072] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for predicting target object attribute information, characterized in that The method for predicting the target object attribute information includes: Obtain the collected image and the problem description information of the attribute to be predicted of the target object in the collected image; Determine the target prompt information of the target object in the collected image according to the collected image and the problem description information; Input the target prompt information, the collected image, and the problem description information into a multimodal large model to obtain the attribute information of the target object predicted by the multimodal large model; The problem description information includes problem text. The step of determining the target prompt information of the target object in the collected image according to the collected image and the problem description information includes: Extract features from the collected image to obtain the prior attribute information of the target object in the collected image, and the prior attribute information includes at least one of the gender, head pose, or occlusion state of the target object; Determine the target prompt information according to the prior attribute information and the problem text; The step of determining the target prompt information according to the prior attribute information and the problem text includes: Determine the target prompt template from a preset prompt template according to the problem text; Obtain the attribute values in the prior attribute information; Fill in the target prompt template with the corresponding attribute values to obtain the target prompt information.

2. The method for predicting target object attribute information according to claim 1, wherein The multimodal large model includes a first sub-model and a second sub-model. The step of inputting the target prompt information, the collected image, and the problem description information into the multimodal large model to obtain the attribute information of the target object predicted by the multimodal large model includes: Input the collected image and the problem description information into the first sub-model to obtain the token information predicted by the first sub-model; Input the token information, the problem description information, and the target prompt information into the second sub-model to obtain the attribute information of the target object predicted by the second sub-model.

3. The method for predicting target object attribute information according to claim 2, wherein The first sub-model includes a visual encoder and a text decoder. The step of inputting the collected image and the problem description information into the first sub-model to obtain the token information predicted by the first sub-model includes: Encode the collected image through the visual encoder to obtain a visual feature vector; Perform text decoding processing on the problem description information and the visual feature vector through the text decoder to obtain a text feature vector; Determine the token information according to the visual feature vector and the text feature vector.

4. The method for predicting target object attribute information according to claim 3, characterized in that The token information includes context tokens and content tokens. The step of determining the token information according to the visual feature vector and the text feature vector includes: Perform compression processing on the text feature vector and the visual feature vector to obtain a compressed feature vector; Perform conversion processing on the compressed feature vector to obtain the context tokens; Perform conversion processing on the visual feature vector to obtain the content tokens.

5. The method for predicting target object attribute information according to claim 1, wherein The step of inputting the target prompt information, the captured image, and the problem description information into the multimodal large model to obtain the attribute information of the target object predicted by the multimodal large model includes: Obtaining the facial pose angle information and / or facial occlusion information in the target prompt information and the age problem in the problem description information; Inputting the facial pose angle information and / or the facial occlusion information, the captured image, and the age problem into the multimodal large model to obtain the age information of the target object predicted by the multimodal large model.

6. The method for predicting target object attribute information according to claim 1, characterized in that, The method further includes: Obtaining a sample image, the prediction instruction data of the sample object in the sample image, and the prompt information about the sample object, where the prediction instruction data includes the problem description information of the attribute to be predicted of the sample object and the answer information of the problem description information; Inputting the sample image, the prediction instruction data of the sample object, and the prompt information of the sample object into the multimodal large model to be trained to obtain the attribute information prediction result output by the multimodal large model to be trained; Adjusting the multimodal large model according to the attribute information prediction result, the prediction instruction data of the sample object, and the prompt information of the sample object until a multimodal large model that meets the requirements is obtained.

7. An electronic device, characterized in that, Including: A memory and a processor, where the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method according to any one of claims 1-6.

8. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by the processor, it is used to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN115393849A

  • Bidding document information extraction method and device, equipment and storage medium

    CN116205212A

  • Question and answer processing method and device, electronic equipment and storage medium

    CN116881427A