Data classification method and device, electronic equipment and storage medium

Through a multimodal big model, the image, question text and answer text are comprehensively evaluated and classified, and the accuracy of triple data evaluation and classification in the prior art is solved, and the data quality and model iteration effect are improved.

CN120162647APending Publication Date: 2025-06-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510259516.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively evaluate and classify triple data containing images, questions and answers, resulting in poor accuracy of data quality and security evaluation in large model training, affecting the model iteration effect.

Method used

The multimodal big model is used to evaluate the image, question text and answer text as a whole, and the image-question text and image-response text are evaluated, and the to be processed data is classified through the evaluation result set.

Benefits of technology

It improves the evaluation accuracy of triple data containing visual question and answer content, enhances the accuracy of data classification, improves the data quality of training samples used to train multimodal large models, and thus improves the model iteration effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162647A_ABST
    Figure CN120162647A_ABST
Patent Text Reader

Abstract

The invention provides a data classification method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the fields of multi-modal large models, large language models, generative models, computer vision and the like. The specific implementation scheme is as follows: acquiring to-be-processed data, wherein the to-be-processed data comprises an image, a question text and an answer text; based on a multi-modal large model, according to the question text, the answer text and the image, evaluating the security and the quality of the to-be-processed data to obtain a first evaluation result; based on a multi-modal large model, according to one of the question text and the answer text and the image, evaluating the quality of the to-be-processed data to obtain a second evaluation result; and according to an evaluation result set, classifying the to-be-processed data, the evaluation result set comprising a first evaluation result and a second evaluation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to fields such as multimodal large models, large language models, generative models, computer vision, etc. More specifically, the present disclosure provides a data classification method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] In related fields such as artificial intelligence and big data processing, sometimes triple data including images, questions, and answers is used to train large models. Summary of the Invention

[0003] The present disclosure provides a data classification method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to one aspect of the present disclosure, there is provided a data classification method, including: obtaining data to be processed, where the data to be processed includes an image, question text, and answer text; based on a multimodal large model, evaluating the security and quality of the data to be processed according to the question text, answer text, and image to obtain a first evaluation result; based on the multimodal large model, evaluating the quality of the data to be processed according to one of the question text and answer text, and the image to obtain a second evaluation result; and classifying the data to be processed according to the evaluation result set; the evaluation result set includes the first evaluation result and the second evaluation result.

[0005] According to another aspect of the present disclosure, there is provided a data classification apparatus, including: an obtaining module, a first evaluation module, a second evaluation module, and a classification module. The obtaining module is used to obtain data to be processed, where the data to be processed includes an image, question text, and answer text. The first evaluation module is used to evaluate the security and quality of the data to be processed according to the question text, answer text, and image based on a multimodal large model to obtain a first evaluation result. The second evaluation module is used to evaluate the quality of the data to be processed according to one of the question text and answer text, and the image based on the multimodal large model to obtain a second evaluation result. The classification module is used to classify the data to be processed according to the evaluation result set; the evaluation result set includes the first evaluation result and the second evaluation result.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the method provided by the present disclosure when executed by a processor.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic diagram of the application scenario of the data classification method and device according to the embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of the data classification method according to the embodiment of the present disclosure;

[0013] Figure 3 is a schematic principle diagram of the data classification method according to the embodiment of the present disclosure;

[0014] Figure 4 is a schematic structural block diagram of the data classification device according to the embodiment of the present disclosure; and

[0015] Figure 5 is a structural block diagram of an electronic device for implementing the data classification method according to the embodiment of the present disclosure. Detailed Embodiments

[0016] The following makes an explanation of the exemplary embodiments of the present disclosure in conjunction with the drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0017] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0018] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.

[0019] In some application scenarios, training large models requires using triple data including images, questions, and answers. The triple data can be evaluated, and then based on the evaluation results, the triple data can be divided into different training sets. Subsequently, according to actual needs, the triple data in each training set can be used to iterate the large model. For example, the triple data can be divided into multiple training sets such as a risk data set and a high-quality data set. The triple data in the risk data set can be used as negative samples, and the triple data in the high-quality data set can be used as positive samples or test samples.

[0020] In the process of evaluating triple data and dividing training sets, if rough strategies such as image resolution and text length are used to filter images and text, the classification quality is poor. If a large language model is used to evaluate question-and-answer pairs, due to the lack of visual information, the image cannot be effectively combined with the question and answer. This results in the inability to effectively focus on the image content, accurately understand the intent of the image and the question, distinguish incorrect image descriptions, and thus makes it difficult to accurately evaluate the true situation of this type of triple data containing visual question-and-answer content. And the inaccurate evaluation results will directly affect the classification results of the triple data, and further affect the iteration effect of the large model.

[0021] This embodiment aims to provide a data classification method. This method comprehensively evaluates images, question texts, and answer texts based on a multimodal model, and also evaluates image-question texts or image-answer texts. In this way, the evaluation is more comprehensive, realizing the effective integration of images, question texts, and answer texts, improving the evaluation accuracy of triple data containing visual question-and-answer content, thereby improving the classification accuracy, enhancing the data quality of training samples for training multimodal large models, and improving the model iteration effect.

[0022] The data classification method provided in this embodiment can be used to generate training samples for each stage of multimodal large models, providing high-quality and highly reliable training samples for multimodal large models.

[0023] The technical solutions provided by the present disclosure will be elaborated in detail below in conjunction with the accompanying drawings and specific embodiments.

[0024] Figure 1 is a schematic diagram of the application scenario of the data classification method and device according to an embodiment of the present disclosure.

[0025] It should be noted that Figure 1 The example shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.

[0026] Such as Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0027] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0028] The server 105 may be a server that provides various services, such as a background management server (only as an example) that supports the websites browsed by users using the terminal devices 101, 102, 103. The background management server can analyze and process data such as received user requests.

[0029] For example, users can input images and question texts through the terminal devices 101, 102, 103, use a large model to process the images and question texts to generate answer texts, and can construct the images, question texts, and answer texts into data to be processed. It should be noted that for the acquisition and use of the information input by users, users are aware and consent, obtain the authorization of users, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. Then, the server 105 can obtain the data to be processed through the terminal devices 101, 102, 103, and then evaluate and classify the data to be processed to obtain an evaluation result and a classification result.

[0030] The system architecture 100 may further include a database 106, and the data to be processed, evaluation results, classification results, etc. can be stored in the database 106. Then, data can be batch obtained from the database 106 to train a multi-modal large model or a large language model.

[0031] It should be noted that the data classification method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the data classification device provided by the embodiments of the present disclosure can generally be set in the server 105. The data classification method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the data classification device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0032] It should be understood that Figure 1 the number of terminal devices, networks, and servers in [[ ]] is merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0033] Figure 2 is a schematic flowchart of the data classification method according to an embodiment of the present disclosure.

[0034] As Figure 2 shown, the data classification method 200 may include operations S210 to S240.

[0035] In operation S210, the data to be processed is obtained, and the data to be processed includes images, question texts, and answer texts.

[0036] For example, the images, question texts, and answer texts may be associated. For example, the question text is based on the image for questioning, such as asking "What objects are there in the image". For example, the answer text may be based on the question and the image for answering, such as answering "There is a dog in the image".

[0037] In practical applications, the images and question texts may be input by the user through the front end, and the answer texts may be generated by the large model based on the images and question texts. However, due to the influence of the accuracy of the large model's generation results or the quality of the user input, the correlation between the images, question texts, and answer texts may be low. It should be noted that for the acquisition and use of the images and question texts input by the user, the user is aware of and agrees, obtains the authorization of the user, and all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0038] In operation S220, based on the multimodal large model, the security and quality of the data to be processed are evaluated according to the question text, answer text, and image, and a first evaluation result is obtained.

[0039] For example, the question text, answer text, and image can be input into the multimodal large model for evaluation by the multimodal large model. The first evaluation result can be used as an evaluation result and added to the evaluation result set. The first evaluation result may include evaluation values. For example, the multimodal large model can output evaluation values for security and evaluation values for quality. In addition, for any one of security and quality, the large model can output multiple evaluation values according to different evaluation dimensions. For example, for the evaluation value for quality, the large model can output evaluation values for different dimensions such as coherence and relevance.

[0040] In operation S230, based on the multimodal large model, the quality of the data to be processed is evaluated according to one of the question text and the answer text, and the image, and a second evaluation result is obtained.

[0041] For example, the image and the question text can be input into the multi-modal large model, or the image and the answer text can be input into the multi-modal large model for evaluation by the multi-modal large model. The second evaluation result can be used as another evaluation result and added to the evaluation result set. The second evaluation result can include an evaluation value. For example, the multi-modal large model can input the evaluation value of the image-text relevance.

[0042] In operation S240, the data to be processed is classified according to the evaluation result set.

[0043] For example, weighted operations or other operation processes can be performed on the evaluation values output by the multi-modal large model to obtain the total evaluation value. The data to be processed with a higher total evaluation value can be used as high-quality data, and the data to be processed with a lower total evaluation value can be used as low-quality data.

[0044] The data classification method provided in this embodiment comprehensively evaluates the image, question text, and answer text based on the multi-modal model, and also evaluates the image and question text, or the image and answer text. Such evaluation is more comprehensive, realizing the effective integration of the image, question text, and answer text, improving the accuracy of visual question-answering data evaluation, further improving the accuracy of classification, improving the data quality of the training samples used to train the multi-modal large model, and improving the model iteration effect.

[0045] According to another embodiment of the present disclosure, before obtaining the data to be processed, multiple candidate data can be obtained first. The candidate data includes candidate images, candidate questions, and candidate answers. Then, some candidate data is selected from the multiple candidate data as the data to be processed. For example, for each candidate data, the reference information of the candidate data can be determined according to the image information of the candidate image and the question information of the candidate question. Then, according to the respective reference information of the multiple candidate data, the multiple candidate data is de-duplicated. Then, the de-duplicated candidate data is determined as the data to be processed. This embodiment de-duplicates the data to be processed with the same question and the same image, which can avoid repeated processing of the same or similar data to be processed, reduce the data processing volume, improve the processing efficiency, and reduce the processing cost. And because there are cases of one image with multiple questions or multiple images corresponding to the same question, single-modal de-duplication is not adopted, which can avoid incorrect de-duplication of different data to be processed.

[0046] In one example, the image information of the candidate image is the image hash value, and the question information of the candidate question is the question hash value. In the process of determining the reference information, the image hash value and the question hash value can be concatenated to obtain the reference information. In other examples, the image hash value, question hash value, and answer hash value can also be concatenated to obtain the reference information. Then, it can be compared whether the reference information of two candidate data is consistent. If they are consistent, one candidate data is deleted for de-duplication.

[0047] In another example, the image information of the candidate image is the image features extracted by the feature extraction network, and the question information of the candidate question is the question features extracted by the feature extraction network. Then, the image features and the question features can be fused. The fusion method can be feature splicing, and the fused features can be used as reference information. Then, the similarity between the reference information of the two candidate data can be calculated. If the similarity is high, one of the candidate data can be deleted to remove duplicates.

[0048] According to another embodiment of the present disclosure, the data to be processed includes an image, a question text, and an answer text. Any one of the image, the question text, and the answer text can be used as the target sub-data, and then the target sub-data can be detected in a single modality and then classified.

[0049] In one example, the security of the target sub-data in the data to be processed can be detected to obtain a security parameter. If the security parameter meets the security exception condition, the category of the data to be processed is determined to be risk data.

[0050] Taking the target sub-data as an image as an example, it can be detected by a classification model whether the image contains content such as inappropriate content or sensitive content. The security parameter can be the category output by the classification model. The security exception condition can be that the category output by the classification model is a predetermined category. The processing method for the target sub-data being the question text or the answer text is similar to that for the target sub-data being an image. Both can output a category through a classification model, and then based on whether the category output by the classification model is a predetermined category, it is determined whether the target sub-data meets the security exception condition. If the security exception condition is met, the category of the data to be processed can be determined to be risk data.

[0051] After determining the category of the data to be processed as risk data, there is no need to use a multi-modal large model to evaluate the data to be processed. Through this pre-processing process, both the classification effect can be achieved and the number of calls to the multi-modal large model can be reduced, thereby reducing the data processing cost.

[0052] In another example, the quality of the target sub-data in the data to be processed can be detected to obtain a quality parameter. If the quality parameter meets the quality exception condition, the category of the data to be processed is determined to be invalid data.

[0053] Taking the target sub - data as an image as an example, it is possible to detect whether the image is blurred through a classification model, or use a predetermined algorithm to detect the resolution, size, etc. of the image. Correspondingly, the quality parameters can include clarity, resolution, size, etc. The quality exception conditions can include at least one of the following: the clarity is blurred, the resolution is less than the resolution threshold, and the size exceeds the predetermined size range. If the image meets the quality exception conditions, it indicates that the image quality is too low. At this time, the category of the data to be processed can be determined as invalid data.

[0054] Taking the target sub - data as a question text or an answer text as an example, it is possible to use a classification model to detect whether the text contains valid content. For example, if the text is "The server is busy. Please try again later.", then this text lacks valid content. Correspondingly, the quality parameter is the output category indicating whether the text contains valid content, and the quality exception conditions can include the output category being a predetermined category. Or, it is possible to detect the number of characters in the text. Correspondingly, the quality parameter includes the number of characters, and the quality exception conditions can include the number of characters being less than a predetermined number.

[0055] After determining that the category of the data to be processed is invalid data, it is not necessary to train the model using the invalid data, nor is it necessary to evaluate the invalid data using the multi - modal large model. Through this pre - processing process, both the classification effect can be achieved and the number of calls to the multi - modal large model can be reduced, thereby reducing the data processing cost.

[0056] According to another embodiment of the present disclosure, the data to be processed includes images, question texts, and answer texts. Next, the evaluation process of the data to be processed will be described.

[0057] In one example, it is possible to evaluate the target sub - data of a single modality in the data to be processed, and the target sub - data can be any one of an image, a question text, and an answer text.

[0058] Taking the target sub - data as an image as an example, it is possible to evaluate the image based on a large model for processing images. The evaluation dimensions can include at least one of image security and image quality, and the evaluation value can be output through the large model. The evaluation value can include the evaluation value of security or the evaluation value of image quality.

[0059] Taking the target sub - data as a text as an example, the text can be a question text or an answer text. It is possible to evaluate the text based on a large language model. The large model can output at least one of the security evaluation value and the quality evaluation value of the text. The security evaluation value can characterize the degree of discriminatory, sensitive, and other bad content contained in the text, and the quality evaluation value can characterize the quality of the valid information contained in the text. The evaluation value can be output through the large language model.

[0060] In another example, the target data pairs in the data to be processed can be evaluated. The data pairs in the data to be processed can include image-question text data pairs, image-answer text data pairs, and question text-answer text data pairs. The target data pair can be any of the above data pairs.

[0061] Taking the target data pair as an image-question text data pair as an example, based on a multimodal large model, according to the second prompt information, the image, and the question text, the second evaluation result in the evaluation result set can be determined. The second prompt information is used to guide the multimodal large model to perform quality evaluation from the dimension of image-text relevance. The second evaluation result can include the evaluation value of image-text relevance, and this evaluation value reflects the quality of the data to be processed. The second prompt information is pre-configured, and this embodiment does not limit the second prompt information.

[0062] Taking the target data pair as an image-answer text data pair as an example, based on a multimodal large model, according to the third prompt information, the image, and the answer text, the second evaluation result in the evaluation result set can be determined; among them, the third prompt information is used to guide the multimodal large model to perform quality evaluation from the dimension of image description correctness. The second evaluation result can include the evaluation value of image description correctness, and this evaluation value reflects the quality of the data to be processed. The third prompt information is pre-configured, and this embodiment does not limit the third prompt information.

[0063] Taking the target data pair as a question text-answer text data pair as an example, the question text-answer text data pair can be input into a large language model, and the large language model can be guided by the fourth prompt information to perform evaluation from the security dimension, so as to output some evaluation results, such as the evaluation value of the security of the question content and the evaluation value of the security of the answer content. The large language model can also be guided by the fourth prompt information to perform evaluation from the quality dimension, so as to output some other evaluation results, such as the evaluation value of the semantic alignment of the question and answer and the evaluation value of the knowledge correctness of the answer.

[0064] In another example, the data to be processed containing triple data can be evaluated. The triple data refers to the triple composed of an image, question text, and answer text.

[0065] For example, based on the multimodal large model, the first evaluation result in the evaluation result set can be determined according to the first prompt information, the image, the question text and the answer text, wherein the first prompt information is used to guide the multimodal large model to evaluate the answer content from the security dimension, for example, the multimodal large model can output a security evaluation value. In addition, the first prompt information can also guide the multimodal large model to evaluate from the quality dimension, for example, output a comprehensive quality evaluation value, or output multiple separate quality evaluation values ​​such as a semantic coherence evaluation value, a paragraph coherence evaluation value, an intention satisfaction evaluation value, and an answer correctness evaluation value. The above-mentioned intention satisfaction is determined based on the intention asked in the question text and the intention answered in the answer text. For example, the question text asks 3 intentions, the answer text answers 6 intentions, and 1 of the intentions is consistent with the intention asked in the text. Figure 1 If the intention is not consistent, the accuracy and comprehensiveness of the intention are poor, so the intention satisfaction is poor. For another example, the question text asks for 3 intentions, and the answer text answers 3 intentions that are consistent with the 3 intentions asked in the text. Figure 1 If the intention is consistent, the accuracy and comprehensiveness of the intention are high, so the intention satisfaction is high.

[0066] It should be noted that the data to be processed includes images, question texts and answer texts. The above introduces a variety of evaluation methods, such as evaluating the target sub-data of any single modality, evaluating any target data pair, and evaluating triple data. In the actual evaluation process, the evaluation process of the data to be processed may include at least one of the above-mentioned multiple evaluation processes.

[0067] It should be noted that in the process of using large models such as multimodal large models and large language models to process the data to be processed and thus obtain evaluation values, the large model used can also be used to generate a reason text, which represents the reason why the data to be processed obtained the evaluation results.

[0068] It should be noted that after evaluating the data to be processed, an evaluation result set can be obtained, and the evaluation result set can include at least one evaluation value, and some or all of the evaluation values ​​can be selected as the evaluation result set from the evaluation values ​​obtained by the above-mentioned various evaluation methods. For example, some or all of the evaluation values ​​can be selected as the evaluation result set from the evaluation value of any target sub-data in the security dimension, the evaluation value of the target sub-data in the quality dimension, the evaluation value of any target data pair in the quality dimension, the evaluation value of any target data pair in the degree of security, the evaluation value of the triple data in the security dimension, and the evaluation value of the triple data in the quality dimension.

[0069] According to another embodiment of the present disclosure, a process of classifying data to be processed according to an evaluation result set is described.

[0070] It should be noted that in this embodiment, the evaluation result set may include a first evaluation result and a second evaluation result. The first evaluation result and the second evaluation result are evaluation results of different categories. The difference between the two lies in the evaluation process. That is, the first evaluation result is an evaluation result obtained by evaluating triple data, while the second evaluation result is another evaluation result obtained by evaluating an image-text data pair (the text in the data pair refers to one of the question text or the answer text). In addition, the evaluation result set may also include evaluation results of other categories. For example, the evaluation result obtained by evaluating a question-answer text pair. After obtaining the evaluation result set, each evaluation result in the evaluation result set can be divided according to quality and security. For example, the evaluation values of the security dimension can be selected from the evaluation result set and used as the security sub-results. The evaluation values of the quality dimension can be selected from the evaluation result set and used as the quality sub-results. The predetermined security condition may be that the security sub-result is less than or equal to the corresponding threshold. The predetermined quality condition may include that each quality sub-result is greater than or equal to the corresponding threshold, or the total quality evaluation value obtained by performing a weighted operation or other operations on each quality sub-result is greater than or equal to the corresponding threshold.

[0071] In one example, if the security sub-result does not meet the predetermined security condition, the category of the data to be processed can be determined as risk data.

[0072] In another example, if the security sub-result meets the predetermined security condition and the quality sub-result meets the predetermined quality condition, the category of the data to be processed can be determined as high-quality data.

[0073] In another example, if the security sub-result meets the predetermined security condition and the quality sub-result does not meet the predetermined quality condition, it can be determined that the data to be processed meets the predetermined modification condition, and the category of the data to be processed can be determined as the category to be modified.

[0074] According to another embodiment of the present disclosure, for the data to be processed in the category to be modified, the following method can be used for modification. For example, based on the question text, the answer text, and the reason text, the answer text can be modified based on a large language model to obtain the modified answer text, and then the data to be processed can be updated according to the image, the question text, and the modified answer text.

[0075] For example, the above-mentioned reason text is generated during the evaluation of the data to be processed by a large model such as a multimodal large model or a large language model, and the reason text represents the reason for the data to be processed to obtain the corresponding evaluation value.

[0076] For example, the process of updating the data to be processed may include: replacing the original response text in the data to be processed with the modified response text. Then, the data to be processed can be re-evaluated and classified to determine the category of the updated data to be processed.

[0077] In some embodiments, if the number of updates of the data to be processed is greater than or equal to a predetermined number, it means that after multiple modifications to the response text, the quality of the data to be processed is still low. Therefore, the category of the data to be processed can be determined as data to be labeled.

[0078] In practical applications, the data to be processed can be divided into multiple categories, which may include invalid data, risk data, high-quality data, data to be labeled, etc. For invalid data, the invalid data can be neither used to train the model nor evaluated by the multimodal large model. For risk data, it can be used as a negative sample to iterate the multimodal large model, large language model and other large models. For high-quality data, it can be used as a positive sample to iterate the multimodal large model, large language model and other large models, or as a test sample to test the training effect of the large model. For the data to be labeled, the data to be labeled can be fed back to the staff, and the staff can label the data to be labeled according to actual needs.

[0079] Figure 3 It is a schematic diagram of the data classification method according to the embodiments of the present disclosure.

[0080] In this embodiment, the data to be processed Data_2 can be screened from multiple candidate data Data_1. For example, each candidate data Data_1 includes a candidate image, a candidate question, and a candidate answer. The image hash value of the candidate image and the question hash value of the candidate question are concatenated to obtain the reference information of the data to be processed Data_2. Then, based on the reference information, the multiple candidate data Data_1 are de-duplicated, and the de-duplicated candidate data Data_1 are determined as the data to be processed Data_2. It can be seen that the candidate image in the candidate data is the image Iamge in the data to be processed Data_2, the candidate question in the candidate data is the question text Text_Q in the data to be processed Data_2, and the candidate answer in the candidate data is the answer text Text_A in the data to be processed Data_2.

[0081] Next, the data to be processed Data_2 can be preliminarily evaluated. For example, any one of the image Iamge, question text Text_Q and answer text Text_A in the data to be processed Data_2 can be used as the target sub-data, and then the target sub-data can be tested for security, such as detecting whether any one of the image Iamge, question text Text_Q and answer text Text_A contains bad content, thereby obtaining security parameters. If the security parameters meet the security abnormality conditions, it is determined that the category of the data to be processed Data_2 is risk data, and the data to be processed Data_2 can be added to the risk data set A_2. The target sub-data can also be tested for quality, such as detecting whether the image Iamge is blurred or has low resolution, and detecting whether the question text Text_Q and answer text Text_A contain valid content, and detecting the number of characters in the question text Text_Q and answer text Text_A, thereby obtaining quality parameters. If the quality parameters meet the quality abnormality conditions, it is determined that the category of the data to be processed Data_2 is invalid data, and the data to be processed Data_2 can be added to the invalid data set A_1.

[0082] The above preliminary evaluation process is based on the single-modal target sub-data for evaluation. After the preliminary evaluation, for the remaining risk-free and high-quality data to be processed Data_2, the large model can be used to evaluate the remaining data to be processed Data_2.

[0083] For example, the first prompt information, image Iamge, question text Text_Q, and answer text Text_A can be input into the multimodal big model, and the multimodal big model can be used to evaluate the answer content in terms of security, semantics, text coherence, satisfaction of question intent, knowledge correctness and other dimensions to obtain evaluation values ​​for each dimension.

[0084] For another example, the second prompt information, the image Iamge, and the question text Text_Q can be input into the multimodal large model, and the relevance between the image and the text can be evaluated through the multimodal large model to obtain the evaluation value of the dimension.

[0085] For another example, the third prompt information, image Iamge, and answer text Text_A can be input into the multimodal large model, and the correctness of the description of the image Iamge can be evaluated by the multimodal large model to obtain the evaluation value of this dimension.

[0086] For another example, the fourth prompt information, the answer text Text_A and the question text Text_Q can be input into the large language model, and the large language model will evaluate the question and answer semantic alignment, question content security, answer content security, and answer knowledge correctness to obtain evaluation values ​​for each dimension.

[0087] Through the above evaluation process, multiple evaluation results can be obtained. Each evaluation result can be added to the evaluation result set, so that the evaluation result set can include the evaluation values obtained from the above various evaluation processes. The evaluation values in the evaluation result set can be divided into two parts according to the evaluation dimension. One part is the security sub-result in the security dimension, and the other part is the quality sub-result in the quality dimension. For example, the fact that the security sub-result of the data to be processed Data_2 meets the predetermined security condition indicates that the data to be processed Data_2 is compliant, and the fact that the quality sub-result of the data to be processed Data_2 meets the predetermined quality condition indicates that the data to be processed Data_2 is qualified.

[0088] Next, classification can be performed according to the evaluation result set.

[0089] If the data to be processed Data_2 is non-compliant, it can be determined that the data to be processed Data_2 is category risk data, and the data to be processed Data_2 can be added to the risk data set A_2.

[0090] If the data to be processed Data_2 is compliant and qualified, it can be determined that the data to be processed Data_2 is category high-quality data, and the data to be processed Data_2 can be added to the high-quality data set A_3.

[0091] If the data to be processed Data_2 is compliant but unqualified, it can be determined that the category of the data to be processed Data_2 is data to be modified, and the data to be processed Data_2 of this category can be added to the data set to be modified.

[0092] For the data to be processed Data_2 in the above data set to be modified, a large language model can be used as the rewriting model, and then based on the rewriting model, the answer text Text_A in the data to be processed Data_2 is modified, so as to update the data to be processed Data_2 with the modified answer text Text_A. If the update times of the data to be processed Data_2 are less than the predetermined times, the updated data to be processed Data_2 can be evaluated and classified. If the update times of the data to be processed Data_2 are greater than or equal to the predetermined times, it can be determined that the category of the data to be processed Data_2 is data to be labeled, and the data to be processed Data_2 can be added to the data set to be labeled A_4, and then the data set to be labeled A_4 can be manually labeled according to actual needs.

[0093] Figure 4 It is a schematic structural block diagram of the data classification device according to an embodiment of the present disclosure.

[0094] Such as Figure 4As shown, the data classification device 400 may include an acquisition module 410, a first evaluation module 420, a second evaluation module 430, and a classification module 440.

[0095] The acquisition module 410 is configured to acquire data to be processed, and the data to be processed includes images, question texts, and answer texts.

[0096] The first evaluation module 420 is configured to evaluate the security and quality of the data to be processed based on a multimodal large model according to the question text, the answer text, and the image, and obtain a first evaluation result.

[0097] The second evaluation module 430 is configured to evaluate the quality of the data to be processed based on a multimodal large model according to one of the question text and the answer text, and the image, and obtain a second evaluation result.

[0098] The classification module 440 is configured to classify the data to be processed according to the evaluation result set, and the evaluation result set includes the first evaluation result and the second evaluation result.

[0099] According to another embodiment of the present disclosure, the first evaluation module includes: a first sub-module, configured to determine a first evaluation result based on a multimodal large model according to the first prompt information, the image, the question text, and the answer text; wherein, the first prompt information is used to guide the multimodal large model to evaluate from at least one of the following dimensions: security, coherence, intention satisfaction, and answer correctness.

[0100] According to another embodiment of the present disclosure, the second evaluation module includes at least one of the following: a second sub-module and a third sub-module. The second sub-module is configured to determine a second evaluation result based on a multimodal large model according to the second prompt information, the image, and the question text; wherein, the second prompt information is used to guide the multimodal large model to evaluate the quality from the dimension of graphic-text relevance. The third sub-module is configured to determine a second evaluation result based on a multimodal large model according to the third prompt information, the image, and the answer text; wherein, the third prompt information is used to guide the multimodal large model to evaluate the quality from the dimension of image description correctness.

[0101] According to another embodiment of the present disclosure, the classification module includes: a fourth sub-module and a fifth sub-module. The fourth sub-module is configured to determine that the category of the data to be processed is high-quality data in response to detecting that the security sub-result in the evaluation result set meets a predetermined security condition and the quality sub-result in the evaluation result set meets a predetermined quality condition. The fifth sub-module is configured to determine that the category of the data to be processed is risk data in response to detecting that the security sub-result in the evaluation result set does not meet the predetermined security condition.

[0102] According to another embodiment of the present disclosure, it further includes: a modification module and an update module. The modification module is configured to, in response to detecting that the evaluation result set meets a predetermined modification condition, modify the answer text based on the large language model according to the question text, the answer text, and the reason text to obtain a modified answer text; wherein, the reason text is generated during the process of the multi-modal large model evaluating the data to be processed, and the reason text represents the reason for the data to be processed to obtain the evaluation result. The update module is configured to update the data to be processed according to the image, the question text, and the modified answer text.

[0103] According to another embodiment of the present disclosure, the predetermined modification condition includes: the security sub-result in the evaluation result set meets the predetermined security condition, and the quality sub-result in the evaluation result set does not meet the predetermined quality condition.

[0104] According to another embodiment of the present disclosure, it further includes: a category determination module, configured to, in response to detecting that the update times of the data to be processed are greater than or equal to a predetermined number of times, determine the category of the data to be processed as data to be labeled.

[0105] According to another embodiment of the present disclosure, it further includes: a candidate data acquisition module, a reference information determination module, a duplicate removal module, and a data to be processed determination module. The candidate data acquisition module is configured to acquire a plurality of candidate data, and each candidate data in the plurality of candidate data includes a candidate image, a candidate question, and a candidate answer. The reference information determination module is configured to, for each candidate data, determine the reference information of the candidate data according to the image information of the candidate image and the question information of the candidate question. The duplicate removal module is configured to remove duplicates from the plurality of candidate data according to the respective reference information of the plurality of candidate data. The processed data determination module is configured to determine the de-duplicated candidate data as the data to be processed.

[0106] According to another embodiment of the present disclosure, the image information of the candidate image is an image hash value, and the question information of the candidate question is a question hash value; the reference information determination module includes: a splicing sub-module, configured to splice the image hash value and the question hash value to obtain the reference information.

[0107] According to another embodiment of the present disclosure, it further includes at least one of the following: a first processing module and a second processing module. The first processing module is configured to perform a security detection on the target sub-data in the data to be processed to obtain a security parameter; in response to detecting that the security parameter meets the security exception condition, determine the category of the data to be processed as risk data. The second processing module is configured to perform a quality detection on the target sub-data in the data to be processed to obtain a quality parameter; in response to detecting that the quality parameter meets the quality exception condition, determine the category of the data to be processed as invalid data. The target sub-data is any one of an image, a question text, and an answer text.

[0108] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above data classification method.

[0109] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above data classification method.

[0110] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program which, when executed by a processor, implements the above data classification method.

[0111] Figure 5 is a structural block diagram of an electronic device for implementing the data classification method of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0112] As Figure 5 shown, the device 500 includes a computing unit 501, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0113] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0114] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the data classification method. For example, in some embodiments, the data classification method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the data classification method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the data classification method in any other suitable manner (e.g., by means of firmware).

[0115] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0118] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0119] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0120] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0122] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data classification method, comprising: Acquire data to be processed, wherein the data to be processed includes an image, a question text, and an answer text; Based on the multimodal large model, the security and quality of the data to be processed are evaluated according to the question text, the answer text and the image to obtain a first evaluation result; Based on the multimodal large model, the quality of the data to be processed is evaluated according to the question text, one of the answer text, and the image to obtain a second evaluation result; as well as The data to be processed is classified according to an evaluation result set; the evaluation result set includes the first evaluation result and the second evaluation result.

2. The method according to claim 1, wherein: Based on the multimodal large model, the security and quality of the data to be processed are evaluated according to the question text, the answer text and the image, and the first evaluation result is obtained, including: Determine the first evaluation result based on the multimodal large model according to the first prompt information, the image, the question text and the answer text; Among them, the first prompt information is used to guide the multimodal large model to evaluate from at least one of the following dimensions: security, coherence, intention satisfaction, and answer correctness.

3. The method according to claim 1, wherein: The method of evaluating the quality of the data to be processed based on the multimodal large model according to one of the question text and the answer text, and the image, and obtaining a second evaluation result includes at least one of the following: Based on the multimodal large model, the second evaluation result is determined according to the second prompt information, the image and the question text; wherein the second prompt information is used to guide the multimodal large model to perform quality evaluation from the dimension of image-text relevance; and Based on the multimodal large model, the second evaluation result is determined according to the third prompt information, the image and the answer text; wherein the third prompt information is used to guide the multimodal large model to perform quality assessment from the dimension of image description correctness.

4. The method according to claim 1, wherein: The classifying the data to be processed according to the evaluation result set includes: In response to detecting that the security sub-result in the evaluation result set satisfies a predetermined security condition and the quality sub-result in the evaluation result set satisfies a predetermined quality condition, determining that the category of the data to be processed is high-quality data; and In response to detecting that the security sub-result in the evaluation result set does not satisfy a predetermined security condition, determining that the category of the data to be processed is risk data.

5. The method according to claim 1, further comprising: In response to detecting that the evaluation result set satisfies a predetermined modification condition, the answer text is modified based on the large language model according to the question text, the answer text and the reason text to obtain a modified answer text; wherein the reason text is generated during the process of evaluating the data to be processed by the multimodal large model, and the reason text represents the reason why the data to be processed obtains the evaluation result; and The data to be processed is updated according to the image, the question text and the modified answer text.

6. The method according to claim 5, wherein: The predetermined modification conditions include: The security sub-result in the evaluation result set meets a predetermined security condition, and the quality sub-result in the evaluation result set does not meet a predetermined quality condition.

7. The method according to claim 5, further comprising: In response to detecting that the update times of the data to be processed is greater than or equal to a predetermined times, it is determined that the category of the data to be processed is data to be labeled.

8. The method according to claim 1, further comprising: Acquire a plurality of candidate data, each of the plurality of candidate data comprising a candidate image, a candidate question, and a candidate answer; For each candidate data, determining reference information of the candidate data according to image information of the candidate image and question information of the candidate question; Deduplication of the plurality of candidate data is performed according to the respective reference information of the plurality of candidate data; as well as The candidate data after deduplication is determined as the data to be processed.

9. The method according to claim 8, wherein: The image information of the candidate image is an image hash value, and the question information of the candidate question is a question hash value; The determining, for each candidate data, according to the image information of the candidate image and the question information of the candidate question, the reference information of the candidate data comprises: The image hash value and the question hash value are concatenated to obtain the reference information.

10. The method of claim 1, further comprising at least one of the following: Performing security detection on target sub-data in the data to be processed to obtain security parameters; In response to detecting that the security parameter satisfies a security abnormality condition, determining that the category of the data to be processed is risk data; Performing quality detection on target sub-data in the data to be processed to obtain quality parameters; In response to detecting that the quality parameter satisfies a quality abnormality condition, determining that the category of the data to be processed is invalid data; The target sub-data is any one of the image, the question text and the answer text.

11. A data classification device, comprising: An acquisition module, used for acquiring data to be processed, wherein the data to be processed includes an image, a question text and an answer text; A first evaluation module is used to evaluate the security and quality of the data to be processed based on the multimodal large model according to the question text, the answer text and the image, and obtain a first evaluation result; A second evaluation module is used to evaluate the quality of the data to be processed based on the multimodal large model according to one of the question text and the answer text, and the image, to obtain a second evaluation result; as well as A classification module is used to classify the data to be processed according to an evaluation result set; the evaluation result set includes the first evaluation result and the second evaluation result.

12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Video evaluation method and device based on artificial intelligence, equipment and medium

    CN120852971A