Method, device, and storage medium for training and using a quality assessment model

By vectorizing and fusion processing of multimodal data, combined with binary classification training of activation functions, the association relationship identification problem in multimodal data compliance detection is solved, and more accurate quality evaluation and detection is achieved.

CN115048996BActive Publication Date: 2025-07-01BEIJING 58 INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210658095.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-07-01
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect the association relationship between multimodal data, resulting in vulnerabilities in multimodal data compliance detection and the overall compliance of data cannot be ensured.

Method used

By vectorized and fusion processing of multimodal data, binary classification is performed using activation functions, and batch training is carried out to obtain general evaluation parameters of quality evaluation model, and quality evaluation is performed based on the association relationship of different modal data.

Benefits of technology

The overall quality evaluation of multimodal data is achieved, the accuracy and universality of detection are improved, and whether the data meets the needs of the scenario is possible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115048996B_ABST
    Figure CN115048996B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, and storage medium for training and using a quality assessment model. Among them, the quality assessment model is trained based on vectors corresponding to data of different modalities in multimodal data. Therefore, the quality assessment model can be used to evaluate the quality of target multimodal data. Moreover, before training the model using vectors of multiple modalities, the vectors of multiple modalities are also fused and the fused vectors are binary-classified, and the model is trained based on the binary-classified fused vectors. In this way, when using the quality assessment model to evaluate the quality of target multimodal data, it is not only applicable to evaluating the quality of data of each modality in the target multimodal data, but also can comprehensively evaluate the quality of the target multimodal data in combination with the correlation between different modalities of data in the target multimodal data, and the obtained quality assessment result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model training, and in particular, to a method, device, and storage medium for training and using a quality assessment model. Background Art

[0002] In the scenario of posting on the Internet, the content of many posts is multi-modal data. For example, in the recruitment scenario, the content of recruitment posts can be divided into text data, user behavior data, and auxiliary data for describing user behavior from the perspective of data modality. For enterprises, from the perspective of data security, it is necessary to perform compliance detection on the published data to determine whether the data has been tampered with or suffered malicious attacks.

[0003] When performing compliance detection on multi-modal data in a target scenario, model training is usually performed based on a large amount of historical multi-modal data in the target scenario to obtain an evaluation model for evaluating whether any multi-modal data in the target scenario is compliant. Although this method can perform compliance detection on each modality of data, in the case where there is an association relationship between different modality data, if each modality of data is compliant but the data between different modalities is not compliant, it is very difficult to detect. Therefore, it is necessary to provide a solution for overall compliance evaluation of multi-modal data. Summary of the Invention

[0004] This application provides a method, device, and storage medium for training and using a quality assessment model from multiple aspects to perform model training and quality assessment on multi-modal data and determine the compliance of multi-modal data.

[0005] An embodiment of this application provides a method for training a quality assessment model for multi-modal data, including: obtaining N multi-modal sample data corresponding to N training tasks, where each multi-modal sample data includes at least two modalities of data; performing vectorization processing and fusion processing on the N multi-modal sample data to obtain N fusion vectors, and performing binary classification on the N fusion vectors using an activation function; based on the initialized evaluation parameters and the N fusion vectors after binary classification, performing the first batch training on the quality assessment model to obtain multiple intermediate state evaluation parameters corresponding to each batch of training; based on the multiple intermediate state evaluation parameters and the N fusion vectors after binary classification, performing the second batch training on the quality assessment model to obtain the total function loss corresponding to each batch of training; determining the general evaluation parameters corresponding to the quality assessment model according to the multiple total function losses obtained from the second batch training.

[0006] In an alternative embodiment, obtaining N multi-modal sample data corresponding to N training tasks includes: obtaining T multi-modal data under multiple scenarios, and randomly sampling N multi-modal sample data from the T multi-modal data as N training tasks; where N is a positive integer greater than 1, and N is less than or equal to T.

[0007] In an alternative embodiment, vectorizing the N multi-modal sample data includes: performing vector calculation on the text data in the N multi-modal sample data using a word vector calculation method to obtain word vectors in the N multi-modal sample data; performing vector processing on the behavior data in the N multi-modal sample data using a graph neural network to obtain behavior vectors in the N multi-modal sample data; performing vector processing on the auxiliary data in the N multi-modal sample data using a one-hot encoding method to obtain encoded vectors in the N multi-modal sample data.

[0008] In an alternative embodiment, fusing the N multi-modal sample data to obtain the N fusion vectors includes: for the N vectors obtained by vectorizing the N multi-modal sample data, taking every two vectors as a group, calculating the outer product of the every two vectors to obtain intermediate vectors; performing a flattening process on the intermediate vectors to obtain the fusion vectors corresponding to the N multi-modal sample data respectively.

[0009] In an alternative embodiment, the N fusion vectors include X support vectors and Y query vectors; based on the initialized evaluation parameters and the N fusion vectors after binary classification, performing the first batch training on the quality evaluation model to obtain multiple intermediate state evaluation parameters corresponding to each batch of training, including: according to the set number of samples per time, batch obtaining the corresponding number of first fusion vectors from the X support vectors, and based on the initialized evaluation parameters and the first loss function, using the batch-obtained first fusion vectors to perform the first gradient descent calculation on the quality evaluation model in sequence to obtain the single-sample-number intermediate state evaluation parameters corresponding to each calculation.

[0010] In an alternative embodiment, based on the multiple intermediate state evaluation parameters and the N fusion vectors after binary classification, performing the second batch training on the quality evaluation model to obtain the total function loss corresponding to each batch of training, including: according to the set number of samples per time, batch obtaining the corresponding number of second fusion vectors from the Y query vectors; based on the single-sample-number intermediate state evaluation parameters obtained each time and the second loss function, using the batch-obtained second fusion vectors to perform the second gradient descent calculation on the quality evaluation model in sequence to obtain the total function loss corresponding to each calculation; where the total function loss is the sum of the function losses of the second fusion vectors corresponding to each calculation under the intermediate state evaluation parameters respectively used by them.

[0011] In an alternative embodiment, determining the general evaluation parameter corresponding to the quality evaluation model according to the sum of multiple function losses obtained from the second batch training includes: determining the intermediate state evaluation parameter corresponding to the minimum sum of multiple function losses obtained from the second batch training; using the intermediate state evaluation parameter corresponding to the minimum sum of function losses as the general evaluation parameter corresponding to the quality evaluation model.

[0012] In an alternative embodiment, it further includes: obtaining target multi-modal data in a target scenario, and inputting the target multi-modal data into the quality evaluation model; inside the quality evaluation model, performing quality evaluation on the target multi-modal data according to the general evaluation parameter, and the quality evaluation result represents the quality of the target multi-modal data.

[0013] In an alternative embodiment, when obtaining T multi-modal data in multiple scenarios, it further includes: randomly sampling M multi-modal sample data from the T multi-modal data as M test tasks, where M is a positive integer greater than 1, and (N + M) is less than or equal to T; performing vectorization processing and fusion processing on the M multi-modal sample data to obtain M fusion vectors, and using an activation function to perform binary classification on the M fusion vectors.

[0014] In an alternative embodiment, the M multi-modal sample data includes the target multi-modal data, and the M fusion vectors obtained after vectorization processing and fusion processing of the M multi-modal sample data include P support vectors and Q query vectors, where P is less than N; before inputting the target multi-modal data into the quality evaluation model, it further includes: performing third batch training on the quality evaluation model using the P support vectors and fine-tuning the general evaluation parameter.

[0015] In an alternative embodiment, performing quality evaluation on the target multi-modal data according to the general evaluation parameter inside the quality evaluation model includes: inside the quality evaluation model, performing vectorization processing and fusion processing on the target multi-modal data to obtain a fusion vector corresponding to the target multi-modal data; using an activation function to perform binary classification on the fusion vector corresponding to the target multi-modal data; predicting the binary-classified fusion vector based on the fine-tuned general evaluation parameter and the Q query vectors to obtain the ratio of the two types of fusion vectors as the quality evaluation result and outputting it.

[0016] The embodiments of the present application further provide a method for using a quality assessment model, including: obtaining target multimodal data in a target scenario, and inputting the target multimodal data into the quality assessment model; inside the quality assessment model, performing quality assessment on the target multimodal data according to general assessment parameters, and the quality assessment result represents the quality of the target multimodal data.

[0017] In an optional embodiment, inside the quality assessment model, performing quality assessment on the target multimodal data according to general assessment parameters includes: inside the quality assessment model, performing vectorization processing and fusion processing on the target multimodal data to obtain a fusion vector corresponding to the target multimodal data, and using an activation function to perform binary classification on the fusion vector corresponding to the target multimodal data; according to the evaluation model parameters, predicting the binary-classified fusion vector to obtain the ratio of the two types of fusion vectors corresponding thereto and using it as the quality assessment result and outputting it.

[0018] The embodiments of the present application further provide a training device for a quality assessment model for multimodal data, including: a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps in any one of the above methods.

[0019] The embodiments of the present application further provide a training device for a quality assessment model for multimodal data, including: a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps in any one of the above methods.

[0020] The embodiments of the present application further provide a computer-readable storage medium storing computer programs / instructions, which, when executed by a processor, cause the processor to implement the steps in any one of the above methods.

[0021] In the embodiments of the present application, the quality assessment model is trained based on vectors corresponding to data of different modalities in multimodal data. Therefore, this quality assessment model can be used to assess the quality of target multimodal data. Moreover, before using vectors of multiple modalities for model training, the vectors of multiple modalities are fused and the fused vectors are binary-classified, and the model is trained based on the binary-classified fused vectors. In this way, when using the quality assessment model to assess the quality of target multimodal data, it is not only applicable to assessing the quality of data of each modality in the target multimodal data, but also can comprehensively assess the quality of the target multimodal data in combination with the correlation between data of different modalities in the target multimodal data, and the obtained quality assessment result is more accurate. In addition, the quality assessment model provided by the embodiments of the present application is trained based on multimodal data in multiple scenarios, and is more universal in use, without any limitation on the scenario corresponding to the target multimodal data to be evaluated. Description of the Drawings

[0022] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0023] Figure 1a It is a flowchart of a method for training a quality assessment model provided by an embodiment of the present application;

[0024] Figure 1b It is a flowchart of a method for using a quality assessment model provided by an embodiment of the present application;

[0025] Figure 1c It is a schematic diagram of the process for assessing the quality of target multimodal data provided by an embodiment of the present application;

[0026] Figure 2 It is a schematic diagram of the structure of a quality assessment model training device provided by an embodiment of the present application. Detailed Embodiments

[0027] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0028] In many scenarios of the Internet, multi-modal data is involved. For example, in the scenario of posting a thread, it involves the text data corresponding to the thread, the behavior data such as the user's browsing, clicking, and favoriting of the thread, and also the auxiliary data for restricting, interpreting, and explaining the content of the thread and the user's behavior. Another example is the customer service scenario, which involves the text data corresponding to the communication between the customer service and the user, the behavior data such as the user's initiation of consultation with the customer service, browsing of the communication content, evaluation of the service content, and the customer service's acceptance of the user's consultation and provision of corresponding services to the user, and also the auxiliary data for restricting, interpreting, and explaining the identity information and behavior of the customer service and the user. However, in Internet scenarios, malicious attacks are often suffered. For example, data is tampered with or abnormally published, etc. Taking the recruitment scenario as an example, for instance, a position is usually posted in one city, but in the case of a malicious attack, there may be a situation where the same position is posted in different cities; another example is that for a certain position, when no suitable candidate has been recruited, the position information is usually updated once a specified period, such as one week, but in the case of a malicious attack, there may be a situation where the position information is continuously updated; another example is that the posted position information is about cleaning services, but in the case of a malicious attack, the position information may be tampered with to be about other content such as maternity nurses, engineers, technicians, etc., or the position information is tampered with to contain bad information such as fraud, gambling, and obscenity. Therefore, for the above similar situations and for the target scenario, it is necessary to perform compliance verification on the data in the target scenario.

[0029] However, there is a certain correlation between different modalities of multi-modal data. Even if any one modality of data meets the scenario requirements, if the correlation relationship between different modalities of data does not meet the scenario requirements, then the multi-modal data cannot meet the scenario requirements. For example, in a recruitment scenario, the content, posting time, and posting city of recruitment post A all meet the requirements. However, if a large number of recruitment posts A are posted at the corresponding posting time in the target city, it can be determined that recruitment post A was maliciously attacked during posting and belongs to abnormal posting. Another example is that the content, posting time, and the number of recruitment posts B posted in the target city of recruitment post B all meet the requirements, that is, only one recruitment post B is posted in the target city. However, if recruitment post B is also posted at the corresponding posting time in other cities except the target city, it can be determined that recruitment post B was maliciously attacked during posting and belongs to abnormal posting. Based on this, in the process of verifying the compliance of multi-modal data, it is necessary to jointly verify the compliance of multi-modal data by combining the correlation relationship between different modalities of data. To this end, the embodiments of the present application provide a method for training a quality evaluation model for multi-modal data. During model training, in order to avoid considering the characteristics of data only from a single modality, the model training method provided by the embodiments of the present application can fuse different modalities of data in multi-modal data and perform model training based on the fused data to train a general model for comprehensive consideration of multi-modal data in different scenarios for quality evaluation.

[0030] Taking the recruitment scenario as an example below, the quality evaluation model training method provided by the embodiments of the present application will be described in conjunction with the accompanying drawings.

[0031] Figure 1a It is a flowchart of the quality evaluation model training method provided by the embodiments of the present application. As Figure 1a shown, the method includes:

[0032] S1a. Obtain N multi-modal sample data corresponding to N training tasks, and each multi-modal sample data includes at least two modalities of data;

[0033] S2a. Perform vectorization processing and fusion processing on the N multi-modal sample data to obtain N fusion vectors, and use an activation function to perform binary classification on the N fusion vectors;

[0034] S3a. Based on the initialized evaluation parameters and the N fusion vectors after binary classification, perform the first batch training on the quality evaluation model to obtain multiple intermediate state evaluation parameters corresponding to each batch of training;

[0035] S4a. Based on the multiple intermediate state evaluation parameters and the N fusion vectors after binary classification, perform the second batch training on the quality evaluation model to obtain the total function loss corresponding to each batch of training;

[0036] S5a. Determine the general evaluation parameters corresponding to the quality evaluation model according to the total sum of function losses obtained from the second batch training.

[0037] In the embodiments of the present application, multi-modal data in multiple scenarios can be used for model training to obtain the quality evaluation model provided by the embodiments of the present application. In the embodiments of the present application, the source of the multi-modal data used is not limited. In an alternative embodiment, in order to adapt to the data characteristics corresponding to the recruitment scenario, the multi-modal data suitable for model training can be the online data corresponding to the recruitment scenario. For example, historical multi-modal data within the last 3 months can be selected for model training. In another alternative embodiment, the multi-modal data for model training can also be the multi-modal data obtained from the business platform. For example, the multi-modal data required for model training can be obtained from the anti-fraud platform, and these data contain labeled data to identify the data that does not conform to the recruitment scenario. For example, sensitive words that cannot be used in the recruitment scenario are marked. In another alternative embodiment, the online historical multi-modal data corresponding to the recruitment scenario and the multi-modal data provided by the business platform can also be obtained simultaneously for model training using these data. In the embodiments of the present application, based on the online historical multi-modal data corresponding to the recruitment scenario and / or the multi-modal data provided by the business platform for model training, the data characteristics corresponding to the recruitment scenario can be learned to obtain a model suitable for quality evaluation of multi-modal data in the recruitment scenario.

[0038] In the embodiments of the present application, each multi-modal data can be used as a training task. Assume that the number of training tasks used in the embodiments of the present application is N. To ensure that the trained quality evaluation model can not only evaluate the quality of each type of data in the multi-modal data but also evaluate the multi-modal data as a whole by combining the correlation relationships between different types of data in the multi-modal data, when N multi-modal sample data corresponding to N training tasks are obtained, the data of each type in the N multi-modal sample data can be respectively vectorized and fused to obtain N fused vectors. Among them, vectorizing the N multi-modal sample data means: respectively vectorizing the data of each type in the N multi-modal sample data; fusing the N multi-modal sample data means: for each multi-modal sample data among the N multi-modal sample data, fusing the vectors obtained by vectorizing the data of each type included therein to obtain N fused vectors.

[0039] Further, an activation function can be used to perform binary classification on the N fusion vectors, so as to perform model training based on the N fusion vectors after binary classification. In the embodiments of the present application, the multimodal data that meets the requirements of the recruitment scenario is called compliant data, and the multimodal data that does not meet the requirements of the recruitment scenario is called non-compliant data. The purpose of binary classification is to distinguish between compliant and non-compliant data in the multimodal sample data. Based on the fusion vectors after binary classification for model training, the characteristics corresponding to compliant and non-compliant data can be learned, and then the quality assessment model obtained by training can be used to perform quality assessment on the target multimodal data. The quality assessment result represents the compliance of the target multimodal data. In the embodiments of the present application, the specific method for performing binary classification on the fusion vectors is not limited. Optionally, the fusion vectors can be distinguished according to the vector values of 0 and 1, and the ratio of the two can be determined according to the number of vectors corresponding to the vector values of 0 and 1 respectively. Furthermore, the compliance of the multimodal data can be determined according to this ratio. For example, if the vector value of 1 indicates that the data is compliant and the vector value of 0 indicates that the data is non-compliant, when the ratio of the vector value of 1 to the vector value of 0 in the fusion vector corresponding to the target multimodal data is greater than 50%, it is considered that the target multimodal data is compliant. Of course, this is not limited thereto, and according to specific requirements, the vector ratios used to represent whether the multimodal data is compliant can also be different.

[0040] In the embodiments of the present application, an initial evaluation parameter can be set for the quality assessment model. When performing model training, the quality assessment model can be subjected to the first batch training based on the initialization evaluation parameter and the N fusion vectors after binary classification, so as to obtain the intermediate state evaluation parameters corresponding to the quality assessment model after each batch of training; further, these intermediate state evaluation parameters are used as new evaluation parameters, and the N fusion vectors after binary classification are used to perform the second batch training on the quality assessment model, so as to obtain the total sum of multiple function losses corresponding to the quality assessment model after the second batch of training, and the general evaluation parameter corresponding to the quality assessment model is determined according to the total sum of these multiple function losses, so as to perform quality assessment on the target multimodal data to be evaluated according to this general evaluation parameter. For the specific processes of the first batch training and the second batch training of the quality assessment model, reference can be made to the following embodiments, which will not be elaborated here for the time being.

[0041] In the embodiments of the present application, the specific implementation manner of the initial evaluation parameter is not limited, and according to the difference of the quality assessment model, the implementation manner of the corresponding initial evaluation parameter can also be different. Further, the specific type of the quality assessment model is not limited in the embodiments of the present application. According to specific scenario requirements and the characteristics of multimodal data, a suitable model can be selected for training. Optionally, the embodiments of the present application take the optimized meta-learning framework (Model-agnostic MetaLearning, MAML) as an example for illustration, and the corresponding initialization evaluation parameter is a function with a normal distribution characteristic. Of course, this is not limited thereto.

[0042] Next, the specific implementation process of each step in the above method will be described in detail.

[0043] In the embodiments of the present application, in order to ensure the universality of the N multi-modal sample data for model training, when obtaining N training tasks for model training, T multi-modal data in multiple scenarios can be obtained first, and then N multi-modal sample data are randomly sampled from the T multi-modal data as N training tasks; where N is a positive integer greater than 1 and N is less than or equal to T. In this way, it is possible to avoid selecting multi-modal data with too single characteristics for model training. For example, if the N multi-modal sample data selected are all recruitment data corresponding to the same type of position, or most of the N multi-modal sample data are recruitment data corresponding to the same or similar types of positions, the quality evaluation model obtained by using these multi-modal sample data is difficult to meet the universality requirements. Further, in the case of obtaining N multi-modal sample data that meet the requirements, the N multi-modal sample data can be vectorized and fused. In the embodiments of the present application, the specific method of vectorizing the N multi-modal sample data is not limited. Since the types of data modalities included in each multi-modal sample data are different, the method of vectorization will also be different.

[0044] In the embodiments of the present application, the modal type corresponding to each type of modal data is not limited. Optionally, taking the example that each multi-modal data includes three types of modal data: text data, behavior data, and auxiliary data, when vectorizing the N multi-modal sample data, the text data, behavior data, and auxiliary data in each multi-modal sample data can be vectorized in an appropriate manner respectively. Optionally, a word vector calculation method can be used to calculate the word vectors in the N multi-modal sample data to obtain the word vectors in the N multi-modal sample data; a graph neural network can be used to process the behavior data in the N multi-modal sample data to obtain the behavior vectors in the N multi-modal sample data; and a one-hot encoding method can be used to process the auxiliary data in the N multi-modal sample data to obtain the encoded vectors in the N multi-modal sample data. For example, a fast text classification model (FastText) can be used to calculate the word vectors corresponding to the text data; for another example, a graph embedding model (Large-scale Information Network Embedding, LINE) can be used to determine the behavior vectors corresponding to each behavior data and their similarities; for another example, a one-hot encoding model (one-hot) can be used to process the auxiliary data and classify it. Of course, the above is only an illustrative description and is not limited thereto.

[0045] Correspondingly, the embodiments of the present application do not limit the specific manner of fusing N multi-modal sample data. Optionally, methods such as point-wise addition or concatenate can be used to fuse the N vectors obtained after vectorizing the N multi-modal sample data. Taking the concatenate method as an example, for each multi-modal sample data, every two vectors can be grouped, and the outer product of every two vectors can be calculated to obtain intermediate vectors; then, the intermediate vectors are flattened to obtain the fusion vectors corresponding to the N multi-modal sample data respectively.

[0046] Based on the above, in the case of obtaining N fusion vectors corresponding to N multi-modal sample data, model training can be performed based on the N fusion vectors to obtain a quality evaluation model that can evaluate the compliance of target multi-modal data. In the embodiments of the present application, the number of training tasks used each time for batch training can be determined according to the set batch size, that is, the number of fusion vectors, and the corresponding number of fusion vectors is obtained in batches according to the set batch size for model training. In the embodiments of the present application, the specific manner of batch training the quality evaluation model is not limited. Optionally, the N fusion vectors can be divided into X support sets and Y query sets; where the numbers of X and Y can be the same or different, and are not limited here. Assuming that the X support sets are divided into K groups according to the set batch size, then batch_size first fusion vectors are obtained from the X support sets in batches, and this is done K times, that is, the number of times of performing the first batch training is K; further, based on the initialized evaluation parameters and the first loss function, the first gradient descent calculation is performed on the quality evaluation model in turn using the batch-obtained first fusion vectors to obtain the batch_size intermediate evaluation parameters corresponding to each calculation, which are respectively: Among them, represents the evaluation parameters before each gradient descent calculation; η represents the learning rate used in the first gradient descent calculation; represents taking the partial derivative of the evaluation parameters; represents the first loss function; θ i represents the evaluation parameters after each first gradient descent calculation; these intermediate evaluation parameters respectively represent the quality evaluation results corresponding to the multi-modal data corresponding to the batch_size fusion vectors used in each calculation under the initial evaluation parameters.

[0047] Further, assume that Y query vectors are divided into L groups according to the set batch_size. Then, batch_size second fusion vectors are obtained in batches from the Y query vectors, and this is done L times, that is, the number of times of the second batch training is L. Further, based on the batch_size intermediate state evaluation parameters and the first loss function obtained by each first gradient descent calculation, the quality evaluation model is successively subjected to second gradient descent calculations using the second fusion vectors obtained in batches to obtain the total function loss corresponding to each calculation: Among them, fi(θi) represents the function loss corresponding to 1 intermediate state evaluation parameter and 1 second fusion vector used in each second gradient descent calculation; represents the sum of the function losses corresponding to the batch_size intermediate state evaluation parameters and the batch_size second fusion vectors used in each second gradient descent calculation respectively; the total function loss corresponding to each second gradient descent calculation represents the quality evaluation result corresponding to the batch_size second fusion vectors used in each second gradient descent calculation under their respective corresponding intermediate state evaluation parameters.

[0048] Further, in the case of obtaining the L total function losses corresponding to the second gradient descent calculations, the batch_size intermediate state evaluation parameters corresponding to the smallest total function loss among these L total function losses can be determined, and these batch_size intermediate state evaluation parameters are used as the general evaluation parameters corresponding to the quality evaluation model. Further optionally, in order to improve the accuracy of the general evaluation parameters, the quality evaluation model can be subjected to a third gradient descent calculation based on the above general evaluation parameters and the third loss function to obtain updated general evaluation parameters: Among them, μ represents the initial evaluation parameter corresponding before the third gradient descent calculation; λ represents the learning rate used in the third gradient descent calculation; represents taking the partial derivative of the evaluation parameter; represents the third loss function; ω represents the general evaluation parameter with higher longitude obtained after the third gradient descent calculation.

[0049] It should be noted that in the embodiments of the present application, the specific types of the first loss function, the second loss function, and the third loss function are not limited, including but not limited to any one of mean squared error loss, mean absolute error loss, cross entropy loss, and hinge loss; correspondingly, the embodiments of the present application also do not limit the relationship between the first loss function, the second loss function, and the third loss function. They can be the same or different, and can be specifically determined according to actual needs.

[0050] Based on the above, when the general evaluation parameters corresponding to the quality evaluation model are obtained, the quality evaluation model and its corresponding general evaluation parameters can be used to evaluate the instructions of the target multi-modal data to predict the compliance of the target multi-modal data. Optionally, before performing quality evaluation on the target multi-modal data, the target multi-modal data in the target scenario can be obtained and input into the quality evaluation model; inside the quality evaluation model, the quality of the target multi-modal data can be evaluated according to the general evaluation parameters, and the quality evaluation result represents the compliance of the target multi-modal data. For example, if you want to evaluate the multi-modal data corresponding to the recruitment scenario, you can obtain the online data corresponding to the recruitment scenario and input the obtained online data as the target multi-modal data corresponding to the recruitment scenario into the quality evaluation model to predict whether the obtained online data meets the requirements of the recruitment scenario according to the quality evaluation result output by the quality evaluation model; for example, the output quality evaluation result is the ratio corresponding to the fusion vector of the target multi-modal data after binary classification. If the ratio of the vector value of 1 is greater than 50%, it is determined that the target multi-modal data meets the requirements of the recruitment scenario. If the ratio of the vector value of 0 is greater than 50%, it is determined that the target multi-modal data does not meet the requirements of the recruitment scenario.

[0051] It should be noted that the quality evaluation model provided in the embodiments of the present application is trained based on multi-modal data in multiple scenarios, including the target scenario corresponding to the target multi-modal data to be evaluated. The quality evaluation model trained in this way is more universal; further, it should be noted that since the amount of data obtained online is relatively small, in order to improve the accuracy of quality evaluation of online data, before performing quality evaluation on online data, a batch of small sample data can also be used to fine-tune the general evaluation parameters corresponding to the quality evaluation model, so as to perform quality evaluation on the target multi-modal data based on the fine-tuned general evaluation parameters to obtain a more accurate quality evaluation result.

[0052] In the embodiments of the present application, the specific method for obtaining small-sample data is not limited. In an alternative embodiment, M preset multi-modal sample data can be used as small-sample data to fine-tune the general evaluation parameters. Also, in the case of obtaining T multi-modal data in multiple scenarios, M multi-modal sample data can be randomly sampled from the T multi-modal data as M test tasks to fine-tune the general evaluation parameters. Here, M is a positive integer greater than 1, and (N + M) is less than or equal to T. Further, in the case of obtaining M multi-modal sample data, the M multi-modal sample data can also be vectorized and fused to obtain M fused vectors, and the activation function is used to perform binary classification on the M fused vectors, and the general evaluation parameters are fine-tuned through the M fused vectors after binary classification. Further optionally, to ensure that the quality evaluation model can perform targeted quality evaluation on the target multi-modal data in the target scenario after fine-tuning, the M multi-modal sample data obtained can include the target multi-modal data, or other multi-modal data in the target scenario, so as to improve the adaptability of the quality evaluation parameters to the target scenario.

[0053] In the embodiments of the present application, the specific method for fine-tuning the quality evaluation model with small-sample data is not limited. Optionally, the M fused vectors obtained by vectorizing and fusing the M multi-modal sample data can be divided into P support vectors and Q query vectors; where P is less than N. Taking these P support vectors as small-sample data, before inputting the target multi-modal data into the quality evaluation model, these P support vectors can be used to perform the third batch training on the quality evaluation model to fine-tune the general evaluation parameters. In the embodiments of the present application, the specific method for the third batch training is not limited, and it can be the same as or different from the above first and second batch training methods, and the specific method can be determined according to actual needs.

[0054] Further, after performing small-sample fine-tuning on the general evaluation model, the quality of the target multi-modal data can be evaluated based on the fine-tuned general evaluation parameters, and the compliance of the target multi-modal data can be predicted. Optionally, when evaluating the quality of the target multi-modal data, inside the quality evaluation model, the target multi-modal data can be vectorized and fused to obtain the fused vector corresponding to the target multi-modal data, and the activation function is used to perform binary classification on the fused vector corresponding to the target multi-modal data. Further, based on the fine-tuned general evaluation parameters and the Q query vectors, the fused vectors after binary classification are predicted to obtain the ratio of the two types of fused vectors corresponding to the target multi-modal data, which is used as the quality evaluation result and output.

[0055] In the embodiments of the present application, the specific manner of adjusting the target multimodal data according to the quality assessment result is not limited. Optionally, when the quality assessment result corresponding to the target multimodal data is obtained, the quality assessment result can be provided to the service terminal responsible for modifying the target multimodal data, so that the service terminal can make corresponding adjustments to the target multimodal data according to the quality assessment result. For example, in the recruitment scenario, when the quality assessment result corresponding to the recruitment data is obtained, the quality assessment result can be provided to the corresponding review end. The review end can determine that there are problems with the recruitment data according to the quality assessment result, and verify and adjust the recruitment data to make the recruitment data meet the requirements of the recruitment scenario. For example, when it is verified that there are problem posts such as posting in multiple cities and posting multiple posts in a single city at the same time, the problem posts can be adjusted to meet the recruitment requirements.

[0056] It should be noted that since the quality assessment model provided in the embodiments of the present application is trained based on multimodal data in multiple scenarios, in actual use, the scenario corresponding to the multimodal data to be quality-assessed is not limited. The above embodiments illustrate the process of quality assessment of multimodal data by taking the recruitment scenario as an example. According to specific requirements, the above process of quality assessment can also be applied to scenarios such as house rental / sale scenarios, vehicle rental / sale scenarios, e-commerce sales and service scenarios, etc., and can be flexibly used according to actual needs. For the specific process of using the above model training method in different scenarios, reference can be made to the above embodiments, which will not be elaborated here.

[0057] Based on the above, the embodiments of the present application further provide a method for using a quality assessment model, Figure 1b which is a flowchart of the method for using the quality assessment model, as Figure 1b shown. The method includes:

[0058] S1b. Obtain target multimodal data in a target scenario, and input the target multimodal data into the quality assessment model;

[0059] S2b. Inside the quality assessment model, perform quality assessment on the target multimodal data according to general assessment parameters, and the quality assessment result represents the quality of the target multimodal data.

[0060] Optionally, when performing quality assessment on the target multimodal data according to general assessment parameters, inside the quality assessment model, the target multimodal data can be vectorized and fused to obtain a fusion vector corresponding to the target multimodal data, and the activation function is used to perform binary classification on the fusion vector corresponding to the target multimodal data; furthermore, according to the evaluation model parameters, the fused vector after binary classification is predicted to obtain the ratio of the two types of fused vectors as the quality assessment result and output.

[0061] It should be noted that the specific process of the user of the quality assessment model can be referred to the above embodiments and will not be elaborated here. In addition, the embodiments of the present application do not limit the type of activation function used for binary classification of the fusion vector. For example, it can be a Sigmoid function, a softmax function or other functions, as long as the activation function that meets the actual needs is applicable to the embodiments of the present application.

[0062] Figure 1c FIG. is a schematic diagram of the process of using the method for using the quality assessment model provided by the embodiments of the present application to perform quality assessment on target multi-modal data. The following combines Figure 1c to briefly describe the process of quality assessment. Among them, the target multi-modal data includes data of three modalities: text data, behavior data, and auxiliary data. The vectorization processing models used are the FastText model, the Line model, and the One-Hot model respectively. The vector fusion method used is the concatenate method, and the activation function used is SoftMax. As Figure 1c shown, in the case of obtaining the target multi-modal data, the FastText model, the Line model, and the One-Hot model can be used to perform vectorization processing on the text data, behavior data, and auxiliary data in the target multi-modal data respectively to obtain corresponding text vectors, behavior vectors, and auxiliary vectors. Further, the vectors of the three modalities are fused by the concatenate method to obtain the fusion vector corresponding to the target multi-modal data, and the SoftMax function is used to perform normalization processing on the fusion vector to perform binary classification on the fusion vector. Finally, the binary-classified fusion vector is input into the quality assessment model to perform quality assessment on the target multi-modal data to obtain the corresponding quality assessment result.

[0063] The quality assessment model provided by the embodiments of the present application is trained based on the vectors corresponding to the data of different modalities in the multi-modal data. Therefore, this quality assessment model can be used to perform quality assessment on the target multi-modal data. Moreover, before using the vectors of multiple modalities for model training, the vectors of multiple modalities are also fused and the fused vectors are binary-classified, and the model is trained based on the binary-classified fusion vectors. In this way, when using the quality assessment model to perform quality assessment on the target multi-modal data, it is not only applicable to performing quality assessment on the data of each modality in the target multi-modal data, but also can comprehensively perform quality assessment on the target multi-modal data in combination with the correlation between different modality data in the target multi-modal data, and the obtained quality assessment result is more accurate. In addition, the quality assessment model provided by the embodiments of the present application is trained based on multi-modal data in multiple scenarios and is more universal in use, and there is no limitation on the scenario corresponding to the target multi-modal data to be evaluated.

[0064] It should be noted that for the specific implementation methods of the steps of the above method, reference can be made to the descriptions of the corresponding parts in the above system embodiments, which will not be elaborated here. The execution subject of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps S1a to S5a can be device A; for another example, the execution subject of step S1a can be device A, and the execution subjects of steps S2a to S5a can be device B; and so on.

[0065] In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as S1a, S1b, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0066] Based on the above, the embodiments of the present application further provide a device for training a quality evaluation model for multimodal data, Figure 2 which is a schematic structural diagram of the device for training the quality evaluation model, as Figure 2 shown, the device for training the quality evaluation model includes: a processor 21 and a memory 22 storing a computer program; wherein, the processor 21 and the memory 22 can be one or more.

[0067] The memory 22 is mainly used for storing computer programs, which can be executed by the processor 21, causing the processor 21 to control the device for training the quality evaluation model to implement corresponding functions, complete corresponding actions or tasks. In addition to storing computer programs, the memory 22 can also be configured to store various other data to support operations on the device for training the quality evaluation model. Examples of these data include instructions for any application program or method operating on the device for training the quality evaluation model.

[0068] The memory 22 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc.

[0069] In the embodiments of the present application, the implementation form of the processor 21 is not limited. For example, it may be, but is not limited to, a CPU, a GPU, an MCU, etc. The processor 21 can be regarded as the control system of the quality evaluation model training device, and can be used to execute the computer program stored in the memory 22 to control the quality evaluation model training device to implement corresponding functions, complete corresponding actions or tasks. It should be noted that according to the implementation form of the quality evaluation model training device and the different scenarios it is in, the functions to be implemented, the actions or tasks to be completed will be different; correspondingly, the computer programs stored in the memory 22 will also be different, and the processor 21 executing different computer programs can control the quality evaluation model training device to implement different functions, complete different actions or tasks.

[0070] In some alternative embodiments, as Figure 2 shown, the quality evaluation model training device may further include: other components such as a display 23, a power supply component 24, and a communication component 25. Figure 2 Only some components are schematically shown, which does not mean that the quality evaluation model training device only includes Figure 2 the components shown. For different application requirements, the quality evaluation model training device may further include other components. For example, in the case of voice interaction requirements, as Figure 2 shown, the quality evaluation model training device may further include an audio component 26. Regarding the components that the quality evaluation model training device may include, it specifically depends on the product form of the quality evaluation model training device and is not limited herein.

[0071] In the embodiments of the present application, when the processor executes the computer program in the memory, it is used for: obtaining N multi-modal sample data corresponding to N training tasks, and each multi-modal sample data includes at least two types of modal data; performing vectorization processing and fusion processing on the N multi-modal sample data to obtain N fusion vectors, and using an activation function to perform binary classification on the N fusion vectors; based on the initialized evaluation parameters and the N fusion vectors after binary classification, performing the first batch training on the quality evaluation model to obtain multiple intermediate state evaluation parameters corresponding to each batch of training; based on the multiple intermediate state evaluation parameters and the N fusion vectors after binary classification, performing the second batch training on the quality evaluation model to obtain the total function loss corresponding to each batch of training; determining the general evaluation parameters corresponding to the quality evaluation model according to the multiple total function losses obtained from the second batch training.

[0072] In an alternative embodiment, when the processor 21 obtains N multi-modal sample data corresponding to N training tasks, it is used for: obtaining T multi-modal data in multiple scenarios, and randomly sampling N multi-modal sample data from the T multi-modal data as N training tasks; where N is a positive integer greater than 1 and N is less than or equal to T.

[0073] In an alternative embodiment, when the processor 21 performs vectorization processing on N multimodal sample data, it is used to: perform vector calculation on the text data in the N multimodal sample data by using a word vector calculation method to obtain word vectors in the N multimodal sample data; perform vector processing on the behavior data in the N multimodal sample data by using a graph neural network to obtain behavior vectors in the N multimodal sample data; perform vector processing on the auxiliary data in the N multimodal sample data by using a one-hot encoding method to obtain encoded vectors in the N multimodal sample data.

[0074] In an alternative embodiment, when the processor 21 performs fusion processing on N multimodal sample data to obtain N fusion vectors, it is used to: for the N vectors obtained by vectorizing the N multimodal sample data, take every two vectors as a group, calculate the outer product of every two vectors to obtain intermediate vectors; perform a flattening process on the intermediate vectors to obtain the fusion vectors corresponding to the N multimodal sample data respectively.

[0075] In an alternative embodiment, among the N fusion vectors, there are X support vectors and Y query vectors; when the processor 21 performs the first batch training on the quality evaluation model based on the initialized evaluation parameters and the N fusion vectors after binary classification to obtain multiple intermediate state evaluation parameters corresponding to each batch of training, it is used to: according to the set number of samples per time, batch obtain the corresponding number of first fusion vectors from the X support vectors, and based on the initialized evaluation parameters and the first loss function, use the batch-obtained first fusion vectors to perform the first gradient descent calculation on the quality evaluation model in sequence to obtain the intermediate state evaluation parameters corresponding to each batch of calculations.

[0076] In an alternative embodiment, when the processor 21 performs the second batch training on the quality evaluation model based on multiple intermediate state evaluation parameters and the N fusion vectors after binary classification to obtain the total function loss corresponding to each batch of training, it is used to: according to the set number of samples per time, batch obtain the corresponding number of second fusion vectors from the Y query vectors; based on the number of intermediate state evaluation parameters of the single-sample quantity obtained each time and the second loss function, use the batch-obtained second fusion vectors to perform the second gradient descent calculation on the quality evaluation model in sequence to obtain the total function loss corresponding to each batch of calculations; where the total function loss is the sum of the function losses of the second fusion vectors corresponding to each second gradient descent calculation under the intermediate state evaluation parameters respectively used by them.

[0077] In an alternative embodiment, when the processor 21 determines the general evaluation parameters corresponding to the quality evaluation model according to the total function losses obtained from the second batch training, it is used to: determine the intermediate state evaluation parameters corresponding to the smallest total function loss among the total function losses obtained from the second batch training; use the intermediate state evaluation parameters corresponding to the smallest total function loss as the general evaluation parameters corresponding to the quality evaluation model.

[0078] In an alternative embodiment, the processor 21 is further configured to: obtain target multimodal data in a target scenario, and input the target multimodal data into a quality evaluation model; inside the quality evaluation model, perform quality evaluation on the target multimodal data according to general evaluation parameters, and the quality evaluation result represents the quality of the target multimodal data.

[0079] In an alternative embodiment, when the processor 21 obtains T multimodal data in multiple scenarios, it is further configured to: randomly sample M multimodal sample data from the T multimodal data as M test tasks, where M is a positive integer greater than 1; perform vectorization processing and fusion processing on the M multimodal sample data to obtain M fusion vectors, and perform binary classification on the M fusion vectors using an activation function; where (N + M) is less than or equal to T.

[0080] In an alternative embodiment, the M multimodal sample data includes the target multimodal data, and the M fusion vectors obtained after the M multimodal sample data are subjected to vectorization processing and fusion processing include P support vectors and Q query vectors; where P is less than N; before inputting the target multimodal data into the quality evaluation model, the processor 21 is further configured to: perform third-batch training on the quality evaluation model using the P support vectors, and fine-tune the general evaluation parameters.

[0081] In an alternative embodiment, inside the quality evaluation model, when the processor 21 performs quality evaluation on the target multimodal data according to general evaluation parameters, it is configured to: perform vectorization processing and fusion processing on the target multimodal data inside the quality evaluation model to obtain a fusion vector corresponding to the target multimodal data; perform binary classification on the fusion vector corresponding to the target multimodal data using an activation function; predict the binary-classified fusion vector based on the fine-tuned general evaluation parameters and the Q query vectors, obtain the ratio of the two types of fusion vectors corresponding thereto as the quality evaluation result, and output it.

[0082] It should be noted that for the specific functions of the processor in the above quality evaluation model training device, reference may be made to the above method embodiments, which will not be elaborated here.

[0083] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps that can be executed by the quality evaluation model training device in the above method embodiments.

[0084] Based on the above, an embodiment of the present application further provides a device for using a quality evaluation model for multimodal data, and its corresponding structure is similar to Figure 2 the structure shown, and reference may be made to Figure 2. In this embodiment, the devices using the quality assessment model include: a processor and a memory storing a computer program; wherein, the processor and the memory can be one or more.

[0085] The memory is mainly used to store computer programs, which can be executed by the processor, causing the processor to control the devices using the quality assessment model to implement corresponding functions, complete corresponding actions or tasks. In addition to storing computer programs, the memory can also be configured to store various other data to support operations on the devices using the quality assessment model. Examples of such data include instructions for any application program or method for operating on the devices using the quality assessment model.

[0086] The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0087] In the embodiments of this application, the implementation form of the processor is not limited. For example, it can be but is not limited to a CPU, a GPU or an MCU, etc. The processor can be regarded as the control system of the devices using the quality assessment model and can be used to execute the computer programs stored in the memory to control the devices using the quality assessment model to implement corresponding functions, complete corresponding actions or tasks. It should be noted that depending on the implementation form of the devices using the quality assessment model and the different scenarios they are in, the functions to be implemented, the actions or tasks to be completed will be different; correspondingly, the computer programs stored in the memory will also be different, and the processor executing different computer programs can control the devices using the quality assessment model to implement different functions, complete different actions or tasks.

[0088] In some alternative embodiments, the devices using the quality assessment model may further include: other components such as a display, a power supply component, and a communication component. These are only schematically given some components, and it does not mean that the devices using the quality assessment model only include these components. For different application requirements, the devices using the quality assessment model may further include other components. For example, in the case of voice interaction requirements, the devices using the quality assessment model may further include an audio component. Regarding the components that the devices using the quality assessment model may include, it can be determined specifically according to the product form of the devices using the quality assessment model and is not limited here.

[0089] In an embodiment of the present application, when a processor executes a computer program in a memory, it is used to: obtain target multimodal data in a target scenario, and input the target multimodal data into a quality evaluation model; inside the quality evaluation model, perform quality evaluation on the target multimodal data according to general evaluation parameters, and the quality evaluation result represents the quality of the target multimodal data.

[0090] In an optional embodiment, inside the quality evaluation model, when the processor performs quality evaluation on the target multimodal data according to general evaluation parameters, it is used to: inside the quality evaluation model, perform vectorization processing and fusion processing on the target multimodal data to obtain a fusion vector corresponding to the target multimodal data; and use an activation function to perform binary classification on the fusion vector corresponding to the target multimodal data; according to the evaluation model parameters, perform prediction on the binary-classified fusion vector to obtain the ratio of the two types of corresponding fusion vectors and output it as the quality evaluation result.

[0091] It should be noted that for the specific functions of the processor in the above quality evaluation model using device, reference may be made to the above method embodiments, which will not be elaborated here.

[0092] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executable by the quality evaluation model using device in the above method embodiments.

[0093] The communication component in the above embodiments is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on technologies such as Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0094] The display in the above embodiments includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensor can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.

[0095] The power supply component in the above embodiments provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0096] The audio component in the above embodiments may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals may be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0097] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0098] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0099] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or multiple processes in the flowchart and / or one block or multiple blocks in the block diagram.

[0101] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0102] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0103] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0104] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0105] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for training a quality assessment model for multimodal data, characterized in that, Including: Obtain N multi-modal sample data corresponding to N training tasks, each multi-modal sample data including data of at least two modalities, where N is a positive integer greater than 1; Perform vectorization processing and fusion processing on the N multi-modal sample data to obtain N fusion vectors, and use an activation function to perform binary classification on the N fusion vectors; wherein, the N fusion vectors include X support vectors and Y query vectors; Based on the initialized evaluation parameters and the N fusion vectors after binary classification, perform the first batch training on the quality evaluation model to obtain multiple intermediate state evaluation parameters corresponding to each batch of training; Based on the multiple intermediate state evaluation parameters and the N fusion vectors after binary classification, perform the second batch training on the quality evaluation model to obtain the total function loss corresponding to each batch of training; Determine the general evaluation parameters corresponding to the quality evaluation model according to the multiple total function losses obtained from the second batch training; Among them, based on the initialized evaluation parameters and the N fusion vectors after binary classification, performing the first batch training on the quality evaluation model to obtain multiple intermediate state evaluation parameters corresponding to each batch of training includes: Obtain the corresponding number of first fusion vectors in batches from the X support vectors according to the set number of samples per time; based on the initialized evaluation parameters and the first loss function, use the first fusion vectors obtained in batches to perform the first gradient descent calculation on the quality evaluation model in turn to obtain the corresponding number of intermediate state evaluation parameters for each calculation; Among them, based on the multiple intermediate state evaluation parameters and the N fusion vectors after binary classification, performing the second batch training on the quality evaluation model to obtain the total function loss corresponding to each batch of training includes: Obtain the corresponding number of second fusion vectors in batches from the Y query vectors according to the set number of samples per time; based on the corresponding number of intermediate state evaluation parameters obtained each time and the second loss function, use the second fusion vectors obtained in batches to perform the second gradient descent calculation on the quality evaluation model in turn to obtain the total function loss corresponding to each calculation; wherein, the total function loss is the sum of the function losses corresponding to the second fusion vectors used in each second gradient descent calculation under their respectively used intermediate state evaluation parameters; Among them, determining the general evaluation parameters corresponding to the quality evaluation model according to the multiple total function losses obtained from the second batch training includes: Determine the intermediate state evaluation parameters corresponding to the smallest total function loss among the multiple total function losses obtained from the second batch training; use the intermediate state evaluation parameters corresponding to the smallest total function loss as the general evaluation parameters corresponding to the quality evaluation model.

2. The method according to claim 1, wherein Obtain N multi-modal sample data corresponding to N training tasks, including: Obtain T multi-modal data in multiple scenarios, and randomly sample N multi-modal sample data from the T multi-modal data as N training tasks; where N is less than or equal to T.

3. The method according to claim 1, characterized in that, Perform vectorization processing on the N multi-modal sample data, including: Perform vector calculation on the text data in the N multimodal sample data using a word vector calculation method to obtain word vectors in the N multimodal sample data; Perform vector processing on the behavior data in the N multimodal sample data using a graph neural network to obtain behavior vectors in the N multimodal sample data; Perform vector processing on the auxiliary data in the N multimodal sample data using a one-hot encoding method to obtain encoded vectors in the N multimodal sample data.

4. The method according to claim 3, wherein Perform fusion processing on the N multimodal sample data to obtain the N fusion vectors, including: For the N vectors obtained by vectorizing the N multimodal sample data, taking every two vectors as a group, calculate the outer product of the every two vectors to obtain intermediate vectors; Perform a flattening process on the intermediate vectors to obtain the fusion vectors corresponding to the N multimodal sample data respectively.

5. The method according to claim 1, wherein Further includes: Obtain target multimodal data in a target scenario, and input the target multimodal data into the quality assessment model; Inside the quality assessment model, perform quality assessment on the target multimodal data according to the general assessment parameters, and the quality assessment result represents the quality of the target multimodal data.

6. The method according to claim 5, wherein In the case of obtaining T multimodal data in multiple scenarios, further includes: Randomly sample M multimodal sample data from the T multimodal data as M test tasks, where M is a positive integer greater than 1, and (N + M) is less than or equal to T; Perform vectorization processing and fusion processing on the M multimodal sample data to obtain M fusion vectors, and perform binary classification on the M fusion vectors using an activation function.

7. The method according to claim 6, characterized in that, The M multimodal sample data includes the target multimodal data, and the M fusion vectors obtained after vectorization processing and fusion processing of the M multimodal sample data include P support vectors and Q query vectors, where P is less than N; Before inputting the target multimodal data into the quality assessment model, further includes: Use the P support vectors to perform third batch training on the quality assessment model and fine-tune the general assessment parameters.

8. The method according to claim 7, further characterized in that, Inside the quality assessment model, performing quality assessment on the target multimodal data according to the general assessment parameters includes: Inside the quality assessment model, perform vectorization processing and fusion processing on the target multimodal data to obtain the fusion vector corresponding to the target multimodal data; Perform binary classification on the fusion vector corresponding to the target multimodal data using an activation function; Based on the fine-tuned general assessment parameters and the Q query vectors, predict the binary-classified fusion vectors to obtain the ratio of the two types of fusion vectors and output it as the quality assessment result.

9. A method for using a quality assessment model, characterized in that, Includes: Obtain target multimodal data in a target scenario, and input the target multimodal data into a quality assessment model; wherein, the input of the target multimodal data into the quality assessment model is trained based on the quality assessment model training method for multimodal data according to any one of claims 1-8. Inside the quality assessment model, the quality of the target multi-modal data is evaluated according to general evaluation parameters, and the quality assessment result represents the quality of the target multi-modal data.

10. The method according to claim 9, wherein Inside the quality assessment model, evaluating the quality of the target multi-modal data according to general evaluation parameters includes: Inside the quality assessment model, performing vectorization processing and fusion processing on the target multi-modal data to obtain a fusion vector corresponding to the target multi-modal data, and using an activation function to perform binary classification on the fusion vector corresponding to the target multi-modal data; According to the evaluation model parameters, predicting the fusion vector after binary classification to obtain the ratio of the two types of fusion vectors corresponding thereto and outputting it as the quality assessment result.

11. A training device for a quality assessment model of multimodal data, characterized in that, including: A memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps in the method according to any one of claims 1-8.

12. An apparatus for using a quality assessment model, characterized in that, including: A memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the steps in the method according to any one of claims 9-10.

13. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the processor is caused to implement the steps in the method according to any one of claims 1-8 or 9-10.

Citation Information

Patent Citations

  • Illegal behavior detection method and device and electronic equipment

    CN113743522A

  • Distributed training of multi-modal machine learning models

    WO2021097494A2