Scoring method and device, equipment, storage medium and program product

By encoding and fusion of multiple modal data of the evaluation object, and combining the functional types of the evaluation scenarios, the evaluation object is scored, which solves the problem that the existing system cannot adapt to multi-dimensional assessment, and realizes a more flexible and accurate scoring method.

CN120235508APending Publication Date: 2025-07-01IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455036.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing intelligent scoring system can only be evaluated based on single-modal data and cannot adapt to evaluation scenarios that perform multi-dimensional assessment.

Method used

By obtaining data of multiple modals of the evaluation object (such as text, audio, and images), the data of each modal are encoded, and embedded features are obtained, and they are fused in the current evaluation scenario to obtain joint embedded features. Then, the joint embedding features are processed using the processing method corresponding to the functional type of the current evaluation scenario to obtain the score of the evaluation object.

Benefits of technology

Scoring based on multimodal data is realized, suitable for evaluation scenarios for multi-dimensional assessment, and improves the adaptability and accuracy of the scoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235508A_ABST
    Figure CN120235508A_ABST
Patent Text Reader

Abstract

The invention discloses a scoring method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining the data of a plurality of modes of an evaluation object; the plurality of modalities comprise at least two of the following modalities: text, audio and image; encoding the obtained data of each mode to obtain an embedding feature of each mode; if the current evaluation scene is a target evaluation scene, fusing the embedding features of each modal according to a mode corresponding to the current evaluation scene to obtain a joint embedding feature; and processing the joint embedded feature by using a processing mode corresponding to the function type of the current evaluation scene to obtain a score of the evaluation object. According to the method and the device, scoring based on multi-modal data is realized, and the method and the device are suitable for evaluation scenes of multi-dimensional assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a scoring method, device, equipment, storage medium, and program product. Background Art

[0002] With the development of artificial intelligence, more and more evaluation scenarios adopt intelligent scoring systems. However, the current intelligent scoring systems can only evaluate based on single-modal data (such as text or voice), and cannot adapt to evaluation scenarios that require multi-dimensional assessment. Summary of the Invention

[0003] In view of the above problems, this application provides a scoring method, device, equipment, storage medium, and program product to adapt to evaluation scenarios that require multi-dimensional assessment. The specific solutions are as follows:

[0004] The first aspect of this application provides a scoring method, including:

[0005] Obtain data of multiple modalities of the evaluation object; the multiple modalities include at least two of the following modalities: text, audio, and image;

[0006] Encode the data of each obtained modality respectively to obtain the embedding features of each modality;

[0007] If the current evaluation scenario is the target evaluation scenario, fuse the embedding features of each modality in the manner corresponding to the current evaluation scenario to obtain a joint embedding feature;

[0008] Process the joint embedding feature using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object.

[0009] In a possible implementation, the fusing the embedding features of each modality in the manner corresponding to the current evaluation scenario includes:

[0010] Fuse the embedding features of each modality based on the weights of the embedding features of each modality;

[0011] Wherein, the weights of the embedding features of each modality correspond to the current evaluation scenario; the weights of the embedding features of different modalities are the same or different.

[0012] In a possible implementation, the fusing the embedding features of each modality based on the weights of the embedding features of each modality includes:

[0013] Based on the weights of the embedding features of each modality, perform weighted splicing on the embedding features of each modality to obtain a first initial fusion feature;

[0014] Encode the embedded features of each modality based on the attention mechanism to obtain the hidden layer features of each modality; concatenate the hidden layer features of each modality to obtain the concatenated hidden layer features; encode the concatenated hidden layer features based on the attention mechanism to obtain the second initial fusion feature;

[0015] Concatenate the first initial fusion feature and the second initial fusion feature to obtain the joint embedded feature.

[0016] In a possible implementation, fusing the embedded features of each modality based on the weights of the embedded features of each modality includes:

[0017] Concatenate the embedded features of each modality to obtain the third initial fusion feature;

[0018] Encode the embedded features of each modality based on the attention mechanism to obtain the hidden layer features of each modality; based on the weights of the embedded features of each modality, perform weighted concatenation on the hidden layer features of each modality to obtain the concatenated hidden layer features; encode the concatenated hidden layer features based on the attention mechanism to obtain the fourth initial fusion feature;

[0019] Concatenate the third initial fusion feature and the fourth initial fusion feature to obtain the joint embedded feature.

[0020] In a possible implementation, fusing the embedded features of each modality based on the weights of the embedded features of each modality includes:

[0021] Based on the weights of the embedded features of each modality, perform weighted concatenation on the embedded features of each modality to obtain the joint embedded feature.

[0022] In a possible implementation, fusing the embedded features of each modality based on the weights of the embedded features of each modality includes:

[0023] Encode the embedded features of each modality based on the attention mechanism to obtain the hidden layer features of each modality;

[0024] Based on the weights of the embedded features of each modality, perform weighted concatenation on the hidden layer features of each modality to obtain the concatenated hidden layer features;

[0025] Encode the concatenated hidden layer features based on the attention mechanism to obtain the joint embedded feature.

[0026] In a possible implementation, it further includes:

[0027] If the current evaluation scenario is a non-target evaluation scenario, for each non-image modality, the embedded features of the non-image modality are processed using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object in the non-image modality;

[0028] Based on the embedded features of the image modality, determine whether the evaluation object has cheated;

[0029] If the evaluation object has cheated, punish the scores of the evaluation object in each non-image modality to obtain the final scores of each non-image modality.

[0030] In a possible implementation, the processing of the joint embedded features using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object includes:

[0031] If the function type of the current evaluation scenario is a pass-type evaluation, perform score prediction based on the joint embedded features to obtain the score of the evaluation object;

[0032] If the function type of the current evaluation scenario is a selection-type evaluation, perform grading prediction based on the joint embedded features to obtain the score grade to which the evaluation object belongs; perform score prediction based on the score grade and the joint embedded features to obtain the score of the evaluation object.

[0033] The second aspect of the present application provides a scoring device, including:

[0034] A multi-modal data acquisition module, configured to acquire data of multiple modalities of an evaluation object; the multiple modalities include at least two of the following modalities: text, audio, image;

[0035] An encoding module, configured to encode the data of each acquired modality respectively to obtain the embedded features of each modality;

[0036] A fusion module, configured to, if the current evaluation scenario is a target evaluation scenario, fuse the embedded features of each modality in the manner corresponding to the current evaluation scenario to obtain joint embedded features;

[0037] A scoring module, configured to process the joint embedded features using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object.

[0038] The third aspect of the present application provides a computer program product, including computer-readable instructions, which when running on an electronic device, cause the electronic device to implement the scoring method of the first aspect or any implementation manner of the first aspect.

[0039] A fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0040] The memory is used to store a computer program;

[0041] The processor is used to execute the computer program so that the electronic device can implement the scoring method of the above first aspect or any implementation manner of the first aspect.

[0042] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the scoring method of the above first aspect or any implementation manner of the first aspect.

[0043] By means of the above technical solutions, the scoring method, device, equipment, storage medium and program product provided by the present application obtain data of multiple modalities of the evaluation object; the multiple modalities include at least two of the following modalities: text, audio, and image; each modality of the obtained data is respectively encoded to obtain embedding features of each modality; if the current evaluation scenario is the target evaluation scenario, the embedding features of each modality are fused in the manner corresponding to the current evaluation scenario to obtain a joint embedding feature; the joint embedding feature is processed by the processing manner corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object. The scoring scheme of the present application can score the evaluation object based on data of multiple modalities of the evaluation object. After obtaining the embedding features of each modality based on the data of each modality, the evaluation scenario is considered when fusing the embedding features of each modality, and the function type of the evaluation scenario is considered when processing the joint embedding feature to obtain the score of the evaluation object, realizing scoring based on multi-modal data and being applicable to evaluation scenarios for multi-dimensional assessment. Description of the Drawings

[0044] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.

[0045] Figure 1 It is a flowchart of an implementation of the scoring method provided by the present application;

[0046] Figure 2 It is a flowchart of an implementation of fusing the embedding features of each modality based on the weights of the embedding features of each modality provided by the present application;

[0047] Figure 3Another implementation flowchart for fusing the embedding features of each modality based on the weights of the embedding features of each modality provided by this application;

[0048] Figure 4 Another implementation flowchart for fusing the embedding features of each modality based on the weights of the embedding features of each modality provided by this application;

[0049] Figure 5 A structural schematic diagram of a scoring model provided by this application;

[0050] Figure 6 A structural schematic diagram of a scoring device provided by this application;

[0051] Figure 7 A structural schematic diagram of an electronic device provided by this application. Detailed implementation manners

[0052] The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application. The terms used in the embodiments part of this application are only used to explain the specific embodiments of this application, rather than intending to limit this application.

[0053] The embodiments of this application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0054] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0055] Currently, more and more evaluation scenarios assess candidates from multiple dimensions. For example, some evaluation scenarios will conduct written tests and interviews on candidates, and then calculate the comprehensive score based on the written test score and the interview score. However, the current intelligent scoring system can only score papers or only score spoken language, and cannot adapt to the scenarios of multi-dimensional assessment. In order to adapt to the scenarios of multi-dimensional assessment, different scoring systems need to be deployed, which will result in cumbersome setup and debugging work. Therefore, it is necessary to provide an integrated scoring system to meet the evaluation needs of different evaluation scenarios.

[0056] To adapt to the evaluation scenario of multi-dimensional assessment, the solution of this application is proposed.

[0057] As Figure 1 shown, it is a flowchart of an implementation of the scoring method provided by an embodiment of this application, which may include:

[0058] Step S101: Obtain data of multiple modalities of the evaluation object (i.e., examinee).

[0059] The above-mentioned multiple modalities include at least two of the following modalities: text, audio, and image. The image here can be a single image of the evaluation object or a continuous multiple images of the evaluation object (i.e., the video collected for the evaluation object).

[0060] Step S102: Encode the data of each obtained modality respectively to obtain the embedding features of each modality.

[0061] Optionally, for text data, a pre-trained text encoder can be used to encode the text data to obtain the embedding features of the text modality. As an example, the text encoder may include, but is not limited to, a text encoder with a transformer architecture. For example, the text encoding module of GPT-2 can be used as the text encoder. Of course, the text encoder can also adopt a text encoder with other architectures. For example, a long short-term memory network can be used as the text encoder, and this application does not make specific limitations on this.

[0062] For each image, a pre-trained image encoder can be used to encode each image to obtain the embedding features of each image. In the case of obtaining multiple images of the evaluation object, the embedding features of the multiple images can be fused to obtain the embedding features of the image modality. As an example, the embedding features of the multiple images can be summed to obtain the embedding features of the image modality; or, the embedding features of the multiple images can be averaged to obtain the embedding features of the image modality. As an example, the image encoder may include, but is not limited to, an image encoder with a transformer architecture. For example, the image encoder can be a vision transformer (ViT). Of course, the image encoder can also adopt an image encoder with other architectures. For example, a convolutional neural network CNN can be used as the image encoder, and this application does not make specific limitations on this.

[0063] For speech data, a pre-trained speech encoder can be used to encode the speech data to obtain the embedded features in the speech modality. As an example, the speech encoder can include, but is not limited to, a speech encoder with a transformer architecture. For example, the speech encoder can be a Conformer. Of course, the speech encoder can also use speech encoders with other architectures. For example, a hybrid architecture of a convolutional neural network and a transformer can be used. The present application does not make specific limitations on this.

[0064] The embedded features of each modality can be a vector, and the lengths of the embedded features of different modalities can be the same or different.

[0065] Step S103: If the current evaluation scenario is the target evaluation scenario, fuse the embedded features of each modality in the manner corresponding to the current evaluation scenario to obtain the joint embedded features.

[0066] Optionally, the current evaluation scenario can be selected by the user according to the actual evaluation requirements. Based on this, a first interaction interface can be provided. If a preset operation (such as a click operation) for this first interaction interface is detected, in response to the preset operation, a test scenario configuration interface is displayed. The test scenario configuration interface displays a list of test scenarios. The user can select the corresponding test scenario according to the actual evaluation requirements. If a selection operation for any test scenario is obtained, in response to the selection operation, the any test scenario is determined as the current evaluation scenario.

[0067] The user can select the evaluation scenario before the evaluation starts (i.e., before step S101).

[0068] In order to determine whether the current evaluation scenario is the target evaluation scenario, a whitelist of evaluation scenarios can be preset. The whitelist of evaluation scenarios records the test scenarios that require multi-modal joint scoring. When it is necessary to determine whether the current evaluation scenario is the target evaluation scenario, the current evaluation scenario can be compared with the evaluation scenarios in the whitelist of evaluation scenarios. If the current evaluation scenario is in the whitelist of evaluation scenarios, it is determined that the current evaluation scenario is the target evaluation scenario. Otherwise, it is determined that the current evaluation scenario is not the target evaluation scenario.

[0069] Optionally, it can be determined whether the current evaluation scenario is the target evaluation scenario after the user selects the evaluation scenario, or it can also be determined whether the current evaluation scenario is the target evaluation scenario when step S103 needs to be executed.

[0070] In addition, in order to adapt to different evaluation scenarios, the present application sets a corresponding embedded feature fusion method for each evaluation scenario. According to actual needs, the embedded feature fusion methods corresponding to different evaluation scenarios may be the same or different.

[0071] Step S104: Process the joint embedding features using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object.

[0072] To adapt to evaluation scenarios of different function types, when predicting scores using joint embedding features, the processing methods (ways of predicting scores) for joint embedding features are different.

[0073] As an example, the function types of evaluation can include but are not limited to the following two: pass - type exams and selection - type exams. Among them, pass - type exams and selection - type exams are different in terms of exam purpose, exam content, evaluation criteria, and social significance, etc. Specifically:

[0074] In terms of exam purpose, the main purpose of pass - type exams is to check whether candidates meet specific basic standards or requirements, focusing on "meeting the standard" and not emphasizing discrimination; selection - type exams aim to select outstanding talents who meet specific conditions from numerous candidates, with a strong discrimination function, and pay attention to the competition and comparison among candidates.

[0075] In terms of exam content, the content of pass - type exams is relatively basic and extensive, usually covering the basic knowledge and basic skills of the subject to ensure that candidates have mastered the necessary knowledge and abilities; the content of selection - type exams is often deeper and more complex. In addition to basic knowledge, it also includes the examination of candidates' abilities to comprehensively apply knowledge, innovative thinking, and problem - solving abilities, etc.

[0076] In terms of evaluation criteria, the evaluation criteria of pass - type exams are relatively unified and fixed, usually based on whether the candidates reach the predetermined score line or grade. As long as the candidates meet the basic requirements, they can pass; the evaluation criteria of selection - type exams pay more attention to the overall performance and relative ranking of candidates, and determine whether they are selected according to the position of the candidates' scores among all candidates. Therefore, it has stronger competitiveness.

[0077] In terms of social significance, pass - type exams more reflect the universality and fairness of education, ensuring that most people can meet the basic learning and ability requirements and laying a foundation for further study and career development; selection - type exams highlight the selection and cultivation of talents, selecting students with potential and specialties for higher education institutions, high - quality high schools or specific professional fields to meet the needs of society for different levels and types of talents.

[0078] The scoring method provided by this application can score the evaluation object based on data of multiple modalities of the evaluation object. After obtaining the embedding features of each modality based on the data of each modality, the evaluation scenario is considered when fusing the embedding features of each modality, and the function type of the evaluation scenario is considered when processing the joint embedding features to obtain the score of the evaluation object later, realizing scoring based on multi-modal data and being applicable to evaluation scenarios for multi-dimensional assessment.

[0079] In an optional embodiment, one implementation manner of fusing the embedding features of each modality in the manner corresponding to the current evaluation scenario may be:

[0080] Fuse the embedding features of each modality based on the weights of the embedding features of each modality;

[0081] Among them, the weights of the embedding features of each modality correspond to the current evaluation scenario; the weights of the embedding features of different modalities may be the same or different.

[0082] That is to say, when this application fuses the embedding features of different modalities, weights are set for different modalities. In the same evaluation scenario, the weights corresponding to different modalities may be the same or different; in different evaluation scenarios, the weights corresponding to the same modality may be the same or different.

[0083] In the same evaluation scenario, the weights corresponding to each modality can be the system default values or can be configured by the user in real time according to the actual evaluation requirements. Based on this, a second interaction interface can be provided. If a preset operation (such as a click operation) for this second interaction interface is detected, in response to this preset operation, a weight configuration interface is displayed. The weight configuration interface displays different modalities and weight configuration interfaces for different modalities. The user can configure the corresponding weights through the weight configuration interfaces corresponding to each modality.

[0084] By assigning different weights to different modalities, while ensuring the accuracy of the evaluation results, the interpretability of the evaluation results is made stronger.

[0085] In an optional embodiment, a flowchart of one implementation of fusing the embedding features of each modality based on the weights of the embedding features of each modality is as Figure 2 shown and may include:

[0086] Step S201: Based on the weights of the embedding features of each modality, perform weighted splicing on the embedding features of each modality to obtain a first initial fusion feature.

[0087] Weighted concatenation of the embedding features of each modality means: multiplying the embedding features of each modality by their corresponding weights and then concatenating them in the length direction. For example, suppose there are two modalities of embedding features. One modality's embedding feature E1 is a vector of length L1, and its corresponding weight is α. The other modality's embedding feature E2 is a vector of length L2, and its corresponding weight is β. Then the first initial fusion feature E3 obtained by weighted concatenation of the embedding features of these two modalities is E3 = concat(α×E1, β×E2), where concat() is the concatenation function, indicating that α×E1 and β×E2 are concatenated into a vector. Based on this, the length of E3 is L1 + L2.

[0088] Step S202: Perform encoding on the embedding features of each modality based on the attention mechanism to obtain the hidden layer features of each modality; concatenate the hidden layer features of each modality to obtain the concatenated hidden layer features; perform encoding on the concatenated hidden layer features based on the attention mechanism to obtain the second initial fusion feature.

[0089] The embedding features of each modality can be encoded based on the attention mechanism through a pre-trained fusion network. Optionally, the fusion network can be a multi-modal fusion network based on the Transformer architecture. For example, it can be a network with the FACT (Full-Attention Crossmodal Transformer) architecture.

[0090] As Figure 3 shown, it is a structural example diagram of the fusion network with the FACT architecture provided by the embodiment of the present application for fusing the embedding features of three modalities. In this example, the structures of the encoding modules of each modality and the cross-modal fusion module are the same. Taking the encoding module of the text modality as an example to illustrate the structure of each module, the encoding module of the text modality consists of multiple levels of networks. Each level of network includes: an attention layer (Attention layer), a residual connection and a normalization layer (Add & Normalize layer), and a fully connected feed-forward network layer (Feed Forward layer), etc. The attention layer therein can be a global attention layer (Full Attention layer) or a self-attention layer (self-Attention layer).

[0091] The encoding modules of different modalities of the fusion network are used to encode the embedding features of different modalities to obtain the hidden layer features of each modality; after directly concatenating or weighted concatenating the hidden layer features of each modality to obtain the concatenated hidden layer features, the cross-modal fusion module encodes the concatenated hidden layer features to obtain the second initial fusion feature.

[0092] The advantages of adopting the FACT architecture are as follows: making full use of the advantages of single-stream and multi-stream structures, separately processing data of different modalities, ensuring cross-modal interaction while retaining the independence of each modality.

[0093] Step S203: Concatenate the first initial fusion feature and the second initial fusion feature to obtain a joint embedding feature. In this way, both the intra-modal features and the inter-modal joint features (also known as the joint representation of multi-modal data) are obtained.

[0094] In an optional embodiment, another implementation flowchart of fusing the embedding features of each modality based on the weights of the embedding features of each modality is as Figure 3 shown, and it may include:

[0095] Step S301: Concatenate the embedding features of each modality to obtain a third initial fusion feature.

[0096] Different from the Figure 2 embodiment shown, Figure 2 in the embodiment shown, the embedding features of each modality are weighted and concatenated, while in this embodiment, the embedding features of each modality are directly concatenated, that is, the weights of the embedding features of each modality are the same.

[0097] Step S302: Encode the embedding features of each modality based on the attention mechanism to obtain the hidden layer features of each modality; based on the weights of the embedding features of each modality, weight and concatenate the hidden layer features of each modality to obtain the concatenated hidden layer features; encode the concatenated hidden layer features based on the attention mechanism to obtain a fourth initial fusion feature.

[0098] Different from the Figure 2 embodiment shown, Figure 2 in the embodiment shown, the hidden layer features of each modality are directly concatenated, that is, the weights of the hidden layer features of each modality are the same, while in this embodiment, the hidden layer features of each modality are weighted and concatenated, that is, the weights of the hidden layer features of each modality may be different.

[0099] Step S303: Concatenate the third initial fusion feature and the fourth initial fusion feature to obtain a joint embedding feature.

[0100] In an optional embodiment, an implementation manner of fusing the embedding features of each modality based on the weights of the embedding features of each modality may be:

[0101] Based on the weights of the embedding features of each modality, weight and concatenate the embedding features of each modality to obtain a joint embedding feature.

[0102] That is to say, the result of directly weighting and concatenating the embedding features of each modality can be used as the joint embedding feature.

[0103] In an alternative embodiment, another implementation flowchart of fusing the embedding features of each modality based on the weights of the embedding features of each modality is as follows Figure 4 shown, and may include:

[0104] Step S401: Encode the embedding features of each modality based on the attention mechanism to obtain the hidden layer features of each modality.

[0105] The specific implementation method can refer to the foregoing step S202 and will not be elaborated here.

[0106] Step S402: Based on the weights of the embedding features of each modality, weight and splice the hidden layer features of each modality to obtain the spliced hidden layer features.

[0107] The specific implementation method can refer to the foregoing step S202 and will not be elaborated here.

[0108] Step S403: Encode the spliced hidden layer features based on the attention mechanism to obtain the joint embedding features.

[0109] The specific implementation method can refer to the foregoing step S202 and will not be elaborated here.

[0110] In an alternative embodiment, the scoring method provided by this application may further include:

[0111] If the current evaluation scenario is a non-target evaluation scenario, for each non-image modality, use the processing method corresponding to the function type of the current evaluation scenario to process the embedding features of the non-image modality to obtain the score of the evaluation object in the non-image modality.

[0112] The non-image modalities may include the text modality and the speech modality. In a non-target evaluation scenario, usually the scoring is based on the data of the non-image modalities, and each non-image modality is scored independently without mixed scoring. Therefore, for each non-image modality, directly use the processing method corresponding to the function type of the current evaluation scenario to process the embedding features of the non-image modality to obtain the score of the evaluation object in the non-image modality.

[0113] Judge whether the evaluation object has cheating behavior based on the embedding features of the image modality.

[0114] The embedding features of the image modality do not directly participate in the scoring, but determine whether the evaluation object has cheating behavior.

[0115] If the evaluation object has cheating behavior, punish the scores of the evaluation object in each non-image modality to obtain the final scores of each non-image modality.

[0116] If it is determined that the evaluation object has cheated, it is necessary to penalize the scores of the evaluation object in each non-image modality to reduce the scores of the evaluation object in each non-image modality. Specifically, the scores of each non-modal can be multiplied by a preset coefficient (the preset coefficient is a non-negative number less than 1) to obtain the final scores of each non-image modality.

[0117] If it is determined that the evaluation object has not cheated, there is no need to penalize the scores of the evaluation object in each non-image modality, and the scores of the evaluation object in each non-image modality are the final scores of the evaluation object in each non-image modality.

[0118] In an optional embodiment, an implementation manner of processing the joint embedding features by using the processing method corresponding to the function type of the current evaluation scenario may be:

[0119] If the function type of the current evaluation scenario is a pass-type evaluation, score prediction is performed based on the joint embedding features to obtain the score of the evaluation object (i.e., the comprehensive score).

[0120] That is to say, when the function type of the current evaluation scenario is a pass-type evaluation, the comprehensive score of the evaluation object is directly predicted by using the joint embedding features.

[0121] If the function type of the current evaluation scenario is a selection-type evaluation, tier prediction is performed based on the joint embedding features to obtain the score tier to which the evaluation object belongs; score prediction is performed based on the score tier and the joint embedding features to obtain the score of the evaluation object.

[0122] That is to say, when the function type of the current evaluation scenario is a selection-type evaluation, first predict which score tier the evaluation object belongs to based on the joint embedding features, and then combine the score tier to which the evaluation object belongs and the joint embedding features to predict the comprehensive score of the evaluation object. Since there needs to be a clear distinction between tiers in a selection-type examination, the score tiers of the evaluation object are first divided, and after determining the score tiers, the scores of the evaluation object within the score tier range are predicted based on the joint embedding features and the predicted score tiers, further ensuring the accuracy of score prediction.

[0123] Which function type the current evaluation scenario belongs to can be determined according to the preset corresponding relationship, that is, the corresponding relationship between the evaluation scenario and the function type is preset. When it is necessary to determine the function type to which the current evaluation scenario belongs, the function type to which the current evaluation scenario belongs can be directly determined according to the corresponding relationship between the evaluation scenario and the function type.

[0124] In an optional embodiment, the data of each obtained modality is separately encoded to obtain the embedding features of each modality; if the current evaluation scenario is the target evaluation scenario, the embedding features of each modality are fused in the manner corresponding to the current evaluation scenario to obtain joint embedding features; the process of processing the joint embedding features by using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object can be implemented by a scoring model. As Figure 5 shown, a schematic structural diagram of a scoring model provided by an embodiment of the present application may include:

[0125] A plurality of single-modal encoding units 501_1 to 501_n, a fusion unit 502, and a scoring unit 503;

[0126] Among them, each single-modal encoding unit is used to encode the data of one modality, and different single-modal encoding units are used to encode the data of different modalities. As an example, the scoring model may include: a text encoding unit 501_1, a voice encoding unit 501_2, and an image encoding unit 501_3.

[0127] The text encoding unit 501_1 is used to encode the text data to obtain the embedding features of the text modality (which can be simply referred to as text embedding features). As an example, the text data can be directly input into the text encoding unit 501_1 to obtain the embedding features of the text modality output by the text encoding unit 501_1.

[0128] The voice encoding unit 501_2 is used to encode the voice data to obtain the embedding features of the voice modality (which can be simply referred to as voice embedding features). As an example, the voice can be directly input into the voice encoding unit 501_2 to obtain the embedding features of the voice modality output by the voice encoding unit 501_2; or, the acoustic features of the voice data can be extracted first, such as MFCC features, etc., and the acoustic features are input into the voice encoding unit 501_2 to obtain the embedding features of the voice modality output by the voice encoding unit 501_2.

[0129] The image encoding unit 501_3 is used to encode the image data to obtain the embedding features of the image modality (which can be simply referred to as image embedding features). As an example, the image data can be divided into multiple image blocks, and each image block is sequentially input into the image encoding unit 501_3 to obtain the embedding features of the image modality output by the image encoding unit 501_3.

[0130] The fusion unit 502 is used to fuse the embedding features of the text modality, the embedding features of the voice modality, and the embedding features of the image modality in the manner corresponding to the current evaluation scenario when the current evaluation scenario is the target evaluation scenario to obtain joint embedding features.

[0131] Optionally, any of the foregoing fusion schemes may be adopted to fuse the embedded features of the text modality, the embedded features of the speech modality, and the embedded features of the image modality. When any of the fusion schemes is adopted to fuse the embedded features of the text modality, the embedded features of the speech modality, and the embedded features of the image modality, the weights corresponding to different modalities may be different. In different evaluation scenarios, the weights corresponding to the same modality may be the same or different.

[0132] As an example, in the scores of some enterprise recruitment examinations, the ratio of the written test to the interview is 1:1. Among them, during the interview, the appearance part accounts for 10% of the interview score. Based on this, the weights of text, speech, and image can be allocated according to the following ratio:

[0133] Text: Speech: Image = 10:9:1.

[0134] For example, the weight corresponding to the text modality is 1, the weight corresponding to the speech modality is 0.9, and the weight corresponding to the image modality is 0.1.

[0135] The scoring unit 503 is used to process the joint embedded features by using the processing method corresponding to the function type of the current evaluation scenario, and obtain the score of the evaluation object.

[0136] For an evaluation scenario with a function type of pass / fail evaluation, the scoring unit 503 predicts the score based on the joint embedded features through the first scoring unit, and obtains the score of the evaluation object.

[0137] For an evaluation scenario with a function type of selection evaluation, the scoring unit 503 first predicts the score range based on the joint embedded features through the second scoring unit, and then predicts the score based on the predicted score range and the joint embedded features, and obtains the score of the evaluation object. The predicted score is within the predicted score range.

[0138] In the case where the current evaluation scenario is a non-target evaluation scenario, such as the IELTS test, the TOEFL test, etc., the abilities of the candidates in the four dimensions of listening, speaking, reading, and writing are examined, and the four dimensions are examined independently. Therefore, it is not necessary to fuse the embedded features of the four dimensions. In addition, the examinations of the listening, reading, and writing dimensions are all realized by relying on the text. Therefore, the listening, reading, and writing can all be simplified to the evaluation of the text modality, while "speaking" (i.e., oral expression) is evaluated by relying on speech and can be simplified to the evaluation of the speech modality. That is to say, for non-target evaluation scenarios, the embedded features of each modality are not sent to the fusion unit 502, but are sent to the scoring unit 503 directly for score prediction.

[0139] Optionally, the scoring model of the present application may adopt a network with a Transformer architecture. Each unimodal encoding unit 501_i (i = 1, 2, 3, ……, n) can be processed using a mature architecture in the field. For example, the image encoding unit can adopt vision transformer (ViT), the text encoding unit can adopt the text encoding module of GPT-2, and the voice encoding unit can adopt conformer. The fusion unit 502 can also adopt a pre-trained fusion network, such as the FACT network, etc. The unimodal encoding unit and the fusion unit 502 can adopt networks that have been trained in the field and do not need to be trained again. Only the scoring unit 503 needs to be trained based on the scoring task. Specifically, it can be trained in the following way:

[0140] Input the training samples into the scoring model to obtain the scores corresponding to the training samples output by the scoring model; take the scores output by the scoring model approaching the labels of the training samples as the goal, and update the parameters of the scoring unit 503.

[0141] Among them, each training sample corresponds to an evaluation object, and each training sample includes data of multiple modalities (at least two modalities among text, voice, and image) of its corresponding evaluation object; the label of the training sample is the score of the evaluation object corresponding to the training sample.

[0142] Corresponding to the method embodiment, the present application also provides a scoring device. A schematic structural diagram of the scoring device provided in the embodiment of the present application is as Figure 6 shown, and may include:

[0143] A multi-modal data acquisition module 601, an encoding module 602, a fusion module 603, and a scoring module 604;

[0144] Among them, the multi-modal data acquisition module 601 is used to acquire data of multiple modalities of the evaluation object; the multiple modalities include at least two of the following modalities: text, audio, and image;

[0145] The encoding module 602 is used to encode the data of each acquired modality respectively to obtain the embedding features of each modality;

[0146] The fusion module 603 is used to fuse the embedding features of each modality in the corresponding manner of the current evaluation scenario if the current evaluation scenario is the target evaluation scenario to obtain the joint embedding features;

[0147] The scoring module 604 is used to process the joint embedding features using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object.

[0148] The scoring device provided by the embodiment of the present application can score an evaluation object based on data of multiple modalities of the evaluation object. After obtaining the embedding features of each modality based on the data of each modality, the evaluation scenario is considered when fusing the embedding features of each modality, and the function type of the evaluation scenario is considered when processing the joint embedding features to obtain the score of the evaluation object later, realizing scoring based on multi-modal data and being applicable to the evaluation scenario for multi-dimensional assessment.

[0149] In an optional embodiment, when the fusion module 603 fuses the embedding features of each modality in the manner corresponding to the current evaluation scenario, it is used for:

[0150] Fusing the embedding features of each modality based on the weights of the embedding features of each modality;

[0151] Wherein, the weights of the embedding features of each modality correspond to the current evaluation scenario; the weights of the embedding features of different modalities are the same or different.

[0152] In an optional embodiment, when the fusion module 603 fuses the embedding features of each modality based on the weights of the embedding features of each modality, it is used for:

[0153] Based on the weights of the embedding features of each modality, weighted splicing the embedding features of each modality to obtain a first initial fusion feature;

[0154] Encoding the embedding features of each modality based on the attention mechanism to obtain the hidden layer features of each modality; splicing the hidden layer features of each modality to obtain the spliced hidden layer features; encoding the spliced hidden layer features based on the attention mechanism to obtain a second initial fusion feature;

[0155] Splicing the first initial fusion feature and the second initial fusion feature to obtain the joint embedding feature.

[0156] In an optional embodiment, when the fusion module 603 fuses the embedding features of each modality based on the weights of the embedding features of each modality, it is used for:

[0157] Splicing the embedding features of each modality to obtain a third initial fusion feature;

[0158] Encoding the embedding features of each modality based on the attention mechanism to obtain the hidden layer features of each modality; based on the weights of the embedding features of each modality, weighted splicing the hidden layer features of each modality to obtain the spliced hidden layer features; encoding the spliced hidden layer features based on the attention mechanism to obtain a fourth initial fusion feature;

[0159] Splicing the third initial fusion feature and the fourth initial fusion feature to obtain the joint embedding feature.

[0160] In an alternative embodiment, when the fusion module 603 fuses the embedded features of each modality based on the weights of the embedded features of each modality, it is used for:

[0161] Based on the weights of the embedded features of each modality, the embedded features of each modality are weighted and concatenated to obtain the joint embedded feature.

[0162] In an alternative embodiment, the fusion module 603 fuses the embedded features of each modality based on the weights of the embedded features of each modality, including:

[0163] Encoding the embedded features of each modality based on the attention mechanism to obtain the hidden layer features of each modality;

[0164] Based on the weights of the embedded features of each modality, the hidden layer features of each modality are weighted and concatenated to obtain the concatenated hidden layer features;

[0165] Encoding the concatenated hidden layer features based on the attention mechanism to obtain the joint embedded feature.

[0166] In an alternative embodiment, the scoring module 604 is further used for:

[0167] If the current evaluation scenario is a non-target evaluation scenario, for each non-image modality, the embedded features of the non-image modality are processed using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object in the non-image modality;

[0168] Judging whether the evaluation object has cheating behavior based on the embedded features of the image modality;

[0169] If the evaluation object has cheating behavior, the scores of the evaluation object in each non-image modality are penalized to obtain the final scores of each non-image modality.

[0170] In an alternative embodiment, when the scoring module 604 processes the joint embedded feature using the processing method corresponding to the function type of the current evaluation scenario to obtain the score of the evaluation object, it is used for:

[0171] If the function type of the current evaluation scenario is a pass / fail evaluation, score prediction is performed based on the joint embedded feature to obtain the score of the evaluation object;

[0172] If the function type of the current evaluation scenario is a selection evaluation, tier prediction is performed based on the joint embedded feature to obtain the score tier to which the evaluation object belongs; score prediction is performed based on the score tier and the joint embedded feature to obtain the score of the evaluation object.

[0173] An embodiment of the present application also provides an electronic device. Refer to Figure 7 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application can be a terminal device (such as a car machine, a large-screen device, a smart home, a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc.), or a server (which can be a single server, a server cluster, or a cloud server, etc.). Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0174] As Figure 7 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0175] Generally, the following devices can be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a memory card, a hard disk, etc.; and a communication device 709. The communication device 709 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0176] An embodiment of the present application also provides a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any one of the scoring methods provided by the embodiments of the present application.

[0177] An embodiment of the present application also provides a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any one of the scoring methods provided by the embodiments of the present application.

[0178] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0179] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0180] In the above embodiments, it can be implemented in whole or in part through software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. Professional technicians can use different methods to implement the described functions for each specific solution, but such implementation should not be considered to exceed the scope of this application.

[0181] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0182] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments may be referred to each other.

[0183] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A scoring method, characterized in that: include: Obtaining data of multiple modalities of the evaluation object; the multiple modalities include at least two of the following modalities: text, audio, and image; Encode the data of each modality separately to obtain the embedded features of each modality; If the current evaluation scenario is the target evaluation scenario, the embedded features of each modality are fused in a manner corresponding to the current evaluation scenario to obtain a joint embedded feature; The joint embedding feature is processed using a processing method corresponding to the functional type of the current evaluation scenario to obtain a score for the evaluation object.

2. The method according to claim 1, characterized in that The embedding features of each modality are fused in a manner corresponding to the current evaluation scenario, including: The embedded features of each modality are fused based on the weights of the embedded features of each modality; Among them, the weight of the embedded features of each modality corresponds to the current evaluation scenario; the weights of the embedded features of different modalities are the same or different.

3. The method according to claim 2, characterized in that The method of fusing the embedding features of each modality based on the weights of the embedding features of each modality includes: Based on the weight of the embedded features of each modality, the embedded features of each modality are weighted and concatenated to obtain the first initial fusion feature; The embedded features of each modality are encoded based on the attention mechanism to obtain the hidden features of each modality; the hidden features of each modality are concatenated to obtain the concatenated hidden features; the concatenated hidden features are encoded based on the attention mechanism to obtain the second initial fusion features; The first initial fusion feature and the second initial fusion feature are concatenated to obtain the joint embedding feature.

4. The method according to claim 2, characterized in that: The method of fusing the embedding features of each modality based on the weights of the embedding features of each modality includes: Concatenate the embedded features of each modality to obtain the third initial fusion feature; The embedded features of each modality are encoded based on the attention mechanism to obtain the hidden features of each modality; based on the weight of the embedded features of each modality, the hidden features of each modality are weighted concatenated to obtain the concatenated hidden features; the concatenated hidden features are encoded based on the attention mechanism to obtain the fourth initial fusion feature; The third initial fusion feature and the fourth initial fusion feature are concatenated to obtain the joint embedding feature.

5. The method according to claim 2, characterized in that: The method of fusing the embedding features of each modality based on the weights of the embedding features of each modality includes: Based on the weight of the embedding feature of each modality, the embedding features of each modality are weighted and concatenated to obtain the joint embedding feature.

6. The method according to claim 2, characterized in that The method of fusing the embedding features of each modality based on the weights of the embedding features of each modality includes: The embedded features of each modality are encoded based on the attention mechanism to obtain the hidden features of each modality; Based on the weight of the embedded features of each modality, the hidden layer features of each modality are weighted concatenated to obtain the concatenated hidden layer features; The concatenated hidden layer features are encoded based on the attention mechanism to obtain the joint embedding features.

7. The method according to claim 1, characterized in that Also includes: If the current evaluation scene is a non-target evaluation scene, corresponding to each non-image modality, the embedded features of the non-image modality are processed using a processing method corresponding to the functional type of the current evaluation scene to obtain a score of the evaluation object in the non-image modality; Determining whether the evaluation object has cheating behavior based on the embedded features of the image modality; If the evaluation object has cheating behavior, the scores of the evaluation object in each non-image modality are penalized to obtain the final scores of each non-image modality.

8. The method according to claim 1, characterized in that: The processing of the joint embedding feature using a processing method corresponding to the functional type of the current evaluation scenario to obtain a score of the evaluation object includes: If the function type of the current evaluation scenario is a qualified evaluation, a score prediction is performed based on the joint embedding feature to obtain a score of the evaluation object; If the functional type of the current evaluation scenario is a selection evaluation, a grade prediction is performed based on the joint embedded feature to obtain the score grade to which the evaluation object belongs; and a score prediction is performed based on the score grade and the joint embedded feature to obtain the score of the evaluation object.

9. A scoring device, characterized in that: include: A multimodal data acquisition module, used to obtain data of multiple modes of the evaluation object; the multiple modes include at least two of the following modes: text, audio, and image; The encoding module is used to encode the data of each modality respectively to obtain the embedding features of each modality; A fusion module is used to fuse the embedded features of each modality in a manner corresponding to the current evaluation scenario to obtain a joint embedded feature if the current evaluation scenario is the target evaluation scenario; The scoring module is used to process the joint embedding feature using a processing method corresponding to the functional type of the current evaluation scenario to obtain a score for the evaluation object.

10. A computer program product, characterized in that The method comprises computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the scoring method according to any one of claims 1 to 8.

11. An electronic device, characterized in that: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the scoring method according to any one of claims 1 to 8.

12. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the scoring method according to any one of claims 1 to 8.