AI dubbing quality automatic scoring method and system based on multi-modal data analysis

Through multimodal data analysis and actual feedback correction, the problems of insufficient coverage and lack of subjective feelings in the AI ​​dubbing quality scoring model were solved, achieving more accurate and flexible scoring results.

CN120748451AInactive Publication Date: 2025-10-03HANGZHOU XIAOZHONGQUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511040853.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing AI dubbing quality scoring model relies on limited training data, resulting in insufficient coverage and an inability to fully reflect the real dubbing situation and the subjective feelings of the audience. The scoring results deviate from the actual level.

Method used

By collecting multimodal data (audio, text, and emotional data) and combining it with pre-trained models for initial scoring, the scoring correction coefficient is calculated using actual performance indicators such as the number of dissemination times, barrage, and playback status, and the monitoring period is dynamically adjusted to optimize the score.

Benefits of technology

It improves the accuracy and authenticity of the scoring, can better reflect the audience's subjective feelings and emotional attitudes, achieves the matching of scoring results with the actual level, and improves the scoring efficiency and dynamic adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748451A_ABST
    Figure CN120748451A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, and particularly discloses an AI dubbing quality automatic scoring method and system based on multi-modal data analysis, and the method comprises the following steps: collecting the multi-modal data of a dubbing segment, scoring the dubbing segment based on the multi-modal data and a pre-trained scoring model, and recording the score as an initial score; acquiring a scoring index in a preset monitoring time period, calculating a scoring correction coefficient based on the scoring index, and correcting the initial score based on the scoring index to obtain a target score; and calculating a time period correction coefficient based on the target score and the initial score, correcting the monitoring time period based on the time period correction coefficient to obtain a target time period T, obtaining a time point t1 at which the monitoring time period ends, and calculating a new target score at the time point t1 + T. According to the invention, the accuracy and reliability of dubbing quality scoring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a method and system for automatically grading AI dubbing quality based on multimodal data analysis. Background Art

[0002] In the field of dubbing, trained AI models can automatically score dubbing videos, accurately detecting voice intonation, emotion, speaking speed, and pronunciation accuracy, providing objective and professional evaluations of dubbing works. This technology not only helps voice actors continuously improve their performances but also provides strong data support for post-production and optimization, driving the entire dubbing industry towards higher quality.

[0003] However, the dubbing quality scores generated by the model are highly dependent on the training data, which is often derived from a limited number of scenarios and samples. This data may not cover all real-world dubbing situations and may not fully reflect the subtle differences in actual performance. Furthermore, the training data often fails to fully reflect the listener's subjective feelings, making it difficult for the model's scores to capture key factors such as emotional expression in the dubbing. Ultimately, the scores deviate from the actual level. In summary, existing technologies suffer from the problem of model-generated scores deviating from reality. Summary of the Invention

[0004] The purpose of this invention is to provide an AI dubbing quality automatic scoring method and system based on multimodal data analysis to solve the following technical problems:

[0005] The dubbing quality ratings generated by the model are highly dependent on the training data, which is typically derived from a limited number of scenarios and samples. This data may not cover all real-world dubbing situations and may not fully reflect the nuances of actual performance. Furthermore, training data often fails to fully capture the listener's subjective feelings, making it difficult for the model's ratings to capture key factors such as emotional expression in the dubbing. This ultimately leads to a deviation between the ratings and the actual performance. In summary, existing technologies suffer from the problem of model-generated ratings deviating from reality.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] The AI ​​dubbing quality automatic scoring method based on multimodal data analysis includes the following steps:

[0008] Collecting multimodal data of the dubbing clip, the multimodal data including audio data, text data, and emotional data, and scoring the dubbing clip based on the multimodal data and a pre-trained scoring model to obtain an initial score;

[0009] Collect scoring indicators within a preset monitoring period, the scoring indicators including the number of disseminations, barrages, and playback status, the dissemination including collections, forwardings, likes, and dislikes, calculate a score correction coefficient based on the scoring indicators, and correct the initial score based on the scoring indicators to obtain a target score;

[0010] A period correction coefficient is calculated based on the target score and the initial score, the monitoring period is corrected based on the period correction coefficient to obtain a target period T, the time point t1 at which the monitoring period ends is obtained, and a new target score is calculated at time point t1+T.

[0011] As a further solution of the present invention, the process of obtaining the initial score specifically includes:

[0012] Establishing a database storing multimodal data with annotated scores, wherein the multimodal data includes audio data, text data, and sentiment data;

[0013] Establishing a scoring model based on a deep learning algorithm, and training and validating the scoring model based on the database to obtain a pre-trained scoring model;

[0014] Multimodal data of the dubbing segment is collected and input into the scoring model, and the score of the dubbing segment is output, which is recorded as the initial score.

[0015] As a further solution of the present invention: the multimodal data stored in the database is scored based on manual annotation.

[0016] As a further solution of the present invention: the process of obtaining the target score includes:

[0017] A score correction coefficient K1 is calculated based on the score index, and a target score P is calculated based on the score correction coefficient K1.

[0018] As a further solution of the present invention, the process of modifying the monitoring period to obtain the target period T specifically includes:

[0019] The correction ratio is calculated based on the initial score and the target score, and the monitoring period is corrected by the correction ratio to obtain the target period T.

[0020] As a further solution of the present invention, the process of modifying the monitoring period to obtain the target period T further includes the following steps:

[0021] Set the target period interval [T1, T2]. When the monitoring period T<T1, set the monitoring period T=T1; when the monitoring period T>T2, set the monitoring period T=T2.

[0022] The AI ​​dubbing quality automatic scoring system based on multimodal data analysis is applied to the AI ​​dubbing quality automatic scoring method based on multimodal data analysis, including an acquisition module, a scoring optimization module and a cycle optimization module. Specifically:

[0023] Acquisition module: collects multimodal data of the dubbing clip, the multimodal data including audio data, text data and emotional data, and scores the dubbing clip based on the multimodal data and a pre-trained scoring model to obtain an initial score;

[0024] Scoring Optimization Module: Collects scoring indicators within a preset monitoring period. The scoring indicators include the number of disseminations, barrages, and playback status. The dissemination includes collections, forwardings, likes, and dislikes. The scoring correction coefficient is calculated based on the scoring indicators. The initial score is corrected based on the scoring indicators to obtain the target score.

[0025] Period optimization module: Calculates a period correction coefficient based on the target score and the initial score, corrects the monitoring period based on the period correction coefficient to obtain a target period T, obtains the time point t1 at which the monitoring period ends, and calculates a new target score at time point t1+T.

[0026] The beneficial effects of the present invention are as follows:

[0027] 1) Although the initial ratings are derived from the results of a trained rating model, the model training data often lacks coverage or fails to fully reflect the audience's subjective feelings. By subsequently monitoring actual performance indicators such as the number of shares, comments, and playback duration, and revising the initial ratings accordingly, we can compensate for the deviations caused by insufficient training data to a certain extent, making the ratings more in line with actual audience feedback.

[0028] 2) This invention incorporates audience feedback, positive and negative comments, and playback preferences (favorites and dislikes) in the comments as correction factors, which can more directly reflect the audience's subjective perception and emotional attitude towards dubbing quality. This helps to compensate for the model's deficiency in relying solely on objective features such as voice and text, while ignoring subtle differences in emotional expression, making the final score more authentic and valuable.

[0029] 3) By adjusting the monitoring period, if the score deviates significantly from the initial value, the next indicator collection period will be shortened or extended accordingly. This adaptive monitoring strategy can collect feedback more frequently when the score changes significantly, and reduce unnecessary monitoring when the score changes tend to be stable, thereby achieving continuous optimization and dynamic updating of the score, improving scoring efficiency and accuracy;

[0030] 4) Since the audience's actual viewing and feedback data are continuously collected and integrated throughout the evaluation process, the audience's true reactions to different dubbing works can be captured in a timely manner, and the rating results can be dynamically adjusted as the audience's preferences and emotional attitudes change. This balances objectivity and real-time performance, helping to improve the overall evaluation level of dubbing quality.

[0031] 5) This approach introduces actual feedback from multimodal data and dynamically adjusts the initial model scores. This not only corrects the problem of insufficient model data coverage, but also allows the scores to more accurately reflect the audience's true feelings and preferences, thereby reflecting the true quality of the dubbing works to a greater extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The present invention will be further described below with reference to the accompanying drawings.

[0033] Figure 1 It is a flow chart of the AI ​​dubbing quality automatic scoring method based on multimodal data analysis of the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0035] See also Figure 1 As shown, the present invention is an AI dubbing quality automatic scoring method based on multimodal data analysis, comprising the following steps:

[0036] Collecting multimodal data of the dubbing clip, the multimodal data including audio data, text data, and emotional data, and scoring the dubbing clip based on the multimodal data and a pre-trained scoring model to obtain an initial score;

[0037] It is understandable that audio data usually refers to the original voice signal or its spectral characteristics and other representational information in the dubbing process, and text data can come from the corresponding dialogue manuscript or the text content obtained by transcribing the audio through a speech recognition system. Emotional data is generally based on the recognition and quantification of the speaker's tone, speaking speed, rhythm, expression and the emotions conveyed by the dubbing, such as manually labeling audio samples, or using existing emotion recognition models to output emotion labels and perform manual verification and correction; in the process of establishing the database, a reference score is also assigned to each multimodal data. The score often comes from the subjective evaluation of the dubbing work by professional reviewers, and the average or unified scale will be calculated within a certain range to make the score more consistent and comparable; the initial score reflects the quality of the dubbing clip. The higher the initial score, the higher the quality of the dubbing clip.

[0038] In a preferred embodiment of the present invention, the process of obtaining the initial score specifically includes:

[0039] Establishing a database storing multimodal data with annotated scores, wherein the multimodal data includes audio data, text data, and sentiment data;

[0040] Establishing a scoring model based on a deep learning algorithm, and training and validating the scoring model based on the database to obtain a pre-trained scoring model;

[0041] Collecting multimodal data of the dubbing clip, inputting it into the scoring model, and outputting a score of the dubbing clip, which is recorded as an initial score;

[0042] In a preferred embodiment of the present invention, the multimodal data stored in the database is scored based on manual annotation;

[0043] It is understandable that by establishing a database and training a scoring model, multimodal data such as audio, text, and emotion can be collected and integrated in the early stage to provide the model with a more comprehensive reference basis for dubbing quality. This step can make the most of the existing annotation information to learn the core features for preliminary judgment of dubbing quality, providing a relatively accurate starting point for subsequent scoring and correction, and also laying a data foundation for the generalization of the scoring model in various dubbing scenarios. Then, when obtaining the initial score, the dubbing clip to be evaluated is scored based on the model that has been trained and verified, and a relatively objective and stable initial result is obtained. This initial result can be used as the core reference for subsequent monitoring and correction, thereby achieving effective control of score changes and deviations.

[0044] Specifically, these multimodal data and their corresponding annotated scores are used to train and verify the scoring model. For example, a multi-input network structure that simultaneously receives speech waveform features, text embeddings, and sentiment label information, or a hybrid model that fuses different feature levels can be used. During the training process, common regression or classification methods are used to learn the comprehensive judgment of dubbing quality. After training, the model is verified, and the prediction accuracy and generalization ability of the model to different scenarios are evaluated through the verification data set. If the verification result meets the expected indicators, the model is regarded as a pre-trained scoring model. When a new dubbing clip needs to be evaluated, the multimodal data corresponding to the dubbing clip is first collected, including its audio file or audio feature extraction results, corresponding text or speech transcription text information, and sentiment data that may be obtained through real-time sentiment analysis or post-annotation. These data are then fed into the trained model as input. The model will score the quality of the dubbing clip based on the fusion of multimodal features and their internal correlation. The output score is the initial score, which lays the foundation for subsequent corrections and updates based on actual feedback.

[0045] Collect scoring indicators within a preset monitoring period, the scoring indicators including the number of disseminations, barrages, and playback status, the dissemination including collections, forwardings, likes, and dislikes, calculate a score correction coefficient based on the scoring indicators, and correct the initial score based on the scoring indicators to obtain a target score;

[0046] It is worth noting that in the process of collecting scoring indicators, actual performance information such as the number of disseminations, barrages, and playback time were introduced. These indicators can intuitively reflect the audience's evaluation and recognition of the dubbing quality in real usage scenarios, providing multi-dimensional feedback from the user side for subsequent corrections. At the same time, real-time data collection also further reduces the model's dependence on limited training samples. It can be understood that the number of disseminations is the total number of times corresponding to all dissemination types (total number of collections, total number of reposts, total number of likes, and total number of dislikes). The types of dissemination are collections, reposts, likes, and dislikes, and the playback situation is the duration of a single playback.

[0047] In another preferred embodiment of the present invention, the process of obtaining the target score includes:

[0048] A score correction coefficient K1 is calculated based on the score index, and a target score P is calculated based on the score correction coefficient K1.

[0049] It should be noted that the specific process includes:

[0050] Obtain the barrage of the dubbing video, determine whether there is a preset positive keyword in the barrage, and if so, use it as the positive barrage; determine whether there is a preset negative keyword in the barrage, and if so, use it as the negative barrage;

[0051] Obtain the single play duration and total play duration of the dubbing clip. If the ratio of the single play duration to the total play duration is less than 0.2, then the play is recorded as a disliked play; if the ratio of the single play duration to the total play duration is greater than 0.8, then the play is recorded as a liked play;

[0052] Calculate the score correction factor η1 is the preset first correction coefficient, X1 and X2 represent the total number of positive and negative barrages, respectively, Y1 and Y2 represent the total number of liked and disliked plays, respectively, and the degree of likeness C = C1-C2, where C1 represents the number of disseminations and C2 represents the number of dislikes.

[0053] Calculate the target score P = (1 + K1) * p, where p represents the initial score;

[0054] It is understandable that by correcting the initial score based on the score correction coefficient and allowing factors such as the number of dissemination, positive and negative comments, and playback preferences to participate in the score calculation, it is possible to more accurately capture the reputation and emotional value of the dubbing work among real audiences, making up for the deviation that may be caused by relying solely on the scoring model, and achieving a balance between subjective feelings and objective characteristics in the score;

[0055] Specifically, in order to revise the initial rating based on actual feedback, it is necessary to collect rating indicators within a preset monitoring period, including the number of disseminations, barrages and playback status. The number of disseminations can be further subdivided into operation data such as collections, forwarding, likes and dislikes. For example, the system will count the number of audiences’ collections of dubbing videos on video platforms or social media over a period of time, whether they forwarded, shared, liked or disliked the videos, etc.; then, for the barrage information, the preset positive keywords (such as “very nice”, “infectious”, etc.) and reverse keywords (such as “too bland”, “inaudible”, etc.) can be retrieved automatically or manually. All barrages that match the positive keywords are counted as positive barrages, and those that match the reverse keywords are counted as reverse barrages; the playback status mainly focuses on the proportion of a single playback time to the total video time. If the proportion of a certain playback time is less than 0.2, it is considered that the playback is not liked. Playback, for example, the audience quits after watching less than one-fifth of the video; if the playback time accounts for more than 0.8, it is determined that the playback is a favorite playback, which often indicates that the audience is willing to watch it for most of the time; after collecting the above data, the first step is to calculate the degree of likes C=C1-C2 based on the number of positive barrages X1, the number of reverse barrages X2, the number of favorite playbacks Y1, the number of dislike playbacks Y2, the number of disseminations C1 and the number of dislikes C2, and combine it with the preset first correction coefficient η1 to obtain the score correction coefficient K1, which can be specifically derived by combining different weight formulas, and finally obtain the target score in the form of P=(1+K1)*p based on the initial score p; in this way, the score given by the original model can be reasonably and dynamically corrected in combination with the audience's real subjective feedback on the dubbing work, so that the score result is closer to the evaluation and recognition obtained by the dubbing work in real scenarios;

[0056] Calculating a time period correction coefficient based on the target score and the initial score, correcting the monitoring period based on the time period correction coefficient to obtain a target time period T, obtaining a time point t1 at which the monitoring period ends, and calculating a new target score at time point t1+T;

[0057] It should be noted that the steps of correcting the monitoring period and calculating the target period T are to dynamically adjust the cycle of the next data collection according to the magnitude of the score change. When the score difference is large or the fluctuation is obvious, more intensive monitoring is carried out to promptly detect and correct the score. When the score change tends to be stable, the monitoring interval is appropriately extended, thereby improving resource utilization and ensuring that the score can be updated at the appropriate time. Finally, by updating the score after the end of the target period, a continuous evaluation of the quality of the dubbing work is achieved, which can not only reflect the audience's latest feedback on the dubbing quality in a timely manner, but also enable the scoring mechanism to continuously evolve and optimize, ultimately achieving the goal of making the dubbing quality score closer to the actual performance and more flexibly adapt to different scenarios and audience needs.

[0058] In a preferred embodiment of the present invention, the process of modifying the monitoring period to obtain the target period T specifically includes:

[0059] The correction ratio is calculated based on the initial score and the target score, and the monitoring period is corrected by the correction ratio to obtain the target period T.

[0060] It should be noted that the specific process includes:

[0061] Calculate the correction ratio A = η2*(Pp) / p, where p represents the initial score and η2 is the preset second correction coefficient;

[0062] Calculate the target period T = (Ays / A)*Tys, where Ays represents the preset correction ratio threshold and Tys represents the length of the monitoring period;

[0063] It is worth noting that the process of modifying the monitoring period to obtain the target period T also includes the following steps:

[0064] Set the target period interval [T1, T2]. When the monitoring period T<T1, set the monitoring period T=T1; when the monitoring period T>T2, set the monitoring period T=T2.

[0065] It should be noted that in order to achieve dynamic adjustment of the monitoring period, the deviation of the current target score P from the initial score p is first measured by calculating the correction ratio A, and the correction ratio is combined with the preset correction ratio threshold Ays to obtain the target period T = Ays / A. In this way, the time interval for the next monitoring can be flexibly determined according to the magnitude of the change in the score; for example, when P is significantly higher than p, A will increase accordingly, causing T to become relatively smaller, thereby more frequent monitoring and correction when the score fluctuates greatly, and vice versa, unnecessary update frequency will be reduced; in addition, in order to avoid T being too small or too large and having a negative impact on system efficiency or real-time performance, a target period interval [T1, T2] will also be set, which not only ensures the timeliness of monitoring, but also takes into account the rational use of resources and moderate control of the score update cycle.

[0066] The AI ​​dubbing quality automatic scoring system based on multimodal data analysis is applied to the AI ​​dubbing quality automatic scoring method based on multimodal data analysis, including an acquisition module, a scoring optimization module and a cycle optimization module. Specifically:

[0067] Acquisition module: collects multimodal data of the dubbing clip, the multimodal data including audio data, text data and emotional data, and scores the dubbing clip based on the multimodal data and a pre-trained scoring model to obtain an initial score;

[0068] Scoring Optimization Module: Collects scoring indicators within a preset monitoring period. The scoring indicators include the number of disseminations, barrages, and playback status. The dissemination includes collections, forwardings, likes, and dislikes. The scoring correction coefficient is calculated based on the scoring indicators. The initial score is corrected based on the scoring indicators to obtain the target score.

[0069] Cycle Optimization Module: Calculates the period correction coefficient based on the target score and the initial score, corrects the monitoring period based on the period correction coefficient to obtain the target period T, obtains the time point t1 at which the monitoring period ends, and calculates the new target score at time point t1+T

[0070] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. An AI dubbing quality automatic scoring method based on multimodal data analysis, characterized in that: The following steps are involved: Collecting multimodal data of the dubbing clip, the multimodal data including audio data, text data, and emotional data, and scoring the dubbing clip based on the multimodal data and a pre-trained scoring model to obtain an initial score; Collect scoring indicators within a preset monitoring period, wherein the scoring indicators include the number of disseminations, barrages, and playback status, wherein the dissemination includes collections, forwardings, likes, and dislikes, and revise the initial score based on the scoring indicators to obtain a target score; A period correction coefficient is calculated based on the target score and the initial score, the monitoring period is corrected based on the period correction coefficient to obtain a target period T, the time point t1 at which the monitoring period ends is obtained, and a new target score is calculated at time point t1+T.

2. The AI ​​dubbing quality automatic scoring method based on multimodal data analysis according to claim 1 is characterized in that: The process of obtaining the initial score specifically includes: Establishing a database storing multimodal data with annotated ratings; Establishing a scoring model based on a deep learning algorithm, and training and validating the scoring model based on the database to obtain a pre-trained scoring model; Multimodal data of the dubbing segment is collected and input into the scoring model, and the score of the dubbing segment is output, which is recorded as the initial score.

3. The AI ​​dubbing quality automatic scoring method based on multimodal data analysis according to claim 2 is characterized in that: The multimodal data stored in the database is scored based on manual annotation.

4. The AI ​​dubbing quality automatic scoring method based on multimodal data analysis according to claim 1 is characterized in that: The process of obtaining a target score includes: A score correction coefficient K1 is calculated based on the score index, and a target score P is calculated based on the score correction coefficient K1.

5. The AI ​​dubbing quality automatic scoring method based on multimodal data analysis according to claim 4 is characterized in that: The process of modifying the monitoring period to obtain the target period T specifically includes: The correction ratio is calculated based on the initial score and the target score, and the monitoring period is corrected by the correction ratio to obtain the target period T.

6. The AI ​​dubbing quality automatic scoring method based on multimodal data analysis according to claim 5 is characterized in that: The process of modifying the monitoring period to obtain the target period T also includes the following steps: Set the target period interval [T1, T2]. When the monitoring period T<T1, set the monitoring period T=T1; when the monitoring period T>T2, set the monitoring period T=T2.

7. An AI dubbing quality automatic scoring system based on multimodal data analysis, applied to the AI ​​dubbing quality automatic scoring method based on multimodal data analysis according to any one of claims 1 to 6, characterized in that: It includes the collection module, the scoring optimization module and the cycle optimization module. Specifically: Acquisition module: collects multimodal data of the dubbing clip, the multimodal data including audio data, text data and emotional data, and scores the dubbing clip based on the multimodal data and a pre-trained scoring model to obtain an initial score; Scoring Optimization Module: Collects scoring indicators within a preset monitoring period. The scoring indicators include the number of disseminations, barrages, and playback status. The dissemination includes collections, forwardings, likes, and dislikes. The scoring correction coefficient is calculated based on the scoring indicators. The initial score is corrected based on the scoring indicators to obtain the target score. Period optimization module: Calculates a period correction coefficient based on the target score and the initial score, corrects the monitoring period based on the period correction coefficient to obtain a target period T, obtains the time point t1 at which the monitoring period ends, and calculates a new target score at time point t1+T.