Sound quality test method

By comparing the intensity of lip-movement units in videos of subjects and references, this study addresses the problem of traditional pronunciation quality testing being affected by environmental noise and individual native languages, achieving a more accurate assessment of pronunciation levels.

CN115831153BActive Publication Date: 2025-12-30ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211159540.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-12-30
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

Traditional methods of testing pronunciation quality are easily affected by environmental noise and individual native language pronunciation characteristics, making it difficult to accurately determine the pronunciation level of the test taker.

Method used

By acquiring videos of subjects and references, visual information is used to detect the intensity of lip-movement units. The intensity of the lip-movement units of subjects and references is compared to determine the pronunciation quality.

Benefits of technology

It improves the accuracy of pronunciation quality testing, reduces the influence of environmental interference, and can more realistically assess the pronunciation level of the test takers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831153B_ABST
    Figure CN115831153B_ABST
Patent Text Reader

Abstract

The application provides a pronunciation quality test method, comprising: acquiring a first video of a subject reading target content and a second video of a target reference reading the target content; performing mouth shape action unit intensity detection on each of a plurality of first images in the first video to obtain mouth shape action unit intensity corresponding to each of the plurality of first images; and performing mouth shape action unit intensity detection on each of a plurality of second images in the second video to obtain mouth shape action unit intensity corresponding to each of the plurality of second images. The mouth shape action unit intensity corresponding to each of the plurality of first images is compared with the mouth shape action unit intensity corresponding to each of the plurality of second images to determine the pronunciation quality of the subject. The present scheme tests the pronunciation quality of the subject, such as aphasia personnel, based on visual information, i.e. the dynamic characteristics of the mouth during pronunciation, is less affected by environmental factors, and can obtain more accurate pronunciation quality test results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet technology, and in particular to a method for testing pronunciation quality. Background Technology

[0002] For example, in application scenarios such as learning a language or understanding the pronunciation of people with aphasia, it is necessary to understand the pronunciation level of the subject.

[0003] The traditional method for assessing pronunciation quality involves having the subject read a passage aloud, recording their pronunciation audio, and then using a phoneme recognition model to identify the subject's actual phoneme sequence. This sequence is then compared with the phoneme sequence identified from a standard pronunciation audio recording to determine the pronunciation quality. The standard pronunciation audio recording can be taken from a person with standard pronunciation skills (referred to as the reference) reading the same passage.

[0004] However, in practice, phoneme-based pronunciation quality testing methods are more susceptible to environmental noise, individual and native language pronunciation characteristics, making it difficult to accurately determine the pronunciation level of the test taker. Summary of the Invention

[0005] This invention provides a method, device, and storage medium for testing pronunciation quality, which can accurately determine the pronunciation level of a subject through visual information.

[0006] In a first aspect, embodiments of the present invention provide a method for testing pronunciation quality, the method comprising:

[0007] Acquire a first video of a subject reading the target content, and a second video of a target reference reading the target content;

[0008] Lip-sync unit intensity detection is performed on multiple frames of first images in the first video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of first images.

[0009] The intensity of lip-sync unit is detected for each of the multiple frames of the second image in the second video to obtain the intensity of the lip-sync unit corresponding to each of the multiple frames of the second image.

[0010] The intensity of the lip-sync unit corresponding to each of the multiple first images is compared with the intensity of the lip-sync unit corresponding to each of the multiple second images to determine the pronunciation quality of the subject;

[0011] The intensity of the lip-shape action unit corresponding to any image refers to the intensity coefficient of various preset lip-shape action units when forming the lip shape in any image.

[0012] Secondly, embodiments of the present invention provide a sound quality testing device, the device comprising:

[0013] The acquisition module is used to acquire a first video of a subject reading the target content and a second video of a target reference reading the target content;

[0014] The detection module is used to perform lip-sync unit intensity detection on multiple frames of first images in the first video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of first images; and to perform lip-sync unit intensity detection on multiple frames of second images in the second video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of second images.

[0015] The determination module is used to compare the intensity of the lip-shape action unit corresponding to each of the multiple first images with the intensity of the lip-shape action unit corresponding to each of the multiple second images to determine the pronunciation quality of the subject; wherein, the intensity of the lip-shape action unit corresponding to any image refers to the intensity coefficient corresponding to multiple preset lip-shape action units when forming the lip shape in any image.

[0016] Thirdly, embodiments of the present invention provide an electronic device, including: a memory, a processor, and a communication interface; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor performs the pronunciation quality testing method as described in the first aspect.

[0017] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, causes the processor to perform the pronunciation quality testing method as described in the first aspect.

[0018] Fifthly, embodiments of the present invention provide a method for testing pronunciation quality, the method comprising:

[0019] Receive a request triggered by a user device through calling a pronunciation quality test service, the request including a first video of a subject reading the target content and a second video of a target reference reading the target content;

[0020] The following steps are performed using the processing resources corresponding to the pronunciation quality testing service:

[0021] Lip-sync unit intensity detection is performed on multiple frames of first images in the first video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of first images.

[0022] The intensity of lip-sync unit is detected for each of the multiple frames of the second image in the second video to obtain the intensity of the lip-sync unit corresponding to each of the multiple frames of the second image.

[0023] The intensity of the lip-sync unit corresponding to each of the multiple first images is compared with the intensity of the lip-sync unit corresponding to each of the multiple second images to determine the pronunciation quality of the subject;

[0024] The intensity of the lip-shape action unit corresponding to any image refers to the intensity coefficient of various preset lip-shape action units when forming the lip shape in any image.

[0025] Sixthly, embodiments of the present invention provide a pronunciation quality testing method applied to an extended reality device, the method comprising:

[0026] Acquire a first video of a subject reading the target content, and a second video of a target reference reading the target content;

[0027] The intensity of lip-sync unit is detected for each of the first frames in the first video to obtain the intensity of the lip-sync unit corresponding to each of the first frames; the intensity of lip-sync unit is detected for each of the second frames in the second video to obtain the intensity of the lip-sync unit corresponding to each of the second frames; wherein, the intensity of the lip-sync unit corresponding to any image refers to the intensity coefficient of each of the various preset lip-sync units when forming the lip shape in any image.

[0028] The intensity of the lip-sync unit corresponding to each of the multiple first images is compared with the intensity of the lip-sync unit corresponding to each of the multiple second images to determine the pronunciation quality of the subject;

[0029] The sound quality is rendered and displayed on the extended reality device screen.

[0030] The pronunciation quality testing scheme provided in this invention is a scheme for testing the pronunciation level of subjects based on visual information. The pronunciation quality test focuses more on the shape of the subject's mouth. Therefore, multiple mouth shape action units are pre-set to reflect the mouth shape. Different mouth shape action units correspond to different mouth shape states. A certain mouth shape currently presented by a person can be represented by a linear combination of multiple mouth shape action units. When performing a linear combination, the intensity (also called intensity coefficient or weighting coefficient) of each mouth shape action unit needs to be known. Based on this, when testing the pronunciation quality of subjects, firstly, a first video of the subject reading the target content and a second video of a target reference reading the target content are acquired. The subject is equivalent to a student, and the target reference is equivalent to a teacher. The subject learns the pronunciation of the target reference to read the same content. Next, the first and second videos were sampled to obtain multiple frames from the first video (referred to as multi-frame first images) and multiple frames from the second video (referred to as multi-frame second images). Lip movement unit intensity was detected for each frame to obtain the lip movement unit intensity corresponding to each of the multi-frame first images and the multi-frame second images. Finally, the lip movement unit intensities corresponding to each of the multi-frame first images and the multi-frame second images were compared to determine the subject's pronunciation quality. Simply put, the closer the lip movement unit intensities corresponding to the multi-frame first images and the multi-frame second images are, the better the pronunciation quality.

[0031] The above-mentioned scheme for testing the pronunciation quality of subjects based on visual information (i.e., the dynamic characteristics of the mouth during pronunciation) is less affected by environmental factors and can obtain more accurate pronunciation quality test results. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 A flowchart of a pronunciation quality testing method provided in an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram of multiple mouth-shaped action units provided in an embodiment of the present invention;

[0035] Figure 3 A flowchart of an optional implementation method for step 103;

[0036] Figure 4A schematic diagram illustrating the training process of the coherence score weight prediction model provided in an embodiment of the present invention;

[0037] Figure 5 This is a schematic diagram illustrating the application of a pronunciation quality testing method provided in an embodiment of the present invention;

[0038] Figure 6 A flowchart of a pronunciation quality testing method provided in an embodiment of the present invention;

[0039] Figure 7 A flowchart illustrating a face detection model training method provided in an embodiment of the present invention;

[0040] Figure 8 for Figure 7 A schematic diagram illustrating the principle of the model training method shown;

[0041] Figure 9 A flowchart illustrating another face detection model training method provided in an embodiment of the present invention;

[0042] Figure 10 for Figure 9 A schematic diagram illustrating the principle of the model training method shown;

[0043] Figure 11 This is a schematic diagram illustrating the application of a pronunciation quality testing method provided in an embodiment of the present invention;

[0044] Figure 12 This is a schematic diagram of the structure of a sound quality testing device provided in an embodiment of the present invention;

[0045] Figure 13 This is a schematic diagram of the structure of an electronic device provided in this embodiment. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.

[0048] Figure 1 A flowchart of a pronunciation quality testing method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the method includes the following steps:

[0049] 101. Obtain a first video of a subject reading the target content and a second video of a target reference reading the target content.

[0050] 102. Perform lip-sync intensity detection on multiple frames of the first image in the first video to obtain the lip-sync intensity of each frame of the first image. Perform lip-sync intensity detection on multiple frames of the second image in the second video to obtain the lip-sync intensity of each frame of the second image.

[0051] 103. Compare the intensity of the lip-sync units corresponding to each of the first frames of multiple images with the intensity of the lip-sync units corresponding to each of the second frames of multiple images to determine the pronunciation quality of the subject.

[0052] The intensity of the lip-shape action unit corresponding to any image refers to the intensity coefficient of various preset lip-shape action units when forming the lip shape in any image.

[0053] The pronunciation quality testing scheme provided in this invention is used to test the spoken pronunciation quality of subjects, such as students, people with aphasia, etc. The test is conducted using a shadowing method. This involves a "teacher" reading a passage of text, and the subject, as the "student," repeating the same passage. In this invention, the "teacher" is referred to as the target reference, and the text being read is called the target content. It is understood that in practical applications, multiple different test texts can be pre-set, and the target content can be one of them.

[0054] While the target reference person is reading the target content, video recordings are made of the person's pronunciation, resulting in the second video mentioned above. Similarly, while the subject is reading the target content, a corresponding video recording is also made, resulting in the first video mentioned above.

[0055] Then, the first video can be sampled at a set sampling frequency to obtain multiple frames of images in the first video, referred to as the multi-frame first images. Similarly, the second video can be sampled at the same sampling frequency to obtain multiple frames of images in the second video, referred to as the multi-frame second images.

[0056] Understandably, the lengths of the first and second videos may differ, resulting in unequal numbers of images sampled. For instance, when a subject has aphasia, the time required for them to read the target content is often longer than that of the target reference.

[0057] Subsequently, lip-sync intensity detection is performed on multiple frames of the first image in the first video to obtain the lip-sync intensity corresponding to each frame. Similarly, lip-sync intensity detection is performed on multiple frames of the second image in the second video to obtain the lip-sync intensity corresponding to each frame. Here, the lip-sync intensity corresponding to any image refers to the intensity coefficient of various preset lip-sync actions used to form the lip shape in that image.

[0058] The test of pronunciation quality focuses more on the shape of the subject's mouth. Therefore, a number of basic mouth action units are pre-set to reflect the shape of the mouth. A certain mouth shape presented by a person can be represented by the linear superposition of multiple mouth action units. When performing linear superposition, it is necessary to know the intensity of each mouth action unit (also known as intensity coefficient or coefficient).

[0059] Similar to facial motion coding systems, mouth movements can be pre-divided into multiple independent yet interconnected action units based on their characteristics. Each action unit has a defined intensity range, typically [0,1]. Different intensities of an action unit often correspond to different motion characteristics, and different action units typically control different mouth regions. Therefore, when someone pronounces a sound, the corresponding mouth shape can be the result of linearly superimposing different action units according to their respective intensities.

[0060] For example, suppose multiple mouth movement units are represented as AU1, AU2, ..., AUn. The intensity of AU1 corresponding to a person's pronunciation of "ah" is about 0.6-0.8, while the intensity of AU1 for another person with a pronunciation disorder may only be 0.3.

[0061] In this embodiment, a face detection model for detecting the intensity of lip movements can be pre-trained. This face detection model is trained in conjunction with a face reconstruction model, and its training process will be described below. Figure 2 As shown, in Figure 2 In this model, the face detection model is represented as encoder, and the multiple lip action units are represented as AU1, AU2, ..., AUn, respectively. The intensity of each of these n lip action units is represented as r1, r2, ..., rn.

[0062] After sampling multiple frames of images from the first video and the second video, the multiple frames of the first image from the first video can be input into the face detection model to obtain the lip-sync unit intensity corresponding to each of the multiple frames of the first image. Similarly, the multiple frames of the second image from the second video can be input into the face detection model to obtain the lip-sync unit intensity corresponding to each of the multiple frames of the second image.

[0063] Then, the pronunciation quality of the subject is determined by comparing the intensity of lip-sync units (LCUs) corresponding to each of the first multiple frames with the intensity of lip-sync units corresponding to each of the second multiple frames. It is understandable that, since multiple (let's say n) lip-sync units are pre-defined, each frame input into the face detection model will yield the intensity of multiple lip-sync units corresponding to that frame. Assuming the first video contains m1 frames, then m1*n lip-sync unit intensities will be obtained for the first video, and they will be sorted chronologically. Similarly, assuming the second video contains m2 frames, then m2*n lip-sync unit intensities will be obtained for the first video, and they will be sorted chronologically. In practical applications, due to differences in speech rate between the subject and the target reference, m1 and m2 may not be equal.

[0064] In this embodiment of the invention, the aim is to analyze the spoken pronunciation quality of a subject by comparing the differences between the lip features of the subject speaking the same speech and a standard level. The lip features are represented by the intensity of the lip movement units corresponding to each of the multiple first images, while the standard level is represented by the intensity of the lip movement units corresponding to each of the multiple second images of the target reference speaking the same speech.

[0065] In an optional embodiment, for any frame of a first image i in the first video, a search time period can be determined based on its corresponding sampling timestamp according to a set time span. From multiple frames of second images in the second video, second images whose sampling timestamps fall within this search time period (assuming there are k such images) are identified. Then, the lip-sync unit intensity corresponding to the first image i is compared with the lip-sync unit intensity corresponding to each of the k second images. If there is a target second image among the k second images whose lip-sync unit intensity matches that of the first image i, the matching score for the first image i is determined to be one point. If there is no target second image among the k second images whose lip-sync unit intensity matches that of the first image i, the matching score for the first image i is determined to be zero points.

[0066] Through the above process, the matching scores corresponding to each of the multiple frames of the first image in the first video can be obtained. The ratio of the sum of the matching scores to the total number of the multiple frames of the first image can be determined as the quality score of the first video. If the quality score is greater than a set threshold, the subject's pronunciation level is considered good; otherwise, it is considered poor.

[0067] Specifically, assuming the sampling timestamp corresponding to the first image i is t1, the corresponding search time period is, for example, frames t1-15 to t1+5. Frames t1-15 represent the second image sampled at time t1 in the second video and the 15 frames sampled before it. Similarly, frames t1+5 represent the second image sampled at time t1 in the second video and the 5 frames sampled after it. In other words, the aforementioned 21 frames in the second video are included within the search time period of the first image i. In practical applications, generally speaking, the pronunciation of the subject is often not better than that of the target reference person; for example, the subject's pronunciation fluency may be worse than that of the target reference person. Therefore, optionally, the above search time period can be set based on the assumption that for the same word in the target content, the subject's pronunciation is more likely to be later than that of the target reference person. Of course, in practical applications, the above search time period setting is not limited to this.

[0068] Specifically, the matching degree between the lip-sync unit intensities corresponding to two frames of images can be determined using the following optional method: Taking the n lip-sync units AU1, AU2, ..., AUn as an example, and taking images a and b as examples, the average value of the lip-sync unit intensities corresponding to each of the two frames can be calculated based on the intensity of each lip-sync unit in each of the two frames. If the difference between the two average values ​​is less than a set threshold, then images a and b are considered to match, and their matching degree score is 1. Otherwise, if the difference between the two average values ​​is greater than or equal to the set threshold, then images a and b are considered to not match, and their matching degree score is 0.

[0069] Alternatively, for the intensities of the n lip-sync units corresponding to each of images a and b, the intensity difference corresponding to the same lip-sync unit can be calculated. If all n intensity differences are less than a set threshold, then images a and b are considered to match, and their matching score is 1. Otherwise, if there is at least one intensity difference among the n intensity differences that is greater than or equal to the set threshold, then images a and b are considered to not match, and their matching score is 0.

[0070] In practical applications, taking the first image i as an example, if there are at least two frames of the second image within its corresponding search time period that satisfy the condition that "the difference of the average value is less than the set threshold" or "the difference of n intensity values ​​is less than the set threshold", then for the first image i, its corresponding matching score is still 1.

[0071] In the solution provided in the above embodiments, by recognizing and comparing the visual features of the mouth of the subject and the target reference respectively, the pronunciation quality of the subject can be tested visually, which has better anti-interference ability and can improve the accuracy of the pronunciation quality test results.

[0072] In this embodiment of the invention, considering the coherence and correctness of speech, and combining the coherence and correctness of specific sentences, a health score for evaluating pronunciation quality is designed. In summary, in determining the pronunciation quality of a subject by comparing the intensity of lip-sync units corresponding to each of multiple first images and multiple second images, the total score of pronunciation coherence and / or the total score of pronunciation correctness during the reading of the target content can be determined. This allows for the combination of the total score of pronunciation coherence and / or the total score of pronunciation correctness to obtain the subject's pronunciation health score. If this health score is greater than a set threshold, the subject's pronunciation quality is considered good; otherwise, it is considered poor. The following is a further explanation... Figure 3 The illustrated embodiment describes this pronunciation quality assessment scheme.

[0073] Figure 3 A flowchart of an optional implementation method for step 103 is shown below. Figure 3 As shown, the method includes the following steps:

[0074] 301. Based on the degree of change in the intensity of the lip-syncing units corresponding to each of the multiple first images, determine multiple pause points of the subject; and based on the degree of change in the intensity of the lip-syncing units corresponding to each of the multiple second images, determine multiple pause points of the target reference.

[0075] 302. Determine multiple speech segments of a subject based on multiple pauses of the subject; and determine multiple speech segments of a target reference based on multiple pauses of the target reference, wherein a speech segment includes multiple frames of images between adjacent pauses.

[0076] 303. Compare multiple pronunciation segments of the subject with multiple pronunciation segments of the target reference to determine the subject's total score for pronunciation coherence and total score for pronunciation accuracy, so as to determine the subject's pronunciation quality based on the subject's total score for pronunciation coherence and total score for pronunciation accuracy.

[0077] Since the target content is often a paragraph consisting of multiple sentences, both the target reference and the subject will experience prolonged pauses (i.e., sentence breaks) during the reading of the target content. In this embodiment of the invention, the pause points of the corresponding reference are identified based on the degree of change in the intensity of lip-syncing units in consecutive frames of images.

[0078] Taking the first video as an example, after sampling the first frame, it's possible to determine whether multiple consecutive frames meet the pause criteria. Assuming the multiple first frames in the first video are represented sequentially as F1-Fm1, and if the lip-sync intensity remains constant for h consecutive frames (e.g., h=10), a pause is considered to exist. If the lip-sync intensity changes for the 10 frames F1-F10, and remains constant for frames F11-F20, then frame F10 is determined to be the end of a pronunciation segment, thus identifying the 10 frames F1-F10 as the first pronunciation segment. Next, assuming the lip-sync intensity differs from that of frame F21 and F20, then frame F21 is determined to be the start of the second pronunciation segment, and the end point of the second pronunciation segment is determined using the same process. Similarly, multiple pronunciation segments of the subject during the pronunciation of the target content can be obtained. Likewise, multiple pronunciation segments of the target reference during the pronunciation of the target content can also be obtained.

[0079] It is understandable that the pause points of the subject and the target reference may not be the same, that is, multiple pronunciation segments of the subject and the target reference may not be the same.

[0080] For example, suppose the target content is: Hello, please read the following content with the teacher, paying attention to the punctuation. The pauses below are indicated by " / ". The target learner's pauses would be: Hello / Please read the following content with the teacher / Pay attention to the punctuation. The subject's pauses would be: Hello / Please read the following content with the teacher / Pay attention to / Punctuation.

[0081] As the examples above illustrate, if the subject is aphasic, it is more difficult for them to read a long sentence fluently, and they may exhibit significantly more pauses compared to the target reference. In the examples above, the target reference had a total of 3 vocal segments, while the subject had 5 vocal segments.

[0082] After the first and second videos, which were collected during the process of the subject and the target reference reading the target content, were divided into multiple pronunciation segments according to the changes in the intensity of the mouth movement units in multiple consecutive frames, the pronunciation coherence and correctness of the two pronunciation segments could be compared. In order to finally determine the subject's performance in pronunciation coherence and correctness relative to the pronunciation of the target reference, a health score reflecting this performance was obtained.

[0083] It should be noted that the health score may consider only pronunciation fluency or pronunciation accuracy, or both, depending on the actual purpose of the test.

[0084] The following sections explain the judgment process for both pronunciation fluency and the correctness of the invention.

[0085] Regarding pronunciation fluency:

[0086] First, the duration of each of the subject's multiple pronunciation segments and the duration of each of the target reference's multiple pronunciation segments can be determined. Then, based on the durations of each of the subject's and target reference's multiple pronunciation segments, multiple pronunciation coherence scores can be determined, where each score corresponds to one of the subject's multiple pronunciation segments. Finally, based on the subject's multiple pronunciation coherence scores, a total pronunciation coherence score can be determined.

[0087] This embodiment provides an optional method for calculating coherence scores, expressed as follows:

[0088] tmp score(i) =count time (xj') / [count time (xi)], if i = j; (1)

[0089] tmp score(i) =count time (xj') / [count time (xi)+δ], if i≠j (2)

[0090] Where xi represents the i-th pronunciation segment of the subject, xi' represents the j-th pronunciation segment of the target reference, and count time (xi) represents the duration of the i-th vocal segment of the subject, count time (xj') represents the duration of the j-th pronunciation segment of the target reference, tmp score(i)δ represents the coherence score corresponding to the i-th pronunciation segment of the subject, and is a preset infinite value. The value of i ranges from [0, n1], and the value of j ranges from [1, n2]. Here, n1 represents the total number of pronunciation segments of the subject, and n2 represents the total number of pronunciation segments of the target reference.

[0091] Understandably, as illustrated in the examples above, the number of pronunciation segments produced by the target follower and the subject when reading the target content may differ. For instance, the target reference may only have the first, second, and third pronunciation segments, while the subject has the first, second, third, fourth, and fifth pronunciation segments. Under this assumption, since both the subject and the target reference have the first three pronunciation segments, the coherence score corresponding to each of the subject's first three pronunciation segments can be determined according to formula (1). Because the target reference does not have the fourth and fifth pronunciation segments, the calculation is based on the formula (2) above. Since count is calculated in this case... time (x4') and count time (x5') takes a value of 0, resulting in the coherence scores for the subject's fourth and fifth pronunciation segments being both 0. Similarly, if the subject has fewer pronunciation segments than the target reference (e.g., the subject has 3 pronunciation segments while the target reference has 4), then the duration count of the subject's fourth pronunciation segment will be lower. time (x4) = 0. Based on formula (2), since the denominator is an infinite value at this time, the continuity score of the subject's fourth pronunciation segment is 0.

[0092] In other words, the above formulas (1) and (2) are meant to express that, for the subject, if the number of their pronunciation segments is different from that of the target reference, the coherence score corresponding to the missing or extra pronunciation segments is 0.

[0093] After obtaining the coherence scores for each of the subject's multiple pronunciation segments, the multiple coherence scores can optionally be summed together to obtain the subject's total coherence score.

[0094] Alternatively, multiple articulation coherence scores can be input into a coherence score weight prediction model to obtain the weights of the multiple articulation coherence scores. The multiple articulation coherence scores can be weighted and summed according to their weights to determine the subject's total articulation coherence score, so as to determine the subject's articulation quality by combining the subject's total articulation coherence score.

[0095] The coherence score weight prediction model is a pre-trained model used to predict the weights corresponding to the multiple pronunciation coherence scores of the input. In terms of implementation structure, it can be implemented as a convolutional neural network composed of multiple convolutional layers.

[0096] Typically, the target content includes multiple sentences of varying lengths. These sentences can be considered as pre-defined segments that should be pronounced according to grammar, semantic groups, and other factors. It can be assumed that all references, including the target reader, should read the target content according to these pre-defined segmentation points. For example, if the target content is: "Hello, please follow the teacher and read the following content, paying attention to the segmentation points," and the pauses are indicated by " / ", then the segmented sentences are as follows: "Hello / Please follow the teacher and read the following content / Pay attention to the segmentation points."

[0097] In practice, for subjects, such as those with aphasia or poor learning abilities, short sentences are often easier to pronounce and pause correctly, while long sentences are more difficult to pause accurately. Therefore, the coherence score weighting prediction model is trained to assign higher weights to long sentences and lower weights to short sentences.

[0098] To train a coherence score weight prediction model with the aforementioned capabilities, a training set is first required. In this embodiment of the invention, for ease of description, only the target content mentioned above is used as the training material as an example. In reality, the training material can include multiple different contents. Videos of multiple references reading the target content aloud can be collected in advance, and the duration of multiple pronunciation segments corresponding to each reference can be obtained according to the method described in the previous embodiment. For ease of description, it is assumed that multiple references read strictly according to the pre-defined punctuation positions (i.e., pause positions), so the number of pronunciation segments of multiple references is the same, and in fact, the start and end words corresponding to each pronunciation segment are also the same. The main difference lies in the fact that the duration of the same pronunciation segment of different references may be different. For example, the duration of the first pronunciation segment of reference a is ta1, the duration of the first pronunciation segment of reference b is tb1, and the duration of the first pronunciation segment of reference c is tc1. The average duration of the same pronunciation segment from multiple references can be used as the reference duration for that segment. For example, (ta1+tb1+tc1) / 3 can be used as the reference duration for the first pronunciation segment. This allows us to obtain the reference durations for multiple pronunciation segments from the same reference subject.

[0099] Subsequently, similar to the subjects mentioned above, videos of any test subject reading the target content can be collected, and the duration of multiple pronunciation segments corresponding to the test subject can be obtained according to the method described in the aforementioned embodiments. Then, the coherence scores corresponding to the multiple pronunciation segments of the test subject can be obtained according to the above formulas (1) and (2).

[0100] The supervision information used during training is then determined as follows: The training of the coherence score weight prediction model requires two types of supervision information: one is the total coherence score corresponding to the tester, and the other is supervision information reflecting the length of sentences.

[0101] Specifically, based on the test taker's subjective feeling about the actual quality of the target content, a total score for the coherence during the reading process can be manually assigned. For example, if the total score is 100 points, the actual score might be 95 or 80 points.

[0102] For monitoring information reflecting sentence length, optionally, after obtaining the duration of each of the test taker's multiple pronunciation segments, the test taker can label each of the multiple pronunciation segments with a corresponding sentence length category according to different preset duration ranges. Several sentence length categories can be preset, such as: very long sentences, medium-long sentences, and short sentences, where each category corresponds to a set duration range. Thus, after obtaining the coherence scores corresponding to each of the test taker's multiple pronunciation segments, the sentence length category corresponding to each of the test taker's multiple pronunciation segments can be used as the sentence length category corresponding to each coherence score.

[0103] For supervisory information reflecting sentence length, optionally, after obtaining the coherence scores for multiple pronunciation segments of the test taker, the test taker can be labeled with the corresponding sentence length category based on the values ​​of the multiple coherence scores. Simply put, for short sentences, the test taker's pronunciation duration is not significantly different from the reference subject's, resulting in a higher coherence score; while for long sentences, the test taker's pronunciation duration differs significantly from the reference subject's, resulting in a lower coherence score. Based on this, different correspondences between score ranges and sentence length categories can be set, thereby enabling the labeling of sentence length categories corresponding to multiple coherence scores.

[0104] The coherence scores, labeled with the above sentence length categories, and the total coherence score are input into the coherence score weight prediction model. The coherence score weight prediction model determines the weight corresponding to each coherence score based on the principle of assigning higher weights to the coherence scores corresponding to long sentences and lower weights to the coherence scores corresponding to short sentences, so that the weighted sum of the multiple coherence scores after the weighting is close to the total coherence score.

[0105] The following is combined Figure 4 The example illustrates the training process of the above coherent score weight prediction model.

[0106] exist Figure 4In this scenario, suppose a test taker pronounces a target content and identifies three pronunciation segments: Y1, Y2, and Y3. Assume these segments correspond to sentence length categories of short, medium-long, and very long sentences, respectively. Also assume the total coherence score, based on the test taker's pronunciation, is 95 points, with coherence scores for the three segments as s1, s2, and s3. These three scores, labeled with their respective sentence length categories, along with the total coherence score, are input into a coherence score weight prediction model. The model outputs weights for these three scores: w1, w2, and w3. These weights are then used to perform a weighted sum. In practice, the order of these weights might be: w1... <w2<w3。

[0107] The above describes the process of judging the fluency of a subject's pronunciation. The following describes the process of judging the correctness of a subject's pronunciation.

[0108] In one optional embodiment, a vision-based pronunciation accuracy assessment scheme is provided. Specifically, firstly, the lip-synchrotron intensity corresponding to a first pronunciation segment of the subject and the lip-synchrotron intensity corresponding to a second pronunciation segment of the target reference are obtained, wherein the sequence number of the first pronunciation segment in the subject's multiple pronunciation segments is the same as the sequence number of the second pronunciation segment in the target reference's multiple pronunciation segments. Then, the lip-synchrotron intensity corresponding to the subject's first pronunciation segment is compared with the lip-synchrotron intensity corresponding to the target reference's second pronunciation segment to determine the pronunciation accuracy score of the subject's first pronunciation segment. Finally, the subject's total pronunciation accuracy score is determined based on the multiple pronunciation accuracy scores corresponding to the subject's multiple pronunciation segments.

[0109] In the above embodiments, the subject's first pronunciation segment and the target reference's second pronunciation segment are the same pronunciation segment, such as their respective first pronunciation segment, their respective second pronunciation segment, and so on. In this embodiment, the correctness score corresponding to a pronunciation segment can be set to 1 or 0.

[0110] As mentioned earlier, a pronunciation segment often corresponds to multiple frames, and each frame corresponds to the intensity of multiple lip-sync units. Optionally, for the first and second pronunciation segments, the average value of the multiple intensities of each lip-sync unit corresponding to each of the multiple frames can be calculated to obtain the average intensity of the multiple lip-sync units corresponding to the first and second pronunciation segments respectively. For example, assuming the first pronunciation segment includes frame P1 and the second pronunciation segment includes frame P2, and assuming there are n lip-sync units: AU1, AU2, ..., AUn, then the average value of the P1 intensities of AU1 corresponding to frame P1 is calculated to obtain the average intensity of AU1 corresponding to the first pronunciation segment. Similarly, the average intensity of each of the n lip-sync units in the first pronunciation segment is obtained. Similarly, the same calculation is performed on frame P2 in the second pronunciation segment to obtain the average intensity of each of the n lip-sync units in the second pronunciation segment.

[0111] Next, the average intensity of each of the n mouth shape action units in the first pronunciation segment is compared with the average intensity of each of the n mouth shape action units in the second pronunciation segment to determine the pronunciation accuracy score of the first pronunciation segment. In summary, the closer the average intensity of each of the n mouth shape action units in the first pronunciation segment is to the average intensity of each of the n mouth shape action units in the second pronunciation segment, the closer the mouth shapes in the first and second pronunciation segments are, and the higher the pronunciation accuracy score; conversely, the closer they are to each other, the lower the score.

[0112] Optionally, the average intensity of each of the n mouth-shaped motion units corresponding to the first pronunciation segment can be calculated, i.e., the mean of these n values ​​A1. The average intensity of each of the n mouth-shaped motion units corresponding to the second pronunciation segment can be calculated, i.e., the mean of these n values ​​A2. If the difference between A1 and A2 is less than a set threshold, the pronunciation of the first and second pronunciation segments is considered very similar, and the pronunciation accuracy score of the first pronunciation segment is determined to be 1; otherwise, it is 0. Alternatively, the difference in average intensity of the same mouth-shaped motion unit in the first and second pronunciation segments can be calculated separately. If the difference in intensity of all n mouth-shaped motion units is less than a set threshold, the pronunciation accuracy score of the first pronunciation segment is determined to be 1; otherwise, if there is a mouth-shaped motion unit with an intensity difference greater than or equal to the set threshold, the pronunciation accuracy score of the first pronunciation segment is determined to be 0.

[0113] In another optional embodiment, an audio-based pronunciation accuracy assessment scheme is provided. Specifically, firstly, text recognition processing is performed on a first pronunciation segment of the subject and a second pronunciation segment of the target reference, respectively, to obtain a first text content corresponding to the first pronunciation segment and a second text content corresponding to the second pronunciation segment. The sequence number of the first pronunciation segment among the subject's multiple pronunciation segments is the same as the sequence number of the second pronunciation segment among the target reference's multiple pronunciation segments. Then, the first text content and the second text content are compared to determine the pronunciation accuracy score of the subject's first pronunciation segment. Finally, the subject's total pronunciation accuracy score is determined based on the multiple pronunciation accuracy scores corresponding to the subject's multiple pronunciation segments.

[0114] The aforementioned text recognition processing is equivalent to speech recognition processing, which yields the text content corresponding to each pronunciation segment. Based on the ability to recognize each individual character within a pronunciation segment, the accuracy score for a pronunciation segment can be defined as the number of correctly pronounced characters within that segment.

[0115] For example, if the first text corresponding to the first pronunciation segment is the same as the second text corresponding to the second pronunciation segment, and this text contains 5 characters, then the subject's pronunciation accuracy score for the first pronunciation segment is determined to be 5. If the first text corresponding to the first pronunciation segment is different from the second text corresponding to the second pronunciation segment, and assuming that the first and third characters in the first text are the same as the first and third characters in the second text, while the other characters are different, then the subject's pronunciation accuracy score for the first pronunciation segment is determined to be 2.

[0116] After determining the pronunciation accuracy scores for multiple pronunciation segments of the subject according to the method provided in the above embodiments, the multiple pronunciation accuracy scores can be summed together to obtain the total pronunciation accuracy score. When the health score only considers the total pronunciation accuracy score, the subject's pronunciation quality can be determined based on the comparison result of the total pronunciation accuracy score and a set threshold.

[0117] Similar to the aforementioned weight prediction of multiple pronunciation coherence scores, in this embodiment of the invention, a correctness score weight prediction model can also be pre-trained to predict the weight corresponding to each pronunciation correctness score, so as to perform a weighted summation of multiple pronunciation correctness scores based on the weight prediction results to obtain the total pronunciation correctness score.

[0118] The correctness score weight prediction model can also be implemented as a convolutional neural network model consisting of multiple convolutional layers. Similar to the coherence score weight prediction model, the correctness score weight prediction model is also trained to assign higher weights to long sentences and lower weights to short sentences.

[0119] To train a correctness score weight prediction model with the aforementioned capabilities, a training set is first required. In this embodiment of the invention, for ease of description, only the target content mentioned above is used as the training material as an example; in reality, the training material can include multiple different contents. Videos of a reference person reading the target content can be collected beforehand, and multiple pronunciation segments corresponding to each reference person are obtained according to the method described in the preceding embodiments. Speech recognition is then performed on each pronunciation segment to obtain the text content contained in each pronunciation segment. Next, videos of any test person reading the target content can be collected, and multiple pronunciation segments corresponding to that test person are obtained according to the method described in the preceding embodiments. Speech recognition is then performed on each pronunciation segment to obtain the text content contained in each pronunciation segment. It should be noted that the video mentioned here refers to a video containing audio. Finally, according to the correctness score calculation method described above, the pronunciation correctness score corresponding to each pronunciation segment of the test person is obtained.

[0120] Training the correctness score weight prediction model also requires the use of two types of supervision information: one is the total correctness score of the tester, and the other is supervision information reflecting the length of the sentences, i.e., the sentence length category.

[0121] The total score for accuracy can be the number of correct words determined subjectively by the test taker after hearing the test taker read the target content aloud.

[0122] For the annotation of sentence length categories, please refer to the relevant introduction to the coherence score weight prediction model. In addition, sentence length categories can also be annotated based on the number of characters corresponding to each pronunciation segment after speech recognition. The correspondence between different numbers of characters and sentence length categories is pre-defined, thereby enabling the annotation of sentence length categories.

[0123] The training process of the correctness score weight prediction model can be referenced from the training process of the coherence score weight prediction model, and will not be elaborated here.

[0124] In summary, after obtaining the total score for pronunciation fluency and the total score for pronunciation accuracy of the subjects, the sum of the two total scores can be determined as the subjects' health score. If the health score is greater than the set threshold, the subjects' pronunciation quality is determined to be good; otherwise, the pronunciation quality is poor.

[0125] As described above, the pronunciation quality testing scheme provided in this embodiment of the invention can be applied to pronunciation testing scenarios such as testing the recovery status of aphasic individuals and testing the learning effectiveness of language learners. Taking the testing of the recovery status of aphasic individuals as an example, combined with... Figure 5 The example illustrates the implementation process of this solution.

[0126] like Figure 5 As shown, taking a user with aphasia as an example, a corresponding testing application can be downloaded and used on the user's terminal. This application provides different test content (or test text). In fact, these test contents can be divided into several categories according to the severity of the subject's aphasia, such as... Figure 5 The text indicates mild, moderate, and severe aphasia. It's understandable that the test content for severely aphasic users will be less difficult to pronounce and shorter than that for moderate and mild aphasic users. Within each category, at least one test item can be set, and for each test item, at least one shadowing video can be provided. The shadowing video refers to a video recording of a person acting as a reference (or imitation target) while pronouncing a particular test item. Figure 5 The text illustrates some test content and the corresponding reading videos for each test content.

[0127] The aphasic user or other personnel conducting the test select a test content and a corresponding video for repetition (corresponding to the second video mentioned above, assuming it's video 1 in the diagram) based on the user's condition. The video is then played so the user can hear how the participant pronounces the content. The user then looks at the test content, imitates the participant's pronunciation, and reads the content aloud. The user's reading process is captured on video, resulting in the first video mentioned above. Figure 5 The text in the image represents the user's video.

[0128] exist Figure 5 In this example, let's assume test content 1 is: "Hello, please follow the teacher and read the following content, paying attention to the punctuation." Based on the processing described in the previous embodiment, and considering the intensity changes of the lip-sync units corresponding to each frame in each video, we determine that multiple pronunciation segments in the follow-along video 1 are Y1, Y2, and Y3, and multiple pronunciation segments in the user video for the aphasic user are Z1, Z2, Z3, Z4, and Z5. For ease of understanding, let's assume the pauses corresponding to Y1, Y2, and Y3 are: "Hello / Please follow the teacher and read the following content / Pay attention to the punctuation." Let's assume the pauses corresponding to Z1, Z2, Z3, Z4, and Z5 are: "Hello / Please follow the teacher / Read the following content / Pay attention / Punchage."

[0129] After obtaining the above-mentioned multiple pronunciation segments, the corresponding pronunciation coherence scores can be calculated separately for the multiple pronunciation segments of the aphasic user. Based on the pronunciation coherence score calculation formula introduced above, the pronunciation coherence scores corresponding to the pronunciation segments Z4 and Z5 are 0, and the pronunciation coherence scores corresponding to the pronunciation segments Z1, Z2, and Z3 are determined according to the aforementioned formula (1). Furthermore, the total pronunciation coherence score SM1 corresponding to the aphasic user and the test content can be obtained.

[0130] After obtaining the above-mentioned multiple pronunciation segments, the corresponding pronunciation correctness scores can also be calculated separately for the multiple pronunciation segments of the aphasic user. Here, it is assumed that the text content recognition results corresponding to Z1, Z2, Z3, Z4, and Z5 are: Hello / Please follow the teacher / Read a content / Pay attention / Sentence break position. That is, three words are mispronounced. According to the pronunciation correctness scores corresponding to each pronunciation segment, the total pronunciation correctness score SM2 corresponding to the aphasic user and the test content can be obtained.

[0131] After that, the sum of the total pronunciation coherence score SM1 and the total pronunciation correctness score SM2 is used as the health score of the aphasic user. If this health score is greater than the set threshold, it is determined that the pronunciation quality of the aphasic user is good.

[0132] Figure 6 It is a flowchart of a pronunciation quality test method provided by an embodiment of the present invention. As Figure 6 shown, the method includes the following steps:

[0133] 601. Obtain multiple third videos of multiple reference persons reading the target content.

[0134] 602. Detect the intensity of oral movement units for multiple third images in the target third video respectively, obtain the intensity of oral movement units corresponding to each of the multiple third images, and determine the oral feature vector corresponding to the target third video according to the intensity of oral movement units corresponding to each of the multiple third images.

[0135] Among them, the target third video is any one of the multiple third videos.

[0136] 603. Cluster the oral feature vectors corresponding to the multiple third videos respectively to obtain multiple clustering results.

[0137] 604. Obtain the first video of the subject reading the target content, detect the intensity of oral movement units for multiple first images in the first video respectively to obtain the intensity of oral movement units corresponding to each of the multiple first images, and determine the oral feature vector corresponding to the first video according to the intensity of oral movement units corresponding to each of the multiple first images.

[0138] 605. Determine the target clustering result corresponding to the lip shape feature vector of the first video from multiple clustering results, and determine the target reference from the references corresponding to the target clustering result.

[0139] 606. Obtain a second video of the target reference reading the target content, and perform lip-sync unit intensity detection on multiple frames of the second image in the second video to obtain the lip-sync unit intensity corresponding to each frame of the second image.

[0140] 607. Compare the intensity of the lip-sync units corresponding to each of the first frames of multiple images with the intensity of the lip-sync units corresponding to each of the second frames of multiple images to determine the pronunciation quality of the subject.

[0141] The pronunciation quality testing scheme provided in this embodiment also considers the personalized pronunciation characteristics of different individuals. Different individuals will have slight differences in mouth shape when reading the same text or pronouncing the same sound. This embodiment fully considers these differences and realizes personalized pronunciation quality testing for the current subject.

[0142] In summary, the core idea of ​​the personalized pronunciation quality testing scheme in this embodiment is as follows: For the same target content, pronunciation videos of multiple references can be collected. Different references have different personalized pronunciation characteristics. Based on these personalized pronunciation characteristics, multiple references are clustered to obtain multiple clustering results. For the current subject, a target clustering result matching their pronunciation characteristics is selected from the multiple clustering results. This target clustering result is then used to complete the pronunciation quality test for the subject. The aforementioned personalized pronunciation characteristics can be reflected by the intensity of lip-sync units corresponding to each frame of the pronunciation video.

[0143] Specifically, multiple third videos (i.e., multiple pronunciation videos) are acquired from multiple references reading the target content. Then, for each third video, samples are taken to obtain multiple frames of images (called multi-frame third images). Taking any target third video as an example, lip-syncing motion unit intensity is detected for each frame of the target third video to obtain the lip-syncing motion unit intensity corresponding to each of the multi-frame third images. Then, based on the lip-syncing motion unit intensities corresponding to each of the multi-frame third images, the lip-syncing feature vector corresponding to the target third video is determined. The lip-syncing feature vector corresponding to the target third video can be the concatenation result of the lip-syncing motion unit intensities corresponding to each of the multi-frame third images in the target third video. Assuming the target third video contains m frames and the number of lip-syncing motion units is n, then the lip-syncing feature vector is an m*n dimensional feature vector.

[0144] Then, a pre-defined clustering algorithm (such as k-means) was used to cluster the lip-sync feature vectors corresponding to the multiple third videos, resulting in multiple clustering results. It is understandable that references clustered together often have similar pronunciation characteristics.

[0145] In practical applications, to ensure broad coverage, the selection of multiple references can fully consider individual differences, such as age, gender, body weight, geographical distribution, and so on.

[0146] Since the dimensionality of the lip-sync feature vectors corresponding to each third video is often high, dimensionality reduction can be performed first during clustering calculations to reduce the lip-sync feature vectors of each third video to a set dimension. After obtaining the above multiple clustering results, the cluster center feature vector corresponding to each clustering result can be determined. For example, the mean of multiple lip-sync feature vectors contained in the same clustering result can be calculated, and the mean result can be determined as the cluster center feature vector.

[0147] For the current subjects, a first video of the subjects reading the same target content is acquired. Lip movement unit intensity (LMU) detection is performed on multiple frames of the first image in the first video to obtain the LMU intensity for each frame. Based on the LMU intensity of each frame, the lip movement feature vector corresponding to the first video is determined. Then, the similarity between the lip movement feature vector corresponding to the first video and the feature vectors of each cluster center can be calculated. The clustering result corresponding to the highest similarity is determined as the target clustering result. The target reference can then be determined from the references corresponding to the target clustering result. The similarity can be represented by some distance, such as cosine distance.

[0148] The process of determining the target reference from the references corresponding to the target clustering results can involve randomly selecting one from among the references corresponding to the target clustering results. In this case, the target reference is chosen because it shares more similar pronunciation characteristics with the subject, allowing the subject to better follow along with the target reference's pronunciation and thus more objectively and accurately assess the subject's pronunciation quality.

[0149] The training scheme for the face detection model mentioned above is described below.

[0150] Figure 7 A flowchart of a face detection model training method provided in an embodiment of the present invention is shown below. Figure 7 As shown, the training method may include the following steps:

[0151] 701. Obtain the first face sample image, input the first face sample image into the first face detection model, and obtain multiple face reconstruction parameters corresponding to the first face sample image. The multiple face reconstruction parameters include the intensity of the lip movement unit.

[0152] 702. Input the various face reconstruction parameters corresponding to the first face sample image into the face reconstruction model to obtain the first face 3D model, and generate the first face reconstruction image based on the various face reconstruction parameters and the first face 3D model.

[0153] 703. Determine the loss function value based on the first face sample image and the first face reconstruction image, and train the first face detection model based on the loss function value.

[0154] In summary, the training of the face detection model is carried out under the task of face reconstruction. In this embodiment, the trained face detection model is called the first face detection model, which can be implemented as an encoder network model composed of multiple feature extraction layers. The face reconstruction model can be a three-dimensional deformable face model (3DMM).

[0155] The first face detection model is used to detect various face reconstruction parameters, including the intensity of lip-sync action units. The face reconstruction model is then used to reconstruct face images based on these parameters.

[0156] Specifically, during training, the first face detection model takes a 2D first face sample image as input and outputs various face reconstruction parameters, including but not limited to: identity parameters, expression parameters, mouth parameters, texture parameters, pose parameters, and lighting parameters. Among these, the identity parameters control the shape of the face and can therefore also be called shape parameters. The mouth parameters represent the intensity of multiple mouth action units. The goal of training the first face detection model is to make the mouth parameters as accurate as possible; the other parameters are necessary for face reconstruction during self-supervised training.

[0157] The first face sample image is only one of several training sample images for the first face detection model, and is used as an example to illustrate the training process. The training cutoff condition for the first face detection model can be that all training sample images have been used, the number of training rounds has reached a set value, or the accuracy of the model has reached a set requirement.

[0158] After obtaining the aforementioned face reconstruction parameters, these parameters can be input into the face reconstruction model to obtain the output result. Specifically, during the face reconstruction process, the face reconstruction model first establishes a 3D face model (i.e., a mesh model) based on the various face reconstruction parameters. The established first 3D face model can present the facial expressions, poses, and lip shapes of the face in the first face sample image, but it does not have accurate texture features. Simply put, it first reconstructs the 3D contour of the face in the first face sample image. Then, based on the first 3D face model and the various face reconstruction parameters, a 2D first face reconstruction image is generated. During the generation of the first face reconstruction image, the aforementioned various face reconstruction parameters, such as texture parameters and lighting parameters, are fully utilized to ensure that the texture and brightness features presented in the first face reconstruction image match those of the first face sample image.

[0159] Additionally, it's important to note that during the face reconstruction process, various deformation bases, or blended shape (BS) models, are used for each face reconstruction parameter. The coefficients of each BS model are adjustable, resulting in different shapes for the corresponding BS models. In fact, a BS model is also a three-dimensional geometric model. Regarding lip shape parameters, i.e., the intensity of multiple lip shape action units, since multiple lip shape action units are set, each lip shape action unit corresponds to a specific lip shape BS model, and the intensity of each lip shape action unit is the coefficient of the corresponding lip shape BS model.

[0160] In this embodiment, each of the above-mentioned BS models is a preset general model, that is, it is independent of the face (or training object) in the face sample image.

[0161] In the process of face reconstruction, the use of BS models corresponding to various face reconstruction parameters can be simply described as follows: the coefficients of the corresponding BS model are determined according to a certain face reconstruction parameter, and the BS model adjusted to the corresponding coefficients is superimposed on the preset neutral face 3D model. Thus, by superimposing the BS models corresponding to various face reconstruction parameters onto the neutral face 3D model, the first face 3D model can be obtained.

[0162] In summary, the output of the face reconstruction model can be considered to consist of two parts: a first 3D face model and a first reconstructed face image. Based on this output, the values ​​of several predetermined loss functions can be calculated, and then backpropagation can be used to adjust the parameters of the trained first face detection model.

[0163] Optionally, the loss function may include at least one of the following: an identity or shape loss function (the first loss function below), a keypoint loss function (the second loss function below), and a photometric loss function (the third loss function below). Based on these loss functions, the face detection model can achieve better performance.

[0164] Specifically, the first face features corresponding to the first face sample image and the second face features corresponding to the first face reconstruction image can be extracted respectively, and the first loss function value can be determined based on the first face features and the second face features.

[0165] Specifically, the first facial key points can be extracted from the first facial sample image, and the second facial key points can be obtained from the first facial 3D model. The second loss function value is determined based on the first facial key points and the second facial key points.

[0166] The pixel values ​​of the first face sample image and the first face reconstruction image can be compared to determine the value of the third loss function.

[0167] Specifically, a pre-trained face feature detection model can be used to extract the aforementioned face features. In this embodiment, face features refer to feature parameters that can distinguish different faces. The first loss function value can be determined by calculating the distance or similarity between the first face feature and the second face feature.

[0168] Similarly, a pre-trained facial landmark detection model can be used to extract the first facial landmark. In practical applications, the first face detection model can also be equipped with the function of detecting this facial landmark, thereby obtaining the first facial landmark detected by the first face detection model from the first face sample image. In fact, several types of facial landmarks can be predefined, such as landmarks of different parts of the face, such as eyebrows, nose, mouth, eyes, ears, and facial contours. The purpose of facial landmark detection is to determine the position coordinates of each facial landmark in the input image (the first face sample image in this embodiment).

[0169] During the creation of the first 3D face model, the face reconstruction model actually generates a first 3D face model composed of several triangular facets. Each triangular facet corresponds to a number, and the first 3D face model can be generated according to a set numbering order and positional arrangement. The numbers of the triangular facets corresponding to multiple pre-defined facial key points can be pre-labeled. Therefore, after generating the first 3D face model, the triangular facets corresponding to multiple facial key points can be found, each triangular facet corresponding to a positional coordinate. This allows the second facial key point to be obtained from the first 3D face model. Then, the difference in positional coordinates between the first and second facial key points can be compared, for example, by calculating a certain distance, to determine the value of the second loss function.

[0170] In simple terms, the third loss function value determines the pixel differences between the first face sample image and the reconstructed first face image. Since both images are the same size, the third loss function value can be determined by calculating the sum of the differences in pixel values ​​at corresponding pixel locations in the two images, or by accumulating the number of pixel locations with inconsistent pixel values.

[0171] When the loss function of the first face detection model adopts at least two of the above three types, the values ​​of the at least two loss functions adopted can be summed to obtain the total loss, and the parameters of the first face detection model can be adjusted based on the total loss value.

[0172] Understandably, after completing the above training process, in the pronunciation quality detection process, only the lip shape parameters detected by the first face detection model for the input image, i.e., the lip shape action unit intensity, need to be obtained, and other face reconstruction parameters do not need to be used.

[0173] To facilitate a more intuitive understanding of the above training process, Figure 8 The above training process is illustrated in the diagram.

[0174] Figure 9 A flowchart of another face detection model training method provided in an embodiment of the present invention is shown below. Figure 9 As shown, the training method may include the following steps:

[0175] 901. Obtain the first face sample image, input the first face sample image into the first face detection model, and obtain multiple face reconstruction parameters corresponding to the first face sample image. The multiple face reconstruction parameters include the intensity of the lip movement unit.

[0176] 902. Input the various face reconstruction parameters corresponding to the first face sample image into the face reconstruction model to obtain the first face 3D model, and generate the first face reconstruction image based on the various face reconstruction parameters and the first face 3D model.

[0177] 903. Determine the loss function value based on the first face sample image and the first face reconstruction image, and train the first face detection model based on the loss function value.

[0178] 904. Obtain a second face sample image, input the second face sample image into the second face detection model, and obtain multiple face reconstruction parameters corresponding to the second face sample image. Among the multiple face reconstruction parameters are identity parameters. The second face detection model is the model obtained by training the first face detection model to meet the set cutoff conditions.

[0179] 905. Input the identity parameters into the face reconstruction model to obtain the second face 3D model. Based on the second face 3D model, the preset neutral face 3D model and the general lip shape deformation base model, determine the lip shape deformation base model of the trainee corresponding to the second face sample image through the preset deformation transfer algorithm.

[0180] 906. Optimize and train the second face detection model based on the trainee's lip shape deformation baseline model.

[0181] In this embodiment, a training set containing several face sample images is pre-constructed, and the first and second face sample images mentioned above are any sample images in the training set.

[0182] In this embodiment, the training process of the face detection model is divided into three stages: the first stage corresponds to steps 901-903 above, which is the aforementioned Figure 7 In the training process of the illustrated embodiment, during the training stage, the initial first face detection model is trained until it meets the set cutoff condition, and the face detection model at this time is called the second face detection model; the second stage corresponds to steps 904-905 above, which is used to obtain a personalized lip shape BS model; the third stage corresponds to step 906 above, which is used to optimize the training of the second face detection model based on the personalized lip shape BS model, so as to obtain a third face detection model that has been optimized and trained to meet the cutoff condition.

[0183] Combination Figure 10 This illustrates the three training phases described above. Figure 10 In this context, we assume that the loss functions used in the first and third stages are the three loss functions exemplified above: identity or shape loss function, keypoint loss function, and photometric loss function.

[0184] As mentioned earlier, the various BS models used in the first stage of training the first face detection model are all general-purpose models, i.e., models independent of individuals. This embodiment fully considers individual differences; that is, different individuals will have slight differences in their lip movements when reading the same text or pronouncing the same sound. Based on this, a personalized lip-shape BS model is introduced to optimize the training process of the face detection model, achieving the goal of predicting lip-shape action unit strengths as consistently as possible for the same pronunciation lip shape from different individuals. Here, the so-called personalized lip-shape BS model refers to acquiring lip-shape BS models for different individuals (i.e., trainees in different face sample images).

[0185] It should be noted that, since only the lip shape parameters output by the face detection model are needed during the pronunciation quality test, this embodiment only needs to obtain a personalized lip shape BS model, and does not need to obtain personalized BS models for other types of parameters.

[0186] Furthermore, it's important to note that the ultimate goal of introducing personalized mouth shape BS models is to ensure that the face detection model predicts consistent lip action unit strengths for the same pronunciation lip shape across different individuals. This allows for more accurate assessment of the subject's pronunciation quality by comparing the lip action unit strengths detected in the pronunciation videos of the subject and the target reference. However, in the first stage of training, the same generic lip shape BS model was used for face reconstruction across all face sample images, failing to reflect individual differences. This resulted in more pronounced inconsistencies in the lip action unit strengths predicted by the trained second face detection model for the same lip shape across different individuals. Therefore, after completing the first stage of training and enabling the second face detection model to possess preliminary lip shape parameter detection capabilities, the second stage utilizes the second face detection model to learn individualized lip shape BS models based on its output.

[0187] Specifically, in the second stage, the second face detection model and face reconstruction model trained in the first stage are used to reconstruct faces from the face sample images in the training set. Taking the second face sample image as an example, the second face sample image is input into the second face detection model to obtain various face reconstruction parameters corresponding to the second face sample image, including identity parameters. Then, the identity parameters are input into the face reconstruction model to obtain the second face 3D model. Subsequently, based on the second face 3D model, a preset neutral face 3D model, and a general lip shape BS model, a preset deformation transfer algorithm is used to determine the trainee's lip shape BS model corresponding to the second face sample image.

[0188] In the second stage of face reconstruction, as mentioned above, optionally, to simplify the process, only the identity parameter from the various face reconstruction parameters output by the second face detection model can be used to reconstruct the face and obtain the second 3D face model. This is because the identity parameter is the main parameter used to distinguish the face shapes of different individuals, and obtaining the personalized lip-shape BS model involves acquiring the corresponding lip-shape BS model for each individual. Of course, using multiple face reconstruction parameters to generate the second 3D face model is also possible. Then, based on a deformation transfer algorithm, the lip-shape BS model corresponding to this second 3D face model is obtained through deformation transfer, serving as the lip-shape BS model corresponding to the trainer in the second face sample image.

[0189] In simple terms, the deformation transfer algorithm requires three types of input information: the source object A, the object A' that undergoes deformation, the target object B, and the object B' that is transformed from the target object B. Specifically, the deformation method that transforms A into A' needs to be applied to the target object B to obtain the result B' after performing that deformation process on the target object B.

[0190] Based on the principle of the deformation transfer algorithm described above, in this embodiment, the source object A is a preset neutral 3D face model, the object A' corresponding to A and undergoing deformation is a general lip-shape BS model, and the target object B is a second 3D face model. The deformation transfer algorithm can obtain B' based on these three input pieces of information: the lip-shape BS model of the trainer corresponding to the second face sample image.

[0191] Assuming the training set contains 500 trainees, the second stage will allow us to learn the lip-sync model for each of these 500 trainees.

[0192] Then, in the third stage, the second face detection model trained in the first stage was fine-tuned using the BS model of each trainee's mouth shape, resulting in the final third face detection model.

[0193] Taking the second face sample image as an example, after obtaining various face reconstruction parameters corresponding to the second face sample image through the second face detection model, in the third stage, the various face reconstruction parameters corresponding to the second face sample image, along with the trainer's lip-shape BS model, are input into the face reconstruction model to obtain the third face 3D model. Based on the third face 3D model and the various face reconstruction parameters, the second face reconstruction image is generated. Then, the loss function value is determined based on the second face sample image and the second face reconstruction image, and the second face detection model is trained based on the determined loss function value.

[0194] In this embodiment, considering the differences in mouth shapes among different individuals, a personalized mouth shape BS model is introduced to optimize and train the face detection model used for mouth shape action unit intensity detection. This allows the optimized face detection model to detect more accurate mouth shape action unit intensity, thereby helping to improve the accuracy of the test results for the subject's pronunciation quality.

[0195] The video recognition method provided in this invention can be executed in the cloud, where multiple computing nodes (cloud servers) can be deployed. Each computing node has processing resources such as computing and storage. In the cloud, multiple computing nodes can be organized to provide a certain service; of course, a single computing node can also provide one or more services. The cloud provides this service by providing an external service interface, which users call to use the corresponding service.

[0196] According to the solution provided in this embodiment of the invention, the cloud can provide a service interface for pronunciation quality testing. Users invoke this service interface through their user devices to trigger a pronunciation quality test request to the cloud. This request includes a first video of a subject reading the target content and a second video of a target reference reading the target content. The cloud determines the computing node that responds to the request and utilizes the processing resources in that computing node to perform the following steps:

[0197] Lip-sync unit intensity detection is performed on multiple frames of first images in the first video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of first images.

[0198] The intensity of lip-sync unit is detected for each of the multiple frames of the second image in the second video to obtain the intensity of the lip-sync unit corresponding to each of the multiple frames of the second image.

[0199] The lip-sync intensity of the subject is determined by comparing the lip-sync intensity of each of the first frames of the multi-frame first image with the lip-sync intensity of each of the second frames of the multi-frame second image; wherein, the lip-sync intensity of any image refers to the intensity coefficient of each of the multiple preset lip-sync units when forming the lip shape in any image.

[0200] The subject's pronunciation quality is fed back to the user's device.

[0201] The above execution process can be referred to the relevant descriptions in the other embodiments mentioned above, and will not be repeated here.

[0202] For ease of understanding, combined with Figure 11 To illustrate with an example. Users can... Figure 11The illustration shows user device E1 calling the pronunciation quality test service to upload a service request containing a first video of a subject reading the target content and a second video of a target reference reading the target content. The service interface for the user to call this service can take the form of a Software Development Kit (SDK) or an Application Programming Interface (API). Figure 11 The diagram illustrates an API interface scenario. In the cloud, as shown, assume that a pronunciation quality testing service is provided by service cluster E2, which includes at least one computing node. Upon receiving the request, service cluster E2 executes the steps described in the previous embodiment to obtain the subject's pronunciation quality and then sends the subject's pronunciation quality to user device E1.

[0203] The following will describe in detail one or more embodiments of the pronunciation quality testing apparatus of the present invention. Those skilled in the art will understand that these apparatuses can all be configured using commercially available hardware components through the steps taught in this solution.

[0204] Figure 12 This is a schematic diagram of the structure of a sound quality testing device provided in an embodiment of the present invention, as shown below. Figure 12 As shown, the device includes: an acquisition module 11, a detection module 12, and a determination module 13.

[0205] The acquisition module 11 is used to acquire a first video of a subject reading the target content and a second video of a target reference reading the target content.

[0206] The detection module 12 is used to perform lip-sync unit intensity detection on multiple frames of first images in the first video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of first images; and to perform lip-sync unit intensity detection on multiple frames of second images in the second video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of second images.

[0207] The determining module 13 is used to compare the intensity of the lip-shape action unit corresponding to each of the multiple first frames with the intensity of the lip-shape action unit corresponding to each of the multiple second frames to determine the pronunciation quality of the subject; wherein, the intensity of the lip-shape action unit corresponding to any image refers to the intensity coefficient corresponding to the multiple preset lip-shape action units when forming the lip shape in any image.

[0208] Optionally, the apparatus further includes: a clustering module, configured to acquire multiple third videos in which multiple references read the target content; perform lip-sync unit intensity detection on multiple frames of third images in the target third video to obtain the lip-sync unit intensity corresponding to each of the multiple frames of third images, wherein the target third video is any one of the multiple third videos; determine the lip-sync feature vector corresponding to the target third video based on the lip-sync unit intensity corresponding to each of the multiple frames of third images; cluster the lip-sync feature vectors corresponding to each of the multiple third videos to obtain multiple clustering results; determine the lip-sync feature vector corresponding to the first video based on the lip-sync unit intensity corresponding to each of the multiple frames of first images; determine the target clustering result corresponding to the lip-sync feature vector corresponding to the first video from the multiple clustering results; and determine the target reference from the references corresponding to the target clustering result.

[0209] Optionally, the determining module 13 is specifically configured to: determine multiple pause points of the subject based on the degree of change in the intensity of the lip-sync units corresponding to each of the multiple first images; determine multiple pause points of the target reference based on the degree of change in the intensity of the lip-sync units corresponding to each of the multiple second images; determine multiple pronunciation segments of the subject based on the multiple pause points of the subject; and determine multiple pronunciation segments of the target reference based on the multiple pause points of the target reference; wherein, a pronunciation segment includes multiple images between adjacent pause points; compare the multiple pronunciation segments of the subject and the multiple pronunciation segments of the target reference to determine the subject's total pronunciation coherence score and / or total pronunciation accuracy score, so as to determine the subject's pronunciation quality based on the total pronunciation coherence score and / or the total pronunciation accuracy score.

[0210] Optionally, the determining module 13 is specifically used to: determine the duration of each of the multiple pronunciation segments of the subject; determine the duration of each of the multiple pronunciation segments of the target reference; determine multiple pronunciation coherence scores of the subject based on the duration of each of the multiple pronunciation segments of the subject and the duration of each of the multiple pronunciation segments of the target reference, wherein the multiple pronunciation coherence scores correspond to the multiple pronunciation segments of the subject; and determine the pronunciation quality of the subject based on the multiple pronunciation coherence scores.

[0211] Optionally, the determining module 13 is specifically used to: input the plurality of pronunciation coherence scores into the coherence score weight prediction model to obtain the weights of the plurality of pronunciation coherence scores; determine the total pronunciation coherence score of the subject based on the weights of the plurality of pronunciation coherence scores; and determine the pronunciation quality of the subject based on the total pronunciation coherence score of the subject, wherein the coherence score weight prediction model is a neural network model.

[0212] Optionally, the determining module 13 is specifically configured to: acquire the lip-sync unit intensity corresponding to the first pronunciation segment of the subject and the lip-sync unit intensity corresponding to the second pronunciation segment of the target reference, wherein the sequence number of the first pronunciation segment in the subject's multiple pronunciation segments is the same as the sequence number of the second pronunciation segment in the target reference's multiple pronunciation segments; compare the lip-sync unit intensity corresponding to the first pronunciation segment of the subject with the lip-sync unit intensity corresponding to the second pronunciation segment of the target reference to determine the pronunciation correctness score of the subject's first pronunciation segment; and determine the pronunciation quality of the subject based on the subject's multiple pronunciation correctness scores, wherein the multiple pronunciation correctness scores correspond to the subject's multiple pronunciation segments.

[0213] Optionally, the determining module 13 is specifically configured to: perform text recognition processing on the first pronunciation segment of the subject and the second pronunciation segment of the target reference, respectively, to obtain the first text content corresponding to the first pronunciation segment and the second text content corresponding to the second pronunciation segment, wherein the sequence number of the first pronunciation segment in the subject's multiple pronunciation segments is the same as the sequence number of the second pronunciation segment in the target reference's multiple pronunciation segments; compare the first text content and the second text content to determine the pronunciation correctness score of the subject's first pronunciation segment; and determine the pronunciation quality of the subject based on the subject's multiple pronunciation correctness scores, wherein the multiple pronunciation correctness scores correspond to the subject's multiple pronunciation segments.

[0214] Optionally, the determining module 13 is specifically used to: input the plurality of pronunciation correctness scores into the correctness score weight prediction model to obtain the weights of the plurality of pronunciation correctness scores; determine the total pronunciation correctness score of the subject based on the weights of the plurality of pronunciation correctness scores; and determine the pronunciation quality of the subject based on the total pronunciation correctness score of the subject, wherein the correctness score weight prediction model is a neural network model.

[0215] Optionally, the device further includes: a training module, configured to acquire a first face sample image; input the first face sample image into a first face detection model to obtain multiple face reconstruction parameters corresponding to the first face sample image, wherein the multiple face reconstruction parameters include lip movement unit intensity; input the multiple face reconstruction parameters corresponding to the first face sample image into a face reconstruction model to obtain a first face 3D model, and generate a first face reconstruction image based on the multiple face reconstruction parameters and the first face 3D model; determine a loss function value based on the first face sample image and the first face reconstruction image; and train the first face detection model based on the loss function value.

[0216] Optionally, the training module is specifically used to: extract a first face feature corresponding to the first face sample image and a second face feature corresponding to the first face reconstruction image; determine a first loss function value based on the first face feature and the second face feature; extract a first face key point from the first face sample image and obtain a second face key point from the first face 3D model; determine a second loss function value based on the first face key point and the second face key point; and compare the pixel values ​​of the first face sample image and the first face reconstruction image to determine a third loss function value.

[0217] Optionally, the training module is further configured to: acquire a second face sample image; input the second face sample image into a second face detection model to obtain multiple face reconstruction parameters corresponding to the second face sample image, wherein the multiple face reconstruction parameters include an identity parameter, wherein the second face detection model is a model obtained by training the first face detection model to meet a set cutoff condition; input the identity parameter into the face reconstruction model to obtain a second face 3D model; determine the lip shape deformation base model of the trainee corresponding to the second face sample image through a preset deformation transfer algorithm based on the second face 3D model, a preset neutral face 3D model, and a general lip shape deformation base model; and optimize the training of the second face detection model based on the lip shape deformation base model of the trainee.

[0218] Optionally, during the optimization training of the second face detection model, the training module is specifically used to: input multiple face reconstruction parameters corresponding to the second face sample image and the lip deformation base model of the trainer into the face reconstruction model to obtain a third face 3D model, and generate a second face reconstruction image based on the third face 3D model and the multiple face reconstruction parameters; determine the loss function value according to the second face sample image and the second face reconstruction image; and train the second face detection model according to the loss function value.

[0219] Figure 12 The device shown can perform the steps in the foregoing embodiments. For detailed execution process and technical effects, please refer to the description in the foregoing embodiments, which will not be repeated here.

[0220] In one possible design, the above Figure 12 The structure of the sound quality testing device shown can be implemented as an electronic device. For example... Figure 13 As shown, the electronic device may include: a processor 21, a memory 22, and a communication interface 23. The memory 22 stores executable code, which, when executed by the processor 21, enables the processor 21 to at least implement the pronunciation quality testing method provided in the foregoing embodiments.

[0221] In one optional embodiment, the electronic device used to perform the video recognition method provided in this embodiment of the invention can be any type of user terminal, such as a mobile phone, laptop computer, or PC, or it can be an extended reality (XR) device. XR is a general term for various forms such as virtual reality and augmented reality.

[0222] In practical applications, a pronunciation quality test program can be run on the XR device. Once the program is started, multiple test contents and corresponding reference videos (i.e., videos of a reference reading the corresponding test content) can be displayed on the relevant program interface. Based on the testing requirements, a target content and a video of the target reference reading that target content (the second video) can be selected. This video is then displayed on the XR device screen so that the subject can watch it and repeat after the subject.

[0223] In an optional embodiment, during the process of judging the pronunciation quality of the first video generated by the subject's repetition, some intermediate results can be produced, such as the pause positions corresponding to multiple pronunciation segments of the subject and the target reference as described in the above embodiment, the coherence score, the accuracy score, and the subject's health score for each pronunciation segment. These intermediate results can also be displayed on the screen of the XR device to more intuitively understand the specific situation of the subject's pronunciation.

[0224] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the pronunciation quality testing method provided in the foregoing embodiments.

[0225] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0226] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of a necessary general-purpose hardware platform, or by a combination of hardware and software. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method of testing a sound production quality, characterized by, The method comprises: obtaining a first video of a subject reading target content and a second video of a target reference reading the target content; performing mouth shape action unit intensity detection on each of a plurality of first images in the first video to obtain a respective mouth shape action unit intensity corresponding to each of the plurality of first images, and performing mouth shape action unit intensity detection on each of a plurality of second images in the second video to obtain a respective mouth shape action unit intensity corresponding to each of the plurality of second images; comparing the respective mouth shape action unit intensities of the plurality of first images with the respective mouth shape action unit intensities of the plurality of second images to determine the pronunciation quality of the subject. The mouth shape action unit intensity corresponding to any image refers to intensity coefficients corresponding to a plurality of preset mouth shape action units when forming a mouth shape in the any image. The plurality of preset mouth shape action units are a plurality of independent mouth shape action units divided in advance according to mouth movement characteristics. Different mouth shape action units control different mouth regions, and the mouth shape is formed by linear superposition of the different mouth shape action units according to respective intensity coefficients.

2. The method of claim 1, wherein, The method further comprises: obtaining a plurality of third videos of a plurality of reference persons reading the target content; performing mouth shape action unit intensity detection on each of a plurality of third images in a target third video to obtain a respective mouth shape action unit intensity corresponding to each of the plurality of third images, the target third video being any one of the plurality of third videos; determining a mouth shape feature vector corresponding to the target third video according to the respective mouth shape action unit intensities of the plurality of third images; performing clustering on the respective mouth shape feature vectors corresponding to the plurality of third videos to obtain a plurality of clustering results; determining a mouth shape feature vector corresponding to the first video according to the respective mouth shape action unit intensities of the plurality of first images; determining a target clustering result corresponding to the mouth shape feature vector corresponding to the first video from the plurality of clustering results; determining the target reference person from reference persons corresponding to the target clustering result.

3. The method of claim 1, wherein, The comparison of the respective mouth shape action unit intensities of the plurality of first images with the respective mouth shape action unit intensities of the plurality of second images to determine the pronunciation quality of the subject comprises: determining a plurality of pause points of the subject according to the degree of change of the respective mouth shape action unit intensities of the plurality of first images; determining a plurality of pause points of the target reference person according to the degree of change of the respective mouth shape action unit intensities of the plurality of second images; determining a plurality of pronunciation segments of the subject according to the plurality of pause points of the subject, and determining a plurality of pronunciation segments of the target reference person according to the plurality of pause points of the target reference person; wherein one pronunciation segment includes a plurality of images between adjacent pause points; comparing the plurality of pronunciation segments of the subject with the plurality of pronunciation segments of the target reference person to determine a total score of pronunciation coherence and / or a total score of pronunciation correctness of the subject, and determining the pronunciation quality of the subject according to the total score of pronunciation coherence and / or the total score of pronunciation correctness.

4. The method of claim 3, wherein, The comparing the plurality of pronunciation segments of the subject and the plurality of pronunciation segments of the target reference person to determine the pronunciation quality of the subject comprises: determining a time length corresponding to each of the plurality of pronunciation segments of the subject; determining a time length corresponding to each of the plurality of pronunciation segments of the target reference person; determining a plurality of pronunciation fluency scores of the subject according to the time length corresponding to each of the plurality of pronunciation segments of the subject and the time length corresponding to each of the plurality of pronunciation segments of the target reference person, wherein the plurality of pronunciation fluency scores correspond to the plurality of pronunciation segments of the subject; determining the pronunciation quality of the subject according to the plurality of pronunciation fluency scores.

5. The method of claim 4, wherein, The determining the pronunciation quality of the subject according to the plurality of pronunciation fluency scores comprises: inputting the plurality of pronunciation fluency scores into a fluency score weight prediction model to obtain weights of the plurality of pronunciation fluency scores, the fluency score weight prediction model being a neural network model; determining a total pronunciation fluency score of the subject according to the weights of the plurality of pronunciation fluency scores; determining the pronunciation quality of the subject according to the total pronunciation fluency score of the subject.

6. The method of claim 3, wherein, The comparing the plurality of pronunciation segments of the subject and the plurality of pronunciation segments of the target reference person to determine the pronunciation quality of the subject comprises: obtaining a mouth shape action unit intensity corresponding to a first pronunciation segment of the subject and a mouth shape action unit intensity corresponding to a second pronunciation segment of the target reference person, wherein a sequence number corresponding to the first pronunciation segment in the plurality of pronunciation segments of the subject is the same as a sequence number corresponding to the second pronunciation segment in the plurality of pronunciation segments of the target reference person; comparing the mouth shape action unit intensity corresponding to the first pronunciation segment of the subject and the mouth shape action unit intensity corresponding to the second pronunciation segment of the target reference person to determine a pronunciation correctness score of the first pronunciation segment of the subject; determining the pronunciation quality of the subject according to a plurality of pronunciation correctness scores of the subject, wherein the plurality of pronunciation correctness scores correspond to the plurality of pronunciation segments of the subject.

7. The method of claim 3, wherein, The comparing the plurality of pronunciation segments of the subject and the plurality of pronunciation segments of the target reference person to determine the pronunciation quality of the subject comprises: respectively performing character recognition processing on a first pronunciation segment of the subject and a second pronunciation segment of the target reference person to obtain first character content corresponding to the first pronunciation segment and second character content corresponding to the second pronunciation segment, wherein a sequence number corresponding to the first pronunciation segment in the plurality of pronunciation segments of the subject is the same as a sequence number corresponding to the second pronunciation segment in the plurality of pronunciation segments of the target reference person; comparing the first character content and the second character content to determine a pronunciation correctness score of the first pronunciation segment of the subject; determining the pronunciation quality of the subject according to a plurality of pronunciation correctness scores of the subject, wherein the plurality of pronunciation correctness scores correspond to the plurality of pronunciation segments of the subject.

8. The method according to claim 6 or 7, characterized in that, The method further comprises: determining the pronunciation quality of the subject according to the plurality of pronunciation correctness scores of the subject, comprising: inputting the plurality of pronunciation correctness scores into a correctness score weight prediction model to obtain weights of the plurality of pronunciation correctness scores, wherein the correctness score weight prediction model is a neural network model; determining a total pronunciation correctness score of the subject according to the weights of the plurality of pronunciation correctness scores; 9. The method of claim 1, wherein, determining the pronunciation quality of the subject according to the total pronunciation correctness score of the subject. The training process of the face detection model for conducting the mouth shape action unit strength detection comprises: obtaining a first face sample image; inputting the first face sample image into a first face detection model to obtain a plurality of face reconstruction parameters corresponding to the first face sample image, wherein the plurality of face reconstruction parameters comprise a mouth shape action unit strength; inputting the plurality of face reconstruction parameters corresponding to the first face sample image into a face reconstruction model to obtain a first face three-dimensional model, and generating a first face reconstruction image based on the plurality of face reconstruction parameters and the first face three-dimensional model; determining a loss function value according to the first face sample image and the first face reconstruction image; 10. The method of claim 9, wherein, training the first face detection model according to the loss function value. The method further comprises: obtaining a second face sample image; inputting the second face sample image into a second face detection model to obtain a plurality of face reconstruction parameters corresponding to the second face sample image, wherein the plurality of face reconstruction parameters comprise an identity parameter, and the second face detection model is obtained by training the first face detection model to meet a set stop condition; inputting the identity parameter into the face reconstruction model to obtain a second face three-dimensional model; determining a trainer's mouth shape deformation base model corresponding to the second face sample image through a preset deformation transfer algorithm according to the second face three-dimensional model, a preset neutral face three-dimensional model, and a general mouth shape deformation base model; 11. The method of claim 9, wherein, optimizing and training the second face detection model according to the trainer's mouth shape deformation base model. The method further comprises: obtaining a second face sample image; inputting the second face sample image into a second face detection model to obtain a plurality of face reconstruction parameters corresponding to the second face sample image, wherein the plurality of face reconstruction parameters comprise an identity parameter, and the second face detection model is obtained by training the first face detection model to meet a set stop condition; inputting the identity parameter into the face reconstruction model to obtain a second face three-dimensional model; determining a trainer's mouth shape deformation base model corresponding to the second face sample image through a preset deformation transfer algorithm according to the second face three-dimensional model, a preset neutral face three-dimensional model, and a general mouth shape deformation base model; 12. The method of claim 11, wherein, optimizing and training the second face detection model according to the trainer's mouth shape deformation base model. The method further comprises: obtaining a second face sample image; inputting the second face sample image into a second face detection model to obtain a plurality of face reconstruction parameters corresponding to the second face sample image, wherein the plurality of face reconstruction parameters comprise an identity parameter, and the second face detection model is obtained by training the first face detection model to meet a set stop condition; inputting the identity parameter into the face reconstruction model to obtain a second face three-dimensional model; determining a trainer's mouth shape deformation base model corresponding to the second face sample image through a preset deformation transfer algorithm according to the second face three-dimensional model, a preset neutral face three-dimensional model, and a general mouth shape deformation base model; optimizing and training the second face detection model according to the trainer's mouth shape deformation base model. inputting the plurality of face reconstruction parameters corresponding to the second face sample image and the lip deformation base model of the trainer into the face reconstruction model to obtain a third face three-dimensional model, and generating a second face reconstruction image based on the third face three-dimensional model and the plurality of face reconstruction parameters; determining a loss function value according to the second face sample image and the second face reconstruction image; training the second face detection model according to the loss function value.

13. A method of testing a sound production quality, characterized by, comprising: receiving a request triggered by a user equipment by calling a pronunciation quality test service, the request including a first video of a subject reading target content and a second video of a target reference reading the target content; using processing resources corresponding to the pronunciation quality test service to perform the following steps: performing lip action unit intensity detection on a plurality of first images in the first video respectively to obtain respective lip action unit intensities corresponding to the plurality of first images; and performing lip action unit intensity detection on a plurality of second images in the second video respectively to obtain respective lip action unit intensities corresponding to the plurality of second images; comparing the respective lip action unit intensities corresponding to the plurality of first images with the respective lip action unit intensities corresponding to the plurality of second images to determine the pronunciation quality of the subject; wherein the lip action unit intensity corresponding to any image refers to intensity coefficients corresponding to a plurality of preset lip action units when forming a lip shape in the any image, the plurality of preset lip action units being a plurality of independent lip action units pre-divided according to mouth movement characteristics, different lip action units controlling different mouth regions, and the lip shape being formed by linear superposition of the different lip action units according to respective intensity coefficients.

14. A method of testing a sound production quality, characterized by, The method is applied to an extended reality device, and the method comprises: obtaining a first video of a subject reading target content and a second video of a target reference reading the target content; performing lip action unit intensity detection on a plurality of first images in the first video respectively to obtain respective lip action unit intensities corresponding to the plurality of first images; and performing lip action unit intensity detection on a plurality of second images in the second video respectively to obtain respective lip action unit intensities corresponding to the plurality of second images; wherein the lip action unit intensity corresponding to any image refers to intensity coefficients corresponding to a plurality of preset lip action units when forming a lip shape in the any image, the plurality of preset lip action units being a plurality of independent lip action units pre-divided according to mouth movement characteristics, different lip action units controlling different mouth regions, and the lip shape being formed by linear superposition of the different lip action units according to respective intensity coefficients; comparing the respective lip action unit intensities corresponding to the plurality of first images with the respective lip action unit intensities corresponding to the plurality of second images to determine the pronunciation quality of the subject; rendering and displaying the pronunciation quality on a screen of the extended reality device.

Citation Information

Patent Citations

  • Pronunciation evaluation method, device, system, medium and computing equipment

    CN111951828A

  • Method for estimating pronunciation using image

    JP2008146268A