Pronunciation correction data analysis method based on three-dimensional dynamic oral modeling

Through function fitting and reference point setting methods, the problem of inaccurate tongue position judgment in three-dimensional dynamic oral modeling is solved, and a higher accuracy of tongue position judgment is achieved.

CN120296641BActive Publication Date: 2025-08-29BEIJING CETEN EDUCATION TECH GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510779091.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-08-29
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the existing pronunciation correction technology, three-dimensional dynamic oral modeling fails to accurately determine whether the tongue position is correct, resulting in a low accuracy in judging the tongue position during pronunciation.

Method used

The normal pronunciation function is obtained by obtaining historical reference points and function fitting. The normal pronunciation range is obtained based on the normal pronunciation function and historical reference points. The reference point is set based on the tongue characteristics to be analyzed to determine whether the tongue position is correct when the person is pronunciating.

Benefits of technology

It improves the accuracy of judging whether the tongue position is correct when pronunciating, and establishes a correct position judgment method based on the characteristics of the tongue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296641B_ABST
    Figure CN120296641B_ABST
Patent Text Reader

Abstract

The present invention discloses a pronunciation correction data analysis method based on three-dimensional dynamic oral modeling, which relates to the technical field of pronunciation correction and comprises the following steps: performing function fitting on historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue portion in the three-dimensional oral modeling of a person to be analyzed when pronouncing and when not pronouncing, and marking them as a detection pronunciation graph and a detection unpronounced graph, respectively; scaling the detection pronunciation graph based on a reference size and the detection unpronounced graph to obtain a detection correction graph; and judging whether pronunciation abnormality occurs based on the detection correction graph and the normal pronunciation range. The present invention is used to solve the problem in existing pronunciation correction technology that the three-dimensional dynamic oral modeling fails to establish a method for judging the correct position of the tongue, resulting in low accuracy in judging whether the tongue position is correct during pronunciation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pronunciation correction, and in particular to a pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling. Background Art

[0002] Traditional pronunciation correction methods rely primarily on auditory feedback. Using auditory feedback alone makes it difficult for speakers to intuitively perceive subtle changes in the oral cavity during pronunciation, particularly whether the tongue position is abnormal. The tongue plays a crucial role in pronunciation, and the accuracy of its position directly affects the quality of pronunciation. Therefore, a pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling was proposed.

[0003] However, since each person's tongue size is different and the tongue has an irregular shape, and the tongue may curl when pronouncing words, it is more difficult to judge whether the tongue position is correct, resulting in an inability to accurately judge whether the tongue position is correct. For example, the patent application with publication number CN112381913A discloses a method for constructing a dynamic pronunciation teaching model based on 3D modeling and oral anatomy. This solution fails to set the tongue position during correct pronunciation to compare and judge whether the pronunciation is abnormal, and only judges whether the tongue position is correct based on experience, resulting in a decrease in the accuracy of the judgment. The three-dimensional dynamic oral modeling in the existing pronunciation correction technology fails to establish a method for judging the correct position of the tongue, resulting in a low accuracy in judging whether the tongue position is correct during pronunciation. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems in the prior art to a certain extent, by performing function fitting on historical reference points to obtain a normal pronunciation function, obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points, obtaining a side view of the tongue part in the three-dimensional oral modeling of the person to be analyzed when pronouncing and when not pronouncing, marking them as a detection pronunciation diagram and a detection unpronounced diagram respectively, scaling the detection pronunciation diagram based on the reference size and the detection unpronounced diagram to obtain a detection correction diagram, and judging whether pronunciation abnormality occurs based on the detection correction diagram and the normal pronunciation range, so as to solve the problem that the three-dimensional dynamic oral modeling in pronunciation correction technology fails to establish a method for judging the correct position of the tongue, resulting in low accuracy in judging whether the tongue position is correct during pronunciation.

[0005] To achieve the above objectives, the present application provides a pronunciation correction data analysis method based on three-dimensional dynamic oral modeling, comprising the following steps:

[0006] Obtaining side views of tongue portions in the three-dimensional oral cavity modeling of a first number of individuals when they pronounce and when they do not pronounce, marking the images as historical pronunciation images and historical unpronounced images, respectively;

[0007] Get reference dimensions based on historical unvoiced figures;

[0008] scaling the historical pronunciation image based on the reference size and the historical unpronounced image to obtain a historical correction image;

[0009] Obtain reference points of all historical correction maps based on a reference point setting method and mark them as historical reference points;

[0010] Perform function fitting on historical reference points to obtain a normal pronunciation function;

[0011] Obtaining a normal pronunciation range based on a normal pronunciation function and historical reference points;

[0012] Obtain side views of the tongue portion of the three-dimensional oral model of the person to be analyzed when the person is speaking and when the person is not speaking, and mark them as a detection pronunciation image and a detection unpronounced image respectively;

[0013] scaling the detected pronunciation image based on the reference size and the detected unpronounced image to obtain a detected correction image;

[0014] Determine whether there is pronunciation abnormality based on the detection correction graph and the normal pronunciation range.

[0015] Furthermore, obtaining the reference size based on the historical unvoiced image includes the following sub-steps:

[0016] Establish a plane rectangular coordinate system, mark it as the length coordinate system, and place the historical unpronounced graph in the length coordinate system;

[0017] Obtain the minimum and maximum values ​​of the horizontal coordinate of the tongue part in each historical unpronounced graph, and mark them as historical minimum value and historical maximum value respectively;

[0018] Calculate the difference between the historical maximum and historical minimum values ​​of each historical unvoiced graph, and mark it as the historical difference;

[0019] Divide the range of historical differences into q equal ranges, marked as difference division ranges;

[0020] Count the frequency of each difference division range and mark it as division frequency;

[0021] A histogram is drawn with the historical difference as the X-axis, the divided frequency as the Y-axis, and the difference division range as the histogram interval, and is marked as a difference histogram.

[0022] Furthermore, obtaining the reference size based on the historical unvoiced image further includes the following sub-steps:

[0023] Get the maximum division frequency in the difference histogram and mark it as the starting frequency;

[0024] The median frequency threshold is calculated as: F1=0.5*D1 / q; where F1 is the median frequency threshold and D1 is the first quantity;

[0025] In the difference histogram, starting from the starting frequency part, determine whether the divided frequency is greater than the median frequency threshold to the right. If so, continue to determine whether the divided frequency is greater than the median frequency threshold to the right, and stop when it is not. Get the maximum value of the horizontal axis of the divided frequency before stopping, and mark it as the maximum value of the screening;

[0026] In the difference histogram, starting from the starting frequency part, determine whether the divided frequency is greater than the median frequency threshold to the left. If so, continue to determine whether the divided frequency is greater than the median frequency threshold to the left until it is not. Get the minimum value of the horizontal axis of the divided frequency before stopping, and mark it as the minimum value of the screening;

[0027] Get the average of the historical differences between the maximum and minimum screening values ​​and mark it as a reference size.

[0028] Furthermore, scaling the historical pronunciation image based on the reference size and the historical unpronounced image to obtain the historical correction image includes the following sub-steps:

[0029] The scaling ratio of each person is calculated as: Sb1=Cc / Cl; where Sb1 is the scaling ratio of each person, Cl is the historical difference, and Cc is the reference size;

[0030] The historical pronunciation graph of each person is horizontally scaled according to its scaling ratio to obtain a historical correction graph.

[0031] Furthermore, obtaining reference points of all historical correction images based on the reference point setting method and marking them as historical reference points includes the following sub-steps:

[0032] Establish a plane rectangular coordinate system, mark it as the historical reference coordinate system, and place the historical correction diagram in the length coordinate system;

[0033] Obtain the tongue contour portion in the historical correction image and mark it as the historical contour;

[0034] The reference points obtained by using the reference point setting method with the historical contour as the target contour are marked as historical reference points.

[0035] Furthermore, the reference point setting method includes:

[0036] Get a coordinate point when the horizontal coordinate of the target contour is the smallest, mark it as the starting coordinate point, get a coordinate point when the horizontal coordinate of the target contour is the largest, mark it as the final coordinate point, take the starting coordinate point as the starting point and move along the upper half of the target contour, get the coordinate point of this position every time you move the first distance, mark it as the search coordinate point, stop until you reach the final coordinate point, start from the starting coordinate point and connect the search coordinate points in the order in which they were obtained, get the connecting line segments, get the slope of the connecting line segments, get the absolute value of the difference between the slope of each connecting line segment and the slope of the next connecting line segment, mark it as the slope change value; get the search coordinate point in the middle of the two connecting line segments when the slope change value is the maximum, mark it as the filtering starting coordinate point, get all the search coordinate points after the filtering starting coordinate point, mark them as reference points.

[0037] Furthermore, obtaining the normal pronunciation range based on the normal pronunciation function and the historical reference point includes the following sub-steps:

[0038] Substituting the horizontal coordinate of the historical reference point into the normal pronunciation function, the value obtained is marked as the function prediction value;

[0039] Calculate the absolute value of the difference between the function prediction value and the ordinate of the historical reference point, and mark it as the fitting difference;

[0040] Sort the fitting differences from small to large and label them as Cz1 to Cz i ;

[0041] Determine whether w1*D1 is an integer. If so, set Cz (w1*D1) As the first position value; if not, get the integers on the left and right sides of w1*D1, mark them as R1 and R2 respectively, and calculate Cz (R1) With Cz (R2) The mean of is taken as the first position value; w1 is the first position coefficient, the range of w1 is: (0, 0.5), and D1 is the number of fitting differences;

[0042] Determine whether w2*D1 is an integer. If so, set Cz (w2*D1) As the second position value; if not, get the integers on the left and right sides of w2*D1, mark them as R3 and R4 respectively, and calculate Cz (R3) With Cz (R4) The mean of is taken as the second position value; where w2 is 1-w1;

[0043] The moving distance threshold is obtained as follows: E2=R2+[(R2-R1) / (w2-w1)]*w1; where E2 is the moving distance threshold, R1 is the first position value, and R2 is the second position value;

[0044] The normal pronunciation function is moved in the positive direction and the negative direction of the vertical axis by a moving distance threshold respectively to obtain an upper threshold function and a lower threshold function, and the range between the upper threshold function and the lower threshold function is marked as the normal pronunciation range.

[0045] Furthermore, scaling the detected pronunciation image based on the reference size and the detected unpronounced image to obtain the detected correction image includes the following sub-steps:

[0046] Establish a plane rectangular coordinate system, mark it as the scaled coordinate system, and place the detected unvoiced graph in the scaled coordinate system;

[0047] Obtain the difference between the maximum and minimum values ​​of the horizontal coordinate of the tongue part in the unpronounced image, and mark it as the detection difference;

[0048] The scaling ratio of the person to be analyzed is calculated as: Sb2=Cc / C2; where Sb2 is the scaling ratio of the person to be analyzed, and C2 is the detection difference;

[0049] The detection pronunciation graph of the person to be analyzed is horizontally scaled according to its scaling ratio to obtain a detection correction graph.

[0050] Furthermore, judging whether pronunciation abnormality occurs based on the detection correction graph and the normal pronunciation range includes the following sub-steps:

[0051] Place the inspection correction diagram in the length coordinate system, ensuring that the placement of the inspection correction diagram is consistent with that of the historical correction diagram;

[0052] Obtain the tongue contour in the detection and correction image, and mark it as the detection contour;

[0053] The reference points obtained by the reference point setting method using the detection contour as the target contour are marked as detection reference points.

[0054] Furthermore, judging whether there is abnormal pronunciation based on the detection correction graph and the normal pronunciation range further includes the following sub-steps:

[0055] The detection reference points are fitted with a function to obtain a detection pronunciation function, and it is determined whether the detection pronunciation function is all within the normal pronunciation range. If not, a voice abnormality signal is issued.

[0056] The beneficial effects of the present invention are as follows: the present invention obtains a normal pronunciation function by performing function fitting on historical reference points, obtains a normal pronunciation range based on the normal pronunciation function and the historical reference points, obtains side views of the tongue portion in a three-dimensional oral modeling of the person to be analyzed when pronouncing and when not pronouncing, and marks them as a detection pronunciation graph and a detection unpronounced graph, respectively; scales the detection pronunciation graph based on a reference size and the detection unpronounced graph to obtain a detection correction graph; and determines whether pronunciation abnormality occurs based on the detection correction graph and the normal pronunciation range. The advantage of the present invention is that a method for determining the correct position based on tongue characteristics is established, thereby increasing the accuracy of determining whether the tongue position is correct during pronunciation;

[0057] The present invention obtains the reference points of all historical correction images based on the reference point setting method. The advantage is that the reference point setting method can be set based on the characteristic of the tongue bending upward when pronouncing words and the characteristics of the tongue shape. The tongue position is judged based on the reference point, which increases the accuracy of judging whether the tongue is in the correct position when pronouncing words. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flow chart of the steps of the method of the present invention;

[0059] Figure 2 is a schematic diagram of a difference histogram of the present invention;

[0060] Figure 3 It is a schematic diagram of the screening maximum value and the screening minimum value of the present invention;

[0061] Figure 4 A schematic diagram of the search coordinate points of the present invention;

[0062] Figure 5 A schematic diagram of a reference point of the present invention;

[0063] Figure 6 is a schematic diagram of a normal pronunciation function of the present invention;

[0064] Figure 7 Schematic diagram of the upper threshold function and the lower threshold function of the present invention;

[0065] Figure 8 Schematic diagram of the pronunciation detection function of the present invention. DETAILED DESCRIPTION

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0067] Example 1, please refer to Figure 1 As shown, the present application provides a pronunciation correction data analysis method based on three-dimensional dynamic oral modeling, comprising the following steps:

[0068] Step S1, obtain side views of the tongue part in the three-dimensional oral modeling of a first number of individuals when pronouncing and not pronouncing, and mark them as historical pronunciation images and historical unpronounced images, respectively. The historical pronunciation images and historical unpronounced images are directly obtained from an existing three-dimensional oral modeling database that has been established.

[0069] Step S2, obtaining a reference size based on the historical unvoiced image; Step S2 further includes the following sub-steps: establishing a plane rectangular coordinate system, marked as a length coordinate system, and placing the historical unvoiced image in the length coordinate system;

[0070] Step S201, obtaining the minimum value and maximum value of the horizontal coordinate of the tongue part in each historical unpronounced graph, marking them as historical minimum value and historical maximum value respectively;

[0071] Step S202, calculating the difference between the historical maximum value and the historical minimum value of each historical unvoiced graph, and marking it as the historical difference;

[0072] Step S203, dividing the range of historical differences into q equal ranges, marked as difference division ranges;

[0073] Step S204, counting the frequency of each difference division range and marking it as division frequency;

[0074] Step S205: draw a histogram with the historical difference as the X-axis, the divided frequency as the Y-axis, and the difference division range as the histogram interval, and mark it as a difference histogram;

[0075] Step S206, obtaining the maximum division frequency in the difference histogram, marking it as the starting frequency;

[0076] Step S207, calculate the median frequency threshold as: F1 = 0.5 * D1 / q; where F1 is the median frequency threshold, D1 is the first quantity; D1 / q is the average frequency of each division, and when it is less than half, it is considered a smaller frequency and is therefore multiplied by 0.5;

[0077] Step S208: In the difference histogram, starting from the initial frequency portion, determine whether the divided frequency is greater than the median frequency threshold. If so, continue to determine whether the divided frequency is greater than the median frequency threshold until it is not. The maximum value of the horizontal coordinate of the divided frequency before stopping is obtained and marked as the maximum value of the screening;

[0078] Step S209: In the difference histogram, starting from the initial frequency portion, determine whether the divided frequency is greater than the median frequency threshold. If so, continue to determine whether the divided frequency is greater than the median frequency threshold until it is not. The minimum value of the horizontal coordinate of the divided frequency before stopping is obtained and marked as the minimum value of the screening.

[0079] Step S210: Obtain the average of the historical differences between the maximum and minimum values ​​of the screening, and mark it as a reference size; the reference size obtained after screening is more universal;

[0080] In practical applications, please refer to Figure 2 and Figure 3 As shown, the range of historical differences is evenly divided into 7 difference division ranges. If the first number is 220, the median frequency threshold is calculated as: F1=0.5*220 / 7=15.7. The calculation result is rounded to one decimal place. Starting from the starting frequency 65, the division frequency 41 is judged to the right to see if it is greater than 15.7. If so, the judgment is continued to the right until the division frequency 2 is less than 15.7 and stops. The maximum value of the horizontal coordinate of the division frequency before stopping is 6.2 cm, so the screening maximum value is 6.2 cm. Similarly, the screening minimum value is 5.2 cm. The average of the historical differences between 6.2 cm and 5.2 cm is 5.66 cm. The calculation result is rounded to two decimal places, so the reference size is 5.66 cm.

[0081] Step S3, scaling the historical pronunciation image based on the reference size and the historical unpronounced image to obtain a historical correction image; Step S3 also includes the following sub-steps:

[0082] Step S301, calculate the scaling ratio of each person as: Sb1=Cc / Cl; where Sb1 is the scaling ratio of each person, Cl is the historical difference, and Cc is the reference size; the scaling ratio obtained before pronunciation is more accurate;

[0083] Step S302: horizontally scaling each person's historical pronunciation graph according to the scaling ratio to obtain a historical correction graph;

[0084] In practical applications, for example, when the historical difference of a person is 5.66 cm and the reference size is 5.66 cm, the scaling ratio of the person is: Sb1=1, and the historical pronunciation graph is horizontally scaled by 1 times to obtain the historical correction graph.

[0085] Step S4, obtaining reference points of all historical correction images based on the reference point setting method and marking them as historical reference points; Step S4 also includes the following sub-steps:

[0086] Step S401: Establish a plane rectangular coordinate system, mark it as the historical reference coordinate system, and place the historical correction graph in the length coordinate system;

[0087] Step S402, obtaining the contour of the tongue in the historical correction image and marking it as the historical contour;

[0088] Step S403: Using the historical contour as the target contour, a reference point obtained by using the reference point setting method is marked as a historical reference point. Step S403 further includes the following sub-steps:

[0089] Step S40301: Obtain a coordinate point where the horizontal coordinate of the target contour is the smallest, mark it as the starting coordinate point, obtain a coordinate point where the horizontal coordinate of the target contour is the largest, mark it as the final coordinate point, move along the upper half of the target contour with the starting coordinate point as the starting point, obtain the coordinate point of this position each time it moves a first distance, mark it as the search coordinate point, stop until it reaches the final coordinate point, start from the starting coordinate point and connect the search coordinate points in the order in which the search coordinate points were obtained, obtain connecting line segments, obtain the slopes of the connecting line segments, obtain the absolute value of the difference between the slope of each connecting line segment and the slope of the next connecting line segment, mark it as the slope change value; obtain the search coordinate point between the two connecting line segments when the slope change value is the maximum, mark it as the screening starting coordinate point, obtain all search coordinate points after the screening starting coordinate point, and mark them as reference points;

[0090] In practical applications, please refer to Figure 4 and Figure 5 As shown, the beneficial effect of this method is that, when the tongue bends, the starting coordinate point is not the tip of the tongue, and the slope change value of the tip of the tongue relative to the rest of the tongue is the maximum point. Therefore, it can be determined that the tip of the tongue is between the two connecting line segments at the maximum slope change value.

[0091] Step S5: Perform function fitting on the historical reference points to obtain a normal pronunciation function.

[0092] In practical applications, please refer to Figure 6 As shown, the historical reference points are for multiple people, and a more universal normal pronunciation function can be obtained.

[0093] Step S6, obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference point; Step S6 also includes the following sub-steps:

[0094] Step S601, substituting the horizontal coordinate of the historical reference point into the normal pronunciation function to obtain a value marked as a function prediction value;

[0095] Step S602, calculating the absolute value of the difference between the function prediction value and the ordinate of the historical reference point, and marking it as the fitting difference;

[0096] Step S603: sort the fitting differences from small to large and label them as Cz1 to Cz i ;

[0097] Step S604: determine whether w1*D1 is an integer. If so, set Cz (w1*D1) As the first position value; if not, get the integers on the left and right sides of w1*D1, mark them as R1 and R2 respectively, and calculate Cz (R1) With Cz (R2) The mean of is taken as the first position value; w1 is the first position coefficient, the range of w1 is: (0, 0.5), D1 is the number of fitting differences; in order to select the value of the first position value in the smaller half range, w1 is set between 0 and 0.5. Selecting an intermediate value can ensure that the selected first position value has the representativeness of the smaller half range. w1 is preferably selected as 0.25;

[0098] Step S605, determine whether w2*D1 is an integer, if so, set Cz (w2*D1) As the second position value; if not, get the integers on the left and right sides of w2*D1, mark them as R3 and R4 respectively, and calculate Cz (R3) With Cz (R4) The mean of the two positions is used as the second position value; where w2 is 1-w1; in order to select the value of the second position value in the larger half range, since the range of w1 is (0, 0.5), w2 can be set to 1-w1. When w1 is 0.25, w2 is 0.75;

[0099] Step S606: Calculate the moving distance threshold as follows: E2 = R2 + [(R2-R1) / (w2-w1)]*w1; where E2 is the moving distance threshold, R1 is the first position value, and R2 is the second position value. When the fitting differences are evenly distributed, the overall distribution difference is calculated based on the first position value and the second position value, and then the fitting difference between the proportions from 1-w1 to w2 is calculated as [(R2-R1) / (w2-w1)]*w1. R2 + [(R2-R1) / (w2-w1)]*w1 is the maximum value when the fitting differences are evenly distributed. For the same pronunciation, the fitting differences are relatively concentrated. Therefore, when the pronunciation is correct, the fitting differences should not move the distance threshold. Therefore, the moving distance threshold is set.

[0100] Step S607, moving the normal pronunciation function in the positive direction of the vertical axis and the negative direction of the vertical axis by a moving distance threshold, respectively, to obtain an upper threshold function and a lower threshold function, and marking the area between the upper threshold function and the lower threshold function as a normal pronunciation range;

[0101] In practical applications, please refer to Figure 7 As shown, when w1 is 0.25, w2 is 0.75, the number of fitting differences is 1540, and it is determined that 0.25*1540=385 is an integer. (385) The corresponding fitting difference is 0.13cm, which is the first position value; judge that 0.75*1540=1155 is an integer, and Cz(1155) The corresponding fitting difference is 0.38cm, which is the second position value. The moving distance threshold is: E2=0.38+[(0.38-0.13) / (0.75-0.25)]*0.25=0.51cm. The calculation result is rounded to two decimal places, so the moving distance threshold is 0.51cm. Move the normal pronunciation function in the positive direction and the negative direction of the vertical axis by 0.51cm respectively to obtain the upper threshold function and the lower threshold function. The range between the upper threshold function and the lower threshold function is marked as the normal pronunciation range.

[0102] Step S7: obtaining side views of the tongue portion in the three-dimensional oral modeling of the person to be analyzed when the person is pronouncing a word and when the person is not pronouncing a word, and marking them as a detection pronunciation image and a detection unpronounced image, respectively.

[0103] Step S8, scaling the detected pronunciation image based on the reference size and the detected unpronounced image to obtain a detected correction image; step S8 also includes the following sub-steps:

[0104] Step S801, establishing a plane rectangular coordinate system, marked as a scaled coordinate system, and placing the detected unpronounced image in the scaled coordinate system;

[0105] Step S802, obtaining the difference between the maximum and minimum values ​​of the horizontal coordinate of the tongue portion in the unpronounced image, and marking it as the detection difference;

[0106] Step S803, calculating the scaling ratio of the person to be analyzed as: Sb2=Cc / C2; where Sb2 is the scaling ratio of the person to be analyzed, and C2 is the detection difference;

[0107] Step S804, horizontally scaling the detected pronunciation graph of the person to be analyzed according to its scaling ratio to obtain a detection correction graph;

[0108] In actual applications, if the detection difference is 5.45cm, the scaling ratio of the person to be analyzed is calculated as: Sb2=5.66 / 5.55. The detection pronunciation diagram of the person to be analyzed is horizontally scaled by 5.66 / 5.55 times to obtain the detection correction diagram. Since different people have different tongue sizes, scaling them to a consistent size facilitates comparison.

[0109] Step S9, judging whether there is abnormal pronunciation based on the detection correction graph and the normal pronunciation range; Step S9 also includes the following sub-steps:

[0110] Step S901: Place the detected correction image in the length coordinate system, ensuring that the placement of the detected correction image is consistent with that of the historical correction image; the consistent placement facilitates comparison;

[0111] Step S902, obtaining the tongue contour portion in the detection and correction image, and marking it as the detection contour;

[0112] Step S903, using the detection contour as the target contour, a reference point is obtained by using a reference point setting method, and marked as a detection reference point;

[0113] Step S904, performing function fitting on the detection reference points to obtain a detection pronunciation function, and determining whether the detection pronunciation function is all within the normal pronunciation range. If not, a sound abnormality signal is issued;

[0114] In practical applications, please refer to Figure 8 As shown, the acquired detection pronunciation function is in the normal pronunciation range between the upper threshold function and the lower threshold function, so there is no need to issue a voice abnormality signal.

[0115] Example 2: The present application also provides an electronic device, which may include: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the pronunciation correction data analysis method based on three-dimensional dynamic oral modeling are performed to achieve the following functions: obtaining side views of the tongue part in the three-dimensional oral modeling of a first number of individuals when they pronounce and do not pronounce, and marking them as historical pronunciation images and historical unpronounced images respectively; obtaining reference dimensions based on the historical unpronounced images; scaling the historical pronunciation images based on the reference dimensions and the historical unpronounced images to obtain historical correction images; obtaining reference points of all historical correction images based on a reference point setting method, and marking them as historical reference points; performing function fitting on the historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue part in the three-dimensional oral modeling of the person to be analyzed when they pronounce and do not pronounce, and marking them as detection pronunciation images and detection unpronounced images respectively; scaling the detection pronunciation images based on the reference dimensions and the detection unpronounced images to obtain detection correction images; and judging whether pronunciation abnormality occurs based on the detection correction images and the normal pronunciation range.

[0116] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0117] Example 3. The present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the pronunciation correction data analysis method based on three-dimensional dynamic oral modeling provided by the above methods, the method including: obtaining side views of the tongue part in the three-dimensional oral modeling of a first number of individuals when they pronounce and do not pronounce, and marking them as historical pronunciation images and historical unpronounced images respectively; obtaining reference dimensions based on the historical unpronounced images; scaling the historical pronunciation images based on the reference dimensions and the historical unpronounced images to obtain historical correction images; obtaining reference points of all historical correction images based on a reference point setting method, and marking them as historical reference points; performing function fitting on the historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue part in the three-dimensional oral modeling of the person to be analyzed when they pronounce and do not pronounce, and marking them as detection pronunciation images and detection unpronounced images respectively; scaling the detection pronunciation images based on the reference dimensions and the detection unpronounced images to obtain detection correction images; judging whether pronunciation abnormality occurs based on the detection correction image and the normal pronunciation range.

[0118] Example 4. The present application also provides a computer-readable storage medium. The present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the pronunciation correction data analysis method based on three-dimensional dynamic oral modeling are executed to achieve the following functions: obtaining side views of the tongue part in the three-dimensional oral modeling of a first number of individuals when they pronounce and do not pronounce, and marking them as historical pronunciation images and historical unpronounced images respectively; obtaining reference dimensions based on the historical unpronounced images; scaling the historical pronunciation images based on the reference dimensions and the historical unpronounced images to obtain historical correction images; obtaining reference points of all historical correction images based on a reference point setting method, and marking them as historical reference points; performing function fitting on the historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue part in the three-dimensional oral modeling of the person to be analyzed when they pronounce and do not pronounce, and marking them as detection pronunciation images and detection unpronounced images respectively; scaling the detection pronunciation images based on the reference dimensions and the detection unpronounced images to obtain detection correction images; judging whether pronunciation abnormality occurs based on the detection correction image and the normal pronunciation range.

[0119] Through the description of the above embodiments, the embodiments of the present invention can be provided as methods, systems, or computer program products. Based on this understanding, the essence of the above technical solutions or the portion that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the methods described in various embodiments or certain portions of the embodiments.

[0120] In the embodiments provided in this application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of systems, modules and units can be electrical, mechanical or other forms.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A pronunciation correction data analysis method based on three-dimensional dynamic oral modeling, characterized in that: The steps include: Obtaining side views of tongue portions in the three-dimensional oral cavity modeling of a first number of individuals when they pronounce and when they do not pronounce, marking the images as historical pronunciation images and historical unpronounced images, respectively; Get reference dimensions based on historical unvoiced figures; scaling the historical pronunciation image based on the reference size and the historical unpronounced image to obtain a historical correction image; Obtaining reference points of all historical correction maps based on the reference point setting method and marking them as historical reference points includes the following sub-steps: Establish a plane rectangular coordinate system, mark it as the historical reference coordinate system, and place the historical correction diagram in the length coordinate system; Obtain the tongue contour portion in the historical correction image and mark it as the historical contour; The reference point obtained by the reference point setting method using the historical contour as the target contour is marked as the historical reference point, wherein the reference point setting method includes: obtaining a coordinate point at which the horizontal coordinate of the target contour is the smallest, marking it as the starting coordinate point, obtaining a coordinate point at which the horizontal coordinate of the target contour is the largest, marking it as the final coordinate point, moving along the upper half of the target contour with the starting coordinate point as the starting point, obtaining the coordinate point of this position each time a first distance is moved, marking it as the search coordinate point, and stopping until the final coordinate point is reached, starting from the starting coordinate point, connecting the search coordinate points in the order in which the search coordinate points are obtained, obtaining the connecting line segments, obtaining the slopes of the connecting line segments, obtaining the absolute value of the difference between the slope of each connecting line segment and the slope of the next connecting line segment, marking it as the slope change value; obtaining the search coordinate point between the two connecting line segments when the slope change value is the maximum, marking it as the screening starting coordinate point, obtaining all the search coordinate points after the screening starting coordinate point, and marking them as reference points; Perform function fitting on historical reference points to obtain a normal pronunciation function; Obtaining a normal pronunciation range based on a normal pronunciation function and historical reference points; Obtain side views of the tongue portion of the three-dimensional oral model of the person to be analyzed when the person is speaking and when the person is not speaking, and mark them as a detection pronunciation image and a detection unpronounced image respectively; scaling the detected pronunciation image based on the reference size and the detected unpronounced image to obtain a detected correction image; Determine whether there is pronunciation abnormality based on the detection correction graph and the normal pronunciation range.

2. The pronunciation correction data analysis method based on three-dimensional dynamic oral modeling according to claim 1, characterized in that: Obtaining reference dimensions based on historical unvoiced images includes the following sub-steps: Establish a plane rectangular coordinate system, mark it as the length coordinate system, and place the historical unpronounced graph in the length coordinate system; Obtain the minimum and maximum values ​​of the horizontal coordinate of the tongue part in each historical unpronounced graph, and mark them as historical minimum value and historical maximum value respectively; Calculate the difference between the historical maximum and historical minimum values ​​of each historical unvoiced graph, and mark it as the historical difference; Divide the range of historical differences into q equal ranges, marked as difference division ranges; Count the frequency of each difference division range and mark it as division frequency; A histogram is drawn with the historical difference as the X-axis, the divided frequency as the Y-axis, and the difference division range as the histogram interval, and is marked as a difference histogram.

3. The pronunciation correction data analysis method based on three-dimensional dynamic oral modeling according to claim 2, characterized in that: Obtaining the reference size based on the historical unvoiced image also includes the following sub-steps: Get the maximum division frequency in the difference histogram and mark it as the starting frequency; The median frequency threshold was calculated as: F1=0.5*D1 / q; Where F1 is the median frequency threshold, D1 is the first quantity; In the difference histogram, starting from the starting frequency part, determine whether the divided frequency is greater than the median frequency threshold to the right. If so, continue to determine whether the divided frequency is greater than the median frequency threshold to the right, and stop when it is not. Get the maximum value of the horizontal axis of the divided frequency before stopping, and mark it as the maximum value of the screening; In the difference histogram, starting from the starting frequency part, determine whether the divided frequency is greater than the median frequency threshold to the left. If so, continue to determine whether the divided frequency is greater than the median frequency threshold to the left until it is not. Get the minimum value of the horizontal axis of the divided frequency before stopping, and mark it as the minimum value of the screening; Get the average of the historical differences between the maximum and minimum screening values ​​and mark it as a reference size.

4. The method for analyzing pronunciation correction data based on three-dimensional dynamic oral modeling according to claim 3, characterized in that: Scaling the historical pronunciation image based on the reference size and the historical unpronounced image to obtain the historical correction image includes the following sub-steps: The scaling ratio of each person is calculated as: Sb1=Cc / Cl; where Sb1 is the scaling ratio of each person, Cl is the historical difference, and Cc is the reference size; The historical pronunciation graph of each person is horizontally scaled according to its scaling ratio to obtain a historical correction graph.

5. The pronunciation correction data analysis method based on three-dimensional dynamic oral modeling according to claim 4, characterized in that: Acquiring a normal pronunciation range based on a normal pronunciation function and historical reference points includes the following sub-steps: Substituting the horizontal coordinate of the historical reference point into the normal pronunciation function, the value obtained is marked as the function prediction value; Calculate the absolute value of the difference between the function prediction value and the ordinate of the historical reference point, and mark it as the fitting difference; Sort the fitting differences from small to large and label them as Cz1 to Cz i ; Determine whether w1*D1 is an integer. If so, set Cz (w1*D1) As the first position value; if not, get the integers on the left and right sides of w1*D1, mark them as R1 and R2 respectively, and calculate Cz (R1) With Cz (R2) The mean of is taken as the first position value; w1 is the first position coefficient, the range of w1 is: (0, 0.5), and D1 is the number of fitting differences; Determine whether w2*D1 is an integer. If so, set Cz (w2*D1) As the second position value; if not, get the integers on the left and right sides of w2*D1, mark them as R3 and R4 respectively, and calculate Cz (R3) With Cz (R4) The mean of is taken as the second position value; where w2 is 1-w1; The moving distance threshold is obtained as follows: E2=R2+[(R2-R1) / (w2-w1)]*w1; where E2 is the moving distance threshold, R1 is the first position value, and R2 is the second position value; The normal pronunciation function is moved in the positive direction and the negative direction of the vertical axis by a moving distance threshold respectively to obtain an upper threshold function and a lower threshold function, and the range between the upper threshold function and the lower threshold function is marked as the normal pronunciation range.

6. The method for analyzing pronunciation correction data based on three-dimensional dynamic oral modeling according to claim 5, characterized in that: Scaling the detected pronunciation image based on the reference size and the detected unpronounced image to obtain the detected correction image includes the following sub-steps: Establish a plane rectangular coordinate system, mark it as the scaled coordinate system, and place the detected unvoiced graph in the scaled coordinate system; Obtain the difference between the maximum and minimum values ​​of the horizontal coordinate of the tongue part in the unpronounced image, and mark it as the detection difference; The scaling ratio of the person to be analyzed is calculated as: Sb2=Cc / C2; Where Sb2 is the scaling ratio of the person to be analyzed, and C2 is the detection difference; The detection pronunciation graph of the person to be analyzed is horizontally scaled according to its scaling ratio to obtain a detection correction graph.

7. The method for analyzing pronunciation correction data based on three-dimensional dynamic oral cavity modeling according to claim 6, characterized in that: Determining whether there is pronunciation abnormality based on the detection correction image and the normal pronunciation range includes the following sub-steps: Place the inspection correction diagram in the length coordinate system, ensuring that the placement of the inspection correction diagram is consistent with that of the historical correction diagram; Obtain the tongue contour in the detection and correction image, and mark it as the detection contour; The reference points obtained by the reference point setting method using the detection contour as the target contour are marked as detection reference points.

8. The method for analyzing pronunciation correction data based on three-dimensional dynamic oral cavity modeling according to claim 7, characterized in that: Determining whether there is pronunciation abnormality based on the detection correction image and the normal pronunciation range also includes the following sub-steps: The detection reference points are fitted with a function to obtain a detection pronunciation function, and it is determined whether the detection pronunciation function is all within the normal pronunciation range. If not, a voice abnormality signal is issued.

Citation Information

Patent Citations

  • Dynamic pronunciation teaching model construction method based on 3D modeling and oral anatomy

    CN112381913A

  • Method for acquiring vocal organ profile in medical image

    CN102831606A

  • Statistical classification method of dysarthria pronunciation movement abnormal distribution

    CN109360645A