Pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling

Through function fitting and scaling technology, the problem of inaccurate tongue position judgment in three-dimensional dynamic oral modeling is solved, and a higher accuracy of tongue position judgment is achieved.

CN120296641AActive Publication Date: 2025-07-11BEIJING CETEN EDUCATION TECH GRP CO LTD

Patent Information

Application Number
CN202510779091.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the existing pronunciation correction technology, three-dimensional dynamic oral modeling fails to accurately determine whether the tongue position is correct, resulting in a low accuracy in judging the tongue position during pronunciation.

Method used

The normal pronunciation function is obtained by obtaining historical reference points and function fitting. The normal pronunciation range is obtained based on the normal pronunciation function and historical reference points. The reference size and detection of unpronunciation map are used for scaling, and to determine whether the detection pronunciation map and the normal pronunciation range are abnormal.

Benefits of technology

A correct position judgment method based on tongue characteristics was established, which improved the accuracy of tongue position judgment during pronunciation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296641A_ABST
    Figure CN120296641A_ABST
Patent Text Reader

Abstract

The invention discloses a pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling, and relates to the technical field of pronunciation correction, comprising the following steps: performing function fitting on historical reference points to obtain a normal pronunciation function; acquiring a normal pronunciation range based on the normal pronunciation function and the historical reference point; obtaining side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed pronounces and does not pronounce, and respectively marking the side views as a detection pronouncing diagram and a detection non-pronouncing diagram; scaling the detection pronunciation diagram based on the reference size and the detection unpronunciation diagram to obtain a detection correction diagram; judging whether abnormal pronunciation occurs based on the detection correction graph and the normal pronunciation range; the method is used for solving the problem that the accuracy of judging whether the position of the tongue is correct or not during pronunciation is low due to the fact that a method for judging the correct position of the tongue cannot be established through three-dimensional dynamic oral cavity modeling in an existing pronunciation correction technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pronunciation correction, and specifically to a data analysis method for pronunciation correction based on three-dimensional dynamic oral cavity modeling. Background Art

[0002] Traditional pronunciation correction methods mainly rely on auditory feedback. When relying solely on auditory feedback, it is difficult for the pronunciation learner to intuitively perceive the subtle changes in the internal structure of the oral cavity during the pronunciation process, especially whether the position of the tongue is abnormal. The tongue plays a crucial role in the pronunciation process, and the accuracy of its position directly affects the quality of pronunciation. Therefore, a data analysis method for pronunciation correction based on three-dimensional dynamic oral cavity modeling is proposed. However, due to the different sizes of each person's tongue, and the tongue being an irregular shape, and the tongue may curl during pronunciation, it is more difficult to determine whether the position of the tongue is correct, resulting in an inability to accurately judge whether the position of the tongue is accurate. For example, in the patent application with the publication number CN112381913A, a method for constructing a dynamic pronunciation teaching model based on 3D modeling and oral anatomy is disclosed. This solution fails to set the correct position of the tongue during correct pronunciation to compare and judge whether the pronunciation is abnormal, and only judges whether the position of the tongue is correct through experience, resulting in a decrease in the judgment accuracy. In the existing pronunciation correction technology, the three-dimensional dynamic oral cavity modeling fails to establish a method for judging the correct position of the tongue, resulting in a low accuracy in judging whether the position of the tongue is correct during pronunciation. Summary of the Invention

[0003] The present invention aims to at least partly solve one of the technical problems in the prior art. By obtaining a normal pronunciation function through function fitting of historical reference points, obtaining a normal pronunciation range based on the normal pronunciation function and historical reference points, obtaining side views of the tongue part in the three-dimensional oral cavity modeling of a person to be analyzed when pronouncing and when not pronouncing, respectively marked as a detected pronunciation diagram and a detected non-pronunciation diagram, scaling the detected pronunciation diagram based on a reference size and the detected non-pronunciation diagram to obtain a detected corrected diagram, and judging whether there is abnormal pronunciation based on the detected corrected diagram and the normal pronunciation range, so as to solve the problem that in the pronunciation correction technology, the three-dimensional dynamic oral cavity modeling fails to establish a method for judging the correct position of the tongue, resulting in a low accuracy in judging whether the position of the tongue is correct during pronunciation.

[0004] To achieve the above object, the present application provides a data analysis method for pronunciation correction based on three-dimensional dynamic oral cavity modeling, including the following steps: Obtain side views of the tongue part in the three-dimensional oral cavity modeling of a first number of individuals when pronouncing and when not pronouncing, respectively marked as historical pronunciation diagrams and historical non-pronunciation diagrams; Obtain a reference size based on the historical non-pronunciation diagram; Scale the historical pronunciation diagram based on the reference size and the historical non-pronunciation diagram to obtain a historical corrected diagram; Obtain the reference points of all historical correction diagrams based on the reference point setting method, and mark them as historical reference points; Perform function fitting on the historical reference points to obtain a normal pronunciation function; Obtain the normal pronunciation range based on the normal pronunciation function and the historical reference points; Obtain the side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed is pronouncing and not pronouncing, and mark them as the detected pronunciation diagram and the detected non - pronunciation diagram respectively; Scale the detected pronunciation diagram based on the reference size and the detected non - pronunciation diagram to obtain the detected correction diagram; Judge whether there is abnormal pronunciation based on the detected correction diagram and the normal pronunciation range.

[0005] Further, obtaining the reference size based on the historical non - pronunciation diagram includes the following sub - steps: Establish a plane rectangular coordinate system, marked as the length coordinate system, and place the historical non - pronunciation diagram in the length coordinate system; Obtain the minimum and maximum abscissas of the tongue part in each historical non - pronunciation diagram, and mark them as the historical minimum and the historical maximum respectively; Calculate the difference between the historical maximum and the historical minimum of each historical non - pronunciation diagram, and mark it as the historical difference; Evenly divide the range of the historical difference into q equal ranges, marked as the difference division range; Count the frequency of each difference division range, marked as the division frequency; Taking the historical difference as the X - axis, the division frequency as the Y - axis, and the difference division range as the histogram interval, draw a histogram, marked as the difference histogram.

[0006] Further, obtaining the reference size based on the historical non - pronunciation diagram also includes the following sub - steps: Obtain the largest division frequency in the difference histogram, marked as the starting frequency; Calculate the median frequency threshold as: F1 = 0.5 * D1 / q; where F1 is the median frequency threshold and D1 is the first quantity; In the difference histogram, starting from the part of the starting frequency, judge whether the division frequency is greater than the median frequency threshold to the right. If so, continue to judge whether the division frequency is greater than the median frequency threshold to the right until it is not, and then stop. Obtain the maximum value of the abscissa of the division frequency before stopping, marked as the screening maximum value; In the difference histogram, starting from the part of the starting frequency, judge whether the division frequency is greater than the median frequency threshold to the left. If so, continue to judge whether the division frequency is greater than the median frequency threshold to the left until it is not, and then stop. Obtain the minimum value of the abscissa of the division frequency before stopping, marked as the screening minimum value; Obtain the mean value of the historical difference between the screening maximum value and the screening minimum value, marked as the reference size.

[0007] Further, obtaining the historical correction map by scaling the historical pronunciation map based on the reference dimension and the historical unpronounced map includes the following sub-steps: Calculate the scaling ratio for each person as: Sb1 = Cc / Cl; where Sb1 is the scaling ratio for each person, Cl is the historical difference, and Cc is the reference dimension; Horizontally scale the historical pronunciation map of each person according to its scaling ratio to obtain the historical correction map.

[0008] Further, obtaining the reference points of all historical correction maps based on the reference point setting method, marked as historical reference points, includes the following sub-steps: Establish a plane rectangular coordinate system, marked as the historical reference coordinate system, and place the historical correction map in the length coordinate system; Obtain the contour part of the tongue in the historical correction map, marked as the historical contour; The reference points obtained by using the reference point setting method with the historical contour as the target contour are marked as historical reference points.

[0009] Further, the reference point setting method includes: Obtain a coordinate point when the abscissa of the target contour is the smallest, marked as the starting coordinate point, obtain a coordinate point when the abscissa of the target contour is the largest, marked as the final coordinate point, start from the starting coordinate point and move along the upper half of the target contour, obtain the coordinate point at this position every time the first distance is moved, marked as the search coordinate point, stop until reaching the final coordinate point, connect the search coordinate points in the order of obtaining the search coordinate points starting from the starting coordinate point to obtain the connection line segment, obtain the slope of the connection line segment, obtain the absolute value of the difference between the slope of each connection line segment and the slope of the next connection line segment, marked as the slope change value; obtain the search coordinate point in the middle of the two connection line segments when the maximum value of the slope change value is obtained, marked as the screening starting coordinate point, and obtain all the search coordinate points after the screening starting coordinate point, marked as the reference points.

[0010] Further, obtaining the normal pronunciation range based on the normal pronunciation function and the historical reference points includes the following sub-steps: The value obtained by substituting the abscissa of the historical reference point into the normal pronunciation function is marked as the function predicted value; Calculate the absolute value of the difference between the function predicted value and the ordinate of the historical reference point, marked as the fitting difference; Sort the fitting differences from smallest to largest and mark them as Cz1 to Cz i ; Judge whether w1 * D1 is an integer. If so, Cz (w1*D1)As the first position value; if not, obtain the integers on both sides of w1*D1, mark them as R1 and R2 respectively, and calculate Cz (R1) and Cz (R2) The average value of is used as the first position value; where w1 is the first position coefficient, and the range of w1 is: (0, 0.5), and D1 is the number of fitting difference values; Judge whether w2*D1 is an integer. If so, use Cz (w2*D1) as the second position value; if not, obtain the integers on both sides of w2*D1, mark them as R3 and R4 respectively, and calculate Cz (R3) and Cz (R4) The average value of is used as the second position value; where w2 is 1 - w1; Calculate the movement distance threshold as: E2 = R2 + [(R2 - R1) / (w2 - w1)]*w1; where E2 is the movement distance threshold, R1 is the first position value, and R2 is the second position value; Move the normal pronunciation function one movement distance threshold in the positive direction and negative direction of the vertical axis respectively to obtain the upper limit threshold function and the lower limit threshold function, and mark the area between the upper limit threshold function and the lower limit threshold function as the normal pronunciation range.

[0011] Furthermore, scaling the detected pronunciation map based on the reference size and the detected non - pronunciation map to obtain the detected correction map includes the following sub - steps: Establish a rectangular coordinate system, marked as the scaling coordinate system, and place the detected non - pronunciation map in the scaling coordinate system; Obtain the difference between the maximum and minimum abscissas of the tongue part in the detected non - pronunciation map, marked as the detection difference; Calculate the scaling ratio of the person to be analyzed as: Sb2 = Cc / C2; where Sb2 is the scaling ratio of the person to be analyzed, and C2 is the detection difference; Horizontally scale the detected pronunciation map of the person to be analyzed by its scaling ratio to obtain the detected correction map.

[0012] Furthermore, judging whether there is abnormal pronunciation based on the detected correction map and the normal pronunciation range includes the following sub - steps: Place the detected correction map in the length coordinate system, ensuring that the placement position of the detected correction map is the same as that of the historical correction map; Obtain the contour part of the tongue in the detected correction map, marked as the detection contour; The reference points obtained by using the reference point setting method with the detection contour as the target contour are marked as the detection reference points.

[0013] Furthermore, judging whether there is abnormal pronunciation based on the detected correction map and the normal pronunciation range also includes the following sub - steps: Perform function fitting on the detection reference points to obtain a detection pronunciation function, and determine whether the detection pronunciation function is entirely within the normal pronunciation range. If not, send out a pronunciation anomaly signal.

[0014] Advantages of the present invention: The present invention obtains a normal pronunciation function by performing function fitting on historical reference points, obtains a normal pronunciation range based on the normal pronunciation function and historical reference points, obtains side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed is pronouncing and when not pronouncing, which are respectively marked as a detection pronunciation diagram and a detection non - pronunciation diagram, scales the detection pronunciation diagram based on the reference size and the detection non - pronunciation diagram to obtain a detection correction diagram, and determines whether there is a pronunciation anomaly based on the detection correction diagram and the normal pronunciation range. The advantage lies in establishing a judgment method based on the correct position of the tongue characteristics, increasing the accuracy of judging whether the tongue position is correct during pronunciation; The present invention obtains the reference points of all historical correction diagrams through a reference point setting method. The advantage is that, based on the characteristic that the tongue bends upward during pronunciation and the characteristics of the tongue shape, this reference point setting method is set. Judging the tongue position based on the reference points increases the accuracy of judging whether the tongue is in the correct position during pronunciation. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of the steps of the method of the present invention; Figure 2 is a schematic diagram of the difference histogram of the present invention; Figure 3 is a schematic diagram of screening the maximum value and screening the minimum value of the present invention; Figure 4 is a schematic diagram of searching for coordinate points of the present invention; Figure 5 is a schematic diagram of the reference points of the present invention; Figure 6 is a schematic diagram of the normal pronunciation function of the present invention; Figure 7 is a schematic diagram of the upper limit threshold function and the lower limit threshold function of the present invention; Figure 8 is a schematic diagram of the detection pronunciation function of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Example 1, please refer toFigure 1 As shown in the figure, the present application provides a pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling, including the following steps: Step S1, obtain side views of the tongue part in the three-dimensional oral cavity modeling when a first number of individuals are pronouncing and not pronouncing, and mark them as historical pronunciation diagrams and historical non-pronunciation diagrams respectively. The historical pronunciation diagrams and historical non-pronunciation diagrams are directly obtained from an existing database of established three-dimensional oral cavity modeling.

[0018] Step S2, obtain a reference size based on the historical non-pronunciation diagram; Step S2 further includes the following sub-steps: establish a plane rectangular coordinate system, marked as the length coordinate system, and place the historical non-pronunciation diagram in the length coordinate system; Step S201, obtain the minimum and maximum abscissas of the tongue part in each historical non-pronunciation diagram, and mark them as historical minimum and historical maximum respectively; Step S202, calculate the difference between the historical maximum and historical minimum of each historical non-pronunciation diagram, and mark it as the historical difference; Step S203, evenly divide the range of the historical difference into q equal ranges, marked as difference division ranges; Step S204, count the frequency of each difference division range, marked as division frequency; Step S205, draw a histogram with the historical difference as the X-axis, the division frequency as the Y-axis, and the difference division range as the histogram interval, marked as the difference histogram; Step S206, obtain the largest division frequency in the difference histogram, marked as the starting frequency; Step S207, calculate the median frequency threshold as: F1 = 0.5 * D1 / q; where F1 is the median frequency threshold, D1 is the first number; D1 / q is the average division frequency in each case, and when it is less than half, it is considered a smaller frequency, so multiply by 0.5; Step S208, in the difference histogram, starting from the part of the starting frequency, judge whether the division frequency is greater than the median frequency threshold to the right. If so, continue to judge whether the division frequency is greater than the median frequency threshold to the right until it is not, and then stop. Obtain the maximum value of the abscissa of the division frequency before stopping, marked as the screening maximum value; Step S209, in the difference histogram, starting from the part of the starting frequency, judge whether the division frequency is greater than the median frequency threshold to the left. If so, continue to judge whether the division frequency is greater than the median frequency threshold to the left until it is not, and then stop. Obtain the minimum value of the abscissa of the division frequency before stopping, marked as the screening minimum value; Step S210, obtain the mean value of the historical difference between the screening maximum value and the screening minimum value, marked as the reference size; The obtained reference size after screening is more universal; In practical applications, please refer toFigure 2 As shown in Figure 3 Figure 3 , the range of historical differences is evenly divided into 7 difference division ranges. When the first quantity is 220, the median frequency threshold is calculated as: F1 = 0.5 * 220 / 7 = 15.7. The calculation result is rounded to one decimal place. Starting from the part with the starting frequency of 65, judge whether the division frequency 41 is greater than 15.7 to the right. If so, continue to judge to the right until the division frequency 2 is less than 15.7 and stop. The maximum value of the division frequency abscissa before stopping is 6.2 cm, so the screening maximum value is 6.2 cm. Similarly, the screening minimum value is obtained as 5.2 cm. The average value of the historical differences between 6.2 cm and 5.2 cm is 5.66 cm. The calculation result is rounded to two decimal places, so the reference size is 5.66 cm.

[0019] Step S3, scale the historical pronunciation map based on the reference size and the historical non - pronunciation map to obtain the historical correction map; Step S3 also includes the following sub - steps: Step S301, calculate the scaling ratio for each person as: Sb1 = Cc / Cl; where Sb1 is the scaling ratio for each person, Cl is the historical difference, and Cc is the reference size; the scaling ratio obtained when there is no pronunciation is more accurate; Step S302, horizontally scale each person's historical pronunciation map according to its scaling ratio to obtain the historical correction map; In practical applications, for example, when a person's historical difference is 5.66 cm and the reference size is 5.66 cm, then the scaling ratio of this person is: Sb1 = 1, and the historical pronunciation map is horizontally scaled by 1 times to obtain the historical correction map.

[0020] Step S4, obtain the reference points of all historical correction maps based on the reference point setting method, and mark them as historical reference points; Step S4 also includes the following sub - steps: Step S401, establish a plane rectangular coordinate system, marked as the historical reference coordinate system, and place the historical correction map in the length coordinate system; Step S402, obtain the contour part of the tongue in the historical correction map, marked as the historical contour; Step S403, the reference points obtained by using the reference point setting method with the historical contour as the target contour are marked as historical reference points; Step S403 also includes the following sub - steps: Step S40301: Obtain a coordinate point when the abscissa of the target contour is the smallest, mark it as the starting coordinate point, obtain a coordinate point when the abscissa of the target contour is the largest, mark it as the final coordinate point, start from the starting coordinate point and move along the upper half of the target contour. When moving a first distance each time, obtain the coordinate point at this position and mark it as the search coordinate point until reaching the final coordinate point and stop. Connect the search coordinate points in the order of obtaining them starting from the starting coordinate point to obtain a connecting line segment. Obtain the slope of the connecting line segment, and obtain the absolute value of the difference between the slope of each connecting line segment and the slope of the next connecting line segment, marked as the slope change value; Obtain the search coordinate point in the middle of the two connecting line segments when the maximum value of the slope change value is obtained, mark it as the screening starting coordinate point, and obtain all the search coordinate points after the screening starting coordinate point, marked as the reference points; In practical applications, please refer to Figure 4 and Figure 5 As shown, the beneficial effect of this method is that when the tongue bends, the starting coordinate point is not the tip of the tongue. The tip of the tongue is the point with the maximum slope change value relative to the rest of the part. Therefore, it can be determined that the tip of the tongue is between the two connecting line segments when the maximum value of the slope change value occurs.

[0021] Step S5: Perform function fitting on the historical reference points to obtain the normal pronunciation function.

[0022] In practical applications, please refer to Figure 6 As shown, since the historical reference points are for multiple people, a more general normal pronunciation function can be obtained.

[0023] Step S6: Obtain the normal pronunciation range based on the normal pronunciation function and the historical reference points; Step S6 also includes the following sub-steps: Step S601: Mark the value obtained by substituting the abscissa of the historical reference point into the normal pronunciation function as the function prediction value; Step S602: Calculate the absolute value of the difference between the function prediction value and the ordinate of the historical reference point, marked as the fitting difference; Step S603: Sort the fitting differences from smallest to largest and mark them as Cz1 to Cz i ; Step S604: Determine whether w1*D1 is an integer. If so, use Cz (w1*D1) as the first position value; if not, obtain the integers on both sides of w1*D1, marked as R1 and R2 respectively, and obtain Cz (R1) and Cz (R2)Take the mean value of as the first position value; where w1 is the first position coefficient, and the range of w1 is: (0, 0.5), and D1 is the number of fitting difference values; in order to select the value of the first position value in the smaller half area, so w1 is set between 0 and 0.5, and selecting an intermediate value can ensure that the selected first position value has the representativeness of the smaller half area of the value. w1 is preferably selected as 0.25; Step S605, determine whether w2*D1 is an integer. If so, take Cz (w2*D1) as the second position value; if not, obtain the integers on the left and right sides of w2*D1, and mark them as R3 and R4 respectively, and calculate the mean value of Cz (R3) and Cz (R4) as the second position value; where w2 is 1 - w1; in order to select the value of the second position value in the larger half area, because the range of w1 is (0, 0.5), so w2 can be set to 1 - w1. When w1 is 0.25, w2 is 0.75; Step S606, calculate the movement distance threshold as: E2 = R2 + [(R2 - R1) / (w2 - w1)]*w1; where E2 is the movement distance threshold, R1 is the first position value, and R2 is the second position value; when the fitting differences are evenly distributed, calculate the overall distribution difference based on the first position value and the second position value, and then calculate that the fitting difference between the ratios of 1 - w1 to w2 is [(R2 - R1) / (w2 - w1)]*w1. R2 + [(R2 - R1) / (w2 - w1)]*w1 is the maximum value when the fitting differences are evenly distributed. Due to the same pronunciation and similarity, the fitting differences are relatively concentrated. Therefore, when pronouncing correctly, the fitting differences should not exceed the movement distance threshold. So the movement distance threshold is set; Step S607, move the normal pronunciation function one movement distance threshold in the positive direction and negative direction of the vertical axis respectively to obtain the upper limit threshold function and the lower limit threshold function, and mark the area between the upper limit threshold function and the lower limit threshold function as the normal pronunciation range; In practical applications, please refer to Figure 7 As shown, when w1 is 0.25, w2 is 0.75, and the number of fitting difference values is 1540. Determine whether 0.25*1540 = 385 is an integer. Take the fitting difference corresponding to Cz (385) which is 0.13 cm as the first position value; determine whether 0.75*1540 = 1155 is an integer. Take Cz (1155)The corresponding fitting difference is 0.38 cm, which is the second position value. The moving distance threshold is obtained as follows: E2 = 0.38 + [(0.38 - 0.13) / (0.75 - 0.25)] * 0.25 = 0.51 cm. The calculation result is rounded to two decimal places, so the moving distance threshold is 0.51 cm. The normal pronunciation function is moved 0.51 cm in the positive direction and negative direction of the vertical axis respectively to obtain the upper limit threshold function and the lower limit threshold function. The area between the upper limit threshold function and the lower limit threshold function is marked as the normal pronunciation range.

[0024] Step S7: Obtain the side views of the tongue part in the three-dimensional oral cavity models when the person to be analyzed is pronouncing and not pronouncing, and mark them as the detected pronunciation diagram and the detected non - pronunciation diagram respectively.

[0025] Step S8: Scale the detected pronunciation diagram based on the reference size and the detected non - pronunciation diagram to obtain the detected corrected diagram. Step S8 also includes the following sub - steps: Step S801: Establish a plane rectangular coordinate system, marked as the scaling coordinate system, and place the detected non - pronunciation diagram in the scaling coordinate system. Step S802: Obtain the difference between the maximum and minimum abscissas of the tongue part in the detected non - pronunciation diagram, marked as the detected difference. Step S803: Calculate the scaling ratio of the person to be analyzed as: Sb2 = Cc / C2; where Sb2 is the scaling ratio of the person to be analyzed, and C2 is the detected difference. Step S804: Horizontally scale the detected pronunciation diagram of the person to be analyzed by its scaling ratio to obtain the detected corrected diagram. In practical applications, if the detected difference is 5.45 cm, calculate the scaling ratio of the person to be analyzed as: Sb2 = 5.66 / 5.55. Horizontally scale the detected pronunciation diagram of the person to be analyzed by 5.66 / 5.55 times to obtain the detected corrected diagram. Since the tongue sizes of different people are different, scaling them to be consistent is convenient for comparison.

[0026] Step S9: Judge whether there is abnormal pronunciation based on the detected corrected diagram and the normal pronunciation range. Step S9 also includes the following sub - steps: Step S901: Place the detected corrected diagram in the length coordinate system, ensuring that the placement position of the detected corrected diagram is the same as that of the historical corrected diagram; the same placement position is convenient for comparison. Step S902: Obtain the contour part of the tongue part in the detected corrected diagram, marked as the detected contour. Step S903: Use the reference point setting method with the detected contour as the target contour to obtain the reference points, marked as the detected reference points. Step S904: Fit the detected reference points to obtain the detected pronunciation function, and judge whether the detected pronunciation function is entirely within the normal pronunciation range. If not, send out an abnormal pronunciation signal. In practical applications, please refer to Figure 8 As shown, the obtained pronunciation detection function is within the normal pronunciation range between the upper limit threshold function and the lower limit threshold function, so there is no need to send out an abnormal pronunciation signal.

[0027] Embodiment 2, the present application also provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling are run to achieve the following functions: obtaining side views of the tongue part in the three-dimensional oral cavity modeling when a first number of individuals pronounce and when they do not pronounce, and respectively marking them as historical pronunciation maps and historical non-pronunciation maps; obtaining a reference size based on the historical non-pronunciation map; scaling the historical pronunciation map based on the reference size and the historical non-pronunciation map to obtain a historical corrected map; obtaining reference points of all historical corrected maps based on the reference point setting method and marking them as historical reference points; performing function fitting on the historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed pronounces and when they do not pronounce, and respectively marking them as detection pronunciation maps and detection non-pronunciation maps; scaling the detection pronunciation map based on the reference size and the detection non-pronunciation map to obtain a detection corrected map; and determining whether there is abnormal pronunciation based on the detection corrected map and the normal pronunciation range.

[0028] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0029] Embodiment 3. The present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling provided by the above-mentioned various methods. The method includes: obtaining side views of the tongue part in the three-dimensional oral cavity modeling when a first number of individuals pronounce and when they do not pronounce, and respectively marking them as historical pronunciation maps and historical non-pronunciation maps; obtaining a reference size based on the historical non-pronunciation map; scaling the historical pronunciation map based on the reference size and the historical non-pronunciation map to obtain a historical corrected map; obtaining reference points of all historical corrected maps based on a reference point setting method, and marking them as historical reference points; performing function fitting on the historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed pronounces and when they do not pronounce, and respectively marking them as detected pronunciation maps and detected non-pronunciation maps; scaling the detected pronunciation map based on the reference size and the detected non-pronunciation map to obtain a detected corrected map; judging whether there is an abnormal pronunciation based on the detected corrected map and the normal pronunciation range.

[0030] Embodiment 4. The present application also provides a computer-readable storage medium. The present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling are run to achieve the following functions: obtaining side views of the tongue part in the three-dimensional oral cavity modeling when a first number of individuals pronounce and when they do not pronounce, and respectively marking them as historical pronunciation maps and historical non-pronunciation maps; obtaining a reference size based on the historical non-pronunciation map; scaling the historical pronunciation map based on the reference size and the historical non-pronunciation map to obtain a historical corrected map; obtaining reference points of all historical corrected maps based on a reference point setting method, and marking them as historical reference points; performing function fitting on the historical reference points to obtain a normal pronunciation function; obtaining a normal pronunciation range based on the normal pronunciation function and the historical reference points; obtaining side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed pronounces and when they do not pronounce, and respectively marking them as detected pronunciation maps and detected non-pronunciation maps; scaling the detected pronunciation map based on the reference size and the detected non-pronunciation map to obtain a detected corrected map; judging whether there is an abnormal pronunciation based on the detected corrected map and the normal pronunciation range.

[0031] Through the description of the above embodiments, the embodiments of the present invention can be provided as a method, a system or a computer program product. Based on such an understanding, the above technical solution, in essence or the part that contributes to the prior art, can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0032] In the embodiments provided in the present application, it should be understood that the disclosed system or method can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces, and the indirect couplings or communication connections of systems, modules and units can be electrical, mechanical or other forms.

[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for analyzing pronunciation correction data based on three-dimensional dynamic oral cavity modeling, characterized in that It includes the following steps: Obtain the side views of the tongue part in the three-dimensional oral cavity modeling when a first number of individuals pronounce and when they do not pronounce, and mark them as the historical pronunciation map and the historical non - pronunciation map respectively; Obtain the reference size based on the historical non - pronunciation map; Scale the historical pronunciation map based on the reference size and the historical non - pronunciation map to obtain the historical corrected map; Obtain the reference points of all historical corrected maps based on the reference point setting method, and mark them as historical reference points; Perform function fitting on the historical reference points to obtain the normal pronunciation function; Obtain the normal pronunciation range based on the normal pronunciation function and the historical reference points; Obtain the side views of the tongue part in the three-dimensional oral cavity modeling when the person to be analyzed pronounces and when they do not pronounce, and mark them as the detected pronunciation map and the detected non - pronunciation map respectively; Scale the detected pronunciation map based on the reference size and the detected non - pronunciation map to obtain the detected corrected map; Judge whether there is abnormal pronunciation based on the detected corrected map and the normal pronunciation range.

2. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 1, characterized in that Obtaining the reference size based on the historical non - pronunciation map includes the following sub - steps: Establish a plane rectangular coordinate system, marked as the length coordinate system, and place the historical non - pronunciation map in the length coordinate system; Obtain the minimum and maximum values of the abscissa of the tongue part in each historical non - pronunciation map, and mark them as the historical minimum value and the historical maximum value respectively; Calculate the difference between the historical maximum value and the historical minimum value of each historical non - pronunciation map, and mark it as the historical difference; Evenly divide the range of the historical difference into q equal ranges, and mark them as the difference division ranges; Count the frequency of each difference division range, and mark it as the division frequency; Draw a histogram with the historical difference as the X - axis, the division frequency as the Y - axis, and the difference division range as the histogram interval, and mark it as the difference histogram.

3. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 2, wherein Obtaining the reference size based on the historical non - pronunciation map further includes the following sub - steps: Obtain the largest division frequency in the difference histogram, and mark it as the starting frequency; Calculate the median frequency threshold as: F1 = 0.5 * D1 / q; Where F1 is the median frequency threshold and D1 is the first number; In the difference histogram, starting from the part of the starting frequency, judge whether the division frequency is greater than the median frequency threshold to the right. If so, continue to judge whether the division frequency is greater than the median frequency threshold to the right until it is not, and then stop. Obtain the maximum value of the abscissa of the division frequency before stopping, and mark it as the screening maximum value; In the difference histogram, starting from the part of the starting frequency, judge whether the division frequency is greater than the median frequency threshold to the left. If so, continue to judge whether the division frequency is greater than the median frequency threshold to the left until it is not, and then stop. Obtain the minimum value of the abscissa of the division frequency before stopping, and mark it as the screening minimum value; Obtain the mean value of the historical difference between the screening maximum value and the screening minimum value, and mark it as the reference size.

4. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 3, characterized in that, Scaling the historical pronunciation map based on the reference size and the historical non - pronunciation map to obtain the historical corrected map includes the following sub - steps: Calculate the scaling ratio for each person as: Sb1 = Cc / Cl; where Sb1 is the scaling ratio for each person, Cl is the historical difference, and Cc is the reference size; Horizontally scale each person's historical pronunciation map according to its scaling ratio to obtain the historical corrected map.

5. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 4, characterized in that, Obtaining the reference points of all historical corrected maps based on the reference point setting method and marking them as historical reference points includes the following sub - steps: Establish a plane rectangular coordinate system, marked as the historical reference coordinate system, and place the historical correction diagram in the length coordinate system; Obtain the contour part of the tongue in the historical correction diagram, marked as the historical contour; The reference points obtained by using the reference point setting method with the historical contour as the target contour are marked as historical reference points.

6. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 5, characterized in that, The reference point setting method includes: Obtain a coordinate point when the abscissa of the target contour is the smallest, marked as the starting coordinate point, obtain a coordinate point when the abscissa of the target contour is the largest, marked as the final coordinate point, start from the starting coordinate point and move along the upper half of the target contour, and obtain the coordinate point at this position every time the first distance is moved, marked as the search coordinate point, until reaching the final coordinate point and stopping. Connect the search coordinate points in the order of obtaining the search coordinate points starting from the starting coordinate point to obtain a connecting line segment. Obtain the slope of the connecting line segment, and obtain the absolute value of the difference between the slope of each connecting line segment and the slope of the next connecting line segment, marked as the slope change value; obtain the search coordinate point in the middle of the two connecting line segments when the maximum value of the slope change value is obtained, marked as the screening starting coordinate point, and obtain all the search coordinate points after the screening starting coordinate point, marked as the reference points.

7. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 6, characterized in that Obtaining the normal pronunciation range based on the normal pronunciation function and the historical reference points includes the following sub-steps: The value obtained by substituting the abscissa of the historical reference point into the normal pronunciation function is marked as the function prediction value; Calculate the absolute value of the difference between the function prediction value and the ordinate of the historical reference point, marked as the fitting difference; Sort the fitting differences in ascending order and label them as Cz1 to Cz respectively i ; Determine whether w1*D1 is an integer. If it is, take Cz (w1*D1) as the first position value; if not, obtain the integers on the left and right sides of w1*D1, mark them as R1 and R2 respectively, and calculate Cz (R1) and Cz (R2) The average value of is used as the first position value; where w1 is the first position coefficient, and the range of w1 is: (0, 0.5), and D1 is the number of fitting difference values; Judge whether w2 * D1 is an integer. If it is, set Cz (w2*D1) as the second position value; if not, obtain the integers on the left and right sides of w2 * D1, mark them as R3 and R4 respectively, and calculate the mean of Cz (R3) and Cz (R4) as the second position value; where w2 is 1 - w1; Obtain the moving distance threshold as: E2 = R2 + [(R2 - R1) / (w2 - w1)] * w1; where E2 is the moving distance threshold, R1 is the first position value, and R2 is the second position value; Move the normal pronunciation function one moving distance threshold in the positive direction and the negative direction of the vertical axis respectively to obtain the upper limit threshold function and the lower limit threshold function, and mark the area between the upper limit threshold function and the lower limit threshold function as the normal pronunciation range.

8. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 7, characterized in that Scaling the detected pronunciation diagram based on the reference size and the detected non-pronunciation diagram to obtain the detected correction diagram includes the following sub-steps: Establish a plane rectangular coordinate system, marked as the scaling coordinate system, and place the detected non-pronunciation diagram in the scaling coordinate system; Obtain the difference between the maximum and minimum values of the abscissa of the tongue part in the detected non-pronunciation diagram, marked as the detected difference; Calculate the scaling ratio of the person to be analyzed as: Sb2 = Cc / C2; Where Sb2 is the scaling ratio of the person to be analyzed, and C2 is the detected difference; Horizontally scale the detected pronunciation diagram of the person to be analyzed by its scaling ratio to obtain the detected correction diagram.

9. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 8, characterized in that Judging whether there is abnormal pronunciation based on the detected correction diagram and the normal pronunciation range includes the following sub-steps: Place the detected correction diagram in the length coordinate system, ensuring that the placement position of the detected correction diagram is the same as that of the historical correction diagram; Obtain the contour part of the tongue in the detected correction diagram, marked as the detected contour; The reference points obtained by using the reference point setting method with the detected contour as the target contour are marked as the detected reference points.

10. The pronunciation correction data analysis method based on three-dimensional dynamic oral cavity modeling according to claim 9, wherein Judging whether there is abnormal pronunciation based on the detected correction diagram and the normal pronunciation range also includes the following sub-steps: Perform a function fitting on the detection reference points to obtain a detection pronunciation function, and determine whether the detection pronunciation function is entirely within the normal pronunciation range. If not, send out a signal indicating abnormal pronunciation.

Citation Information

Patent Citations

  • Method for acquiring vocal organ profile in medical image

    CN102831606A

  • Statistical classification method of dysarthria pronunciation movement abnormal distribution

    CN109360645A

  • Tongue Imaging in Medical Diagnostic Ultrasound

    US20130281856A1

Cited By

  • Basin ecological flow accounting method and system

    CN120542748A

  • Operation error monitoring method of electric energy meter and electric energy meter

    CN121385776A

  • Underground water balance calculation simulation method based on three-dimensional geological data

    CN121480095A