Speaking state analysis method, computer equipment and computer readable storage medium

By analyzing the mapping of facial status and content data in the teacher's teaching videos and generating state analysis information, the analysis problem of the relationship between teaching content and facial status is solved, and the teacher's emotional control ability on specific content is improved.

CN120372049APending Publication Date: 2025-07-25GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410072495.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology cannot effectively analyze the relationship between teachers' teaching content and facial status during teaching, making it difficult for novice teachers to improve their emotional control ability in different teaching contents.

Method used

By obtaining target video data, extracting image and audio data, determining facial state data and mapping it with content data, generating state analysis information, comparing it with reference map information, providing differential feedback to assist in training.

Benefits of technology

It improves teachers' ability to control facial status on specific teaching content, enhances the robustness and accuracy of state analysis, and helps teachers better adjust facial status to match teaching content needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372049A_ABST
    Figure CN120372049A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of intelligent teaching, in particular to a speaking state analysis method, computer equipment and a computer readable storage medium. The method comprises the following steps: acquiring target video data, extracting image data and audio data from the target video data, determining face state data according to the image data, converting the audio data into content data, determining target mapping information of the face state data and the content data, and generating state analysis information according to the target mapping information and preset reference mapping information. According to the embodiment, the face state data can be mapped to the content data, and the face state is associated with the content, so that the change condition of the face state along with the content can be represented, the state analysis information of the same content in the aspect of the face state can be obtained, and the user experience is improved. Therefore, people can adjust the face state on the same content according to the state analysis information, and the ability of people to control the face state on the content can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of smart teaching technology, and in particular to a speaking state analysis method, a computer device, and a computer-readable storage medium. Background Art

[0002] Facial expressions and facial emotions are important ways of communication, which can directly convey emotions, intentions, and cognition. In the field of education, teachers' emotions are closely related to the quality of teaching. Teachers can not only convey knowledge through expressions and emotions, but also convey attitudes towards life and learning. Excellent teachers can reasonably control their expressions and emotions according to the teaching content to create an appropriate atmosphere and allow students to better integrate into the classroom. However, in reality, many preparatory teachers and novice teachers find it difficult to analyze the teaching emotions required for the teaching content, and it is not easy for them to freely control and express their emotions according to the teaching content during teaching.

[0003] Related technologies mainly focus on how to analyze the emotional state of teachers during teaching to generate teaching quality reports. However, different teachers or the same teacher in different teaching periods teach different teaching contents. The generated teaching quality reports can only help teachers improve teaching quality in a macro sense, but cannot output the relationship between teachers' teaching content and emotions, and thus cannot help teachers improve their emotional control ability in the corresponding teaching content. Summary of the invention

[0004] One purpose of the embodiments of the present application is to provide a speaking state analysis method, a computer device, and a computer-readable storage medium to solve the technical problem that the related technology cannot analyze the connection between content and facial state.

[0005] In a first aspect, an embodiment of the present application provides a method for analyzing a speaking state, comprising:

[0006] Get target video data;

[0007] Extracting image data and audio data from the target video data;

[0008] determining facial status data based on the image data;

[0009] converting the audio data into content data;

[0010] Determining target mapping information between the facial state data and the content data;

[0011] State analysis information is generated according to the target mapping information and preset reference mapping information.

[0012] Optionally, the target video data is configured with a video time sequence. The facial state data includes multiple facial state values, and the content data includes multiple text contents. Determining the target mapping information between the facial state data and the content data includes: generating a state time sequence distribution diagram according to the facial state data, where the state time sequence distribution diagram is used to represent the distribution of each facial state value along the video time sequence; generating a content time sequence distribution diagram according to the content data, where the content time sequence distribution diagram is used to represent the distribution of each text content along the video time sequence; and generating a state-content distribution diagram according to the content time sequence distribution diagram and the state time sequence distribution diagram, where the state-content distribution diagram is used to represent the target mapping information between the facial state value and the text content. This embodiment can fuse the state time sequence distribution diagram and the content time sequence distribution diagram, convert it to the content dimension to represent the change of the facial state data, filter out the interference of time factors, and eliminate the influence caused by the difference in speech speed when different individuals express the same content.

[0013] Optionally, generating the state-content distribution diagram according to the content time sequence distribution diagram and the state time sequence distribution diagram includes: sequentially selecting a text content as the target text content on the content time sequence distribution diagram, determining the target state value corresponding to the target text content according to the state time sequence distribution diagram, and generating the state-content distribution diagram according to the target text content and the target state value. This embodiment uses the target text content of the content time sequence distribution diagram as an index, and can find out the target state value corresponding to the target text content one by one, so as to be able to draw an accurate and reliable state-content distribution diagram.

[0014] Optionally, determining the target state value corresponding to the target text content according to the state time sequence distribution diagram includes: determining the time of the target text content on the video time sequence as the target time, determining the facial state value in the facial state data whose time matches the target time as the reference state value according to the state time sequence distribution diagram, and determining the target state value corresponding to the target text content according to at least one of the reference state values. This embodiment uses the target time as an intermediate value, and can effectively transition to find out the facial state value corresponding to each text content, so as to promote the generation of the state-content distribution diagram.

[0015] Optionally, determining the target state value corresponding to the target text content according to at least one of the reference state values includes: obtaining the average value of multiple reference state values to obtain the target state value corresponding to the target text content. This embodiment can avoid the error caused by the fluctuation of individual reference state values by obtaining the average value of multiple reference state values, effectively smooth the target state value of each text content, and make the target state value of each text content more reliable and accurate.

[0016] Optionally, the form of the text content may be in the form of characters, sentences, or paragraphs.

[0017] Optionally, the generating the status analysis information according to the target mapping information and the preset reference mapping information includes: determining the difference information between the target mapping information and the reference mapping information, and generating the status analysis information according to the difference information. In this embodiment, by generating the difference information, it is beneficial for the user to intuitively understand the gap between himself and the famous teacher, so that the user can conduct repeated training and adjustment to achieve the purpose of learning and growth. Optionally, the target mapping information includes a plurality of target mapping relationships, the target mapping relationship is a mapping relationship between a target state value and a target text content, the reference mapping information includes a plurality of reference mapping relationships, the reference mapping relationship is a mapping relationship between a reference state value and a reference text content, and the determining the difference information between the target mapping information and the reference mapping information includes: calculating the difference between the first state value and the second state value, the first state value is the target state value of the first text content, the first text content is one of the plurality of target text contents, the second state value is the reference state value of the second text content, the second text content is one of the plurality of reference text contents, and generating the difference information according to the differences corresponding to each of the target text contents. In this embodiment, the target state value of each target text content can be subtracted from the reference state value of the reference text content, deeply analyzing the difference between the facial state of the user on each text content and the facial state required by the target content, so that the user can more pertinently adjust the facial state of the text content that does not meet the requirements, so as to quickly help the user improve his own deficiencies at a faster speed, and thus be able to grow faster.

[0018] Optionally, the facial state data includes positivity and / or vitality.

[0019] In a second aspect, an embodiment of the present application provides an analysis device for speech states, including:

[0020] A data acquisition module, configured to acquire target video data;

[0021] A data extraction module, configured to extract image data and audio data from the target video data;

[0022] A state determination module, configured to determine facial state data according to the image data;

[0023] An audio conversion module, configured to convert the audio data into content data;

[0024] An information mapping module, configured to determine target mapping information between the facial state data and the content data;

[0025] A status analysis module, configured to generate status analysis information according to the target mapping information and preset reference mapping information.

[0026] In some embodiments, the target video data is configured with video timing, the facial status data includes multiple facial status values, the content data includes multiple text contents, and the information mapping module is specifically configured to: generate a status timing distribution diagram according to the facial status data, where the status timing distribution diagram is used to represent the distribution of each facial status value along the video timing; generate a content timing distribution diagram according to the content data, where the content timing distribution diagram is used to represent the distribution of each text content along the video timing; and generate a status content distribution diagram according to the content timing distribution diagram and the status timing distribution diagram, where the status content distribution diagram is used to represent the target mapping information between the facial status values and the text contents.

[0027] In some embodiments, the information mapping module is specifically configured to: sequentially select a text content as the target text content on the content timing distribution diagram, determine the target status value corresponding to the target text content according to the status timing distribution diagram, and generate a status content distribution diagram according to the target text content and the target status value.

[0028] In some embodiments, the information mapping module is specifically configured to: determine the time of the target text content in the video timing as the target time, according to the status timing distribution diagram, determine the facial status value in the facial status data whose time matches the target time as the reference status value, and determine the target status value corresponding to the target text content according to at least one reference status value.

[0029] In some embodiments, the information mapping module is specifically configured to: obtain the average value of multiple reference status values to obtain the target status value corresponding to the target text content.

[0030] In some embodiments, the form of the text content may be in the form of characters, sentences, or paragraphs.

[0031] In some embodiments, the status analysis module is specifically configured to: determine the difference information between the target mapping information and the reference mapping information, and generate status analysis information according to the difference information.

[0032] In some embodiments, the target mapping information includes multiple target mapping relationships, where the target mapping relationship is a mapping relationship between a target state value and target text content, and the reference mapping information includes multiple reference mapping relationships, where the reference mapping relationship is a mapping relationship between a reference state value and reference text content. The state analysis module is specifically configured to: calculate the difference between the first state value and the second state value, where the first state value is the target state value of the first text content, the first text content is one of the multiple target text contents, the second state value is the reference state value of the second text content, and the second text content is one of the multiple reference text contents, and generate difference information based on the differences corresponding to each target text content.

[0033] In a third aspect, an embodiment of the present application provides a computer device, including a memory and a processor, the memory is connected to the processor, and the processor is configured to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device implements the above method.

[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the above method.

[0035] The embodiments of the present application can achieve the following technical effects: In the method for analyzing the speaking state provided by the embodiments of the present application, target video data is obtained, image data and audio data are extracted from the target video data, facial state data is determined according to the image data, the audio data is converted into content data, target mapping information between the facial state data and the content data is determined, and state analysis information is generated based on the target mapping information and preset reference mapping information. This embodiment can map the facial state data onto the content data, associate the facial state with the content, and thus can characterize the change of the facial state with the content. After the facial state data and the content data are associated, this embodiment can compare the target mapping information with the reference mapping information, which is beneficial to obtaining state analysis information on the facial state of the same content, so that people can adjust their facial state on the same content according to the state analysis information, and further beneficial to improving people's control ability of their own facial state on this content. In addition, this embodiment maps the facial state data onto the content data. Compared with the method of mapping the facial state data onto time, this embodiment can use the content data as a reference object, eliminating the influence caused by the difference in speech rate when different individuals express the same content, thereby improving the robustness and accuracy of state analysis. Description of the Drawings

[0036] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0037] Figure 1 Schematic diagram of the system architecture of a speech state analysis system provided by an embodiment of the present application;

[0038] Figure 2 Schematic diagram of the flow of an analysis method for speech state provided by an embodiment of the present application;

[0039] Figure 3 Schematic diagram of a state time series distribution diagram provided by an embodiment of the present application;

[0040] Figure 4 Schematic diagram of a content time series distribution diagram provided by an embodiment of the present application;

[0041] Figure 5 Schematic diagram of a state content distribution diagram provided by an embodiment of the present application;

[0042] Figure 6 Schematic diagram of the comparison between the target mapping information and the reference mapping information provided by an embodiment of the present application;

[0043] Figure 7 Schematic diagram of the structure of an analysis device for speech state provided by an embodiment of the present application;

[0044] Figure 8 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0046] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart. Furthermore, the terms "first", "second", "third", etc. used in the present application do not limit the data and the execution order, but only distinguish the same items or similar items with basically the same functions and roles.

[0047] In a related art, the related art includes an image acquisition device and an analysis device. The image acquisition device acquires the human image of a teacher during teaching, transmits the human image to the analysis device, and the analysis device analyzes the human image based on an emotion model to obtain the emotion state of the teacher during teaching, and then generates a teaching quality report according to the emotion state of the teacher. As described above, the related art does not pay attention to what kind of teaching content the teacher has during teaching, and only needs to obtain the human image of the teacher during teaching to analyze the emotion state.

[0048] For example, whether teacher A teaches students to recite the ancient poem "Thoughts in the Silent Night" or teaches students to recite the ancient poem "Leaving Baidi City at Dawn" in class, as long as the image acquisition device obtains the human image of teacher A in class to generate a teaching quality report, the related art does not care whether the ancient poem recited by teacher A currently is "Thoughts in the Silent Night" or "Leaving Baidi City at Dawn". Thus, after obtaining the teaching quality report, teacher A can only improve their own teaching quality as a whole, and will not consciously adjust their emotion state according to the corresponding teaching content, such as adjusting their emotion state to match the expression required for "Thoughts in the Silent Night", or matching the expression required for "Leaving Baidi City at Dawn".

[0049] In another related art, the emotion state of the teacher will be mapped onto the time axis so that the teacher can know the emotion changes when teaching the corresponding teaching content. For example, the related art maps the emotion state of teacher A reciting the ancient poem "Thoughts in the Silent Night" onto the time axis in chronological order to obtain an emotion-time distribution diagram. Through the emotion-time distribution diagram, teacher A can know the relationship between the emotion of teacher A changing with time when teaching "Thoughts in the Silent Night". However, there are differences in the speaking speeds among different individuals. When an individual with a faster speaking speed recites "Thoughts in the Silent Night", the amount of emotion change is less, and when an individual with a slower speaking speed recites "Thoughts in the Silent Night", the amount of emotion change is more. However, both of them meet the expression requirements required for "Thoughts in the Silent Night" in terms of emotion performance. However, the related art cannot construct a benchmark to compatibly compare the emotion-time distribution diagrams of those with a faster speaking speed and those with a slower speaking speed.

[0050] In the embodiment of the present application, the facial state data is mapped to the content data, and the facial state is associated with the content, so as to be able to characterize the change of the facial state with the content. After the facial state data is associated with the content data, this embodiment can compare the target mapping information with the reference mapping information, which is conducive to obtaining the state analysis information of the same content in terms of facial state, so that people can adjust the facial state on the same content according to the state analysis information, and then is conducive to improving people's control ability of their own facial state on this content.

[0051] In the embodiment of the present application, the facial state of a famous teacher teaching the same content is used as the preset reference mapping information to facilitate comparing the facial state of a user expressing the same content, so as to obtain a difference feedback, which is conducive to the user's repeated training based on this to achieve the purpose of self-learning and growth, and helps the user improve the control ability of their own facial state on this content.

[0052] The embodiment of the present application can use the content data as a reference object, eliminating the influence brought by the difference in speech rate when different individuals express the same content, so as to improve the robustness and accuracy of state analysis.

[0053] The embodiment of the present application provides a speech state analysis system. Please refer to Figure 1 , the speech state analysis system includes a camera 11 and a computer device 12, and the camera 11 is communicatively connected to the computer device 12.

[0054] The camera 11 is used to collect the target video data of the user 10. Among them, the camera 11 can be set at any position such as a classroom, a training room, or a room. The user can stand in front of the camera 11 and recite the ancient poem "Thoughts in the Silent Night", and the camera 11 records the person image of the user during the speech process to obtain the target video data.

[0055] The computer device 12 retrieves the target video data from the camera 11 and executes the speech state analysis method provided in each of the following embodiments according to the target video data, so as to obtain the state analysis information.

[0056] The embodiment of the present application provides a speech state analysis method. It can be understood that the speech state analysis method can be applied not only in the teacher teaching training scenario, but also in the crosstalk training scenario, the speech training scenario, the defense training scenario, or the training scenario for self-improving the expression ability, etc.

[0057] Please refer to Figure 2 , the speech state analysis method includes the following steps:

[0058] S21: Obtain the target video data.

[0059] In this step, the target video data is the video data collected from the user's verbal expression of the target content. In some embodiments, the target content may be a paradigmatic content. For example, the ancient poems "Thoughts in the Silent Night" or "Leaving White Emperor City at Dawn" mentioned above both belong to paradigmatic content, or the text "The Background" also belongs to paradigmatic content. In some embodiments, the target content may be the content agreed upon by people. For example, singer B writes a song C, and fan D knows song C, then song C may be the target content.

[0060] It can be understood that the forms of expression of the target content are quite diverse, and this article does not make any limitations on the specific forms of expression of the target content.

[0061] It can be understood that when the user verbally expresses the target content, they need to control their facial expressions or emotions in accordance with the artistic conception of the target content. For example, when the user recites the ancient poem "Thoughts in the Silent Night", they need to show an expression of deep longing for their hometown.

[0062] Obtaining the target video data includes: in response to the analysis command received by the computer device, controlling the computer device to obtain the target video data. In some embodiments, the analysis command can be triggered by the user operating the physical buttons of the computer device to cause the computer device to issue the analysis command. In some embodiments, the analysis command can be triggered by the user operating the UI analysis button on the interaction interface provided by the computer device to cause the computer device to issue the analysis command. In some embodiments, the analysis command can be sent by an external device to the computer device.

[0063] In some embodiments, controlling the computer device to obtain the target video data includes: controlling the computer device to retrieve the target video data from the local. In some embodiments, controlling the computer device to obtain the target video data includes: controlling the computer device to access the camera or intermediate device at a preset frequency to obtain the target video data, where the intermediate device is a device that stores the target video data transmitted by the camera, such as the intermediate device is a USB flash drive, etc.

[0064] S22: Extract the image data and audio data from the target video data.

[0065] In this step, the image data is the data of the person's image when the user verbally expresses the target content, where the image data includes at least one frame of the person's image.

[0066] In some embodiments, extracting the image data from the target video data includes: extracting at least one frame of the person's image from the target video data according to a preset image extraction algorithm, and at least one frame of the person's image can form the image data. For example, in this embodiment, the VideoCapture function provided by OpenCV is used to read the target video data, so as to continuously read out multiple frames of the person's image in a loop.

[0067] In some embodiments, extracting image data from the target video data includes: inputting the target video data into a preset image extraction tool to obtain at least one frame of a human image, and the at least one frame of human image can form the image data.

[0068] The audio data is the data of the voice when collecting the user's verbal expression of the target content. In some embodiments, extracting audio data from the target video data includes: extracting the audio data from the target video data according to a preset audio separation algorithm. The preset audio separation algorithm can be, for example, based on the non-negative matrix factorization algorithm, etc.

[0069] In some embodiments, extracting audio data from the target video data includes: inputting the target video data into a preset audio separation application program for processing to obtain the audio data. The preset audio separation application program can use the audio separation APP provided by related technologies, which will not be elaborated here.

[0070] S23: Determine the facial state data according to the image data.

[0071] In this step, the facial state data is the data of the facial state when the user verbally expresses the target content. The facial state includes the expression state and / or the emotional state, and the facial state includes states such as neutral, happy, sad, surprised, frightened, delayed, angry, contemptuous, etc.

[0072] In some embodiments, the facial state data includes a type of facial state value. For example, the facial state value is the positive degree, and the positive degree is used to represent the degree of the user's investment in the expression and / or emotion of the artistic conception required by the target content. Among them, the positive degree can be a normalized value, that is, the value range of the positive degree is [0, 1]. The degree of investment is positively correlated with the positive degree, that is, the greater the degree of investment, the greater the positive degree, and the smaller the degree of investment, the smaller the positive degree.

[0073] For another example, the facial state value is the vitality degree, and the vitality degree is used to represent the active degree of the user's expression and / or emotion of the artistic conception required by the target content. Similarly, the vitality degree can be a normalized value, that is, the value range of the vitality degree is [0, 1]. The active degree is positively correlated with the vitality degree, that is, the greater the active degree, the greater the vitality degree, and the smaller the active degree, the smaller the vitality degree.

[0074] In some embodiments, the facial state data includes two types of facial state values. For example, the two types of facial state values are the positive degree and the vitality degree respectively.

[0075] The image data includes at least one frame of a human image. Determining the facial state data according to the image data includes the following steps: extracting the face region image from the human image according to a preset face analysis algorithm, and processing the face region image according to a preset facial recognition model to obtain the facial state data.

[0076] The preset face analysis algorithm can be a traditional face recognition algorithm, a neural network algorithm, or a support vector machine classification algorithm. The traditional face recognition algorithms include algorithms based on facial feature points, algorithms based on the entire face image, algorithms based on templates, etc. The neural network algorithms include the YOLO series algorithms, the VGG series algorithms, the ResNet series algorithms, etc.

[0077] The preset face recognition model is trained using a neural network algorithm. Among them, the neural network algorithms supported by the preset face recognition model can also include the YOLO series algorithms, the VGG series algorithms, the ResNet series algorithms, etc.

[0078] S24: Convert the audio data into content data.

[0079] In this step, the content data is a set of texts obtained by converting the audio data into texts. Among them, when the user can accurately express the target content, the content data is consistent with the target content. For example, when the user recites the ancient poem "Thoughts in the Silent Night", when the user does not make mistakes or omit sentences, the content data obtained by converting the audio data is consistent with the full text of the ancient poem "Thoughts in the Silent Night". In other words, the content data is as follows: Thoughts in the Silent Night, Before my bed a pool of light, I wonder if it's frost aground. Looking up, I find the moon bright; Bowing, in homesickness I'm drowned.

[0080] Converting the audio data into content data includes the following steps: Input the audio data into a preset speech recognition model for processing to obtain content data. Among them, the speech recognition model is an ASR (Automatic Speech Recognition) model.

[0081] S25: Determine the target mapping information between the facial state data and the content data.

[0082] In this step, the target mapping information is the information obtained by mapping the facial state data to the content data. The target mapping information can adopt various data representation methods to represent the correlation between the facial state data and the content data.

[0083] Before determining the target mapping information between the facial state data and the content data, the analysis method of the speaking state further includes: judging whether the content data is consistent with the preset target content. If it is consistent, then enter the step of determining the target mapping information between the facial state data and the content data. If it is not consistent, then generate a prompt message, and the prompt message is used to prompt the user to reread the target content.

[0084] S26: Generate state analysis information according to the target mapping information and the preset reference mapping information.

[0085] In this step, the reference mapping information is the mapping information between the reference state data and the reference text data. The reference state data is the data of the facial state when an authoritative person expresses the target content in language, and the reference text data is the data corresponding to the target content. The authoritative person can be a person recognized by the user. For example, the authoritative person is a famous teacher, etc. In the embodiment of the present application, the facial state of a famous teacher when teaching the same content is used as the preset reference mapping information, so as to facilitate comparing the facial state of the user when expressing the same content, so as to obtain a difference feedback. In this way, it is beneficial for the user to perform repeated training accordingly to achieve the purpose of self-learning and growth, and help the user improve the control ability of his own facial state in this content.

[0086] The state analysis information is used to represent the difference between the target mapping information and the reference mapping information. In the embodiment of the present application, the facial state data is mapped to the content data, and the facial state is associated with the content, so as to be able to characterize the change of the facial state with the content. After the facial state data is associated with the content data, this embodiment can compare the target mapping information with the reference mapping information. In this way, it is beneficial to obtain the state analysis information of the same content in terms of facial state, so that people can adjust the facial state on the same content according to the state analysis information, and further help to improve people's control ability of their own facial state in this content.

[0087] In some embodiments, the target video data is configured with a video time sequence, and the video time sequence is the time sequence for collecting the target video data. The time sequences of the facial state data and the content data are both consistent with the video time sequence.

[0088] The facial state data includes multiple facial state values, and the facial state values are used to represent the expression and / or emotional state of the user when expressing the target content in language. All the facial state values of the facial state data are distributed according to the video time sequence.

[0089] The content data includes multiple text contents, and the text contents are the expression contents of the user. All the text contents of the content data are distributed according to the video time sequence. It can be understood that the form of the text content can be in the form of characters or sentences or paragraphs.

[0090] When the form of the text content is in the form of characters, the text content is a single character. For example, when the user recites the ancient poem "Thoughts in the Silent Night", the character "bed" is a text content, the character "front" is another text content, the character "bright" is yet another text content, and so on.

[0091] When the form of the text content is in the form of sentences, the text content is a sentence. For example, "In front of the bed, the moonlight is bright" is a text content, "I wonder if it's frost on the ground" is another text content, "Raising my head, I look at the bright moon" is yet another text content, and so on.

[0092] When the form of the text content is in paragraph form, the text content is a single paragraph. For example, when the user reads aloud "The Background", the paragraph "When I arrived in Nanjing, a friend invited me to go sightseeing and I stayed for a day... It wouldn't be good for them to go!" is one piece of text content, and the paragraph "We crossed the Yangtze River and entered the railway station... How smart we were at that time!" is another piece of text content, and so on.

[0093] In some embodiments, this embodiment divides the content data into multiple pieces of text content that match the specified text form.

[0094] When the specified text form is character form, dividing the content data into multiple pieces of text content that match the specified text form includes: dividing the content data into multiple characters, with each character being a piece of text content, and the multiple characters forming multiple pieces of text content.

[0095] In some embodiments, the content data is configured with multiple pause identifiers. The pause identifiers are used to mark the end of a sentence. It can be understood that the pause identifiers can be represented by timestamps. For example, when audio data is input into an ASR model for processing to obtain the content data, in the content data, a timestamp is inserted at the beginning of each sentence, and this timestamp can be used as a pause identifier.

[0096] When the specified text form is sentence form, in some embodiments, the content data includes multiple characters. Dividing the content data into multiple pieces of text content that match the specified text form includes: determining the set of characters between two pause identifiers as a sentence, with one sentence being a piece of text content, and multiple sentences forming multiple pieces of text content.

[0097] In some embodiments, the content data includes multiple characters. Dividing the content data into multiple pieces of text content that match the specified text form includes: according to a preset clustering algorithm, classifying the multiple characters of the content data into multiple sets of character classes, with each set of character classes including at least one character, and forming sentences with the characters in the set of character classes, with one sentence being a piece of text content, and multiple sentences forming multiple pieces of text content.

[0098] Classifying the multiple characters of the content data into multiple sets of character classes according to a preset clustering algorithm includes: according to a preset clustering algorithm, classifying multiple characters that are time-intensive into the same set of character classes. The preset clustering algorithm includes hierarchical clustering algorithm or K-means algorithm, etc.

[0099] When the specified text form is paragraph form, dividing the content data into multiple pieces of text content that match the specified text form includes: determining each sentence according to the content data, merging a specified number of sentences into a paragraph, with one paragraph being a piece of text content, and multiple paragraphs being multiple pieces of text content.

[0100] In some embodiments, determining each statement according to the content data includes: determining a character set between two pause identifiers as a statement.

[0101] In some embodiments, determining each statement according to the content data includes: classifying multiple characters of the content data into multiple character sets according to a preset clustering algorithm, each character set including at least one character, and forming a statement with the characters of the character set.

[0102] In some embodiments, determining the target mapping information between the facial state data and the content data includes the following steps:

[0103] S251: Generate a state time series distribution diagram according to the facial state data, and the state time series distribution diagram is used to represent the distribution of each facial state value along the video time series.

[0104] S252: Generate a content time series distribution diagram according to the content data, and the content time series distribution diagram is used to represent the distribution of each text content along the video time series.

[0105] S253: Generate a state-content distribution diagram according to the content time series distribution diagram and the state time series distribution diagram, and the state-content distribution diagram is used to represent the target mapping information between the facial state value and the text content.

[0106] In S251, please refer to Figure 3 , the horizontal axis x of the state time series distribution diagram is time, the vertical axis y is the facial state value, and the video time series can be mapped on the time axis. In the state time series distribution diagram, the facial state value corresponds to a time point. For example, the facial state value at time point t1 is 0.6, the facial state value at time point t2 is 0.8, and the facial state value at time point t3 is 0.2.

[0107] In some embodiments, the facial state data includes positivity and vitality. The state time series distribution diagram includes a positivity time series distribution diagram and a vitality time series distribution diagram. The horizontal axis x of the positivity time series distribution diagram is the time axis, the vertical axis y is the positivity, and the positivity corresponds to a time point. The horizontal axis x of the vitality time series distribution diagram is the time axis, the vertical axis y is the vitality, and the vitality corresponds to a time point.

[0108] In S252, please refer to Figure 4 , the horizontal axis x of the content time series distribution diagram is time, and the vertical axis y is the text content. Each corresponding time point or time interval on the time axis can be corresponding to the corresponding text content. For example, the time interval t1-t5 corresponds to the mth text content p m , the time interval t5-t8 corresponds to the (m + 1)th text content p m+1 , the time interval t8-t12 corresponds to the (m + 2)th text content p m+2 .

[0109] In S253, refer to Figure 5 , the horizontal axis of the state content distribution graph is the text content, and the vertical axis is the facial state value. For example, the facial state value corresponding to the m-th text content p m is 0.5, the facial state value corresponding to the (m + 1)-th text content p m+1 is 0.4, and the facial state value corresponding to the (m + 2)-th text content p m+2 is 0.6.

[0110] In some embodiments, the facial state data includes positiveness and vitality. The state time-series distribution graph includes a positiveness time-series distribution graph and a vitality time-series distribution graph. The state content distribution graph includes a positiveness content distribution graph and a vitality content distribution graph. Generating the state content distribution graph according to the content time-series distribution graph and the state time-series distribution graph includes the following steps: generating the positiveness content distribution graph according to the content time-series distribution graph and the positiveness time-series distribution graph, and generating the vitality content distribution graph according to the content time-series distribution graph and the vitality time-series distribution graph.

[0111] This embodiment can fuse the state time-series distribution graph and the content time-series distribution graph, convert it to represent the change of facial state data in the content dimension, filter out the interference of time factors, and eliminate the influence brought by the difference in speech rate when different individuals express the same content. For example, when users D and E recite the ancient poem "Thoughts in the Silent Night", user D has a faster speech rate and finishes reciting in 15 seconds. User E has a slower speech rate and finishes reciting in 20 seconds. If only considering the state time-series distribution graph without considering the content time-series distribution graph, comparing the state time-series distribution graph e1 of user E with the reference state time-series distribution graph F of a famous teacher to obtain the first comparison result, and comparing the state time-series distribution graph e2 of user E with the reference state time-series distribution graph F of a famous teacher to obtain the second comparison result, neither the first comparison result nor the second comparison result can accurately and reliably reflect whether the facial states of users D and E truly match the facial states required for the ancient poem "Thoughts in the Silent Night".

[0112] However, this embodiment represents the change of facial state data in the content dimension, and can be compatible with judging whether the facial state data of users with various speech rates match the facial state data of a famous teacher. In this way, not only can the robustness and accuracy of state analysis be improved, but also the compatibility can be improved.

[0113] In some embodiments, generating the state content distribution graph according to the content time-series distribution graph and the state time-series distribution graph includes the following steps:

[0114] S2531: Sequentially select a text content on the content time-series distribution graph as the target text content.

[0115] S2532: Determine the target state value corresponding to the target text content according to the state time-series distribution graph.

[0116] S2533: Generate a status content distribution map based on the target text content and the target status value.

[0117] In S2531, in this embodiment, the first text content p1 is first selected as the target text content. After determining the target status value of the first text content p1, the second text content p2 is then selected as the target text content. After determining the target status value of the second text content p2, the third text content p3 is then selected as the target text content, and so on, until all the content data is traversed.

[0118] In S2532, the target status value is the facial status value corresponding to the target text content.

[0119] In S2533, one target status value maps to one target text content to form one target mapping relationship, and multiple target mapping relationships constitute all the content of the status content distribution map.

[0120] In this embodiment, using the target text content of the content time sequence distribution map as an index, the target status value corresponding to the target text content can be found one by one, so that an accurate and reliable status content distribution map can be drawn.

[0121] In some embodiments, determining the target status value corresponding to the target text content according to the status time sequence distribution map includes the following steps:

[0122] S25321: Determine the time of the target text content in the video time sequence as the target time.

[0123] S25322: According to the status time sequence distribution map, determine the facial status value in the facial status data whose time matches the target time as the reference status value.

[0124] S25323: Determine the target status value corresponding to the target text content according to at least one reference status value.

[0125] In S25321, the target time can be a time point or a time interval.

[0126] In some embodiments, when the form of the text content is a character form, determining the time of the target text content in the video time sequence as the target time includes: directly using the time point of the target text content in the video time sequence as the target time.

[0127] In some embodiments, when the form of the text content is in the form of a sentence or a paragraph, determining the time of the target text content in the video time sequence as the target time includes: determining the time point of the first character of the target text content in the video time sequence as the first time point, determining the time point of the last character of the target text content in the video time sequence as the second time point, and subtracting the first time point from the second time point to obtain the target time.

[0128] In S25322, in some embodiments, when the form of the text content is in the form of characters, determining the facial state value in the facial state data whose time matches the target time as the reference state value according to the state time sequence distribution diagram includes the following steps: traversing the facial state values whose time is consistent with the target time on the time axis of the state time sequence distribution diagram, and taking the facial state values whose time is consistent with the target time as the reference state values.

[0129] In some embodiments, when the form of the text content is in the form of a sentence or a paragraph, determining the facial state value in the facial state data whose time matches the target time as the reference state value according to the state time sequence distribution diagram includes the following steps: traversing the facial state values whose time falls within the target time on the time axis of the state time sequence distribution diagram, and taking the facial state values whose time falls within the target time as the reference state values.

[0130] In S25323, in some embodiments, determining the target state value corresponding to the target text content according to at least one reference state value includes: obtaining the average value of multiple reference state values to obtain the target state value corresponding to the target text content. In this embodiment, by obtaining the average value of multiple reference state values, the error caused by the fluctuation of individual reference state values can be avoided, effectively smoothing the target state value of each text content, and making the target state value of each text content more reliable and accurate.

[0131] For example, the time interval t1 - t5 corresponds to the mth text content p m , the target time is t1 - t5, and the facial state values (i.e., reference state values) whose time falls within the target time are 0.5, 0.6, 0.5, 0.7, 0.6, 0.5, 0.4, 0.6 respectively.

[0132] 0.5 + 0.6 + 0.5 + 0.7 + 0.6 + 0.5 + 0.4 + 0.6 = 4.4, the number of reference state values is 8, and the average value is 0.55. That is, the target state value corresponding to the target text content is 0.55. Therefore, when the user recites the mth text content p m the target state value is 0.55.

[0133] In this embodiment, taking the target time as the intermediate value can effectively transition to find the facial state value corresponding to each text content, thereby promoting the generation of the state content distribution diagram.

[0134] In some embodiments, generating status analysis information based on target mapping information and preset reference mapping information includes the following steps: determining the difference information between the target mapping information and the reference mapping information, and generating status analysis information according to the difference information. The difference information is information on the degree of difference between the target mapping information and the reference mapping information in the same text content. By generating the difference information in this embodiment, it is beneficial for users to intuitively understand the gap between themselves and the famous teachers, so that users can conduct repeated training and adjustment to achieve the purpose of learning and growth.

[0135] In some embodiments, the target mapping information includes a plurality of target mapping relationships, where the target mapping relationship is the mapping relationship between the target status value and the target text content, the reference mapping information includes a plurality of reference mapping relationships, the reference mapping relationship is the mapping relationship between the reference status value and the reference text content, the reference status value is the facial status value of an authoritative person when expressing the reference text content, and the reference text content is part or all of the target content.

[0136] Determining the difference information between the target mapping information and the reference mapping information includes the following steps: calculating the difference between the first status value and the second status value, where the first status value is the target status value of the first text content, the first text content is one of the plurality of target text contents, the second status value is the reference status value of the second text content, the second text content is one of the plurality of reference text contents, and generating difference information according to the differences corresponding to each target text content.

[0137] For example, as described above, please refer to Figure 6 , in the target mapping information, the time interval t1 - t5 corresponds to the first text content p j , and the target status value of the user when reciting the first text content p j is 0.55.

[0138] In the reference mapping information, the time interval t1 - t5 corresponds to the i-th text content p i , and the reference status value of the user when reciting the i-th text content p i is 0.6. In this embodiment, the difference between the first status value and the second status value, 0.55 - 0.6 = -0.05, is calculated. By analogy, this embodiment generates difference information according to the differences corresponding to each target text content.

[0139] In some embodiments, generating status analysis information according to the difference information includes: directly using the difference information as the status analysis information.

[0140] In some embodiments, generating status analysis information based on the difference information includes: generating analysis conclusion information based on the difference information, and packaging the difference information and the analysis conclusion information into status analysis information.

[0141] This embodiment can subtract the target status value of each target text content from the reference status value of the reference text content, deeply analyze the difference between the user's facial status on each text content and the facial status required by the target content, so that the user can more specifically adjust the facial status of the text content that does not meet the requirements. In this way, it can quickly help the user improve their own deficiencies at a faster speed, and thus can grow faster.

[0142] It should be noted that in the above various embodiments, there is not necessarily a certain order between the above steps. Those of ordinary skill in the art can understand from the description of the embodiments of the present application that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel or exchanged, etc.

[0143] As another aspect of the embodiments of the present application, the embodiments of the present application provide an analysis device for speech status. Among them, the analysis device for speech status can be a software module. The software module includes several instructions, which are stored in a memory. The processor can access this memory and call the instructions for execution to complete the analysis method of speech status described in the above various embodiments.

[0144] In some embodiments, the analysis device for speech status can also be built by hardware devices. For example, the analysis device for speech status can be built by one or more than two chips. Each chip can work in coordination with each other to complete the analysis method of speech status described in the above various embodiments. For another example, the analysis device for speech status can also be built by various logic devices, such as being built by a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0145] Please refer to Figure 7 , the analysis device 700 for speech status includes a data acquisition module 71, a data extraction module 72, a status determination module 73, an audio conversion module 74, an information mapping module 75, and a status analysis module 76.

[0146] The data acquisition module 71 is used to acquire target video data. The data extraction module 72 is used to extract image data and audio data from the target video data. The status determination module 73 is used to determine facial status data according to the image data. The audio conversion module 74 is used to convert the audio data into content data. The information mapping module 75 is used to determine the target mapping information between the facial status data and the content data. The status analysis module 76 is used to generate status analysis information according to the target mapping information and the preset reference mapping information.

[0147] This embodiment can map the facial status data to the content data, associate the facial status with the content, so as to be able to characterize the change of the facial status with the content. After the facial status data is associated with the content data, this embodiment can compare the target mapping information with the reference mapping information, which is beneficial to obtain the status analysis information of the same content in terms of facial status, so that people can adjust their facial status on the same content according to the status analysis information, and then is beneficial to improve people's control ability of their own facial status on this content. In addition, this embodiment maps the facial status data to the content data. Compared with the method of mapping the facial status data to time, this embodiment can use the content data as a reference object, eliminating the influence brought by the difference in speech rate when different individuals express the same content, so as to improve the robustness and accuracy of status analysis.

[0148] In some embodiments, the target video data is configured with a video time sequence. The facial status data includes multiple facial status values. The content data includes multiple text contents. The information mapping module 74 is specifically used for: generating a status time sequence distribution diagram according to the facial status data, where the status time sequence distribution diagram is used to represent the distribution of each facial status value along the video time sequence; generating a content time sequence distribution diagram according to the content data, where the content time sequence distribution diagram is used to represent the distribution of each text content along the video time sequence; generating a status-content distribution diagram according to the content time sequence distribution diagram and the status time sequence distribution diagram, where the status-content distribution diagram is used to represent the target mapping information between the facial status value and the text content.

[0149] In some embodiments, the information mapping module 75 is specifically used for: sequentially selecting a text content as the target text content on the content time sequence distribution diagram, determining the target status value corresponding to the target text content according to the status time sequence distribution diagram, and generating a status-content distribution diagram according to the target text content and the target status value.

[0150] In some embodiments, the information mapping module 75 is specifically used for: determining the time of the target text content on the video time sequence as the target time, according to the status time sequence distribution diagram, determining the facial status value in the facial status data whose time matches the target time as the reference status value, and determining the target status value corresponding to the target text content according to at least one reference status value.

[0151] In some embodiments, the information mapping module 75 is specifically configured to: obtain the average value of multiple reference state values to obtain a target state value corresponding to the target text content.

[0152] In some embodiments, the form of the text content may be in the form of characters, sentences, or paragraphs.

[0153] In some embodiments, the state analysis module 76 is specifically configured to: determine the difference information between the target mapping information and the reference mapping information, and generate state analysis information according to the difference information.

[0154] In some embodiments, the target mapping information includes multiple target mapping relationships, the target mapping relationship is the mapping relationship between the target state value and the target text content, the reference mapping information includes multiple reference mapping relationships, the reference mapping relationship is the mapping relationship between the reference state value and the reference text content, and the state analysis module 75 is specifically configured to: calculate the difference between the first state value and the second state value, the first state value is the target state value of the first text content, the first text content is one of the multiple target text contents, the second state value is the reference state value of the second text content, the second text content is one of the multiple reference text contents, and generate difference information according to the differences corresponding to each target text content.

[0155] It should be noted that the above-mentioned speech state analysis device can execute the speech state analysis method provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in the embodiments of the speech state analysis device, reference can be made to the speech state analysis method provided by the embodiments of the present application.

[0156] See Figure 8 , Figure 8 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 800 includes one or more processors 81 and a memory 82. The memory 82 is connected to one or more processors 81, for example, connected to the processor 81 through a bus.

[0157] The processor 81 is configured to support the computer device in performing the corresponding functions in the methods in the above method embodiments. The processor 81 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The above hardware chip may be an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0158] The memory 82 is used to store program codes and the like. The memory 82 may include a volatile memory (VM), such as a random access memory (RAM); the memory may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory may further include a combination of the above types of memories.

[0159] The memory 82 can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the method for analyzing the speaking state in the embodiments of the present application. By running the non-volatile software programs, instructions, and modules stored in the memory, the processor executes various functional applications and data processing of the method for analyzing the speaking state and the apparatus for analyzing the speaking state, that is, implements the functions of each module or unit of the method for analyzing the speaking state and the apparatus for analyzing the speaking state provided in the above method embodiments.

[0160] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the speech state analysis device, etc. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories may be connected to the speech state analysis device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0161] The one or more modules are stored in the memory and, when executed by the one or more processors, execute the speech state analysis method in any of the above method embodiments. For example, the method steps described in the above method embodiments are executed to implement the functions of the modules described in the above device embodiments.

[0162] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, where the computer program includes program instructions, and the program instructions, when executed by a computer device, cause the computer to execute the method as described in the foregoing embodiments.

[0163] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0164] The foregoing disclosure is only for the preferred embodiments of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A method for analyzing a speaking state, characterized in that, Including: Obtain target video data; Extract image data and audio data from the target video data; Determine facial state data according to the image data; Convert the audio data into content data; Determine the target mapping information between the facial state data and the content data; Generate state analysis information according to the target mapping information and the preset reference mapping information.

2. The analysis method according to claim 1, characterized in that, The target video data is configured with a video time sequence, the facial state data includes multiple facial state values, and the content data includes multiple text contents. The determining the target mapping information between the facial state data and the content data includes: Generate a state time sequence distribution diagram according to the facial state data, and the state time sequence distribution diagram is used to represent the distribution of each facial state value along the video time sequence; Generate a content time sequence distribution diagram according to the content data, and the content time sequence distribution diagram is used to represent the distribution of each text content along the video time sequence; Generate a state-content distribution diagram according to the content time sequence distribution diagram and the state time sequence distribution diagram, and the state-content distribution diagram is used to represent the target mapping information between the facial state value and the text content.

3. The analysis method according to claim 2, wherein The generating the state-content distribution diagram according to the content time sequence distribution diagram and the state time sequence distribution diagram includes: Sequentially select a text content as the target text content on the content time sequence distribution diagram; Determine the target state value corresponding to the target text content according to the state time sequence distribution diagram; Generate a state-content distribution diagram according to the target text content and the target state value.

4. The analysis method according to claim 3, characterized in that, The determining the target state value corresponding to the target text content according to the state time sequence distribution diagram includes: Determine the time of the target text content on the video time sequence as the target time; According to the state time sequence distribution diagram, determine the facial state value in the facial state data whose time matches the target time as the reference state value; Determine the target state value corresponding to the target text content according to at least one of the reference state values.

5. The analysis method according to claim 4, characterized in that The determining the target state value corresponding to the target text content according to at least one of the reference state values includes: obtaining the average value of multiple reference state values to obtain the target state value corresponding to the target text content.

6. The analysis method according to claim 2, wherein The form of the text content can be in the form of characters, sentences or paragraphs.

7. The analysis method according to any one of claims 1 to 6, characterized in that The generating the state analysis information according to the target mapping information and the preset reference mapping information includes: Determine the difference information of the target mapping information relative to the reference mapping information; Generate state analysis information according to the difference information.

8. The analysis method according to claim 7, wherein The target mapping information includes multiple target mapping relationships, the target mapping relationship is the mapping relationship between the target state value and the target text content, the reference mapping information includes multiple reference mapping relationships, and the reference mapping relationship is the mapping relationship between the reference state value and the reference text content. The determining the difference information of the target mapping information relative to the reference mapping information includes: Calculate the difference between the first state value and the second state value, where the first state value is the target state value of the first text content, the first text content is one of the multiple target text contents, the second state value is the reference state value of the second text content, and the second text content is one of the multiple reference text contents; Generate difference information according to the differences corresponding to each of the target text contents.

9. The analysis method according to any one of claims 1 to 6, characterized in that The facial state data includes positivity and / or vitality.

10. An analysis device for speech states, characterized in that, Comprising: A data acquisition module for acquiring target video data; A data extraction module for extracting image data and audio data from the target video data; A state determination module for determining facial state data according to the image data; An audio conversion module for converting the audio data into content data; An information mapping module for determining the target mapping information between the facial state data and the content data; A state analysis module for generating state analysis information according to the target mapping information and the preset reference mapping information.

11. A computer device, characterized in that, Comprising a memory and a processor, the memory is connected to the processor, the processor is configured to execute one or more computer programs stored in the memory, and when the processor executes the one or more computer programs, the computer device implements the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to any one of claims 1-9.