Role determination method, role determination apparatus and electronic device in teaching scenario
By performing speech activity detection and voiceprint feature analysis on audio data in teaching scenarios, the problem of high cost and large computational load in existing technologies for role separation is solved, achieving low-cost and accurate role determination and supporting automated teaching data analysis.
Patent Information
- Application Number
- CN202310749876.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2043-06-21
AI Technical Summary
In existing technologies, role separation in teaching scenarios requires the activation of a pre-set user data model, resulting in high costs and computational demands, and failing to fully utilize existing classroom hardware.
By performing voice activity detection on the first and second audio data, multiple audio segments are obtained respectively. The role information of each audio segment is determined based on voiceprint features and role information. Finally, the target role information of the target audio data is determined by combining the role information of multiple audio segments, without the need to start the user data model.
It reduces equipment requirements and overall costs, while improving the accuracy of role separation, enabling automated teaching data analysis, and providing high-quality structured data support.
Smart Images

Figure CN116778933B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition, and more specifically, to a method, device, computer-readable storage medium, and electronic device for role determination in a teaching scenario. Background Technology
[0002] Modern classrooms are typically equipped with audio acquisition devices, such as lavalier microphones worn by teachers, microphones built into surveillance cameras, or omnidirectional microphones installed in the classroom. Therefore, making full use of existing technologies and equipment to conduct intelligent teaching analysis and automatically obtain teaching data corresponding to a lesson is an essential and necessary technology.
[0003] For speech recognition and role separation of audio data generated during teaching, current technologies typically employ the VPR (Voice Print Recognition) algorithm for role analysis. However, VPR algorithms often require the pre-construction of a user data model before role analysis can be performed. This not only makes system startup cumbersome and maintenance costs high due to the need to start the user data model, but also results in relatively low separation accuracy and fails to fully utilize the existing hardware in modern classrooms. Summary of the Invention
[0004] The main objective of this application is to provide a method, device, computer-readable storage medium, and electronic device for role determination in a teaching scenario, so as to at least solve the problems of high cost and large computational load caused by the need to start a preset user data model for role separation in a teaching scenario in the prior art.
[0005] To achieve the above objectives, according to one aspect of this application, a method for role determination in a teaching scenario is provided, comprising: performing speech activity detection on first audio data to obtain multiple first audio segments; and performing speech activity detection on second audio data to obtain multiple second audio segments, wherein the multiple first audio segments and the multiple second audio segments are arranged in ascending chronological order, and the audio acquisition devices for the first audio data and the second audio data are different; determining the role information of multiple other first audio segments and the role information of the first second audio segment based on the voiceprint features and role information of the first first audio segment, and determining the role information of multiple other second audio segments based on the voiceprint features and role information of the first second audio segment. The first other audio segment is the first audio segment other than the first first audio segment, and the second other audio segment is the second audio segment other than the first second audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment. Based on the role information of multiple first audio segments and multiple second audio segments, the target role information of the target audio data is determined. The target audio data is obtained by performing sound quality evaluation on multiple first audio segments and multiple second audio segments. The target audio data is converted into target text information, and the target role information is added to the corresponding target text information according to the target role information of the target audio data.
[0006] Optionally, based on the voiceprint features and role information of the first audio segment, role information of multiple first other audio segments is determined, including: extracting the voiceprint features of the first audio segment to obtain a first voiceprint feature, and extracting the voiceprint features of multiple first other audio segments to obtain multiple first other voiceprint features; determining the similarity scores between the first first voiceprint feature and each of the first other voiceprint features to obtain multiple first score values; and determining the role information of the first other audio segments corresponding to the first score values that are greater than or equal to a first predetermined threshold as first role information, wherein the first role information is the role information of the first audio segment.
[0007] Optionally, determining the role information of the first second audio segment based on the voiceprint features and role information of the first first audio segment includes: determining a first sub-audio segment in the first second audio segment that has the same timestamp information as the first first audio segment, and determining the role information of the first sub-audio segment as the first role information, wherein the first role information is the role information of the first first audio segment; determining the role information of the second sub-audio segment as the second role information, wherein the second sub-audio segment is any other audio segment in the first second audio segment besides the first sub-audio segment, and the second role information is different from the first role information.
[0008] Optionally, based on the voiceprint features and role information of the first second audio segment, role information of multiple second other audio segments is determined, including: extracting the voiceprint features of the first sub-audio segment to obtain a first sub-voiceprint feature, and extracting the voiceprint features of multiple second other audio segments to obtain multiple second other voiceprint features; determining the similarity scores between the first sub-voiceprint feature and each of the second other voiceprint features to obtain multiple second score values; determining the audio segment in the second other audio segment corresponding to the second score value that is greater than or equal to a first predetermined threshold as a third sub-audio segment, and determining the role information of the third sub-audio segment as the first role information; determining the audio segment in the multiple second other audio segments that does not have the first role information as a fourth sub-audio segment, and determining the role information of the fourth sub-audio segment as the second role information.
[0009] Optionally, determining the target role information of the target audio data based on the role information of multiple first audio segments and multiple second audio segments includes: determining a first target audio segment and a second target audio segment with target timestamp information, wherein the first target audio segment is one of multiple first audio segments and the second target audio segment is one of multiple second audio segments; when the role information of the first target audio segment and the second target audio segment is the same, determining the role information of the first target audio segment and the second target audio segment as the target role information of the audio segment in the target audio data corresponding to the target timestamp information; when the first score value of the first target audio segment is less than the first predetermined threshold and greater than or equal to the second predetermined threshold, and the role information of the second target audio segment is the first role information, determining the first role information as the target role information of the audio segment in the target audio data corresponding to the target timestamp information.
[0010] Optionally, the process of evaluating the sound quality of multiple first audio segments and multiple second audio segments to obtain the target audio data includes: determining a third target audio segment and a fourth target audio segment with the same timestamp information, wherein the third target audio segment is one of the multiple first audio segments and the fourth target audio segment is one of the multiple second audio segments; determining a first sound intensity of the third target audio segment and a second sound intensity of the fourth target audio segment; determining the third target audio segment as the target audio segment when the first sound intensity is greater than the second sound intensity, and determining the fourth target audio segment as the target audio segment when the first sound intensity is less than the second sound intensity; and combining the multiple target audio segments to obtain the target audio data.
[0011] Optionally, converting the target audio data into target text information, and adding the target role information to the corresponding target text information based on the target role information in the target audio data, includes: using an automatic speech recognition algorithm to perform speech recognition on the target audio data to obtain the target text information; and adding the target role information to the corresponding target text information based on the target role information corresponding to the target audio data.
[0012] According to another aspect of this application, a role determination device for a teaching scenario is provided, comprising: a detection unit, configured to perform speech activity detection on first audio data to obtain multiple first audio segments, and to perform speech activity detection on second audio data to obtain multiple second audio segments, wherein the multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices for the first audio data and the second audio data are different; and a first determination unit, configured to determine the role information of multiple first other audio segments and the role information of the first second audio segment based on the voiceprint features and role information of the first first audio segment, and to determine the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment. The first audio segment is the first audio segment excluding the first first audio segment, and the second other audio segment is the second audio segment excluding the first second audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment. The second determining unit is used to determine the target role information of the target audio data based on the role information of the multiple first audio segments and the multiple second audio segments. The target audio data is obtained by performing sound quality evaluation on the multiple first audio segments and the multiple second audio segments. The execution unit is used to convert the target audio data into target text information and add the target role information to the corresponding target text information according to the target role information of the target audio data.
[0013] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to execute any of the aforementioned role determination methods in the teaching scenario.
[0014] According to another aspect of this application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including methods for performing any of the described role determination methods in the teaching scenario.
[0015] Applying the technical solution of this application, firstly, speech activity detection is performed on the first audio data and the second audio data respectively to obtain multiple first audio segments and multiple second audio segments; then, based on the voiceprint features and role information of the first first audio segment, the role information of multiple other first audio segments and the role information of the first second audio segment are determined, and based on the voiceprint features and role information of the first second audio segment, the role information of multiple other second audio segments is determined; then, based on the role information of the multiple first audio segments and the role information of the multiple second audio segments, the target role information of the target audio data is determined; finally, the target audio data is converted into target text information, and the target role information is added to the corresponding target text information according to the target role information of the target audio data. Compared to existing technologies that determine target role information in target audio data by activating a user data model, this solution eliminates the need to activate a user data model. Instead, knowing the role information and voiceprint features of the first audio segment, it determines the role information of multiple first other audio segments and the first second audio segment based on the voiceprint features and role information of the first first audio segment. Furthermore, it determines the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment. This ensures lower equipment requirements and thus lower overall costs. Moreover, this solution determines the target role information of the target audio data based on the role information of multiple first and second audio segments, ensuring relatively accurate target role information. This solves the problem of high cost and computational burden caused by the need to activate a pre-set user data model for role separation in teaching scenarios in existing technologies. Attached Figure Description
[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0017] Figure 1 A hardware structure block diagram of a mobile terminal for performing a role determination method in a teaching scenario, according to an embodiment of this application, is shown.
[0018] Figure 2 A flowchart illustrating a role determination method in a teaching scenario according to an embodiment of this application is shown.
[0019] Figure 3 A flowchart illustrating another role determination method in a teaching scenario provided by an embodiment of this application is shown.
[0020] Figure 4A schematic diagram of a role determination device in a teaching scenario provided according to an embodiment of this application is shown.
[0021] The above figures include the following reference numerals:
[0022] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed Implementation
[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] As described in the background section, the existing technology for role separation in teaching scenarios is costly and computationally intensive due to the need to activate a preset user data model. To address these issues, embodiments of this application provide a role determination method, a role determination device, a computer-readable storage medium, and an electronic device for teaching scenarios.
[0027] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0028] The methods and embodiments provided in this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1This is a hardware structure block diagram of a mobile terminal for a role determination method in a teaching scenario according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0029] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the role determination method in the teaching scenario in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned networks may include wireless networks provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0030] This embodiment provides a role determination method for a teaching scenario running on a mobile terminal, computer terminal, or similar computing device. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0031] Figure 2 This is a flowchart of a role determination method in a teaching scenario according to an embodiment of this application. For example... Figure 2 As shown, the method includes the following steps:
[0032] Step S201: Perform speech activity detection on the first audio data to obtain multiple first audio segments, and perform speech activity detection on the second audio data to obtain multiple second audio segments. The multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices of the first audio data and the second audio data are different.
[0033] In practical applications, the first and second audio data acquired by audio acquisition devices often contain noise, meaning that not all of them consist of valid speech segments. Therefore, performing speech activity detection on the first and second audio data can filter out this noise, ensuring high overall system stability. Specifically, a speech activity detection algorithm (VAD) can be used to detect speech activity in both the first and second audio data. However, this is not limited to using a VAD algorithm; other feasible noise filtering algorithms can also be employed, as long as they effectively filter out the noise in both the first and second audio data.
[0034] In step S201 above, the audio acquisition devices for the first audio data and the second audio data are different. In a specific embodiment of this application, in a teaching scenario, the first audio data can be audio data acquired by a lavalier microphone worn by the teacher. The second audio data can be audio data acquired by an omnidirectional microphone installed in the classroom or by a microphone integrated into a monitoring device installed in the classroom. Furthermore, the overall duration of the first and second audio data can be the same. For example, the duration of the first audio data can be 40 minutes, and the duration of the second audio data can also be 40 minutes.
[0035] In step S201 above, the multiple first audio segments and the multiple second audio segments are arranged in ascending order of time. Specifically, the multiple first audio segments are arranged in ascending order of time, and the multiple second audio segments are arranged in ascending order of time. For simplicity, this application uses the arrangement of multiple first audio segments in ascending order of time as an example. In one specific embodiment, the duration of the first audio data is 40 minutes. The timestamps of the multiple first audio segments are 0-5 minutes, 8-15 minutes, 10-20 minutes, etc., and the multiple first audio segments are arranged in the order of 0-5 minutes, 8-15 minutes, 10-20 minutes, etc.
[0036] Step S202: Based on the voiceprint features and role information of the first audio segment, determine the role information of multiple other first audio segments and the role information of the first second audio segment. Based on the voiceprint features and role information of the first second audio segment, determine the role information of multiple other second audio segments. The first other audio segments are the first audio segments other than the first audio segment. The second other audio segments are the second audio segments other than the first audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment.
[0037] In practical applications, if the first audio data is collected by the teacher's lapel microphone, then the role information of the first audio segment must be the teacher. Therefore, knowing the voiceprint characteristics and role information of the first audio segment, the role information of multiple first other audio segments and the role information of the first second audio segment can be determined based on the voiceprint characteristics and role information of the first audio segment.
[0038] In step S202 above, the timestamp information of any one of the multiple first other audio segments (i.e., including the start time and end time of the first other audio segment) is later than the timestamp information of the first first audio segment. The timestamp information of any one of the multiple second other audio segments (i.e., including the start time and end time of the second other audio segment) is later than the timestamp information of the first second audio segment.
[0039] Step S203: Based on the role information of the multiple first audio segments and the multiple second audio segments, determine the target role information of the target audio data. The target audio data is obtained by performing sound quality evaluation on the multiple first audio segments and the multiple second audio segments.
[0040] Specifically, given the role information of multiple first audio segments and multiple second audio segments, the role information of the multiple first audio segments and the role information of the multiple second audio segments are cross-verified. This ensures that the target role information of the obtained target audio data is relatively accurate, thereby improving the overall accuracy.
[0041] Specifically, the duration of the target audio data is the same as the duration of the first audio data and the second audio data.
[0042] Step S204: Convert the target audio data into target text information, and add the target role information to the corresponding target text information based on the target role information in the target audio data.
[0043] In practical applications, any feasible speech conversion algorithm in the existing technology can be used to convert the target audio data into the target text information.
[0044] In this embodiment, firstly, speech activity detection is performed on the first audio data and the second audio data respectively to obtain multiple first audio segments and multiple second audio segments; then, based on the voiceprint features and role information of the first first audio segment, role information of multiple other first audio segments and the first second audio segment are determined, and based on the voiceprint features and role information of the first second audio segment, role information of multiple other second audio segments is determined; then, based on the role information of the multiple first audio segments and the multiple second audio segments, target role information of the target audio data is determined; finally, the target audio data is converted into target text information, and target role information is added to the corresponding target text information according to the target role information of the target audio data. Compared to existing technologies that determine target role information in target audio data by activating a user data model, this solution eliminates the need to activate a user data model. Instead, knowing the role information and voiceprint features of the first audio segment, it determines the role information of multiple first other audio segments and the first second audio segment based on the voiceprint features and role information of the first first audio segment. Furthermore, it determines the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment. This ensures lower equipment requirements and thus lower overall costs. Moreover, this solution determines the target role information of the target audio data based on the role information of multiple first and second audio segments, ensuring relatively accurate target role information. This solves the problem of high cost and computational burden caused by the need to activate a pre-set user data model for role separation in teaching scenarios in existing technologies.
[0045] In the role determination method for teaching scenarios described in this application, the existing lavalier microphones and omnidirectional microphones in the classroom are utilized. First audio data and second audio data are used respectively (since the second audio data is collected by the omnidirectional microphone or the microphone attached to the monitoring system, the role information in the second audio data includes both teachers and students). Speech activity detection is then performed on the first and second audio data respectively, resulting in multiple first audio segments and multiple second audio segments. Then, based on the voiceprint features and role information of the first first audio segment, the role information of multiple other first audio segments and the role information of the first second audio segment are determined. Furthermore, based on the voiceprint features and role information of the first second audio segment, the role information of multiple other second audio segments is determined. Finally, based on the role information of the multiple first and second audio segments, the target role information of the target audio data is determined. This ingenious method achieves role determination of the target audio data without needing to activate the user data model, thus significantly reducing the manpower and time costs of maintenance. Finally, the target audio data is converted into target text information, and target role information is added to the corresponding target text information based on the target role information in the target audio data. This automatically generates structured data (teacher / student - speech content - timestamp) for teachers and students at the end of a teaching activity, providing high-quality data support for subsequent archiving, performance spot checks, and teaching quality evaluation.
[0046] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0047] In specific implementation, step S202 can be achieved through the following steps: Based on the voiceprint features and role information of the first audio segment, determine the role information of multiple other audio segments, including: extracting the voiceprint features of the first audio segment to obtain a first voiceprint feature, and extracting the voiceprint features of multiple other audio segments to obtain multiple other voiceprint features; determining the similarity scores between the first voiceprint feature and each other voiceprint feature to obtain multiple first score values; and determining the role information of the other audio segments corresponding to the first score values that are greater than or equal to a first predetermined threshold as first role information, wherein the first role information is the role information of the first audio segment. In this scheme, by calculating the similarity between the other voiceprint features and the first voiceprint feature, multiple first score values are obtained, and then the other audio segments with first score values higher than the first predetermined threshold are determined and their role information is determined as the role information of the first audio segment, i.e., the first role information. This further achieves a more ingenious determination of the first role information in the first audio data, i.e., the audio belonging to the teacher.
[0048] In practical applications, the first predetermined threshold can be 90%. Furthermore, any feasible voiceprint feature extraction algorithm in the prior art can be used to extract the voiceprint features of the first audio segment and multiple first other audio segments respectively. This application does not limit the specific algorithm used for voiceprint feature extraction. In a specific embodiment, the VPR (Voice Print Recognition) algorithm can be used to extract the voiceprint features of the first audio segment and multiple first other audio segments.
[0049] Secondly, considering that even when the distance between students and teachers is small, the teacher's lavalier microphone may still capture students' audio. In this case, the similarity between the student's and teacher's voiceprint features is low. Therefore, other audio segments corresponding to first score values below the second predetermined threshold can be left without adding role information. The student's role in the above situation can then be obtained using the second audio data. In one specific embodiment of this application, the second predetermined threshold can be 60%.
[0050] Meanwhile, for other audio segments with a first score of 60% to 90%, temporary character information can be added to these segments. Subsequently, reverse verification can be performed based on the character information of multiple second audio segments, which further ensures that the target character information of the target audio data obtained later is relatively accurate.
[0051] To more easily determine the role information in the second audio data, step S202 of this application can be implemented through the following steps: based on the voiceprint features and role information of the first audio segment, determine the role information of the first second audio segment, including: determining a first sub-audio segment in the first second audio segment that has the same timestamp information as the first audio segment, and determining the role information of the first sub-audio segment as the first role information, wherein the first role information is the role information of the first audio segment; determining the role information of a second sub-audio segment as the second role information, wherein the second sub-audio segment is any other audio segment in the first second audio segment besides the first sub-audio segment, and the second role information is different from the first role information. In practical applications, since the audio acquisition device for the second audio data can be an omnidirectional microphone in the classroom or a microphone built into the monitoring equipment, the second audio data must contain the first audio data. That is, the timestamp information of the first second audio segment and the first first audio segment at least partially overlaps, that is, the timestamp of the first second audio segment is the same as that of the first first audio segment, or the timestamp of the first first audio segment is within the range of the timestamp of the first second audio segment. Therefore, based on the voiceprint features and role information of the first first audio segment, the role information of the first second audio segment can be determined relatively easily.
[0052] Furthermore, if the timestamp of the first audio segment falls within the range of the timestamp of the first audio segment, for example, if the timestamp of the first audio segment is 0-5 minutes and the timestamp of the first audio segment is 0-10 minutes, where 0-5 minutes represents the teacher lecturing and 5-10 minutes represents the student answering questions, then the role information of the first audio segment can be determined as the role information of the audio data corresponding to the 0-5 minute segment (i.e., the first sub-audio segment) of the first second audio segment, i.e., the first role information (teacher). Therefore, the role information of the audio data corresponding to the 5-10 minute segment (i.e., the second sub-audio segment) of the first second audio segment can also be determined as the second role information (i.e., student).
[0053] If the timestamps of the first second audio segment and the first first audio segment are the same, for example, if the timestamps of the first first audio segment are both 0-5 minutes and the timestamps of the first second audio segment are both 0-5 minutes, then the role information of the first first audio segment can be used as the role information of the first second audio segment. In other words, if the timestamps of the first second audio segment and the first first audio segment are the same, the first sub-audio segment can be the entire first second audio segment, and the second sub-audio segment can be nonexistent.
[0054] Step S202 above can also be implemented in the following way: based on the voiceprint features and role information of the first second audio segment, determine the role information of multiple second other audio segments, including: extracting the voiceprint features of the first sub-audio segment to obtain the first sub-voiceprint feature, and extracting the voiceprint features of multiple second other audio segments to obtain multiple second other voiceprint features; determining the similarity scores between the first sub-voiceprint feature and each of the second other voiceprint features to obtain multiple second score values; determining the audio segment in the second other audio segment corresponding to the second score value that is greater than or equal to a first predetermined threshold as the third sub-audio segment, and determining the role information of the third sub-audio segment as the first role information; determining the audio segment in the multiple second other audio segments that does not have the first role information as the fourth sub-audio segment, and determining the role information of the fourth sub-audio segment as the second role information. In this scheme, knowing the role information of the first second audio segment, the similarity score between the second other voiceprint features and the first sub-voiceprint features can be calculated to obtain multiple second score values. Audio segments with second score values greater than or equal to the first predetermined threshold are identified as third sub-audio segments. The role information of the first sub-audio segment is identified as the role information of the third sub-audio segment, i.e., the first role information. The role information of the fourth sub-audio segment is identified as the second role information (i.e., student).
[0055] The following is a specific embodiment to illustrate the above embodiment. Specifically, in the case where the first second audio segment is 0-5 minutes long, it is assumed that all the audio data in this 0-5 minute segment is the teacher's voice, that is, the role information of this 0-5 minute audio data is all the first role information. In this case, the first second audio segment can be the first sub-audio segment mentioned above. For a second other audio segment, such as 6-10 minutes, the second other voiceprint features of the second other audio segment are compared with the voiceprint features of the second audio segment (first sub-audio segment) for similarity scoring. If there is a second score value between the voiceprint features in the 6-8 minute period and the voiceprint features of the second audio segment (first sub-audio segment) that is greater than or equal to a first predetermined threshold, then the audio segment in the 6-8 minute period is determined as the third sub-audio segment, and the role information of the third sub-audio segment is determined as the first role information. At the same time, the 8-10 minute segment, for which the role information has not yet been determined, is determined as the second role information. If the second score of the voiceprint feature in the 6-10 minute period is greater than or equal to the second score of the voiceprint feature of the second audio segment (first sub-audio segment), then the audio segment in the 6-10 minute period is determined as the third sub-audio segment, and the role information of the third sub-audio segment is determined as the first role information. In this case, there is no fourth sub-audio segment.
[0056] In practical applications, the first predetermined threshold can be 90%. Furthermore, any feasible voiceprint feature extraction algorithm in the prior art can be used to extract the voiceprint features of the second audio segment and multiple other second audio segments respectively. This application does not limit the specific algorithm used for voiceprint feature extraction. In one specific embodiment, the VPR (Voice Print Recognition) algorithm can be used to extract the voiceprint features of the second audio segment and multiple other second audio segments.
[0057] To further and more accurately determine the target role information of the target audio data, in some embodiments, step S203 can be implemented through steps S2031, S2032, and S2033. Step S2031: Determine a first target audio segment and a second target audio segment with target timestamp information, wherein the first target audio segment is one of a plurality of first audio segments, and the second target audio segment is one of a plurality of second audio segments; Step S2032: If the role information of the first target audio segment and the second target audio segment is the same, determine the role information of the first target audio segment and the second target audio segment as the target role information of the audio segment in the target audio data corresponding to the target timestamp information; Step S2033: If the first score value of the first target audio segment is less than the first predetermined threshold and greater than or equal to the second predetermined threshold, and the role information of the second target audio segment is the first role information, determine the first role information as the target role information of the audio segment in the target audio data corresponding to the target timestamp information. In other words, if audio segments at the same timestamp in the first and second audio data correspond to the same role information, then that same role information is identified as the role information at the same timestamp in the target audio data. For example, if the role information for the 0-5 minute segment in the first audio data and the 0-5 minute segment in the second audio data is the same—both are first role information—then the role information for the audio segments in the 0-5 minute segment of the target audio data is identified as the first role information. If audio segments at the same timestamp in the first and second audio data correspond to different role information, for example, if the first score value of the 0-5 minute segment in the first audio data falls within the range formed by the second predetermined threshold and the first predetermined threshold, then other role information can be temporary role information, while the role information for the 0-5 minute segment in the second audio data is first role information. In this case, the role information for the audio segments in the 0-5 minute segment of the target audio data is identified as the first role information.
[0058] In addition, apart from the above situations, the character information of audio segments in the target audio data that do not have first character information or second character information is determined as second character information. That is, there are two types of character information in the target audio data: first character information and second character information.
[0059] In some specific implementation processes, step S203 can also be implemented through steps S2031, S2032, S2033, and S2034. Step S2031 involves identifying a third target audio segment and a fourth target audio segment with the same timestamp information. The third target audio segment is one of multiple first audio segments, and the fourth target audio segment is one of multiple second audio segments. Step S2032 involves determining the first sound intensity of the third target audio segment and the second sound intensity of the fourth target audio segment. Step S2033 involves determining the third target audio segment as the target audio segment if the first sound intensity is greater than the second sound intensity, and determining the fourth target audio segment as the target audio segment if the first sound intensity is less than the second sound intensity. Step S2034 involves combining multiple target audio segments to obtain the target audio data. This relatively simple method of determining target audio data with better sound quality further ensures that the subsequent speech-to-text conversion using the target audio data yields more accurate target text information.
[0060] Of course, in practical applications, the target audio data is not limited to obtaining it by evaluating the sound intensity of multiple first and second audio segments. It can also be obtained by establishing a sound quality assessment model based on neural networks or similar technologies to evaluate the sound intensity of multiple first and second audio segments.
[0061] The following specific example illustrates the above embodiments. Example 1: If there is a third target audio segment of 0-5 minutes and a fourth target audio segment of 0-5 minutes, and the sound intensity of the third target audio segment is greater than that of the fourth target audio segment, then the third target audio segment is identified as the target audio segment; if the sound intensity of the fourth target audio segment is greater than that of the third target audio segment, then the fourth target audio segment is identified as the target audio segment. Example 2: If there is a third target audio segment of 0-5 minutes and a fourth target audio segment of 0-10 minutes, audio segments within the same 0-5 minute time interval with the same timestamp can be identified using the method mentioned in Example 1. Since the third target audio segment has no sound within the 5-10 minute time interval, while the fourth target audio segment has sound within the 5-10 minute time interval, the sound intensity of the fourth target audio segment within the 5-10 minute time interval must be greater than that of the third target audio segment within the 5-10 minute time interval. Therefore, the audio segment within the 5-10 minute time interval of the fourth target audio segment can be identified as the target audio segment.
[0062] In some implementations, step S204 can also be achieved through steps S2041 and S2042. Step S2041 involves using an Automatic Speech Recognition (ASR) algorithm to perform speech recognition on the target audio data to obtain the target text information. Step S2042 involves adding the target role information to the corresponding target text information based on the target role information corresponding to the target audio data. In this scheme, the use of an automatic speech recognition algorithm to perform speech recognition on the target audio data ensures that the obtained target text information is relatively accurate and that the speech recognition effect is good, further ensuring that the addition of target role information to the corresponding target text information is relatively accurate.
[0063] To enable those skilled in the art to better understand the technical solution of this application, the implementation process of the role determination method in the teaching scenario of this application will be described in detail below with reference to specific embodiments.
[0064] This embodiment relates to a specific role determination method in a teaching scenario. This role determination method is applied in a server, such as... Figure 3 As shown, it includes the following steps:
[0065] Step S1: Collect first audio data using a lavalier microphone worn by the teacher, and collect second audio data using an omnidirectional microphone.
[0066] Step S2: For the server, it receives the first audio data and the second audio data. The VAD algorithm is then used to perform speech activity detection on the first and second audio data respectively, resulting in multiple first audio segments corresponding to the first audio data and multiple second audio segments corresponding to the second audio data.
[0067] Step S3: Since the multiple first audio segments and multiple second audio segments are arranged in ascending chronological order, the role information of the first first audio segment is automatically registered as the first role information, and the role information of the first second audio segment is automatically registered as the mixed role information.
[0068] Step S4: Perform VPR voiceprint feature extraction on the first audio segment to obtain the first voiceprint feature, and store the first voiceprint feature in the database (i.e., first voiceprint feature storage). Perform VPR voiceprint feature extraction on multiple other audio segments to obtain multiple other voiceprint features. Determine the similarity score between the first voiceprint feature and each other other voiceprint feature to obtain multiple first score values. If the first score value is greater than or equal to a first predetermined threshold, the role information of the corresponding other audio segment is determined as first role information; no role information is added to the other audio segments with first score values less than a second predetermined threshold; the other audio segments with first score values less than the first predetermined threshold but greater than or equal to the second predetermined threshold are determined as temporary role information.
[0069] Step S5: For cases where the first second audio segment is identical to the first first audio segment, but also includes the first first audio segment (the first second audio segment overlaps with the first first audio segment), the first role information of the first first audio segment can be used to modify the role information of the first second audio segment. This allows the role information of the first second audio segment to be either solely the first role information or a mixture of first and second role information. Specifically, the process is as follows: Identify the first sub-audio segment in the first second audio segment that has the same timestamp information as the first first audio segment, and assign the role information of this first sub-audio segment as the first role information. The first role information is the role information of the first first audio segment. Then, assign the role information of the second sub-audio segment as the second role information. The second sub-audio segment is any audio segment in the first second audio segment other than the first sub-audio segment, and the second role information is different from the first role information.
[0070] Step S6: Perform VPR voiceprint feature extraction on the first sub-audio segment and multiple second other audio segments to obtain the first sub-voiceprint feature corresponding to the first sub-audio segment and multiple second other voiceprint features corresponding to the multiple second other audio segments; determine the similarity score between the first sub-voiceprint feature and the multiple second other voiceprint features to obtain multiple second score values; if the second score value is greater than or equal to a first predetermined threshold, determine the audio segment in the corresponding second other audio segment as the third sub-audio segment, and determine the role information of the third sub-audio segment as the first role information; determine the audio segment in the multiple second other audio segments that does not have the first role information as the fourth sub-audio segment, and determine the role information of the fourth sub-audio segment as the second role information.
[0071] Step S7: Evaluate the sound quality of multiple first audio segments and multiple second audio segments to obtain target audio data.
[0072] Step S8: Based on the role information of multiple first audio segments and multiple second audio segments, determine the target role information of the target audio data.
[0073] Step S9: Convert the target audio data into target text information, and add target role information to the corresponding target text information based on the target role information in the target audio data.
[0074] Using the aforementioned role determination method, the first and second audio data can be converted into a composition of target role, target audio data, and target text information, and then archived. This allows for the automated acquisition of complete teaching data for a lesson, providing strong data support for subsequent integration with other comprehensive intelligent educational processing systems, including those handling text or audio.
[0075] This application also provides a role determination device for a teaching scenario. It should be noted that the role determination device for a teaching scenario in this application can be used to execute the role determination method for a teaching scenario provided in this application. This device is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0076] The following describes the role determination device for teaching scenarios provided in the embodiments of this application.
[0077] Figure 4 This is a structural schematic diagram of a role determination device in a teaching scenario according to an embodiment of this application. Figure 4 As shown, the role determination device includes:
[0078] The detection unit 10 is used to perform speech activity detection on the first audio data to obtain multiple first audio segments, and to perform speech activity detection on the second audio data to obtain multiple second audio segments. The multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices of the first audio data and the second audio data are different.
[0079] In practical applications, the first and second audio data acquired by audio acquisition devices often contain noise, meaning that not all of them consist of valid speech segments. Therefore, performing speech activity detection on the first and second audio data can filter out this noise, ensuring high overall system stability. Specifically, a speech activity detection algorithm (VAD) can be used to detect speech activity in both the first and second audio data. However, this is not limited to using a VAD algorithm; other feasible noise filtering algorithms can also be employed, as long as they effectively filter out the noise in both the first and second audio data.
[0080] In the aforementioned detection unit, the audio acquisition devices for the first audio data and the second audio data are different. In one specific embodiment of this application, in a teaching scenario, the first audio data can be audio data acquired by a lavalier microphone worn by the teacher. The second audio data can be audio data acquired by an omnidirectional microphone installed in the classroom or by a microphone integrated into a monitoring device installed in the classroom. Furthermore, the overall duration of the first and second audio data can be the same. For example, the duration of the first audio data can be 40 minutes, and the duration of the second audio data can also be 40 minutes.
[0081] In the aforementioned detection unit, the multiple first audio segments and the multiple second audio segments are arranged in ascending order of time. Specifically, the multiple first audio segments are arranged in ascending order of time, and the multiple second audio segments are arranged in ascending order of time. For simplicity, this application uses the arrangement of multiple first audio segments in ascending order of time as an example. In one specific embodiment, the duration of the first audio data is 40 minutes. The timestamps of the multiple first audio segments are 0-5 minutes, 8-15 minutes, 10-20 minutes, etc., and the multiple first audio segments are arranged in the order of 0-5 minutes, 8-15 minutes, 10-20 minutes, etc.
[0082] The first determining unit 20 is configured to determine the role information of multiple first other audio segments and the role information of the first second audio segment based on the voiceprint features and role information of the first audio segment, and to determine the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment, wherein the first other audio segments are the first audio segments other than the first audio segment, and the second other audio segments are the second audio segments other than the first second audio segment, and the timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment;
[0083] In practical applications, if the first audio data is collected by the teacher's lapel microphone, then the role information of the first audio segment must be the teacher. Therefore, knowing the voiceprint characteristics and role information of the first audio segment, the role information of multiple first other audio segments and the role information of the first second audio segment can be determined based on the voiceprint characteristics and role information of the first audio segment.
[0084] In the aforementioned first determining unit, the timestamp information (i.e., including the start time and end time of the first other audio segment) of any one of the plurality of first other audio segments is later than the timestamp information of the first first audio segment. The timestamp information (i.e., including the start time and end time of the second other audio segment) of any one of the plurality of second other audio segments is later than the timestamp information of the first second audio segment.
[0085] The second determining unit 30 is used to determine the target role information of the target audio data based on the role information of the multiple first audio segments and the multiple second audio segments. The target audio data is obtained by performing a sound quality assessment on the multiple first audio segments and the multiple second audio segments.
[0086] Specifically, given the role information of multiple first audio segments and multiple second audio segments, the role information of the multiple first audio segments and the role information of the multiple second audio segments are cross-verified. This ensures that the target role information of the obtained target audio data is relatively accurate, thereby improving the overall accuracy.
[0087] Specifically, the duration of the target audio data is the same as the duration of the first audio data and the second audio data.
[0088] The execution unit 40 is used to convert the target audio data into target text information, and add the target role information to the corresponding target text information according to the target role information of the target audio data.
[0089] In practical applications, any feasible speech conversion algorithm in the existing technology can be used to convert the target audio data into the target text information.
[0090] In the aforementioned role determination device, the detection unit is used to perform speech activity detection on the first audio data and the second audio data respectively, to obtain multiple first audio segments and multiple second audio segments; the first determination unit is used to determine the role information of multiple first other audio segments and the role information of the first second audio segment based on the voiceprint features and role information of the first first audio segment, and to determine the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment; the second determination unit is used to determine the target role information of the target audio data based on the role information of the multiple first audio segments and the role information of the multiple second audio segments; the execution unit is used to convert the target audio data into target text information, and to add target role information to the corresponding target text information according to the target role information of the target audio data. Compared to existing technologies that determine target role information in target audio data by activating a user data model, this solution eliminates the need to activate a user data model. Instead, knowing the role information and voiceprint features of the first audio segment, it determines the role information of multiple first other audio segments and the first second audio segment based on the voiceprint features and role information of the first first audio segment. Furthermore, it determines the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment. This ensures lower equipment requirements and thus lower overall costs. Moreover, this solution determines the target role information of the target audio data based on the role information of multiple first and second audio segments, ensuring relatively accurate target role information. This solves the problem of high cost and computational burden caused by the need to activate a pre-set user data model for role separation in teaching scenarios in existing technologies.
[0091] In the role determination device for teaching scenarios described in this application, existing classroom microphones (including lavalier and omnidirectional microphones) are used. First and second audio data are collected (since the second audio data is collected by the omnidirectional microphone or the microphone attached to the monitoring system, the role information in the second audio data includes both teachers and students). Voice activity detection is then performed on the first and second audio data to obtain multiple first audio segments and multiple second audio segments. Based on the voiceprint features and role information of the first first audio segment, the role information of multiple other first audio segments and the role information of the first second audio segment are determined. Similarly, based on the voiceprint features and role information of the first second audio segment, the role information of multiple other second audio segments is determined. Finally, based on the role information of the multiple first and second audio segments, the target role information of the target audio data is determined. This ingenious method achieves role determination of the target audio data without activating a user data model, significantly reducing maintenance manpower and time costs. Finally, the target audio data is converted into target text information, and target role information is added to the corresponding target text information based on the target role information in the target audio data. This automatically generates structured data (teacher / student - speech content - timestamp) for teachers and students at the end of a teaching activity, providing high-quality data support for subsequent archiving, performance spot checks, and teaching quality evaluation.
[0092] In the specific implementation process, the aforementioned first determining unit includes a first extraction module, a first determining module, and a second determining module. The first extraction module is used to extract the voiceprint features of the first audio segment to obtain a first voiceprint feature, and to extract the voiceprint features of multiple other audio segments to obtain multiple other voiceprint features. The first determining module is used to determine the similarity scores between the first voiceprint feature and each of the other other voiceprint features to obtain multiple first score values. The second determining module is used to determine the role information of the other audio segments corresponding to the first score values that are greater than or equal to a first predetermined threshold as first role information, which is the role information of the first audio segment. In this scheme, by calculating the similarity between the other voiceprint features and the first voiceprint feature, multiple first score values are obtained. Then, the other audio segments with first score values higher than the first predetermined threshold are determined, and their role information is determined as the role information of the first audio segment, i.e., the first role information. This further achieves a more ingenious determination of the first role information in the first audio data, i.e., audio belonging to the teacher.
[0093] In practical applications, the first predetermined threshold can be 90%. Furthermore, any feasible voiceprint feature extraction algorithm in the prior art can be used to extract the voiceprint features of the first audio segment and multiple first other audio segments respectively. This application does not limit the specific algorithm used for voiceprint feature extraction. In a specific embodiment, the VPR (Voice Print Recognition) algorithm can be used to extract the voiceprint features of the first audio segment and multiple first other audio segments.
[0094] Secondly, considering that even when the distance between students and teachers is small, the teacher's lavalier microphone may still capture students' audio. In this case, the similarity between the student's and teacher's voiceprint features is low. Therefore, other audio segments corresponding to first score values below the second predetermined threshold can be left without adding role information. The student's role in the above situation can then be obtained using the second audio data. In one specific embodiment of this application, the second predetermined threshold can be 60%.
[0095] Meanwhile, for other audio segments with a first score of 60% to 90%, temporary character information can be added to these segments. Subsequently, reverse verification can be performed based on the character information of multiple second audio segments, which further ensures that the target character information of the target audio data obtained later is relatively accurate.
[0096] To more easily determine the role information in the second audio data, the first determining unit of this application includes a third determining module and a fourth determining module. The third determining module is used to determine a first sub-audio segment in the first second audio segment that has the same timestamp information as the first first audio segment, and to determine the role information of the first sub-audio segment as the first role information. The first role information is the role information of the first first audio segment. The fourth determining module is used to determine the role information of a second sub-audio segment as the second role information. The second sub-audio segment is any audio segment in the first second audio segment other than the first sub-audio segment, and the second role information is different from the first role information. In practical applications, since the audio acquisition device for the second audio data can be an omnidirectional microphone in the classroom or a microphone built into the monitoring equipment, the second audio data must contain the first audio data. That is, the timestamp information of the first second audio segment and the first first audio segment at least partially overlaps, that is, the timestamp of the first second audio segment is the same as that of the first first audio segment, or the timestamp of the first first audio segment is within the range of the timestamp of the first second audio segment. Therefore, based on the voiceprint features and role information of the first first audio segment, the role information of the first second audio segment can be determined relatively easily.
[0097] Furthermore, if the timestamp of the first audio segment falls within the range of the timestamp of the first audio segment, for example, if the timestamp of the first audio segment is 0-5 minutes and the timestamp of the first audio segment is 0-10 minutes, where 0-5 minutes represents the teacher lecturing and 5-10 minutes represents the student answering questions, then the role information of the first audio segment can be determined as the role information of the audio data corresponding to the 0-5 minute segment (i.e., the first sub-audio segment) of the first second audio segment, i.e., the first role information (teacher). Therefore, the role information of the audio data corresponding to the 5-10 minute segment (i.e., the second sub-audio segment) of the first second audio segment can also be determined as the second role information (i.e., student).
[0098] If the timestamps of the first second audio segment and the first first audio segment are the same, for example, if the timestamps of the first first audio segment are both 0-5 minutes and the timestamps of the first second audio segment are both 0-5 minutes, then the role information of the first first audio segment can be used as the role information of the first second audio segment. In other words, if the timestamps of the first second audio segment and the first first audio segment are the same, the first sub-audio segment can be the entire first second audio segment, and the second sub-audio segment can be nonexistent.
[0099] The first determining unit includes a second extraction module, a fifth determining module, a sixth determining module, and a seventh determining module. The second extraction module is used to extract the voiceprint features of the first sub-audio segment to obtain a first sub-voiceprint feature, and to extract the voiceprint features of multiple second other audio segments to obtain multiple second other voiceprint features. The fifth determining module is used to determine the similarity scores between the first sub-voiceprint feature and each of the second other voiceprint features to obtain multiple second score values. The sixth determining module is used to determine audio segments in the second other audio segments corresponding to second score values greater than or equal to a first predetermined threshold as third sub-audio segments, and to determine the role information of the third sub-audio segment as the first role information. The seventh determining module is used to determine audio segments in the multiple second other audio segments that do not have the first role information as fourth sub-audio segments, and to determine the role information of the fourth sub-audio segment as the second role information. In this scheme, knowing the role information of the first second audio segment, the similarity score between the second other voiceprint features and the first sub-voiceprint features can be calculated to obtain multiple second score values. Audio segments with second score values greater than or equal to the first predetermined threshold are identified as third sub-audio segments. The role information of the first sub-audio segment is identified as the role information of the third sub-audio segment, i.e., the first role information. The role information of the fourth sub-audio segment is identified as the second role information (i.e., student).
[0100] The following is a specific embodiment to illustrate the above embodiment. Specifically, in the case where the first second audio segment is 0-5 minutes long, it is assumed that all the audio data in this 0-5 minute segment is the teacher's voice, that is, the role information of this 0-5 minute audio data is all the first role information. In this case, the first second audio segment can be the first sub-audio segment mentioned above. For a second other audio segment, such as 6-10 minutes, the second other voiceprint features of the second other audio segment are compared with the voiceprint features of the second audio segment (first sub-audio segment) for similarity scoring. If there is a second score value between the voiceprint features in the 6-8 minute period and the voiceprint features of the second audio segment (first sub-audio segment) that is greater than or equal to a first predetermined threshold, then the audio segment in the 6-8 minute period is determined as the third sub-audio segment, and the role information of the third sub-audio segment is determined as the first role information. At the same time, the 8-10 minute segment, for which the role information has not yet been determined, is determined as the second role information. If the second score of the voiceprint feature in the 6-10 minute period is greater than or equal to the second score of the voiceprint feature of the second audio segment (first sub-audio segment), then the audio segment in the 6-10 minute period is determined as the third sub-audio segment, and the role information of the third sub-audio segment is determined as the first role information. In this case, there is no fourth sub-audio segment.
[0101] In practical applications, the first predetermined threshold can be 90%. Furthermore, any feasible voiceprint feature extraction algorithm in the prior art can be used to extract the voiceprint features of the second audio segment and multiple other second audio segments respectively. This application does not limit the specific algorithm used for voiceprint feature extraction. In one specific embodiment, the VPR (Voice Print Recognition) algorithm can be used to extract the voiceprint features of the second audio segment and multiple other second audio segments.
[0102] To further and more accurately determine the target role information of the target audio data, in some embodiments, the second determining unit includes an eighth determining module, a ninth determining module, and a tenth determining module. The eighth determining module is used to determine a first target audio segment and a second target audio segment having target timestamp information. The first target audio segment is one of a plurality of first audio segments, and the second target audio segment is one of a plurality of second audio segments. The ninth determining module is used to determine the role information of the first target audio segment and the second target audio segment as the target role information of the audio segment in the target audio data corresponding to the target timestamp information when the role information of the first target audio segment and the second target audio segment are the same. The tenth determining module is used to determine the first role information as the target role information of the audio segment in the target audio data corresponding to the target timestamp information when the first score value of the first target audio segment is less than the first predetermined threshold and greater than or equal to the second predetermined threshold, and the role information of the second target audio segment is the first role information. In other words, if audio segments at the same timestamp in the first and second audio data correspond to the same role information, then that same role information is identified as the role information at the same timestamp in the target audio data. For example, if the role information for the 0-5 minute segment in the first audio data and the 0-5 minute segment in the second audio data is the same—both are first role information—then the role information for the audio segments in the 0-5 minute segment of the target audio data is identified as the first role information. If audio segments at the same timestamp in the first and second audio data correspond to different role information, for example, if the first score value of the 0-5 minute segment in the first audio data falls within the range formed by the second predetermined threshold and the first predetermined threshold, then other role information can be temporary role information, while the role information for the 0-5 minute segment in the second audio data is first role information. In this case, the role information for the audio segments in the 0-5 minute segment of the target audio data is identified as the first role information.
[0103] In addition, apart from the above situations, the character information of audio segments in the target audio data that do not have first character information or second character information is determined as second character information. That is, there are two types of character information in the target audio data: first character information and second character information.
[0104] In some specific implementation processes, the aforementioned second determining unit further includes an eleventh determining module, a twelfth determining module, a thirteenth determining module, and a combining module. The eleventh determining module is used to determine a third target audio segment and a fourth target audio segment with the same timestamp information. The third target audio segment is one of multiple first audio segments, and the fourth target audio segment is one of multiple second audio segments. The twelfth determining module is used to determine the first sound intensity of the third target audio segment and the second sound intensity of the fourth target audio segment. The thirteenth determining module is used to determine the third target audio segment as the target audio segment when the first sound intensity is greater than the second sound intensity, and to determine the fourth target audio segment as the target audio segment when the first sound intensity is less than the second sound intensity. The combining module is used to combine multiple target audio segments to obtain the target audio data. This relatively simple method of determining target audio data with better sound quality further ensures that the target text information obtained from subsequent speech conversion using the target audio data is more accurate.
[0105] Of course, in practical applications, the target audio data is not limited to obtaining it by evaluating the sound intensity of multiple first and second audio segments. It can also be obtained by establishing a sound quality assessment model based on neural networks or similar technologies to evaluate the sound intensity of multiple first and second audio segments.
[0106] The following specific example illustrates the above embodiments. Example 1: If there is a third target audio segment of 0-5 minutes and a fourth target audio segment of 0-5 minutes, and the sound intensity of the third target audio segment is greater than that of the fourth target audio segment, then the third target audio segment is identified as the target audio segment; if the sound intensity of the fourth target audio segment is greater than that of the third target audio segment, then the fourth target audio segment is identified as the target audio segment. Example 2: If there is a third target audio segment of 0-5 minutes and a fourth target audio segment of 0-10 minutes, audio segments within the same 0-5 minute time interval with the same timestamp can be identified using the method mentioned in Example 1. Since the third target audio segment has no sound within the 5-10 minute time interval, while the fourth target audio segment has sound within the 5-10 minute time interval, the sound intensity of the fourth target audio segment within the 5-10 minute time interval must be greater than that of the third target audio segment within the 5-10 minute time interval. Therefore, the audio segment within the 5-10 minute time interval of the fourth target audio segment can be identified as the target audio segment.
[0107] In some implementations, the aforementioned execution unit further includes a recognition module and an adding module. The recognition module uses an Automatic Speech Recognition (ASR) algorithm to perform speech recognition on the target audio data to obtain the target text information. The adding module adds the target role information to the corresponding target text information based on the target role information corresponding to the target audio data. In this scheme, the use of an automatic speech recognition algorithm to perform speech recognition on the target audio data ensures that the obtained target text information is relatively accurate and that the speech recognition effect is good, further ensuring the accuracy of adding target role information to the corresponding target text information.
[0108] The role determination device in the aforementioned teaching scenario includes a processor and a memory. The detection unit, the first determination unit, the second determination unit, and the execution unit are all stored as program units in the memory. The processor executes the program units stored in the memory to achieve the corresponding functions. All of the above modules are located in the same processor; alternatively, the above modules may be located in different processors in any combination.
[0109] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and adjusting kernel parameters can address the high cost and computational demands of role separation in teaching scenarios, which requires launching a pre-defined user data model.
[0110] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0111] This invention provides a computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device containing the computer-readable storage medium to perform the role determination method in the teaching scenario.
[0112] Specifically, methods for determining roles in teaching scenarios include:
[0113] Step S201: Perform speech activity detection on the first audio data to obtain multiple first audio segments, and perform speech activity detection on the second audio data to obtain multiple second audio segments. The multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices of the first audio data and the second audio data are different.
[0114] Step S202: Based on the voiceprint features and role information of the first audio segment, determine the role information of multiple other first audio segments and the role information of the first second audio segment. Based on the voiceprint features and role information of the first second audio segment, determine the role information of multiple other second audio segments. The first other audio segments are the first audio segments other than the first audio segment. The second other audio segments are the second audio segments other than the first audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment.
[0115] Step S203: Based on the role information of the multiple first audio segments and the multiple second audio segments, determine the target role information of the target audio data. The target audio data is obtained by performing sound quality evaluation on the multiple first audio segments and the multiple second audio segments.
[0116] Step S204: Convert the target audio data into target text information, and add the target role information to the corresponding target text information based on the target role information in the target audio data.
[0117] This invention provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor is configured to execute the role determination method in the teaching scenario described above through the computer program.
[0118] Specifically, methods for determining roles in teaching scenarios include:
[0119] Step S201: Perform speech activity detection on the first audio data to obtain multiple first audio segments, and perform speech activity detection on the second audio data to obtain multiple second audio segments. The multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices of the first audio data and the second audio data are different.
[0120] Step S202: Based on the voiceprint features and role information of the first audio segment, determine the role information of multiple other first audio segments and the role information of the first second audio segment. Based on the voiceprint features and role information of the first second audio segment, determine the role information of multiple other second audio segments. The first other audio segments are the first audio segments other than the first audio segment. The second other audio segments are the second audio segments other than the first audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment.
[0121] Step S203: Based on the role information of the multiple first audio segments and the multiple second audio segments, determine the target role information of the target audio data. The target audio data is obtained by performing sound quality evaluation on the multiple first audio segments and the multiple second audio segments.
[0122] Step S204: Convert the target audio data into target text information, and add the target role information to the corresponding target text information based on the target role information in the target audio data.
[0123] This invention provides a device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs at least the following steps:
[0124] Step S201: Perform speech activity detection on the first audio data to obtain multiple first audio segments, and perform speech activity detection on the second audio data to obtain multiple second audio segments. The multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices of the first audio data and the second audio data are different.
[0125] Step S202: Based on the voiceprint features and role information of the first audio segment, determine the role information of multiple other first audio segments and the role information of the first second audio segment. Based on the voiceprint features and role information of the first second audio segment, determine the role information of multiple other second audio segments. The first other audio segments are the first audio segments other than the first audio segment. The second other audio segments are the second audio segments other than the first audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment.
[0126] Step S203: Based on the role information of the multiple first audio segments and the multiple second audio segments, determine the target role information of the target audio data. The target audio data is obtained by performing sound quality evaluation on the multiple first audio segments and the multiple second audio segments.
[0127] Step S204: Convert the target audio data into target text information, and add the target role information to the corresponding target text information based on the target role information in the target audio data.
[0128] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.
[0129] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program having at least the following method steps:
[0130] Step S201: Perform speech activity detection on the first audio data to obtain multiple first audio segments, and perform speech activity detection on the second audio data to obtain multiple second audio segments. The multiple first audio segments and the multiple second audio segments are arranged in ascending order of time, and the audio acquisition devices of the first audio data and the second audio data are different.
[0131] Step S202: Based on the voiceprint features and role information of the first audio segment, determine the role information of multiple other first audio segments and the role information of the first second audio segment. Based on the voiceprint features and role information of the first second audio segment, determine the role information of multiple other second audio segments. The first other audio segments are the first audio segments other than the first audio segment. The second other audio segments are the second audio segments other than the first audio segment. The timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment.
[0132] Step S203: Based on the role information of the multiple first audio segments and the multiple second audio segments, determine the target role information of the target audio data. The target audio data is obtained by performing sound quality evaluation on the multiple first audio segments and the multiple second audio segments.
[0133] Step S204: Convert the target audio data into target text information, and add the target role information to the corresponding target text information based on the target role information in the target audio data.
[0134] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0135] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0136] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0137] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0138] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0139] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0140] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0141] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0142] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0143] As can be seen from the above description, the embodiments of this application achieve the following technical effects:
[0144] 1) Compared with the prior art, which determines the target role information in the target audio data by activating a user data model, the role determination method of this application does not require activating a user data model. Instead, knowing the role information and voiceprint features of the first audio segment, it determines the role information of multiple first other audio segments and the role information of the first second audio segment based on the voiceprint features and role information of the first first audio segment. Furthermore, it determines the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment. This ensures lower equipment requirements and thus lower overall costs. Moreover, this solution determines the target role information of the target audio data based on the role information of multiple first and second audio segments, ensuring relatively accurate target role information. This solves the problem of high cost and large computational load caused by the need to activate a preset user data model for role separation in teaching scenarios in the prior art.
[0145] 2) Compared with existing technologies that determine target role information in target audio data by activating a user data model, this solution eliminates the need to activate a user data model. Instead, knowing the role information and voiceprint features of the first audio segment, it determines the role information of multiple first other audio segments and the first second audio segment based on the voiceprint features and role information of the first first audio segment. Furthermore, it determines the role information of multiple second other audio segments based on the voiceprint features and role information of the first second audio segment. This ensures lower equipment requirements and thus lower overall costs. Moreover, this solution determines the target role information of the target audio data based on the role information of multiple first and second audio segments, ensuring relatively accurate target role information. This solves the problem of high cost and large computational load caused by the need to activate a preset user data model for role separation in teaching scenarios in existing technologies.
[0146] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A role determination method in a teaching scenario, characterized in that, The method comprises the following steps: voice activity detection is performed on first audio data to obtain a plurality of first audio segments, and voice activity detection is performed on second audio data to obtain a plurality of second audio segments, the plurality of first audio segments and the plurality of second audio segments are arranged in order of time from small to large respectively, and the audio acquisition devices of the first audio data and the second audio data are different; based on the voiceprint feature and the role information of the first first audio segment, the role information of a plurality of first other audio segments and the role information of the first second audio segment are determined, and based on the voiceprint feature and the role information of the first second audio segment, the role information of a plurality of second other audio segments is determined, the first other audio segment is the first audio segment other than the first first audio segment, the second other audio segment is the second audio segment other than the first second audio segment, and the timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment; based on the role information of the plurality of first audio segments and the role information of the plurality of second audio segments, target role information of target audio data is determined, the target audio data is obtained by performing sound quality evaluation on the plurality of first audio segments and the plurality of second audio segments; the target audio data is converted into target text information, and the target role information is added to the corresponding target text information according to the target role information of the target audio data.
2. The role determination method of claim 1, wherein, based on the voiceprint feature and the role information of the first first audio segment, the role information of a plurality of first other audio segments is determined, comprising: extracting the voiceprint feature of the first first audio segment to obtain a first first voiceprint feature, and extracting the voiceprint feature of a plurality of first other audio segments to obtain a plurality of first other voiceprint features; determining the similarity score of the first first voiceprint feature and each first other voiceprint feature to obtain a plurality of first score values; the role information of the first other audio segment corresponding to the first score value greater than or equal to the first predetermined threshold is determined as the first role information, and the first role information is the role information of the first first audio segment.
3. The role determination method of claim 1, wherein, based on the voiceprint feature and the role information of the first first audio segment, the role information of the first second audio segment is determined, comprising: determining a first sub-audio segment in the first second audio segment that has the same timestamp information as the first first audio segment, and determining the role information of the first sub-audio segment as the first role information, the first role information being the role information of the first first audio segment; the role information of a second sub-audio segment is determined as second role information, the second sub-audio segment being an audio segment other than the first sub-audio segment in the first second audio segment, and the second role information being different from the first role information.
4. The role determination method of claim 3, wherein, Determine the role information of the plurality of second other audio segments based on the voiceprint features and the role information of the first second audio segment, including: Extract the voiceprint features of the first sub-audio segment to obtain first sub-voiceprint features, and extract the voiceprint features of the plurality of second other audio segments to obtain a plurality of second other voiceprint features; Determine the similarity scores of the first sub-voiceprint features and each of the second other voiceprint features to obtain a plurality of second score values; Determine the audio segment in the second other audio segment corresponding to the second score value greater than or equal to the first predetermined threshold as a third sub-audio segment, and determine the role information of the third sub-audio segment as the first role information; Determine the audio segment in the plurality of second other audio segments that does not have the first role information as a fourth sub-audio segment, and determine the role information of the fourth sub-audio segment as the second role information.
5. The role determination method of claim 2, wherein, Determine the target role information of the target audio data based on the role information of the plurality of first audio segments and the role information of the plurality of second audio segments, including: Determine a first target audio segment and a second target audio segment with target timestamp information, the first target audio segment being one of the plurality of first audio segments, and the second target audio segment being one of the plurality of second audio segments; In the case where the role information of the first target audio segment and the second target audio segment is the same, determine the role information of the first target audio segment and the second target audio segment as the target role information of the audio segment in the target audio data corresponding to the target timestamp information; In the case where the first score value of the first target audio segment is less than the first predetermined threshold and greater than or equal to a second predetermined threshold, and the role information of the second target audio segment is the first role information, determine the first role information as the target role information of the audio segment in the target audio data corresponding to the target timestamp information.
6. The role determination method according to any one of claims 1 to 5, characterized by, The process of performing sound quality evaluation on the plurality of first audio segments and the plurality of second audio segments to obtain the target audio data includes: Determine a third target audio segment and a fourth target audio segment with the same timestamp information, the third target audio segment being one of the plurality of first audio segments, and the fourth target audio segment being one of the plurality of second audio segments; Determine a first sound intensity of the third target audio segment, and determine a second sound intensity of the fourth target audio segment; In the case where the first sound intensity is greater than the second sound intensity, determine the third target audio segment as the target audio segment, and in the case where the first sound intensity is less than the second sound intensity, determine the fourth target audio segment as the target audio segment; Combine the plurality of target audio segments to obtain the target audio data.
7. The role determination method according to any one of claims 1 to 5, characterized by, The target audio data is converted into target text information, and the target character information of the target audio data is added to the corresponding target text information, comprising: The target audio data is subjected to speech recognition by using an automatic speech recognition algorithm to obtain the target text information. The target character information of the target audio data is added to the corresponding target text information according to the target character information of the target audio data.
8. A role-determination device for a teaching scenario, characterized in that, Comprise: The detection unit is used for detecting the first audio data by voice activity detection to obtain a plurality of first audio segments, and detecting the second audio data by voice activity detection to obtain a plurality of second audio segments, the plurality of first audio segments and the plurality of second audio segments are arranged in order from small to large according to time, and the audio acquisition devices of the first audio data and the second audio data are different; The first determination unit is used for determining the character information of a plurality of first other audio segments and the character information of a first second audio segment based on the voiceprint features and character information of the first first audio segment, and determining the character information of a plurality of second other audio segments based on the voiceprint features and character information of the first second audio segment, the first other audio segment is the first audio segment except the first first audio segment, the second other audio segment is the second audio segment except the first second audio segment, and the timestamp information of the first second audio segment at least partially overlaps with the timestamp information of the first first audio segment; The second determination unit is used for determining the target character information of the target audio data based on the character information of the plurality of first audio segments and the character information of the plurality of second audio segments, the target audio data being obtained by sound quality evaluation on the plurality of first audio segments and the plurality of second audio segments; The execution unit is used for converting the target audio data into target text information, and adding the target character information of the target audio data to the corresponding target text information according to the target character information of the target audio data.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored program, wherein the program controls the device where the computer readable storage medium is located to execute the role determination method in the teaching scene according to any one of claims 1 to 7 when the program is running.
10. An electronic device, comprising: Comprise: One or more processors, memories, and one or more programs, wherein the one or more programs are stored in the memories and configured to be executed by the one or more processors, and the one or more programs comprise a program for executing the role determination method in the teaching scene according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and apparatus for training voiceprint recognition system
CN106297807A
Spokesman role determining method in multi-people session context, intelligent conference method and system
CN107993665A