A method, apparatus, device, and storage medium for speech segmentation
By monitoring VAD service failures in the speech recognition system and switching to a new service for voice clip detection and segmentation, the problem of muted clips in long speech affecting speech recognition efficiency is solved, and the fault tolerance of speech segmentation and the accuracy of recognition are improved.
Patent Information
- Application Number
- CN202210488588.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-05-06
AI Technical Summary
During the speech recognition process, mute clips often exist in long-term speeches, resulting in low efficiency of speech recognition. Long speeches need to be segmented to eliminate mute clips.
By continuously receiving the voice packet of the target voice and cache it locally, forwarding the voice packet to the target VAD service to monitor whether the VAD service has failed. If it fails, switch to the new VAD service for voice segment detection and segmentation.
The fault tolerance of voice segmentation is improved, ensuring that even if the target VAD service fails, the accuracy of voice recognition does not decrease, and the missed unsegmented voice data is reduced through breakpoint continuous transmission.
Smart Images

Figure CN114822513B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technologies, and in particular to the field of voice technologies. Background Art
[0002] As voice applications such as live broadcasts become more and more popular, the demand for speech recognition is increasing, especially the demand for speech recognition of long-duration speech. Among them, the above-mentioned long-duration speech can be referred to as long speech.
[0003] Since the duration of long speech is relatively long, there are often silent segments in long speech, and silent segments do not need to be recognized. Therefore, in order to improve the speech recognition efficiency, it is necessary to segment the long speech before speech recognition to obtain speech segments, and then remove the silent segments in the long speech, that is, remove the non-speech segments in the long speech. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and storage medium for speech segmentation.
[0005] According to one aspect of the present disclosure, a speech segmentation method is provided, including:
[0006] Continuously receive speech packets of the target speech, cache the received speech packets locally, and forward the received speech packets to the target voice activity detection (VAD) service;
[0007] Receive information of the speech segments detected by the target VAD service and the segmented speech segments, where the information of the speech segments includes: the start position and end position of the speech segments;
[0008] Monitor whether the target VAD service fails;
[0009] If it is monitored that the target VAD service fails, based on the information of the latest received speech segment, determine the speech data to be re-segmented in the speech packets cached locally;
[0010] Forward the speech data to be re-segmented and the received speech packets to a new VAD service, so that the new VAD service detects speech segments and segments speech segments for the received data.
[0011] According to another aspect of the present disclosure, a speech segmentation apparatus is provided, including:
[0012] A speech packet receiving module, configured to continuously receive speech packets of the target speech, cache the received speech packets locally, and forward the received speech packets to the target voice activity detection (VAD) service;
[0013] A voice information receiving module, configured to receive information of a voice segment detected by the target VAD service and the segmented voice segment, where the information of the voice segment includes: a start position and an end position of the voice segment;
[0014] A fault monitoring module, configured to monitor whether a fault occurs in the target VAD service;
[0015] A voice data determining module, configured to, if it is monitored that a fault occurs in the target VAD service, determine voice data to be re-segmented in a locally cached voice packet based on information of the latest received voice segment;
[0016] A voice data forwarding module, configured to forward the voice data to be re-segmented and the received voice packet to a new VAD service, so that the new VAD service performs voice segment detection on the received data and segments the voice segment.
[0017] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the above voice segmentation method.
[0021] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above voice segmentation method.
[0022] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program implements the above voice segmentation method when executed by a processor
[0023] As can be seen from the above, during the process of voice segmentation using the solution provided by the embodiments of the present disclosure, it is monitored whether the target VAD (Voice Activity Detect) service fails. Once the target VAD service fails, a new VAD service will be enabled to perform voice segment detection and voice segment segmentation. Also, in the solution provided by the embodiments of the present disclosure, not only continuously receives voice packets of the target voice, but also caches the received voice packets locally. In this way, when switching to the new VAD service for voice segment detection and segmentation, not only can the new VAD service perform voice segment detection and segmentation on the subsequently received voice packets, but also, based on the locally cached voice data, send the voice data that has been received but not detected and segmented when the target VAD service fails to the new VAD service for voice segment detection and segmentation. In summary, using the solution provided by the embodiments of the present disclosure for voice segmentation can improve the fault tolerance rate of voice segmentation.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0026] Figure 1 is a schematic structural diagram of a voice recognition system provided by an embodiment of the present disclosure;
[0027] Figure 2 is a waveform schematic diagram of a voice data provided by an embodiment of the present disclosure;
[0028] Figure 3 is a schematic flowchart of a voice segmentation method provided by an embodiment of the present disclosure;
[0029] Figure 4 is a signaling schematic diagram of a voice segmentation method provided by an embodiment of the present disclosure;
[0030] Figure 5 is a schematic flowchart of another voice segmentation method provided by an embodiment of the present disclosure;
[0031] Figure 6 is a signaling schematic diagram of another voice segmentation method provided by an embodiment of the present disclosure;
[0032] Figure 7 is a schematic diagram of a voice state transition provided by an embodiment of the present disclosure;
[0033] Figure 8It is a schematic structural diagram of a voice segmentation device provided by an embodiment of the present disclosure;
[0034] Figure 9 It is a schematic structural diagram of another voice segmentation device provided by an embodiment of the present disclosure;
[0035] Figure 10 It is a block diagram of an electronic device for implementing the voice segmentation method of an embodiment of the present disclosure. Detailed implementation manners
[0036] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0037] First, the application scenarios of the solutions provided by the embodiments of the present disclosure are described.
[0038] In one case, the solutions provided by the embodiments of the present disclosure can be applied to voice recognition scenarios.
[0039] According to the duration, voice can be divided into multiple voice frames. One or more voice frames form a voice packet, and the duration of each voice frame is generally equal, for example, 10 milliseconds, 15 milliseconds, etc. Each voice packet may contain a silent segment or a voice segment. When performing voice recognition, since the silent segment does not contain sound information, voice recognition is not required. Therefore, it is necessary to remove the silent segment from the voice packet and then perform voice recognition only on the voice segment to improve the efficiency of voice recognition.
[0040] In addition, the above voice recognition can be performed on the voice included in live broadcasts, short videos, etc. Live broadcast and short video applications often face a large number of users. Therefore, in one case, the system for implementing voice recognition can be implemented based on a microservice architecture.
[0041] In an embodiment of the present disclosure, refer to Figure 1 , a schematic structural diagram of a voice recognition system is provided. The system includes an access service, a control service, a VAD service, and a voice recognition engine. These services can run on the same server or on different servers, and the embodiments of the present disclosure do not limit this. In addition, multiple instances of the access service, the control service, the VAD service, and the voice recognition engine may exist in the voice recognition system.
[0042] Specifically, the above access service is used to receive each voice packet of the voice to be recognized and forward the received voice packet to the control service. It can be considered that the access service provides the API (Application Programming Interface) for voice recognition to users and can receive voice packets based on the WebSocket protocol.
[0043] The control service is used to forward the voice packet to the VAD service, receive the information and voice segment of the voice segment feedback by the VAD service, and forward the voice segment to the voice recognition engine. In addition, the control service also caches the received voice packet, monitors whether the VAD service fails, and performs fault tolerance processing when the VAD service fails to ensure the accurate progress of the voice recognition process for the voice to be recognized.
[0044] The VAD service is used to receive the voice packet sent by the control service, detect the voice segment in the voice packet, feedback the information of the detected voice segment to the control service, and segment the voice segment according to the information of the voice segment and feedback the voice segment to the control service.
[0045] The voice recognition engine is used to perform voice recognition on the voice segment sent by the control service.
[0046] The above briefly describes Figure 1 the functions of various services in the voice recognition system shown. Among them, the detailed functions of the control service will be described in detail in the following embodiments.
[0047] Next, the voice segment and the silent segment in the voice packet will be described.
[0048] The voice segment refers to the segment with sound, and the silent segment refers to the segment without sound.
[0049] Those skilled in the art can understand that the segment with sound in the voice packet has higher energy in the frequency domain, while the segment without sound has lower energy in the frequency domain, almost approaching zero. For example, referring to Figure 2 the schematic diagram of the frequency domain waveform corresponding to the voice packet shown, it can be seen from this figure that the voice segment has energy fluctuations, and the energy is high and not zero, while the silent segment, that is Figure 2 the non-voice segment shown in, has almost no energy fluctuations and the energy value is basically zero.
[0050] Based on this, in one implementation, the VAD service can determine the silent segment and the voice segment based on the energy of the voice packet in the frequency domain. When determining the voice segment, it is mainly to determine the start time and end time of the voice segment, and then take the segment from the start time to the adjacent end time as the voice segment. The segment in the voice packet except the voice segment is determined as the silent segment.
[0051] For example, if the VAD service detects that the energy in the frequency domain at the start time of the voice packet is not zero, it determines that the start time is the start time of the voice segment, that is, it detects the start of the voice segment.
[0052] If the VAD service detects that the energy in the frequency domain is not zero at the end time of the voice packet, since the voice packet has ended, it can be confirmed that the voice segment has also ended, that is, it determines that the end time of the voice segment is the end time of the voice packet.
[0053] In addition, the start time and end time of the voice segment can also be determined according to the change in the frequency domain energy at the previous and subsequent times in the voice packet. Specifically, if it is detected that the frequency domain energy changes from non-existent to existent at adjacent times, it determines that the time when the frequency domain energy appears is the start time of the voice segment; if it is detected that the voice energy at adjacent times changes from existent to non-existent, it determines that the time when the frequency domain energy becomes non-existent is the end time of the voice segment.
[0054] The execution subject of the solution provided in the embodiments of the present disclosure can be the above control service. Taking the execution subject as the control service as an example, the voice segmentation solution provided in the embodiments of the present disclosure will be described in detail through specific embodiments.
[0055] In an embodiment of the present disclosure, referring to Figure 3 and Figure 4 , Figure 3 shows a schematic flowchart of a voice segmentation method, Figure 4 shows a signaling schematic diagram of a voice segmentation method. The following will describe the above voice segmentation method in combination with Figure 3 and Figure 4 .
[0056] Specifically, the above voice segmentation method includes the following steps S301 - step S305.
[0057] Step S301: Continuously receive voice packets of the target voice, cache the received voice packets locally, and forward the received voice packets to the target VAD service.
[0058] The above target voice can be the voice to be subjected to voice recognition. From the perspective of duration, the target voice can be a long voice, such as a voice with a duration of 20 seconds, 10 minutes, or even 1 hour.
[0059] The target voice can be divided into multiple voice packets, and these voice packets can be continuously sent to the control service at preset time intervals. Therefore, the control service needs to continuously receive voice packets to ensure that all voice packets of the target voice are obtained. Each voice packet is sent to the control service at a preset time interval, which can create a time interval between the sending times of adjacent voice packets, enabling the control service to have sufficient time to forward the voice packet to the target VAD service and allowing the target VAD service to have sufficient time to perform voice segment detection and segmentation on the voice packet.
[0060] For example, the above time interval can be 160 milliseconds, 200 milliseconds, etc.
[0061] In addition, after receiving a voice packet, in addition to caching the voice packet locally, the control service also forwards the received voice packet to the target VAD service, which not only ensures that the target VAD service can receive the latest voice packet and ensures that the latest voice packet is promptly detected and segmented for voice segments by the target VAD service, but also ensures that the control service locally caches the latest received voice packet in a timely manner.
[0062] Step S302: Receive the information of the voice segments detected by the target VAD service and the segmented voice segments.
[0063] The information of the voice segment includes: the start position and the end position of the voice segment.
[0064] After receiving a voice packet, the target VAD service performs voice segment detection and segmentation on the voice packet. Specifically, when detecting a voice segment, the target VAD service first detects the start position of the voice segment in sequence, then detects the end position of the voice segment, and then segments the voice segment from the start position to the end position. From the above process, it can be seen that the acquisition of the start position, end position, and voice segment has a sequential order. Therefore, in one implementation, the target VAD service can respectively feedback the above start position, end position, and voice segment to the control service, that is, the target VAD service feeds back the start position to the control service immediately after detecting the start position without waiting for the end position and the voice segment, feeds back the end position to the control service immediately after detecting the end position without waiting for the voice segment, and feeds back the voice segment to the control service after segmenting the voice segment. This can enable the control service to promptly receive the start position, end position, and segmented voice segment detected by the target VAD service, so as to promptly grasp the progress of the target VAD service in performing voice segment detection and segmentation.
[0065] In one embodiment of the present disclosure, to facilitate distinguishing whether the target VAD service feedbacks the start position or the end position of the voice segment to the control service, the target VAD service may set different identifiers for the above start position and end position, and then feedback the specific values of the above identifiers and positions to the control service together.
[0066] For example, set the identifier of the start position as Vad Start and the identifier of the end position as Vad End. In this way, when the target VAD service feedbacks the start position to the control service, the feedback information includes: Vad Start and the specific value of the start position. When the target VAD service feedbacks the end position to the control service, the feedback information includes: Vad End and the specific value of the end position.
[0067] Specifically, there can be multiple representation methods for the above start position and end position. For example, in one implementation, it can be represented by the offset relative to the start position of the target voice. In another implementation, it can also be represented by the identifier of the voice packet to which it belongs and the offset relative to the start position of the voice packet to which it belongs.
[0068] Step S303: Monitor whether the target VAD service fails. If it is detected that the target VAD service fails, execute the following step S304.
[0069] In one implementation, the control service can monitor whether the target VAD service fails within the fault monitoring period. Among them, the above fault monitoring period can be set by the developer according to experience, and of course, it can also be set based on other information.
[0070] In addition, the control service can monitor whether the target VAD service fails by sending a fault detection packet to the target VAD service. The specific implementation method can be referred to in the subsequent embodiments and will not be elaborated here for the time being.
[0071] It should be noted that the step of the control service monitoring whether the target VAD service fails can continue throughout the entire working process of the target VAD service. Therefore, step S303 can be executed in parallel or serially with the above step S301 and step S302. The embodiments of the present disclosure do not limit this.
[0072] Step S304: Based on the information of the latest received voice segment, determine the voice data to be re - segmented in the locally cached voice packet.
[0073] Since the information of the voice segment includes the start position and the end position of the voice segment, therefore, the information of the latest received voice segment may be the start position of the voice segment or the end position of the voice segment.
[0074] In one implementation, when determining the voice data to be re-segmented, the position corresponding to the information of the latest received voice segment can be directly searched in the locally cached data packet, and the searched position is used as the start position of the voice data to be re-segmented. Then, the data from the start position to the end of the cached voice packet in the locally cached voice packet is determined as the voice data to be re-segmented.
[0075] For other implementations of determining the voice data to be re-segmented, reference can be made to the embodiments shown in the following Figure 5 and Figure 6 which will not be elaborated here for the time being.
[0076] Step S305: Forward the voice data to be re-segmented and the received voice packet to the new VAD service, so that the new VAD service performs voice segment detection and voice segment segmentation on the received data.
[0077] The above new VAD can be the VAD service selected by the control service from the VAD services in the idle state. Of course, it can also be the VAD service selected from the VAD services with a busyness degree less than the preset busyness degree.
[0078] In chronological order, the control service first receives the above voice data to be re-segmented, and then the subsequently received voice packet. Therefore, when the control service forwards the voice data to the target VAD service, it first forwards the voice data to be re-segmented, and then the subsequently received voice packet.
[0079] After receiving the data, the new VAD service can detect and segment the voice segments in the manner described in step S302 above, with the only difference being the difference in name between the new VAD and the target VAD.
[0080] In addition, after detecting the start position of the voice segment, the new VAD service will also feedback the above start position to the control service, and after detecting the end position of the voice segment, it will feedback the above end position to the control service, and will feedback the voice segment to the control service after segmenting the voice segment.
[0081] Furthermore, the control service will monitor whether the new VAD service fails in the same way as it monitors whether the target VAD service fails.
[0082] As can be seen from the above, during the process of voice segmentation using the solution provided by the embodiments of the present disclosure, whether the target VAD service fails will be monitored. Once the target VAD service fails, a new VAD service will be enabled to perform voice segment detection and voice segment segmentation. Also, in the solution provided by the embodiments of the present disclosure, not only are voice packets of the target voice continuously received, but the received voice packets are also cached locally. In this way, when switching to the new VAD service for voice segment detection and segmentation, not only can the new VAD service perform voice segment detection and segmentation on the subsequently received voice packets, but also, based on the locally cached voice data, the voice data that has been received but not detected and segmented when the target VAD service fails can be sent to the new VAD service for voice segment detection and segmentation. In summary, using the solution provided by the embodiments of the present disclosure for voice segmentation can improve the fault tolerance rate of voice segmentation.
[0083] In addition, this can enable the new VAD service to take over the work of the target VAD service in a "resume interrupted transfer" manner, reducing the omission of unsegmented data in the target voice. Thus, when performing voice recognition, it is ensured that even if the target VAD service fails, the accuracy of voice recognition will not decrease.
[0084] In one embodiment of the present disclosure, when locally caching the received voice packets in the above step S301, only the voice packets received most recently for a preset duration can be kept cached locally, where the preset duration is greater than or equal to a preset fault monitoring period.
[0085] After the control service receives a voice packet, it will cache the received voice packet locally. However, the more voice packets cached, the larger the cache occupied. Therefore, all received voice packets cannot be cached indefinitely. In view of the above situation, in the solution provided in this embodiment, the control service only caches the voice packets received most recently for a preset duration. In this way, when a new voice packet is received, the control service deletes the oldest existing voice data in the local cache and caches the above new voice packet, where the duration of the above existing voice data is equal to the duration of the new voice packet. In this way, as the control service continuously receives voice packets, the data with a long cache time in the local cache can be continuously updated by the newly received voice packets, but the duration of all cached data still remains the above preset duration.
[0086] In one case, the above fault monitoring period can be a duration set by a developer based on experience, or can be determined according to a specific fault monitoring method.
[0087] As can be seen from the above, when the control service caches the received voice packets, it only keeps the locally cached voice packets of the latest received preset duration. This not only ensures that the latest received voice packets are cached, but also takes into account the occupancy of the cache space by the voice packets, so that the cache space occupied by the cached voice packets will not be too large, balancing the occupancy of the cache resources. In addition, since the above preset duration is greater than the fault monitoring period, if the target VAD fails within the fault monitoring period, it can effectively ensure that the voice that has not been successfully segmented still exists in the cache.
[0088] The following details the method by which the control service monitors whether the target VAD service fails.
[0089] In an embodiment of the present disclosure, the control service may send a fault detection packet to the target VAD service according to a preset sending period. For example, the above fault detection packet may be a keepalive packet. In addition, the above fault detection packet may carry a timestamp of the sent fault detection packet. After receiving the fault detection packet, the target VAD service will immediately respond to the fault detection packet and feedback a response packet to the control service. After receiving the response packet of the above fault detection packet fed back by the target VAD service, the control service calculates the time difference between sending the above fault detection packet and receiving the response packet. If the number of consecutive occurrences of the time difference being greater than a preset duration threshold is greater than a preset number, it indicates that the time for the target VAD service to respond to the fault detection packet becomes longer and it is difficult to respond to the fault detection packet in a timely manner. At this time, it can be determined that the target VAD service has failed. Otherwise, it is determined that the target VAD service has not failed. In this way, it is possible to quickly and timely monitor whether the target VAD service has failed, and then switch from the target VAD service to a new VAD service in a timely manner, effectively ensuring that voice segmentation can proceed smoothly, thereby further improving the fault tolerance rate of voice segmentation.
[0090] For example, the above preset period may be 1 second, 2 seconds, etc. The above preset duration threshold may be 2 seconds, 3 seconds, etc. The above preset number may be 2 times, 3 times, etc.
[0091] To ensure that the control service can accurately calculate the above time difference after receiving the response packet, the response packet may contain the timestamp carried in the fault detection packet. In this way, when calculating the above time difference, obtain the time of receiving the response packet and the timestamp carried in the response packet, and take the difference to obtain the above time difference.
[0092] Based on the above method for monitoring whether the target VAD fails, in an embodiment of the present disclosure, the above fault monitoring period may be a duration determined according to the above preset duration threshold and preset number.
[0093] For example, the product of the above-mentioned preset duration threshold and the preset number of times can be calculated as the above-mentioned fault monitoring period. In this case, the above-mentioned preset duration can be set to a value greater than the fault monitoring period, that is, a value greater than the above-mentioned product. For example, when the preset duration threshold is 2 seconds and the preset number of times is 2 times, the above-mentioned product is 4 seconds, then the above-mentioned fault monitoring period is 4 seconds. In this case, the above-mentioned preset duration can be 10 seconds, 12 seconds, etc.
[0094] In another embodiment of the present disclosure, when setting the above-mentioned fault monitoring period, in addition to considering the above-mentioned preset duration threshold and the preset number of times, the duration of the voice packet, especially the maximum duration of the voice packet, can also be considered.
[0095] For example, after calculating the product of the above-mentioned preset duration threshold and the preset number of times, the sum of the above-mentioned product and the maximum duration of the voice packet can also be calculated, and the above-mentioned sum value can be used as the fault monitoring period. In this case, the above-mentioned preset duration can be set to a value greater than the above-mentioned fault monitoring period, that is, a value greater than the above-mentioned sum value. For example, on the basis of the above-mentioned example, if the maximum duration of the voice packet is 3 seconds and the above-mentioned sum value is 4 seconds + 3 seconds = 7 seconds, in this case, the above-mentioned preset duration can be a value greater than 7 seconds, for example, the above-mentioned preset duration can be 10 seconds, 12 seconds, etc.
[0096] As can be seen from the above, setting the above-mentioned fault monitoring period according to the above-mentioned preset duration threshold and the preset number of times, and then setting the preset duration for caching the voice packet can not only ensure that the control service caches enough voice packets locally and can locate the voice data that has not been successfully segmented from the local cache of the control service when the target VAD service fails, but also effectively prevent too many voice packets from being cached locally and reduce the occupation of the cache.
[0097] It can be known from the previous description that the voice data to be re-segmented can be determined in different ways. On this basis, another voice segmentation method is provided in the embodiments of the present disclosure.
[0098] In one embodiment of the present disclosure, see Figure 5 and Figure 6 , Figure 5 A flowchart of another voice segmentation method is provided, Figure 6 A signaling diagram of another voice segmentation method is provided. The above-mentioned voice segmentation method will be described below with reference to Figure 5 and Figure 6 .
[0099] Specifically, the above-mentioned voice segmentation method includes the following steps S501-S507.
[0100] Step S501: Continuously receive the voice packets of the target voice, keep the locally cached voice packets of the preset duration that are newly received, and forward the received voice packets to the target VAD service.
[0101] It should be noted that this step S501 is the same as the aforementioned step S301 and the steps regarding caching voice packets, and will not be elaborated here.
[0102] Step S502: Receive the frame number offset of the voice packet feedback by the target VAD service.
[0103] After the control service sends the voice packet to the target VAD service, the target VAD service receives the voice packet, can obtain the frame number offset of the voice packet from the voice packet, and feedback the above frame number offset to the control service. On the one hand, it informs the control service that the voice packet has been received, and on the other hand, it informs the control service which voice packet will be subjected to voice segment detection and segmentation.
[0104] In one case, the above frame number offset can be the offset of the first frame in the voice packet relative to the target voice.
[0105] In one implementation, the target VAD service can feedback the above frame number offset to the control service by sending an ACK (Acknowledge character) message to the control service. In this case, the above frame number offset is carried in the ACK message.
[0106] Step S503: Receive the information of the voice segment detected by the target VAD service and the segmented voice segment.
[0107] Step S504: Monitor whether the target VAD service fails. If it is detected that the target VAD service fails, execute the following step S505 or S506.
[0108] It should be noted that the above step S503 and step S504 are the same as the aforementioned step S302 and step S303 respectively, and will not be elaborated here.
[0109] Step S505: If the information of the newly received voice segment is the starting position of the voice segment, determine the data starting from the above starting position in the voice packets cached locally as the voice data to be re-segmented.
[0110] Since the voice packet may contain two types of segments, namely silent segments and voice segments, it can be considered that the voice packet may be in two states, the speech state and the non-speech state, and these two states can be converted. Specifically, reference can be made to Figure 7How to convert between the two states depends on the information of the speech segment detected by the target VAD service. When currently in the speech state, if the target VAD service detects the end position of the speech segment, that is, Vad End, then it switches from the speech state to the non-speech state; when currently in the non-speech state, if the target VAD service detects the start position of the speech segment, that is, Vad Start, then it switches from the non-speech state to the speech state.
[0111] Based on the above, if the information of the latest received speech segment is the start position of the speech segment, it can be considered that it is in the speech state at this time, and the target VAD service has detected the speech segment, that is, it is detecting and segmenting the speech segment. Since the control service locally caches speech packets for a preset duration, and the preset duration is greater than or equal to the fault monitoring period, therefore, after detecting that the target VAD service has failed within the fault monitoring period, the speech packet that the target VAD service is currently processing is still cached in the local cache of the control service. Thus, in this case, the data corresponding to the above start position can be determined from the speech packets cached locally, and then the data from the determined data to the end of the data cached locally is determined as the speech data to be re-segmented.
[0112] Step S506: If the information of the latest received speech segment is the end position of the speech segment, then based on the latest received frame number offset, determine the speech data to be re-segmented in the speech packets cached locally.
[0113] Based on the description in the foregoing step S505, it can be known that if the information of the latest received speech segment is the end position of the speech segment, it can be considered that it is in the non-speech state at this time, and what the target VAD service detects is not a speech segment but a silent segment. However, the reason why the segment detected by the target VAD service is a silent segment may, on the one hand, be really because the current speech packet is a silent segment, and on the other hand, it may be because the target VAD service has failed some time ago.
[0114] In view of the above situation, in an embodiment of the present disclosure, if there is a target position in the speech packets cached locally, it means that there is a speech packet that the target VAD service is currently processing in the data cached locally by the control service. At this time, the data from the target position in the speech packets cached locally is determined as the speech data to be re-segmented. Among them, the target position is: the position corresponding to the latest received frame number offset.
[0115] If the target location does not exist in the locally cached speech packets, it indicates that the target VAD has experienced a fault for some time. However, since the control service has been performing fault monitoring according to the fault monitoring period, it can be ensured that the speech packets that no longer exist in the local cache of the control service have been successfully segmented by the target VAD service. To prevent omission of the unsegmented data in the local cache of the control service, in this case, all the speech packets cached locally in this embodiment are determined as the speech data to be re-segmented.
[0116] It can be seen that in the solution provided in this embodiment, when the information of the latest received speech segment is the termination position of the speech segment, in combination with the target position corresponding to the frame number offset of the latest received frame, the speech data to be re-segmented is determined in different cases, which is beneficial to improving the accuracy of the determined speech data to be re-segmented and reducing the situation of omitted segmentation in the data cached locally by the control service.
[0117] Step S507: Forward the speech data to be re-segmented and the received speech packets to a new VAD service, so that the new VAD service performs speech segment detection on the received data and segments the speech segments.
[0118] It should be noted that this step S507 is the same as the foregoing step S305 and will not be elaborated here.
[0119] As can be seen from the above, in the solution provided in this embodiment, according to whether the information of the latest received speech segment is the start position or the end position of the speech segment, different methods are respectively used to determine the speech data to be re-segmented, which effectively takes into account the characteristics of the data cached locally by the control service in different cases, making the determined data to be re-segmented more accurate, thereby further effectively improving the fault tolerance rate of speech segmentation.
[0120] Corresponding to the above speech segmentation method, an embodiment of the present disclosure also provides a speech segmentation device.
[0121] In an embodiment of the present disclosure, referring to Figure 8 , a schematic structural diagram of a speech segmentation device is provided, including:
[0122] A speech packet receiving module 801, configured to continuously receive speech packets of a target speech, cache the received speech packets locally, and forward the received speech packets to a target voice activity detection VAD service;
[0123] A speech information receiving module 802, configured to receive the information of the speech segments detected by the target VAD service and the segmented speech segments, where the information of the speech segments includes: the start position and the end position of the speech segment;
[0124] A fault monitoring module 803, configured to monitor whether a fault occurs in the target VAD service;
[0125] A voice data determination module 804, configured to, if it is monitored that a fault occurs in the target VAD service, determine voice data to be re-segmented in the locally cached voice packets based on the information of the latest received voice segment;
[0126] A voice data forwarding module 805, configured to forward the voice data to be re-segmented and the received voice packets to a new VAD service, so that the new VAD service performs voice segment detection and voice segment segmentation on the received data.
[0127] As can be seen from the above, in the process of performing voice segmentation by applying the solution provided in the embodiments of the present disclosure, it is monitored whether a fault occurs in the target VAD service. Once a fault occurs in the target VAD service, a new VAD service will be enabled to perform voice segment detection and voice segment segmentation. Also, since in the solution provided in the embodiments of the present disclosure, not only the voice packets of the target voice are continuously received, but also the received voice packets are cached locally. In this way, when switching to the new VAD service for voice segment detection and segmentation, not only can the new VAD service perform voice segment detection and segmentation on the subsequently received voice packets, but also based on the locally cached voice data, the voice data that has been received but not detected and segmented when the target VAD service fails can be sent to the new VAD service for voice segment detection and segmentation. In summary, applying the solution provided in the embodiments of the present disclosure for voice segmentation can improve the fault tolerance rate of voice segmentation.
[0128] In addition, this can enable the new VAD service to take over the work of the target VAD service in a "resume interrupted transfer" manner, reducing the omission of unsegmented data in the target voice. Thus, when performing voice recognition, it is ensured that even if the target VAD service fails, the accuracy of voice recognition will not decrease.
[0129] In an embodiment of the present disclosure, the voice packet receiving module 801 is specifically configured to continuously receive voice packets of a target voice, keep the locally cached voice packets of the latest received preset duration, and forward the received voice packets to a target voice activity detection (VAD) service, where the preset duration is greater than or equal to a preset fault monitoring period.
[0130] As can be seen from the above, when the control service caches the received voice packets, it only keeps the voice packets received most recently within a preset duration in the local cache. This not only ensures that the most recently received voice packets are cached, but also takes into account the occupancy of the cache space by the voice packets, so that the cache space occupied by the cached voice packets will not be too large, thus balancing the occupancy of the cache resources. In addition, since the above-mentioned preset duration is greater than the fault monitoring period, if the target VAD fails within the fault monitoring period, it can effectively ensure that the voice that has not been successfully segmented still exists in the cache.
[0131] In one embodiment of the present disclosure, referring to Figure 9 , a structural schematic diagram of another voice segmentation device is provided. The device includes:
[0132] A voice packet receiving module 901, specifically configured to continuously receive voice packets of a target voice, keep the voice packets received most recently within a preset duration in the local cache, and forward the received voice packets to a target voice activity detection (VAD) service, where the preset duration is greater than or equal to a preset fault monitoring period.
[0133] A frame number offset receiving module 902, configured to receive the frame number offset of the voice packet fed back by the target VAD service after the voice packet receiving module forwards the received voice packet to the target VAD service;
[0134] A voice information receiving module 903, configured to receive the information of the voice segment detected by the target VAD service and the segmented voice segment, where the information of the voice segment includes: the start position and the end position of the voice segment;
[0135] A fault monitoring module 904, configured to monitor whether the target VAD service fails;
[0136] A first voice data determination unit 905, configured to, when it is monitored that the target VAD service fails, if the information of the most recently received voice segment is the start position of the voice segment, determine the data starting from the start position in the voice packets cached locally as the voice data to be re-segmented;
[0137] A second voice data determination unit 906, configured to, when it is monitored that the target VAD service fails, if the information of the most recently received voice segment is the end position of the voice segment, determine the voice data to be re-segmented in the voice packets cached locally based on the most recently received frame number offset.
[0138] A voice data forwarding module 907, configured to forward the voice data to be re-segmented and the received voice packets to a new VAD service, so that the new VAD service performs voice segment detection on the received data and segments the voice segment.
[0139] As can be seen from the above, in the solution provided by this embodiment, according to the information of the latest received speech segment as the start position and end position of the speech segment, different methods are respectively used to determine the speech data to be re-segmented, which effectively takes into account the characteristics of the data cached locally by the control service in different situations, making the determined data to be re-segmented more accurate, thereby further effectively improving the error tolerance rate of speech segmentation.
[0140] In one embodiment of the present disclosure, the second speech data determination unit 906 is specifically configured to, in the case of detecting a failure of the target VAD service, if there is a target position in the locally cached speech packets, determine the data starting from the target position in the locally cached speech packets as the speech data to be re-segmented, where the target position is: the position corresponding to the frame number offset of the latest received; if the target position does not exist in the locally cached speech packets, determine all the locally cached speech packets as the speech data to be re-segmented.
[0141] It can be seen that in the solution provided by this embodiment, when the information of the latest received speech segment is the end position of the speech segment, in combination with the target position corresponding to the frame number offset of the latest received, the speech data to be re-segmented is determined in different cases, which is beneficial to improving the accuracy of the determined speech data to be re-segmented and reducing the situation of missed segmentation in the data cached locally by the control service.
[0142] In one embodiment of the present disclosure, the fault monitoring module is specifically configured to send a fault detection packet to the target VAD service according to a preset sending period; receive the response packet of the fault detection packet fed back by the target VAD service; calculate the time difference between sending the fault detection packet and receiving the response packet; if the number of consecutive occurrences of the time difference being greater than a preset duration threshold is greater than a preset number, determine that the target VAD service has failed.
[0143] In this way, it is possible to quickly and timely monitor whether the target VAD service has failed, and then timely switch from the target VAD service to a new VAD service, effectively ensuring that speech segmentation can proceed smoothly, thereby further improving the error tolerance rate of speech segmentation.
[0144] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0145] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0146] In one embodiment of the present disclosure, an electronic device is provided, including:
[0147] at least one processor; and
[0148] a memory communicatively connected to the at least one processor; wherein
[0149] the memory stores instructions executable by the at least one processor, and when executed by the at least one processor, the instructions enable the at least one processor to perform the voice segmentation method described in the above method embodiments.
[0150] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the voice segmentation method described in the above method embodiments.
[0151] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the voice segmentation method described in the above method embodiments is implemented.
[0152] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0153] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0154] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as a keyboard, mouse, etc.; output unit 1007, such as various types of displays, speakers, etc.; storage unit 1008, such as a disk, optical disc, etc.; and communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0155] Computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 1001 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1001 executes the various methods and processes described above, such as the speech segmentation method. For example, in some embodiments, the speech segmentation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the speech segmentation method described above can be executed. Alternatively, in other embodiments, computing unit 1001 can be configured to execute the speech segmentation method in any other suitable manner (e.g., by means of firmware).
[0156] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0157] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0158] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0160] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0161] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server combined with a blockchain.
[0162] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0163] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A voice segmentation method, comprising: Continuously receiving voice packets of the target voice, caching the received voice packets locally, and forwarding the received voice packets to the target voice activity detection (VAD) service; Receiving information of voice segments detected by the target VAD service and the segmented voice segments, wherein the information of the voice segments includes: the start position and the end position of the voice segment; Monitoring whether the target VAD service fails; If it is monitored that the target VAD service fails, determining the voice data to be re-segmented in the voice packets cached locally based on the information of the latest received voice segment; Forwarding the voice data to be re-segmented and the received voice packets to a new VAD service, so that the new VAD service performs voice segment detection and segments voice segments on the received data; The caching the received voice packets locally includes: Keeping the latest received voice packets of a preset duration cached locally, wherein the preset duration is greater than or equal to a preset fault monitoring period.
2. The method according to claim 1, wherein After forwarding the received voice packets to the target VAD service, further comprising: Receiving the frame number offset of the voice packets fed back by the target VAD service; The determining the voice data to be re-segmented in the voice packets cached locally based on the information of the latest received voice segment includes: If the information of the latest received voice segment is the start position of the voice segment, determining the data starting from the start position in the voice packets cached locally as the voice data to be re-segmented; If the information of the latest received voice segment is the end position of the voice segment, determining the voice data to be re-segmented in the voice packets cached locally based on the latest received frame number offset.
3. The method according to claim 2, wherein The determining the voice data to be re-segmented in the voice packets cached locally based on the latest received frame number offset includes: If there is a target position in the voice packets cached locally, determining the data starting from the target position in the voice packets cached locally as the voice data to be re-segmented, wherein the target position is: the position corresponding to the latest received frame number offset; If the target position does not exist in the voice packets cached locally, determining all the voice packets cached locally as the voice data to be re-segmented.
4. The method according to any one of claims 1-3, wherein, The monitoring whether the target VAD service fails includes: Sending a fault detection packet to the target VAD service according to a preset sending period; Receiving a response packet of the fault detection packet fed back by the target VAD service; Calculating the time difference between sending the fault detection packet and receiving the response packet; If the number of consecutive occurrences of the time difference being greater than a preset duration threshold is greater than a preset number, determining that the target VAD service fails.
5. A voice segmentation device, comprising: A voice packet receiving module, configured to continuously receive voice packets of the target voice, cache the received voice packets locally, and forward the received voice packets to the target voice activity detection (VAD) service; A voice information receiving module, configured to receive information of a voice segment detected by the target VAD service and the segmented voice segment, where the information of the voice segment includes: a start position and an end position of the voice segment; A fault monitoring module, configured to monitor whether a fault occurs in the target VAD service; A voice data determining module, configured to, if it is monitored that a fault occurs in the target VAD service, determine voice data to be re-segmented in a locally cached voice packet based on information of the latest received voice segment; A voice data forwarding module, configured to forward the voice data to be re-segmented and the received voice packet to a new VAD service, so that the new VAD service performs voice segment detection on the received data and segments the voice segment; The voice packet receiving module is specifically configured to continuously receive voice packets of a target voice, keep a locally cached voice packet of a preset duration of the latest received voice packet, and forward the received voice packet to a target voice activity detection (VAD) service, where the preset duration is greater than or equal to a preset fault monitoring period.
6. The apparatus according to claim 5, wherein The apparatus further includes: A frame number offset receiving module, configured to receive a frame number offset of the voice packet fed back by the target VAD service after the voice packet receiving module forwards the received voice packet to the target VAD service; The determination of the voice data includes: A first voice data determining unit, configured to, when it is monitored that a fault occurs in the target VAD service, if the information of the latest received voice segment is the start position of the voice segment, determine the data starting from the start position in the locally cached voice packet as the voice data to be re-segmented; A second voice data determining unit, configured to, when it is monitored that a fault occurs in the target VAD service, if the information of the latest received voice segment is the end position of the voice segment, determine the voice data to be re-segmented in the locally cached voice packet based on the latest received frame number offset.
7. The apparatus according to claim 6, wherein The second voice data determining unit is specifically configured to, when it is monitored that a fault occurs in the target VAD service, if there is a target position in the locally cached voice packet, determine the data starting from the target position in the locally cached voice packet as the voice data to be re-segmented, where the target position is: the position corresponding to the latest received frame number offset; if the target position does not exist in the locally cached voice packet, determine all the locally cached voice packets as the voice data to be re-segmented.
8. The apparatus according to any one of claims 5-7, wherein The fault monitoring module is specifically configured to send a fault detection packet to the target VAD service according to a preset sending period; receive a response packet of the fault detection packet fed back by the target VAD service; calculate a time difference between sending the fault detection packet and receiving the response packet; If the number of consecutive occurrences of the time difference being greater than a preset duration threshold is greater than a preset number, determine that a fault occurs in the target VAD service.
9. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1-4.
11. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-4.
Citation Information
Patent Citations
Voice wake-up optimization method, device and system, storage medium and electronic equipment
CN112687298A