A speech-to-text method and related apparatus, electronic device

By dynamically adjusting the priority and response requirements of online and offline recognition services, the problems of low recognition efficiency and poor real-time performance of speech-to-text under unstable network conditions are solved, achieving efficient and accurate speech-to-text conversion under unstable network conditions.

CN116052679BActive Publication Date: 2026-05-19IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-12-19
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing speech-to-text technology suffers from low recognition efficiency and poor real-time performance when the network is unstable, failing to meet the requirement of combining online and offline recognition methods.

Method used

By acquiring target audio data, and based on the network status information and service selection strategy of the online recognition service, the priority and response requirements of the recognition service are dynamically adjusted, and a suitable recognition service is selected for speech-to-text conversion, including multi-level degradation and upgrade strategies for online and offline recognition services.

Benefits of technology

It improves the efficiency and accuracy of speech recognition, ensuring that accurate text recognition results can be provided quickly even when the network is unstable, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052679B_ABST
    Figure CN116052679B_ABST
Patent Text Reader

Abstract

The application discloses a speech-to-text method and related device and electronic equipment, and the method comprises the following steps: obtaining target audio data; based on network state information of an online recognition service and a service selection strategy, at least one target recognition service is enabled from a plurality of recognition services, and the plurality of recognition services comprise at least one online recognition service; the target audio data is recognized by using the target recognition service, and a text recognition result corresponding to the target audio data is obtained; and based on at least one adjustment reference factor, the service selection strategy is adjusted, and the adjusted service selection strategy is used for selecting the target recognition service of the next target audio data, and the adjustment reference factor comprises at least one of the network state information of the online recognition service and operation information of a user in an audio recognition process. The above scheme can improve the efficiency and accuracy of audio recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition, and in particular to a speech-to-text method and related apparatus and electronic equipment. Background Technology

[0002] Currently, most speech-to-text conversion solutions on the market utilize online network recognition. However, these methods cannot select the appropriate recognition method based on the network status information of the online network, resulting in low speech recognition efficiency. Another solution on the market is to use offline recognition methods, but their recognition performance is inferior to online network recognition. To improve recognition efficiency, another approach is to first use offline recognition methods to recognize the speech data, and then, if the network status is good later, use online network recognition methods to re-recognize the speech data. This reactive approach cannot meet the requirements of timeliness and is unusable in real-time scenarios. Summary of the Invention

[0003] This application provides at least one speech-to-text method and related apparatus and electronic device, which can improve the efficiency and accuracy of audio recognition.

[0004] The first aspect of this application provides a speech-to-text method, comprising: acquiring target audio data; enabling at least one target recognition service from a plurality of recognition services based on network status information of an online recognition service and a service selection strategy, wherein the plurality of recognition services includes at least one online recognition service; recognizing the target audio data using the target recognition service to obtain a text recognition result corresponding to the target audio data; and adjusting the service selection strategy based on at least one adjustment reference factor, wherein the adjusted service selection strategy is used to select a target recognition service for the next target audio data, and the adjustment reference factor includes at least one of network status information of the online recognition service and user operation information during the audio recognition process.

[0005] Among them, several identification services have different priorities, and the service selection strategy is as follows: each identification service is activated sequentially according to priority as the target identification service, and the activation trigger condition for identification services other than the highest priority is that the network status information of the reference identification service after activation meets the preset response requirements of the reference identification service; the reference identification service is an online identification service with a higher priority than the identification service; the service selection strategy is adjusted based on at least one adjustment reference factor, including at least one of the following steps: adjusting the priority of several identification services based on at least one adjustment reference factor; adjusting the network response requirements corresponding to the identification service based on at least one adjustment reference factor.

[0006] The network status information of the reference recognition service after it is enabled includes at least one network status factor. The preset response requirements corresponding to the reference recognition service include: the network status factor of the reference recognition service exceeds the reference threshold of the reference recognition service regarding the network status factor; the network status factor includes at least one of the following: the connection establishment time of the reference recognition service, and the time interval between receiving two adjacent text recognition results obtained by the reference recognition service.

[0007] The reference identification service for the online identification service is another online identification service with a higher priority than the online identification service; and / or, several identification services also include offline identification services, with the offline identification service having the lowest priority. The reference identification service for the offline identification service is the overall online identification service. The network state factors of the overall online identification service include at least one of the following: the connection establishment time of each online identification service and the time interval between receiving two adjacent online text identification results. The online text identification results are obtained by any enabled online identification service. The preset response requirements corresponding to the overall online identification service include at least one of the following: the connection establishment time of each online identification service exceeds the first reference threshold of the overall online identification service regarding the connection establishment time, and the time interval between receiving two adjacent online text identification results exceeds the second reference threshold of the overall online identification service regarding the time interval.

[0008] Among them, several identification services also include offline identification services, which have the lowest priority. The reference identification service for offline identification services is the overall online identification service. Based on at least one adjustment reference factor, the network response requirements corresponding to the identification service are adjusted, including: in response to the network state factor of the overall online identification service exceeding the reference threshold of the overall online identification service regarding the network state factor, the reference threshold of the overall online identification service regarding the network state factor is lowered.

[0009] The network state factors include at least one of the following: connection establishment time and the time interval between receiving two adjacent text recognition results, wherein the text recognition result recognized by the online recognition service is the online text recognition result; based on at least one adjustment reference factor, the network response requirements corresponding to the recognition service are adjusted, including at least one of the following steps: in response to the user canceling the target audio data recognition service during the initial text waiting period, obtaining the first time interval between the start point of the current speech and the cancellation time, reducing the first reference threshold of each recognition service regarding the connection establishment time to less than the first time interval, wherein the initial text waiting period is the time interval between the start point of the current speech and the first receipt of the online text recognition result; in response to the user canceling the target audio data recognition service during the initial text waiting period, obtaining the first ... text waiting period and the first receipt of the online text recognition result; in response to the user canceling the target audio data recognition service during the initial text waiting period, obtaining the first time interval between the start point of the current text waiting period and the first receipt of the online text recognition result; in response to the user canceling the target audio data recognition service during the initial text waiting period, obtaining the first time interval between the start point of the current text waiting period If a user cancels the target audio data recognition service during a preset text reception period, the second reference threshold of each recognition service regarding the reception time interval is adjusted to be less than the preset text reception period, which is the second time interval after the most recent online text recognition result is received. The third time interval between the start point of the current speech and the first online text recognition result is obtained, and the first reference threshold of each online recognition service is adjusted to be no greater than the initial reference threshold and the third time interval. The reception time interval between every two adjacent online text recognition results is counted, and the second reference threshold of each online recognition service is adjusted to be between the largest received time interval and the initial reference threshold.

[0010] Specifically, based on at least one adjustment reference factor, the priorities of several recognition services are adjusted, including: in response to the network status information of the online recognition service not meeting the network response requirements corresponding to the online recognition service, the selection priority of the online recognition service is downgraded; in response to the number of times the text recognition result corresponding to the online recognition service is used by the user reaching the preset number of times corresponding to the online recognition service, the selection priority of the online recognition service is upgraded.

[0011] The target audio data is an audio segment in the current speech; the execution device of the method sequentially takes each audio segment in the current speech as target audio data according to the order of audio acquisition and executes the method based on the target audio data to obtain the text recognition result corresponding to each audio segment; obtaining the first target audio data of the current speech includes: in response to detecting the start endpoint of the current speech, obtaining the audio segment located in the first time range of the start endpoint as the first target audio data; obtaining the last target audio data of the current speech includes: in response to detecting the end endpoint of the current speech, obtaining the audio segment located in the second time range of the end endpoint as the last target audio data.

[0012] A second aspect of this application provides a speech-to-text device, comprising: an acquisition module for acquiring target audio data; a selection module for enabling at least one target recognition service from a plurality of recognition services based on network status information of an online recognition service and a service selection strategy, wherein the plurality of recognition services includes at least one online recognition service; a recognition module for recognizing the target audio data using the target recognition service to obtain a text recognition result corresponding to the target audio data; and an adjustment module for adjusting the service selection strategy based on at least one adjustment reference factor, wherein the adjusted service selection strategy is used to select a target recognition service for the next target audio data, and the adjustment reference factor includes at least one of network status information of the online recognition service and user operation information during the audio recognition process.

[0013] A third aspect of this application provides an electronic device including a memory and a processor coupled to each other, the processor being configured to execute program instructions stored in the memory to implement the speech-to-text method described in the first aspect above.

[0014] The above scheme, when recognizing target audio data, activates at least one target recognition service from several recognition services to recognize the target audio data based on the network status information and service selection strategy of the online recognition service, obtaining the text recognition result corresponding to the target audio data. Then, it adjusts the service selection strategy based on at least one of the network status information of the online recognition service and the user's operation information during the audio recognition process, so as to select a suitable target recognition service for subsequent target audio data recognition. This enables dynamic adjustment of the service selection strategy for audio recognition, and allows the dynamically adjusted service selection strategy to select a target recognition service that is more in line with the network conditions of the online recognition service or the user's intention, thereby helping to improve the efficiency and accuracy of audio recognition.

[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0017] Figure 1 This is a flowchart illustrating an embodiment of the speech-to-text method of this application;

[0018] Figure 2 This is a flowchart illustrating another embodiment of the speech-to-text method of this application;

[0019] Figure 3This is a schematic diagram of the framework of an embodiment of the speech-to-text device of this application;

[0020] Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of this application;

[0021] Figure 5 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0022] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0023] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0024] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0025] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the speech-to-text method of this application. Specifically, it may include the following steps:

[0026] Step S110: Obtain the target audio data.

[0027] The speech-to-text method presented in this paper can be applied to transcription scenarios such as voice recorders and translation machines. The target audio data in this paper is speech data recorded while a person is speaking; this can be real-time recorded speech data or stored speech data. The recorded speech data can be a partial speech segment or a complete speech segment.

[0028] In some embodiments, the target audio data is an audio segment in the current speech; the execution device of this speech-to-text method sequentially takes each audio segment in the current speech as the target audio data according to the order of audio acquisition and executes the method based on the target audio data to obtain the text recognition result corresponding to each audio segment.

[0029] In some specific embodiments, when a user initiates a speech-to-text request, an endpoint detection method is used to obtain the first target audio data of the current speech. Based on the detected start endpoint of the current speech, an audio segment within a first time range of the start endpoint is obtained as the first target audio data. Specifically, the first target audio data can be obtained starting from a first preset moment before the start endpoint, with an audio segment of a preset duration. The first time range includes both the first preset moment and the preset duration. For example, the first preset moment can be set to 1 second, and the preset duration to 10 seconds; then, the speech data from 1 second before the start endpoint to 9 seconds after the start endpoint would be the first target audio data. The first preset moment and preset duration can be dynamically adjusted and are not specifically limited here.

[0030] In other specific embodiments, the user can also use an endpoint detection method to obtain the last target audio data of the current speech. Based on the detected end endpoint of the current speech, the audio segment located within a second time range of the end endpoint is obtained as the last target audio data. Specifically, the audio segment of a preset time length after the end endpoint is obtained as the last target audio data. Between the start and end endpoints of the current speech, the execution device can extract audio segments of any length or any time period as target audio data. It is understood that, in addition to the endpoint detection method, other methods can be used to monitor the start and end endpoints of the current speech, and no specific limitation is made here.

[0031] Step S120: Based on the network status information of the online identification service and the service selection strategy, enable at least one target identification service from a number of identification services, including at least one online identification service.

[0032] The identification services described in this document may include only multiple online identification services, or they may include both online and offline identification services. Different priorities can be set for the multiple online and offline identification services. The service selection strategy is to activate each identification service sequentially according to its priority as the target identification service. The activation trigger condition for identification services other than the highest priority is that the network status information of the reference identification service after activation meets the preset response requirements of the reference identification service; the reference identification service is an online identification service with a higher priority than the target identification service.

[0033] The network status information of the reference identification service after activation includes at least one network status factor. The preset response requirements for the reference identification service include: the network status factor of the reference identification service exceeding a reference threshold for the network status factor. The network status factor includes at least one of the following: the connection establishment time of the reference identification service, and the time interval between receiving two adjacent text recognition results obtained by the reference identification service. The reason why the network status factor of the reference identification service exceeds the reference threshold for the network status factor may be due to poor network conditions or equipment malfunction, etc.

[0034] In some embodiments, the reference recognition service for the online recognition service is another online recognition service with a higher priority than the current online recognition service. For example, several recognition services include a first online recognition service, a second online recognition service, and a third online recognition service. The first online recognition service has a higher priority than the second online recognition service, and the second online recognition service has a higher priority than the third online recognition service. The first online recognition service serves as the reference recognition service for the online recognition service. If the first online recognition service cannot be enabled, then the second online recognition service becomes the reference recognition service. Specifically, the first online recognition service is enabled sequentially and used as the target recognition service to recognize the target audio data. If the time interval between two adjacent text recognition results obtained after the first online recognition service is enabled exceeds its reference threshold, the second online recognition service is enabled as the target recognition service to recognize the target audio data again. At this time, the first online recognition service is still recognizing the target audio data. The recognition processes of the first and second online recognition services for the same target audio data do not interfere with each other. It is understandable that the preset response requirements for the reference identification service can be either that only one network state factor exceeds its reference threshold, or that all network state factors exceed their reference thresholds, depending on the specific circumstances, and no specific limitation is made here.

[0035] In other embodiments, the recognition services include a first online recognition service and a second online recognition service, with the first online recognition service having a higher priority than the second online recognition service. Before enabling the first online recognition service to recognize the target audio data, the executing device attempts to establish a connection with the first online recognition service. If the actual connection establishment time exceeds its reference threshold, the second online recognition service is then enabled. Alternatively, after the executing device successfully connects to the first online recognition service, if the time interval between two consecutive text recognition results returned by the first online recognition service exceeds its reference threshold, the second online recognition service is then enabled.

[0036] In some embodiments, the identification services may further include offline identification services, which have the lowest priority. The reference identification service for the offline identification service is the overall online identification service. The network state factors of the overall online identification service include at least one of the following: the connection establishment time of each online identification service and the time interval between receiving two adjacent online text recognition results. The online text recognition results are obtained by any enabled online identification service. The preset response requirements corresponding to the overall online identification service include at least one of the following: the connection establishment time of each online identification service exceeds a first reference threshold for the overall online identification service regarding the connection establishment time, and the time interval between receiving two adjacent online text recognition results exceeds a second reference threshold for the overall online identification service regarding the time interval.

[0037] Specifically, the identification services include a first online identification service, a second online identification service, and an offline identification service. The first online identification service has a higher priority than the second online identification service, and the offline identification service has the lowest priority. The execution device first establishes a connection with the first online identification service. If the connection establishment time exceeds a first reference threshold for the overall online identification service regarding connection establishment time, the execution device begins to establish a connection with the second online identification service. At this time, the execution device is still attempting to connect with the first online identification service. If the connection establishment time with the second online identification service also exceeds the first reference threshold for the overall online identification service regarding connection establishment time, the offline identification service is activated as the target identification service. At this time, the execution device is still attempting to connect with the second online identification service without interruption.

[0038] For example, if the execution device successfully connects to the first online recognition service, and the first online recognition service identifies the target audio data to obtain a text recognition result, and if the time interval between two consecutive text recognition results returned by the first online recognition service exceeds the second reference threshold for the overall online recognition service regarding the time interval, but does not exceed the reference threshold for the time interval of the first online recognition service, then the offline service is activated. Alternatively, if the execution device successfully connects to the first online recognition service, and the time interval between two consecutive text recognition results returned by the first online recognition service exceeds its reference threshold, then the execution device establishes a connection with the second online recognition service. Once the connection is established, the second online recognition service acts as the target recognition service and re-identifies the target audio data that the first online recognition service failed to recognize. If the time between the last text recognition result successfully returned by the first online recognition service and the time between the second online recognition service recognizing the target audio data that the first online recognition service failed to recognize and returning a text result exceeds the second reference threshold for the overall online recognition service regarding the time interval, then the offline recognition service is activated as the target recognition service.

[0039] In other embodiments, the identification services include at least two online identification services and an offline identification service, with the offline identification service having the lowest priority. When the device is detected to be connected to the network, the highest-priority online identification service is activated as the target identification service. Each time a preset online identification service is activated as the target identification service, at least one network status factor of the preset online identification service is monitored. If the network status factor of the preset online identification service exceeds a reference threshold for the preset online identification service regarding network status factors, the next-highest priority online identification service is activated. This process of monitoring at least one network status factor of the preset online identification service and subsequent steps is repeated each time a preset online identification service is activated as the target identification service until all online identification services are activated as target identification services. The preset online identification services are the remaining online identification services other than the lowest-priority among the at least two online identification services. When the device is detected to be connected to the network, at least one network status factor of the overall online identification service is detected. If all network status factors of the overall online identification service exceed the reference threshold for the overall online identification service regarding network status factors, an offline identification service is activated as the target identification service. The overall online identification service represents all online identification services. When the device is detected to be offline, an offline identification service is activated as the target identification service.

[0040] Step S130: Use the target recognition service to recognize the target audio data and obtain the text recognition result corresponding to the target audio data.

[0041] In some embodiments, target audio data is identified using various target recognition services. If the target audio data is not the last target audio data, the target recognition service provides candidate text recognition results for the target audio data. If the target audio data is the last target audio data, the target recognition service provides candidate text recognition results for all audio data in this speech.

[0042] In some specific embodiments, a candidate text recognition result from a target recognition service is received for the first time. If the target audio data is not the last target audio data, the first received candidate text recognition result is used as the text recognition result of the target audio data and provided to the user. Subsequently, whenever a new candidate text recognition result from another target recognition service is received, if the newly received candidate text recognition result is longer than the current text recognition result of the target audio data, the text recognition result of the target audio data is updated using the newly received candidate text recognition result and provided to the user.

[0043] For example, several recognition services include a first online recognition service and a second online recognition service. The first online recognition service is activated to recognize target audio data. When recognizing target audio data one, the first online recognition service successfully recognizes it and returns a text recognition result. When recognizing target audio data two, which is adjacent to target audio data one, if the time interval between the reception of the text recognition results obtained by the first online recognition service for these two adjacent target audio data exceeds its reference threshold, the second online recognition service is activated to recognize target audio data two. At this time, the first online recognition service is still recognizing target audio data two. The second online recognition service completes the recognition of target audio data two first and provides the second text recognition result of target audio data two to the user. Subsequently, the first online recognition service also completes the recognition of target audio data two and also returns its first text recognition result for target audio data two to the user. At this point, the first text recognition result and the second text recognition result of the target audio data received by the user are compared. If the number of characters in the first text recognition result is more than the number of characters in the second text recognition result, the first text recognition result replaces the second text recognition result, and the first text recognition result is the final text recognition result of the target audio data two and is provided to the user. If the number of characters in the first text recognition result is less than the number of characters in the second text recognition result, the second text recognition result is the final text recognition result of the target audio data two and is provided to the user.

[0044] In other specific embodiments, a candidate text recognition result is initially received from a target recognition service. If the target audio data is the last target audio data, and the initially received candidate text recognition result was obtained by an offline recognition service, then the initially received candidate text recognition result is used as the final text recognition result for this speech and provided to the user. If a candidate text recognition result from an online recognition service is received within a preset waiting time after the initially received candidate text recognition result, the candidate text recognition result from the online recognition service replaces the final text recognition result for this speech and is provided to the user.

[0045] For example, several recognition services include online recognition services and offline recognition services. The online recognition service is activated to recognize target audio data. The online recognition service successfully recognizes target audio data one and returns a text recognition result. When the online recognition service recognizes adjacent target audio data two, if the time interval between the online recognition service receiving the text recognition result of this adjacent target audio data exceeds its reference threshold, the offline recognition service is activated to recognize target audio data two. At this time, the online recognition service still recognizes target audio data two. The offline recognition service first completes the recognition of target audio data two and returns its second text recognition result to the user, then waits for a preset waiting time. If the online recognition service completes the recognition of target audio data two and returns its first text recognition result to the user within the preset waiting time, the first text recognition result replaces the second text recognition result, becoming the final text recognition result of target audio data two and providing it to the user. If the online recognition service has not completed the recognition of target audio data two within the preset waiting time, the second text recognition result becomes the final text recognition result of target audio data two.

[0046] Step S140: Adjust the service selection strategy based on at least one adjustment reference factor. The adjusted service selection strategy is used to select the target recognition service for the next target audio data.

[0047] The adjustment reference factors include at least one of the network status information of the online recognition service and the user's operational information during the audio recognition process. This allows the dynamically adjusted service selection strategy to choose a target recognition service that is more aligned with the network conditions of the online recognition service or the user's intentions, thereby helping to improve the efficiency and accuracy of audio recognition.

[0048] In some embodiments, the service selection strategy can be adjusted by adjusting the priorities of several recognition services based on at least one adjustment reference factor. Specifically, if the network status information of an online recognition service does not meet the network response requirements corresponding to that service, the selection priority of the online recognition service is downgraded. For example, when using an online recognition service to establish a connection and recognize target audio data, if the connection establishment time of the online recognition service exceeds its reference threshold, the online recognition service is downgraded to another online recognition service with a lower priority. As another example, when using an online recognition service to recognize target audio data, if the time interval between two consecutive text recognition results obtained by the online recognition service exceeds its reference threshold, the online recognition service is downgraded to another online recognition service with a lower priority.

[0049] Furthermore, the selection priority of an online recognition service can be upgraded based on the number of times its text recognition results are used by users, reaching a preset number. The text recognition result identified by the online recognition service is the final text recognition result, and is considered to have been used by the user if it is provided to the user. Specifically, the number of times it is used by users can be counted locally by the online recognition service. If an online recognition service's local count of the number of times its recognized text recognition results are used by users reaches its preset number, then that online recognition service is upgraded to another online recognition service with a higher priority.

[0050] In other embodiments, the network response requirements corresponding to the identification service can be adjusted based on at least one adjustment reference factor to adjust the service selection strategy. The identification services include online identification services and offline identification services, with offline identification services having the lowest priority. The reference identification service for offline identification services is the overall online identification service. If the network state factor of the overall online identification service exceeds a reference threshold for the overall online identification service regarding the network state factor, the reference threshold for the overall online identification service regarding the network state factor is lowered.

[0051] Specifically, network state factors include connection establishment time and the time interval between two consecutive text recognition results. If the connection establishment time for the overall online recognition service exceeds a first reference threshold for connection establishment time, the first reference threshold for the overall online recognition service is reduced. For example, if the sum of the connection establishment times for each online recognition service exceeds the first reference threshold for connection establishment time during the recognition of target audio data, then the first reference threshold is subtracted by one unit of time. Similarly, if the time interval between two consecutive text recognition results exceeds the second reference threshold for the overall online recognition service, the second reference threshold is reduced. For example, if the time interval between two consecutive text recognition results exceeds the second reference threshold for the overall online recognition service, then the second reference threshold is subtracted by one unit of time. It is understood that the methods for reducing the first and second reference thresholds can include subtracting a unit of time, proportional reduction, etc., and are not specifically limited here.

[0052] For example, if the receiving time interval corresponding to the overall online identification service exceeds the second reference threshold of the overall online identification service regarding the receiving time interval, and the second reference threshold of the overall online identification service is reduced, if the second reference threshold of the overall online identification service is reduced to a preset interval value, then the second reference threshold of the overall online identification service is restored to the initial reference threshold of the overall online identification service regarding the receiving time interval. If the receiving time interval corresponding to the overall online identification service exceeds the second reference threshold in the future, the first reference threshold of the overall online identification service is reduced, while the second reference threshold can remain unchanged.

[0053] In some embodiments, when using several recognition services to recognize target audio data, the service selection strategy can be adjusted based on the user's operation information during the audio recognition process. Specifically, if the user feels that the connection establishment time of the recognition service is too long, they can cancel the target audio data recognition service during the initial text waiting period. To reduce the user's waiting time and improve the user experience, the time between the start point of the current speech and the cancellation time can be set as the first time interval. The first reference threshold for the connection establishment time of each recognition service is reduced to less than the first time interval, which is more in line with the user's intention. The initial text waiting period is the time period between the start point of the current speech and the first receipt of the online text recognition result. The text recognition result recognized by the online recognition service is the online text recognition result.

[0054] For example, several identification services include not only multiple online identification services but also offline identification services, with offline identification services having the lowest priority. The first reference threshold of the offline identification service is reduced to be less than a first time interval, that is, the first reference threshold of the overall online identification service is reduced to be less than the first time interval. Furthermore, the first reference thresholds of each online identification service are reduced, wherein the sum of the reduced first reference thresholds of all online identification services is less than the first reference threshold of the offline identification service.

[0055] In other embodiments, when using several recognition services to recognize target audio data, the service selection strategy can be adjusted based on the user's operation information during the audio recognition process. Specifically, if the user feels that the waiting time for the text recognition result is too long, they can cancel the target audio data recognition service during a preset text reception period. To reduce the user's waiting time for the text recognition result and improve the user experience, the second reference threshold of each recognition service regarding the reception time interval can be adjusted to be less than the preset text reception period, which is the second time interval after the most recent online text recognition result is received.

[0056] In other embodiments, the time between the start point of the current speech and the first receipt of the online text recognition result is set as the third time interval, and the first reference threshold of each online recognition service is adjusted to be no greater than the initial reference threshold and the third time interval.

[0057] In other embodiments, the receiving time interval between every two adjacent online text recognition results is counted, and the second reference threshold of each online recognition service is adjusted to be between the maximum receiving time interval obtained by the statistics and the initial reference threshold.

[0058] In a specific application scenario, this method can be used in a translator that is connected to a network to perform real-time translation. First, the user operates the translator to enable real-time speech-to-text conversion. The translator then starts collecting audio data and translates the collected audio data in real time.

[0059] When a user initiates real-time speech-to-text conversion, the translator uses endpoint detection to monitor whether the user has started speaking. Only the last small segment of audio before speaking begins is retained (this can be dynamically adjusted, for example, retaining the audio from the first second before speaking). The remaining audio before speaking is considered invalid and discarded. The audio point at the start of speaking is recorded as the start endpoint. Upon detecting the start endpoint, the recognition service is initiated. The translator begins caching audio data and starting recognition upon detecting the start endpoint, continuing until the endpoint detection method detects the end endpoint. At this point, the caching of audio data ends, and the next front-end monitoring process begins. Specifically, when the endpoint detection method detects the end endpoint, if invalid audio persists for a period of time (this can be dynamically adjusted, for example, two seconds of silence), it determines that speaking has ended and generates a continuous segment of valid audio data.

[0060] Upon discovering the starting endpoint, the translator checks its network connection. If connected, it prioritizes online recognition services. Based on its location and carrier information, the translator identifies the nearest data center as the primary data center, the second nearest as the secondary data center, and the third nearest as the backup data center. The online recognition service includes primary, secondary, and backup data center recognition services. The translator also includes a local offline recognition service. Primary data center recognition has higher priority than secondary data center recognition, which in turn has higher priority than backup data center recognition, with offline recognition having the lowest priority. The translator first sends segments of valid audio data as target audio data to the primary data center for recognition. If the primary data center fails to establish a connection within a specified timeframe (C1), or if the time interval between two consecutive text recognition results exceeds a specified timeframe (B1), the translator accesses the secondary data center and initiates a recognition service there. Upon successful connection establishment, the target audio data is sent, and the translator listens for the return of the text recognition result. If the secondary data center fails to establish a connection within the reference threshold C2 of its connection establishment time, or if the time interval between two consecutive text recognition results obtained by the secondary data center exceeds the reference threshold B2 of its reception time interval, then it accesses the backup data center and initiates a recognition service to the backup data center. After the connection is successfully established, it sends the target audio data and listens for the return of the text recognition result.

[0061] If the translator is not connected to the internet, the online recognition service will not be activated; instead, the offline recognition service will be used directly. If the device is connected to the internet, and the entire online recognition service fails to establish a connection within its first reference threshold R1 timeframe, or the time interval between receiving two adjacent text recognition results for the overall online recognition service exceeds its second reference threshold R2 time interval, then the local offline recognition service will be activated. After activating the local offline recognition service, the target audio data will be sent to the local offline recognition service, and the return of the text recognition results will be monitored.

[0062] If the target audio data is not the last target audio data, the first received candidate text recognition result is used as the text recognition result of the target audio data and provided to the user. Subsequently, whenever a new candidate text recognition result from another recognition service is received, if the number of characters in the new candidate text recognition result is greater than the number of characters in the current text recognition result of the target audio data, the text recognition result of the target audio data is updated using the new candidate text recognition result and provided to the user. If the target audio data is the last target audio data, and the first received candidate text recognition result was obtained by an offline recognition service, then the first received candidate text recognition result is used as the final text recognition result of this speech and provided to the user. If a candidate text recognition result from an online recognition service is received within a waiting time D1 after the first received candidate text recognition result, then the candidate text recognition result from the online recognition service replaces the final text recognition result of this speech and is provided to the user. If the target audio data is the last target audio data, and the first received candidate text recognition result was obtained by an online recognition service, then the candidate text recognition result is considered the final text recognition result of this speech and provided to the user.

[0063] The appropriate target recognition service is selected based on the adjusted service selection strategy. The adjustment of the service selection strategy includes a multi-level degradation strategy and a multi-level upgrade strategy.

[0064] The multi-level degradation strategy is as follows: Upon discovering the starting endpoint, the target audio data is identified using the online recognition service. During the recognition process, if a poor network condition or equipment failure in the data center causes the connection establishment time of the Level 1 data center recognition service to exceed C1 or the reception time interval to exceed B1, the Level 1 data center is downgraded to a Level 2 data center; or if the connection establishment time of the Level 2 data center recognition service exceeds C2 or the reception time interval exceeds B2, the Level 2 data center is downgraded to a backup data center; or if the connection establishment time of the overall online recognition service exceeds the first reference threshold R1 for connection establishment time, the time unit R1 is reduced.

[0065] If the overall online identification service successfully establishes a connection within time R1, and the corresponding reception time interval of the overall online identification service exceeds the second reference threshold R2 for the reception time interval, then the time unit R2 is reduced. When the time unit R2 is reduced to the preset interval value F1, the R2 value is restored to the initial reference threshold, and then the time unit R1 is reduced. When the time unit R1 is reduced to zero, the offline identification service is immediately activated after the start endpoint is detected.

[0066] If, after the start endpoint appears, the user cancels the target audio data recognition service during the initial text waiting period, the first time interval U1 between the start endpoint of this speech and the cancellation time is calculated, and the time units of C1, C2, and R1 are reduced to ensure that C1 + C2 is less than R1 and R1 is less than U1. If the user cancels the target audio data recognition service during a preset text reception period, which is within the second time interval T1 after the most recent online text recognition result is received, the time units of B1, B2, and R2 are reduced to ensure that B1, B2, and R2 are less than T1.

[0067] The multi-level upgrade strategy is as follows: The translator counts the number of text recognition results returned by the final secondary server room. If the count reaches its preset number N1, the secondary server room is upgraded to a primary server room. The translator also counts the number of text recognition results returned by the final backup server room. If the count reaches its preset number N2, the backup server room is upgraded to a secondary server room. If the first text recognition result is submitted to the user, the third time interval U2 between the start of the current speech and the first received online text recognition result is calculated. C1 and C2 are relaxed until they reach the initial reference threshold but do not exceed U2. The reception time interval between any two adjacent online text recognition results is counted, and the maximum value is recorded as T2. B1 and B2 are relaxed until they are greater than T2 but less than their corresponding initial reference thresholds.

[0068] This application uses an audio endpoint detection method to remove invalid audio for segmentation recognition, and employs a multi-level degradation and upgrade strategy to intelligently schedule data center services and local offline engine recognition services, so as to provide accurate results quickly even when the network is poor, balancing real-time performance and accuracy.

[0069] Please see Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the speech-to-text method of this application. The specific steps are as follows:

[0070] Step S210: Obtain audio data and analyze it to obtain target audio data.

[0071] This step is the same as step S110 above, and will not be described in detail here.

[0072] Step S220: Determine the device's network connection status. If the device is connected to the network, proceed to step S240; if the device is not connected to the network, proceed to step S230.

[0073] In some embodiments, if the device is not connected to the network or the network conditions are unstable or poor, the offline recognition service is used directly to recognize the target audio data; if the device is connected to the network and the network conditions are stable, the online recognition service is used to recognize the target audio data.

[0074] Step S230: Use offline recognition service to perform target recognition service and recognize the target audio data.

[0075] In some embodiments, the target audio data is recognized using an offline recognition service, and the resulting text recognition result is directly returned to the user. If, within a time period D1 after the offline recognition service is enabled as the target recognition service and the text recognition result is provided to the user, the device successfully connects to the network again and obtains a new text recognition result, the new text recognition result replaces the text recognition result obtained from the offline recognition service. Subsequently, to prioritize the online recognition service, step S250 is executed, adjusting the service selection strategy of the online recognition service based on the adjustment reference factor, and selecting the target online recognition service to recognize the next target audio data according to the adjusted service selection strategy.

[0076] In other embodiments, the device is never connected to the network and always uses an offline recognition service to identify the target audio data, so step S250 is not executed.

[0077] Step S240: Using the network status information and service selection strategy of the online identification service, enable at least one target online identification service to identify the target audio data.

[0078] This step is the same as steps S120 and S130 above, and will not be described in detail here.

[0079] Step S250: Adjust the service selection strategy based on at least one adjustment reference factor. The adjusted service selection strategy is used to select the target online recognition service for the next target audio data.

[0080] This step is the same as step S140 above, and will not be described in detail here.

[0081] Please see Figure 3 , Figure 3This is a schematic diagram of the framework of an embodiment of the speech-to-text device of this application, including: an acquisition module 310, a selection module 320, a recognition module 330, and an adjustment module 340. The acquisition module 310 is used to acquire target audio data; the selection module 320 is used to enable at least one target recognition service from several recognition services based on the network status information of the online recognition service and a service selection strategy, wherein the several recognition services include at least one online recognition service; the recognition module 330 is used to recognize the target audio data using the target recognition service to obtain the text recognition result corresponding to the target audio data; the adjustment module 340 is used to adjust the service selection strategy based on at least one adjustment reference factor, wherein the adjusted service selection strategy is used to select the target recognition service for the next target audio data, and the adjustment reference factor includes at least one of the network status information of the online recognition service and the user's operation information during the audio recognition process.

[0082] In some embodiments, the selection module 320 has different priorities for several identification services. The service selection strategy is as follows: each identification service is enabled sequentially according to its priority as the target identification service, and the activation trigger condition for identification services other than the highest priority is that the network status information of the reference identification service after activation meets the preset response requirements corresponding to the reference identification service; the reference identification service is an online identification service with a higher priority than the identification service; the adjustment module 340 performs the service selection strategy adjustment based on at least one adjustment reference factor, including at least the following steps: adjusting the priority of several identification services based on at least one adjustment reference factor; adjusting the network response requirements corresponding to the identification service based on at least one adjustment reference factor.

[0083] In some embodiments, the network state information of the reference recognition service after it is enabled by the selection module 320 includes at least one network state factor. The preset response requirements corresponding to the reference recognition service include: each network state factor of the reference recognition service exceeds the reference threshold of the reference recognition service regarding the network state factor. The network state factor includes at least one of the following: the connection establishment time of the reference recognition service, and the time interval between receiving two adjacent text recognition results obtained by the reference recognition service.

[0084] In some embodiments, the reference identification service for the online identification service performed by the selection module 320 is another online identification service with a higher priority than the online identification service; and / or, the identification services further include an offline identification service, the offline identification service having the lowest priority, and the reference identification service for the offline identification service is the overall online identification service. The network state factors of the overall online identification service include at least one of the following: the connection establishment time of each online identification service, and the time interval between receiving two adjacent online text identification results, wherein the online text identification results are obtained by any enabled online identification service; the preset response requirements corresponding to the overall online identification service include at least one of the following: the connection establishment time of each online identification service exceeds a first reference threshold for the overall online identification service regarding the connection establishment time, and the time interval between receiving two adjacent online text identification results exceeds a second reference threshold for the overall online identification service regarding the time interval.

[0085] In some embodiments, the plurality of identification services include at least two online identification services and an offline identification service, with the offline identification service having the lowest priority. The selection module 320 executes network status information and a service selection strategy based on the online identification services to select at least one target identification service from the plurality of identification services for activation, including: when the device is detected to be connected to the network, activating the highest-priority online identification service as the target identification service; each time a preset online identification service is activated as the target identification service, monitoring at least one network status factor of the preset online identification service; in response to the network status factor of the preset online identification service exceeding a reference threshold for the network status factor of the preset online identification service, activating the next-priority online identification service of the preset online identification service, and repeating this process. Each time a preset online identification service is enabled as the target identification service, at least one network status factor of the preset online identification service and subsequent steps are monitored until all online identification services are enabled as target identification services. The preset online identification service is any online identification service other than the lowest priority among at least two online identification services. If the current device is detected to be connected to the network, at least one network status factor of the overall online identification service is detected. In response to each network status factor of the overall online identification service exceeding the reference threshold of the overall online identification service regarding network status factors, an offline identification service is enabled as the target identification service. The overall online identification service represents all online identification services. If the current device is detected to be offline, an offline identification service is enabled as the target identification service.

[0086] In some embodiments, the identification services further include an offline identification service, which has the lowest priority, and the reference identification service for the offline identification service is the overall online identification service; the adjustment module 340 performs adjustments to the network response requirements corresponding to the identification service based on at least one adjustment reference factor, including: in response to the network state factor of the overall online identification service exceeding the reference threshold of the overall online identification service with respect to the network state factor, reducing the reference threshold of the overall online identification service with respect to the network state factor.

[0087] In some embodiments, the network state factor includes the connection establishment time and the time interval between two adjacent text recognition results; the adjustment module 340 performs actions in response to the network state factor of the overall online recognition service exceeding a reference threshold of the overall online recognition service regarding the network state factor, the reference threshold of the overall online recognition service regarding the network state factor including: in response to the connection establishment time corresponding to the overall online recognition service exceeding a first reference threshold of the overall online recognition service regarding the connection establishment time, lowering the first reference threshold of the overall online recognition service; in response to the time interval corresponding to the overall online recognition service exceeding a second reference threshold of the overall online recognition service regarding the time interval, lowering the second reference threshold of the overall online recognition service; and in response to the second reference threshold of the overall online recognition service decreasing to a preset interval value, restoring the second reference threshold of the overall online recognition service to the initial reference threshold of the overall online recognition service regarding the time interval, and lowering the first reference threshold of the overall online recognition service.

[0088] In some embodiments, the network state factors include at least one of the following: connection establishment duration and the time interval between two adjacent text recognition results, wherein the text recognition result recognized by the online recognition service is the online text recognition result; the adjustment module 340 performs adjustments to the network response requirements corresponding to the recognition service based on at least one adjustment reference factor, including at least one of the following steps: in response to the user canceling the target audio data recognition service during the initial text waiting period, obtaining a first time interval between the start point of the current speech and the cancellation time, reducing the first reference threshold of the reference recognition service for each recognition service with respect to the connection establishment duration to less than the first time interval, wherein the initial text waiting period is the time interval between the start point of the current speech and the first receipt of the online text recognition result. Interval; In response to the user canceling the target audio data recognition service during the preset text reception period, the second reference threshold of the reference recognition service for each recognition service with respect to the reception time interval is adjusted to be less than the preset text reception period, which is the second time interval after the most recent online text recognition result is received; the third time interval between the start point of this speech and the first online text recognition result is obtained, and the first reference threshold of the reference recognition service for each online recognition service is adjusted to be no greater than the initial reference threshold and the third time interval; the reception time interval between every two adjacent online text recognition results is counted, and the second reference threshold of the reference recognition service for each online recognition service is adjusted to be between the largest received time interval obtained from the statistics and the initial reference threshold.

[0089] In some embodiments, the identification services further include an offline identification service, which has the lowest priority; the adjustment module 340 performs the action of reducing the first reference threshold of the reference identification service of each identification service with respect to the connection establishment time to less than the first time interval, including: reducing the first reference threshold of the reference identification service of the offline identification service to less than the first time interval; and reducing the first reference threshold of the reference identification service of each online identification service, wherein the sum of the first reference thresholds of the reference identification services of all online identification services after the reduction is less than the first reference threshold of the reference identification service of the offline identification service.

[0090] In some embodiments, the adjustment module 340 performs an adjustment of the priority of several recognition services based on at least one adjustment reference factor, including: downgrading the selection priority of the online recognition service in response to the network status information of the online recognition service not meeting the network response requirements corresponding to the online recognition service; and upgrading the selection priority of the online recognition service in response to the number of times the text recognition result corresponding to the online recognition service is used by the user reaching a preset number of times corresponding to the online recognition service.

[0091] In some embodiments, the acquisition module 310 executes an audio segment in the current speech as the target audio data; the method execution device sequentially takes each audio segment in the current speech as target audio data according to the order of audio acquisition and executes the method based on the target audio data to obtain the text recognition result corresponding to each audio segment; acquiring the first target audio data of the current speech includes: in response to detecting the start endpoint of the current speech, acquiring the audio segment located in a first time range of the start endpoint as the first target audio data; acquiring the last target audio data of the current speech includes: in response to detecting the end endpoint of the current speech, acquiring the audio segment located in a second time range of the end endpoint as the last target audio data.

[0092] In some embodiments, the acquisition module 310 performs the following: acquiring audio segments within a first time range from the start endpoint as the first target audio data, including: acquiring audio segments of a preset time length as the first target audio data starting from a first preset time before the start endpoint; acquiring audio segments within a second time range from the end endpoint as the last target audio data, including: acquiring audio segments of a preset time length after the end endpoint as the last target audio data; and / or, the recognition module 330 performs the following: using a target recognition service to recognize the target audio data and obtain the text recognition result corresponding to the target audio data, including: using each target recognition service to recognize the target audio data, wherein if the target audio data is not the last target audio data, the target recognition service returns candidate text recognition results for the target audio data; if the target audio data is the last target audio data, the target recognition service returns candidate text recognition results for all audio data in this speech; receiving a first audio segment within a second time range as the last target audio data, including: acquiring audio segments within a first time range from the start endpoint as the last target audio data, including: acquiring audio segments of a preset time length after the end ... The target recognition service provides candidate text recognition results. If the target audio data is not the last target audio data, the first received candidate text recognition result is used as the text recognition result of the target audio data and provided to the user. Subsequently, whenever a new candidate text recognition result from another target recognition service is received, if the newly received candidate text recognition result is longer than the current text recognition result of the target audio data, the newly received candidate text recognition result is used to update the text recognition result of the target audio data and provided to the user. If the target audio data is the last target audio data, and the first received candidate text recognition result was obtained by the offline recognition service, the first received candidate text recognition result is used as the final text recognition result of this speech and provided to the user. If a candidate text recognition result from the online recognition service is received within a preset waiting time after the first received candidate text recognition result, the candidate text recognition result from the online recognition service is used to replace the final text recognition result of this speech and provided to the user.

[0093] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0094] Please see Figure 4 , Figure 4 This is a schematic diagram of a framework of an embodiment of the electronic device 40 of this application. The electronic device 40 includes a memory 41 and a processor 42 coupled to each other. The processor 42 is used to execute program instructions stored in the memory 41 to implement the steps in any of the above-described speech-to-text method embodiments. In a specific implementation scenario, the electronic device 40 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 40 may also include mobile devices such as laptops and tablets, which are not limited here.

[0095] Specifically, processor 42 controls itself and memory 41 to implement the steps in any of the above-described speech-to-text method embodiments. Processor 42 can also be referred to as a CPU (Central Processing Unit). Processor 42 may be an integrated circuit chip with signal processing capabilities. Processor 42 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 42 can be implemented using integrated circuit chips.

[0096] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 50 of this application. The computer-readable storage medium 50 stores program instructions 501 that can be executed by a processor. The program instructions 501 are used to implement the steps in any of the above-described embodiments of the speech-to-text method.

[0097] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0098] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0100] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A speech-to-text method, characterized in that, include: Acquire target audio data; Based on the network status information of the online identification service and the service selection strategy, at least one target identification service is activated from a plurality of identification services, wherein the plurality of identification services include at least one online identification service, and the plurality of identification services have different priorities. The service selection strategy is as follows: each identification service is activated sequentially according to its priority as the target identification service, and the activation trigger condition for the identification service other than the highest priority is that the network status information of the reference identification service of the identification service after activation meets the preset response requirements corresponding to the reference identification service; the reference identification service of the identification service is the online identification service and has a higher priority than the identification service. The target audio data is identified using the target recognition service to obtain the text recognition result corresponding to the target audio data; and The service selection strategy is adjusted based on at least one adjustment reference factor. The adjusted service selection strategy is used to select the target recognition service for the next target audio data. The adjustment reference factor includes at least one of the network status information of the online recognition service and the user's operation information during the audio recognition process. The step of adjusting the service selection strategy based on at least one adjustment reference factor includes at least one of the following steps: adjusting the priority of the plurality of identification services based on the at least one adjustment reference factor; and adjusting the network response requirements corresponding to the identification service based on the at least one adjustment reference factor.

2. The method according to claim 1, characterized in that, The network state information of the reference identification service after it is enabled includes at least one network state factor, and the preset response requirements corresponding to the reference identification service include: the network state factor of the reference identification service exceeds the reference threshold of the reference identification service regarding the network state factor; The network state factors include at least one of the following: the connection establishment time of the reference recognition service, and the time interval between receiving two adjacent text recognition results obtained by the reference recognition service.

3. The method according to claim 2, characterized in that, The reference identification service for the online identification service is another online identification service with a higher priority than the online identification service. And / or, the plurality of identification services further includes an offline identification service, the offline identification service having the lowest priority, the reference identification service for the offline identification service being the overall online identification service, the network state factors of the overall online identification service including at least one of the following: the connection establishment time of each of the online identification services, and the time interval between receiving two adjacent online text recognition results, the online text recognition results being obtained by any enabled online identification service; the preset response requirements corresponding to the overall online identification service include at least one of the following: the connection establishment time of each of the online identification services exceeds a first reference threshold of the overall online identification service regarding the connection establishment time, and the time interval between receiving two adjacent online text recognition results exceeds a second reference threshold of the overall online identification service regarding the time interval.

4. The method according to claim 2, characterized in that, The plurality of identification services also includes an offline identification service, which has the lowest priority and the reference identification service for the offline identification service is the overall online identification service. The adjustment of the network response requirements corresponding to the identification service based on the at least one adjustment reference factor includes: In response to the network state factor of the overall online identification service exceeding the reference threshold of the overall online identification service with respect to the network state factor, the reference threshold of the overall online identification service with respect to the network state factor is reduced.

5. The method according to claim 2, characterized in that, The network state factors include at least one of the following: the connection establishment time and the time interval between receiving two adjacent text recognition results, wherein the text recognition result recognized by the online recognition service is the online text recognition result; The step of adjusting the network response requirements corresponding to the identification service based on the at least one adjustment reference factor includes at least one of the following steps: In response to a user canceling the target audio data recognition service during the initial text waiting period, a first time interval between the start point of the current speech and the cancellation time is obtained, and the first reference threshold of each recognition service with respect to the connection establishment time is reduced to less than the first time interval. The initial text waiting period is the time interval between the start point of the current speech and the first receipt of the online text recognition result. In response to a user canceling the target audio data recognition service during a preset text reception period, the second reference threshold of each recognition service with respect to the reception time interval is adjusted to be less than the preset text reception period, where the preset text reception period is the second time interval after the most recent online text recognition result is received. Obtain the third time interval between the start point of this speech and the first receipt of the online text recognition result, and adjust the first reference threshold of each of the online recognition services to be no greater than the initial reference threshold and the third time interval; The receiving time interval between each two adjacent online text recognition results is counted, and the second reference threshold of each online recognition service is adjusted to be between the largest received time interval obtained by statistics and the initial reference threshold.

6. The method according to claim 1, characterized in that, The adjustment of the priority of the plurality of identification services based on the at least one adjustment reference factor includes: If the network status information of the online identification service does not meet the network response requirements corresponding to the online identification service, the selection priority of the online identification service will be downgraded. In response to the number of times the text recognition result corresponding to the online recognition service is used by the user reaching the preset number of times corresponding to the online recognition service, the selection priority of the online recognition service is upgraded.

7. The method according to claim 1, characterized in that, The target audio data is an audio segment in the current speech; the execution device of the method sequentially takes each audio segment in the current speech as the target audio data according to the order of audio acquisition and executes the method based on the target audio data to obtain the text recognition result corresponding to each audio segment; Obtain the first target audio data of this speech, including: In response to detecting the start point of the current speech, the audio segment located within the first time range of the start point is acquired as the first target audio data; Obtain the last target audio data of this speech, including: In response to detecting the end point of the current speech, an audio segment located within a second time range of the end point is acquired as the last target audio data.

8. A speech-to-text device, characterized in that, include: The acquisition module is used to acquire target audio data; A selection module is used to enable at least one target identification service from a plurality of identification services based on network status information of the online identification service and a service selection strategy. The plurality of identification services include at least one online identification service. The plurality of identification services have different priorities. The service selection strategy is as follows: each identification service is enabled sequentially according to its priority as the target identification service. The activation trigger condition for an identification service other than the one with the highest priority is that the network status information of the reference identification service after activation meets the preset response requirements corresponding to the reference identification service. The reference identification service is the online identification service with a higher priority than the identification service. The recognition module is used to recognize the target audio data using the target recognition service to obtain the text recognition result corresponding to the target audio data; An adjustment module is used to adjust the service selection strategy based on at least one adjustment reference factor. The adjusted service selection strategy is used to select a target recognition service for the next target audio data. The adjustment reference factor includes at least one of the network status information of the online recognition service and the user's operation information during the audio recognition process. The adjustment of the service selection strategy based on at least one adjustment reference factor includes at least the following steps: adjusting the priority of the plurality of recognition services based on the at least one adjustment reference factor; and adjusting the network response requirements corresponding to the recognition service based on the at least one adjustment reference factor.

9. An electronic device, characterized in that, It includes a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory to implement the speech-to-text method according to any one of claims 1 to 7.