Audio data transcription method and electronic device

By locally cached when the mobile terminal receives the audio data cache instruction, and sending the cached audio data to the mobile terminal for conversion under appropriate conditions, the problem of excessive amount of information during audio data conversion in the prior art is solved, and the smoothness and accuracy of transcription are improved.

CN113903341BActive Publication Date: 2025-05-20ANKER INNOVATIONS TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111167964.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-05-20
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

In the prior art, when the audio data collected by the mobile terminal is converted into text, the external expansion sound causes the amount of information entered by the recording device to be too large, affecting the accuracy and reliability of the conversion.

Method used

When the mobile terminal receives the audio data cache instruction, the collected audio data is cached locally, and when the audio data return instruction is received, the cached audio data is sent to the mobile terminal for conversion.

Benefits of technology

It effectively solves the conversion blocking problem caused by large amount of information, and improves the smoothness and accuracy of audio data transcription.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113903341B_ABST
    Figure CN113903341B_ABST
Patent Text Reader

Abstract

The disclosed embodiments disclose a method and electronic device for transcribing audio data. The method comprises: upon receiving an audio data caching instruction sent by a mobile terminal, locally caching the audio data received from the mobile terminal and collected by the mobile terminal; upon receiving an audio data return instruction sent by the mobile terminal, sending the locally cached audio data to the mobile terminal so that the mobile terminal transcribes the received audio data into text. The disclosed embodiments cache the audio data received from the mobile terminal and collected by the mobile terminal, thereby achieving audio data transcription of the mobile terminal in the event of congestion due to a large amount of information, and improving the smoothness and accuracy of audio data transcription.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for transcribing audio data and an electronic device. Background Art

[0002] In the prior art, if one wants to convert audio data (such as call data) collected by a mobile terminal into text, it is usually necessary to separately amplify the sound and use a voice recorder to record the sound, and then convert it into text.

[0003] However, since the voice recorder and other recording devices are in an open environment when amplifying the sound, they can record not only the amplified sound, but also the user's own voice, environmental noise, etc., resulting in an abnormal increase in the amount of information. The conversion processing speed of the mobile terminal cannot handle the large amount of information, affecting the accuracy and reliability of the conversion. Summary of the Invention

[0004] In view of this, to solve the above partial or all technical problems, embodiments of the present disclosure provide a method for transcribing audio data and an electronic device to improve the smoothness and accuracy of audio data transcription.

[0005] In a first aspect, embodiments of the present disclosure provide a method for transcribing audio data, the method including:

[0006] When receiving an audio data caching instruction sent by a mobile terminal, locally cache the audio data collected by the mobile terminal and received from the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold;

[0007] When receiving an audio data return instruction sent by the mobile terminal, send the locally cached audio data to the mobile terminal so that the mobile terminal can convert the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold.

[0008] Optionally, in the method of any embodiment of the present disclosure, the audio data collected by the mobile terminal includes voice audio of at least two speakers; and

[0009] The locally caching the audio data collected by the mobile terminal and received from the mobile terminal includes any one of the following:

[0010] Taking the speech audio of a single target speaker in the audio data collected by the mobile terminal received above as one piece of audio data for local caching; wherein, the speech audio of the target speaker is: the speech audio of the speaker selected from the speech audio of at least two speakers, or the speech audio of a preset speaker;

[0011] Taking the speech audio of each speaker in the speech audio of at least two speakers selected from the audio data collected by the mobile terminal received above as one piece of audio data for local caching respectively;

[0012] Taking the speech audio of at least two speakers selected from the audio data collected by the mobile terminal received above as one piece of audio data for local caching.

[0013] Optionally, in the method of any embodiment of the present disclosure, the sending the locally cached audio data to the mobile terminal above includes:

[0014] Sending at least one piece of audio data in the locally cached audio data to the mobile terminal above.

[0015] Optionally, in the method of any embodiment of the present disclosure, the method further includes:

[0016] When receiving an audio data confirmation instruction sent by the mobile terminal above, deleting the audio data corresponding to the audio data confirmation instruction from the local cache, wherein the audio data confirmation instruction indicates that the mobile terminal has received the audio data or the mobile terminal has completed the conversion of the audio data.

[0017] Optionally, in the method of any embodiment of the present disclosure, the method is applied to a Bluetooth adapter; and

[0018] the method further includes:

[0019] When the mobile terminal above receives a Bluetooth connection request from the Bluetooth headset, or the Bluetooth adapter receives a Bluetooth connection request from the Bluetooth headset, establishing a Bluetooth connection with the Bluetooth headset;

[0020] Sending the audio data collected by the mobile terminal received from the mobile terminal to the Bluetooth headset through the Bluetooth connection above.

[0021] Optionally, in the method of any embodiment of the present disclosure, the method is applied to a Bluetooth adapter, and the Bluetooth adapter establishes a connection with the mobile terminal through a connection port, and the connection port is used for transmitting the audio data between the Bluetooth adapter and the mobile terminal.

[0022] Optionally, in the method of any embodiment of the present disclosure, the mobile terminal has no conversion permission for the audio data collected by the mobile terminal.

[0023] In a second aspect, an embodiment of the present disclosure provides a transcription device for audio data, the device includes:

[0024] A cache unit, configured to locally cache the audio data collected by the mobile terminal received from the mobile terminal when receiving an audio data caching instruction sent by the mobile terminal, wherein the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold;

[0025] A first sending unit, configured to send the locally cached audio data to the mobile terminal when receiving an audio data return instruction sent by the mobile terminal, so that the mobile terminal converts the received audio data into text, wherein the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold.

[0026] Optionally, in the device of any embodiment of the present disclosure, the audio data collected by the mobile terminal includes voice audio of at least two speakers; and

[0027] The cache unit includes any one of the following:

[0028] A first cache subunit, configured to locally cache the voice audio of a single target speaker in the audio data collected by the mobile terminal received from the mobile terminal as one piece of audio data; wherein the voice audio of the target speaker is: the voice audio of the speaker selected from the voice audio of the at least two speakers, or the voice audio of a preset speaker;

[0029] A second cache subunit, configured to locally cache the voice audio of each speaker in the voice audio of at least two speakers selected from the audio data collected by the mobile terminal received from the mobile terminal as one piece of audio data respectively;

[0030] A third cache subunit, configured to locally cache the voice audio of at least two speakers selected from the audio data collected by the mobile terminal received from the mobile terminal as one piece of audio data.

[0031] Optionally, in the device of any embodiment of the present disclosure, the first sending unit includes:

[0032] A sending subunit, configured to send at least one piece of audio data in locally cached audio data to the mobile terminal.

[0033] Optionally, in the device according to any embodiment of the present disclosure, the device further includes:

[0034] A deleting unit, configured to delete, from the local cache, the audio data corresponding to the audio data confirmation instruction when receiving the audio data confirmation instruction sent by the mobile terminal, where the audio data confirmation instruction indicates that the mobile terminal has received the audio data or the mobile terminal has completed the conversion of the audio data.

[0035] Optionally, in the device according to any embodiment of the present disclosure, the device further includes:

[0036] A connection establishing unit, configured to establish a Bluetooth connection with the Bluetooth headset when the mobile terminal receives a Bluetooth connection request from the Bluetooth headset, or when the Bluetooth adapter receives a Bluetooth connection request from the Bluetooth headset;

[0037] A second sending unit, configured to send, through the Bluetooth connection, the audio data collected by the mobile terminal received from the mobile terminal to the Bluetooth headset.

[0038] Optionally, in the device according to any embodiment of the present disclosure, the Bluetooth adapter establishes a connection with the mobile terminal through a connection port, and the connection port is used for transmitting the audio data between the Bluetooth adapter and the mobile terminal.

[0039] Optionally, in the device according to any embodiment of the present disclosure, the mobile terminal has no conversion permission for the audio data collected by the mobile terminal.

[0040] In a third aspect, an embodiment of the present disclosure provides a method for transcribing audio data, the method including:

[0041] When the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold, sending an audio data caching instruction of the locally collected audio data to a target device, so that the target device caches the locally collected audio data;

[0042] When the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, sending an audio data return instruction to the target device, and receiving the audio data corresponding to the audio data return instruction sent by the target device;

[0043] Convert the received audio data into text, where the second preset data volume threshold is less than the first preset data volume threshold.

[0044] Optionally, in the method of any embodiment of the present disclosure, the collected audio data includes voice audio of at least two speakers; and

[0045] The target device caches the locally collected audio data in any of the following ways:

[0046] Use the voice audio of a single target speaker in the audio data collected by the mobile terminal received from the mobile terminal as an audio data for local caching; where the voice audio of the target speaker is: the voice audio of the speaker selected from the voice audio of the at least two speakers, or the voice audio of a preset speaker;

[0047] Use the voice audio of each speaker in the at least two speakers' voice audio selected from the audio data collected by the mobile terminal received from the mobile terminal as an audio data for local caching respectively;

[0048] Use the at least two speakers' voice audio selected from the audio data collected by the mobile terminal received from the mobile terminal as an audio data for local caching.

[0049] Optionally, in the method of any embodiment of the present disclosure, after sending the audio data caching instruction of the locally collected audio data to the target device, the target device deletes the audio data corresponding to the audio data confirmation instruction from the cache, where the audio data confirmation instruction indicates that the mobile terminal has received the audio data or the mobile terminal has completed the conversion of the audio data.

[0050] Optionally, in the method of any embodiment of the present disclosure, when the method is applied to a mobile terminal and the mobile terminal receives a Bluetooth connection request from a Bluetooth headset, or when the target device receives a Bluetooth connection request from the Bluetooth headset, the target device establishes a Bluetooth connection with the Bluetooth headset; and

[0051] The method further includes:

[0052] Send the audio data collected by the mobile terminal to the Bluetooth headset via the target device through the Bluetooth connection.

[0053] Fourthly, an embodiment of the present disclosure provides an electronic device, including:

[0054] A memory for storing a computer program;

[0055] A processor for executing a computer program stored in the memory, and when the computer program is executed, implementing the method of any one of the embodiments of the audio data transcription method in the first aspect or the third aspect of the present disclosure.

[0056] In a fifth aspect, an embodiment of the present disclosure provides a computer-readable medium, and when the computer program is executed by a processor, implementing the method of any one of the embodiments of the audio data transcription method in the first aspect or the third aspect as described above.

[0057] In a sixth aspect, an embodiment of the present disclosure provides a computer program, which includes computer-readable code, and when the computer-readable code runs on a device, causing the processor in the device to execute instructions for implementing each step in the method of any one of the embodiments of the audio data transcription method in the first aspect or the third aspect as described above.

[0058] Based on the audio data transcription method provided by the above embodiments of the present disclosure, in the case of receiving an audio data caching instruction sent by a mobile terminal, local caching can be performed on the audio data collected by the mobile terminal and received from the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold. After that, in the case of receiving an audio data return instruction sent by the mobile terminal, the locally cached audio data is sent to the mobile terminal so that the mobile terminal can convert the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold. Thus, the embodiments of the present disclosure achieve the conversion of audio data by the mobile terminal in the case of blockage due to a large amount of information by caching the audio data collected by the mobile terminal and received from the mobile terminal, improving the smoothness and accuracy of audio data transcription.

[0059] Next, through the drawings and embodiments, the technical solutions of the present disclosure will be further described in detail. Description of the Drawings

[0060] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent:

[0061] Figure 1 is an exemplary system architecture diagram of an audio data transcription method or an audio data transcription device provided by an embodiment of the present disclosure;

[0062] Figure 2It is a flowchart of a method for transcribing audio data provided by an embodiment of the present disclosure;

[0063] Figure 3 It is for Figure 2 a schematic diagram of an application scenario of an embodiment;

[0064] Figure 4A It is a schematic diagram of an interaction process of a method for transcribing audio data provided by an embodiment of the present disclosure;

[0065] Figure 4B It is a schematic diagram of the connection manner of a Bluetooth adapter, a Bluetooth headset, and a mobile terminal involved in a method for transcribing audio data provided by an embodiment of the present disclosure;

[0066] Figure 5 It is a schematic diagram of the structure of a device for transcribing audio data provided by an embodiment of the present disclosure;

[0067] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed Embodiments

[0068] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and values set forth in these embodiments do not limit the scope of the present disclosure.

[0069] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices, or modules, etc., without representing any specific technical meaning and without indicating their logical order.

[0070] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two, or more.

[0071] It should also be understood that for any component, data, or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it is generally understood as one or more.

[0072] In addition, the term "and / or" in the present disclosure is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0073] It should also be understood that the descriptions of the various embodiments in this disclosure emphasize the differences between the various embodiments, and the similarities or resemblances between them can be referred to each other. For the sake of brevity, they will not be elaborated one by one.

[0074] The following description of at least one exemplary embodiment is actually merely illustrative and in no way a limitation on this disclosure or its application or use.

[0075] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.

[0076] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0077] It should be noted that, without conflict, the embodiments in this disclosure and the features in the embodiments can be combined with each other. The following will detail this disclosure with reference to the drawings and in conjunction with the embodiments.

[0078] Figure 1 It is an exemplary system architecture diagram of a method for transcribing audio data or a device for transcribing audio data provided by an embodiment of this disclosure.

[0079] As Figure 1 shown, the system architecture 100 may include electronic devices 102, 103, a network 106, and a mobile terminal 104. Optionally, the system architecture 100 may further include a Bluetooth device 101 and a network 105. Among them, the network 106 can provide a medium for the communication chain between the electronic devices 102, 103 and the mobile terminal 104. The network 106 may include various connection types, such as wired, wireless communication chains, or fiber optic cables, etc. The electronic devices 102, 103 can interact with the mobile terminal 104 through the network 106 to receive the audio data sent by the mobile terminal 104 and send the locally cached audio data to the above-mentioned mobile terminal 104.

[0080] In addition, when the system architecture 100 includes a Bluetooth device 101 and a network 105, the network 105 can provide a medium for the communication chain between the electronic devices 102, 103 and the Bluetooth device 101. As an example, the network 106 can be a Bluetooth network. In this case, the electronic devices 102, 103 can send the audio data to the mobile terminal 104.

[0081] Here, at least one of the electronic devices 102 and 103 can be the execution subject of the audio data transcription method provided by the embodiments of the present disclosure. However, it should be noted that the execution subject of the audio data transcription method provided by the embodiments of the present disclosure can be hardware or software, and no specific limitation is made here.

[0082] As an example, the electronic devices 102 and 103 can be a Bluetooth adapter, another mobile terminal different from the mobile terminal 104, or a server, etc. The mobile terminal 104 can be various mobile terminals with the function of audio data collection (such as the recording function). For example, a mobile phone, a computer, etc. The Bluetooth device 101 can be a Bluetooth speaker, a Bluetooth headset, etc.

[0083] It should be understood that Figure 1 the number of electronic devices and networks in

[0084] Continuing to refer to Figure 2 , a flowchart 200 of an embodiment of the audio data transcription method according to the present disclosure is shown. The audio data transcription method includes the following steps:

[0085] Step 201, in the case of receiving an audio data caching instruction sent by the mobile terminal, locally cache the audio data received from the mobile terminal and collected by the mobile terminal.

[0086] In this embodiment, in the case of receiving an audio data caching instruction sent by the mobile terminal, the execution subject of the audio data transcription method (such as Figure 1 the electronic devices 102 and 103 shown) can locally cache the audio data received from the mobile terminal and collected by the mobile terminal through a wired connection method or a wireless connection method.

[0087] Among them, the above audio data caching instruction is sent to the above execution subject by the above mobile terminal when the data volume of the audio data to be converted by the above mobile terminal is greater than or equal to the first preset data volume threshold.

[0088] In practice, the mobile terminal can be connected to the above-mentioned execution entity in a wired or wireless manner. The mobile terminal can have an audio data acquisition function. After the mobile terminal acquires audio data each time, it can send the acquired audio data to the above-mentioned execution entity. After receiving the audio data, if the data volume of the audio data to be converted by the mobile terminal is less than the first preset data volume threshold, then the above-mentioned execution entity can send (i.e., backhaul) the received audio data to the mobile terminal; if the data volume of the audio data to be converted by the mobile terminal is greater than or equal to the first preset data volume threshold, then the above-mentioned execution entity can locally cache the received audio data. The first preset data volume threshold can be used to determine whether the mobile terminal is blocked.

[0089] Optionally, after receiving the audio data, the above-mentioned execution entity can also forward the received audio data to other electronic devices communicatively connected to the execution entity.

[0090] As an example, in some optional implementation manners of this embodiment, the above method is applied to a Bluetooth adapter, that is, the above-mentioned execution entity is a Bluetooth adapter. On this basis, the above-mentioned execution entity can also perform the following steps:

[0091] First, when the above-mentioned mobile terminal receives a Bluetooth connection request sent by a Bluetooth headset, or when the above-mentioned Bluetooth adapter receives a Bluetooth connection request sent by the above-mentioned Bluetooth headset, the above-mentioned execution entity can establish a Bluetooth connection with the above-mentioned Bluetooth headset.

[0092] Among them, the above-mentioned execution entity can detect a Bluetooth connection request sent by the Bluetooth headset through Bluetooth scanning. If the Bluetooth connection request indicates that the Bluetooth headset requests to connect to the Bluetooth adapter, then the Bluetooth headset can directly establish a Bluetooth connection with the Bluetooth adapter; if the Bluetooth connection request indicates that the Bluetooth headset requests to connect to the mobile terminal, then the mobile terminal can send information for instructing the Bluetooth headset to establish a Bluetooth connection with the Bluetooth adapter to the Bluetooth headset, so that the Bluetooth headset can establish a Bluetooth connection with the Bluetooth adapter.

[0093] After that, through the above Bluetooth connection, send the audio data collected by the above-mentioned mobile terminal received from the mobile terminal to the above-mentioned Bluetooth headset.

[0094] It can be understood that in the above optional implementation manner, the Bluetooth connection established between the Bluetooth adapter and the Bluetooth headset is used to transmit audio data, which can avoid information interference caused by the mobile terminal sending audio data to the Bluetooth headset, thereby improving the accuracy of audio data transcription.

[0095] In some alternative implementation manners of this embodiment, the audio data collected by the mobile terminal includes the voice audio of at least two speakers. On this basis, the execution entity may adopt any of the following methods to locally cache the audio data collected by the mobile terminal and received from the mobile terminal:

[0096] First, use the voice audio of a single target speaker in the audio data collected by the mobile terminal and received from the mobile terminal as one piece of audio data for local caching.

[0097] Among them, the voice audio of the target speaker is: the voice audio of the speaker selected from the voice audio of the at least two speakers, or the voice audio of a preset speaker.

[0098] Here, the target speaker can be selected or specified by the user, or can also be filtered based on preset conditions. The preset conditions may include: among the at least two speakers, the speaker with the loudest voice audio, or among the at least two speakers, the speaker with the same voice color as the predetermined speaker.

[0099] As an example, if the mobile terminal collects the audio data of target speaker A and the audio data of target speaker B, and the execution entity receives the audio data of target speaker A and the audio data of target speaker B from the mobile terminal, then the execution entity can use the audio data of target speaker A as one piece of audio data for local caching, and use the audio data of target speaker B as one piece of audio data for local caching. That is, the audio data of target speaker A and the audio data of target speaker B are respectively used as two pieces of audio data for local caching respectively.

[0100] Second, use the voice audio of each speaker in the at least two speakers' voice audio selected from the audio data collected by the mobile terminal and received from the mobile terminal as one piece of audio data for local caching respectively.

[0101] As an example, if the above-mentioned mobile terminal collects the audio data of the target speaker A, the audio data of the target speaker B, and the audio data of the target speaker C, and the above-mentioned execution entity receives the audio data of the target speaker A, the audio data of the target speaker B, and the audio data of the target speaker C from the above-mentioned mobile terminal, then the data identifiers of the above-mentioned respective audio data (used to indicate the audio data of the target speaker A, the audio data of the target speaker B, and the audio data of the target speaker C) can be displayed to the user, so that the user can select at least two data identifiers from the data identifiers of the above-mentioned respective audio data, and further determine the voice audio of the selected at least two speakers. For example, the voice audio of the selected at least two speakers may include: the audio data of the target speaker A and the audio data of the target speaker B. Then, the above-mentioned execution entity can use the audio data of the target speaker A as one piece of audio data for local caching, and use the audio data of the target speaker B as one piece of audio data for local caching. That is, the audio data of the target speaker A and the audio data of the target speaker B are respectively used as two pieces of audio data for separate local caching.

[0102] Thirdly, use the voice audio of at least two speakers selected from the audio data collected by the above-mentioned mobile terminal and received by the above-mentioned execution entity as one piece of audio data for local caching.

[0103] As an example, if the above-mentioned mobile terminal collects the audio data of the target speaker A, the audio data of the target speaker B, and the audio data of the target speaker C, and the above-mentioned execution entity receives the audio data of the target speaker A, the audio data of the target speaker B, and the audio data of the target speaker C from the above-mentioned mobile terminal, then the data identifiers of the above-mentioned respective audio data (used to indicate the audio data of the target speaker A, the audio data of the target speaker B, and the audio data of the target speaker C) can be displayed to the user, so that the user can select at least two data identifiers from the data identifiers of the above-mentioned respective audio data, and further determine the voice audio of the selected at least two speakers. For example, the voice audio of the selected at least two speakers may include: the audio data of the target speaker A and the audio data of the target speaker B. Then, the above-mentioned execution entity can use the audio data of the target speaker A and the audio data of the target speaker B as one piece of audio data for local caching of the audio data of the target speaker A and the audio data of the target speaker B together.

[0104] Among them, different pieces of audio data can be identified or stored separately in different ways. As an example, in the case where the above-mentioned mobile terminal is a mobile phone and the audio data includes uplink data (including the audio data of the user of the mobile phone) and downlink data (including the audio data of the other party) during a call, the above-mentioned execution entity can only use the downlink data as one piece of audio data for local caching, or can use the uplink data and the downlink data as one piece of data respectively and cache them separately locally.

[0105] One piece of audio data can include one frame or multiple frames of audio data. Optionally, the audio data sent by the mobile terminal each time it is packed can also be used as one piece of audio data.

[0106] It can be understood that in the above optional implementation methods, the voice audio of one or more specific speakers among the voice audio of at least two speakers can be used as one piece of audio data, and then it can be locally cached. Thus, through subsequent steps, the above-mentioned execution entity can send a single piece of data to the mobile terminal for it to perform conversion, thereby improving the accuracy of transcription.

[0107] In some application scenarios of the above optional implementation methods, the above-mentioned execution entity can use the following method to send the locally cached audio data to the above-mentioned mobile terminal:

[0108] Send at least one piece of the locally cached audio data to the above-mentioned mobile terminal.

[0109] As an example, the above-mentioned execution entity can send all the locally cached audio data to the above-mentioned mobile terminal in the order of caching time, one by one. Optionally, the above-mentioned execution entity can also send one or more pieces of audio data selected by the user among the locally cached audio data to the above-mentioned mobile terminal.

[0110] It can be understood that in the above optional implementation methods, the audio data is sent to the mobile terminal in units of pieces. For example, the audio data sent to the mobile terminal each time is one piece of audio data. In this way, the mobile terminal can perform the conversion of the audio data in units of pieces, and thus can perform the conversion with reference to the semantic information in a single piece of audio data, thereby further improving the smoothness and accuracy of transcription.

[0111] In some optional implementation methods of this embodiment, the above method is applied to a Bluetooth adapter, that is, the above-mentioned execution entity is a Bluetooth adapter. The Bluetooth adapter establishes a connection with the above-mentioned mobile terminal through a connection port, and the connection port is used for the transmission of the above-mentioned audio data between the Bluetooth adapter and the above-mentioned mobile terminal.

[0112] It can be understood that in the above optional implementation manners, the Bluetooth adapter establishes a connection with the above mobile terminal through a connection port, which can reduce the error rate of audio data transmission and improve the security of audio data transmission.

[0113] Step 202, in the case of receiving an audio data return instruction sent by the above mobile terminal, send the audio data cached locally to the above mobile terminal, so that the above mobile terminal can convert the received audio data into text.

[0114] In this embodiment, in the case of receiving an audio data return instruction sent by the above mobile terminal, the above execution entity can send the audio data cached locally to the above mobile terminal, so that the above mobile terminal can convert the received audio data into text.

[0115] Here, the transcription of audio data can include processes such as recording of audio data and converting audio data into text.

[0116] Among them, the above audio data return instruction is sent by the above mobile terminal when the data volume of the audio data to be converted by the above mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold.

[0117] In some optional implementation manners of this embodiment, in the case of receiving an audio data confirmation instruction sent by the above mobile terminal, the above execution entity can also delete the audio data corresponding to the above audio data confirmation instruction from the local cache.

[0118] Among them, the above audio data confirmation instruction indicates that the above mobile terminal has received the above audio data or the above mobile terminal has completed the conversion of the above audio data.

[0119] The audio data corresponding to the above audio data confirmation instruction can be the audio data that the above mobile terminal has received as indicated by the audio data confirmation instruction, or it can also be the audio data that the above mobile terminal has completed the conversion as indicated by the audio data confirmation instruction.

[0120] For example, if the above execution entity locally caches audio data 1, audio data 2, and audio data 3. In this case, if the audio data confirmation instruction sent by the above mobile terminal indicates the deletion of locally cached audio data 1 (that is, the audio data corresponding to the audio data confirmation instruction is audio data 1), then the above execution entity can delete audio data 1 from the local cache.

[0121] It can be understood that after ensuring that the mobile terminal has received the above audio data or the mobile terminal has completed the conversion of the above audio data, the above execution entity can delete the audio data cached locally, thereby improving the security of audio data storage, saving the storage space of the above execution entity, and improving the conversion success rate.

[0122] In some optional implementation manners of this embodiment, the mobile terminal has no conversion permission for the audio data collected by the mobile terminal.

[0123] Here, generally, for reasons such as information security, the mobile terminal may have no conversion permission for the audio data collected by itself. For example, a mobile phone itself cannot convert audio data such as call records collected by the mobile phone into text.

[0124] It can be understood that in some application scenarios, the mobile terminal has no conversion permission for the audio data it collects. In the above optional implementation manner, the audio data obtained from the mobile terminal can be transmitted back to the mobile terminal in a way of backhaul, so as to realize the conversion of the audio data by the mobile terminal.

[0125] Continue to refer to Figure 3 , Figure 3 is a schematic diagram of an application scenario of the audio data transcription method according to this embodiment. In the Figure 3 application scenario, when receiving the audio data caching instruction sent by the mobile terminal 320, the Bluetooth adapter 310 locally caches the audio data received from the mobile terminal 320 and collected by the mobile terminal 320. Among them, the above audio data caching instruction is sent by the above mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to the first preset data volume threshold. After that, when receiving the audio data return instruction sent by the mobile terminal 320, the Bluetooth adapter 310 sends the locally cached audio data to the mobile terminal 320, so that the mobile terminal 320 can convert the received audio data into text. Among them, the above audio data return instruction is sent by the mobile terminal 320 when the data volume of the audio data to be converted by the mobile terminal is less than or equal to the second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold.

[0126] The method provided in the above embodiments of the present disclosure, when receiving an audio data caching instruction sent by a mobile terminal, locally caches the audio data received from the mobile terminal and collected by the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold. After that, when receiving an audio data return instruction sent by the mobile terminal, the locally cached audio data is sent to the mobile terminal so that the mobile terminal can convert the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold. Thus, the embodiments of the present disclosure achieve the conversion of audio data by the mobile terminal in the case of blocking due to a large amount of information by caching the audio data received from the mobile terminal and collected by the mobile terminal, improving the smoothness and accuracy of audio data transcription.

[0127] Further referring to Figure 4A , Figure 4A is a schematic diagram of an interaction process of a method for transcribing audio data provided by an embodiment of the present disclosure.

[0128] In Figure 4A shown in step 401, the mobile terminal sends a recording start instruction to the Bluetooth adapter.

[0129] In some application scenarios, the Bluetooth adapter, the mobile terminal, and the Bluetooth headset can be connected. Among them, the mobile terminal can be used to collect audio data, and the user can use the Bluetooth headset to listen to the audio corresponding to the audio data. As an example, the connection method among the Bluetooth adapter, the mobile terminal, and the Bluetooth headset can be as Figure 4B shown.

[0130] Exemplarily, the Bluetooth adapter can be connected to the mobile terminal through USB (Universal Serial Bus), and the mobile terminal can send a recording start instruction to the Bluetooth adapter through USB.

[0131] In step 402, the mobile terminal sends audio data to the Bluetooth adapter.

[0132] Here, the mobile terminal can send audio data to the Bluetooth adapter through USB.

[0133] In addition, the audio data sent by the mobile terminal to the Bluetooth adapter can only include unidirectional data (such as downstream data), such as the audio data of the other party obtained by the mobile terminal during a call.

[0134] Optionally, the audio data sent by the mobile terminal to the Bluetooth adapter may also include bidirectional data, such as downlink data and uplink data.

[0135] In step 403, the mobile terminal sends an audio data caching instruction to the Bluetooth adapter.

[0136] Here, if the data volume of the audio data to be converted by the mobile terminal is greater than or equal to the first preset data volume threshold, then the mobile terminal may send an audio data caching instruction to the Bluetooth adapter via USB to instruct the Bluetooth adapter to cache the audio data.

[0137] It should be noted that in some cases, step 403 may be executed first, and then step 402 may be executed.

[0138] In step 404, the Bluetooth adapter performs local caching of the audio data.

[0139] Here, if the mobile terminal is blocked due to a large amount of information, it may send an audio data caching instruction to the Bluetooth adapter to notify the Bluetooth adapter to perform caching. The Bluetooth adapter packs and caches the audio data (such as downlink data), and waits for the mobile terminal to become unblocked before sending an audio data return instruction to the Bluetooth adapter.

[0140] In step 405, the mobile terminal sends an audio data return instruction to the Bluetooth adapter.

[0141] Here, if the data volume of the audio data to be converted by the mobile terminal is less than or equal to the second preset data volume threshold, then the mobile terminal may send an audio data return instruction to the Bluetooth adapter to enable the Bluetooth adapter to send the locally cached audio data to the mobile terminal. Among them, the second preset data volume threshold is less than the first preset data volume threshold.

[0142] In step 406, the Bluetooth adapter sends the locally cached audio data to the mobile terminal.

[0143] Here, the Bluetooth adapter may pack and compress the locally cached audio data (such as the downlink data of the mobile terminal) and forward it to the mobile terminal via USB.

[0144] Optionally, the Bluetooth adapter may also pack the locally cached audio data (such as the downlink data of the mobile terminal) into two paths. One path is distributed to other Bluetooth devices (such as the above-mentioned Bluetooth headset) connected to the Bluetooth adapter, and the other path of data is packed and compressed and forwarded to the mobile terminal via USB.

[0145] Optionally, if the locally cached audio data of the Bluetooth adapter contains uplink data, then the uplink data may also be sent to the mobile terminal.

[0146] In step 407, the mobile terminal converts the received audio data into text.

[0147] Here, after receiving the audio data, the mobile terminal decompresses it, sends a confirmation command to the Bluetooth adapter, and parses the audio data in real time to convert the audio data into text.

[0148] Currently, the development of online teaching is becoming more and more popular, and the scope of teaching is wider. More comprehensive teaching content needs to be recorded during the learning process; in a voice conference or during a voice presentation, participants need to record the content of the conference or presentation; when watching or listening to video and audio in daily life, some text records need to be made for the video and audio. In view of the above actual scenarios in daily life where the unidirectional audio frequency is relatively high, if the device can record the unidirectional audio and convert it into text, it can bring great convenience to users.

[0149] Under the existing technical background, if you want to meet the above requirements and convert the voice of the other party during a call, you need to separately expand the sound and use a voice recorder to record the sound, and then convert it into text. However, since the voice recorder is in an open environment, not only the voice of the other party during the call emitted from the earphone or speaker can be recorded, but also the user's own voice can be recorded. Therefore, it is easily interfered by external sounds, affecting the accuracy and reliability of transcription.

[0150] In this embodiment, the transcription method of audio data can forward the audio data collected by the mobile terminal through the Bluetooth adapter. Since the Bluetooth adapter can directly send the data of the other party during the call back to the mobile terminal, the mobile terminal can thus obtain only the data of the other party during the call, which is independently separated from the user's voice data, thereby improving the accuracy of transcription. And by caching the audio data collected by the mobile terminal received from the mobile terminal, the mobile terminal can convert the audio data in the case of information congestion, improving the smoothness of audio data transcription.

[0151] In addition, the transcription method of the audio data of the present disclosure may further include: when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to the first preset data volume threshold, sending an audio data caching instruction of the locally collected audio data to the target device, so that the target device caches the locally collected audio data; when the data volume of the audio data to be converted by the mobile terminal is less than or equal to the second preset data volume threshold, sending an audio data return instruction to the target device, and receiving the audio data corresponding to the audio data return instruction sent by the target device; converting the received audio data into text, where the second preset data volume threshold is less than the first preset data volume threshold.

[0152] In some alternative implementation manners of this embodiment, the collected audio data includes the voice audios of at least two speakers; and

[0153] The above target device caches the locally collected audio data in any of the following manners:

[0154] Taking the voice audio of a single target speaker in the audio data collected by the above mobile terminal received from the above mobile terminal as an audio data for local caching; wherein, the voice audio of the above target speaker is: the voice audio of the speaker selected from the voice audios of the above at least two speakers, or the voice audio of a preset speaker;

[0155] Taking the voice audio of each speaker in the at least two speakers' voice audios selected from the audio data collected by the above mobile terminal received from the above mobile terminal as an audio data for local caching respectively;

[0156] Taking the at least two speakers' voice audios selected from the audio data collected by the above mobile terminal received from the above mobile terminal as an audio data for local caching.

[0157] In some alternative implementation manners of this embodiment, after sending the audio data caching instruction of the locally collected audio data to the target device, the target device deletes the audio data corresponding to the audio data confirmation instruction from the cache, wherein the audio data confirmation instruction indicates that the above mobile terminal has received the above audio data or the above mobile terminal has completed the conversion of the above audio data.

[0158] In some alternative implementation manners of this embodiment, the above method is applied to a mobile terminal. When the mobile terminal receives a Bluetooth connection request from a Bluetooth headset, or when the above target device receives the Bluetooth connection request from the above Bluetooth headset, the above target device establishes a Bluetooth connection with the above Bluetooth headset; and, the above method further includes:

[0159] Sending the audio data collected by the above mobile terminal to the above Bluetooth headset via the above target device through the above Bluetooth connection.

[0160] It should be noted that the execution manners and the effects generated by the above-described steps of the above embodiment can be referred to the relevant descriptions of the above Figure 2 、 Figure 3 、 Figure 4A and Figure 4B . For the sake of brief description, no further details are given here.

[0161] For further reference Figure 5, as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a transcription device for audio data. This embodiment of the transcription device for audio data corresponds to the method embodiments described above. Except for the features described below, this embodiment of the transcription device for audio data may also include the same or corresponding features as the method embodiments described above, and produce the same or corresponding effects as the method embodiments described above.

[0162] As Figure 5 shown, the transcription device 500 for audio data in this embodiment. The above device 500 includes: a buffer unit 501 and a first sending unit 502. Among them, the buffer unit 501 is configured to locally cache the audio data collected by the mobile terminal received from the mobile terminal when receiving an audio data caching instruction sent by the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold; the first sending unit 502 is configured to send the locally cached audio data to the mobile terminal when receiving an audio data return instruction sent by the mobile terminal, so that the mobile terminal converts the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold.

[0163] In this embodiment, when receiving an audio data caching instruction sent by the mobile terminal, the buffer unit 501 of the transcription device 500 for audio data can locally cache the audio data collected by the mobile terminal received from the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold.

[0164] In this embodiment, when receiving the audio data return instruction sent by the mobile terminal, the first sending unit 502 can send the locally cached audio data to the mobile terminal, so that the mobile terminal converts the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold.

[0165] In some optional implementation manners of this embodiment, the audio data collected by the mobile terminal includes the voice audio of at least two speakers; and

[0166] The above cache unit includes any one of the following:

[0167] A first cache subunit (not shown in the figure), configured to locally cache, as one path of audio data, the speech audio of a single target speaker in the audio data collected by the above mobile terminal and received from the above mobile terminal; wherein, the speech audio of the above target speaker is: the speech audio of the speaker selected from the speech audio of the above at least two speakers, or the speech audio of a preset speaker;

[0168] A second cache subunit (not shown in the figure), configured to locally cache, as one path of audio data respectively, the speech audio of each speaker in the speech audio of at least two speakers selected from the audio data collected by the above mobile terminal and received from the above mobile terminal;

[0169] A third cache subunit (not shown in the figure), configured to locally cache, as one path of audio data, the speech audio of at least two speakers selected from the audio data collected by the above mobile terminal and received from the above mobile terminal.

[0170] In some optional implementation manners of this embodiment, the above first sending unit 502 includes:

[0171] A sending subunit (not shown in the figure), configured to send at least one path of audio data in the locally cached audio data to the above mobile terminal.

[0172] In some optional implementation manners of this embodiment, the above device 500 further includes:

[0173] A deletion unit (not shown in the figure), configured to delete, from the local cache, the audio data corresponding to the above audio data confirmation instruction when receiving the audio data confirmation instruction sent by the above mobile terminal, wherein the above audio data confirmation instruction indicates that the above mobile terminal has received the above audio data or the above mobile terminal has completed the conversion of the above audio data.

[0174] In some optional implementation manners of this embodiment,

[0175] The above device 500 further includes:

[0176] A connection establishment unit (not shown in the figure), configured to establish a Bluetooth connection with the above Bluetooth headset when the above mobile terminal receives a Bluetooth connection request from the above Bluetooth headset, or when the above Bluetooth adapter receives a Bluetooth connection request from the above Bluetooth headset;

[0177] A second sending unit (not shown in the figure) is configured to send, through the above Bluetooth connection, the audio data collected by the mobile terminal and received from the mobile terminal to the above Bluetooth headset.

[0178] In some optional implementation manners of this embodiment, the above device is disposed in a Bluetooth adapter, and the Bluetooth adapter establishes a connection with the mobile terminal through a connection port, where the connection port is used for transmitting the above audio data between the Bluetooth adapter and the mobile terminal.

[0179] In some optional implementation manners of this embodiment, the mobile terminal has no conversion permission for the audio data collected by the mobile terminal.

[0180] In the device 500 provided in the above embodiment of the present disclosure, the cache unit 501 can locally cache the audio data collected by the mobile terminal and received from the mobile terminal when receiving an audio data cache instruction sent by the mobile terminal, where the audio data cache instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold; the first sending unit 502 can send the locally cached audio data to the mobile terminal when receiving an audio data return instruction sent by the mobile terminal, so that the mobile terminal converts the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold. Thus, the embodiment of the present disclosure caches the audio data collected by the mobile terminal and received from the mobile terminal, realizes the conversion of the audio data by the mobile terminal in case of blockage due to a large amount of information, and improves the smoothness and accuracy of audio data transcription.

[0181] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Figure 6 The shown electronic device 600 includes: at least one processor 601, a memory 602, at least one network interface 604, and other user interfaces 603. Each component in the electronic device 600 is coupled together through a bus system 605. It can be understood that the bus system 605 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 605 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 6 all kinds of buses are labeled as the bus system 605.

[0182] Among them, the user interface 603 may include a display, a keyboard, or a pointing device (such as a mouse, a trackball, a touchpad, or a touch screen, etc.).

[0183] It can be understood that the memory 602 in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 602 described herein is intended to include but not be limited to these and any other suitable types of memory.

[0184] In some embodiments, the memory 602 stores the following elements, executable units, or data structures, or subsets thereof, or extended sets thereof: an operating system 6021 and application programs 6022.

[0185] Among them, the operating system 6021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs 6022 include various application programs, such as a media player and a browser, etc., for implementing various application services. The program for implementing the method of the embodiments of the present disclosure may be included in the application programs 6022.

[0186] In the embodiments of the present disclosure, by invoking the programs or instructions stored in the memory 602, specifically, the programs or instructions stored in the application program 6022, the processor 601 is configured to execute the method steps provided in each method embodiment, for example, including: when receiving an audio data caching instruction sent by the mobile terminal, locally caching the audio data collected by the mobile terminal and received from the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold; when receiving an audio data return instruction sent by the mobile terminal, sending the locally cached audio data to the mobile terminal so that the mobile terminal can convert the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold. Or, when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to a first preset data volume threshold, sending an audio data caching instruction of the locally collected audio data to the target device to enable the target device to cache the locally collected audio data; when the data volume of the audio data to be converted by the mobile terminal is less than or equal to a second preset data volume threshold, sending an audio data return instruction to the target device, and receiving the audio data corresponding to the audio data return instruction sent by the target device; converting the received audio data into text, where the second preset data volume threshold is less than the first preset data volume threshold.

[0187] The method disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 601. The processor 601 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 601 or instructions in software form. The above-mentioned processor 601 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or completed by a combination of hardware and software units in the decoding processor. The software units may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory 602, and the processor 601 reads the information in the memory 602 and combines its hardware to complete the steps of the above method.

[0188] It can be understood that these embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or a combination thereof.

[0189] For software implementation, the techniques described herein can be implemented by units that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented within the processor or external to the processor.

[0190] The electronic device provided in this embodiment may be the electronic device shown in Figure 6 and can execute all steps of the audio data transcription method shown in Figure 2 , thereby achieving the technical effects of the audio data transcription method shown in Figure 2 . For specific details, please refer to Figure 2 for relevant descriptions. For the sake of brevity, it will not be elaborated here.

[0191] The embodiments of the present disclosure also provide a storage medium (computer-readable storage medium). The storage medium stores one or more programs. Among them, the storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk or solid-state drive; the memory may also include a combination of the above types of memory.

[0192] When one or more programs in the storage medium can be executed by one or more processors to implement the above-mentioned audio data transcription method executed on the electronic device side.

[0193] The processor is used to execute the communication program stored in the memory to implement the following steps of the audio data transcription method executed on the electronic device side: when receiving an audio data caching instruction sent by the mobile terminal, locally cache the audio data received from the mobile terminal and collected by the mobile terminal, where the audio data caching instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to the first preset data volume threshold; when receiving the audio data return instruction sent by the mobile terminal, send the locally cached audio data to the mobile terminal so that the mobile terminal can convert the received audio data into text, where the audio data return instruction is sent by the mobile terminal when the data volume of the audio data to be converted by the mobile terminal is less than or equal to the second preset data volume threshold, and the second preset data volume threshold is less than the first preset data volume threshold. Or, when the data volume of the audio data to be converted by the mobile terminal is greater than or equal to the first preset data volume threshold, send an audio data caching instruction of the locally collected audio data to the target device to enable the target device to cache the locally collected audio data; when the data volume of the audio data to be converted by the mobile terminal is less than or equal to the second preset data volume threshold, send an audio data return instruction to the target device, and receive the audio data corresponding to the audio data return instruction sent by the target device; convert the received audio data into text, where the second preset data volume threshold is less than the first preset data volume threshold.

[0194] Those skilled in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0195] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0196] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of this disclosure. It should be understood that the above description is only for the specific embodiments of this disclosure and is not used to limit the protection scope of this disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for converting audio data, characterized in that: The method is applied to a Bluetooth adapter, and the method comprises: Upon receiving an audio data caching instruction sent by a mobile terminal, locally caching the audio data received from the mobile terminal and collected by the mobile terminal, wherein the audio data caching instruction is sent by the mobile terminal when the amount of audio data to be converted by the mobile terminal is greater than or equal to a first preset data amount threshold, the mobile terminal has no conversion authority for the audio data collected by the mobile terminal, and the mobile terminal has conversion authority for the audio data transmitted back to the mobile terminal; When an audio data return instruction is received from the mobile terminal, locally cached audio data is sent to the mobile terminal so that the mobile terminal converts the received audio data into text, wherein the audio data return instruction is sent via the mobile terminal when the amount of audio data to be converted at the mobile terminal is less than or equal to a second preset data amount threshold, and the second preset data amount threshold is less than the first preset data amount threshold.

2. The method according to claim 1, characterized in that The audio data collected by the mobile terminal includes voice audio of at least two speakers; and The locally caching the audio data received from the mobile terminal and collected by the mobile terminal includes any of the following: The speech audio of a single target speaker in the audio data received from the mobile terminal and collected by the mobile terminal is used as an audio data and cached locally; wherein the speech audio of the target speaker is: the speech audio of the speaker selected from the speech audios of the at least two speakers, or the speech audio of a preset speaker; The speech audio of each speaker in the speech audio of at least two speakers selected from the audio data received from the mobile terminal and collected by the mobile terminal is respectively taken as an audio data and locally cached; The speech audios of at least two speakers selected from the audio data received from the mobile terminal and collected by the mobile terminal are taken as one piece of audio data and cached locally.

3. The method according to claim 2, characterized in that The sending the locally cached audio data to the mobile terminal includes: At least one piece of audio data among the locally cached audio data is sent to the mobile terminal.

4. The method according to claim 1, characterized in that: The method further comprises: Upon receiving an audio data confirmation instruction sent by the mobile terminal, the audio data corresponding to the audio data confirmation instruction is deleted from the local cache, wherein the audio data confirmation instruction indicates that the mobile terminal has received the audio data or the mobile terminal has completed conversion of the audio data.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: When the mobile terminal receives a Bluetooth connection request from the Bluetooth headset, or when the Bluetooth adapter receives a Bluetooth connection request from the Bluetooth headset, establishing a Bluetooth connection with the Bluetooth headset; The audio data received from the mobile terminal and collected by the mobile terminal is sent to the Bluetooth headset through the Bluetooth connection.

6. A method for transcribing audio data, characterized in that: The method is applied to a mobile terminal, and the method comprises: When the amount of audio data to be converted by the mobile terminal is greater than or equal to a first preset data amount threshold, an audio data caching instruction of locally collected audio data is sent to the Bluetooth adapter, so that the Bluetooth adapter caches the locally collected audio data, wherein the mobile terminal has no conversion authority for the audio data collected by the mobile terminal, and the mobile terminal has conversion authority for the audio data transmitted back to the mobile terminal; When the amount of audio data to be converted by the mobile terminal is less than or equal to a second preset data amount threshold, sending an audio data return instruction to the Bluetooth adapter, and receiving audio data corresponding to the audio data return instruction sent by the Bluetooth adapter; The received audio data is converted into text, wherein the second preset data volume threshold is less than the first preset data volume threshold.

7. The method according to claim 6, characterized in that The collected audio data includes speech audio of at least two speakers; and The Bluetooth adapter caches the locally collected audio data in any of the following ways: The speech audio of a single target speaker in the audio data received from the mobile terminal and collected by the mobile terminal is used as an audio data and cached locally; wherein the speech audio of the target speaker is: the speech audio of the speaker selected from the speech audios of the at least two speakers, or the speech audio of a preset speaker; The speech audio of each speaker in the speech audio of at least two speakers selected from the audio data received from the mobile terminal and collected by the mobile terminal is respectively taken as an audio data and locally cached; The speech audios of at least two speakers selected from the audio data received from the mobile terminal and collected by the mobile terminal are taken as one piece of audio data and cached locally.

8. The method according to claim 6 or 7, characterized in that: After the audio data caching instruction of sending the locally collected audio data to the Bluetooth adapter, the Bluetooth adapter deletes the audio data corresponding to the audio data confirmation instruction from the cache, wherein the audio data confirmation instruction indicates that the mobile terminal has received the audio data or the mobile terminal has completed the conversion of the audio data.

9. The method according to claim 6 or 7, characterized in that: The method is applied to a mobile terminal, and when the mobile terminal receives a Bluetooth connection request from a Bluetooth headset, or when the Bluetooth adapter receives a Bluetooth connection request from the Bluetooth headset, the Bluetooth adapter establishes a Bluetooth connection with the Bluetooth headset; as well as The method further comprises: The audio data collected by the mobile terminal is sent to the Bluetooth headset via the Bluetooth adapter through the Bluetooth connection.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute a computer program stored in the memory, and when the computer program is executed, implement the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Play control method and equipment

    CN104867513A

  • Data transmission method and device, electronic device and storage medium

    CN109445741A

  • Large-data-volume audio Bluetooth real-time transmission method for equipment with recording function

    CN111432384A

  • Information processing method and device, electronic equipment and medium

    CN111755008A

  • Recording control method and recording control device

    CN113055529A