Voice data transmission method and device, apparatus, computer program product
By converting voice data into text data and combining it with voice features for transmission, the problem of inflexible voice data transmission in existing technologies is solved, achieving high-quality and efficient voice data transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing voice data transmission methods are inflexible, resulting in poor voice data transmission quality, especially in satellite communication and Bluetooth wireless communication, where there are problems such as insufficient bandwidth and degraded voice quality.
By converting voice data into target text data and transmitting it, the receiving device obtains target voice data based on the target text data and target voice features, effectively reducing the probability of bandwidth limitation and achieving highly flexible voice data transmission.
It improves the flexibility and quality of voice data transmission, reduces the probability of bandwidth limitations, and achieves efficient voice data transmission.
Smart Images

Figure CN122290598A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice data transmission technology, and in particular to a voice data transmission method, device, apparatus, and computer program product. Background Technology
[0002] With the continuous development of smart electronic devices, satellite communication and Bluetooth wireless communication technologies have become popular communication methods. Currently, voice data transmission is particularly popular among the application scenarios of these communication technologies. However, the inflexibility of these voice data transmission methods leads to problems with poor voice data transmission quality. Summary of the Invention
[0003] This application provides a voice data transmission method, device, apparatus, and computer program product, which realizes highly flexible voice data transmission and improves the transmission effect of voice data.
[0004] In a first aspect, a voice data transmission method is provided, applied to a voice data transmission system, the voice data transmission system further including a receiving device. The method includes: acquiring target text data corresponding to voice data; transmitting the target text data to the receiving device, so that the receiving device obtains target voice data based on the target text data and target voice features.
[0005] Secondly, a voice data transmission method is provided, applied to a voice data transmission system, the voice data transmission system further including a transmitting device, the method comprising: receiving target text data from the transmitting device; and obtaining target voice data based on the target text data and target voice features.
[0006] In this application, the transmitting device in the voice data transmission system can send target text data corresponding to the acquired voice data to the receiving device, so that the receiving device can obtain target voice data based on the target text data and target voice features. Since the amount of target text data is much smaller than the amount of audio data, this application effectively reduces the probability of bandwidth limitations. The transmitting device can efficiently transmit high-quality target text data to the receiving device, and correspondingly, the receiving device can obtain high-quality target voice data based on the received target text data. In other words, this application achieves highly flexible voice data transmission and improves the transmission effect of voice data.
[0007] Thirdly, a transmitting device is provided for use in a voice data transmission system. This voice data transmission system further includes a receiving device. The transmitting device comprises an acquisition module and a transmission module. The acquisition module is used to acquire target text data corresponding to the voice data; the transmission module is used to transmit the target text data to the receiving device, so that the receiving device obtains target voice data based on the target text data and target voice features.
[0008] Fourthly, a receiving device is provided for use in a voice data transmission system, the voice data transmission system further including a transmitting device. The receiving device includes a receiving module and a processing module. The receiving module is used to receive target text data from the transmitting device; the processing module is used to obtain target voice data based on the target text data and target voice features.
[0009] Fifthly, another voice data transmission device is provided, including a processor coupled to a memory, which can be used to execute instructions in the memory to implement the method in any of the possible implementations of the first or second aspect described above. Optionally, the voice data transmission device further includes a memory. Optionally, the voice data transmission device further includes a communication interface, to which the processor is coupled.
[0010] A sixth aspect provides a processor, comprising: an input circuit, an output circuit, and a processing circuit. The processing circuit is used to receive signals through the input circuit and transmit signals through the output circuit, causing the processor to execute the method in any possible implementation of the first or second aspect described above.
[0011] In specific implementation, the processor can be a chip, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, gate circuit, flip-flop, and various logic circuits. The input signal received by the input circuit can be received and input by, for example, but not limited to, a receiver, and the signal output by the output circuit can be output to, for example, but not limited to, a transmitter and transmitted by the transmitter. Furthermore, the input circuit and the output circuit can be the same circuit, which is used as the input circuit and the output circuit at different times. This application does not limit the specific implementation of the processor and various circuits.
[0012] In a seventh aspect, a processing apparatus is provided, including a processor and a memory. The processor is configured to read instructions stored in the memory and to receive signals via a receiver and transmit signals via a transmitter to execute the method in any of the possible implementations of the first or second aspect described above.
[0013] Optionally, there may be one or more processors and one or more memories.
[0014] Alternatively, the memory can be integrated with the processor, or the memory can be set up separately from the processor.
[0015] In specific implementation, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same chip or set on different chips. The embodiments of this application do not limit the type of memory or the way the memory and processor are set.
[0016] It should be understood that the relevant data interaction process, such as sending indication information, can be the process of outputting indication information from the processor, and receiving capability information can be the process of the processor receiving input capability information. Specifically, the processed output data can be output to the transmitter, and the input data received by the processor can come from the receiver. Here, the transmitter and receiver can be collectively referred to as a transceiver.
[0017] The processing device in the seventh aspect above can be a chip. The processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. The memory can be integrated into the processor or located outside the processor and exist independently.
[0018] Eighthly, a computer program product is provided, the computer program product comprising: a computer program (also referred to as code or instructions), which, when run, causes a computer to perform the method in any possible implementation of the first or second aspect described above.
[0019] Ninthly, a computer-readable storage medium is provided that stores a computer program (also referred to as code or instructions) that, when executed on a computer, causes the computer to perform the methods in any possible implementation of the first or second aspect described above. Attached Figure Description
[0020] Figure 1 This is a schematic diagram illustrating the application scenario provided in the embodiments of this application;
[0021] Figure 2 This is a schematic diagram of the system architecture of the electronic device provided in the embodiments of this application;
[0022] Figure 3 This is a schematic flowchart illustrating a voice data transmission method provided in an embodiment of this application;
[0023] Figure 4This is a schematic flowchart illustrating a first specific example of the voice data transmission method provided in the embodiments of this application;
[0024] Figure 5 This is a schematic flowchart illustrating a second specific example of the voice data transmission method provided in the embodiments of this application;
[0025] Figure 6 This is a schematic flowchart illustrating a third specific example of the voice data transmission method provided in the embodiments of this application;
[0026] Figure 7 This is a schematic flowchart illustrating a fourth specific example of the voice data transmission method provided in the embodiments of this application;
[0027] Figure 8 This is a schematic block diagram of a transmitting device provided in an embodiment of this application;
[0028] Figure 9 This is a schematic block diagram of a receiving device provided in an embodiment of this application;
[0029] Figure 10 This is a schematic block diagram of a voice data transmission device provided in an embodiment of this application. Detailed Implementation
[0030] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0031] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0032] It should be noted that, in this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0033] Furthermore, "at least one" refers to one or more, while "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0034] To make the objectives and technical solutions of this application clearer and more intuitive, the voice data transmission method, device, apparatus, and computer program product provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0035] With the continuous development of smart electronic devices, satellite communication and Bluetooth wireless communication technologies have become popular communication methods. Currently, voice data transmission is particularly popular among the application scenarios of these communication technologies.
[0036] It should be understood that voice data transmission has high bandwidth requirements, which are mainly affected by the sampling rate, sampling bit depth, and data compression algorithm.
[0037] Sampling rate, commonly used for speech audio, ranges from 8kHz (kilohertz) to 16kHz. A higher sampling rate means clearer sound quality, but also requires more data. For example, an 8kHz sampling rate is suitable for telephone-grade audio, while a 16kHz sampling rate provides superior audio quality.
[0038] Sampling bit depth, commonly 8-bit or 16-bit, the higher the bit depth, the better the voice quality, but the larger the data volume.
[0039] Data compression algorithms: Different compression algorithms have a significant impact on bandwidth requirements. Commonly used codecs such as G.729 and Opus can compress voice data to a lower bit rate while retaining high sound quality.
[0040] For example, the G.729 codec has a bit rate of approximately 8 kbps (kilobits per second), while Opus can adjust its bit rate between 6 kbps and 16 kbps for voice data transmission. Taking typical voice call data as an example, the commonly used sampling rate is 8 kHz, the bit depth is 16 bits, and it uses mono with AMR compression (compression ratio approximately 1 / 16). Therefore, the required data per second = sampling rate (8,000 Hz) × bit depth (16 bits) / 16 = 8 kbps.
[0041] Figure 1 This is a schematic diagram illustrating application scenario 100 provided in an embodiment of this application. For example... Figure 1 As shown, application scenario 100 includes a transmitting device 101 and a receiving device 102. The transmitting device 101 and the receiving device 102 can establish a connection using the aforementioned Bluetooth wireless communication technology and transmit voice data. This Bluetooth wireless communication technology can establish point-to-point communication within a small range and is increasingly widely used in scenarios requiring instant communication, such as outdoor exploration and rescue. Alternatively, the transmitting device 101 and the receiving device 102 can also establish a connection using satellite communication technology and transmit voice data. This satellite communication technology relies on satellite signals and can achieve long-distance communication even in remote areas without network coverage.
[0042] However, the above-mentioned voice data transmission methods are inflexible, resulting in poor voice data transmission quality.
[0043] For example, satellite communication has limited frequency band resources, making it difficult to meet the needs of large-scale, high-bandwidth voice communication. Real-time voice data transmission requires stable bandwidth, but satellite channels often face bandwidth competition and congestion problems, which can easily lead to a decline in voice quality.
[0044] For example, although Bluetooth wireless communication technology improves the transmission rate of voice data, it is still difficult to cope with the continuous transmission of high-quality voice data, especially in complex scenarios, where there may be insufficient bandwidth, resulting in low clarity of the transmitted voice data.
[0045] This application provides a voice data transmission method, device, apparatus, and computer program product. In the voice data transmission system, a transmitting device can send target text data corresponding to acquired voice data to a receiving device, enabling the receiving device to obtain target voice data based on the target text data and target voice features. Since the amount of target text data is much smaller than the amount of audio data, this application effectively reduces the probability of bandwidth limitations. The transmitting device can efficiently transmit high-quality target text data to the receiving device, and correspondingly, the receiving device can obtain high-quality target voice data based on the received target text data. In other words, this application achieves highly flexible voice data transmission and improves the transmission effect of voice data.
[0046] The electronic devices involved in the embodiments of this application may be mobile phones, watches, laptops, handheld computers, mobile internet devices (MIDs), personal computers (PCs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in self-driving vehicles, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, personal digital assistants (PDAs), etc., but the embodiments of this application are not limited to these.
[0047] For example, Figure 2 This is a schematic diagram of the system architecture of an electronic device provided in an embodiment of this application.
[0048] like Figure 2 As shown, the electronic device includes a processor 210, a transceiver 220, and a display unit 270. The display unit 270 may include a display screen.
[0049] Optionally, the electronic device may also include a memory 230. The processor 210, transceiver 220, and memory 230 can communicate with each other via internal connections to transfer data. The memory 230 stores computer programs, and the processor 210 retrieves and runs the computer programs from the memory 230. The processor 210 and memory 230 can be combined into a single processing device, but more commonly they are independent components. The processor 210 executes the program code stored in the memory 230 to achieve the aforementioned functions. In specific implementations, the memory 230 can be integrated into the processor 210, or it can be independent of the processor 210.
[0050] In addition, to further enhance the functionality of the electronic device, it may also include one or more of an input unit 260, an audio circuit 280, and a sensor 201.
[0051] Optionally, the above-mentioned electronic device may also include a power supply 250 for providing power to various devices or circuits in the electronic device.
[0052] Understandable, Figure 2 The operation and / or function of each module in the illustrated electronic device are respectively for implementing the corresponding processes in the following method embodiments. For details, please refer to the descriptions in the following method embodiments; detailed descriptions are omitted here to avoid repetition.
[0053] Understandable, Figure 2 The processor 210 in the illustrated electronic device may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0054] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0055] Understandable, Figure 2The power supply 250 shown provides power to the processor 210, memory 230, display unit 270, input unit 260, and transceiver 220. The transceiver 220 can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The transceiver 220 can be one or more devices integrating at least one communication processing module. The display unit 270 is used to display images, videos, etc. The display unit 270 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini-LED, a micro-LED, a quantum dot light-emitting diode (QLED), etc. The memory 230 can be used to store computer executable program code, which includes instructions. The memory 230 can include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc. The data storage area can store data created during the use of the electronic device, etc. Furthermore, the memory 230 can include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 210 executes various functional applications and data processing of the electronic device by running instructions stored in the memory 230 and / or instructions stored in the memory configured within the processor. The electronic device can implement audio functions, such as music playback and recording, through the audio circuit 280 and the application processor.
[0056] Figure 3 This is a schematic flowchart illustrating a voice data transmission method 300 provided in an embodiment of this application. This method 300 can be applied to... Figure 1 The application scenario 300 is shown. Besides this, the method 300 can also be applied to other application scenarios, which are not limited in this application. For example... Figure 3 As shown, the method 300 may include the following steps:
[0057] S301, The transmitting device acquires the target text data corresponding to the voice data.
[0058] The transmitting device can convert voice data into text data and process the text data to obtain the aforementioned target text data.
[0059] Speech recognition technology can convert speech data into text information, i.e., the text data mentioned above. This process not only significantly reduces the amount of data transmitted but also improves data transmission efficiency. For example, the speech data mentioned above may require tens of kilobytes (KB) of storage space, while the same text data only requires hundreds of bytes. This conversion is extremely advantageous in bandwidth-constrained scenarios.
[0060] In one embodiment, the transmitting device can convert the aforementioned voice data into text data using a speech-to-text application.
[0061] In other embodiments, the transmitting device may also convert the aforementioned voice data into text data using a speech recognition model.
[0062] For example, in the process of converting the above-mentioned voice data into text data, the voice data can be collected through a microphone, and then converted into text data through the above-mentioned voice data to text application or voice recognition model.
[0063] The converted text data can be further processed to obtain target text data, so as to facilitate efficient transmission and ensure semantic integrity.
[0064] In some embodiments, the transmitting device may use techniques such as variable-length encoding or symbol encoding to compress the text data to obtain the target text data.
[0065] Variable-length coding is an encoding method in which different characters are represented using codewords of different lengths. Its core idea is to assign short codewords to high-frequency characters and long codewords to low-frequency characters, thereby achieving a better compression ratio than fixed-length coding. Prefix coding is a typical example of variable-length coding; in this encoding method, no codeword is a prefix of any other codeword, which simplifies the decoding process.
[0066] Symbol encoding typically refers to assigning specific encoding values to various symbols. For example, Unicode assigns a unique numerical code to each character in various languages around the world, including letters, numbers, symbols, punctuation marks, and Chinese characters and Arabic letters in various languages. Encoding tables for special characters include mathematical formula symbols, geometric symbols, weather symbols, currency symbols, etc., and each symbol has a corresponding encoding form such as UNICODE, HEX CODE, and HTML CODE.
[0067] It should be understood that the encoding methods shown above are merely exemplary. In addition, other encoding methods can be used to obtain the target text data, and this application does not limit this.
[0068] In some embodiments, the transmitting device may also add specific contextual markers or semantic auxiliary data to the text data to obtain the target text data.
[0069] Contextual markers can be words or phrases in a text that help explain the relationships between utterances and between utterances and context. For example, certain discourse markers (such as “because,” “but,” and “therefore”) can indicate logical relationships between sentences.
[0070] Semantic auxiliary data includes information that enhances the understanding of text meaning, such as the semantic roles of words, the grammatical structure of sentences, and the relationships between entities and concepts in discourse. By integrating this data, it is possible to better predict and understand words and phrases in text.
[0071] For example, converting one minute of speech into 300 words of text results in an average of 2 bytes per character:
[0072] The amount of text data in one minute (i.e., the amount of text data mentioned above) = 300 × 2 bytes = 600 bytes.
[0073] For example, the above text data processing operation increased the data volume by 40%, and the final data volume per minute (i.e. the data volume of the target text data) can be: 600 bytes × 1.4 = 840 bytes.
[0074] Therefore, the bandwidth requirement for transmitting one minute of target text data is 0.84 * 8 / 60 kbps = 0.112 kbps (kilobits per second), which is far less than the bandwidth requirement of 8 kbps for traditional voice data transmission, effectively reducing the amount of data transmitted.
[0075] S302, the transmitting device transmits target text data to the receiving device. Correspondingly, the receiving device receives the target text data from the transmitting device.
[0076] In some embodiments, the sending device may transmit the target text data to the receiving device via Bluetooth wireless communication technology.
[0077] For example, the sending device can search for nearby Bluetooth devices, such as the receiving device, and then establish a connection with the receiving device to transmit the target text data to the receiving device through the connection.
[0078] In some embodiments, the transmitting device may also transmit the aforementioned target text data to the receiving device via satellite communication technology.
[0079] For example, the transmitting device can transmit the target text data to a satellite so that the satellite can forward the target text data to the receiving device.
[0080] It should be understood that the communication methods shown above are merely exemplary. In addition, other communication methods may be adopted depending on the specific scenario of voice data transmission, and this application does not limit them.
[0081] S303, the receiving device obtains target speech data based on target text data and target speech features.
[0082] Corresponding to S301 above, after receiving the target text data, the receiving end can decode the target text data to obtain the text data corresponding to the voice data, and use the target voice features and the text data to obtain the target voice data, that is, restore the above voice data.
[0083] For example, the target speech features may include at least one of voiceprint, intonation, speech rate, and timbre.
[0084] In some embodiments, target speech features and text data can be combined with an input speech model (such as Tacotron, WaveNet, etc.) to output target speech data.
[0085] After obtaining the target voice data, the receiving device can save the target voice data and / or play the target voice data to achieve the effect of voice data transmission between devices.
[0086] In this application, the transmitting device can send the target text data corresponding to the acquired voice data to the receiving device, so that the receiving device can obtain the target voice data based on the target text data and target voice features. Since the amount of the target text data is much smaller than the amount of the audio data, this application effectively reduces the probability of bandwidth limitations. The transmitting device can efficiently transmit high-quality target text data to the receiving device, and correspondingly, the receiving device can obtain high-quality target voice data based on the received target text data. In other words, this application achieves highly flexible voice data transmission and improves the transmission effect of voice data.
[0087] In some embodiments, the transmitting device may transmit the target speech features of the aforementioned speech data to the receiving device, so that the receiving device can obtain the target speech data based on the target speech features and target file data of the speech data.
[0088] Figure 4 This is a schematic flowchart of a voice data transmission method 400 provided in an embodiment of this application. Figure 4 As shown, the method 400 may include the following steps:
[0089] S401, The transmitting device acquires the target text data and target speech features corresponding to the voice data.
[0090] The acquisition of target text data can be found in the detailed description of the above embodiments, and will not be repeated here to avoid repetition.
[0091] The transmitting device can collect the speech features of the aforementioned speech data and encode the speech features to obtain the aforementioned target speech features.
[0092] S402, the transmitting device transmits target text data and target speech features to the receiving device. Correspondingly, the receiving device receives the target text data and target speech features from the transmitting device.
[0093] In some embodiments, the transmitting device may transmit the aforementioned target text data and target voice features to the receiving device via Bluetooth wireless communication technology.
[0094] In some embodiments, the transmitting device may also transmit the aforementioned target text data and target speech features to the receiving device via satellite communication technology.
[0095] It should be understood that the communication methods shown above are merely exemplary. In addition, other communication methods may be adopted depending on the specific scenario of voice data transmission, and this application does not limit them.
[0096] Optionally, the target file data and the target speech features may be transmitted to the receiving device simultaneously or separately; this application does not limit this.
[0097] In one possible implementation, the transmitting device can determine whether to transmit the target speech features and the target text data to the receiving device simultaneously, based on transmission status information.
[0098] For example, the transmitting device can simultaneously transmit the target file data and target voice features to the receiving device when the transmission bandwidth is sufficient.
[0099] For example, the transmitting device may also transmit the target file data and target voice features to the receiving device in sequence when the transmission bandwidth is insufficient.
[0100] In another possible implementation, the transmitting device may also determine whether to transmit the target speech features and the target text data to the receiving device simultaneously, based on the type of audio data.
[0101] For example, if the aforementioned voice data is real-time, the transmitting device can simultaneously transmit the aforementioned target voice features and target text data to the receiving device.
[0102] For example, if the aforementioned voice data is non-real-time, the transmitting device may also transmit the aforementioned target voice features and target text data to the receiving device in sequence.
[0103] Optionally, the transmitting device may also combine the aforementioned transmission status information and the type of voice data to determine whether to simultaneously transmit the target voice features and the target text data to the receiving device.
[0104] S403, the receiving device obtains target speech data based on target text data and target speech features.
[0105] The receiving device can decode the target text data to obtain the text data corresponding to the aforementioned speech data, decode the target speech features to obtain the speech features corresponding to the aforementioned speech data, and obtain the target speech data based on the speech features and the text data.
[0106] For example, converting one minute of speech into 300 words of text results in an average of 2 bytes per character:
[0107] The amount of text data per minute (i.e., the amount of text data mentioned above) = 300 × 2 bytes = 600 bytes.
[0108] For example, the above text data processing operation increased the data volume by 40%, and the final data volume per minute (i.e. the data volume of the target text data) can be: 600 bytes × 1.4 = 840 bytes.
[0109] To recreate the personalized characteristics of the speech corresponding to the above speech data, it is assumed that approximately 1-5 kB of speech feature data is required per minute. For simplified calculation, we consider 5 kB: the amount of speech feature data per minute = 5 kB / min.
[0110] Therefore, the bandwidth requirement for transmitting one minute of target text data and target speech features is (0.84+5)*8 / 60kbps = 0.79kbps, which is only 1 / 10 of the bandwidth requirement of traditional speech transmission of 8kbps, effectively reducing the amount of data transmitted.
[0111] In this embodiment of the application, the transmitting device in the voice data transmission system can send the target text data and target voice features corresponding to the acquired voice data to the receiving device, so that the receiving device can obtain the target voice data based on the target voice features and target text data. Wherein, the amount of the target text data is much smaller than the amount of the audio data, effectively reducing the probability of bandwidth limitations, and in order to enable the receiving device to obtain target voice data with personalized features (such as voiceprint, intonation, etc.) of the voice data, the transmitting device can transmit the target voice features of the voice data together with the target text data to the receiving device. The receiving device can efficiently obtain high-quality target voice data with the closest voice effect to the aforementioned voice data based on the target voice features and target text data from the transmitting device, thereby achieving more flexible voice data transmission and further improving the transmission effect of voice data.
[0112] It should be understood that, in addition to transmitting target voice features to the receiving device as shown above, to further reduce the amount of data transmitted, the transmitting device in this application can also transmit voice data while only transmitting target text data. See the following embodiments for specific details.
[0113] In some embodiments, the receiving device locally stores a speech feature library, which may include at least one speech feature. The receiving device can determine target speech features based on the speech feature library, and based on the target speech features, the receiving device obtains target speech data based on target text data and the target speech features.
[0114] Figure 5 This is a schematic flowchart of a voice data transmission method 500 provided in an embodiment of this application. Figure 5 As shown, the method 500 may include the following steps:
[0115] S501, the transmitting device acquires the target text data corresponding to the voice data.
[0116] The specific details and parameters are described in the above embodiments, and will not be repeated here to avoid repetition.
[0117] S502, the transmitting device transmits target text data to the receiving device. Correspondingly, the receiving device receives the target text data from the transmitting device.
[0118] Similarly, the sending device can transmit the target text data to the receiving device via Bluetooth wireless communication technology or satellite communication technology.
[0119] S503, the receiving device determines the target speech features based on the speech feature library, which may include at least one speech feature.
[0120] In one possible implementation, the receiving device can identify at least one speech feature from the aforementioned speech feature library as the target speech feature.
[0121] Table 1 is a speech feature library provided in this application.
[0122] Table 1
[0123] speech features Speech Feature 1 Speech Feature 2
[0124] As shown in Table 1, the speech feature library includes at least one speech feature, such as speech feature 1 and speech feature 2.
[0125] Upon receiving the aforementioned target text data, the receiving device can select at least one speech feature from its local speech feature library as the target speech feature, thereby efficiently obtaining the target speech data based on the target speech feature and the target text data. By performing the target speech data acquisition operation based on different target speech features, target speech data with different speech effects can be obtained.
[0126] For example, the transmitting device can determine the speech feature 1 in Table 1 as the target speech feature, so as to obtain the target speech data of the first speech effect based on the speech feature 1 and the target text data.
[0127] For example, the transmitting device can determine the speech feature 2 in Table 1 as the target speech feature, so as to obtain the target speech data of the second speech effect based on the speech feature 2 and the target text data.
[0128] For example, the transmitting device can determine speech feature 1 and speech feature 2 in Table 1 as target speech features, so as to obtain target speech data for the third speech effect based on speech feature 1, speech feature 2 and target text data.
[0129] In another possible implementation, the receiving device may also respond to the user's selection operation of speech features in the speech feature library, and determine at least one speech feature in the speech feature library as the target speech feature to meet the user's custom needs for different speech effects of the target speech data.
[0130] For example, in conjunction with the speech feature library shown in Table 1 above, a user can perform a selection operation (such as clicking, long pressing, etc.) on speech feature 1. The receiving device can respond to the selection operation, determine speech feature 1 as the target speech feature, and obtain the target speech data of the first speech effect based on speech feature 1 and target text data.
[0131] For example, in conjunction with the speech feature library shown in Table 1 above, the user can perform a selection operation on speech feature 2. The receiving device can respond to the selection operation, determine speech feature 2 as the target speech feature, and obtain the target speech data of the second speech effect based on speech feature 2 and target text data.
[0132] For example, in conjunction with the speech feature library shown in Table 1 above, the user can perform a selection operation on speech feature 1 and speech authentication 2. In response to the selection operation, the receiving device can determine speech feature 1 and speech feature 2 as target speech features, and obtain target speech data for the third speech effect based on speech feature 1, speech feature 2 and target text data.
[0133] S504, the receiving device obtains target speech data based on target text data and target speech features.
[0134] In this embodiment, the transmitting device in the voice data transmission system can send the target text data corresponding to the acquired voice data to the receiving device, so that the receiving device can obtain the target voice data based on the target text data and target voice features. Wherein, while the amount of the target text data is much smaller than the amount of the audio data, this application does not require any bandwidth to transmit the target voice features to the receiving device. The receiving device can determine the target voice features based on a locally stored voice feature library, thus further reducing the probability of bandwidth constraints. With the transmitting device able to efficiently transmit high-quality target text data to the receiving device, and the receiving device able to efficiently obtain high-quality target voice data with different voice effects based on the received target text data and the locally determined target voice features, this application achieves more flexible voice data transmission.
[0135] In some embodiments, both the transmitting device and the receiving device may store the aforementioned voice feature library. In addition to including at least one voice feature, the voice feature library may also include a voice feature identifier corresponding to at least one voice feature. The transmitting device may transmit the voice feature identifier to the receiving device based on the voice features of the voice data and the voice feature library, so that the receiving device can further determine the target voice feature based on the voice feature identifier and the voice feature library, and obtain the voice data to be received based on the target voice feature and the target text data.
[0136] Figure 6 This is a schematic flowchart of a voice data transmission method 600 provided in an embodiment of this application. Figure 6 As shown, the method 600 may include the following steps:
[0137] S601, The transmitting device acquires the target text data corresponding to the voice data.
[0138] For specific details, please refer to the description in the above embodiments. To avoid repetition, it will not be repeated here.
[0139] S602, the transmitting device obtains a voice feature identifier based on the voice feature library and the voice features of the aforementioned voice data.
[0140] It should be understood that the transmitting device may also store a voice feature library, which stores at least one voice feature and a voice feature identifier corresponding to the at least one voice feature.
[0141] Table 2 is a speech feature library provided in this application.
[0142] Table 2
[0143] speech features Voice feature identification Speech Feature 1 1 Speech Feature 2 2
[0144] As shown in Table 2, the speech feature database includes multiple speech features and their corresponding speech feature identifiers. For example, the speech feature identifier for speech feature 1 is "1", and the speech feature identifier for speech feature 2 is "2".
[0145] The transmitting device can acquire the voice features of the voice data, and based on the voice feature library, determine the voice feature identifier corresponding to the voice features of the voice data.
[0146] For example, in conjunction with the speech feature library shown in Table 2, if the speech feature of the speech data is determined to be speech feature 1 in Table 2, the transmitting device can determine that the speech feature identifier of the speech feature of the speech data is "1" and transmit the "1" to the receiving device to indicate that the speech feature of the speech data is speech feature 1, so that the receiving device can obtain the target speech data of the first speech effect based on the speech feature 1 and the target text data.
[0147] For example, in conjunction with the speech feature library shown in Table 2, if the speech feature of the speech data is determined to be speech feature 2 in Table 2, the transmitting device can determine the speech feature identifier of the speech feature of the speech data as "2" and transmit the "2" to the receiving device to indicate that the speech feature of the aforementioned speech data is speech feature 2, so that the receiving device can obtain the target speech data of the second speech effect based on the speech feature 2 and the target text data.
[0148] For example, in conjunction with the speech feature library shown in Table 2, if the speech features of the speech data are determined to be speech feature 1 and speech feature 2 in Table 2, the transmitting device can determine the speech feature identifiers of the speech features of the speech data as "1" and "2", and transmit "1" and "2" to the receiving device to indicate that the speech features of the aforementioned speech data are speech feature 1 and speech feature 2, so that the receiving device can obtain the target speech data of the third speech effect based on the speech feature 1, speech feature 2 and target text data.
[0149] Alternatively, the aforementioned speech feature identifiers can also be determined based on the user's selection of speech features in the speech feature library shown in Table 2 above.
[0150] It should be understood that the speech feature identifiers shown above are merely exemplary. In addition, the speech feature identifiers for each speech feature may be in other forms, and this application does not limit them.
[0151] S603, the transmitting device transmits target text data and voice feature identifier to the receiving device. Correspondingly, the receiving device receives the target text data and voice feature identifier from the transmitting device.
[0152] Similarly, the sending device can transmit the aforementioned target text data and voice feature identifiers to the receiving device via Bluetooth wireless communication technology or satellite communication technology.
[0153] Optionally, the target file data and voice feature identifier can be transmitted to the receiving device together or at the same time; this application does not limit this.
[0154] S604, the receiving device determines the target speech feature based on the speech feature library and the speech feature identifier. The speech feature library may include at least one speech feature and a speech feature identifier corresponding to the at least one speech feature.
[0155] It should be understood that the receiving device also stores a speech feature database, as detailed in Table 2 above. After receiving the target text data and speech feature identifier, the receiving device can determine the target speech features using the speech feature identifier and the speech feature database.
[0156] For example, in conjunction with the speech feature library shown in Table 2 above, if the received speech feature is identified as “1” in Table 2, the receiving device can determine that the target speech feature is speech feature 1 corresponding to “1”, and obtain the target speech data of the first speech effect based on the speech feature 1 and the target text data.
[0157] For example, in conjunction with the speech feature library shown in Table 2 above, if the received speech feature is identified as “2” in Table 2, the receiving device can determine that the target speech feature is speech feature 2 corresponding to “2”, and obtain the target speech data of the second speech effect based on the speech feature 2 and the target text data.
[0158] For example, in conjunction with the speech feature library shown in Table 2 above, if the received speech feature identifier is determined to be “1” and “2” in Table 2, the receiving device can determine that the target speech feature is speech feature 1 corresponding to “1” and speech feature 2 corresponding to “2”, and obtain the target speech data of the first speech effect based on the speech feature 1, speech feature 2 and target text data.
[0159] S605, the receiving device obtains target speech data based on target text data and target speech features.
[0160] In this embodiment, the transmitting device in the voice data transmission system can send the target text data corresponding to the acquired voice data and the voice feature identifier corresponding to the voice feature to the receiving device. This allows the receiving device to determine the target voice feature based on the voice feature identifier and a voice feature library, and to obtain the target voice data based on the target text data and the target voice feature. Since the amount of the target text data is much smaller than the amount of the audio data, compared to consuming excessive bandwidth to transmit a large data stream of target voice features to the receiving device, the transmitting device in this application can transmit only the smaller voice feature identifier. This allows the receiving device to determine the target voice feature based on a locally stored voice feature library and the voice feature identifier from the transmitting device. Therefore, the probability of bandwidth limitations is further reduced. With the transmitting device efficiently transmitting high-quality target text data to the receiving device, and the receiving device efficiently obtaining high-quality target voice data with a voice effect most closely similar to the received target text data and the target voice feature determined based on the voice feature identifier, this application achieves more flexible voice data transmission.
[0161] It should be understood that the voice feature library shown above may be set in the receiving device and / or transmitting device before leaving the factory, or it may be set in the transmitting device at the factory and transmitted from the transmitting device to the receiving device. This application does not limit this.
[0162] Optionally, to improve user experience, the aforementioned voice feature library can also be determined based on the user's voice characteristics. This user can be a user on the sending device side and / or the receiving device side.
[0163] In some embodiments, the transmitting device may also transmit different amounts of data to the receiving device based on transmission status information.
[0164] Figure 7 This is a schematic flowchart of a voice data transmission method 700 provided in an embodiment of this application. Figure 7 As shown, the method 700 may include the following steps:
[0165] S701, the transmitting device acquires the target text data corresponding to the voice data.
[0166] S702, the transmitting device obtains transmission status information, which is used to characterize the maximum amount of data transmitted per unit time.
[0167] S703, determine whether the amount of transmitted data represented by the transmission status information is less than the transmission data amount threshold.
[0168] In some embodiments, the transmission status information may be transmission bandwidth information. Whether the amount of data transmitted is less than the transmission data threshold can be understood as insufficient transmission bandwidth, that is, the maximum data transmission rate of the network connection cannot meet the current or expected data transmission needs.
[0169] S704, when the amount of transmitted data represented by the transmission status information is less than the transmission data amount threshold, the transmitting device transmits the target text data to the receiving device so that the receiving device can obtain the target speech data based on the target text data and the target speech features.
[0170] It should be understood that when the amount of data transmitted is less than the data transmission threshold, i.e., the transmission bandwidth is insufficient and the maximum data transmission rate of the network connection cannot meet the current or expected data transmission needs, the sending device can transmit only the target text data to the receiving device to reduce the amount of data transmitted, alleviate bandwidth pressure, and avoid bandwidth obstruction during data transmission, which would affect the transmission quality of audio data.
[0171] Alternatively, upon receiving the above S703, the transmitting device may also perform the following operations:
[0172] S705, when the amount of transmitted data represented by the transmission status information is greater than or equal to the threshold of the amount of transmitted data, the transmitting device transmits the target text data and the target speech features corresponding to the speech data to the receiving device, so that the receiving device can obtain the target speech data based on the target text data and the target speech features.
[0173] When the transmission status information is transmission bandwidth information, whether the amount of data transmitted here is greater than or equal to the transmission data volume threshold can be understood as sufficient transmission bandwidth, meaning that the maximum data transmission rate of the network connection can meet the current or expected data transmission needs. Therefore, in addition to the target text data, the sending device can also transmit the actual voice characteristics of the voice data to the receiving device to improve the transmission effect of the voice data.
[0174] In some embodiments, the speech features corresponding to the target speech data include at least one of voiceprint, tone, speech rate, and timbre. The larger the amount of transmitted data (the more sufficient the transmission bandwidth), the more speech features the transmitting device can select to be transmitted to the receiving device.
[0175] For example, if the bandwidth is sufficient to divide into multiple levels, such as the first level and the second level, and the transmission bandwidth corresponding to the first level is more sufficient than that corresponding to the second level, if it is determined that the target voice feature of the above voice data includes multiple voice features, the receiving device can transmit part of the voice feature of the target voice feature to the receiving device in the second level, and in the first level, it can transmit all the voice features of the target voice feature to the receiving device.
[0176] In the embodiments of this application, the transmitting device in the voice data transmission system can flexibly adjust the data transmission strategy by transmitting status information. For example, when the amount of data transmitted as represented by the transmission status information is greater than or equal to the data transmission volume threshold, i.e., when bandwidth is sufficient, more data, such as target speech features and target text data, can be transmitted to improve the quality of the target speech data acquired by the receiving device. When the amount of data transmitted as represented by the transmission status information is less than the data transmission volume threshold, i.e., when bandwidth is limited, the amount of data transmitted is reduced, such as transmitting only the target text data, so that the receiving device can obtain highly coherent and real-time target speech data based on the speech features in the locally stored speech feature library and the target text data.
[0177] It should also be understood that the various embodiments described above can be coupled to each other, and this application does not limit this. Furthermore, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0178] The above text combines Figures 1 to 7 The present application describes in detail the voice data transmission method according to the embodiments of this application. The following will be combined with... Figures 8 to 10 This application describes in detail the voice data transmission apparatus according to embodiments thereof.
[0179] Figure 8This application illustrates a transmitting device 800, applied in a voice data transmission system, which also includes a receiving device. The transmitting device 800 includes an acquisition module 801 and a transmission module 802. The acquisition module 801 is used to acquire target text data corresponding to the voice data; the transmission module 802 is used to transmit the target text data to the receiving device, so that the receiving device obtains target voice data based on the target text data and target voice features.
[0180] Optionally, the transmission module 802 is configured to: transmit the speech feature library to the receiving device, the speech feature library including at least one speech feature, so that the receiving device obtains the target speech feature based on the speech feature library, and obtains the target speech data based on the target text data and the target speech feature.
[0181] Optionally, the above-mentioned speech feature library also includes a speech feature identifier corresponding to at least one of the above-mentioned speech features. The transmission module 802 is used to: transmit the target speech feature identifier to the above-mentioned receiving device so that the above-mentioned receiving device can obtain the target speech feature based on the above-mentioned speech feature library and the above-mentioned target speech feature identifier, and obtain the target speech data based on the above-mentioned target text data and the above-mentioned target speech feature.
[0182] Optionally, the aforementioned voice feature identifier is determined by the transmitting device based on the voice features of the aforementioned voice data and the aforementioned voice feature database.
[0183] Optionally, the aforementioned speech feature identifier is determined based on the user's selection operation of speech features in the aforementioned speech feature library.
[0184] Optionally, the acquisition module 801 is used to: acquire the target speech features corresponding to the aforementioned speech data; the transmission module 802 is used to: transmit the aforementioned target speech features to the aforementioned receiving device, so that the aforementioned receiving device can obtain the aforementioned target speech data based on the aforementioned target text data and the aforementioned target speech features.
[0185] Optionally, the acquisition module 801 is configured to: acquire transmission status information, wherein the transmission status information is used to characterize the maximum amount of transmitted data per unit time; the transmission module 802 is configured to: transmit the target text data to the receiving device when the amount of transmitted data characterized by the transmission status information is less than the transmission data amount threshold, so that the receiving device can obtain the target speech data based on the target text data and the target speech features.
[0186] Optionally, the transmission module 802 is configured to: transmit the target text data and the target speech features corresponding to the voice data to the receiving device when the amount of transmitted data represented by the transmission status information is greater than or equal to the threshold of the amount of transmitted data, so that the receiving device can obtain the target voice data based on the target text data and the target speech features transmitted by the transmitting device.
[0187] Optionally, the target speech features corresponding to the aforementioned speech data include at least one of voiceprint, intonation, speech rate, and timbre; the larger the amount of transmitted data, the more speech features the target speech features include.
[0188] It should be understood that the transmitting device 800 here is embodied in the form of a functional module. The term "module" here can refer to an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor, etc.) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the transmitting device 800 can be specifically the transmitting device in the above embodiments, or the functions of the transmitting device in the above embodiments can be integrated into the transmitting device 800. The transmitting device 800 can be used to execute the various processes and / or steps corresponding to the transmitting device in the above method embodiments; to avoid repetition, these will not be described again here. The transmitting device 800 has the function of implementing the corresponding steps executed by the transmitting device in the above method; the above functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In the embodiments of this application, Figure 8 The transmitting device 800 can also be a chip or a chip system, such as a system on chip (SoC).
[0189] Figure 9 A receiving device 900 according to an embodiment of this application is shown, applied to a voice data transmission system, which also includes a transmitting device. The receiving device 900 includes a receiving module 901 and a processing module 902. The receiving module 901 is used to receive target text data from the transmitting device; the processing module 902 is used to obtain target voice data based on the target text data and target voice features.
[0190] Optionally, the receiving device stores a speech feature library, which includes at least one speech feature. The processing module 902 is used to: determine at least one speech feature in the speech feature library as the target speech feature.
[0191] Optionally, the processing module 902 is configured to: in response to a user's selection operation of speech features in the speech feature library, determine at least one speech feature in the speech feature library as the target speech feature.
[0192] Optionally, the above-mentioned speech feature library also includes a speech feature identifier corresponding to the above-mentioned at least one speech feature. The receiving module 901 is used to: receive the speech feature identifier from the above-mentioned transmitting device; the processing module 902 is used to: determine at least one speech feature in the above-mentioned speech feature library corresponding to the above-mentioned speech feature identifier as the above-mentioned target speech feature.
[0193] Optionally, the aforementioned voice feature database is transmitted from the aforementioned transmitting device to the aforementioned receiving device.
[0194] Optionally, the receiving module 901 is configured to: receive target speech features transmitted from the transmitting device; and the processing module 902 is configured to: obtain the target speech data based on the target text data and the target speech features from the transmitting device.
[0195] Optionally, the processing module 902 is used to: play the aforementioned target voice data.
[0196] It should be understood that the receiving device 900 here is embodied in the form of a functional module. The term "module" here can refer to an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor, etc.) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the receiving device 900 can be specifically the receiving device in the above embodiments, or the functions of the receiving device in the above embodiments can be integrated into the receiving device 900. The receiving device 900 can be used to execute the various processes and / or steps corresponding to the receiving device in the above method embodiments; to avoid repetition, these will not be elaborated further here. The above-described receiving device 900 has the function of implementing the corresponding steps executed by the receiving device in the above method; the above functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In the embodiments of this application, Figure 9 The receiving device 900 in the text can also be a chip or a chip system, such as a system on chip (SoC).
[0197] Figure 10A voice data transmission system 1000 according to an embodiment of this application is shown, including a transmitting device and a receiving device. The voice data transmission system 1000 includes a processor 1001, a memory 1002, a communication interface 1003, and a bus 1004. The memory 1002 stores instructions, and the processor 1001 executes the instructions stored in the memory 1002. The processor 1001, memory 1002, and communication interface 1003 are interconnected via the bus 1004.
[0198] In the first implementation, the voice data transmission system 1000 can specifically be the transmitting device in the above embodiments. The processor 1001 is configured to: acquire target text data corresponding to the voice data; and transmit the target text data to the receiving device, so that the receiving device obtains target voice data based on the target text data and target voice features.
[0199] In the second implementation, the voice data transmission system 1000 can specifically be the receiving device in the above embodiments. The processor 1001 is configured to: receive target text data from the aforementioned transmitting device;
[0200] Based on the aforementioned target text data and target speech features, target speech data is obtained.
[0201] It should be understood that the voice data transmission system 1000 may specifically be the transmitting device or receiving device in the above embodiments, or the functions of the transmitting device or receiving device in the above embodiments may be integrated into the voice data transmission system 1000. The voice data transmission system 1000 may be used to execute the various steps and / or processes corresponding to the transmitting device or receiving device in the above method embodiments. Optionally, the memory 1002 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1001. A portion of the memory 1002 may also include non-volatile random access memory. For example, the memory 1002 may also store device type information. The processor 1001 may be used to execute the instructions stored in the memory, and when the processor executes the instructions, the processor 1001 may execute the various steps and / or processes corresponding to the transmitting device or receiving device in the above method embodiments. It should be understood that in the embodiments of this application, the processor can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor executes the instructions in the memory, combining with its hardware to complete the steps of the above method. To avoid repetition, detailed descriptions are not provided here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways.For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces, indirect couplings, or communication connections between devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Additionally, the functional units in the various embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. If the function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. The above descriptions are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A voice data transmission method, characterized in that, A transmitting device used in a voice data transmission system, the voice data transmission system further including a receiving device, the method comprising: Obtain the target text data corresponding to the voice data; The target text data is transmitted to the receiving device so that the receiving device obtains target speech data based on the target text data and target speech features.
2. The method according to claim 1, characterized in that, The method further includes: The speech feature library, which includes at least one speech feature, is transmitted to the receiving device so that the receiving device obtains the target speech feature based on the speech feature library and obtains the target speech data based on the target text data and the target speech feature.
3. The method according to claim 2, characterized in that, The speech feature library also includes a speech feature identifier corresponding to the at least one speech feature, and the method further includes: The target speech feature identifier is transmitted to the receiving device so that the receiving device can obtain the target speech feature based on the speech feature library and the target speech feature identifier, and obtain the target speech data based on the target text data and the target speech feature.
4. The method according to claim 3, characterized in that, The voice feature identifier is determined by the transmitting device based on the voice features of the voice data and the voice feature database.
5. The method according to claim 3, characterized in that, The voice feature identifier is determined based on the user's selection operation of voice features in the voice feature library.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the target speech features corresponding to the speech data; The target speech features are transmitted to the receiving device so that the receiving device obtains the target speech data based on the target text data and the target speech features.
7. The method according to claim 1, characterized in that, Before transmitting the target text data to the receiving device, the method further includes: Obtain transmission status information, which is used to characterize the maximum amount of data transmitted per unit time. The transmission of the target text data to the receiving device includes: If the amount of transmitted data represented by the transmission status information is less than the transmission data amount threshold, the target text data is transmitted to the receiving device so that the receiving device can obtain the target speech data based on the target text data and the target speech features.
8. The method according to claim 7, characterized in that, The method further includes: If the amount of transmitted data represented by the transmission status information is greater than or equal to the threshold of the amount of transmitted data, the target text data and the target speech features corresponding to the speech data are transmitted to the receiving device so that the receiving device can obtain the target speech data based on the target text data and the target speech features.
9. The method according to claim 8, characterized in that, The target speech features corresponding to the speech data include at least one of voiceprint, intonation, speech rate, and timbre; the larger the amount of transmitted data, the more speech features the target speech features include.
10. A method for transmitting voice data, characterized in that, A receiving device for a voice data transmission system, the voice data transmission system further including a transmitting device, the method comprising: Receive target text data from the transmitting device; Based on the target text data and target speech features, target speech data is obtained.
11. The method according to claim 10, characterized in that, The receiving device stores a speech feature library, which includes at least one speech feature. Before obtaining the target speech data based on the target text data and the target speech feature, the method further includes: At least one speech feature in the speech feature library is identified as the target speech feature.
12. The method according to claim 11, characterized in that, The step of determining at least one speech feature from the speech feature library as the target speech feature includes: In response to the user's selection operation of speech features in the speech feature library, at least one speech feature in the speech feature library is identified as the target speech feature.
13. The method according to claim 11, characterized in that, The speech feature library also includes a speech feature identifier corresponding to the at least one speech feature. Determining at least one speech feature in the speech feature library as the target speech feature includes: Receive voice feature identifiers from the transmitting device; At least one speech feature in the speech feature library that corresponds to the speech feature identifier is identified as the target speech feature.
14. The method according to claim 12, characterized in that, The voice feature database is transmitted from the sending device to the receiving device.
15. The method according to claim 10, characterized in that, The method further includes: Receive target voice features transmitted from the transmitting device; Obtaining the target speech data based on the target text data and the target speech features includes: The target speech data is obtained based on the target text data and the target speech features transmitted by the transmitting device.
16. A voice data transmission system, characterized in that, Includes transmitting and receiving devices, wherein: The transmitting device is used to acquire target text data corresponding to the voice data; and to transmit the target text data to the receiving device. The receiving device is used to receive target text data from the sending device; and to obtain target speech data based on the target text data and target speech features.
17. A transmitting device, characterized in that, This is applied to a voice data transmission system, which further includes a receiving device, and the transmitting device includes: The acquisition module is used to acquire the target text data corresponding to the voice data; A transmission module is used to transmit the target text data to the receiving device so that the receiving device can obtain target speech data based on the target text data and target speech features.
18. A receiving device, characterized in that, This is applied to a voice data transmission system, which further includes a transmitting device, and the receiving device includes: The receiving module is used to receive target text data from the sending device; The processing module is used to obtain target speech data based on the target text data and target speech features.
19. A voice data transmission device, characterized in that, It includes a processor and a memory, the memory being used to store code instructions; the processor being used to execute the code instructions to perform the method as described in any one of claims 1 to 16.
20. A computer program product, the computer program product comprising computer program code, characterized in that, When the computer program code is run on a computer, it causes the computer to implement the method as described in any one of claims 1 to 16.