A voice transmission method, system and device

By extracting and transmitting voiceprint and semantic information from the sending end, and converting it at the receiving end, the problems of high bandwidth consumption and poor quality in voice data transmission are solved, achieving efficient and low-latency voice transmission, which is suitable for 5G networks and IoT narrowband scenarios.

CN115223566BActive Publication Date: 2026-04-14CHINA TELECOM CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, voice data transmission suffers from high bandwidth consumption due to the large data volume, making it prone to packet loss and stuttering, resulting in poor transmission quality. This is especially true in 5G network fallback and IoT narrowband applications, where it cannot meet the demands.

Method used

The sending end extracts the voiceprint feature information and semantic information of the speech data and sends them to the receiving end. The receiving end converts the speech data based on the voiceprint feature information, transmitting only the voiceprint feature information and semantic information rather than the complete speech data, and uses a shared trained speech data processing and recovery model for conversion.

Benefits of technology

It reduces the amount of data and latency in voice transmission, improves transmission quality, meets the needs of narrow bandwidth scenarios, and achieves efficient and low-latency voice transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223566B_ABST
    Figure CN115223566B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a voice transmission method, system and device, which are applied to the technical field of data communication. The scheme is applied to a sending end in a voice transmission system, the voice transmission system further comprises a receiving end, and the scheme comprises the following steps: receiving voice data to be sent; extracting voiceprint feature information and semantic information of the voice data to be sent; and sending the extracted voiceprint feature information and semantic information to the receiving end, so that the receiving end converts the received semantic information into voice data based on the received voiceprint feature information after receiving the voiceprint feature information and the semantic information sent by the sending end. Through the scheme, the quality of voice transmission can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data communication technology, and in particular to a voice transmission method, system and device. Background Technology

[0002] With the popularization and widespread application of Internet products such as video and voice calls, online live streaming, and video conferencing, the transmission of high-efficiency and low-latency voice data has become increasingly important.

[0003] In related technologies, after receiving the voice data to be sent, the sending end directly sends the voice data to the receiving end. Since the voice data is large in size, transmitting the voice data requires a lot of bandwidth, which can easily lead to packet loss, stuttering, and poor voice transmission quality. Summary of the Invention

[0004] The purpose of this invention is to provide a voice transmission method, system, and apparatus to improve the quality of voice transmission. The specific technical solution is as follows:

[0005] In a first aspect, embodiments of the present invention provide a voice transmission method applied at a transmitting end of a voice transmission system, the voice transmission system further comprising a receiving end, the method comprising:

[0006] Receive voice data to be sent;

[0007] Extract the voiceprint feature information and semantic information of the voice data to be sent;

[0008] The extracted voiceprint feature information and semantic information are sent to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end converts the received semantic information into speech data based on the received voiceprint feature information.

[0009] Optionally, before extracting the voiceprint feature information and semantic information of the speech data to be sent, the method further includes:

[0010] Determine whether the current session in which the voice data to be sent is located is a newly established session;

[0011] If the current session in which the voice data to be sent is located is a newly established session, then the step of extracting the voiceprint feature information and semantic information of the voice data to be sent is performed.

[0012] Optionally, the method further includes:

[0013] If the current session in which the voice data to be sent is located is not a newly established session, then the semantic information of the voice data to be sent is extracted.

[0014] The extracted semantic information is sent to the receiving end so that the receiving end receives the semantic information sent by the sending end, and reads the voiceprint feature information corresponding to the received semantic information from the historically stored voiceprint feature information, and converts the received semantic information into speech data based on the read voiceprint feature information.

[0015] The historically saved voiceprint feature information includes: the voiceprint feature information first sent by the sending end and received by the receiving end in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the speech data to which the semantic information belongs.

[0016] Optionally, the extraction of voiceprint feature information and semantic information from the voice data to be sent includes:

[0017] The speech data to be sent is input into a pre-trained speech data processing model to obtain the semantic information and voiceprint feature information output by the speech data processing model.

[0018] Optionally, the speech data processing model and the speech data recovery model share training data; wherein, the speech data recovery model is the model used by the receiving end when converting the received semantic information into speech data based on the received voiceprint feature information.

[0019] Secondly, embodiments of the present invention provide a voice transmission method applied to a receiving end in a voice transmission system, wherein the voice transmission system further includes a transmitting end, and the method includes:

[0020] The system receives voiceprint feature information and semantic information sent by the sending end; wherein the voiceprint feature information and semantic information sent by the sending end are: voiceprint feature information and semantic information extracted by the sending end from the received voice data to be sent;

[0021] Based on the received voiceprint feature information, the received semantic information is converted into speech data.

[0022] Optionally, after receiving the voiceprint feature information and semantic information sent by the transmitting end, the method further includes:

[0023] Within the current session period, the received voiceprint feature information is saved.

[0024] Optionally, the method further includes:

[0025] If only the semantic information sent by the sending end is received, then the voiceprint feature information corresponding to the received semantic information is read from the historically saved voiceprint feature information; wherein, the historically saved voiceprint feature information is: the voiceprint feature information first sent by the sending end received by the receiving end in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the voice data to which the semantic information belongs.

[0026] Based on the read voiceprint feature information, the received semantic information is converted into speech data.

[0027] Optionally, the step of converting the received semantic information into speech data based on the received voiceprint feature information includes:

[0028] The received voiceprint feature information and semantic information are input into a pre-trained speech data recovery model to obtain the speech data output by the speech data recovery model.

[0029] Optionally, the speech data recovery model and the speech data processing model share training data; wherein, the speech data processing model is the model used by the sending end when extracting the voiceprint feature information and semantic information of the speech data to be sent.

[0030] Thirdly, embodiments of the present invention provide a voice transmission system, characterized in that the system includes a transmitting end and a receiving end, wherein:

[0031] The sending end is used to perform the steps of the method described in the first aspect;

[0032] The receiving end is used to perform the steps of the method described in the second aspect.

[0033] Optionally, the system further includes:

[0034] A cloud server is used to train a speech data processing model and a speech data recovery model. After training is completed, the speech data processing model is sent to the sending end, and the speech data recovery model is sent to the receiving end.

[0035] The speech data recovery model and the speech data processing model share training data. The speech data processing model is the model used by the sending end to extract the voiceprint feature information and semantic information of the speech data to be sent. The speech data recovery model is the model used by the receiving end to convert the received semantic information into speech data based on the received voiceprint feature information.

[0036] Fourthly, embodiments of the present invention provide a voice transmission device applied to a transmitting end in a voice transmission system, wherein the voice transmission system further includes a receiving end;

[0037] The first receiving module is used to receive the voice data to be sent;

[0038] The information extraction module is used to extract the voiceprint feature information and semantic information of the voice data to be sent;

[0039] The information sending module is used to send the extracted voiceprint feature information and semantic information to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end converts the received semantic information into speech data based on the received voiceprint feature information.

[0040] Fifthly, embodiments of the present invention provide a voice transmission device applied to a receiving end in a voice transmission system, wherein the voice transmission system further includes a transmitting end;

[0041] The information receiving module is used to receive voiceprint feature information and semantic information sent by the sending end; wherein, the voiceprint feature information and semantic information sent by the sending end are: voiceprint feature information and semantic information extracted by the sending end from the received voice data to be sent;

[0042] The data conversion module is used to convert the received semantic information into speech data based on the received voiceprint feature information.

[0043] In a sixth aspect, embodiments of the present invention provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0044] Memory, used to store computer programs;

[0045] A processor, when executing a program stored in memory, implements the steps of the method described in either the first or second aspect.

[0046] In a seventh aspect, embodiments of the present invention provide a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in either the first or second aspect.

[0047] Beneficial effects of the embodiments of the present invention:

[0048] This invention provides a voice transmission method, system, and apparatus. The transmitting end can receive voice data to be transmitted, extract voiceprint feature information and semantic information from the voice data, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the transmitting end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is decreased, and thus the quality of voice transmission is improved.

[0049] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0051] Figure 1 This is a schematic diagram of the structure of a voice transmission system provided in an embodiment of the present invention;

[0052] Figure 2 This is a schematic diagram illustrating the extraction of voiceprint feature information and semantic information according to an embodiment of the present invention;

[0053] Figure 3 This is another structural schematic diagram of a voice transmission system provided in an embodiment of the present invention;

[0054] Figure 4 This is a schematic diagram of the architecture of a voice transmission system provided in an embodiment of the present invention;

[0055] Figure 5 A flowchart of a voice transmission method provided in an embodiment of the present invention;

[0056] Figure 6 A flowchart of another voice transmission method provided in an embodiment of the present invention;

[0057] Figure 7 This is a schematic diagram of the structure of a voice transmission device provided in an embodiment of the present invention;

[0058] Figure 8 This is a schematic diagram of another voice transmission device provided in an embodiment of the present invention;

[0059] Figure 9 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of the present invention.

[0061] With the widespread adoption of internet products such as video and voice calls, live streaming, and video conferencing, efficient and low-latency voice data transmission has become increasingly important. In related technologies, the sending end directly transmits the received voice data to the receiving end. However, due to the large volume of voice data, transmitting it requires significant bandwidth, which can easily lead to packet loss, stuttering, and poor voice transmission quality.

[0062] In addition, with the increasing complexity of communication networks, there is often a need for 5G (5th Generation Mobile Communication Technology) network fallback and voice transmission in IoT narrowband. However, transmitting voice data through related technologies requires a large amount of bandwidth, which cannot meet the needs of such scenarios.

[0063] To address the technical problems existing in related technologies, embodiments of the present invention provide a voice transmission method, system, and apparatus.

[0064] It should be noted that the voice transmission system provided in this embodiment of the invention can be based on a 5G architecture. The sending end and receiving end in this embodiment only represent their roles in a single data transmission; their roles can be interchanged throughout the entire session. For example, the voice transmission system includes communication device A and communication device B. When communication device A needs to send data to communication device B, communication device A is the sending end, and communication device B is the receiving end; conversely, when communication device B needs to send data to communication device A, communication device B is the sending end, and communication device A is the receiving end. The aforementioned receiving end and sending end can be various communication devices such as base stations and edge servers.

[0065] To more clearly illustrate the technical solutions of the embodiments of the present invention, the voice transmission system provided by the embodiments of the present invention will first be introduced. The voice transmission system provided by the embodiments of the present invention may include a transmitting end and a receiving end, wherein:

[0066] The sending end is used to receive voice data to be sent; extract the voiceprint feature information and semantic information of the voice data to be sent; and send the extracted voiceprint feature information and semantic information to the receiving end.

[0067] The receiving end is used to convert the received semantic information into speech data based on the received voiceprint feature information when it receives voiceprint feature information and semantic information sent by the sending end.

[0068] In the above-described scheme of this invention, the sending end can receive the voice data to be sent, extract the voiceprint feature information and semantic information of the voice data to be sent, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the sending end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is reduced, and thus the quality of voice transmission is improved. At the same time, since the amount of data transmitted in voice transmission is small, it can meet the needs of scenarios such as narrow bandwidth, realizing voice transmission in scenarios such as narrow bandwidth.

[0069] The voice transmission system provided in the embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0070] like Figure 1 As shown, this embodiment of the invention provides a voice transmission system, including a transmitter 101 and a receiver 102, wherein:

[0071] The transmitting end 101 is used to receive voice data to be transmitted; extract the voiceprint feature information and semantic information of the voice data to be transmitted; and send the extracted voiceprint feature information and semantic information to the receiving end 102.

[0072] The receiver 102 is used to convert the received semantic information into speech data based on the received voiceprint feature information when it receives the voiceprint feature information and semantic information sent by the sender 101.

[0073] The voice data received by the transmitting end 101 can be voice data sent by the client. For example, one scenario includes communication device A, communication device B, client C, and client D. Communication device A is the edge communication device for client C, and communication device D is the edge communication device for client B. When client C and client D conduct voice communication, client C needs to send voice data to communication device A. Communication device A transmits it to communication device B, and communication device B then transmits it to client D. In this process, communication device A is the transmitting end as defined in this invention, and the voice data it receives from client C is the voice data to be sent. Alternatively, if client D needs to send voice data, client D can send the voice data to communication device B, which transmits it to communication device A, and then communication device A transmits it to client C. In this process, communication device B is the transmitting end as defined in this invention, and the voice data it receives from client D is the voice data to be sent.

[0074] After receiving the voice data to be sent, the sending end 101 does not directly package and send the voice data. Instead, it extracts voiceprint feature information and semantic information from the voice data. Voiceprint feature information refers to the voice feature parameters uniquely representing and identifying the speaker, and the voice model built based on these feature parameters. Simply put, voiceprint feature information can characterize the voice characteristics of a speaker. In this embodiment, the extracted voiceprint feature information can be a voiceprint feature information vector. A voiceprint feature information extraction model, such as one trained using the X-VECTORS (a voiceprint recognition algorithm) model framework, can be used to extract voiceprint feature information from the voice data. The semantic information can be obtained by segmenting the voice data into frame segments and extracting basic acoustic features according to the non-linear frequency perception characteristics of the human ear. These features are then input into an acoustic model and a language model to obtain the final semantic information. In one implementation, the semantic information can be a linguistic feature sequence generated through text preprocessing, such as linguistic information in the voice data, for example, "How are you?".

[0075] In one implementation, the transmitting end 101 can input the speech data to be transmitted into a pre-trained speech data processing model to obtain semantic information and voiceprint feature information output by the speech data processing model. The aforementioned speech data processing model can be a neural network model used to extract both voiceprint feature information and semantic information, or it can include a voiceprint extraction model for extracting voiceprint feature information and a semantic extraction model for extracting semantic information; all of these are possible.

[0076] For example, such as Figure 2As shown in the figure, this embodiment of the invention provides a schematic diagram of voiceprint feature information and semantic information extraction. Voiceprint features and semantic information are extracted through different neural network models. For example, voiceprint features are extracted using a voiceprint extraction model as voiceprint feature information, and semantic information is extracted using a semantic information extraction model. Then, the extracted voiceprint feature information and semantic information are superimposed and encoded for transmission. The above-mentioned voiceprint extraction model can extract features for each frame of data in the speech data to obtain frame features (frame-level), or it can extract features for each frame of data in the speech data to obtain multi-frame features (segmental level). Then, the extracted frame features and multi-frame features are merged as voiceprint features.

[0077] After receiving the extracted voiceprint feature information and semantic information, the sending end 101 sends the extracted voiceprint feature information and semantic information to the receiving end 102. After receiving the voiceprint feature information and semantic information sent by the sending end 101, the receiving end 102 can convert the received semantic information into speech data based on the received voiceprint feature information. For example, the semantic information can be superimposed on the voiceprint and the output spectrum of the acoustic feature generation network can be used as speech data.

[0078] When the transmitting end 101 extracts voiceprint feature information and semantic information using a speech data processing model, the receiving end 102 can input the received voiceprint feature information and semantic information into a pre-trained speech data recovery model to obtain the speech data output by the speech data recovery model.

[0079] Optionally, the aforementioned speech data processing model and speech data recovery model share training data. Simply put, the speech data processing model is trained based on sample speech data and corresponding sample information. During the training process, the sample speech data serves as the input data, and the sample information serves as the ground truth, including the semantic information ground truth and the speaker signature feature information ground truth corresponding to the sample data. Similarly, during the training process of the speech data recovery model, the sample information serves as the input data, and the sample speech data serves as the ground truth. In this approach, by sharing training data between the speech data processing model and the speech data recovery model, the extraction and recovery of speaker signature feature information and speech information can be made more unified and accurate.

[0080] like Figure 3 As shown, this embodiment of the invention also provides a voice transmission system, which further includes a cloud server 103, wherein:

[0081] The cloud server 103 is used to train the speech data processing model and the speech data recovery model. After the training is completed, the speech data processing model is sent to the sending end, and the speech data recovery model is sent to the receiving end.

[0082] The speech data recovery model and the speech data processing model share training data. The speech data processing model is the model used by the sending end 101 when extracting the voiceprint feature information and semantic information of the speech data to be sent. The speech data recovery model is the model used by the receiving end 102 when converting the received semantic information into speech data based on the received voiceprint feature information.

[0083] For example, such as Figure 4 As shown in the figure, this embodiment of the invention also provides a schematic diagram of the architecture of a voice transmission system, including an edge server 401, an edge server 402, and a cloud server 403 that serves as the main anchor point in a 5G communication network. Figure 4 In this architecture, the UPF POOL (User Plane Function Pool) is responsible for routing and forwarding user plane data packets in the 5G core network, identifying data and services, and executing actions and policies. Edge server 401 is located on the sending side of the network, while edge server 402 is located on the receiving side. Cloud server 403 performs cloud model iteration, training a voice data processing model and a voice data recovery model in the cloud, and then distributes these models to edge servers 401 and 402. When a client on the sending side needs to transmit voice data, the voice data is first forwarded through the base station to the UPF POOL on the sending side. The UPF POOL then forwards the data to edge server 401. After receiving the voice data from the client, edge server 401 uses the voice data processing model distributed by cloud server 403 to extract the voiceprint feature information and semantic information (only semantic information needs to be extracted in subsequent sessions), and then sends the extracted voiceprint feature information and semantic information to edge server 402 on the receiving side. After receiving the voiceprint feature information and semantic information, the edge server 402 inputs the received voiceprint feature information and semantic information into the voice data recovery model sent by the cloud server 403 to obtain voice data. Then, the obtained voice data is transmitted to the UPF POOL on the receiving side. The UPF POOL sends the data to the base station on the receiving side, and the base station sends it to the client on the receiving side.

[0084] In practical applications, since the voice data processing model and the voice data recovery model can be trained uniformly in the cloud and then distributed to the communication devices on each edge side after training is completed, there is no need to synchronize training data between the communication devices on each edge side, resulting in higher training efficiency.

[0085] After receiving the voice data, the receiving end 102 can transmit the recovered voice data to the corresponding terminal device, and the terminal device can input the voice data into a vocoder to synthesize voice.

[0086] In the above-described scheme of this invention, the sending end can receive the voice data to be sent, extract the voiceprint feature information and semantic information of the voice data to be sent, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the sending end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is reduced, and thus the quality of voice transmission is improved. At the same time, since the amount of data transmitted in voice transmission is small, it can meet the needs of scenarios such as narrow bandwidth, realizing voice transmission in scenarios such as narrow bandwidth.

[0087] In one embodiment, since the voiceprint feature information contained in the voice data of the same person is the same, in order to further reduce the amount of data transmitted by voice, the present invention embodiment can extract and transmit the voiceprint feature information only once within a session cycle. Thus, in the later transmission process of the session cycle, only the semantic information of the voice data needs to be transmitted, without transmitting the voiceprint feature information.

[0088] In this scenario, in one implementation, the sending end 101 is further configured to determine whether the current session of the voice data to be sent is a newly established session before extracting the voiceprint feature information and semantic information of the voice data to be sent. If it is a newly established session, it means that the voiceprint feature information of the object to which the voice data belongs has not yet been extracted. Therefore, it is necessary to extract the voiceprint feature information. In this case, the sending end 101 is specifically configured to perform the step of extracting the voiceprint feature information and semantic information of the voice data to be sent if the current session of the voice data to be sent is a newly established session. In this scenario, the receiving end 102 is further configured to save the received voiceprint feature information within the current session period. Taking the example where both the receiving end and the sending end are edge servers, the received voiceprint feature information can be cached in the edge server. If the current session in which the voice data to be sent is located is not a newly established session, it means that before receiving the current voice data to be sent, the sending end 101 has already extracted and sent the voiceprint feature information, and the receiving end 102 has also saved the corresponding voiceprint feature information. In this case, in order to further reduce the amount of data transmitted by voice, the sending end 101 can extract only the semantic information of the voice data to be sent. That is, if the current session in which the voice data to be sent is located is not a newly established session, the semantic information of the voice data to be sent is extracted and sent to the receiving end 102. In this case, the data received by the receiving end contains only semantic information and not voiceprint feature information. Therefore, after receiving the semantic information sent by the sending end, it can read the voiceprint feature information corresponding to the received semantic information from the historically saved voiceprint feature information. Based on the read voiceprint feature information, it converts the received semantic information into speech data. The historically saved voiceprint feature information is the voiceprint feature information that the receiving end received first sent by the sending end in the current session. The voiceprint feature information corresponding to each semantic information is the voiceprint feature information contained in the speech data to which the semantic information belongs.

[0089] In the above-described solution of the present invention, the quality of voice transmission can be improved and the requirements of scenarios such as narrow bandwidth can be met, realizing voice transmission in scenarios such as narrow bandwidth. Furthermore, by determining whether the current session in which the voice data to be sent is located is a newly established session, the step of extracting the semantic information of the voice data to be sent is performed only when the current session in which the voice data to be sent is located is a newly established session, thereby further reducing the amount of data transmitted by voice.

[0090] Based on the above solutions, this invention provides a voice transmission system integrated with a 5G framework. At the sending edge, speech semantics and voiceprint feature information are extracted. Semantic information is superimposed and encoded with voiceprint feature information for transmission. At the receiving end, the semantic and voiceprint information are recovered and used for speech synthesis. According to the 5G cloud-edge-device network architecture, the speech codec model is centrally trained in the cloud and distributed to the edge-side communication devices, eliminating the need for synchronized knowledge bases between the sending and receiving ends. This results in more unified and accurate voiceprint and semantic recovery. The edge-side speech data processing model extracts voiceprint and semantic information from the transmitted speech data, requiring no computational power from the device terminal. User voiceprint feature information data is sent and cached at the initial connection establishment and remains valid throughout the entire session, making the session more flexible and secure for both parties.

[0091] Corresponding to the voice transmission system provided in the above embodiments of the present invention, such as Figure 5 As shown, this embodiment of the invention also provides a voice transmission method, applied to a transmitting end in a voice transmission system. The voice transmission system also includes a receiving end, and the method includes steps S501-S503.

[0092] S501 receives voice data to be sent;

[0093] S502, extract the voiceprint feature information and semantic information of the voice data to be sent;

[0094] In one implementation, before extracting the voiceprint feature information and semantic information of the voice data to be sent, it can be determined whether the current session of the voice data to be sent is a newly established session. If it is a newly established session, then step S502 is executed.

[0095] If the current session in which the voice data to be sent is located is not a newly established session, then the semantic information of the voice data to be sent is extracted, and then the extracted semantic information is sent to the receiving end so that the receiving end can receive the semantic information sent by the sending end, and read the voiceprint feature information corresponding to the received semantic information from the historically saved voiceprint feature information, and convert the received semantic information into voice data based on the read voiceprint feature information; wherein, the historically saved voiceprint feature information is: the voiceprint feature information first sent by the sending end received by the receiving end in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the voice data to which the semantic information belongs.

[0096] Optionally, in one information extraction method, the speech data to be sent can be input into a pre-trained speech data processing model to obtain the semantic information and voiceprint feature information output by the speech data processing model.

[0097] The speech data processing model can share training data with the speech data recovery model. The speech data recovery model is the model used by the receiving end to convert the received semantic information into speech data based on the received voiceprint feature information.

[0098] S503, the extracted voiceprint feature information and semantic information are sent to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end converts the received semantic information into speech data based on the received voiceprint feature information.

[0099] In the above-described scheme of this invention, the sending end can receive the voice data to be sent, extract the voiceprint feature information and semantic information of the voice data to be sent, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the sending end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is reduced, and thus the quality of voice transmission is improved. At the same time, since the amount of data transmitted in voice transmission is small, it can meet the needs of scenarios such as narrow bandwidth, realizing voice transmission in scenarios such as narrow bandwidth.

[0100] Corresponding to the voice transmission system provided in the above embodiments of the present invention, such as Figure 6 As shown, this embodiment of the invention also provides a voice transmission method, applied to a receiving end in a voice transmission system. The voice transmission system also includes a sending end, and the method includes steps S601-S602.

[0101] S601, receive voiceprint feature information and semantic information sent by the sending end; wherein, the voiceprint feature information and semantic information sent by the sending end are: voiceprint feature information and semantic information extracted by the sending end from the received voice data to be sent.

[0102] In one implementation, after receiving voiceprint feature information and semantic information sent by the sending end, the received voiceprint feature information can be saved within the current session period. Therefore, if the receiving end only receives the semantic information sent by the sending end, it can read the voiceprint feature information corresponding to the received semantic information from the historically saved voiceprint feature information, and convert the received semantic information into speech data based on the read voiceprint feature information.

[0103] Among them, the historically saved voiceprint feature information is: the voiceprint feature information that the receiver receives first sent by the sender in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the speech data to which the semantic information belongs.

[0104] S602, based on the received voiceprint feature information, converts the received semantic information into speech data.

[0105] In one implementation, the received voiceprint feature information and semantic information can be input into a pre-trained speech data recovery model to obtain the speech data output by the speech data recovery model.

[0106] Optionally, the aforementioned speech data recovery model can share training data with the speech data processing model; wherein, the speech data processing model is the model used by the sending end to extract the voiceprint feature information and semantic information of the speech data to be sent.

[0107] In the above-described scheme of this invention, the sending end can receive the voice data to be sent, extract the voiceprint feature information and semantic information of the voice data to be sent, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the sending end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is reduced, and thus the quality of voice transmission is improved. At the same time, since the amount of data transmitted in voice transmission is small, it can meet the needs of scenarios such as narrow bandwidth, realizing voice transmission in scenarios such as narrow bandwidth.

[0108] Corresponding to the voice transmission method applied to the transmitting end provided in the above embodiments of the present invention, such as Figure 7 As shown, this embodiment of the invention also provides a voice transmission device, applied to the transmitting end of a voice transmission system, wherein the voice transmission system further includes a receiving end, and the device includes:

[0109] The first receiving module 701 is used to receive voice data to be sent;

[0110] The information extraction module 702 is used to extract the voiceprint feature information and semantic information of the voice data to be sent;

[0111] The information sending module 703 is used to send the extracted voiceprint feature information and semantic information to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end converts the received semantic information into speech data based on the received voiceprint feature information.

[0112] Optionally, the information extraction module is further configured to determine whether the current session in which the voice data to be sent is located is a newly established session before extracting the voiceprint feature information and semantic information of the voice data to be sent; if the current session in which the voice data to be sent is located is a newly established session, then the step of extracting the voiceprint feature information and semantic information of the voice data to be sent is performed.

[0113] Optionally, the information extraction module is further configured to: extract semantic information of the voice data to be sent if the current session in which the voice data to be sent is located is not a newly established session; send the extracted semantic information to the receiving end so that the receiving end receives the semantic information sent by the sending end; read the voiceprint feature information corresponding to the received semantic information from the historically saved voiceprint feature information; and convert the received semantic information into voice data based on the read voiceprint feature information; wherein the historically saved voiceprint feature information is: the voiceprint feature information first sent by the sending end that the receiving end receives in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the voice data to which the semantic information belongs.

[0114] Optionally, the information extraction module specifically inputs the voice data to be sent into a pre-trained voice data processing model to obtain the semantic information and voiceprint feature information output by the voice data processing model.

[0115] Optionally, the speech data processing model and the speech data recovery model share training data; wherein, the speech data recovery model is the model used by the receiving end when converting the received semantic information into speech data based on the received voiceprint feature information.

[0116] In the above-described scheme of this invention, the sending end can receive the voice data to be sent, extract the voiceprint feature information and semantic information of the voice data to be sent, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the sending end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is reduced, and thus the quality of voice transmission is improved. At the same time, since the amount of data transmitted in voice transmission is small, it can meet the needs of scenarios such as narrow bandwidth, realizing voice transmission in scenarios such as narrow bandwidth.

[0117] Corresponding to the voice transmission method applied to the receiving end provided in the above embodiments of the present invention, such as Figure 8As shown, this embodiment of the invention also provides a voice transmission device, applied to a receiving end in a voice transmission system, wherein the voice transmission system further includes a transmitting end, and the device includes:

[0118] The information receiving module 801 is used to receive voiceprint feature information and semantic information sent by the sending end; wherein, the voiceprint feature information and semantic information sent by the sending end are: voiceprint feature information and semantic information extracted by the sending end from the received voice data to be sent.

[0119] The data conversion module 802 is used to convert the received semantic information into speech data based on the received voiceprint feature information.

[0120] Optionally, the device further includes:

[0121] The feature storage module is used to store the received voiceprint feature information within the current session period after the information receiving module performs the receiving of voiceprint feature information and semantic information sent by the sending end.

[0122] Optionally, the data conversion module is further configured to, if only semantic information sent by the sending end is received, read the voiceprint feature information corresponding to the received semantic information from the historically saved voiceprint feature information; wherein, the historically saved voiceprint feature information is: the voiceprint feature information first sent by the sending end received by the receiving end in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the voice data to which the semantic information belongs; and convert the received semantic information into voice data based on the read voiceprint feature information.

[0123] Optionally, the data conversion module specifically inputs the received voiceprint feature information and semantic information into a pre-trained speech data recovery model to obtain the speech data output by the speech data recovery model.

[0124] Optionally, the speech data recovery model and the speech data processing model share training data; wherein, the speech data processing model is the model used by the sending end when extracting the voiceprint feature information and semantic information of the speech data to be sent.

[0125] In the above-described scheme of this invention, the sending end can receive the voice data to be sent, extract the voiceprint feature information and semantic information of the voice data to be sent, and send the extracted voiceprint feature information and semantic information to the receiving end. When the receiving end receives the voiceprint feature information and semantic information sent by the sending end, it can convert the received semantic information into voice data based on the received voiceprint feature information. Since only the voiceprint feature information and semantic information of the voice data need to be transmitted during voice transmission, instead of transmitting the complete voice data, the amount of data transmitted is reduced, the latency of voice transmission is reduced, and thus the quality of voice transmission is improved. At the same time, since the amount of data transmitted in voice transmission is small, it can meet the needs of scenarios such as narrow bandwidth, realizing voice transmission in scenarios such as narrow bandwidth.

[0126] This invention also provides an electronic device, such as... Figure 9 As shown, it includes a processor 901, a communication interface 902, a memory 903, and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904.

[0127] Memory 903 is used to store computer programs;

[0128] When processor 901 executes a program stored in memory 903, it performs the following steps:

[0129] The system receives voice data to be sent; extracts voiceprint feature information and semantic information from the voice data to be sent; and sends the extracted voiceprint feature information and semantic information to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end converts the received semantic information into voice data based on the received voiceprint feature information.

[0130] or,

[0131] The system receives voiceprint feature information and semantic information sent by the sending end; wherein the voiceprint feature information and semantic information sent by the sending end are: voiceprint feature information and semantic information extracted by the sending end from the received voice data to be sent; and based on the received voiceprint feature information, the system converts the received semantic information into voice data.

[0132] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0133] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0134] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0135] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0136] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described voice transmission methods.

[0137] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the voice transmission methods described above.

[0138] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0139] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0140] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of methods, apparatus, electronic devices, computer-readable storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0141] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A voice transmission method, characterized in that, A transmitting end applied in a voice transmission system, the voice transmission system further including a receiving end and a cloud server, the method comprising: Receive voice data to be sent; If the current session in which the voice data to be sent is located is a newly established session, the voiceprint feature information and semantic information of the voice data to be sent are extracted using the voice data processing model issued by the cloud server. The extracted voiceprint feature information and semantic information are sent to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end can convert the received semantic information into voice data based on the received voiceprint feature information and use the voice data recovery model issued by the cloud server. The cloud server is used to train the speech data processing model and the speech data recovery model, and the speech data recovery model shares training data with the speech data processing model.

2. The method according to claim 1, characterized in that, Before extracting the voiceprint feature information and semantic information of the voice data to be sent, the method further includes: Determine whether the current session in which the voice data to be sent is located is a newly established session; If the current session in which the voice data to be sent is located is a newly established session, then the step of extracting the voiceprint feature information and semantic information of the voice data to be sent is performed.

3. The method according to claim 2, characterized in that, The method further includes: If the current session in which the voice data to be sent is located is not a newly established session, then the semantic information of the voice data to be sent is extracted. The extracted semantic information is sent to the receiving end so that the receiving end receives the semantic information sent by the sending end, and reads the voiceprint feature information corresponding to the received semantic information from the historically stored voiceprint feature information, and converts the received semantic information into speech data based on the read voiceprint feature information. The historically saved voiceprint feature information includes: the voiceprint feature information first sent by the sending end and received by the receiving end in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the speech data to which the semantic information belongs.

4. The method according to claim 1, characterized in that, The extraction of voiceprint feature information and semantic information from the voice data to be sent includes: The speech data to be sent is input into a pre-trained speech data processing model to obtain the semantic information and voiceprint feature information output by the speech data processing model.

5. The method according to claim 4, characterized in that, The speech data processing model and the speech data recovery model share training data; wherein, the speech data recovery model is the model used by the receiving end when converting the received semantic information into speech data based on the received voiceprint feature information.

6. A voice transmission method, characterized in that, A receiving end applied in a voice transmission system, the voice transmission system further including a sending end and a cloud server, the method comprising: The system receives voiceprint feature information and semantic information sent by the sending end; wherein the voiceprint feature information and semantic information sent by the sending end are: if the current session of the voice data to be sent is a newly established session, the sending end extracts the voiceprint feature information and semantic information from the received voice data to be sent using the voice data processing model issued by the cloud server. Based on the received voiceprint feature information, the received semantic information is converted into voice data using the voice data recovery model issued by the cloud server. The cloud server is used to train the speech data processing model and the speech data recovery model, and the speech data recovery model shares training data with the speech data processing model.

7. The method according to claim 6, characterized in that, After receiving the voiceprint feature information and semantic information sent by the transmitting end, the method further includes: Within the current session period, the received voiceprint feature information is saved.

8. The method according to claim 6, characterized in that, The method further includes: If only the semantic information sent by the sending end is received, then the voiceprint feature information corresponding to the received semantic information is read from the historically saved voiceprint feature information; wherein, the historically saved voiceprint feature information is: the voiceprint feature information first sent by the sending end received by the receiving end in the current session; the voiceprint feature information corresponding to each semantic information is: the voiceprint feature information contained in the voice data to which the semantic information belongs. Based on the read voiceprint feature information, the received semantic information is converted into speech data.

9. The method according to claim 6, characterized in that, The step of converting the received semantic information into speech data based on the received voiceprint feature information includes: The received voiceprint feature information and semantic information are input into a pre-trained speech data recovery model to obtain the speech data output by the speech data recovery model.

10. The method according to claim 9, characterized in that, The speech data recovery model and the speech data processing model share training data; wherein, the speech data processing model is the model used by the sending end when extracting the voiceprint feature information and semantic information of the speech data to be sent.

11. A voice transmission system, characterized in that, The system includes a transmitter, a receiver, and a cloud server, wherein: The sending end is an edge server, used to execute the steps of the method according to any one of claims 1-5; The receiving end is an edge server, used to execute the steps of the method according to any one of claims 6-10.

12. The voice transmission system according to claim 11, characterized in that, The system also includes: A cloud server is used to train a speech data processing model and a speech data recovery model. After training is completed, the speech data processing model is sent to the sending end, and the speech data recovery model is sent to the receiving end. The speech data recovery model and the speech data processing model share training data. The speech data processing model is the model used by the sending end to extract the voiceprint feature information and semantic information of the speech data to be sent. The speech data recovery model is the model used by the receiving end to convert the received semantic information into speech data based on the received voiceprint feature information.

13. A voice transmission device, characterized in that, A transmitting end used in a voice transmission system, the voice transmission system further including a receiving end and a cloud server, the device comprising: The first receiving module is used to receive the voice data to be sent; The information extraction module is used to extract the voiceprint feature information and semantic information of the voice data to be sent using the voice data processing model issued by the cloud server if the current session in which the voice data to be sent is in a newly established session. The information sending module is used to send the extracted voiceprint feature information and semantic information to the receiving end, so that after receiving the voiceprint feature information and semantic information sent by the sending end, the receiving end can convert the received semantic information into voice data based on the received voiceprint feature information and use the voice data recovery model issued by the cloud server. The cloud server is used to train the speech data processing model and the speech data recovery model, and the speech data recovery model shares training data with the speech data processing model.

14. A voice transmission device, characterized in that, A receiving end used in a voice transmission system, the voice transmission system further including a transmitting end and a cloud server, the device comprising: The information receiving module is used to receive voiceprint feature information and semantic information sent by the sending end; wherein, the voiceprint feature information and semantic information sent by the sending end are: if the current session of the voice data to be sent is a newly established session, the sending end extracts the voiceprint feature information and semantic information from the received voice data to be sent using the voice data processing model issued by the cloud server. The data conversion module is used to convert the received semantic information into speech data based on the received voiceprint feature information and using the speech data recovery model issued by the cloud server. The cloud server is used to train the speech data processing model and the speech data recovery model, and the speech data recovery model shares training data with the speech data processing model.

15. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method according to any one of claims 1-5 or 6-10.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described in any one of claims 1-5 or 6-10.

Citation Information

Patent Citations

  • Underwater sound voice digital communication method based on voiceprint features and semantic compression

    CN114387976A