Coding and decoding method and coding and decoding device
By encoding the identifier of the interactive data into a high- and low-frequency combined signal at the encoding end and transmitting it to the decoding end, the problem of the decoding end being unable to identify the sound source information is solved, and accurate identification of user categories and effective processing of interactive data are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-10
AI Technical Summary
In intelligent voice interaction products, the decoding end cannot effectively identify the sound source information corresponding to the voice, such as whether it is the voice of a nearby user or a distant user, which affects the functional processing of the application end.
By acquiring the identifier of the interactive data at the encoding end and encoding it into a high-low frequency combination signal, which is then sent to the decoding end, the decoding end identifies the user category through the high-low frequency combination signal, thereby realizing the transmission and processing of user categories.
The decoding end can accurately identify the user category corresponding to the interactive data, and then effectively process the interactive data based on the user category, thereby improving the application-side functionality of voice interaction.
Smart Images

Figure CN121838780A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of encoding and decoding processing technology, and in particular to an encoding and decoding method and an encoding and decoding device. Background Technology
[0002] In two-way voice conversations, distinguishing sound sources from different directions mainly relies on microphone arrays or beamforming methods, or other complex signal processing algorithms.
[0003] One approach is based on microphone arrays. This method involves arranging multiple microphones in a specific layout and using the time difference of arrival (TDOA) or intensity difference of arrival (IDOA) between the microphones to determine the direction of the sound source. This method is suitable for both indoor and outdoor environments, can handle multiple sound sources, and is relatively low in cost and implementation complexity, but may require a large number of microphones to achieve high-precision positioning.
[0004] The second method is based on beamforming. By adjusting the weights of the microphone array, a beam pointing in a specific direction is formed, thereby enhancing the sound signal from that direction and suppressing interference from other directions. This method is suitable for scenarios requiring high directional gain, such as video conferencing and hearing aids. This method can provide higher directional gain, but its computational complexity is high.
[0005] Thirdly, the high-resolution spectral estimation method analyzes the spectrum of the signal received by the microphone and uses the phase information of the signal to determine the direction of the sound source. This method is suitable for scenarios requiring high-precision sound source localization, such as sound source separation and speech recognition. Theoretically, this method can provide very high localization accuracy, but in practical applications it may be affected by noise and reverberation.
[0006] In related technologies, the differentiation of sound sources from different directions is mainly achieved on the hardware side of intelligent interactive products. However, the application side (or cloud side) of intelligent interactive products cannot effectively identify sound sources from different directions, which in turn affects various functions of the application side. Summary of the Invention
[0007] In view of this, embodiments of this application provide at least one encoding / decoding method and encoding / decoding apparatus.
[0008] The technical solution of this application embodiment is implemented as follows:
[0009] On one hand, embodiments of this application provide an encoding method applied to an encoding end. The method includes: acquiring user interaction data and determining an identifier corresponding to the interaction data; the identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user; encoding the identifier to obtain a high-low frequency combined signal; the high-low frequency combined signal is used to characterize the identifier through a combination of at least two signals of different frequencies; and sending the interaction data and the high-low frequency combined signal to a decoding end.
[0010] On the other hand, embodiments of this application provide a decoding method applied at a decoding end. The method includes: receiving interactive data and a high-low frequency combination signal sent by an encoding end; the high-low frequency combination signal is used to represent an identifier through the combination of at least two signals of different frequencies; the identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user; and decoding the high-low frequency combination signal to obtain the identifier corresponding to the interactive data.
[0011] In another aspect, embodiments of this application provide an encoding device, the device comprising: a determining module, configured to acquire user interaction data and determine an identifier corresponding to the interaction data; the identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user; an encoding module, configured to encode the identifier to obtain a high-low frequency combined signal; the high-low frequency combined signal is used to characterize the identifier through a combination of at least two signals of different frequencies; and a transmitting module, configured to transmit the interaction data and the high-low frequency combined signal to a decoding end.
[0012] In another aspect, embodiments of this application provide a decoding device, the device comprising: a receiving module, configured to receive interactive data and a high-low frequency combined signal sent by an encoding end; the high-low frequency combined signal is used to represent an identifier through the combination of at least two signals of different frequencies; the identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user; and a decoding module, configured to decode the high-low frequency combined signal to obtain the identifier corresponding to the interactive data.
[0013] In this embodiment, user interaction data is acquired, and an identifier corresponding to the interaction data is determined. The identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user. The identifier is encoded to obtain a high-low frequency combined signal. The high-low frequency combined signal is used to characterize the identifier through a combination of at least two signals of different frequencies. The interaction data and the high-low frequency combined signal are sent to the decoding end. In this way, the encoding end sends the user category corresponding to the interaction data to the decoding end at the same time as sending the interaction data, so that the decoding end can obtain the interaction data and the user category, and process the interaction data based on the user category.
[0014] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this application. Attached Figure Description
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0016] Figure 1 A schematic diagram of the implementation process of an encoding method provided in this application embodiment. Figure 1 ;
[0017] Figure 2 A schematic diagram of the implementation process of an encoding method provided in this application embodiment. Figure 2 ;
[0018] Figure 3 A schematic diagram of the implementation process of an encoding method provided in this application embodiment. Figure 3 ;
[0019] Figure 4 A schematic diagram of the implementation process of an encoding method provided in this application embodiment. Figure 4 ;
[0020] Figure 5 A schematic diagram of the implementation process of a decoding method provided in this application embodiment. Figure 1 ;
[0021] Figure 6 A schematic diagram of the implementation process of a decoding method provided in this application embodiment. Figure 2 ;
[0022] Figure 7 A schematic diagram of the implementation process of a decoding method provided in this application embodiment. Figure 3 ;
[0023] Figure 8 A schematic diagram of the implementation process of a decoding method provided in this application embodiment. Figure 4 ;
[0024] Figure 9 The real-time audio data transmission process provided in the embodiments of this application;
[0025] Figure 10 A flowchart of a voice interaction based on real-time translation is provided for an embodiment of this application;
[0026] Figure 11 This is a schematic diagram of the composition structure of an encoding device provided in an embodiment of this application;
[0027] Figure 12 This is a schematic diagram of the composition structure of a decoding device provided in an embodiment of this application;
[0028] Figure 13 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.
[0032] In related technologies, intelligent voice interaction products (smart headphones, smart glasses, smart hearing aids) collect users' voices through microphones, encode the voices directly at the encoding end to obtain encoded data, and send the encoded data to the decoding end. The decoding end decodes the received encoded data to obtain the user's voice. In this process, the decoding end can only receive the user's voice but cannot obtain the corresponding sound source information, such as whether the voice belongs to a nearby or distant user. Therefore, the decoding end cannot further process the voice based on the sound source information.
[0033] This application provides an encoding method that can be executed by a processor of a computer device. The computer device refers to a device with data processing capabilities, such as a server, laptop, tablet, desktop computer, smart TV, set-top box, or mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device).
[0034] Figure 1 A schematic diagram of the implementation process of an encoding method provided in this application embodiment. Figure 1 The method is applied to the encoding end, such as... Figure 1 As shown, the method includes the following steps S101 to S103:
[0035] Step S101: Obtain user interaction data and determine the identifier corresponding to the interaction data; the identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user.
[0036] Interaction data refers to data transmitted between a user and a smart interactive product, or between users themselves. For example, during interaction between a user and a smart interactive product (smart glasses, smart hearing aid), the interaction data can be images or videos captured by the camera in the smart interactive product, or infrared data captured by the infrared thermal imaging system in the smart interactive product. During conversations between users, the interaction data can be the user's audio (speaking voice).
[0037] In some embodiments, when both parties are communicating via voice, the smart interactive product is worn by one of the users. The encoding end of the smart interactive product can capture the voices emitted by both parties in real time through a microphone, encode the captured voices, and send them to the decoding end. The decoding end decodes the received voices. The encoding end refers to the hardware or local end of the smart interactive product, while the decoding end refers to the application end or cloud end of the smart interactive product.
[0038] For the encoding end, users wearing smart interactive products are considered near-end users, while users in voice communication who are not wearing smart interactive products are considered far-end users.
[0039] Each piece of interactive data corresponds to an identifier, which reflects whether the user corresponding to the interactive data is a near-end user or a far-end user. For example, the identifier is a combination of numbers; for voice output by a near-end user, the identifier could be "1 2 5"; for voice output by a far-end user, the identifier could be "1 35". Alternatively, the identifier can be a letter or symbol; for voice output by a near-end user, the identifier could be the letter A; for voice output by a far-end user, the identifier could be the letter B.
[0040] In some embodiments, the users corresponding to the interactive data can be further categorized into users at a distance of 1 meter, users at a distance of 3 meters, etc. A user at a distance of 1 meter can be considered to be approximately 1 meter away from the intelligent interactive product. Each category of user is represented by a corresponding identifier. For example, the voice emitted by a user at a distance of 1 meter is identified as "1 45"; the voice emitted by a user at a distance of 3 meters is identified as "12 3".
[0041] In some embodiments, the identifier is also used to characterize the data category of the interactive data. When the identifier is a combination of numbers, the identifier corresponding to the voice emitted by the near-end user can be "1 2 5", and the identifier corresponding to the video emitted by the near-end user can be "1 5 2".
[0042] In some embodiments, after acquiring user interaction data, the data can be processed using common data processing methods to determine whether the user corresponding to the interaction data belongs to a near-end user or a far-end user, and then a corresponding identifier is determined based on the determined user category. There is a one-to-one correspondence between the identifier and the user category. When the identifier represents the data category of the interaction data, there is a one-to-one correspondence between the identifier and the combination of user category and data category.
[0043] In some embodiments, when the interaction data is audio, after acquiring the interaction data, the direction of the sound source corresponding to the audio (near-end user or far-end user) can be determined based on commonly used microphone array methods, beamforming methods, or high-resolution spectral estimation methods.
[0044] In some embodiments, when the interaction data is video or image, after acquiring the interaction data, the video and image can be processed based on commonly used image processing algorithms or video processing algorithms to determine the user category (near-end user or far-end user).
[0045] In some embodiments, when the interaction data is audio, the user can be distinguished from a remote user by a manual button on the smart interactive product or by the sound intensity of the audio. For example, a button can be manually pressed when a near-end user is speaking, while no button needs to be pressed when a remote user is speaking. When the encoder detects sound, it determines the user category by detecting whether the button has been pressed. For example, the intensity of the collected sound is detected; if the intensity exceeds a threshold, it is determined to be the voice of a near-end user; if the intensity does not exceed the threshold, it is determined to be the voice of a remote user.
[0046] Step S102: Encode the identifier to obtain a high-low frequency combined signal; the high-low frequency combined signal is used to characterize the identifier by combining signals of at least two different frequencies.
[0047] After determining the identifier corresponding to the interactive data, the identifier needs to be encoded, and the encoded signal is sent to the decoding end.
[0048] In some embodiments, the identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol. Each symbol is encoded to obtain the combined signal corresponding to each symbol.
[0049] In some embodiments, the identifier includes a first identifier and a second identifier, wherein the first identifier is used to indicate that the transmission of interactive data has started and the second identifier is used to indicate that the transmission of interactive data has stopped; the first identifier includes at least one symbol and the second identifier includes at least one symbol; the high- and low-frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier; the first identifier is encoded to obtain the first combined signal; the second identifier is encoded to obtain the second combined signal.
[0050] Each symbol corresponds to two different frequencies. A signal of the corresponding frequency is generated based on each frequency. The signals of the two frequencies are combined to obtain the combined signal corresponding to the symbol.
[0051] In some embodiments, the identifier includes a symbol that is characterized by a combination of two signals of different frequencies, i.e., a high-low frequency combination signal is characterized by a combination of two signals of different frequencies.
[0052] In some embodiments, the identifier includes two symbols, each symbol being represented by a combination of two signals of different frequencies. The frequencies corresponding to different symbols are different, that is, the high and low frequency combination signal is represented by a combination of four signals of different frequencies.
[0053] Step S103: Send the interactive data and the high- and low-frequency combined signal to the decoding end.
[0054] The process involves acquiring user interaction data and high- and low-frequency combined signals at the encoding end, and then sending both to the decoding end. This allows the decoding end to determine the content of the data sent by the user based on the interaction data, and to determine the user category (near-end user or far-end user) based on the high- and low-frequency combined signals. The interaction data and high- and low-frequency combined signals can be transmitted to the decoding end via wireless transmission technologies (Bluetooth, Wi-Fi).
[0055] In some embodiments, the interactive data is raw data, which needs to be encoded to obtain encoded data before being sent to the decoding end. For example, the interactive data is the user's audio (speaking voice), which is encoded using a common audio encoding algorithm to obtain encoded data; the interactive data is the user's image, which is encoded using a common image encoding algorithm to obtain encoded data.
[0056] In some embodiments, the interaction data is encoded data obtained by encoding the original data. After obtaining the user's interaction data, the interaction data is directly sent to the decoding end.
[0057] In some embodiments, the interactive data and the high- and low-frequency combined signals are sent to the decoding end in the order of interaction data and high- and low-frequency combined signals. Alternatively, the high- and low-frequency combined signals and the interactive data are sent to the decoding end in the order of interaction data and high- and low-frequency combined signals.
[0058] In some embodiments, the identifier includes a first identifier and a second identifier, and the high-low frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier. The first combined signal, the interactive data, and the second combined signal are sent to the decoding end in the order of the first combined signal, the interactive data, and the second combined signal.
[0059] For example, the interaction data is the audio of the near-end user, with the corresponding identifier being "1 2 5". Each number in the identifier is encoded to obtain the combined signal corresponding to each number. First, the combined signal corresponding to each number is sent to the decoding end, and then the encoded data corresponding to the interaction data is sent to the decoding end.
[0060] In this embodiment, user interaction data is acquired, and an identifier corresponding to the interaction data is determined. The identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user. The identifier is encoded to obtain a high-low frequency combined signal. The high-low frequency combined signal is used to characterize the identifier through a combination of at least two signals of different frequencies. The interaction data and the high-low frequency combined signal are sent to the decoding end. In this way, the encoding end sends the user category corresponding to the interaction data to the decoding end at the same time as sending the interaction data, so that the decoding end can obtain the interaction data and the user category, and process the interaction data based on the user category.
[0061] Figure 2 This is a schematic diagram of the implementation flow of an encoding method provided in an embodiment of this application. Figure 2 This method can be executed by the processor of a computer device. Based on Figure 1 The identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol; Figure 1 S102 in the middle can be updated to S201, which will combine Figure 2 The steps shown are explained.
[0062] Step S201: Encode each symbol to obtain the combined signal corresponding to each symbol.
[0063] The identifier includes at least one symbol, which can be a number, letter, or other arbitrary character. For example, if the interaction data is audio from a nearby user, the corresponding identifier is the number "1," meaning one symbol indicates that the audio belongs to a nearby user. Alternatively, if the interaction data is audio from a nearby user, the corresponding identifier is the numbers "1 2 5," meaning three symbols indicate that the audio belongs to a nearby user.
[0064] Each symbol is encoded individually to obtain a combined signal corresponding to each symbol. All the combined signals together constitute the high and low frequency combined signal corresponding to the identifier.
[0065] In some embodiments, the identifier is divided into two symbols, which can be arranged and combined in four ways to represent four different scenarios of the interactive data, such as audio from the near-end user, image from the near-end user, audio from the far-end user, and image from the far-end user. However, if the identifier is represented by a single symbol, four different symbols are needed to represent these four scenarios. Therefore, the more symbols the identifier is divided into, the more scenarios of the interactive data can be represented.
[0066] In some embodiments, each symbol corresponds to two frequencies. A coded signal corresponding to each frequency is generated, and the coded signals corresponding to each frequency are combined to obtain a combined signal corresponding to the symbol.
[0067] In this embodiment, the identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol. Each symbol is encoded to obtain the combined signal corresponding to each symbol. Thus, by dividing the identifier into at least one symbol and encoding each symbol separately, the high- and low-frequency combined signal corresponding to the identifier can be effectively obtained. Furthermore, using at least one symbol to represent the user corresponding to the interactive data improves the flexibility and diversity of identifier settings.
[0068] Figure 3 This is a schematic diagram of the implementation flow of an encoding method provided in an embodiment of this application. Figure 3 This method can be executed by the processor of a computer device. Based on Figure 2 , Figure 2 S201 in the middle can be updated to S301 to S303, which will combine Figure 3 The steps shown are explained.
[0069] Step S301: Obtain the two frequencies corresponding to the symbol.
[0070] Each symbol corresponds to two frequencies: a high-frequency frequency and a low-frequency frequency, with the high-frequency frequency being greater than the low-frequency frequency. The mapping relationship between each symbol and its two corresponding frequencies is predefined. For example, the high-frequency group contains three high-frequency frequencies, and the low-frequency group contains three low-frequency frequencies. One high-frequency frequency is selected from the high-frequency group, and one low-frequency frequency is selected from the low-frequency group, together corresponding to one symbol. Therefore, a total of nine symbols can be represented.
[0071] In some embodiments, different symbols correspond to different frequencies. For example, a high-frequency group contains three high-frequency frequencies, a low-frequency group contains three low-frequency frequencies, a high-frequency frequency is selected from the high-frequency group, and a low-frequency frequency is selected from the low-frequency group to correspond to one symbol. Therefore, a total of three symbols can be represented.
[0072] Step S302: Generate the coded signal corresponding to each frequency based on each frequency.
[0073] In some embodiments, the encoded signal is a sine wave signal, and a sine wave signal corresponding to the frequency can be generated based on a sine wave generation function.
[0074] In some embodiments, the encoded signal is a cosine wave signal, and the cosine wave signal corresponding to the frequency can be generated based on the cosine wave generation function.
[0075] The duration of the encoded signal can be set according to requirements. For example, the duration can be set to one signal cycle or two signal cycles.
[0076] Step S303: Combine the encoded signals corresponding to each frequency to obtain the combined signal corresponding to the symbol.
[0077] In some embodiments, after obtaining the coded signals of the two frequencies corresponding to the symbol, the two coded signals can be directly superimposed to obtain the combined signal corresponding to the symbol. For example, the coded signal of the first frequency is y = sin(2πf1*t), and the coded signal of the second frequency is y = sin(2πf2*t). Superimposing the two coded signals yields the combined signal sin(2πf1*t) + sin(2πf2*t).
[0078] In some embodiments, after obtaining the encoded signals of the two frequencies corresponding to the symbol, the two encoded signals are directly combined to obtain the combined signal corresponding to the symbol. For example, the first encoded signal and the second encoded signal are concatenated end to end to obtain the combined signal corresponding to the symbol.
[0079] In this embodiment, two frequencies corresponding to a symbol are obtained; based on each frequency, an encoded signal corresponding to each frequency is generated; the encoded signals corresponding to each frequency are combined to obtain a combined signal corresponding to the symbol. In this way, the symbol is encoded using a combined signal of two frequency signals, resulting in a signal with high stability and reliability, and maintaining strong anti-interference capability.
[0080] Figure 4 This is a schematic diagram of the implementation flow of an encoding method provided in an embodiment of this application. Figure 4 This method can be executed by the processor of a computer device. Based on Figure 1 The identifier includes a first identifier and a second identifier, wherein the first identifier is used to indicate that the interactive data transmission has started and the second identifier is used to indicate that the interactive data transmission has stopped; the high- and low-frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier; Figure 1 S103 in the middle can be updated to S401, which will combine Figure 4 The steps shown are explained.
[0081] Step S401: Send the first combined signal, the interactive data, and the second combined signal to the decoding end in the order of the first combined signal, the interactive data, and the second combined signal.
[0082] In some embodiments, for any interactive data, its corresponding identifier can be divided into a first identifier and a second identifier. The first identifier indicates that the interactive data has started transmission, and the second identifier indicates that the interactive data has stopped transmission. For example, if the interactive data is audio from a nearby user, its corresponding first identifier is "12 5", indicating that the transmission of the nearby user's audio has started; its corresponding second identifier is "5 2 1", indicating that the transmission of the nearby user's audio has stopped.
[0083] The first identifier includes at least one symbol, and the second identifier includes at least one symbol. The symbols in the first identifier and the symbols in the second identifier may be the same or different. If the symbols in the first identifier and the second identifier are the same, the order of the symbols in the first identifier and the order of the symbols in the second identifier must not be the same.
[0084] In this system, the first identifier corresponds to the first combined signal, and the second identifier corresponds to the second combined signal. Since the first identifier indicates the start of interactive data transmission, the first combined signal should be sent to the decoding end before the interactive data. Upon receiving the first combined signal, the decoding end parses it to determine that the subsequent data is interactive data. At this point, the interactive data is sent to the decoding end immediately following the first combined signal. The second combined signal is sent to the decoding end after the interactive data. After parsing the second combined signal, the decoding end determines that the interactive data transmission has stopped. Therefore, based on the order of the first combined signal, the interactive data, and the second combined signal, the first combined signal, the interactive data, and the second combined signal are sent to the decoding end sequentially.
[0085] In this embodiment, the high- and low-frequency combined signal includes a first combined signal corresponding to a first identifier and a second combined signal corresponding to a second identifier. The first combined signal, interactive data, and second combined signal are sent to the decoding end in the order of their sequence. In this way, the user category corresponding to the interactive data can be sent to the decoding end via the first and second combined signals, and the decoding end can determine the transmission progress of the interactive data by parsing the first and second combined signals, thus improving parsing efficiency.
[0086] Figure 5 This is a schematic diagram of the implementation flow of a decoding method provided in an embodiment of this application. Figure 1 The method is applied to the decoding end, such as... Figure 5 As shown, the method includes the following steps S501 to S502.
[0087] Step S501: Receive interactive data and high-low frequency combination signal sent by the encoding end; the high-low frequency combination signal is used to represent the identifier through the combination of at least two signals of different frequencies; the identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user.
[0088] Among them, the decoding end refers to the application end or cloud of intelligent interactive products.
[0089] Interaction data refers to data transmitted between a user and a smart interactive product, or between users themselves. For example, during interaction between a user and a smart interactive product (smart glasses, smart hearing aid), the interaction data can be images or videos captured by the camera in the smart interactive product, or infrared data captured by the infrared thermal imaging system in the smart interactive product. During conversations between users, the interaction data can be the user's audio (speaking voice).
[0090] Among them, users wearing smart interactive products are considered near-end users, while users who are not wearing smart interactive products in the voice communication are considered far-end users.
[0091] The process involves decoding the high- and low-frequency combined signals to obtain a decoded identifier. Each piece of interactive data corresponds to an identifier, which reflects whether the user corresponding to the interactive data is a near-end user or a far-end user. For example, the identifier can be a combination of numbers; for voice output by a near-end user, the identifier could be "1 25"; for voice output by a far-end user, the identifier could be "1 3 5". Alternatively, the identifier can be a letter or symbol; for voice output by a near-end user, the identifier could be the letter A; for voice output by a far-end user, the identifier could be the letter B.
[0092] In some embodiments, the decoding end receives interactive data and a high- and low-frequency combined signal sent by the encoding end, determines the data content sent by the user based on the interactive data, and determines the user category (near-end user or far-end user) corresponding to the interactive data based on the high- and low-frequency combined signal.
[0093] In some embodiments, interactive data and high- and low-frequency combined signals are received sequentially, in the order of interaction data and high- and low-frequency combined signals. Alternatively, high- and low-frequency combined signals and interactive data are received sequentially, in the order of high- and low-frequency combined signals and interaction data.
[0094] In some embodiments, the identifier includes a first identifier and a second identifier, wherein the first identifier indicates that the interactive data transmission has begun and the second identifier indicates that the interactive data transmission has stopped; the high- and low-frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier; receiving the interactive data and the high- and low-frequency combined signal sent by the encoding end includes: receiving the first combined signal, the interactive data, and the second combined signal in the order of the first combined signal, the interactive data, and the second combined signal.
[0095] In some embodiments, the interactive data is the original data, and the decoding end receives the encoded data obtained after encoding the interactive data. The encoded data is then decoded to obtain the interactive data.
[0096] In some embodiments, the interactive data is encoded data obtained by encoding the original data, and the interactive data is decoded to obtain the corresponding original data.
[0097] In some embodiments, the identifier includes a symbol that is characterized by a combination of two signals of different frequencies, i.e., a high-low frequency combination signal is characterized by a combination of two signals of different frequencies.
[0098] In some embodiments, the identifier includes two symbols, each symbol being represented by a combination of two signals of different frequencies. The frequencies corresponding to different symbols are different, that is, the high and low frequency combination signal is represented by a combination of four signals of different frequencies.
[0099] Step S502: Decode the high- and low-frequency combined signal to obtain the identifier corresponding to the interactive data.
[0100] In some embodiments, the identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol. Decoding the combined signal corresponding to each symbol yields each symbol.
[0101] Each symbol corresponds to two different frequencies. A signal of the corresponding frequency is generated based on each frequency. The signals of the two frequencies are combined to obtain the combined signal corresponding to the symbol.
[0102] In some embodiments, the identifier includes a first identifier and a second identifier, wherein the first identifier is used to indicate that the interactive data transmission has started and the second identifier is used to indicate that the interactive data transmission has stopped; the first identifier includes at least one symbol and the second identifier includes at least one symbol; the high- and low-frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier; the first combined signal is decoded to obtain the first identifier; the second combined signal is decoded to obtain the second identifier.
[0103] In this embodiment, the receiving end sends interactive data and a high-low frequency combined signal. The high-low frequency combined signal is used to represent an identifier through the combination of at least two signals of different frequencies. The identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user. The high-low frequency combined signal is decoded to obtain the identifier corresponding to the interactive data. In this way, the decoding end also receives the user category corresponding to the interactive data when receiving the interactive data, and can then process the interactive data based on the user category.
[0104] Figure 6 This is a schematic diagram of the implementation flow of a decoding method provided in an embodiment of this application. Figure 2 This method can be executed by the processor of a computer device. Based on Figure 5 The identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol; Figure 5 The S502 in the middle can be updated to S601, which will combine Figure 6 The steps shown are explained.
[0105] Step S601: Decode the combined signal corresponding to each symbol to obtain each symbol.
[0106] The identifier includes at least one symbol, which can be a number, letter, or other arbitrary character. For example, if the interaction data is audio from a nearby user, the corresponding identifier is the number "1," meaning one symbol indicates that the audio belongs to a nearby user. Alternatively, if the interaction data is audio from a nearby user, the corresponding identifier is the numbers "1 2 5," meaning three symbols indicate that the audio belongs to a nearby user.
[0107] In this process, the combined signal corresponding to each symbol is decoded to obtain each symbol, and all symbols together constitute the identifier.
[0108] In some embodiments, each symbol corresponds to two frequencies. The combined signal corresponding to the symbol is split into two encoded signals. The corresponding frequency is determined based on each encoded signal, and the corresponding symbol is determined based on the two frequencies.
[0109] In this embodiment, the identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol. Decoding the combined signal corresponding to each symbol yields each symbol. Thus, by dividing the identifier into at least one symbol and decoding the combined signal corresponding to each symbol, the identifier can be effectively obtained. Furthermore, using at least one symbol to represent the user corresponding to the interactive data improves the flexibility and diversity of identifier settings.
[0110] Figure 7 This is a schematic diagram of the implementation flow of a decoding method provided in an embodiment of this application. Figure 3 This method can be executed by the processor of a computer device. Based on Figure 6 , Figure 6 The S601 in the middle can be updated to S701 to S703, which will combine Figure 7 The steps shown are explained.
[0111] Step S701: Decompose the combined signal corresponding to the symbol to obtain two encoded signals.
[0112] In some embodiments, the combined signal is obtained by superimposing two coded signals based on two frequencies. Therefore, based on the inverse operation of superposition, two coded signals are obtained by splitting the combined signal.
[0113] In some embodiments, the combined signal is obtained by splicing the two encoded signals end to end. Therefore, the combined signal can be split into two encoded signals according to the duration of each encoded signal.
[0114] Step S702: Determine the frequency corresponding to each encoded signal based on each encoded signal.
[0115] In some embodiments, the encoded signal is a sine wave signal, and the frequency corresponding to the sine wave signal can be identified based on the sine wave generation function.
[0116] In some embodiments, the encoded signal is a cosine wave signal, and the frequency corresponding to the cosine wave signal can be identified based on the cosine wave generation function.
[0117] The duration of the encoded signal can be set according to requirements. For example, the duration can be set to one signal cycle or two signal cycles.
[0118] Step S703: Determine the symbols corresponding to the two frequencies based on the two frequencies.
[0119] Each symbol corresponds to two frequencies: a high-frequency frequency and a low-frequency frequency, with the high-frequency frequency being greater than the low-frequency frequency. The mapping relationship between each symbol and its two corresponding frequencies is predefined. For example, the high-frequency group contains three high-frequency frequencies, and the low-frequency group contains three low-frequency frequencies. One high-frequency frequency is selected from the high-frequency group, and one low-frequency frequency is selected from the low-frequency group, together corresponding to one symbol. Therefore, a total of nine symbols can be represented.
[0120] In some embodiments, different symbols correspond to different frequencies. For example, a high-frequency group contains three high-frequency frequencies, a low-frequency group contains three low-frequency frequencies, a high-frequency frequency is selected from the high-frequency group, and a low-frequency frequency is selected from the low-frequency group to correspond to one symbol. Therefore, a total of three symbols can be represented.
[0121] The correspondence between preset frequencies and symbols can be obtained by directly looking up the table after determining two frequencies.
[0122] In this embodiment, the combined signal corresponding to the symbol is split into two coded signals; based on each coded signal, the frequency corresponding to each coded signal is determined; based on the two frequencies, the symbols corresponding to the two frequencies are determined. Thus, by decoding the combined signal of the two frequency signals to obtain the symbol, the signal has high stability and reliability, and maintains strong anti-interference capability.
[0123] Figure 8 This is a schematic diagram of the implementation flow of a decoding method provided in an embodiment of this application. Figure 4 ,based on Figure 5 The interactive data is audio, and the method further includes steps S801 to S803.
[0124] Step S801: Based on the current user corresponding to the identifier, determine the translation language of the audio; the translation language is the language of another user different from the current user.
[0125] After decoding and obtaining the identifier, the decoding end can determine whether the current user corresponding to the interaction data is a near-end user or a far-end user based on the identifier.
[0126] In some embodiments, in a face-to-face communication scenario between a near-end user and a far-end user, the near-end user wears a smart interactive product, while the far-end user does not. The smart interactive product can collect the voices of both the near-end and far-end users in real time and transmit them to a decoding end. Different users may speak different languages; after decoding the audio and its identifier, the decoding end can translate the audio based on the identifier. For example, in the communication, the near-end user speaks Chinese, and the far-end user speaks English.
[0127] The translation language is the text language resulting from the audio translation. If the current user is identified as a near-end user, the audio needs to be translated into the language of the far-end user; conversely, if the current user is identified as a far-end user, the audio needs to be translated into the language of the near-end user.
[0128] In some embodiments, when two users are communicating, the translation language is the language of the other user, which is different from the current user's language.
[0129] In some embodiments, when at least three users are communicating, a target user is determined based on preset rules, and the translation language is different from that of the current user. For example, the target user could be the user closest to the current user.
[0130] Step S802: Translate the audio based on the translation language to obtain the translated text.
[0131] After determining the target language, the audio is translated into corresponding text using common audio translation methods. For example, an audio translation model is pre-trained, and the target language and audio are input into the model; the model outputs the translated text.
[0132] Step S803: Perform speech synthesis based on the translated text to obtain synthesized speech, and send the synthesized speech to the hardware terminal.
[0133] In this context, the synthesized speech is the speech corresponding to the translated language. For example, the translated language is Chinese, and the language corresponding to the synthesized speech is Chinese.
[0134] The hardware terminal is the local terminal of the intelligent interactive product, through which voice can be played.
[0135] In some embodiments, the decoding end first translates the audio into translated text based on the target language, then performs speech synthesis based on the translated text to obtain synthesized speech, which is then sent to the hardware end. The hardware end can then play the synthesized speech to the corresponding user. For example, if the language of the near-end user is Chinese and the language of the far-end user is English, and the interaction data is Chinese audio spoken by the near-end user, the far-end user cannot understand the Chinese audio spoken by the near-end user. By transmitting the interaction data to the decoding end, the decoding end translates the near-end user's Chinese audio into English audio and sends it to the hardware end. The hardware end can then play the English audio to the far-end user, enabling the far-end user to understand the near-end user's audio and achieving barrier-free communication.
[0136] This can be achieved by using common speech synthesis technology to convert translated text into synthesized speech. The synthesized speech can then be transmitted to the hardware via wireless transmission technologies (Bluetooth, Wi-Fi).
[0137] In this embodiment, the translation language of the audio is determined based on the current user corresponding to the identifier; the translation language is the language of another user different from the current user; the audio is translated based on the translation language to obtain translated text; speech synthesis is performed based on the translated text to obtain synthesized speech, and the synthesized speech is sent to the hardware. In this way, the decoding end can determine the translation language based on the identifier, translate the interactive data, and synthesize the corresponding synthesized speech, enabling barrier-free communication between users speaking different languages.
[0138] The following describes the application of the encoding / decoding method provided in the embodiments of this application in a real-world scenario.
[0139] This application provides a novel method for distinguishing sound sources from different directions, applicable to intelligent voice interaction applications (such as translation software, online real-time machine translation, chatGPT real-time translation, etc.). During voice communication, the user's smart wearable device (smart headphones, augmented reality glasses, smart hearing aids) collects sounds from different directions. The collected sound signals are encoded to identify the source information. When the user speaks, the received signal is a near-field sound signal, while when the other party speaks, the received signal is a far-field sound signal. The source information of the sound signals is encoded into a dual-tone multi-frequency (DTMF) signal and transmitted to the application. The application then decodes the DTMF signal, transforming it from the time domain to the frequency domain using mathematical transformations to obtain the digitally encoded information corresponding to the different voices in the dialogue scenario, thus distinguishing sound sources from different directions.
[0140] Before encoding sound signals from different sources, it's necessary to distinguish sound sources from different directions to encode the information from each source separately. Methods for this distinction include: a simple method is to manually distinguish one's own voice from a distant voice using a button, or by judging sound intensity; generally, nearby sounds are louder than distant sounds. More common methods involve using microphone arrays and sound source localization algorithms or beamforming to determine the direction of the sound source.
[0141] In some embodiments, the encoded information is sent in the form of frames, as shown in Table 1. A frame consists of a frame header and frame data. The frame header is 3 bytes long and is of command type. The frame header not only represents the beginning of the data but also determines the length and type of the subsequent data. Different data types determine the length and format of the frame data. Data types can be audio data, image data, or video data. The frame data length can be 0 bytes.
[0142] Table 1
[0143] Frame header Command type, 3 bytes Frame data length Data length, can be 0 bytes Frame data Data, can be 0 bytes
[0144] The frame header instruction word table segment is shown in Table 2, which defines four frame instruction names: start transmitting near-end audio, stop transmitting near-end audio, start transmitting far-end audio, and stop transmitting far-end audio. Each frame instruction has a corresponding command type, command data, and description.
[0145] Table 2
[0146]
[0147]
[0148] Figure 9 This describes the real-time audio data transmission process provided in the embodiments of this application. For example... Figure 9As shown, before transmitting real-time audio data, the start transmission audio command corresponding to the sound source information is transmitted first. After transmitting real-time audio data, the end transmission audio command corresponding to the sound source information is transmitted next. Each command type is transmitted through a DTMF-like signal, which is a composite signal composed of two audio signals of different frequencies superimposed, capable of transmitting information such as numbers, letters, and special symbols. A DTMF-like signal consists of a high-frequency group and a low-frequency group, with each number or symbol represented by a specific combination of a high-frequency signal and a low-frequency signal. For example, the high-frequency group includes three frequencies: 1209Hz, 1336Hz, and 1477Hz, and the low-frequency group includes three frequencies: 697Hz, 770Hz, and 852Hz. The command type for starting transmission is "1, 3, 5," where "1" can be represented by signals at frequencies of 1209Hz and 697Hz, "3" by signals at frequencies of 1336Hz and 770Hz, and "5" by signals at frequencies of 1477Hz and 852Hz. Figure 9 Solid lines in the diagram represent high frequencies, while dashed lines represent low frequencies.
[0149] Figure 10 A flowchart illustrating a real-time translation-based voice interaction process provided for an embodiment of this application. Figure 10 As shown, the embodiment of this application, used for voice interaction in real-time translation, can be divided into a voice collection module, an encoding / decoding module, a multilingual translation module, and a voice synthesis module. The specific process is as follows:
[0150] Step S1001: The microphone collects voices from different directions.
[0151] The system uses the smart wearable product's own microphone to collect the user's voice and the other party's voice respectively; during the voice interaction between the two parties, the microphone can collect the user's voice or the other party's voice in real time.
[0152] Step S1002: Common methods to distinguish sound sources from different directions.
[0153] Common sound source localization methods are used to determine the direction of sound origin and whether the sound is from a nearby user or a distant party. Common sound source localization methods include microphone array-based methods, beamforming-based methods, and high-resolution spectral estimation-based methods. A simple method is to manually distinguish between one's own near-end speech and a distant party's speech using a button; for example, a button can be pressed manually when speaking, but not when someone else is speaking. Smart wearable products determine the sound source information by detecting whether a button has been pressed. Another simple method is to detect the intensity of the collected sound; if the intensity exceeds a threshold, it is identified as the user's near-end speech, and if the intensity does not exceed the threshold, it is identified as a distant party's speech.
[0154] Step S1003: Encode the instruction according to the definition.
[0155] Specifically, the near-end user's voice and the far-end other party's voice are encoded according to the definitions in Table 2. After determining the sound source information of the current voice, before encoding the real-time audio data, the command information to start transmitting audio is encoded based on the predefined command type (start transmitting near-end audio or start transmitting far-end audio). After encoding the real-time audio data, the command information to stop transmitting audio is encoded based on the command type (stop transmitting near-end audio or stop transmitting far-end audio).
[0156] The specific encoding process may include: 1. Defining frequency pairs: determining the high and low frequencies corresponding to each digit; 2. Generating sine waves: determining the sampling rate for generating sine waves and the duration of each digit, and using a sine wave generation function (such as sin(2*pi*f*t), where f is the frequency and t is the time) to generate low-frequency and high-frequency sine waves respectively; the duration of the sine waves should be equal to or slightly longer than one complete cycle of the signal; 3. Synthesizing signals: superimposing the low-frequency and high-frequency sine waves to form a DTMF-like signal; during superposition, the amplitudes of the two sine waves can be equal or adjusted as needed to meet specific signal strength requirements; 4. Adjusting signal parameters: Frequency accuracy: ensuring that the frequency accuracy of the low-frequency and high-frequency components is within the allowable range, typically requiring a frequency deviation of no more than ±1.5%. Signal duration: the duration of each DTMF-like signal should meet standard requirements, typically between 40ms and 55ms for each digit or character. 5. Silence interval: Insert appropriate silence intervals between consecutive numbers or characters to ensure clear distinction between signals; 6. Output DTMF-like signal: Output the synthesized DTMF-like signal to the target device or system.
[0157] Step S1004: Bluetooth transmission encoded signal.
[0158] The process involves transmitting all encoded information to the application via Bluetooth; collecting sound locally using a microphone; and then transmitting the collected sound information and sound source information to the mobile application via Bluetooth.
[0159] Step S1005: Mathematical transformation and decoding to parse out the instruction.
[0160] The application decodes the received encoded information to obtain the command type corresponding to the sound, and determines the direction of the received sound source based on the command information corresponding to the command type. The decoding process of the sound source information may include: 1. Signal processing: Preprocessing the received audio signal, converting the analog signal into a digital signal, dividing the digital signal into multiple frames for processing, and performing spectrum analysis on each frame signal. A commonly used method is Fast Fourier Transform. 2. Frequency identification and conversion: Finding the two peaks with the highest energy in the frequency spectrum, matching the found peak frequencies with standard frequencies similar to DTMF signals, determining each digit, and obtaining the command type.
[0161] Step S1006: Real-time multilingual translation.
[0162] Based on the sound source information received by the encoding and decoding method, the application can select the appropriate translation language for the user or the other party in real time for real-time translation. That is, when the sound source is from the user, the speech is translated into the language used by the other party, and when the sound source is from the other party, the speech is translated into the language used by the user.
[0163] Step S1007: Perform speech synthesis based on the user's auditory characteristics.
[0164] In this process, based on the language of the matched user or other party, and combined with the user's auditory characteristics, the translated text is synthesized into TTS (Text to Speech) on the application side. Auditory characteristics can include hearing sensitivity, which refers to an individual's ability to perceive sounds of different frequencies and intensities; and sound preferences, which refer to an individual's degree of liking for different types of sounds, including music style, rhythm, volume, and sounds in the daily environment (such as natural sounds, urban noise, etc.).
[0165] Step S1008: Transmit synthesized speech via WiFi or Bluetooth.
[0166] The synthesized voice is transmitted to the hardware via Wi-Fi or Bluetooth.
[0167] Step S1009: Play the translated and synthesized speech through the loudspeaker.
[0168] The synthesized speech is played to the user through a loudspeaker.
[0169] The speech collection module corresponds to steps S1001 to S1002, the encoding and decoding module corresponds to steps S1003 to S1005, the multilingual translation module corresponds to step S1006, and the speech synthesis module corresponds to steps S1007 to S1009.
[0170] In some embodiments, for speech sounds from different directions, only the source information of a fixed single direction sound (such as the speech sound of a distant party) is encoded. After decoding, if the encoded information of the corresponding sound source is detected, the collected speech sound is determined to be the speech sound of a distant party. If no encoded information is detected after decoding, the collected speech sound is determined to be the speech sound of a nearby user.
[0171] In some embodiments, as shown in Table 3, multiple voices from different directions (audio from 1 meter away and audio from 3 meters away) are encoded differently, and the encoded information of the corresponding sound source is detected after decoding, and the direction of the sound source is determined accordingly. For smart wearable products equipped with infrared thermal imaging or cameras, in addition to encoding the sound source information of audio data, the infrared sensor data or image or video data can also be encoded and decoded (for example, it can distinguish between near-end infrared data, far-end infrared data, near-end image data, far-end image data, near-end video data, and far-end video data), thereby marking different related data for intelligent applications. For example, the marking can be used by the application to distinguish the infrared / image / video data of different people for individual artificial intelligence applications. When a smart wearable product issues a specified instruction through an application app or cloud, such as prompting the user to complete the "take a picture" action in front of the installed camera, or to complete the "smile" expression in front of the installed camera, the action control instruction or expression control instruction can be encoded so that the application can determine that the action instruction or expression instruction has been completed. For situations where there are multiple distinct directions of speech or other abnormal situations, custom encoding that differs from fixed encoding can be used.
[0172] Table 3
[0173]
[0174]
[0175] This application embodiment differs from general methods for distinguishing sound sources from different directions. It employs a novel encoding and decoding method similar to DTMF for differentiation. The smart devices used are not limited to known smart headphones, smart glasses, and smart hearing aids, but include any other smart wearable products that support voice interaction, motion or facial expression control.
[0176] The embodiments of this application are low in cost and implementation complexity, and have no obvious requirements on the number of microphones; the voice signal is encoded into a DTMF-like signal, which has high anti-interference capability and transmission stability, and is suitable for various communication environments.
[0177] Based on the foregoing embodiments, this application provides an encoding device and a decoding device. The device includes the included units and the modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0178] Figure 11 This is a schematic diagram of the composition structure of an encoding device provided in an embodiment of this application, as shown below. Figure 11 As shown, the encoding device 1100 includes: a determining module 1110, an encoding module 1120, and a transmitting module 1130, wherein:
[0179] The determination module 1110 is used to acquire user interaction data and determine the identifier corresponding to the interaction data; the identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user.
[0180] Encoding module 1120 is used to encode the identifier to obtain a high-low frequency combined signal; the high-low frequency combined signal is used to represent the identifier by a combination of at least two signals of different frequencies;
[0181] The transmitting module 1130 is used to transmit the interactive data and the high- and low-frequency combined signal to the decoding end.
[0182] In some embodiments, the identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol; the encoding module 1120 is further configured to encode each of the symbols to obtain a combined signal corresponding to each of the symbols.
[0183] In some embodiments, the encoding module 1120 is further configured to acquire two frequencies corresponding to the symbol; generate an encoded signal corresponding to each frequency based on each frequency; and combine the encoded signals corresponding to each frequency to obtain a combined signal corresponding to the symbol.
[0184] In some embodiments, the identifier includes a first identifier and a second identifier, wherein the first identifier is used to indicate that the interactive data has started to be transmitted and the second identifier is used to indicate that the interactive data has stopped being transmitted; the high- and low-frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier; the transmitting module 1130 is further configured to transmit the first combined signal, the interactive data and the second combined signal to the decoding end in the order of the first combined signal, the interactive data and the second combined signal.
[0185] Figure 12 This is a schematic diagram of the composition structure of a decoding device provided in an embodiment of this application, as shown below. Figure 12 As shown, the decoding device 1200 includes: a receiving module 1210 and a decoding module 1220, wherein:
[0186] The receiving module 1210 is used to receive interactive data and high-low frequency combined signals sent by the encoding end; the high-low frequency combined signals are used to represent an identifier through the combination of signals of at least two different frequencies; the identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user;
[0187] The decoding module 1220 is used to decode the high- and low-frequency combined signal to obtain the identifier corresponding to the interactive data.
[0188] In some embodiments, the identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol; the decoding module 1220 is further configured to decode the combined signal corresponding to each symbol to obtain each symbol.
[0189] In some embodiments, the decoding module 1220 is further configured to split the combined signal corresponding to the symbol to obtain two encoded signals; determine the frequency corresponding to each encoded signal based on each encoded signal; and determine the symbol corresponding to the two frequencies based on the two frequencies.
[0190] In some embodiments, the interactive data is audio; the decoding module 1220 is further configured to determine the translation language of the audio based on the current user corresponding to the identifier; the translation language is the language of another user different from the current user; translate the audio based on the translation language to obtain translated text; perform speech synthesis based on the translated text to obtain synthesized speech, and send the synthesized speech to the hardware terminal.
[0191] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this application can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0192] It should be noted that, in the embodiments of this application, if the above-described encoding or decoding methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0193] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.
[0194] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.
[0195] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.
[0196] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0197] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0198] Figure 13 This application provides a hardware entity diagram of a computer device as an embodiment of the present application, such as... Figure 13 As shown, the hardware entity of the computer device 1300 includes a processor 1301 and a memory 1302, wherein the memory 1302 stores a computer program that can run on the processor 1301, and the processor 1301 executes the program to implement the steps in the method of any of the above embodiments.
[0199] The memory 1302 stores computer programs that can run on the processor. The memory 1302 is configured to store instructions and applications that can be executed by the processor 1301. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1301 and various modules in the computer device 1300. It can be implemented by flash memory or random access memory (RAM).
[0200] The processor 1301 executes the steps of any of the above methods when executing a program. The processor 1301 typically controls the overall operation of the computer device 1300.
[0201] This application provides a computer storage medium that stores one or more programs, which can be executed by one or more processors to implement the steps of the encoding / decoding method as described in any of the above embodiments.
[0202] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0203] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.
[0204] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0205] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. An encoding method, characterized in that, The method is applied at the encoding end, and the method includes: Acquire user interaction data and determine the identifier corresponding to the interaction data; the identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user; The identifier is encoded to obtain a high-low frequency combined signal; the high-low frequency combined signal is used to characterize the identifier by a combination of at least two signals of different frequencies. The interactive data and the high- and low-frequency combined signal are sent to the decoding end.
2. The method according to claim 1, characterized in that, The identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol; encoding the identifier to obtain the high- and low-frequency combined signal includes: Each of the symbols is encoded to obtain the combined signal corresponding to each symbol.
3. The method according to claim 2, characterized in that, The process of encoding each symbol to obtain a combined signal corresponding to each symbol includes: Obtain the two frequencies corresponding to the symbol; Based on each frequency, generate the coded signal corresponding to each frequency; The encoded signals corresponding to each frequency are combined to obtain the combined signal corresponding to the symbol.
4. The method according to any one of claims 1 to 3, characterized in that, The identifier includes a first identifier and a second identifier, wherein the first identifier is used to indicate that the interactive data transmission has started and the second identifier is used to indicate that the interactive data transmission has stopped; the high- and low-frequency combined signal includes a first combined signal corresponding to the first identifier and a second combined signal corresponding to the second identifier. Sending the interactive data and the high- and low-frequency combined signal to the decoding end includes: The first combined signal, the interactive data, and the second combined signal are sent to the decoding end in the order of the first combined signal, the interactive data, and the second combined signal.
5. A decoding method, characterized in that, The method is applied at the decoding end, and the method includes: The system receives interactive data and a high-low frequency combination signal sent by the encoding end; the high-low frequency combination signal is used to represent an identifier through the combination of at least two signals of different frequencies; the identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user. The high- and low-frequency combined signals are decoded to obtain the identifier corresponding to the interactive data.
6. The method according to claim 5, characterized in that, The identifier includes at least one symbol, and the high- and low-frequency combined signal includes a combined signal corresponding to each symbol; decoding the high- and low-frequency combined signal to obtain the identifier corresponding to the interactive data includes: Decode the combined signal corresponding to each symbol to obtain each symbol.
7. The method according to claim 6, characterized in that, Decoding the combined signal corresponding to each symbol to obtain each symbol includes: The combined signal corresponding to the symbol is split to obtain two coded signals; Based on each coded signal, determine the frequency corresponding to each coded signal; Based on the two frequencies, determine the symbols corresponding to the two frequencies.
8. The method according to any one of claims 5 to 7, characterized in that, The interactive data is audio, and the method further includes: Based on the current user corresponding to the identifier, the translation language of the audio is determined; the translation language is the language of another user, different from the current user. The audio is translated based on the specified translation language to obtain the translated text. Based on the translated text, speech synthesis is performed to obtain synthesized speech, which is then sent to the hardware.
9. An encoding device, characterized in that, The device includes: The determination module is used to acquire user interaction data and determine the identifier corresponding to the interaction data; the identifier is used to characterize whether the user corresponding to the interaction data is a near-end user or a far-end user. An encoding module is used to encode the identifier to obtain a high-low frequency combined signal; the high-low frequency combined signal is used to represent the identifier by a combination of at least two signals of different frequencies; The transmitting module is used to send the interactive data and the high- and low-frequency combined signal to the decoding end.
10. A decoding device, characterized in that, The device includes: The receiving module is used to receive interactive data and high-low frequency combined signals sent by the encoding end; the high-low frequency combined signals are used to represent an identifier through the combination of at least two signals of different frequencies; the identifier is used to represent whether the user corresponding to the interactive data is a near-end user or a far-end user; The decoding module is used to decode the high- and low-frequency combined signals to obtain the identifier corresponding to the interactive data.