Audio processing method and system, terminal equipment and storage medium
By compressing, encrypting, and encoding the voice signal into a voice-like signal at a low bit rate, the problem of maintaining confidentiality in existing links is solved, thus enabling low-cost and efficient secure communication.
Patent Information
- Application Number
- CN202510999771.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies are unable to achieve confidentiality of voice calls in existing links, resulting in high investment and maintenance costs.
The voice signal is compressed and encrypted at a low bit rate, and encoded into a voice-like signal using a codebook method or modulation method. After adding a synchronization mark, it is transmitted in the existing call link; the receiving end decodes and decrypts it to recover the voice data.
It achieves low-cost, efficient and confidential calls in existing call links, avoids problems such as encrypted data being filtered and long delays, and ensures the security and real-time nature of information.
Smart Images

Figure CN120808797A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of audio processing, and particularly relates to an audio processing method and system, a terminal device, and a storage medium. BACKGROUND
[0002] In the field of communication with increasing demand for information security, the confidentiality of language communication becomes a key technical problem. In a communication scenario, the transmission of voice information mainly relies on existing links such as ordinary telephone networks, mobile telephone networks, or network telephones (such as WeChat voice).
[0003] In the related art, the confidentiality of voice communication cannot be achieved by using existing links. In the implementation process, a reliable transmission path and the account and dialing call system of a user need to be established, resulting in high cost investment and maintenance cost. SUMMARY
[0004] The voice audio processing method provided in the embodiments of the present application can achieve the confidentiality of voice communication in the existing link of voice communication, and reduce the investment and maintenance cost.
[0005] In a first aspect, the embodiments of the present application provide a voice audio processing method applied to a sending end of a first communication link, the first communication link being used for transmitting voice signals; the method comprises the following steps:
[0006] obtaining encrypted data of a first voice signal; wherein the encrypted data is a non-voice signal;
[0007] encoding the encrypted data to obtain a second voice signal; wherein the second voice signal is a voice-like signal;
[0008] sending the second voice signal to a receiving end of the first communication link through the first communication link.
[0009] In the embodiments of the present application, non-voice data processed by encryption is first obtained. These data do not have voice characteristics after encryption, and then the encrypted data is encoded into a voice-like signal (second voice signal) with natural voice spectrum characteristics, so that it is similar to real voice in waveform and spectrum. Finally, the voice-like signal is sent to the receiving end through the ordinary communication link (first communication link). Since the encrypted data is converted into a voice-like signal, the filtering of non-voice signals by the communication link is avoided, thereby realizing the secret transmission of encrypted information. In the above method, a transmission path for secret data does not need to be established. Only the encrypted data needs to be converted into a voice-like signal, and then the encrypted data can be transmitted by using the existing link, thereby reducing the investment and maintenance cost.
[0010] In a possible implementation manner of the first aspect, the encrypted data of the first voice signal is obtained, comprising:
[0011] Compress the first voice signal to obtain a compressed third voice signal, wherein a bandwidth of the third voice signal conforms to a bandwidth of the first call link.
[0012] Encrypt the third voice signal to obtain encrypted data.
[0013] In the embodiments of the present application, the encrypted data is obtained by first compressing the first voice signal to adapt to the bandwidth of the link and then encrypting the compressed signal, which not only ensures that the signal can be transmitted in the first call link, but also guarantees the security of the voice information.
[0014] In a possible implementation of the first aspect, the second voice information is encoded according to the encrypted data, comprising:
[0015] Obtain a preset codebook, wherein the preset codebook includes a plurality of preset code words each corresponding to a voice segment;
[0016] For any one data byte in the encrypted data, a first code word matching the data byte is obtained from the preset code words in the preset codebook;
[0017] The first voice segment corresponding to the first code word is recorded as a fourth voice signal corresponding to the data byte;
[0018] The second voice signal corresponding to the encrypted data is spliced according to the fourth voice signal corresponding to each data byte in the encrypted data.
[0019] In the embodiments of the present application, by calling the preset codebook containing the preset code words and the corresponding voice segments, each data byte of the encrypted data is matched to the corresponding first code word and the first voice segment (i.e. the fourth voice signal), and then the second voice signal is spliced. This process realizes the conversion of encrypted data to voice-like signal with the help of the codebook, which not only utilizes the transmission capacity of the voice link, but also hides the real form of the encrypted data in the form of voice-like signal, ensuring the feasibility and concealment of data transmission.
[0020] In a possible implementation of the first aspect, the voice-like signal is sent to the receiving end of the first call link through the first call link, comprising:
[0021] A preset mark is added to a preset position of the second voice signal; wherein the preset mark is used to indicate that the second voice signal is an encrypted voice signal;
[0022] The second voice signal and the preset mark are sent to the receiving end of the first call link through the first call link.
[0023] In the embodiment of the present application, the preset mark indicating that the second voice signal is an encrypted voice signal is added at a preset position of the second voice signal and is sent through the first call link, so that the receiving end can identify the encrypted signal and the encrypted information can be effectively transmitted.
[0024] In a second aspect, the embodiment of the present application provides an audio processing method applied to a receiving end of a first call link, the first call link being used for transmitting voice signals; the method comprising:
[0025] receiving a second voice signal sent through the first call link; wherein the second voice signal is a voice-like signal;
[0026] decoding encrypted data from the second voice signal; wherein the encrypted data is a non-voice signal;
[0027] decrypting the first voice signal from the encrypted data.
[0028] In the embodiment of the present application, after receiving the second voice signal which is a voice-like signal, the non-voice encrypted data is obtained by decoding and decryption, and finally the first voice signal is recovered, so that the encrypted voice signal can be accurately received and restored.
[0029] In a possible implementation manner of the second aspect, the encrypted data is decoded from the second voice signal, comprising:
[0030] detecting whether a preset mark exists at a preset position in the second voice signal;
[0031] if the preset mark exists at the preset position of the second voice signal, decoding the encrypted data from the second voice signal.
[0032] In the embodiment of the present application, by detecting whether the preset mark exists at the preset position of the second voice signal, the encrypted signal can be accurately identified and decoding is triggered to obtain the encrypted data, so that the pertinence and accuracy of decoding are ensured.
[0033] In a possible implementation manner of the second aspect, the encrypted data is decoded from the second voice signal, comprising:
[0034] obtaining a preset codebook, wherein the preset codebook comprises a plurality of preset code words each corresponding to a first voice segment;
[0035] dividing the second voice signal into a plurality of fourth voice signals;
[0036] for any one of the fourth voice signals in the second voice signal, obtaining a second voice segment matched with the fourth voice signal from the first voice segment of the preset codebook;
[0037] taking the second code word corresponding to the second voice segment as a data byte corresponding to the fourth voice signal.
[0038] According to the combination of the data bytes corresponding to each fourth voice signal in the second voice signal, the encrypted data corresponding to the second voice signal is obtained.
[0039] In the embodiment of the present application, the fourth voice signals divided from the second voice signal are matched to the corresponding voice segments and code words by the preset codebook, and then converted into data bytes to be combined into encrypted data, so that the accurate restoration of the voice-like signal to the encrypted data is realized.
[0040] In a third aspect, the embodiment of the present application provides an audio processing system, comprising: a sending end, a receiving end and a first call link;
[0041] The sending end is configured to implement the audio processing method of the first aspect described above;
[0042] The receiving end is configured to implement the audio processing method of any one of the second aspect described above.
[0043] In a fourth aspect, the embodiment of the present application provides a terminal device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the audio processing method of any one of the first aspect described above when executing the computer program.
[0044] In a fifth aspect, the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the audio processing method of any one of the first aspect described above.
[0045] In a sixth aspect, the embodiment of the present application provides a computer program product, which, when executed on a terminal device, causes the terminal device to execute the audio processing method of any one of the first aspect described above.
[0046] It can be understood that the beneficial effects of the third aspect to the sixth aspect described above can be referred to the related description of the first aspect and the second aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0048] Figure 1 is the flowchart of the audio processing method provided by the embodiment of the present application Figure 1 ;
[0049] Figure 2 is a flowchart of a process for obtaining encrypted data provided by an embodiment of the present application Figure 1 ;
[0050] Figure 3 is a flowchart of a process for obtaining a second voice signal provided by an embodiment of the present application
[0051] Figure 4 is a flowchart of a process for sending a second voice signal provided by an embodiment of the present application
[0052] Figure 5 is a flowchart of a process for obtaining encrypted data provided by an embodiment of the present application Figure 2 ;
[0053] Figure 6 is a flowchart of a process for obtaining encrypted data provided by an embodiment of the present application Figure 3 ;
[0054] Figure 7 is a flowchart of a process for obtaining encrypted data provided by an embodiment of the present application Figure 4 ;
[0055] Figure 8 is a structural diagram of an audio processing method provided by an embodiment of the present application
[0056] Figure 9 is a general flowchart of an audio processing method provided by an embodiment of the present application
[0057] Figure 10 is a structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0058] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular sequences of steps, techniques, etc., in order to provide a thorough understanding of the present embodiments. However, it will be apparent to those skilled in the art that the present embodiments can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, and circuits are omitted so as not to obscure the description of the present embodiments.
[0059] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or "comprising", when used in this specification and in the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0060] It is also to be understood that the terminology "and / or" when used in this specification and in the following claims, refers to at least one of the items, or any combination of one or more of the items, and includes all possible combinations of one or more of the items.
[0061] As used in the description and the appended claims of the application, the term "if' can be interpreted as meaning "when" or "upon" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [the described condition or event] is detected" can be interpreted as meaning "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.
[0062] In addition, in the description of the application and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0063] In the present application, the reference "one embodiment" or "some embodiments" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in further some embodiments" and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized.
[0064] In the field of communication with increasing demand for information security, the confidentiality of language calls becomes a key technical problem. In the communication scenario, the transmission of voice information mainly relies on existing links such as ordinary telephone network, mobile telephone network or network telephone (such as WeChat voice).
[0065] In the related art, the confidentiality of voice calls cannot be realized by using existing links. In the implementation process, a reliable transmission path and the user's account and dialing call system need to be established, resulting in high cost investment and maintenance cost.
[0066] In order to solve the problems in the above related art, the application embodiment provides an audio processing method, system, terminal device and storage medium. In the present application, secure communication is realized by multiplexing existing call links (such as ordinary telephone network, mobile telephone network, network telephone such as WeChat). At the sending end, the audio data obtained by the microphone is compressed at a low code rate, encrypted (using symmetric or asymmetric algorithm), and a check bit is added. Then the coded voice signal is transmitted by codebook method or modulation method and a synchronization marker is added. At the receiving end, the synchronization signal is identified first, and then the voice signal is decoded, decrypted and decompressed after the check bit, and finally the voice data is recovered. The whole process increases the processing module on the basis of the original call link, solves the problems of encrypted data being filtered and large delay, and realizes low-cost and efficient secure communication.
[0067] Referring to Figure 1 is a flowchart of an audio processing method provided by an embodiment of the present application Figure 1 The method can include the following steps, applied to a sending end of a first call link used to transmit voice signals. By way of example but not limitation, the method can include the following steps:
[0068] S101, obtaining encrypted data of a first voice signal; wherein the encrypted data is a non-voice signal.
[0069] In an embodiment of the present application, the "first voice signal" refers to raw human voice data collected by a microphone at the sending end, which can have undergone preliminary processing (such as removing local loudspeaker echo), and is the core content that needs to be transmitted securely. In the process of securing it, symmetric encryption algorithms (such as AES, DES) or asymmetric encryption algorithms (such as RSA, ECC) can be used for encryption.
[0070] The "encrypted data" obtained after encryption processing is in the form of binary ciphertext, completely losing the time domain (such as waveform fluctuations) and frequency domain (such as 300-3400Hz spectral distribution) characteristics of natural speech, and belongs to a "non-voice signal". This characteristic makes it impossible for the human auditory system to identify it as speech, and distinguishes it from the signal properties of the original speech, but lays the foundation for subsequent conversion into a speech-like signal through encoding and adaptation to the original speech transmission channel, ensuring the confidentiality of the information and avoiding being filtered as noise by the link when directly transmitted.
[0071] In an embodiment, referring to Figure 2 is a flowchart of an embodiment of the present application for obtaining encrypted data Figure 1 As shown in Figure 2 , step S101 includes:
[0072] S1011, compressing the first voice signal to obtain a compressed third voice signal, wherein the third voice signal conforms to the bandwidth of the first call link.
[0073] In an embodiment of the present application, after the sending end obtains the first voice signal (such as human voice collected by a microphone), a vocoder algorithm or deep learning network technology is used for low-code rate compression to obtain a compressed third voice signal.
[0074] The core purpose of this step is to reduce the voice data volume to adapt to the bandwidth limit of the original voice transmission channel (the first call link), while reducing the data load for subsequent encryption, coding and other processing to ensure low latency of real-time communication and avoid the inability to transmit in the existing link due to excessive data volume. The third compressed voice signal still belongs to voice-related data and retains the key information that can be used for subsequent restoration of the original voice, but the code stream is significantly reduced, providing a foundation for the entire secure transmission process.
[0075] S1012, data encryption is performed on the third voice signal to obtain encrypted data.
[0076] In the embodiments of the present application, the "third voice signal" is voice data compressed at a low code rate (compressed from the first voice signal by a vocoder algorithm or a deep learning network), the code stream rate of which has been reduced to the level of adapting to the existing voice transmission channel, retaining the key features of the original voice but greatly reducing the data volume, and is the direct processing object of the encryption operation. When encrypting the third voice signal, mature encryption algorithms can be used to process the third voice signal, including symmetric encryption algorithms (such as AES, DES) or asymmetric encryption algorithms (such as RSA, ECC). Symmetric encryption uses the same key for encryption and decryption, has high operation efficiency and is suitable for real-time communication; asymmetric encryption uses public key encryption and private key decryption, has higher security and more flexible key distribution, and can be selected according to the scene requirements.
[0077] It should be noted that the key required for encryption is stored in a secure area and strictly prohibited from illegal access to ensure the security of the encryption process. If a symmetric algorithm is used, the key can be set manually by the user and synchronized with the call opposite end; if an asymmetric algorithm is used, the key can be automatically generated by the chip or generated through a public key infrastructure (PKI), wherein the public key can be publicly obtained through a certificate or pre-implanted in the system, and the private key is exclusively held by the receiving end, ensuring the reliability of key distribution.
[0078] In addition, after generating the encrypted data, a check bit is generated according to the encrypted data and is supplemented to a preset position of the encrypted data.
[0079] In the above method, the first voice signal is compressed to adapt to the link bandwidth and then encrypted to obtain encrypted data, which not only ensures that the signal can be transmitted in the first call link, but also guarantees the security of the voice information.
[0080] S102, encoding the encrypted data to obtain a second voice signal; wherein the second voice signal is a voice-like signal.
[0081] In the embodiments of the present application, the encrypted data is the result of processing the compressed third voice signal (low-code flow voice data) through a cryptography algorithm (such as AES, RSA, etc.). The essence of the encryption algorithm is to convert the original data (compressed voice in this case) into disordered binary ciphertext through mathematical transformation, which completely destroys the time domain waveform (such as the periodic fluctuations of voice) and frequency domain characteristics (such as the distribution of voice frequency band of 300-3400 Hz) of the original voice signal. The encrypted data directly through the original voice transmission channel will be filtered out as noise or changed to cause the receiving end to be unable to decode and recover.
[0082] In order to make the encrypted data pass through the voice transmission channel, the encrypted data needs to be encoded into a second voice signal through a specific technology, and the second voice signal is a signal with speech-like characteristics. The encoded speech-like signal can be transmitted safely through the original transmission channel. Specifically, the encoding method includes codebook method and modulation method. The specific process of encoding the encrypted data (containing check bits at the preset positions) into a speech-like signal by using the codebook method is shown in the following steps S1021-S1024.
[0083] In one embodiment, referring to Figure 3 , the flowchart for obtaining the second voice signal provided by the embodiments of the present application, as shown in Figure 3 , step S102 includes:
[0084] S1021, obtaining a preset codebook, wherein the preset codebook includes a plurality of preset code words each corresponding to a voice segment.
[0085] In the embodiments of the present application, before encoding the encrypted data, a codebook set containing a plurality of preset code words and corresponding voice segments needs to be constructed and obtained in advance. The "preset code word" refers to a predefined binary data unit (such as byte, bit stream), and the "voice segment" is an equal-length audio segment with speech-like characteristics (such as 20ms of "ah" "oh" or other ambiguous voice or simulated voice waveform signals). Each code word has a one-to-one correspondence with a specific voice segment.
[0086] The codebook needs to be constructed in advance before communication, and is usually agreed by the sending end and the receiving end. When generating, it needs to be ensured that: the consistency of the voice segment (the same length and sampling rate, such as 16kHz sampling, 20ms / segment, facilitating splicing); the speech-like characteristics of the features (the spectrum is close to natural voice, avoiding being filtered by the transmission link); the uniqueness of the mapping (a code word corresponds to only one segment, and vice versa, ensuring that the bidirectional mapping is unambiguous).
[0087] S1022, for any data byte in the encrypted data, obtaining a first code word matched with the data byte from the preset code words of the preset codebook.
[0088] In the embodiment of the present application, for any one data byte (binary data unit) split out of the encrypted data, a first code word (i.e., the corresponding data unit) that completely matches the data byte is found from all the preset code words in the preset codebook.
[0089] In the preset codebook, the preset code words can be defined in advance in hexadecimal form, that is, each hexadecimal value (such as 00 to FF) corresponds to a unique preset code word and an associated voice-like segment. Each data byte in the encrypted data is essentially an 8-bit binary number (such as 10010110), which is split into groups of 4 bits (1001 and 0110) when converted to hexadecimal, corresponding to hexadecimal 9 and 6, that is, the byte is converted to hexadecimal 96, and then the hexadecimal numbers 9 and 6 are matched to the preset code words in the preset codebook, respectively, so that the data byte in the encrypted data can be matched to the first code word.
[0090] The core of this process is to establish the correspondence between the encrypted data and the codebook - in the preset codebook, each preset code word is pre-associated with a specific voice-like segment, and after the data byte is matched to the first code word, the voice segment corresponding to the code word can be called to provide a basis for subsequent splicing to form a continuous voice-like signal (second voice signal).
[0091] S1023 marks the first voice segment corresponding to the first code word as the fourth voice signal corresponding to the data byte.
[0092] In the embodiment of the present application, after finding the first code word that matches a certain data byte in the encrypted data, the first voice segment (equal-length audio segment with voice-like characteristics) corresponding to the first code word in the preset codebook is marked as the fourth voice signal corresponding to the data byte.
[0093] S1024, according to the fourth voice signal corresponding to each data byte in the encrypted data, the second voice signal corresponding to the encrypted data is spliced and generated.
[0094] In the embodiment of the present application, at the sending end, when the encrypted data is split into multiple data bytes, each byte has been mapped to the corresponding fourth voice signal (i.e., the voice-like segment matching the byte, such as a 20ms equal-length audio segment) through the preset codebook. At this time, only the original order of the bytes in the encrypted data is needed to sequentially splice the corresponding fourth voice signals end to end, and a complete continuous signal can be formed - this signal is the second voice signal and belongs to the voice-like signal.
[0095] For example, the encrypted data is split into byte A, byte B, and byte C, which correspond to fourth voice signals S_A, S_B, and S_C (all 20ms voice-like segments), respectively, and the second voice signal generated after splicing is S_A+S_B+S_C (total duration 60ms).
[0096] The splicing process needs to ensure smooth transition between segments (such as matching the spectral characteristics of adjacent segments by designing a preset codebook) to avoid abrupt signal mutations and enhance the "speech-like" properties of the second speech signal, so that it can be recognized as a normal speech signal by the first call link (such as a telephone network), thereby achieving covert transmission of encrypted data. The receiving end can restore the original encrypted data from the second speech signal by reverse splitting and decoding.
[0097] In the above method, each data byte of the encrypted data is matched to the corresponding first code word and first speech segment (i.e., the fourth speech signal) by calling a preset codebook containing preset code words and corresponding speech segments, and then spliced to generate a second speech signal. This process converts encrypted data into speech-like signals with the help of a codebook, which not only utilizes the transmission capacity of a speech link, but also hides the true form of encrypted data in the form of speech-like signals, ensuring the feasibility and concealment of data transmission.
[0098] S103, sending the second speech signal to the receiving end of the first call link through the first call link.
[0099] In the embodiments of the present application, after the encrypted data is converted into a second speech signal (speech-like signal) by the codebook method at the sending end, the signal is sent to the receiving end through the first call link (such as an existing speech call link such as a regular telephone network, a mobile telephone network, WeChat, etc.).
[0100] Specifically, the second speech signal generated by splicing has speech-like characteristics, with a spectral distribution in the speech frequency band of 300-3400 Hz, which can adapt to the transmission characteristics of the first call link and will not be filtered out as noise or changed due to incompatible characteristics. The sending process relies on the transmission mechanism of the original call link, without the need to establish a new transmission path, which realizes the reuse of existing resources, reduces costs, and ensures real-time performance.
[0101] In one embodiment, referring to Figure 4 is a flowchart of sending a second speech signal provided by an embodiment of the present application, as shown in Figure 4 S103 includes:
[0102] S1031, adding a preset marker at a preset position of the second speech signal; wherein the preset marker is used to indicate that the second speech signal is an encrypted speech signal.
[0103] In the embodiments of the present application, the marker is a pre-defined specific (synchronous) signal (such as a special speech-like signal with a fixed length, or a silence, or a Tone signal with a fixed frequency), and its core functions include two aspects: one is identification, which explicitly informs the receiving end that the current transmitted second speech signal is an encrypted signal, rather than an ordinary speech signal; the other is synchronization calibration, which serves as a "time reference" for the receiving end processing, equivalent to a synchronization signal, helping the receiving end to accurately locate the starting point of the signal, the splitting boundary (such as the splicing position of each fourth speech signal), and ensuring that the subsequent decoding can be performed in the correct order.
[0104] The preset marker is inserted into the "preset position" of the second speech signal, which is usually the beginning of the signal (such as the first 100 ms) or a fixed interval point (such as once every 5 seconds). For example, before splicing to generate the second speech signal, a 100 ms preset marker (such as a sine wave signal with a frequency of 1800 Hz) is added at the very beginning, and then the fourth speech signals are spliced to form a complete second speech signal of "preset marker + S1 + S2 + … + S n ”.
[0105] In S1032, the second speech signal and the preset marker are sent to the receiving end of the first call link through the first call link.
[0106] In the embodiments of the present application, the sending end combines the second speech signal (speech-like signal) containing encrypted data information with the preset marker for identification and synchronization, and then sends them to the receiving end of the first call link (such as a normal telephone network, a mobile voice channel, etc.) through the first call link.
[0107] Specifically, the combination form is usually a continuous signal stream of "preset marker + second speech signal" (for example, a 100 ms preset marker signal is sent first, and then the spliced second speech signal is sent immediately after). During transmission, the existing transmission mechanism of the first call link is relied on, and no additional channel needs to be built - since the preset marker and the second speech signal both have the characteristics of a speech frequency band (300-3400 Hz), they can be recognized as "speech signals" by the link and transmitted normally, avoiding being filtered or modified.
[0108] In the above method, the preset marker identifying the encrypted speech signal is added to the second speech signal at a preset position, and then sent through the first call link, which not only facilitates the receiving end to identify the encrypted signal, but also realizes the effective transmission of the encrypted information.
[0109] Referring to Figure 5 , Fig. 1 is a flowchart of an audio processing method provided by an embodiment of the present application. Figure 2 For example, Figure 5As shown, the method applied to a receiving end of a first call link for transmitting voice signals may include the following steps:
[0110] S201, receiving a second voice signal transmitted through the first call link; wherein the second voice signal is a voice-like signal.
[0111] In the embodiment of the present application, the receiving end receives the second voice signal transmitted by the sending end through the first call link (such as a general telephone network, a mobile voice channel, etc.), and the signal is a voice-like signal with voice characteristics.
[0112] Specifically, the hardware of the receiving end (such as a receiver, a voice collection module) will capture the electrical signal transmitted through the link and convert it into a processable audio signal, i.e. the second voice signal. Since the signal is a voice-like signal generated by the sending end by concatenating the codebook, it retains the time-domain waveform characteristics (such as continuous undulating waveform) and frequency-domain distribution (300-3400Hz voice frequency band) of the voice, and therefore can be recognized as a "voice signal" by the link interface of the receiving end and normally received, and will not be discarded as invalid noise.
[0113] S202, decoding the second voice signal to obtain encrypted data; wherein the encrypted data is a non-voice signal.
[0114] In the embodiment of the present application, after the receiving end obtains the second voice signal (voice-like signal), it extracts the original encrypted data through decoding operation, and the encrypted data is still a non-voice signal. The core of decoding is to rely on the preset codebook identical to the sending end to split the voice-like signal into fragments and restore it to the original data bytes (encrypted data). Details are shown in steps S2021-S202 and steps S20221-S20225.
[0115] Referring to Figure 6 , the flowchart for obtaining encrypted data provided by the embodiment of the present application is shown in Figure 3 , for example Figure 6 As shown, step S202 includes:
[0116] S2021, detecting whether a preset mark exists in a preset position of the second voice signal.
[0117] In the embodiment of the present application, after the receiving end receives the voice-like signal (second voice signal), it searches for the synchronization signal (preset mark) first, and if the synchronization signal cannot be found, it does not play sound (mute).
[0118] The search algorithm is: calculating the distance between the audio segment in the second voice signal and a preset segment (a special voice-like signal of a fixed length, or a segment of silence, or a segment of a fixed frequency Tone sound), or analyzing the frequency spectrum. When the distance or the frequency spectrum attribute of the two reaches a certain preset value, it is considered that the synchronization signal and its position are found. To increase confidence, multiple synchronization signals can be searched in succession, and when a plurality of preset number of synchronization signals are identified, it is confirmed that the synchronization signal and its position are found.
[0119] S2022If it is detected that the preset mark exists at the preset position of the second voice signal, the encrypted data is obtained by decoding the second voice signal.
[0120] In the embodiment of the present application, if the preset mark is detected, its position can be used as a time reference (e.g. the end point of the mark is the starting point of the voice-like segment in the second voice signal), helping the receiving end to accurately locate the boundaries of the subsequent split segments and ensuring the timing consistency during decoding.
[0121] After confirming the existence of the preset mark, the end position of the preset mark is taken as the time starting point, and the second voice signal subsequent to the starting point is decoded to obtain the encrypted data.
[0122] In the above method, by detecting whether the preset mark exists at the preset position of the second voice signal, the encrypted signal can be accurately identified and decoding is triggered to obtain the encrypted data, ensuring the pertinence and accuracy of decoding.
[0123] Referring to Figure 7 , the flowchart of obtaining encrypted data by data provided by the embodiment of the present application Figure 4 , as shown in Figure 7 , step S2022 includes:
[0124] S20221, obtaining a preset codebook, wherein the preset codebook includes a plurality of preset code words each corresponding to a first voice segment.
[0125] In the embodiment of the present application, the core of the decoding operation is to rely on the preset codebook which is completely consistent with the sending end. The preset codebook is a structured database which internally stores the one-to-one correspondence between "preset code words" and "first voice segments". For details, see step S1021 described above.
[0126] S20222, dividing the second voice signal into a plurality of fourth voice signals.
[0127] In the embodiment of the present application, after confirming the presence of the preset mark in the second voice signal, the receiving end divides the continuous second voice signal into multiple independent fourth voice signals (i.e., the speech-like segments mapped by the sending end) according to the consistent rule as when the sending end encodes. The key to the division is to follow the "splicing granularity" of the sending end, i.e., the duration of the fourth voice signal, the division boundary and the sending end are strictly consistent.
[0128] When encoding, the sending end defines the fourth voice signal corresponding to each data byte as a speech-like segment of a fixed duration (such as 20 ms), and continuously splices to form the second voice signal in byte order. Therefore, the receiving end needs to divide in units of the same duration (20 ms), and the starting point of the division is usually based on the preset mark, for example, the preset mark is located at the first 100 ms of the second voice signal, and the end of the mark is the starting point of the first fourth voice signal, and a new fourth voice signal is divided every 20 ms.
[0129] Specifically, the receiving end determines the starting time of the division through the position of the preset mark (such as the end of the mark is at the 100th ms, then the division starts from the 100th ms), and ensures the alignment of the starting point of the splicing with the sending end. From the starting point, the second voice signal is uniformly divided according to the preset segment duration (such as 20 ms). For example, if the total duration of the second voice signal is 500 ms (400 ms is left after deducting the 100 ms mark), then 20 fourth voice signals can be divided (400 ms ÷ 20 ms per = 20). The multiple fourth voice signals obtained after the division need to strictly preserve the original time sequence (such as the 1st segment corresponds to the 1st byte of the encrypted data, and the 2nd segment corresponds to the 2nd byte), and the splicing order is completely consistent with the sending end.
[0130] S20223, for any fourth voice signal in the second voice signal, obtaining a second voice segment matched with the fourth voice signal from the first voice segments of the preset codebook.
[0131] In the embodiment of the present application, the preset codebook stores the "feature templates" (feature parameters pre-stored when the sending end generates the segments) of all the first voice segments. The receiving end compares the feature parameters of the fourth voice signal with the feature templates of each first voice segment in the codebook one by one, calculates the similarity (such as through the cosine similarity, Euclidean distance, etc. Algorithm), and if the similarity of the feature template of a certain first voice segment and the feature parameters of the fourth voice signal exceeds a preset threshold (such as 95%), it is determined that the first voice segment is "a second voice segment matched with the fourth voice signal".
[0132] It should be noted that the "second voice segment" here is not a newly generated signal, but a functional naming of the "first voice segment matching the fourth voice signal" in the codebook - in essence, it is the original voice-like segment corresponding to the data byte during encoding at the sending end (i.e., the "first voice segment" at the sending end, which is called the "second voice segment" after matching at the receiving end, only to distinguish the processing object at the receiving end). For example, the first voice segment S96 at the sending end becomes the fourth voice signal S'96 after transmission, and the receiving end finds S96 in the codebook that matches the characteristics of S'96, at which point S96 is the "second voice segment".
[0133] S20224, the second code word corresponding to the second voice segment is recorded as the data byte corresponding to the fourth voice signal.
[0134] In the embodiments of the present application, after finding the second voice segment matching the fourth voice signal, the second code word corresponding to the second voice segment in the preset codebook is marked as the original data byte corresponding to the current fourth voice signal.
[0135] In the preset codebook, each second voice segment (i.e., the first voice segment at the sending end) is fixedly associated with a unique second code word (i.e., the preset code word at the sending end), and the second code word directly corresponds to the original data byte (such as an 8-bit binary value or a hexadecimal value). For example, there is an entry in the codebook: second voice segment S45 → second code word C45 → data byte 0x2D (binary 00101101).
[0136] After the receiving end determines that the fourth voice signal matches the second voice segment (such as S45), it can obtain the second code word (C45) corresponding to the segment by querying the codebook, and according to the mapping relationship of "second code word-data byte" in the codebook, mark C45 as the original data byte (binary 00101101) corresponding to the current fourth voice signal.
[0137] S20225, according to the data byte corresponding to each fourth voice signal in the second voice signal, the encrypted data corresponding to the second voice signal is combined.
[0138] In the embodiments of the present application, the receiving end has already split the second voice signal into multiple fourth voice signals (such as S1, S2, S3…S n ), and each fourth voice signal has restored the corresponding original data byte (such as B1, B2, B3…B n), because the sending end is in the encoding process is according to the "data byte order → speech segment order → second speech signal" flow generated signal, so the receiving end needs to be restored in the opposite order: the timing sequence of the fourth speech signal = the original order of the data byte. For example, the first S1 in the second speech signal corresponds to the first B1 in the encrypted data, and the second S2 corresponds to B2, and so on.
[0139] The receiving end splices all data bytes B1, B2, B3…B n According to the appearance order of the corresponding fourth speech signal, splice in turn (such as B1+B2+B3+…+B n ), forming a continuous binary data stream, which is the complete encrypted data corresponding to the second speech signal.
[0140] In the above method, the fourth speech signal divided from the second speech signal by the preset codebook is matched as the corresponding speech segment and code word, and then converted into data bytes to combine encrypted data, realizing accurate restoration of the encrypted data from the speech signal.
[0141] S203, decrypting the first speech signal according to the encrypted data.
[0142] In the embodiments of the present application, the check bits are separated from the decoded data. The check value such as the hash value of the decoded encrypted data is calculated, and the calculated check value is compared with the separated check bits. If they are consistent, decryption is performed; if they are not consistent, they are discarded, and logs are recorded for subsequent analysis and optimization.
[0143] The receiving end needs to use the same decryption algorithm as the sending end (such as AES decryption corresponding to AES encryption), and input the correct key (the same key as the sending end in symmetric encryption, or the private key of the receiving end in asymmetric encryption). The key is the core of decryption, and if the key is wrong or the algorithm is not matched, the correct original speech signal cannot be restored. The receiving end inputs the complete encrypted data obtained by combining into the decryption algorithm, and performs inverse operation (such as AES round decryption operation) on the ciphertext through the key, gradually removes the encryption layer, and finally restores the original data before encryption - the third speech signal compressed by the sending end (if compressed before encryption), and then obtains the first speech signal through decompression processing (such as symmetric audio decompression algorithm with the sending end).
[0144] In the above method, after receiving the second speech signal of the receiving speech, the encrypted data is obtained by decoding and decryption, and the first speech signal is finally restored, realizing accurate receiving and restoration of the encrypted speech signal.
[0145] The present application also provides an audio processing system, comprising a first call link, the first call link being used for transmitting speech signals; the first call link comprising a sending end and a receiving end;
[0146] The sending end is configured to implement the audio processing method of any one of steps S101-S103, S1011-S1012, S1021-S1024, and steps S1031-S1032.
[0147] The receiving end is configured to implement the audio processing method of any one of steps S201-S203, S2021-S2022, and steps S20221-S20225.
[0148] Referring to Figure 8 , which is a structural schematic diagram of the audio processing method provided in the present application, as shown in Figure 8 , the part in the green box is the secret communication processing added on the existing communication link. At the speaking end (sending end), the present scheme adds a voice compression, encryption, and encoding module between the microphone and the uplink voice processing unit (generally a DSP). At the receiving end, a voice decoding, decryption, and decompression module is added between the loudspeaker and the downlink voice processing unit. When it is a duplex voice communication, only one reverse peer processing path needs to be added. That is, both ends of the communication have the two added modules.
[0149] Referring to Figure 9 , which is a structural schematic diagram of the audio processing method provided in the present application, as shown in Figure 9 , the steps of the secret audio processing method include:
[0150] The sending end:
[0151] ① The sending end obtains the microphone data (first voice signal).
[0152] Or the microphone data is removed from the data of the local loudspeaker echo. The audio data is compressed and encoded at a low code rate. The algorithm adopted here is some vocoder algorithm or deep learning network, which compresses the original audio from a high code rate to a low code rate (third voice signal).
[0153] ② Encryption and addition of check bits.
[0154] The encryption algorithm is a symmetric algorithm or an asymmetric encryption algorithm. The symmetric algorithm is, for example, AES or DES, and the asymmetric algorithm is, for example, RSA or ECC. The key library is stored in a secure area and is prohibited from being accessed illegally. The key in the symmetric algorithm can be set manually by the user. The key in the asymmetric encryption algorithm can be generated by the chip itself or generated and implanted through the public key infrastructure (PKI). The public key is used for encryption, and the private key is used for decryption. The public key of the communication partner can be published in the form of a certificate in a public way for the speaker to obtain; or it can be pre-implanted into the system of the speaker. A check bit is generated for the encrypted data and supplemented to a preset position of the data.
[0155] ③ Encode into a voice-like signal.
[0156] The entire encrypted data and the check bit are encoded into a voice-like signal (second voice signal) using a codebook method and a modulation method. The codebook method is to generate a set of codebooks in advance, which are equal-length segments of voice-like signals. The data bytes are mapped into these voice segments one by one, and spliced together. The modulation method uses phase and amplitude modulation methods to modulate the data into a voice frequency band (300-3400 Hz) to generate a voice-like signal.
[0157] ④ Add a synchronization marker to generate an audio code stream.
[0158] The generated voice-like signal data is supplemented with a synchronization signal (preset marker) in front or behind. The synchronization signal can be a special voice-like signal of a fixed length, or a segment of silence, or a segment of a tone of a fixed frequency.
[0159] ⑤ Transmit to the receiving end through a transmission channel.
[0160] Finally, the entire signal data (second voice signal and synchronization signal) is transmitted out through the original audio transmission channel (first call link).
[0161] Receiving end:
[0162] ① Receive the audio code stream and identify the synchronization marker.
[0163] The receiving end receives the second voice signal and the synchronization signal, and first searches for the synchronization signal. If the synchronization signal cannot be found, no sound is played (silence). The search algorithm is to calculate the distance between the audio segment of the second voice signal and the preset segment (a special voice-like signal of a fixed length, or a segment of silence, or a segment of a tone of a fixed frequency), or to analyze the frequency spectrum. When the distance or the frequency spectrum attribute of the two reaches a certain preset value, it is considered that the synchronization signal and its position have been found. Multiple synchronization signals can be searched continuously, and when a plurality of preset numbers of synchronization signals are identified, it is considered that the synchronization signal and its position have been found.
[0164] ② Decode data from the speech-like signal.
[0165] After identifying the synchronization signal, the speech-like signal is separated. It is decoded. If the codebook method, the identification is cut into multiple equal segments again, and for each segment, it is found which segment in the preset codebook set it is most similar to, and then the data corresponding to the codebook is decoded. If the modulation method, the phase and amplitude are demodulated, and the data is decoded.
[0166] ③ Check the parity, and if legal, decrypt the data.
[0167] From the decoded data, the data and the parity are separated. The parity value generated by the parity and the data of the parity are consistent. If consistent, decrypt; if not, discard and record the log. Facilitate subsequent analysis and tuning.
[0168] The decryption algorithm is the corresponding processing of the encryption algorithm. When symmetric encryption, the key is managed by the sending end user and shared with the opposite end, and set into the system by the receiving end user. The decryption key (private key) of asymmetric encryption already exists in the system. The encryption and decryption algorithm is the common security algorithm, and the decrypted data (first speech signal) is obtained.
[0169] ④ Decompression.
[0170] The decrypted data is the compressed speech signal data (third speech signal). The data is decompressed, and the corresponding vocoder recovery speech algorithm or deep neural network algorithm is used to recover the speech data (first speech signal).
[0171] ⑤ The receiving end obtains the speech signal.
[0172] The speech data (after the first speech signal) is sent to the downstream speech processing end. According to the needs, it can be enhanced again, or directly sent to the D / A conversion and power amplifier of the loudspeaker to drive the loudspeaker to play the speech.
[0173] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0175] Figure 10 This is a schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 10 As shown, the terminal device 10 of this embodiment includes: at least one processor 100 ( Figure 10 Only one is shown in the figure) a processor, a memory 101, and a computer program 102 stored in the memory 101 and capable of running on at least one processor 100. When the processor 100 executes the computer program 102, the steps in any of the above-mentioned audio processing method embodiments are implemented.
[0176] The terminal device can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 10 This is merely an example of the terminal device 10 and does not constitute a limitation on the terminal device 10 . The terminal device 10 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device 10 may also include input and output devices, network access devices, etc.
[0177] The processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0178] The memory 101 can be an internal storage unit of the terminal device 10, such as a hard disk or a memory of the terminal device 10 in some embodiments. The memory 101 can also be an external storage device of the terminal device 10, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like equipped on the terminal device 10 in other embodiments. Further, the memory 101 can include both an internal storage unit and an external storage device of the terminal device 10. The memory 101 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of computer programs, and the like. The memory 101 can also be used to temporarily store data that has been output or is to be output.
[0179] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above-mentioned various method embodiments.
[0180] The embodiments of the present application provide a computer program product. When the computer program product is run on a terminal device, the terminal device is caused to implement the steps in the above-mentioned various method embodiments.
[0181] The integrated units, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the embodiments of the present application can implement all or part of the above-mentioned method embodiments by a computer program to instruct related hardware to complete. The computer program can be stored in a computer readable storage medium and can implement the steps in the above-mentioned various method embodiments when executed by a processor. The computer program includes computer program codes, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer readable medium at least includes any entity or device capable of carrying the computer program codes to the apparatus / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunications signal.
[0182] In the above embodiments, the description of each embodiment focuses on different aspects, and the parts not described in detail or recorded in a certain embodiment can be referred to the relevant description of other embodiments.
[0183] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0184] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented by other ways. For example, the apparatus / terminal device embodiments described above are only schematic, and the division of the modules or units is only a logical function division, and there can be another division way in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0185] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected to achieve the purpose of the embodiments according to actual needs.
[0186] The above embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An audio processing method, characterized in that: The method is applied to a transmitting end of a first call link, where the first call link is used to transmit a voice signal; the method includes: Acquire encrypted data of a first voice signal; wherein the encrypted data is a non-voice signal; A second voice signal is obtained by encoding the encrypted data; wherein the second voice signal is a voice-like signal; The second voice signal is sent to a receiving end of the first call link through the first call link.
2. The audio processing method according to claim 1, wherein: The step of obtaining encrypted data of the first voice signal includes: compressing the first voice signal to obtain a compressed third voice signal; wherein the bandwidth of the third voice signal matches the bandwidth of the first call link; The third voice signal is encrypted to obtain the encrypted data.
3. The audio processing method according to claim 1, wherein: The step of encoding the encrypted data to obtain the second voice information includes: Obtaining a preset codebook, wherein the preset codebook includes speech segments corresponding to respective preset codewords; For any data byte in the encrypted data, obtaining a first codeword matching the data byte from preset codewords in the preset codebook; Recording the first voice segment corresponding to the first codeword as the fourth voice signal corresponding to the data byte; A second voice signal corresponding to the encrypted data is generated by splicing the fourth voice signal corresponding to each data byte in the encrypted data.
4. The audio processing method according to claim 1, wherein: The sending the voice-like signal to a receiving end of the first call link through the first call link includes: Adding a preset mark at a preset position of the second voice signal; wherein the preset mark is used to indicate that the second language signal is an encrypted voice signal; The second voice signal and the preset mark are sent to a receiving end of the first call link through the first call link.
5. An audio processing method, characterized in that: The method is applied to a receiving end of a first call link, where the first call link is used to transmit a voice signal; the method includes: receiving a second voice signal sent through the first call link; wherein the second voice signal is a voice-like signal; Decoding the second voice signal to obtain encrypted data; wherein the encrypted data is a non-voice signal; The first voice signal is obtained by decrypting the encrypted data.
6. The audio processing method according to claim 5, wherein: The decoding the second voice signal to obtain encrypted data includes: detecting whether a preset mark exists at a preset position in the second voice signal; If it is detected that the preset mark exists at the preset position of the second voice signal, the encrypted data is obtained by decoding the second voice signal.
7. The audio processing method according to claim 5, wherein: The decoding the second voice signal to obtain encrypted data includes: Obtaining a preset codebook, wherein the preset codebook includes a first speech segment corresponding to each of a plurality of preset codewords; dividing the second speech signal into a plurality of fourth speech signals; For any fourth voice signal in the second voice signal, obtaining a second voice segment matching the fourth voice signal from the first voice segment in the preset codebook; Recording the second codeword corresponding to the second voice segment as the data byte corresponding to the fourth voice signal; The encrypted data corresponding to the second voice signal is obtained by combining the data bytes corresponding to each of the fourth voice signals in the second voice signal.
8. An audio processing system, characterized in that: The first call link includes a first call link for transmitting a voice signal; the first call link includes a transmitting end and a receiving end; The transmitting end is used to implement the audio processing method according to any one of claims 1 to 4 above; The receiving end is used to implement the audio processing method as described in any one of claims 5 to 7 above.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.