Audio file processing method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202311734028.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-12-15
AI Technical Summary
[0004]本发明实施例提供了一种音频文件处理方法,以解决音频文件需要长期存储,将占用大量存储空间,增加了运营商的成本的问题
[0025]在本发明实施例中,获取人机交互过程中产生的音频文件和音频文件对应的描述文件,其中,描述文件中至少包括人工合成音频的描述信息,根据描述信息将音频文件拆分为人工合成音频和非人工合成音频,并将人工合成音频转换为文本信息,然后,若识别到非人工合成音频中包括敏感信息,则将非人工合成音频拆分为第一音频和包含敏感信息第二音频,并在描述文件中记录第一音频和第二音频分别对应的描述信息,对第二音频进行加密得到加密后的第二音频,然后,保存文本信息、第一音频和加密后的第二音频以及描述文件,如此,后续可以根据文本信息、第一音频和加密后的第二音频以及描述文件追溯音频文件。本发明实施例对于人机交互过程中产生的音频文件中的人工合成音频采用文本信息来进行记录,因此可以减少音频文件长期存储时所需要占用的存储空间,降低了运营商的成本。
Smart Images

Figure CN117877473B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to an audio file processing method and apparatus, an electronic device, and a storage medium. Background Technology
[0002] In various industries that provide telecommunications and internet services, such as telecom operators, frequent communication with users is necessary to understand their needs and provide corresponding services. Voice communication is the most common method, such as human customer service and AI (Artificial Intelligence) intelligent outbound calls. During voice communication, there are often situations where users need to provide sensitive information such as mobile phone numbers or ID card numbers. The audio files (or voice files) containing sensitive information generated during voice communication will be recorded and stored on the operator's servers for centralized management. They can be retrieved and listened to again when complaints, obstacles, or disputes arise. However, the retrieval process involves multiple layers of transmission, which may lead to security issues and the leakage of users' sensitive information.
[0003] However, audio file storage itself puts enormous pressure on servers. According to statistics, audio files alone require more than 10TB (Terabyte, a unit of computer storage capacity) of storage space per day. Moreover, these audio files need to be stored for a long time, which will occupy a lot of storage space and increase the cost for operators. Summary of the Invention
[0004] This invention provides an audio file processing method to solve the problem that audio files need to be stored for a long time, which will occupy a lot of storage space and increase the cost of operators.
[0005] Accordingly, embodiments of the present invention also provide an audio file processing device, an electronic device, and a storage medium to ensure the implementation and application of the above methods.
[0006] To address the aforementioned problems, this invention discloses an audio file processing method, the method comprising:
[0007] Obtain the audio file generated during human-computer interaction and the corresponding description file of the audio file; the description file includes at least the description information of the artificially synthesized audio.
[0008] Based on the description information, the audio file is split into artificially synthesized audio and non-artificially synthesized audio;
[0009] Convert the artificially synthesized audio into text information;
[0010] If sensitive information is detected in the non-synthesized audio, the non-synthesized audio is split into a first audio and a second audio; the second audio is the audio containing the sensitive information.
[0011] The description file records the description information corresponding to the first audio and the second audio respectively;
[0012] The second audio is encrypted to obtain the encrypted second audio.
[0013] The text information, the first audio, the encrypted second audio, and the description file are saved to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file.
[0014] This invention also discloses an audio file processing apparatus, the apparatus comprising:
[0015] The acquisition module is used to acquire audio files generated during human-computer interaction and description files corresponding to the audio files; the description files include at least descriptive information of artificially synthesized audio.
[0016] The first splitting module is used to split the audio file into artificially synthesized audio and non-artificially synthesized audio according to the description information;
[0017] A conversion module is used to convert the artificially synthesized audio into text information;
[0018] The second splitting module is used to split the non-synthesized audio into a first audio and a second audio if sensitive information is detected in the non-synthesized audio; the second audio is the audio containing the sensitive information.
[0019] The recording module is used to record the description information corresponding to the first audio and the second audio respectively in the description file;
[0020] An encryption module is used to encrypt the second audio to obtain the encrypted second audio.
[0021] A storage module is used to store the text information, the first audio, the encrypted second audio, and the description file, so as to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file.
[0022] This invention also discloses an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform one or more audio file processing methods as described in this invention.
[0023] This invention also discloses one or more machine-readable media storing executable code thereon, which, when executed, causes a processor to perform one or more audio file processing methods as described in this invention.
[0024] The embodiments of the present invention have the following advantages:
[0025] In this embodiment of the invention, audio files and corresponding description files generated during human-computer interaction are obtained. The description file includes at least descriptive information about artificially synthesized audio. Based on the description information, the audio file is split into artificially synthesized audio and non-artificially synthesized audio. The artificially synthesized audio is converted into text information. If sensitive information is detected in the non-artificially synthesized audio, it is split into a first audio and a second audio containing the sensitive information. The description information corresponding to the first and second audios is recorded in the description file. The second audio is encrypted to obtain an encrypted second audio. Then, the text information, the first audio, the encrypted second audio, and the description file are saved. Thus, the audio file can be traced later based on the text information, the first audio, the encrypted second audio, and the description file. This embodiment of the invention uses text information to record artificially synthesized audio in audio files generated during human-computer interaction, thereby reducing the storage space required for long-term storage of audio files and lowering the operator's costs.
[0026] It should also be noted that since artificially synthesized audio is generated through artificial intelligence technology, the accuracy of recognizing artificially synthesized audio as text information and then restoring it back to artificially synthesized audio from the text information can reach 100%, thus ensuring the accuracy of audio file restoration. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating the steps of an embodiment of an audio file processing method according to the present invention;
[0028] Figure 2 This is a schematic diagram of a second audio desensitization process provided in an embodiment of the present invention;
[0029] Figure 3 This is a schematic diagram of a second audio restoration process provided in an embodiment of the present invention;
[0030] Figure 4 This is a schematic diagram of the main process of audio file processing provided in an embodiment of the present invention;
[0031] Figure 5 This is a structural block diagram of an embodiment of an audio file processing device according to the present invention;
[0032] Figure 6 This is a schematic diagram of the structure of a device provided in an embodiment of the present invention. Detailed Implementation
[0033] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] In relevant technical solutions, the following three methods are typically used to process audio files:
[0035] 1. De-identify the audio files, and then compress and store the original audio files and the de-identified audio files together. This method can solve the problems of information security and business availability, but it requires twice the storage space, which will significantly increase server costs.
[0036] 2. Only saving the anonymized audio files will result in the inability to restore the complete audio file when the original audio file is needed, which will not meet business requirements.
[0037] 3. Saving only the original audio file results in the user's sensitive information not being effectively controlled, which may lead to the leakage of sensitive information and bring security risks.
[0038] It can be seen that none of the above three methods can simultaneously meet the requirements of security, business availability, and cost control.
[0039] To address the problems existing in related technical solutions, this invention provides an audio file processing method that solves various problems in the storage and transmission of audio files generated during human-computer interaction, and can simultaneously meet the requirements of security, business availability, and cost control.
[0040] Reference Figure 1 This is a flowchart illustrating the steps of an embodiment of an audio file processing method according to the present invention, including the following steps:
[0041] Step 101: Obtain the audio file generated during the human-computer interaction process and the description file corresponding to the audio file; the description file includes at least the description information of the artificially synthesized audio.
[0042] During human-computer interaction, operators can use TTS (Text-to-Speech) technology and other artificial intelligence technologies to convert text into synthesized audio for voice communication with users. After the voice communication ends, an audio file containing both synthesized and non-synthetic audio (the user's voice) will be obtained, along with a corresponding description file. The description file must include at least descriptive information for both the synthesized and non-synthetic audio, including at least the start time and duration of each audio segment.
[0043] For example, in a telecom operator's intelligent outbound calling system, at 15:00 on January 1, 2020, the operator initiated an intelligent outbound call to user Zhang San, mobile number 18900000000. During the voice communication, the outbound AI (Artificial Intelligence) could not answer user Zhang San's question, so the AI was automatically connected to conduct the voice communication. The entire voice communication took 10 seconds. After the voice communication ended, an audio file was obtained, with a total size of approximately 1.67MB. The artificial voice (artificially synthesized audio) occupied about 500KB of storage space in the first 3 seconds of the audio file. In the last 2 seconds of the audio file, the user recited his ID card number. Since the ID card number is sensitive information, the audio file was judged to contain sensitive information and needed to be de-identified to ensure the user's information security.
[0044] After the voice communication is completed, the system obtains an audio file audio_temp.mp3 and a corresponding description file. The description file can contain description information corresponding to the artificially synthesized audio. The description information corresponding to the artificially synthesized audio can be TTS_audio_1=[0,3000] (meaning it starts at second 0 and lasts for 3 seconds). The audio file is temporarily stored in the system cache.
[0045] It should be noted that the embodiments of the present invention may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).
[0046] Step 102: Based on the description information, split the audio file into artificially synthesized audio and non-artificially synthesized audio.
[0047] Step 103: Convert the artificially synthesized audio into text information.
[0048] In this embodiment of the invention, the description information includes description information corresponding to artificially synthesized audio. Specifically, the description information includes the start time and duration of the artificially synthesized audio in the entire audio file. Therefore, by using the description information corresponding to the artificially synthesized audio, the audio file can be accurately divided into artificially synthesized audio and non-artificially synthesized audio.
[0049] For example, the audio file audio_temp.mp3 is read from the system cache. The description information in the description file of audio_temp.mp3 indicates that the first 3 seconds of the entire audio file are artificially synthesized audio. At this time, the audio file needs to be split into two parts: the first 3 seconds of audio and the last 7 seconds of audio. The first 3 seconds of audio is directly converted into text by ASR (Automatic Speech Recognition) (since artificially synthesized audio is generated by artificial intelligence, its recognition rate can reach 100%). The corresponding text information TTS_audio_1.txt is then generated based on the text.
[0050] Step 104: If sensitive information is detected in the non-synthesized audio, the non-synthesized audio is split into a first audio and a second audio; the second audio is the audio containing the sensitive information.
[0051] Sensitive information may include ID card numbers, user account passwords, or bank card passwords, etc.
[0052] In this embodiment of the invention, if sensitive information is detected in the non-synthesized audio within an audio file, the non-synthesized audio is split into a first audio file and a second audio file containing the sensitive information. The second audio file is then encrypted. This eliminates the need to encrypt (de-sensitize) the entire audio file. Therefore, when tracing the original audio file later, only the second audio file needs to be decrypted, improving the processing efficiency of audio files. Of course, if the non-synthesized audio does not contain sensitive information, then splitting and encrypting the non-synthesized audio is unnecessary; direct compression is sufficient.
[0053] For example, for the audio file audio_temp.mp3, the remaining 7 seconds of audio are parsed using ASR. Based on the dialogue text information parsed by ASR, the first 5 seconds of audio are identified as the first audio of ordinary dialogue information, and the last 2 seconds of audio are the second audio involving sensitive information, namely, the sensitive information of user Zhang San's ID card number. Therefore, the audio file audio_temp.mp3 will be split, the first 5 seconds of audio will be temporarily stored as audio_1.mp3, and the last 2 seconds of audio will be temporarily stored as audio_2.mp3.
[0054] Step 105: Record the description information corresponding to the first audio and the second audio in the description file.
[0055] In this embodiment of the invention, after splitting the audio file to obtain the first audio and the second audio, it is also necessary to record the description information corresponding to the first audio and the second audio respectively in the description file. When it is necessary to restore the audio later, it can be restored according to the description information corresponding to the first audio and the second audio.
[0056] For example, after obtaining the first audio file audio_1.mp3 and the second audio file audio_2.mp3, the description information for the first audio file audio_1.mp3 is recorded in the description file as audio_1 = [3001, 4998] (meaning it starts at 3.001 seconds and lasts for 4.998 seconds), and the description information for the second audio file audio_2.mp3 is recorded in the description file as audio_2 = [8000, 2000] (meaning it starts at 8 seconds and lasts for 2 seconds).
[0057] Step 106: Encrypt the second audio to obtain the encrypted second audio.
[0058] In this embodiment of the invention, to meet security requirements, the second audio file needs to be encrypted to obtain the encrypted second audio file, thus achieving desensitization of sensitive information in the audio file. In an optional example, the encryption method for the second audio file can use a byte shift algorithm. Specifically, the byte shift algorithm is a shift operation algorithm used for encrypting and decrypting data. The byte shift algorithm shifts each bit of each byte of the encrypted object by a specified shift amount. Of course, other encryption methods besides the byte shift algorithm can also be used in practical applications, and this embodiment of the invention does not impose any restrictions on this.
[0059] Step 107: Save the text information, the first audio, the encrypted second audio, and the description file to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file.
[0060] In this embodiment of the invention, by processing the audio file, the cache should store one text message, two split audio files, and one description file, which are:
[0061] Text information: TTS_audio_1.txt;
[0062] The split audio files are: audio_1.mp3 and audio_2.mp3.
[0063] The description file contains the following content: TTS_audio_1 = [0, 3000], audio_1 = [3001, 4998], audio_2 = [8000, 2000].
[0064] Furthermore, in this embodiment of the invention, the description file can be named 20200101150000_18900000000_Zhang San.txt. Then, the text information and the split audio files are packaged and compressed, named 20200101150000_18900000000_Zhang San.tar.gz. Finally, the two files, 20200101150000_18900000000_Zhang San.txt and 20200101150000_18900000000_Zhang San.tar.gz, are stored on the disk, and a sequence is created for them. Subsequently, if user Zhang San has a dispute over the order, he can obtain two files from the disk: 20200101150000_18900000000_Zhang San.txt and 20200101150000_18900000000_Zhang San.tar.gz. Based on these two files, he can reconstruct the audio file, which can then be used as evidence to resolve the dispute.
[0065] It should be noted that, since this embodiment of the invention converts the artificially synthesized audio (TTS) in the audio file into text information before storage, it solves the problem of audio files occupying a lot of storage space. The text information of the artificially synthesized audio, TTS_audio_1.txt, occupies only 45 bytes of storage space (about 15 Chinese characters), which is more than 11,000 times less than the 500KB space occupied by video storage, and 30% less than the 1.67MB storage space of the entire audio file. Therefore, in scenarios containing a large number of TTS synthesized audio files, it significantly solves the problem of storage space occupation and saves storage space.
[0066] In the above-described audio file processing method, an audio file generated during human-computer interaction and its corresponding description file are obtained. The description file includes at least descriptive information about artificially synthesized audio. Based on this description information, the audio file is split into artificially synthesized audio and non-artificially synthesized audio. The artificially synthesized audio is converted into text information. If sensitive information is detected in the non-artificially synthesized audio, it is split into a first audio and a second audio containing the sensitive information. The description information corresponding to the first and second audios is recorded in the description file. The second audio is encrypted to obtain an encrypted second audio. Finally, the text information, the first audio, the encrypted second audio, and the description file are saved. This allows for subsequent tracing of the audio file based on the text information, the first audio, the encrypted second audio, and the description file. This embodiment of the invention uses text information to record artificially synthesized audio in audio files generated during human-computer interaction, thus reducing the storage space required for long-term storage of audio files and lowering the operator's costs. It should also be noted that since artificially synthesized audio is generated through artificial intelligence technology, the accuracy of recognizing artificially synthesized audio as text information and then restoring it back to artificially synthesized audio from the text information can reach 100%, thus ensuring the accuracy of audio file restoration.
[0067] In one embodiment of the present invention, before step 104, which involves splitting the non-synthesized audio into a first audio and a second audio if sensitive information is identified in the non-synthesized audio, the method further includes:
[0068] The non-synthesized audio is converted into dialogue text information, and it is then identified whether the dialogue text information contains sensitive information.
[0069] In this embodiment of the invention, the non-synthesized audio in the audio file is parsed using ASR technology to obtain the corresponding dialogue text information. Then, sensitive information is identified through the dialogue text information to determine whether the non-synthesized audio contains sensitive information. For example, if the dialogue text information includes sensitive information such as an ID card number, it can be determined that there is sensitive information in the audio file, and the second audio in the audio file containing sensitive information needs to be encrypted. If the dialogue text information does not include sensitive information such as an ID card number, it can be determined that there is no sensitive information in the audio file, and there is no need to encrypt the audio file, nor is it necessary to split the audio file.
[0070] In one embodiment of the present invention, step 106, encrypting the second audio to obtain the encrypted second audio, includes:
[0071] Obtain the user information of the user corresponding to the audio file, and generate a key string using the user information; wherein, the user information includes at least the user's name and the user's incoming mobile phone number;
[0072] Convert the second audio file into a byte array;
[0073] The key string is identified as a value in a specified base and then converted into a binary value with a specified number of bits; wherein, the specified base includes at least base 65 and the specified number of bits includes at least eight bits;
[0074] A new byte array is obtained by performing data processing operations on the byte array using the binary values; wherein the data processing operations include at least modulo, XOR, and reversal;
[0075] Reassemble the new byte array into a byte matrix;
[0076] The square root of the byte matrix is then rounded up to obtain an integer value.
[0077] Adjust the byte matrix to a target byte matrix where the number of groups in each row is the integer value;
[0078] The target byte matrix is rotated to obtain a new byte matrix; wherein the rotation includes at least a clockwise rotation of the boundary.
[0079] Transform the new byte matrix into a target byte array;
[0080] The target byte array is then encrypted to obtain the second audio.
[0081] The specified base must include at least base-65 (base-65), and the specified number of bits must include at least eight bits (8 bits).
[0082] In this embodiment of the invention, the second audio is encrypted using a byte shifting algorithm. The specific process may include: converting the second audio into a byte array; using user information such as the phone number of the caller as characters for AES (Advanced Encryption Standard) encryption to obtain ciphertext (key string); converting the ciphertext into a string of base-65 numbers and performing transcoding to obtain a new byte array (binary value) of 8 bits per group; performing basic encryption processing such as modulo and XOR operations on the byte array of the second audio using the binary value to obtain a new byte array; then reassembling the new byte array into a byte matrix; performing shifting operations on the byte matrix, rotating the boundary binary numbers clockwise by 1 bit to obtain the target byte matrix; finally converting the target byte matrix back into an array to obtain the target byte array; and then converting the target byte array into an audio stream, which is the encrypted second audio.
[0083] For example, refer to Figure 2 This is a schematic diagram of a second audio desensitization process provided by an embodiment of the present invention. The encryption process for the second audio containing sensitive information in this embodiment of the present invention can be as follows:
[0084] Using the user's name "Zhang San" as the key, the incoming caller's mobile phone number "18900000000" is encrypted with AES to obtain the key string: U2FsdGVkX18EyiIm7sMcfqPEpnlVHpNWln8fHU+rKRE=;
[0085] The second audio file, audio_2.mp3, which contains sensitive information, is converted into a byte array. A byte array can be represented as:
[0086] audioBytes[]=[10101100][10111100][10101010][00010000][00100100][10001100][10101111][11010001];
[0087] After recognizing the above key string as a base-65 number, it is then converted into an 8-bit binary value, as shown below:
[0088] U=01011100 2=00000010F=01000101s=01100100d=01100101
[0089] G=01000110V=01010101k=01101010X=01011000 1=00000001
[0090] 8=00001000E=00100101y=01111001i=01101001I=01001001
[0091] m=01101101 7=00000111s=01100100M=01001101c=01100011
[0092] f=01100110q=01110001P=01010000E=00100101p=01110000
[0093] n=01101110l=01101100V=01010101H=01001000p=01110000
[0094] N=01001110W=01010111l=01101100n=01101110 8=00001000
[0095] f=01100110H=01001000U=01011100+=00111111r=01110010
[0096] K = 01001011 R = 01010010 E = 00100101 == 01000000 == 01000000 Combining the above binary values, the byte array audioBytes[] is processed through modulo, XOR, and reverse operations to obtain a new byte array audioBytes_n[], that is,
[0097] audioBytes[] = [10101100][10111100][10101010][00010000][00100100][10001100][10101111][11010001], which can be converted into a new byte array audioBytes_n[] = [00001111][01111101][11110111][00101110][10000010][01010011][01011111][11011101].
[0098] Reassemble audioBytes_n[] into a byte matrix. Using Math.sqrt(audioBytes_n[].length) = 2.828, round up to the nearest integer 3. This results in 3 groups of binary numbers in each row, forming the target byte matrix Byte[][]matrix as follows:
[0099] Byte[][]matrix=
[0100] {[00001111][01111101][11110111]},
[0101] {[00101110][10000010][01010011]},
[0102] {[01011111][11011101]}
[0103] The target byte matrix Byte[][] matrix is rotated clockwise around its boundaries as follows:
[0104]
[0105]
[0106]
[0107]
[0108] After the above rotations, a new byte matrix Byte[][]matrix_n can be obtained:
[0109] Byte[][]matrix_n=
[0110] {[00000111][10111110][11111011]},
[0111] {[00101110][10000010][01010011]},
[0112] {[10111111][10111011]};
[0113] The new byte matrix Byte[][]matrix_n is converted back into an array, becoming the target byte array audioBytes_n2[] = [00000111][10111110][11111011][00101110][10000010][01010011][10111111][10111011]. Finally, the target byte array audioBytes_n2[] is converted back into an audio stream, which is used as the encrypted second audio file and overwrites the unencrypted second audio file audio_2.mp3. The description information corresponding to the encrypted second audio file, audio_2 = [8000, 2000] (starting from the 8th second and lasting for 2 seconds), is recorded in the description file.
[0114] In this embodiment of the invention, the use of a byte shifting algorithm to encrypt the second audio file solves the security problem of sensitive information in the audio file, allowing the sensitive information in the audio file to be stored in an encrypted manner without occupying additional storage space, thus achieving the effects of improved security and saving storage space.
[0115] In one embodiment of the present invention, after step 107, saving the text information, the first audio, the encrypted second audio, and the description file to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file, the method further includes:
[0116] In response to a query request from a user with access rights for the audio file, the system obtains the text information corresponding to the audio file, the first audio file, the encrypted second audio file, and the description file, and determines the user's role and permissions.
[0117] The artificially synthesized audio is obtained by converting the text information.
[0118] Based on the role permissions and the description file, the artificially synthesized audio, the first audio, and the encrypted second audio are combined into the audio file.
[0119] In one embodiment of the present invention, the step of combining the artificially synthesized audio, the first audio, and the encrypted second audio into the audio file according to the role permissions and the description file includes:
[0120] When the role permission is not allowed to query sensitive information, the artificially synthesized audio, the first audio, and the decrypted second audio are combined into the audio file according to the description file;
[0121] When the role permission is set to allow querying sensitive information, the encrypted second audio is decrypted to obtain the decrypted second audio, and the artificially synthesized audio, the first audio, and the decrypted second audio are combined into the audio file according to the description file.
[0122] In this embodiment of the invention, when a user with access rights (such as a user on a conversation or an operator's staff member) needs to trace an audio file, the audio file can be reconstructed based on the text information, the first audio, the encrypted second audio, and the description file, thereby achieving the purpose of tracing the audio file.
[0123] It should be noted that, to ensure user information security, the role permissions of the user querying the audio file are used to determine whether decryption of the encrypted second audio file is necessary. (Refer to...) Figure 3 This is a schematic diagram of a second audio restoration process provided by an embodiment of the present invention. It involves obtaining persistent files (i.e., text information, a description file, a first audio file, and a decrypted second audio file). When a user's role permission is set to "not allow querying sensitive information," the artificially synthesized audio, the first audio file, and the decrypted second audio file are combined into an audio file according to the description file. Therefore, the user with the permission to play this audio file will not hear the user's sensitive information. When a user's role permission is set to "allow querying sensitive information," the encrypted second audio file is decrypted to obtain the decrypted second audio file. The artificially synthesized audio, the first audio file, and the decrypted second audio file are then combined into an audio file according to the description file. Therefore, the user with the permission to play this audio file can hear the user's sensitive information. This embodiment of the present invention hierarchically classifies the role permissions of authorized users to query audio files, which can protect user privacy, enhance system security and manageability, and increase system credibility by ensuring that authorized users can only access the audio files they are authorized to access.
[0124] For example, suppose on August 21, 2023 at 10:00 AM, user Zhang San has a dispute with the operator regarding an order. The content of the dispute needs to be corroborated by the call record from January 1, 2020 at 3:00 PM. To this end, Wang Wu, a user with high privileges (who can access the original audio file, i.e., the second audio in the audio file is decrypted), asks Li Si, a user with low privileges (who can only access the un-de-identified audio file, i.e., the second audio in the audio file is still encrypted), to retrieve the communication record at that time.
[0125] User Li Si found the description file 20200101150000_18900000000_Zhang San.txt, along with text information and the split audio file 20200101150000_18900000000_Zhang San.tar.gz, and attempted to listen to it. The system retrieved the description information from 20200101150000_18900000000_Zhang San.txt and extracted the description information from 20200101150000_18900000000_Zhang San. The contents of the .tar.gz file are combined, and a synthesized audio segment is generated using TTS technology based on the text information TTS_audio_1.txt. This is artificially synthesized audio. Through role-based access control, the first audio file, audio_1.mp3, and the encrypted second audio file, audio_2.mp3, are obtained. Audio_2.mp3 contains sensitive information; therefore, the system will not restore it if the user's access level is too low. The system then uses the text information TTS_audio_1 = [0, 3000] to obtain the audio file. =1 = [3001, 4998], audio_2 = [8000, 2000]. These three descriptive information segments are combined to obtain a complete audio file. At this point, the audio file is an anonymized version, and Li Si cannot obtain the key information. To obtain the key information, the high-privilege user Wang Wu personally retrieves the communication records from that time. Through role permission judgment, the first audio file, audio_1.mp3, and the encrypted second audio file, audio_2.mp3, are obtained. Audio_2.mp3 contains sensitive information, and due to the high role permission, the system will automatically restore the original audio file. The encrypted second audio file, audio_2.mp3, is restored using a reverse byte shift algorithm to obtain the original second audio file, audio_2.mp3. TTS_audio_1.txt is reconverted into a composite audio file using TTS and combined with the first audio file, audio_1.mp3, and the second audio file, audio_2.mp3, to obtain the original audio file. Finally, an agreement can be reached with user Zhang San through the original audio file, resolving the dispute.
[0126] In this embodiment of the invention, a strategy of setting different role permissions for users with different permissions is adopted, so that users with different permissions can obtain different audio versions of audio files under different role permissions, thus meeting the security and business requirements of audio files.
[0127] In one embodiment of the present invention, the decryption of the encrypted second audio includes:
[0128] Obtain the user information of the user corresponding to the audio file, and generate a key string using the user information; wherein, the user information includes at least the user's name and the user's incoming mobile phone number;
[0129] The key string is identified as a value in a specified base and then converted into a binary value with a specified number of bits; wherein, the specified base includes at least base 65 and the specified number of bits includes at least eight bits;
[0130] Convert the encrypted second audio into a target byte array;
[0131] The target byte array is rotated to obtain the target byte matrix; wherein the rotation includes at least a counterclockwise rotation of the boundary.
[0132] The target byte matrix is processed using the binary values to obtain the byte array;
[0133] The byte array is converted into the decrypted second audio.
[0134] For example, the specific decryption process for the encrypted second audio file, audio_2.mp3, is as follows:
[0135] The anonymized second audio file, audio_2.mp3, is then parsed back into a binary target byte array, audioBytes_n2[], via the audio stream:
[0136] audioBytes_n2[]=[00000111][10111110][11111011][00101110][10000010][01010011][10111111][10111011];
[0137] Convert the target byte array audioBytes_n2[] into a target byte matrix, and reverse it by rotating the boundary counterclockwise to obtain the target byte matrix audioBytes_n[] = [00001111][01111101][11110111][00101110][10000010][01010011][01011111][11011101];
[0138] Using the same method as encrypting the second audio, the key (key string) is obtained again based on the user information of the user in the conversation. Then, the key string is processed by modulo, reversal, XOR and other data processing methods to restore the target byte matrix audioBytes_n[], and the byte array audioBytes[] is obtained again.
[0139] audioBytes[]=[10101100][10111100][10101010][00010000][00100100][10001100][10101111][11010001];
[0140] Finally, convert the byte array audioBytes[] back into an audio stream to obtain the decrypted second audio, which is the un-de-sensitized second audio.
[0141] In this embodiment of the invention, the method of reverse-engineering the boundary displacement of the array matrix is used, which solves the problem of restoring the obfuscated byte array of the second audio, so that the invalidated audio file can be played normally, and has the effect of restoring the desensitized data.
[0142] In summary, the embodiments of the present invention have at least the following advantages: 1) By using a byte shift algorithm to encrypt audio containing sensitive information in audio files, audio file desensitization can be achieved without additionally occupying system storage resources, and the original audio file can be restored by reverse shifting when needed. 2) When desensitizing audio files, an AES key string is generated using user information, eliminating the need to store the key string separately. By using number base conversion and effectively combining the audio stream array with the key string, the security of audio files is improved. 3) Utilizing the special characteristics of artificially synthesized audio, the artificially synthesized audio in the audio file is converted into text information for storage, which can significantly save server disk space. 4) The audio file is stored in segments, reducing the risk of audio file leakage.
[0143] To enable those skilled in the art to better understand the embodiments of the present invention, a complete example is used for illustration below. Specifically, refer to... Figure 4 This is a schematic diagram of the main process of audio file processing provided by an embodiment of the present invention. The steps of audio file desensitization and audio file restoration in this embodiment of the present invention may specifically include:
[0144] Step 1: The TTS synthesized audio (artificially synthesized audio / TTS content) synthesized by the operator's intelligent outbound calling system during outbound calls is combined into complete voice interaction content (audio file). At the same time, a description file is generated, which describes all relevant information of the audio file. By default, the description file includes relevant timing information of TTS (description information). The audio file contains complete interactive content that conforms to the timing.
[0145] Step 2: Split the audio file and convert the TTS synthesized audio in the audio file into text (i.e., text information), which greatly reduces the storage space occupied by the audio file. Record relevant descriptive information in the description file of the audio file, such as the time period of the TTS synthesized audio (start time and duration).
[0146] Step 3: Parse the non-TTS content (pure human-interactive audio) using ASR (Automatic Speech Recognition) to convert it into text dialogue information, and determine whether the dialogue text information contains sensitive information, such as phone numbers, ID card numbers, etc.
[0147] Step 4: If there is no sensitive information, record the relevant timing information (description information) in the description file and skip the desensitization process.
[0148] Step 5: If sensitive information is contained, the sensitive information is split to desensitize the part of the audio file containing sensitive information, and the relevant descriptive information is recorded in the description file;
[0149] Step 6: Desensitize the sensitive information portion (audio file containing sensitive information) using a sensitive information desensitization method (e.g., using a byte shifting algorithm). Specifically, by parsing the non-TTS content using ASR, the portion containing sensitive information to be desensitized is obtained and converted into a byte array. The caller's mobile phone number is used as the character for AES encryption to obtain ciphertext (key string). The ciphertext is then converted into a base-65 number and encoded to obtain a new byte array (binary value) of 8 bits per group. The aforementioned binary value is used to perform modulo and XOR operations on the byte array (byte array) containing sensitive information to complete the basic encryption process and obtain a new byte array. Then, the new byte array is reassembled into a byte matrix (byte matrix). The byte matrix is shifted, and the boundary binary numbers are rotated clockwise by 1 bit to obtain a new byte array. Finally, the new byte array is converted back to the target byte array, and the target byte array is converted into an audio stream and merged with the original audio file to form the desensitized portion containing sensitive information.
[0150] Step 7: After the sensitive information is desensitized, a desensitized audio file will be generated. Then, the desensitized audio file will be recorded in the description file. At this point, the description file should contain the text information corresponding to the TTS content, the undesensitized audio file or the desensitized audio file, and relevant data describing these files, such as their timing, binding relationship, desensitization time period, user information, etc.
[0151] Step 8: Compress the text information corresponding to the TTS content, the un-anonymized audio file, or the anonymized audio file, and mark them according to the description file;
[0152] Step 9: Persistently store the description file and related data as persistent files, with the description file serving as the index address;
[0153] Step 10: Locate the relevant persistent files using the description file;
[0154] Step 11: If the description file contains an anonymized audio file, it means that the audio file contains anonymized information;
[0155] Step 12: If the description file contains text information converted from TTS content, it indicates that the persistent file involves TTS speech synthesis.
[0156] Step 13: Without involving TTS and desensitization, the persistent file contains only a compressed audio file, which can be directly decompressed and extracted;
[0157] Step 14: Without desensitization, the persistent file contains an undesensitized audio file and a text message converted from TTS content. At this point, the text message can be converted back into a TTS synthesized audio message using TTS technology. Then, the audio file and the TTS synthesized audio message regenerated from the text message can be combined according to the description information in the description file to obtain the original audio file and restore the real communication process.
[0158] Step 15: When both TTS and desensitization are involved, the persistent file contains the text information after TTS content conversion, corresponding to the desensitized audio file. When processing the desensitized information, the role permissions of the user who extracts the file determine whether it is necessary to restore the desensitized audio file. If not, the audio file is obtained directly by merging through the description file. If so, the key string composed of user information in the description file is used for processing, the desensitized audio file is restored, and then merged to obtain the audio file.
[0159] Step 16, Anti-desensitization method: The anti-desensitization method refers to obtaining user information and the address of the file to be desensitized from the description file, obtaining the key string through the user information, reading the desensitized audio file as a byte array, converting it into a byte matrix, and then shifting the boundary of the byte matrix counterclockwise by 1 bit; using the same key string, the anti-desensitization operation is performed to restore the desensitized audio file; through the above operations, the desensitized audio file is restored, and the restored audio file can play the original sound normally;
[0160] Step 17: At this point, users with the necessary permissions can obtain different audio files based on the specific details of the audio file and their own role-based access restrictions.
[0161] Through the above 17 steps, it can be seen that in step 1, converting the TTS synthesized audio back into text information for storage can significantly save storage space, as text storage occupies much less storage space than audio. In step 6, using a byte shift algorithm to desensitize sensitive information in the audio file makes the audio file containing sensitive information more secure, and this method hardly increases storage space. In step 15, using a key string formed from user information to restore the audio file according to role permissions ensures that only users with certain permissions can obtain the original audio file before desensitization. This not only meets the availability requirements of the business but also ensures data security. In summary, the embodiments of the present invention solve the problems of data security, business availability, and cost control simultaneously by splitting, texturing, and reversibly desensitizing audio files.
[0162] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0163] Based on the above embodiments, this embodiment also provides an audio file processing device, which is applied in electronic devices such as terminal devices and servers.
[0164] Reference Figure 5 The diagram illustrates a structural block diagram of an embodiment of an audio file processing device according to the present invention, which may specifically include the following modules:
[0165] The acquisition module 501 is used to acquire an audio file generated during human-computer interaction and a description file corresponding to the audio file; the description file includes at least descriptive information of artificially synthesized audio.
[0166] The first splitting module 502 is used to split the audio file into artificially synthesized audio and non-artificially synthesized audio according to the description information;
[0167] Conversion module 503 is used to convert the artificially synthesized audio into text information;
[0168] The second splitting module 504 is used to split the non-synthesized audio into a first audio and a second audio if sensitive information is detected in the non-synthesized audio; the second audio is the audio containing the sensitive information.
[0169] Recording module 505 is used to record description information corresponding to the first audio and the second audio respectively in the description file;
[0170] Encryption module 506 is used to encrypt the second audio to obtain the encrypted second audio;
[0171] The storage module 507 is used to store the text information, the first audio, the encrypted second audio, and the description file, so as to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file.
[0172] In one embodiment of the present invention, the description information includes at least the start time and duration of the corresponding audio.
[0173] In one embodiment of the present invention, the apparatus further includes: a text parsing module, used for:
[0174] The non-synthesized audio is converted into dialogue text information, and it is then identified whether the dialogue text information contains sensitive information.
[0175] In one embodiment of the present invention, the encryption module 506 is specifically used for:
[0176] Obtain the user information of the user corresponding to the audio file, and generate a key string using the user information; wherein, the user information includes at least the user's name and the user's incoming mobile phone number;
[0177] Convert the second audio file into a byte array;
[0178] The key string is identified as a value in a specified base and then converted into a binary value with a specified number of bits; wherein, the specified base includes at least base 65 and the specified number of bits includes at least eight bits;
[0179] A new byte array is obtained by performing data processing operations on the byte array using the binary values; wherein the data processing operations include at least modulo, XOR, and reversal;
[0180] Reassemble the new byte array into a byte matrix;
[0181] The square root of the byte matrix is then rounded up to obtain an integer value.
[0182] Adjust the byte matrix to a target byte matrix where the number of groups in each row is the integer value;
[0183] The target byte matrix is rotated to obtain a new byte matrix; wherein the rotation includes at least a clockwise rotation of the boundary.
[0184] Transform the new byte matrix into a target byte array;
[0185] The target byte array is then encrypted to obtain the second audio.
[0186] In one embodiment of the present invention, the device further includes: a decryption module, used for:
[0187] In response to a query request from a user with access rights for the audio file, the system obtains the text information corresponding to the audio file, the first audio file, the encrypted second audio file, and the description file, and determines the user's role and permissions.
[0188] The artificially synthesized audio is obtained by converting the text information.
[0189] Based on the role permissions and the description file, the artificially synthesized audio, the first audio, and the encrypted second audio are combined into the audio file.
[0190] In one embodiment of the present invention, the step of combining the artificially synthesized audio, the first audio, and the encrypted second audio into the audio file according to the role permissions and the description file includes:
[0191] When the role permission is not allowed to query sensitive information, the artificially synthesized audio, the first audio, and the decrypted second audio are combined into the audio file according to the description file;
[0192] When the role permission is set to allow querying sensitive information, the encrypted second audio is decrypted to obtain the decrypted second audio, and the artificially synthesized audio, the first audio, and the decrypted second audio are combined into the audio file according to the description file.
[0193] In one embodiment of the present invention, the decryption module is specifically used for:
[0194] Obtain the user information of the user corresponding to the audio file, and generate a key string using the user information; wherein, the user information includes at least the user's name and the user's incoming mobile phone number;
[0195] The key string is identified as a value in a specified base and then converted into a binary value with a specified number of bits; wherein, the specified base includes at least base 65 and the specified number of bits includes at least eight bits;
[0196] Convert the encrypted second audio into a target byte array;
[0197] The target byte array is rotated to obtain the target byte matrix; wherein the rotation includes at least a counterclockwise rotation of the boundary.
[0198] The target byte matrix is processed using the binary values to obtain the byte array;
[0199] The byte array is converted into the decrypted second audio.
[0200] In this embodiment of the invention, audio files and corresponding description files generated during human-computer interaction are obtained. The description file includes at least descriptive information about artificially synthesized audio. Based on the description information, the audio file is split into artificially synthesized audio and non-artificially synthesized audio. The artificially synthesized audio is converted into text information. If sensitive information is detected in the non-artificially synthesized audio, it is split into a first audio and a second audio containing the sensitive information. The description information corresponding to the first and second audios is recorded in the description file. The second audio is encrypted to obtain an encrypted second audio. Then, the text information, the first audio, the encrypted second audio, and the description file are saved. Thus, the audio file can be traced later based on the text information, the first audio, the encrypted second audio, and the description file. This embodiment of the invention uses text information to record artificially synthesized audio in audio files generated during human-computer interaction, thereby reducing the storage space required for long-term storage of audio files and lowering the operator's costs. It should also be noted that since artificially synthesized audio is generated through artificial intelligence technology, the accuracy of recognizing artificially synthesized audio as text information and then restoring it back to artificially synthesized audio from the text information can reach 100%, thus ensuring the accuracy of audio file restoration.
[0201] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0202] This invention also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute instructions for the method steps in this invention.
[0203] This invention provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this invention, the electronic device includes various types of devices such as terminal devices and servers (clusters).
[0204] The embodiments of this disclosure can be implemented as an apparatus configured as desired using any suitable hardware, firmware, software, or any combination thereof, including electronic devices such as terminal devices, servers (clusters), etc. Figure 6 An exemplary apparatus 600 is schematically shown that can be used to implement the various embodiments described in this invention.
[0205] In one embodiment, Figure 6 An exemplary device 600 is shown, which includes one or more processors 602, a control module (chipset) 604 coupled to at least one of the processors 602, a memory 606 coupled to the control module 604, a non-volatile memory (NVM) / storage device 608 coupled to the control module 604, one or more input / output devices 610 coupled to the control module 604, and a network interface 612 coupled to the control module 604.
[0206] Processor 602 may include one or more single-core or multi-core processors, and processor 602 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 600 can serve as a terminal device, server (cluster), or other device as described in the embodiments of the present invention.
[0207] In some embodiments, the apparatus 600 may include one or more computer-readable media (e.g., memory 606 or NVM / storage device 608) having instructions 614 and one or more processors 602 that are combined with the one or more computer-readable media and configured to execute the instructions 614 to implement the module and thus perform the actions described in this disclosure.
[0208] In one embodiment, the control module 604 may include any suitable interface controller to provide any suitable interface to at least one of the processors 602 and / or any suitable device or component communicating with the control module 604.
[0209] The control module 604 may include a memory controller module to provide an interface to the memory 606. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0210] Memory 606 may be used, for example, to load and store data and / or instructions 614 for device 600. In one embodiment, memory 606 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 606 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0211] In one embodiment, the control module 604 may include one or more input / output controllers to provide an interface to the NVM / storage device 608 and (one or more) input / output devices 610.
[0212] For example, NVM / storage device 608 may be used to store data and / or instructions 614. NVM / storage device 608 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).
[0213] NVM / storage device 608 may include storage resources that are physically part of a device on which device 600 is mounted, or that are accessible to the device but do not necessarily have to be part of the device. For example, NVM / storage device 608 may be accessed via a network through one or more input / output devices 610.
[0214] One or more input / output devices 610 may provide an interface for device 600 to communicate with any other suitable device. Input / output devices 610 may include communication components, audio components, sensor components, etc. A network interface 612 may provide an interface for device 600 to communicate via one or more networks. Device 600 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.
[0215] In one embodiment, at least one of the processors 602 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 604. In one embodiment, at least one of the processors 602 may be logically packaged with one or more controllers of the control module 604 to form a system-in-package (SiP). In one embodiment, at least one of the processors 602 may be integrated with the logic of one or more controllers of the control module 604 on the same die. In one embodiment, at least one of the processors 602 may be integrated with the logic of one or more controllers of the control module 604 on the same die to form a system-on-a-chip (SoC).
[0216] In various embodiments, device 600 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop, handheld computing device, tablet, netbook, etc.). In various embodiments, device 600 may have more or fewer components and / or different architectures. For example, in some embodiments, device 600 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0217] The detection device may use a main control chip as a processor or control module, and sensor data, position information, etc. may be stored in a memory or NVM / storage device. The sensor group may be used as an input / output device, and the communication interface may include a network interface.
[0218] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0219] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable audio file processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable audio file processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0220] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable audio file processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0221] These computer program instructions can also be loaded onto a computer or other programmable audio file processing terminal device, causing a series of operational steps to be performed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0222] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0223] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0224] The above provides a detailed description of an audio file processing method and apparatus, an electronic device, and a storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An audio file processing method, characterized in that, The method includes: Obtain the audio file generated during human-computer interaction and the corresponding description file of the audio file; the description file includes at least the description information of the artificially synthesized audio. Based on the description information, the audio file is split into artificially synthesized audio and non-artificially synthesized audio; Convert the artificially synthesized audio into text information; If sensitive information is detected in the non-synthesized audio, the non-synthesized audio is split into a first audio and a second audio; the second audio is the audio containing the sensitive information. The description file records the description information corresponding to the first audio and the second audio respectively; The second audio is encrypted to obtain the encrypted second audio. The text information, the first audio, the encrypted second audio, and the description file are saved to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file.
2. The method according to claim 1, characterized in that, The description information includes at least the start time and duration of the corresponding audio.
3. The method according to claim 1, characterized in that, Before splitting the non-synthetic audio into a first audio and a second audio if sensitive information is detected in the non-synthetic audio, the method further includes: The non-synthesized audio is converted into dialogue text information, and it is then identified whether the dialogue text information contains sensitive information.
4. The method according to claim 1, characterized in that, The process of encrypting the second audio to obtain the encrypted second audio includes: Obtain the user information of the user corresponding to the audio file, and generate a key string using the user information; wherein, the user information includes at least the user's name and the user's incoming mobile phone number; Convert the second audio file into a byte array; The key string is identified as a value in a specified base and then converted into a binary value with a specified number of bits; wherein, the specified base includes at least base 65 and the specified number of bits includes at least eight bits; A new byte array is obtained by performing data processing operations on the byte array using the binary values; wherein the data processing operations include at least modulo, XOR, and reversal; Reassemble the new byte array into a byte matrix; The square root of the byte matrix is then rounded up to obtain an integer value. Adjust the byte matrix to a target byte matrix where the number of groups in each row is the integer value; The target byte matrix is rotated to obtain a new byte matrix; wherein the rotation includes at least a clockwise rotation of the boundary. Transform the new byte matrix into a target byte array; The target byte array is converted to obtain the encrypted second audio.
5. The method according to claim 1, characterized in that, After saving the text information, the first audio, the encrypted second audio, and the description file to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file, the method further includes: In response to a query request from a user with access rights for the audio file, the system obtains the text information corresponding to the audio file, the first audio file, the encrypted second audio file, and the description file, and determines the user's role and permissions. The artificially synthesized audio is obtained by converting the text information. Based on the role permissions and the description file, the artificially synthesized audio, the first audio, and the encrypted second audio are combined into the audio file.
6. The method according to claim 5, characterized in that, The step of combining the artificially synthesized audio, the first audio, and the encrypted second audio into the audio file according to the role permissions and the description file includes: When the role permission is not allowed to query sensitive information, the artificially synthesized audio, the first audio, and the decrypted second audio are combined into the audio file according to the description file; When the role permission is set to allow querying sensitive information, the encrypted second audio is decrypted to obtain the decrypted second audio, and the artificially synthesized audio, the first audio, and the decrypted second audio are combined into the audio file according to the description file.
7. The method according to claim 6, characterized in that, The process of decrypting the encrypted second audio includes: Obtain the user information of the user corresponding to the audio file, and generate a key string using the user information; wherein, the user information includes at least the user's name and the user's incoming mobile phone number; The key string is identified as a value in a specified base and then converted into a binary value with a specified number of bits; wherein, the specified base includes at least base 65 and the specified number of bits includes at least eight bits; Convert the encrypted second audio into a target byte array; The target byte array is rotated to obtain the target byte matrix; wherein the rotation includes at least a counterclockwise rotation of the boundary. The target byte matrix is processed using the binary values to obtain the byte array; The byte array is converted into the decrypted second audio.
8. An audio file processing device, characterized in that, The device includes: The acquisition module is used to acquire audio files generated during human-computer interaction and description files corresponding to the audio files; the description files include at least descriptive information of artificially synthesized audio. The first splitting module is used to split the audio file into artificially synthesized audio and non-artificially synthesized audio according to the description information; A conversion module is used to convert the artificially synthesized audio into text information; The second splitting module is used to split the non-synthesized audio into a first audio and a second audio if sensitive information is detected in the non-synthesized audio; the second audio is the audio containing the sensitive information. The recording module is used to record the description information corresponding to the first audio and the second audio respectively in the description file; An encryption module is used to encrypt the second audio to obtain the encrypted second audio. A storage module is used to store the text information, the first audio, the encrypted second audio, and the description file, so as to trace the audio file based on the text information, the first audio, the encrypted second audio, and the description file.
9. An electronic device, characterized in that, include: processor; and A memory having executable code stored thereon, which, when executed, causes the processor to perform the audio file processing method as described in any one of claims 1-7.
10. One or more machine-readable media having executable code stored thereon, which, when executed, causes a processor to perform the audio file processing method as claimed in any one of claims 1-7.
Citation Information
Patent Citations
Audio file processing method and equipment
CN104517068A
Machine learning for microphone style transfer
CN116472579A