Audio data processing methods, devices, and servers

CN116665668BActive Publication Date: 2026-05-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2023-05-30
Publication Date
2026-05-26

Smart Images

  • Figure CN116665668B_ABST
    Figure CN116665668B_ABST
Patent Text Reader

Abstract

This specification provides a method, apparatus, and server for processing audio data, applicable to the financial field. Based on this method, after receiving target audio data and obtaining the corresponding first data group through speech recognition, the first data group is first split according to a preset splitting rule to obtain multiple character matrices. Then, based on the pinyin data of the characters, the characters in the character matrices are mapped to corresponding letter character combinations to obtain the corresponding first letter character matrix. According to a preset exchange rule, letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined. The initial consonants and / or final vowels of the letter character combinations that satisfy the exchange conditions are then exchanged to obtain the corresponding second letter character matrix. Based on the second letter character matrix, the corresponding second data group is obtained. This method can efficiently hide relevant information carried in the audio data, protecting the data security of that information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual pertains to the field of Internet technology, and in particular relates to methods, devices, and servers for processing audio data. Background Technology

[0002] When conducting financial transactions, users often need to record and upload their audio data. This audio data often contains personal information about the user.

[0003] Based on existing methods, the user-related personal information carried in the audio data is easily leaked when using and processing it. Furthermore, audio data is more complex and cumbersome to process than other forms of data, and the amount of computer data processing required by existing methods is relatively large, resulting in low overall processing efficiency.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This specification provides a method, apparatus, and server for processing audio data, which can efficiently hide relevant information carried in audio data, better protect the data security of relevant information carried in audio data, and prevent the relevant information carried in audio data from being leaked.

[0006] This manual provides a method for processing audio data, including:

[0007] Receive target audio data;

[0008] Speech recognition is performed on the target audio data to obtain the corresponding first data group; wherein, the first data group includes multiple text characters arranged in order;

[0009] According to the preset splitting rules, the first data group is split to obtain multiple text character matrices; wherein, the text character matrices use text characters as matrix elements;

[0010] Based on the pinyin data of the text characters, the text characters in the text character matrix are mapped to the corresponding letter character combinations to obtain the corresponding first letter character matrix; wherein, the first letter character matrix uses letter character combinations as matrix elements;

[0011] According to the preset exchange rules, the letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined; and the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions are exchanged to obtain the corresponding second letter character matrix;

[0012] Based on the second letter character matrix, the corresponding second data group is obtained.

[0013] In one embodiment, the target audio data includes audio data involving sensitive information provided by the user when conducting a transaction.

[0014] In one embodiment, after receiving the target audio data, the method further includes:

[0015] The target audio data is preprocessed; wherein the preprocessing includes at least one of the following: background noise filtering, dialect speech correction, and business scenario matching.

[0016] In one embodiment, when the preprocessing includes business scenario matching, speech recognition is performed on the target audio data to obtain a corresponding first data set, including:

[0017] Based on the matching results of the business scenario, the matching preset speech recognition model is determined from multiple preset speech recognition models as the target speech recognition model;

[0018] The target audio data is processed using a target speech recognition model to obtain the corresponding target speech recognition result;

[0019] Based on the target speech recognition results, text characters involving sensitive information are extracted and combined to obtain the first data group.

[0020] In one embodiment, after obtaining the corresponding second data set, the method further includes:

[0021] The second data group is sent to the target database for storage; and the target audio data and the first data group are deleted; wherein the target database is a blockchain-based database.

[0022] In one embodiment, the first data group is split according to a preset splitting rule to obtain multiple character matrices, including:

[0023] Based on the preset splitting rules and the number of characters in the text characters in the first data group, multiple blank symmetric matrices are created; the multiple blank symmetric matrices are arranged in order, and the number of rows of the symmetric matrix that is sorted first is greater than or equal to the number of rows of the symmetric matrix that is sorted later.

[0024] Based on the sorting information of the text characters in the first data group, the text characters in the first data group are assigned to the corresponding blank symmetric matrices, resulting in multiple sub-data groups corresponding to multiple blank symmetric matrices respectively.

[0025] Based on the sorting information of the text characters in the first data group, the blank matrix elements in the corresponding blank symmetric matrix are replaced with the text characters in the sub-data group to obtain multiple text character matrices.

[0026] In one embodiment, according to a preset swapping rule, the letter character combinations in the first letter character matrix that satisfy the swapping conditions are determined, including:

[0027] The matrix coordinates of the letter character combination are determined based on the row and column number of the letter character combination in the first letter character matrix; wherein, the matrix coordinates include row coordinates and column coordinates;

[0028] According to the preset exchange rules, the first matrix coordinates and the second matrix coordinates that satisfy the preset data relationship are determined by retrieving the matrix coordinates of letter character combinations in the same first letter character matrix;

[0029] The combination of first-letter characters indicated by the first matrix coordinate and the combination of second-letter characters indicated by the second matrix coordinate in the same first-letter character matrix are determined as letter character combinations that satisfy the exchange condition.

[0030] In one embodiment, the preset data relationship includes: the row coordinate of the first matrix coordinate is equal to the column coordinate of the second matrix coordinate, and the column coordinate of the first matrix coordinate is equal to the row coordinate of the second matrix coordinate; and / or, the row coordinate and column coordinate of the first matrix coordinate are equal, the row coordinate and column coordinate of the second matrix coordinate are equal, and the sum of the row coordinate of the first matrix coordinate and the row coordinate of the second matrix coordinate is equal to the row number of the first letter character matrix.

[0031] In one embodiment, the process of swapping initial consonants and / or final vowels in letter character combinations that satisfy the swapping conditions includes:

[0032] Detect whether the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination satisfy a pinyin combination relationship; Detect whether the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination satisfy a pinyin combination relationship;

[0033] If the initial consonant in the first letter character group and the final vowel in the second letter character group satisfy a phonetic combination relationship, and the final vowel in the first letter character group and the initial consonant in the second letter character group satisfy a phonetic combination relationship, then swap the initial consonant in the first letter character group with the initial consonant in the second letter character group.

[0034] In one embodiment, after detecting whether the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination satisfy a pinyin combination relationship; and after detecting whether the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination satisfy a pinyin combination relationship, the method further includes:

[0035] If it is determined that the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination do not satisfy the pinyin combination relationship, and / or, the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination do not satisfy the pinyin combination relationship, the first letter character combination and the second letter character combination shall not be swapped.

[0036] In one embodiment, the method further includes:

[0037] Detect whether there is a consonant character in the first letter character combination and the second letter character combination;

[0038] If it is determined that there is no initial consonant character in the first letter character combination and / or the second letter character combination, the first letter character combination and the second letter character combination shall not be swapped.

[0039] This specification also provides an audio data processing apparatus, including:

[0040] The receiving module is used to receive target audio data;

[0041] The speech recognition module is used to perform speech recognition on the target audio data to obtain the corresponding first data group; wherein, the first data group includes multiple text characters arranged in order;

[0042] The splitting module is used to split the first data group according to a preset splitting rule to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements.

[0043] The mapping module is used to acquire and map the characters in the character matrix to corresponding letter character combinations based on the pinyin data of the characters, thereby obtaining the corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements.

[0044] The exchange module is used to determine the letter character combinations in the first letter character matrix that satisfy the exchange conditions according to the preset exchange rules; and to exchange the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain the corresponding second letter character matrix.

[0045] The processing module is used to obtain the corresponding second data group based on the second letter character matrix.

[0046] This specification also provides a server, including a processor and a memory for storing processor-executable instructions, wherein the processor, when executing the instructions, implements the relevant steps of the method for processing the audio data.

[0047] This specification also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, perform the following steps: receiving target audio data; performing speech recognition on the target audio data to obtain a corresponding first data group; wherein the first data group includes multiple sequentially arranged text characters; splitting the first data group according to a preset splitting rule to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements; acquiring and mapping the text characters in the text character matrices to corresponding letter character combinations based on the pinyin data of the text characters, to obtain a corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements; determining the letter character combinations in the first letter character matrix that satisfy the exchange conditions according to a preset exchange rule; and exchanging the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain a corresponding second letter character matrix; and obtaining a corresponding second data group based on the second letter character matrix.

[0048] This specification also provides a computer program product comprising a computer program that, when executed by a processor, implements the relevant steps of the audio data processing method.

[0049] Based on the audio data processing method, apparatus, and server provided in this specification, after receiving target audio data and obtaining the corresponding first data group through speech recognition, the first data group is first split according to a preset splitting rule to obtain multiple character matrices; then, based on the pinyin data of the characters, the characters in the character matrices are mapped to corresponding letter character combinations to obtain the corresponding first letter character matrix; according to a preset exchange rule, the letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined; and the initial consonants and / or final vowels of the letter character combinations that satisfy the exchange conditions are exchanged to obtain the corresponding second letter character matrix; based on the second letter character matrix, the corresponding second data group is obtained. This fully utilizes the spatial characteristics and computational advantages of the matrix structure, as well as the relevant characteristics of pinyin data, to efficiently hide the real data information carried in the audio data through relevant data processing, effectively preventing the leakage of data information carried in the audio data and better protecting the data security of the audio data. Attached Figure Description

[0050] To more clearly illustrate the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart illustrating an embodiment of an audio data processing method provided in this specification.

[0052] Figure 2 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0053] Figure 3 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0054] Figure 4 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0055] Figure 5 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0056] Figure 6 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0057] Figure 7 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0058] Figure 8 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0059] Figure 9 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0060] Figure 10 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0061] Figure 11 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0062] Figure 12 This is a schematic diagram of an embodiment of the audio data processing method provided in the embodiments of this specification, applied in a scenario example.

[0063] Figure 13 This is a schematic diagram of the structural composition of a server provided in one embodiment of this specification;

[0064] Figure 14 This is a schematic diagram of the structure of an audio data processing device provided in one embodiment of this specification. Detailed Implementation

[0065] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0066] It should be noted that all user-related information and data mentioned in this manual were obtained and used with the user's knowledge and consent; and the acquisition, storage, use, and processing of the aforementioned information and data comply with the relevant provisions of national laws and regulations.

[0067] See Figure 1 As shown in the embodiments of this specification, an audio data processing method is provided, wherein the method is specifically applied to the server side. In specific implementation, the method may include the following:

[0068] S101: Receive target audio data;

[0069] S102: Perform speech recognition on the target audio data to obtain the corresponding first data group; wherein, the first data group includes multiple text characters arranged in order;

[0070] S103: According to the preset splitting rules, the first data group is split to obtain multiple text character matrices; wherein, the text character matrices use text characters as matrix elements;

[0071] S104: Obtain and map the characters in the character matrix to corresponding letter character combinations based on the pinyin data of the characters, to obtain the corresponding first letter character matrix; wherein, the first letter character matrix uses letter character combinations as matrix elements;

[0072] S105: Based on the preset exchange rules, determine the letter character combinations in the first letter character matrix that satisfy the exchange conditions; and exchange the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain the corresponding second letter character matrix;

[0073] S106: Obtain the corresponding second data group based on the second letter character matrix.

[0074] Based on the above embodiments, after the server obtains the corresponding first data group by performing speech recognition on the target audio data, it can first obtain multiple first letter character matrices with letter character combinations as matrix elements according to preset splitting rules and related pinyin data; then, by exchanging the initial consonant characters and / or final vowel characters of letter character combinations that meet the exchange conditions in the first letter character matrices, the corresponding second data group is obtained. This can fully utilize the spatial characteristics and computational advantages of the matrix structure, as well as the relevant characteristics of pinyin data, to efficiently perform related data processing, reliably hide the real data information, effectively prevent the data information carried in the audio data from being leaked, and better protect the data security of the audio data.

[0075] In some embodiments, the target audio data includes audio data involving sensitive information provided by the user when conducting a transaction.

[0076] Specifically, the aforementioned sensitive information can be understood as information data related to the user's personal information, involving user privacy, and requiring protection. Examples include user information such as the user's name, identity information, and address.

[0077] Based on the above embodiments, the audio data processing method provided in this specification can be extended to transaction business scenarios to effectively protect the audio data provided by users when handling transaction business.

[0078] In some embodiments, the above-described audio data processing method can be specifically applied to the server side.

[0079] In transaction scenarios, when users want to conduct specific transactions, they often need to record relevant audio data using their own terminal devices according to relevant instructions, and then send the target audio data to the server for verification.

[0080] Among them, see Figure 2 As shown, the aforementioned server (e.g., a front-end server) may specifically include a back-end server applied to a financial trading platform (e.g., XX Bank, etc.) capable of data transmission, data processing, and other functions. Specifically, the server may be, for example, an electronic device with data processing, storage, and network interaction capabilities. Alternatively, the server may be a software program running on the electronic device, providing support for data processing, storage, and network interaction. In this embodiment, the number of servers is not specifically limited. The server may be a single server, several servers, or a server cluster formed by several servers.

[0081] The aforementioned terminal device may specifically include a front-end applied to the user side, capable of data collection, data transmission, and other functions. Specifically, the terminal device may be an electronic device such as a smartphone, desktop computer, tablet computer, or laptop computer. Alternatively, the terminal device may also be a software application that can run on the aforementioned electronic device, such as a client app of XX Bank running on a smartphone.

[0082] In practice, the terminal device can collect the user's target audio data and send the target audio data to the front-end server through a secure communication channel between the terminal device and the front-end server.

[0083] After receiving the target audio data, the front-end server first performs speech recognition on the target audio data to obtain a first data group containing multiple sequentially arranged text characters; wherein, the first data group contains sensitive information related to the user.

[0084] Next, the front-end server can process the first data group according to the preset processing rules to obtain a second data group that hides the real relevant information; then send the second data group to the verification server for further verification; at the same time, the second data group can replace the target audio data and be stored in the corresponding target database for evidence preservation, and the original target audio data and the first data group can be deleted.

[0085] When processing the first data group, the front-end server can first split the first data group according to a preset splitting rule to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements; then, based on the pinyin data of the text characters, the text characters in the text character matrices are mapped to corresponding letter character combinations to obtain the corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements; according to a preset exchange rule, the letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined; and the initial consonants and / or final vowels of the letter character combinations that satisfy the exchange conditions are exchanged to obtain the corresponding second letter character matrix; finally, the pinyin data indicated by each letter character combination in the second letter character matrix is ​​mapped sequentially to obtain the corresponding second data group. This effectively hides the real relevant information carried in the first data group and obtains the corresponding second data group.

[0086] Subsequently, the server can use the second data group, which hides the true relevant information, to replace the original target audio data and the first data group for related data interaction and data storage. This way, even if the second data group is leaked during data interaction or data storage, the true relevant information carried by the original target audio data cannot be obtained solely based on the second data group. This effectively prevents the leakage of the relevant information carried by the target audio data and better protects user information security.

[0087] Accordingly, after receiving the second data set, the verification server can perform restoration processing on the second data set in a trusted execution environment according to the preset processing rules agreed upon with the front-end server in advance, so as to obtain the first data set containing real relevant information; then, based on the first data set, the transaction business applied for by the user is verified to determine whether the user meets the requirements, and the final verification result is fed back to the front-end server.

[0088] Specifically, the aforementioned trusted execution environment can include a high-security area within the server (e.g., a security level that meets preset security requirements). More specifically, the trusted execution environment can be a hardware area separated from commonly used, relatively open environment areas (e.g., Rich Execution Environment, REE, etc.) by means of hardware configuration or other methods.

[0089] In this example scenario, the Trust Execution Environment (TEE) described above can run a complete operating system, which can be understood as the Secure World within the server. Unlike the Normal World (e.g., the REE within a server), the TEE typically has a relatively small memory footprint, perhaps only 100MB. Within the server, only a portion of the data with high security requirements is usually processed within the TEE; the majority of the data is processed in the Normal World, such as the REE. Of course, the TEE described above is merely illustrative. In practice, depending on the specific application scenario and server configuration, other high-security areas within the server can be selected to replace the TEE.

[0090] The front-end server receives the verification results. If the user meets the requirements, the server can proceed with the transaction. Conversely, if the user does not meet the requirements, the server generates a failure message indicating that the transaction cannot be completed and sends the message to the terminal device.

[0091] In some embodiments, after receiving the target audio data, the method may further include the following: preprocessing the target audio data; wherein the preprocessing may specifically include at least one of the following: background noise filtering, dialect speech correction, business scenario matching, etc.

[0092] In practice, the server can invoke a background noise reduction model to process the target audio data, automatically identifying and filtering out background noise. Specifically, this background noise reduction model can be a statistical algorithm model pre-established using a large amount of sample background noise audio data.

[0093] In practice, the server can call a preset dialect accent correction model to process the target audio data, automatically finding audio data segments containing dialect accents within the target audio data; and then making targeted adjustments to the frequency bands of the aforementioned audio data to obtain target audio data that conforms to the standard Mandarin. Specifically, the preset dialect accent correction model can be a neural network model pre-trained using a large amount of audio data containing dialect accents.

[0094] In practice, the server can also collect associated feature information of the target audio data currently being recorded by the user, such as the user's current location information, the interface information of the business interface currently displayed to the user by the terminal device when the user is recording the target audio data, and the data information of the user's interaction with the front-end server at the previous time point; at the same time, the server will also perform business keyword retrieval on the target audio data to extract business keywords related to the business; and then combine the above associated feature information and business keywords to perform business scenario matching in order to accurately determine the user's current business scenario.

[0095] Based on the above embodiments, the target audio data can be preprocessed in various ways in advance so that the subsequent speech recognition of the target audio data can be more accurate and reduce speech recognition errors.

[0096] In some embodiments, where the preprocessing includes business scenario matching, see [reference needed]. Figure 3 As shown, speech recognition is performed on the target audio data to obtain the corresponding first data set. In specific implementation, this may include the following:

[0097] S1: Based on the matching results of the business scenario, determine the matching preset speech recognition model from multiple preset speech recognition models as the target speech recognition model;

[0098] S2: Process the target audio data using the target speech recognition model to obtain the corresponding target speech recognition result;

[0099] S3: Based on the target speech recognition result, extract the text characters involving sensitive information and combine them to obtain the first data group.

[0100] Based on the above embodiments, the target speech recognition model can be selected and used to more accurately recognize the target audio data according to the matching results of the business scenario, thereby further reducing speech recognition errors in speech recognition.

[0101] In some embodiments, before implementation, multiple preset speech recognition models can be trained using audio sample data from different business scenarios; each preset speech recognition model corresponds to a business scenario.

[0102] In some specific implementations, the server can directly arrange and combine the text characters contained in the target speech recognition result in order as the first data group.

[0103] The server can also first retrieve text characters containing sensitive information from the target speech recognition results based on a preset key character table; then, it extracts only the retrieved text characters containing sensitive information and arranges them in order to obtain data containing only the more important sensitive information, which is used as the first data group. This can effectively reduce the amount of data in the first data group, reduce the amount of subsequent data processing, and improve the overall data processing efficiency.

[0104] Specifically, the first data group mentioned above may include multiple text characters arranged in sequence, such as "a text encryption and decryption algorithm".

[0105] In some embodiments, see Figure 4 As shown, the first data group is split according to the preset splitting rules to obtain multiple character matrices. In specific implementation, the following may be included:

[0106] S1: Based on the preset splitting rules and the number of characters in the text characters in the first data group, create multiple blank symmetric matrices; wherein, the multiple blank symmetric matrices are arranged in order, and the number of rows of the symmetric matrix that is sorted first is greater than or equal to the number of rows of the symmetric matrix that is sorted later.

[0107] S2: Based on the sorting information of the text characters in the first data group, assign the text characters in the first data group to the corresponding blank symmetric matrices to obtain multiple sub-data groups corresponding to multiple blank symmetric matrices respectively;

[0108] S3: Based on the sorting information of the text characters in the first data group, replace the blank matrix elements in the corresponding blank symmetric matrix with the text characters in the sub-data group to obtain multiple text character matrices.

[0109] Based on the above embodiments, the first data group can be split into multiple character matrices according to the preset splitting rules, so that the relevant data processing can be carried out more efficiently and securely by taking advantage of the characteristics and computational advantages of the matrix structure based on the character matrices.

[0110] It should be noted that converting the first data group into a matrix structure is due to two considerations: First, matrix structure data is more suitable for computer processing and has better computational advantages. Subsequent data processing in matrix form can effectively improve the overall data processing efficiency. Second, splitting the originally conventional and simple string structure of the first data group into multiple matrix structures can utilize the polygons of the matrix structure to increase the difficulty of cracking and improve data security.

[0111] In some embodiments, for example, one may refer to Figure 5 As shown, the aforementioned blank symmetric matrix can specifically refer to a matrix whose initial matrix elements are 0 (i.e., blank matrix elements) and whose number of rows and columns are the same (e.g., q rows and q columns).

[0112] Specifically, when creating multiple blank symmetric matrices based on the preset splitting rules and the number of characters in the first data group, the server can create the first blank symmetric matrix in the following way: determine the number of characters currently contained in the first data group as the first quantity (e.g., n), and determine the number of rows and columns of the first blank symmetric matrix (e.g., m) by taking the square root of the first quantity and rounding it down; and create a symmetric matrix with m rows and m columns, with the initial matrix elements being 0 as the first blank symmetric matrix.

[0113] After creating the first blank symmetric matrix, the initial detection and judgment can be performed as follows: Calculate the number of matrix elements contained in the first blank symmetric matrix (e.g., m). 2 ); the first difference between the first quantity in the first data set and the first number of matrix elements contained in the first blank matrix (e.g., nm). 2 The threshold is less than or equal to a preset first threshold (e.g., 4) and greater than a preset second threshold (e.g., 0).

[0114] Specifically, when the first difference is determined to be greater than the preset second threshold and less than or equal to the preset first threshold, it is sufficient to create a second blank symmetric matrix with 2 rows and 2 columns and 0 initial matrix elements to complete the creation of multiple blank symmetric matrices and end the creation of blank symmetric matrices.

[0115] When the first difference is determined to be less than or equal to the preset second threshold, the creation of the second blank symmetric matrix can be terminated without creating a second blank symmetric matrix.

[0116] When it is determined that the first difference is greater than a preset first threshold, the first difference (e.g., nm) can be adjusted. 2 Perform a square root and round down to determine the number of rows and columns (e.g., p) of the second blank symmetric matrix; then create a symmetric matrix with p rows and p columns, initially containing 0 elements, as the second blank symmetric matrix. Then perform a second detection and judgment.

[0117] Continue in the same manner until the creation of the blank symmetric matrix is ​​complete.

[0118] Specifically, for example, the first data set "Mr. Ming plans to take the train to a certain city with his college classmate Mr. Wang tomorrow" contains 22 characters. Based on the above method, refer to... Figure 6 As shown, we can first create a blank symmetric matrix with 4 rows and 4 columns. The first difference is 2² - 16 = 6. After the first check, this is greater than the preset first threshold. Then, based on the first difference of 6, we can create a second blank symmetric matrix with 2 rows and 2 columns. The second difference is 6 - 4 = 2. After the second check, this is greater than the preset second threshold and less than or equal to the preset first threshold. At this point, we only need to create a third blank symmetric matrix with 2 rows and 2 columns to complete the creation of blank symmetric matrices, ultimately obtaining 3 blank symmetric matrices.

[0119] In some embodiments, when implementing the text characters, a specified number of text characters that are ranked first in the first data group can be extracted according to the sorting information of the text characters in the first data group and assigned to the first blank symmetric matrix, the second blank symmetric matrix, and so on, until the last blank symmetric matrix, as multiple sub-data groups corresponding to each blank symmetric matrix, thus completing the allocation of text characters.

[0120] In some embodiments, when replacing the blank matrix elements in the corresponding blank symmetric matrix with text characters from the sub-data group, the text characters in each sub-data group can be used to replace the blank matrix elements in the corresponding blank matrix in a left-to-right, top-to-bottom order.

[0121] For the last blank matrix, if the number of text characters contained in the sub-data group is less than the total number of matrix elements in the blank matrix, the blank matrix elements at the corresponding positions can be replaced by the text characters in the sub-data group in order from left to right and from top to bottom. The remaining blank matrix elements that cannot be replaced by text characters are left in their original positions as placeholders.

[0122] Thus, one or more matrices of literal characters can be obtained. For the blank matrix elements used for placeholder in the matrix of literal characters, they can be recognized subsequently and no processing will be performed.

[0123] Specifically, for example, refer to Figure 7 As shown, there is only one matrix of literal characters corresponding to the first data group "a literal encryption and decryption algorithm".

[0124] Again, for example, refer to Figure 8 As shown, the first data group "Someone Ming plans to take the train to a certain city to play with his college classmate Wang tomorrow" includes 22 sorted literal characters. According to the above method, the first 16 sorted literal characters "Someone Ming plans to take the train to a certain city to play with his college classmate Wang tomorrow" in the first data group can be used to sequentially replace each blank matrix element in the first blank symmetric matrix to obtain the first matrix of literal characters. Then, the next 4 sorted literal characters "take the train to a certain" in the first data can be used to sequentially replace each blank matrix element in the second blank symmetric matrix to obtain the second matrix of literal characters. Finally, the last remaining literal characters "city to play" in the first data are used to replace the blank matrix elements at the positions of the first row and first column, and the first row and second column in the third blank symmetric matrix in sequence, and the blank matrix elements at the positions of the second row and first column, and the second row and second column are retained to obtain the third matrix of literal characters.

[0125] In some embodiments, the pinyin data of literal characters is obtained and based on it, the literal characters in the matrix of literal characters are mapped to corresponding combinations of alphabetic characters to obtain a corresponding first matrix of alphabetic characters; wherein, the first matrix of alphabetic characters uses combinations of alphabetic characters as matrix elements.

[0126] Specifically, taking the current matrix of literal characters in multiple matrices of literal characters as an example, for any current literal character in the current matrix of literal characters, first, based on the pinyin data of the current literal character, the alphabetic character group of the current literal character and the tone information can be determined; then, according to the preset conversion rules, the tone identifier corresponding to the tone information of the current literal character can be determined; the alphabetic character group of the current literal character and the tone identifier are combined to obtain the corresponding combination of alphabetic characters. Thus, the current matrix of literal characters can be mapped to the corresponding first matrix of alphabetic characters.

[0127] Among them, the above alphabetic character group can specifically be a combination of initial consonant characters and final consonant characters, or can only contain final consonant characters. For example, the alphabetic character group of the literal character "zhong" can be expressed as "zhong", and the alphabetic character group of the literal character "e" can be expressed as "e".

[0128] Specifically, the tone identifier may be an identifier corresponding to tone information based on a preset conversion rule. For example, based on a preset conversion rule, the tone identifier corresponding to the first tone may be denoted as "1", the tone identifier corresponding to the second tone may be denoted as "2", the tone identifier corresponding to the third tone may be denoted as "3", the tone identifier corresponding to the fourth tone may be denoted as "4", and the tone identifier corresponding to the light tone may be denoted as "0".

[0129] During specific implementation, the alphabetic character group and the tone identifier of the text characters may be combined in the order of the alphabetic character group first and then the tone identifier to obtain the corresponding intermediate data. For example, the intermediate data group corresponding to the text character "声" may be denoted as: "sheng1", and the intermediate data group corresponding to the text character "汉" may be denoted as: "han4".

[0130] Specifically, referring to Figure 9 as shown, the first data group "a text encryption and decryption algorithm" may be first converted into a corresponding text character matrix; then the text character matrix may be further mapped into a corresponding first alphabetic character matrix.

[0131] In some embodiments, referring to Figure 10 as shown, the above-mentioned determination of the alphabetic character combinations satisfying the exchange condition in the first alphabetic character matrix according to the preset exchange rule includes:

[0132] S1: Determine the matrix coordinates of the alphabetic character combination according to the row number and column number of the alphabetic character combination in the first alphabetic character matrix; wherein, the matrix coordinates include row coordinates and column coordinates;

[0133] S2: According to the preset exchange rule, determine the first matrix coordinates and the second matrix coordinates that mutually satisfy the preset data relationship by retrieving the matrix coordinates of the alphabetic character combinations in the same first alphabetic character matrix;

[0134] S3: Determine the first alphabetic character combination indicated by the first matrix coordinates and the second alphabetic character combination indicated by the second matrix coordinates in the same first alphabetic character matrix as the alphabetic character combinations satisfying the exchange condition.

[0135] Based on the above embodiments, two alphabetic character combinations that satisfy the exchange condition and need to be exchanged with each other can be quickly found in the first alphabetic character matrix.

[0136] In some embodiments, the preset data relationship may specifically include: the row coordinate of the first matrix coordinate is equal to the column coordinate of the second matrix coordinate, and the column coordinate of the first matrix coordinate is equal to the row coordinate of the second matrix coordinate; and / or, the row coordinate and column coordinate of the first matrix coordinate are equal, the row coordinate and column coordinate of the second matrix coordinate are equal, and the sum of the row coordinate of the first matrix coordinate and the row coordinate of the second matrix coordinate is equal to the row number of the first letter character matrix, etc.

[0137] Based on the above embodiments, the letter character combinations that meet the exchange conditions can be determined by using the preset data relationships. In this way, subsequent processing only needs to be performed on the letter character combinations that meet the exchange conditions, rather than all letter character combinations, which can effectively hide the real data information and reduce the amount of related data processing.

[0138] In some embodiments, the above-described process of exchanging initial consonants and / or final vowels in letter character combinations that satisfy the exchange conditions may include the following:

[0139] S1: Detect whether the initial consonant in the first letter character combination and the final vowel in the second letter character combination satisfy the pinyin combination relationship; Detect whether the final vowel in the first letter character combination and the initial consonant in the second letter character combination satisfy the pinyin combination relationship;

[0140] S2: If the initial consonant in the first letter character combination and the final vowel in the second letter character combination satisfy the pinyin combination relationship, and the final vowel in the first letter character combination and the initial consonant in the second letter character combination satisfy the pinyin combination relationship, then swap the initial consonant in the first letter character combination with the initial consonant in the second letter character combination.

[0141] Based on the above embodiments, the true relevant information can be hidden by accurately and effectively exchanging the initial consonant characters and / or final vowel characters of letter character combinations that meet the exchange conditions.

[0142] In some embodiments, after detecting whether the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination satisfy a pinyin combination relationship; and after detecting whether the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination satisfy a pinyin combination relationship, the method may further include the following:

[0143] If it is determined that the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination do not satisfy the pinyin combination relationship, and / or, the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination do not satisfy the pinyin combination relationship, the first letter character combination and the second letter character combination shall not be swapped.

[0144] Based on the above embodiments, when it is found that the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination do not satisfy the pinyin combination relationship, and / or the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination do not satisfy the pinyin combination relationship, they can be left unexchanged to avoid errors and affect subsequent processing.

[0145] In some embodiments, the above-mentioned satisfaction of the pinyin combination relationship can be specifically understood as, based on the pinyin combination rules, the initial consonant character can be combined with the final vowel character to obtain a reasonable pinyin letter combination.

[0146] For example, in the first letter character combination "cheng1", the initial consonant is "ch" and the final vowel is "eng". In the same first-letter character matrix, the second letter character combination that satisfies the commutative condition is "zhong3", with the initial consonant "zh" and the final vowel "ong". Based on the pinyin combination rules, since the initial consonant "zh" can combine with the final vowel "eng" to obtain the reasonable pinyin combination "zheng", they satisfy the pinyin combination relationship. Similarly, the initial consonant "ch" can also combine with the final vowel "ong" to obtain the reasonable pinyin combination "chong". Therefore, they also satisfy the pinyin combination relationship.

[0147] At this point, the initial consonants of “cheng1” and “zhong3” in the first letter character matrix that satisfy the exchange condition can be swapped, resulting in the processed letter character combinations “zheng1” and “chong3”.

[0148] For example, consider the first-letter character combination "zhong3" and the second-letter character combination "ji1" in the same first-letter character matrix that satisfy the exchange condition. Based on the rules of pinyin combination, the initial consonant "j" cannot be combined with the final vowel "ong." The resulting "jong" is illogical and does not exist; therefore, the two do not satisfy the pinyin combination relationship. In this case, the first-letter character combination "zhong3" and the second-letter character combination "ji1" that satisfy the exchange condition are not exchanged and remain unchanged.

[0149] Specifically, for example, see Figure 11 As shown, with Figure 9Taking the first letter character matrix as an example, firstly, the matrix coordinates (which can be denoted as the first matrix coordinates) of the first letter character combination "zhong3" are 1 in the row and 2 in the column; the matrix coordinates (which can be denoted as the second matrix coordinates) of the second letter character combination "zi4" are 2 in the row and 1 in the column. It can be seen that the row coordinate of the first matrix is ​​equal to the column coordinate of the second matrix, and the column coordinate of the first matrix is ​​equal to the row coordinate of the second matrix. Therefore, the first matrix coordinates and the second matrix coordinates satisfy the preset data relationship; correspondingly, the first letter character combination "zhong3" corresponding to the first matrix coordinates and the second letter character combination "zi4" corresponding to the second matrix coordinates satisfy the exchange condition.

[0150] Furthermore, it checks whether the initial consonant "zh" in the first letter character combination and the final vowel "i" in the second letter character combination satisfy a pinyin combination relationship; it also checks whether the final vowel "ong" in the first letter character combination and the initial consonant "z" in the second letter character combination satisfy a pinyin combination relationship. If both satisfy the pinyin combination relationship, the initial consonants of the first and second letter character combinations can be swapped. Accordingly, the swapped first letter character combination becomes "zhi4", and the swapped second letter character combination becomes "zong3".

[0151] In some embodiments, the method may further include the following:

[0152] S1: Detect whether there is a consonant character in the first letter character combination and the second letter character combination;

[0153] S2: If it is determined that there is no initial consonant character in the first letter character combination and / or there is no initial consonant character in the second letter character combination, the first letter character combination and the second letter character combination shall not be swapped.

[0154] Based on the above embodiments, if it is found that there is no initial consonant character in the first letter character combination and / or the second letter character combination does not contain an initial consonant character, no swapping is required to avoid errors that could affect subsequent processing.

[0155] In some embodiments, after obtaining the corresponding second data group, the method may further include the following:

[0156] The second data group is sent to the target database for storage; and the target audio data and the first data group are deleted; wherein the target database is a blockchain-based database.

[0157] Based on the above embodiments, after obtaining the second data group, the target audio data and the first data group that directly contain real relevant information can be eliminated in a timely manner, and only the second data group that hides the real relevant information can be retained to avoid the leakage of real relevant information; at the same time, the immutability of blockchain can be used to securely save the second data group for evidence storage, so as to facilitate subsequent retrospective query.

[0158] In some embodiments, obtaining the corresponding second data group based on the second letter character matrix can specifically include: first, mapping the letter character combinations in the second letter character matrix to corresponding text characters according to the pinyin data indicated by the letter character combinations, to obtain a processed text character matrix; then, sequentially extracting the corresponding text characters from the processed text character matrix and concatenating them in a left-to-right, top-to-bottom order to obtain a second data group that hides the real relevant information.

[0159] In some embodiments, when mapping letter character combinations in the second letter character matrix to corresponding text characters based on the pinyin data indicated by the letter character combinations, sometimes a single letter character combination may indicate a pinyin data that corresponds to multiple text characters (e.g., multiple homophones). In this case, multiple text characters can be treated as undetermined characters for mapping other letter character combinations. After completing the mapping of the letter character combinations and concatenating the mapped text characters in sequence, the semantic correlation between each undetermined text character and its adjacent text characters can be calculated sequentially. Then, based on the semantic correlation, the undetermined character with the highest semantic correlation is selected from the multiple undetermined characters as the text character corresponding to the restored intermediate data.

[0160] In some embodiments, see Figure 12 As shown, in specific implementations, the method may also include the following:

[0161] S1: Obtain the second data set;

[0162] S2: According to the preset splitting rules, the second data group is split to obtain multiple text character matrices; wherein, the text character matrices use text characters as matrix elements;

[0163] S3: Obtain the pinyin data of the characters and map the characters in the character matrix to the corresponding letter character combinations to obtain the corresponding second letter character matrix; wherein, the second letter character matrix uses letter character combinations as matrix elements;

[0164] S4: Based on the preset exchange rules, determine the letter character combinations in the second letter character matrix that satisfy the exchange conditions; and exchange the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain the corresponding first letter character matrix;

[0165] S5: Based on the first letter character matrix, obtain the corresponding first data group.

[0166] Based on the above embodiments, when target text data is needed, the first data group containing real relevant information can be obtained again by performing corresponding restoration processing on the second data group according to preset processing rules.

[0167] In practice, the first data set can be used for verification to determine whether the user meets the requirements; and then to determine whether to provide the relevant transaction services to the user.

[0168] In practice, the second data group can be restored in a trusted execution environment according to the preset processing rules in the manner described above, so as to avoid the leakage of the real relevant information carried in the first data group.

[0169] As can be seen from the above, based on the audio data processing method provided in this specification, after receiving the target audio data and obtaining the corresponding first data group through speech recognition, the first data group is first split according to a preset splitting rule to obtain multiple character matrices; then, based on the pinyin data of the characters, the characters in the character matrix are mapped to corresponding letter character combinations to obtain the corresponding first letter character matrix; according to a preset exchange rule, the letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined; and the initial consonants and / or final vowels of the letter character combinations that satisfy the exchange conditions are exchanged to obtain the corresponding second letter character matrix; based on the second letter character matrix, the corresponding second data group is obtained. This fully utilizes the spatial characteristics and computational advantages of the matrix structure, as well as the relevant characteristics of pinyin data, to efficiently process related data, reliably hide the true data information, effectively prevent the leakage of data information carried in the audio data, and better protect the data security of the audio data.

[0170] See Figure 13 As shown in the embodiments of this specification, a specific server is also provided, wherein the server includes a network communication port 1301, a processor 1302 and a memory 1303, and the above structures are connected by internal cables so that the various structures can perform specific data interaction.

[0171] Specifically, the network communication port 1301 can be used to receive target audio data.

[0172] The processor 1302 is specifically used to perform speech recognition on target audio data to obtain a corresponding first data group; wherein the first data group includes multiple sequentially arranged text characters; according to a preset splitting rule, the first data group is split to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements; acquire and map the text characters in the text character matrices to corresponding letter character combinations based on the pinyin data of the text characters, to obtain a corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements; according to a preset exchange rule, determine the letter character combinations in the first letter character matrix that satisfy the exchange conditions; and exchange the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain a corresponding second letter character matrix; and obtain a corresponding second data group based on the second letter character matrix.

[0173] The memory 1303 can be used to store the corresponding instruction program.

[0174] In this embodiment, the network communication port 1301 can be a virtual port bound to different communication protocols, thereby enabling the sending or receiving of different data. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. Furthermore, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM or CDMA; it can also be a Wi-Fi chip; or it can be a Bluetooth chip.

[0175] In this embodiment, the processor 1302 can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. This specification is not limiting.

[0176] In this embodiment, the memory 1303 may include multiple layers. In a digital system, anything that can store binary data can be a memory. In an integrated circuit, a circuit with storage function but no physical form is also called a memory, such as RAM, FIFO, etc. In a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.

[0177] This specification also provides a computer-readable storage medium based on the above-described audio data processing method. The computer-readable storage medium stores computer program instructions that, when executed, implement the following: receiving target audio data; performing speech recognition on the target audio data to obtain a corresponding first data group; wherein the first data group includes multiple sequentially arranged text characters; splitting the first data group according to a preset splitting rule to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements; acquiring and mapping the text characters in the text character matrices to corresponding letter character combinations based on the pinyin data of the text characters, to obtain a corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements; determining letter character combinations in the first letter character matrix that satisfy the exchange conditions according to a preset exchange rule; and exchanging the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain a corresponding second letter character matrix; and obtaining a corresponding second data group based on the second letter character matrix.

[0178] In this embodiment, the storage medium includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions. The network communication unit can be an interface configured according to standards specified in the communication protocol for network connection communication.

[0179] In this embodiment, the specific functions and effects implemented by the program instructions stored in the computer-readable storage medium can be explained in comparison with other embodiments, and will not be repeated here.

[0180] This specification also provides a computer program product comprising a computer program that, when executed by a processor, performs the following steps: receiving target audio data; performing speech recognition on the target audio data to obtain a corresponding first data group; wherein the first data group includes multiple sequentially arranged text characters; splitting the first data group according to a preset splitting rule to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements; acquiring and mapping the text characters in the text character matrices to corresponding letter character combinations based on the pinyin data of the text characters, to obtain a corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements; determining the letter character combinations in the first letter character matrix that satisfy the exchange conditions according to a preset exchange rule; and exchanging the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain a corresponding second letter character matrix; and obtaining a corresponding second data group based on the second letter character matrix.

[0181] See Figure 14 As shown, at the software level, this specification also provides an audio data processing apparatus, which may specifically include the following structural modules:

[0182] The receiving module 1401 can be used to receive target audio data;

[0183] The speech recognition module 1402 is specifically used to perform speech recognition on target audio data to obtain a corresponding first data group; wherein, the first data group includes multiple text characters arranged in sequence;

[0184] The splitting module 1403 can be used to split the first data group according to a preset splitting rule to obtain multiple text character matrices; wherein the text character matrices use text characters as matrix elements.

[0185] The mapping module 1404 is specifically used to acquire and map the text characters in the text character matrix into corresponding letter character combinations based on the pinyin data of the text characters, thereby obtaining the corresponding first letter character matrix; wherein, the first letter character matrix uses letter character combinations as matrix elements;

[0186] The exchange module 1405 is specifically used to determine the letter character combinations in the first letter character matrix that satisfy the exchange conditions according to preset exchange rules; and to exchange the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain the corresponding second letter character matrix.

[0187] The processing module 1406 can be used to obtain the corresponding second data group based on the second letter character matrix.

[0188] In some embodiments, the target audio data may specifically include audio data involving sensitive information provided by the user when conducting a transaction.

[0189] In some embodiments, the apparatus may further include a preprocessing module, specifically used to preprocess the target audio data after receiving the target audio data; wherein the preprocessing includes at least one of the following: background noise filtering, dialect speech correction, and business scenario matching.

[0190] In some embodiments, where the preprocessing includes business scenario matching, the speech recognition module 1402 can specifically perform speech recognition on the target audio data in the following manner to obtain the corresponding first data group: based on the business scenario matching result, determine the matching preset speech recognition model from multiple preset speech recognition models as the target speech recognition model; process the target audio data using the target speech recognition model to obtain the corresponding target speech recognition result; and extract and combine text characters involving sensitive information based on the target speech recognition result to obtain the first data group.

[0191] In some embodiments, after obtaining the corresponding second data group, the device may also be used to send the second data group to a target database for storage; and delete the target audio data and the first data group; wherein the target database is a blockchain-based database.

[0192] In some embodiments, when the splitting module 1403 is specifically implemented, it can split the first data group according to a preset splitting rule in the following manner to obtain multiple text character matrices: Multiple blank symmetric matrices are created according to the preset splitting rule and the number of characters in the first data group; wherein the multiple blank symmetric matrices are arranged in order, and the number of rows of the symmetric matrix that is sorted first is greater than or equal to the number of rows of the symmetric matrix that is sorted last; according to the sorting information of the text characters in the first data group, the text characters in the first data group are assigned to the corresponding blank symmetric matrices to obtain multiple sub-data groups corresponding to the multiple blank symmetric matrices; according to the sorting information of the text characters in the first data group, the blank matrix elements in the corresponding blank symmetric matrices are replaced with the text characters in the sub-data groups to obtain multiple text character matrices.

[0193] In some embodiments, when the above-described exchange module 1405 is specifically implemented, it can determine the letter character combinations that satisfy the exchange conditions in the first letter character matrix according to the following method based on the preset exchange rules: determine the matrix coordinates of the letter character combinations based on the row and column numbers of the letter character combinations in the first letter character matrix; wherein, the matrix coordinates include row coordinates and column coordinates; determine the first matrix coordinates and the second matrix coordinates that satisfy the preset data relationship by searching the matrix coordinates of letter character combinations in the same first letter character matrix according to the preset exchange rules; and determine the first letter character combination indicated by the first matrix coordinates and the second letter character combination indicated by the second matrix coordinates in the same first letter character matrix as the letter character combinations that satisfy the exchange conditions.

[0194] In some embodiments, the preset data relationship may specifically include: the row coordinate of the first matrix coordinate is equal to the column coordinate of the second matrix coordinate, and the column coordinate of the first matrix coordinate is equal to the row coordinate of the second matrix coordinate; and / or, the row coordinate and column coordinate of the first matrix coordinate are equal, the row coordinate and column coordinate of the second matrix coordinate are equal, and the sum of the row coordinate of the first matrix coordinate and the row coordinate of the second matrix coordinate is equal to the row number of the first letter character matrix, etc.

[0195] In some embodiments, when the above-described exchange module 1405 is specifically implemented, it can perform exchange processing on the initial consonant characters and / or final vowel characters of letter character combinations that meet the exchange conditions in the following manner: detecting whether the initial consonant characters in the first letter character combination and the final vowel characters in the second letter character combination satisfy a pinyin combination relationship; detecting whether the final vowel characters in the first letter character combination and the initial consonant characters in the second letter character combination satisfy a pinyin combination relationship; if it is determined that the initial consonant characters in the first letter character combination and the final vowel characters in the second letter character combination satisfy a pinyin combination relationship, and the final vowel characters in the first letter character combination and the initial consonant characters in the second letter character combination satisfy a pinyin combination relationship, then exchange the initial consonant characters in the first letter character combination with the initial consonant characters in the second letter character combination.

[0196] In some embodiments, when the above-described exchange module 1405 is specifically implemented, after detecting whether the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination satisfy the pinyin combination relationship; and after detecting whether the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination satisfy the pinyin combination relationship, it can also be used to not exchange the first letter character combination and the second letter character combination if it is determined that the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination do not satisfy the pinyin combination relationship, and / or, the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination do not satisfy the pinyin combination relationship.

[0197] In some embodiments, when the above-described exchange module 1405 is specifically implemented, it can also detect whether there is a consonant character in the first letter character combination and the second letter character combination; if it is determined that there is no consonant character in the first letter character combination and / or the second letter character combination, the first letter character combination and the second letter character combination are not exchanged.

[0198] It should be noted that the units, devices, or modules described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. For ease of description, the above devices are described by dividing them into various modules according to their functions. Of course, in implementing this specification, the functions of each module can be implemented in one or more software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection between the devices or units shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0199] As can be seen from the above, the audio data processing device provided in the embodiments of this specification can make full use of the spatial characteristics and computational advantages of the matrix structure, as well as the relevant characteristics of the pinyin data, to efficiently process the relevant data, reliably hide the real data information, effectively prevent the data information carried in the audio data from being leaked, and better protect the data security of the audio data.

[0200] While this specification provides the steps of operation for the methods described in the embodiments or flowcharts, more or fewer steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or client product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded. The terms "first," "second," etc., are used to denote names and do not indicate any particular order.

[0201] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer-readable storage media, including storage devices.

[0202] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. This specification can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0203] Although this specification has been described by way of examples, those skilled in the art will recognize that many variations and modifications are possible without departing from the spirit of this specification, and it is intended that the appended claims cover such variations and modifications without departing from the spirit of this specification.

Claims

1. A method of processing audio data, characterized by, include: Receive target audio data; Speech recognition is performed on the target audio data to obtain the corresponding first data group; wherein, the first data group includes multiple text characters arranged in order; According to a preset splitting rule, the first data group is split to obtain multiple character matrices, including: creating multiple blank symmetric matrices according to the preset splitting rule and the number of characters in the first data group; wherein the multiple blank symmetric matrices are arranged in order, and the number of rows of the symmetric matrix that is sorted first is greater than or equal to the number of rows of the symmetric matrix that is sorted later; according to the sorting information of the characters in the first data group, the characters in the first data group are assigned to the corresponding blank symmetric matrices to obtain multiple sub-data groups corresponding to the multiple blank symmetric matrices; according to the sorting information of the characters in the first data group, the characters in the sub-data groups are used to replace the blank matrix elements in the corresponding blank symmetric matrices to obtain multiple character matrices, wherein the character matrices use characters as matrix elements; Based on the pinyin data of the text characters, the text characters in the text character matrix are mapped to the corresponding letter character combinations to obtain the corresponding first letter character matrix; wherein, the first letter character matrix uses letter character combinations as matrix elements; According to the preset exchange rules, the letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined; and the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions are exchanged to obtain the corresponding second letter character matrix; Based on the second letter character matrix, the corresponding second data group is obtained.

2. The method according to claim 1, characterized in that, The target audio data includes audio data involving sensitive information provided by users when conducting transactions.

3. The method according to claim 1, characterized in that, After receiving the target audio data, the method further includes: The target audio data is preprocessed; wherein the preprocessing includes at least one of the following: background noise filtering, dialect speech correction, and business scenario matching.

4. The method according to claim 3, characterized in that, In the case that the preprocessing includes business scenario matching, speech recognition is performed on the target audio data to obtain the corresponding first data group, including: Based on the matching results of the business scenario, the matching preset speech recognition model is determined from multiple preset speech recognition models as the target speech recognition model; The target audio data is processed using a target speech recognition model to obtain the corresponding target speech recognition result; Based on the target speech recognition results, the text characters involving sensitive information are extracted and combined to obtain the first data group.

5. The method according to claim 1, characterized in that, After obtaining the corresponding second data set, the method further includes: The second data group is sent to the target database for storage; and the target audio data and the first data group are deleted; wherein the target database is a blockchain-based database.

6. The method according to claim 1, characterized in that, Based on the preset exchange rules, the letter character combinations in the first letter character matrix that satisfy the exchange conditions are determined, including: The matrix coordinates of the letter character combination are determined based on the row and column number of the letter character combination in the first letter character matrix; wherein, the matrix coordinates include row coordinates and column coordinates; According to the preset exchange rules, the first matrix coordinates and the second matrix coordinates that satisfy the preset data relationship are determined by retrieving the matrix coordinates of letter character combinations in the same first letter character matrix; The combination of first-letter characters indicated by the first matrix coordinate and the combination of second-letter characters indicated by the second matrix coordinate in the same first-letter character matrix are determined as letter character combinations that satisfy the exchange condition.

7. The method according to claim 6, characterized in that, The preset data relationships include: the row coordinate of the first matrix coordinate is equal to the column coordinate of the second matrix coordinate, and the column coordinate of the first matrix coordinate is equal to the row coordinate of the second matrix coordinate; and / or, the row coordinate and column coordinate of the first matrix coordinate are equal, the row coordinate and column coordinate of the second matrix coordinate are equal, and the sum of the row coordinate of the first matrix coordinate and the row coordinate of the second matrix coordinate is equal to the row number of the first letter character matrix.

8. The method according to claim 6, characterized in that, The process involves swapping initial consonants and / or final vowels in letter character combinations that meet the swapping criteria, including: Detect whether the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination satisfy a pinyin combination relationship; Detect whether the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination satisfy a pinyin combination relationship; If the initial consonant in the first letter character group and the final vowel in the second letter character group satisfy a phonetic combination relationship, and the final vowel in the first letter character group and the initial consonant in the second letter character group satisfy a phonetic combination relationship, then swap the initial consonant in the first letter character group with the initial consonant in the second letter character group.

9. The method according to claim 8, characterized in that, The test checks whether the initial consonant in the first letter character combination and the final vowel in the second letter character combination satisfy a phonetic combination relationship. After detecting whether the vowel character in the first letter character combination and the initial consonant character in the second letter character combination satisfy a pinyin combination relationship, the method further includes: If it is determined that the initial consonant character in the first letter character combination and the final vowel character in the second letter character combination do not satisfy the pinyin combination relationship, and / or, the final vowel character in the first letter character combination and the initial consonant character in the second letter character combination do not satisfy the pinyin combination relationship, the first letter character combination and the second letter character combination shall not be swapped.

10. The method according to claim 8, characterized in that, The method further includes: Detect whether there is a consonant character in the first letter character combination and the second letter character combination; If it is determined that there is no initial consonant character in the first letter character combination and / or the second letter character combination, the first letter character combination and the second letter character combination shall not be swapped.

11. An audio data processing apparatus, characterized in that, include: The receiving module is used to receive target audio data; The speech recognition module is used to perform speech recognition on the target audio data to obtain the corresponding first data group; wherein, the first data group includes multiple text characters arranged in order; The splitting module is used to split the first data group according to a preset splitting rule to obtain multiple text character matrices. This includes: creating multiple blank symmetric matrices according to the preset splitting rule and the number of characters in the first data group; wherein the multiple blank symmetric matrices are arranged in order, and the number of rows in the first symmetric matrix is ​​greater than or equal to the number of rows in the second symmetric matrix; assigning the text characters in the first data group to the corresponding blank symmetric matrices according to the order of the text characters in the first data group, obtaining multiple sub-data groups corresponding to the multiple blank symmetric matrices; and replacing the blank matrix elements in the corresponding blank symmetric matrices with the text characters from the sub-data groups according to the order of the text characters in the first data group, obtaining multiple text character matrices, wherein the text character matrices use text characters as matrix elements. The mapping module is used to acquire and map the characters in the character matrix to corresponding letter character combinations based on the pinyin data of the characters, thereby obtaining the corresponding first letter character matrix; wherein the first letter character matrix uses letter character combinations as matrix elements. The exchange module is used to determine the letter character combinations in the first letter character matrix that satisfy the exchange conditions according to the preset exchange rules; and to exchange the initial consonant characters and / or final vowel characters of the letter character combinations that satisfy the exchange conditions to obtain the corresponding second letter character matrix. The processing module is used to obtain the corresponding second data group based on the second letter character matrix.

12. A server, characterized in that, It includes a processor and a memory for storing processor-executable instructions, wherein the processor, when executing the instructions, implements the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.