Travel voice privacy protection method and device and storage medium
By combining BERT embedding, Transformer neural network and graph convolution network methods, the detection lag problem in sensitive text detection is solved, and efficient sensitive word recognition and real-time privacy protection in complex scenarios is achieved, which is suitable for voice privacy protection for online car-hailing and taxi.
Patent Information
- Application Number
- CN202510384724.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the face of complex scenarios such as emerging hot meme variants and cross-language expressions, sensitive text detection has problems such as lag in detection and high false alarm rate, making it difficult to achieve real-time privacy protection.
BERT embedding and Transformer neural network encoding combined with graph convolution network are used to construct a text graph structure based on PMI and TF-IDF to achieve deep interaction between word-level and word-level features, and filter sensitive words using the cross-attention mechanism to improve detection efficiency and recall.
It significantly improves the accuracy and detection efficiency of sensitive words in complex contexts, and is suitable for real-time privacy protection in travel scenarios, especially voice privacy protection in online car-hailing and taxis.
Smart Images

Figure CN120256602A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the technical field of data processing, and particularly relates to a method, device and storage medium for protecting trip voice privacy. Background Art
[0002] In daily travel scenarios such as online car-hailing and taxis, passengers and drivers will have a large amount of communication during the trip, and these conversation voices are very likely to contain sensitive words. If these sensitive words are improperly spread or used, it will not only violate the privacy of both parties, but also may cause a series of security problems and disputes. In order to protect the rights and interests of both drivers and passengers and maintain a good travel communication environment, it has become extremely urgent to accurately identify sensitive words in the conversation voices of drivers and passengers. Since there is relatively little research on privacy protection AI models for sensitive information filtering, the related task is the sensitive text detection task. And the related technologies of sensitive text detection have developed rapidly, which can be roughly divided into rule-based matching, machine learning and deep learning methods, as follows:
[0003] (1) Rule-based matching methods, whose rules are manually set for detection work. Common techniques include single-pattern string and multi-pattern string matching algorithms. First, there are various single-pattern string matching algorithms. The edit distance algorithm quantifies the similarity between the target and sensitive words by calculating the number of letter operations at the pinyin level, and determines sensitivity when the threshold is reached. The KMP algorithm constructs a partial match table and adjusts the matching position according to the table when scanning the text, efficiently finding variants of sensitive words formed by splitting Chinese characters. The approximate string matching algorithm mines the mapping of special characters similar in shape and sound to English letters, and judges vulgar and sensitive information by comparing the matching score with the threshold. The association rule-based algorithm analyzes the text based on common variant rules and associated word libraries, expands and associates suspected sensitive word fragments, and improves the recognition accuracy of complex variants. Second, the multi-pattern string matching algorithm is typically the AC-Trie prefix tree. In the scenario of social network text streams, hot sensitive phrases are built into a prefix tree, and nodes store characters and the occurrence frequencies of phrases. During detection, the text character sequence matches along the branches, and the phrase frequency is obtained at the leaf node, and various heat calculation methods are used to evaluate the sensitive heat. In the verification of the Sina Weibo dataset, this algorithm has excellent performance and can match multiple hot sensitive phrases in milliseconds, efficiently screening sensitive Weibo content. However, in the detection system based on rule matching, due to the manual preset and relative static nature of the rules, in the face of complex network language environments such as emerging hot meme variants and cross-lingual sensitive expressions, there are problems such as detection lag and rising false negative and false positive rates, and more intelligent detection technology innovation is urgently needed.
[0004] (2) The machine learning-based method, compared with the rule-based matching method, the advantage of machine learning algorithms lies in their ability to automatically learn the feature patterns of sensitive and non-sensitive words from a large amount of labeled data. By extracting features such as text word frequency and part of speech and converting them into vectors, the model can learn the association between features and sensitive words, greatly improving the detection efficiency and enabling the processing of larger-scale text data. Common machine learning algorithms include: support vector machines, naive Bayes, decision trees, etc. Support vector machines distinguish sensitive and non-sensitive texts effectively in the feature space by finding an optimal hyperplane. It is good at dealing with non-linear classification problems and has good adaptability to complex text feature distributions. In some scenarios of detecting sensitive words in enterprise internal documents with high requirements for detection accuracy, support vector machines can accurately identify documents containing sensitive information and reduce the occurrence of false positives and false negatives. The naive Bayes algorithm calculates the probability that a text belongs to the sensitive category based on Bayes' theorem and the assumption of feature conditional independence. It shows the characteristics of high computational efficiency and simple model in text classification tasks and can quickly judge whether a new text is sensitive or not. The decision tree algorithm divides text features layer by layer in a tree structure and determines the classification direction of the text according to different feature values. Its advantage is that the model is interpretable and easy to understand and analyze. However, the machine learning method depends on manually extracted features, which not only requires a large amount of human and time costs, but also the quality of feature extraction directly affects the detection effect. In the face of the ever-changing network language and the endless emergence of new sensitive expressions, it is difficult for manually extracted features to adapt quickly, resulting in a decline in detection performance.
[0005] (3) The deep learning-based method, deep learning methods such as recurrent neural networks (RNN) and its variants long short-term memory networks (LSTM), convolutional neural networks (CNN), Transformer, etc., can automatically learn rich feature representations from text data without the need for complex feature engineering by humans. RNN and LSTM are particularly suitable for processing data with sequential features such as text. They can capture long-distance semantic dependencies in text and are very effective for understanding sensitive information in the context. The LSTM model can accurately judge whether subsequent ambiguous expressions are sensitive content based on the semantic information of the previous text. CNN extracts local text features by sliding convolutional kernels over the text and has high computational efficiency when dealing with large-scale text data. The Transformer model can quickly and accurately grasp the semantic context of the whole text and accurately identify sensitive information hidden in complex contexts. It extracts and fuses text features from different perspectives through the multi-head attention mechanism, making the model's understanding of the text more comprehensive and in-depth.
[0006] However, the existing research has the following deficiencies in solving the sensitive text detection task:
[0007] (1) The existing detection methods have poor effects. For the rule-based matching method, the rules are manually set and static. Facing emerging hot meme variants and cross-language sensitive expressions, the detection is lagging, with many false positives and false negatives. Moreover, in actual detection, the combination of rule matching and semantic understanding is not flexible enough, resulting in insufficient ability to detect sensitive information with complex semantics. For the machine learning-based method, it relies on manual feature extraction, which is costly and difficult to adapt to new sensitive expressions, leading to performance degradation.
[0008] (2) The existing sensitive word detection schemes similar in deep learning technology have defects in input form and information fusion. Most are based on a single input form combined with a deep learning model, and only adopt the scheme of word-level input. Although it can learn the semantic relationship between words with the help of Transformer, it ignores the character-level information and is prone to losing key details such as spelling variants of sensitive words. The scheme that focuses on character-level input can capture character information, but it is not as good as the word-level input scheme in understanding the overall semantic structure. And the scheme that tries to fuse word and character information only simply extracts features independently and then concatenates them, without fully considering the feature interaction relationship between the two, making it difficult to effectively utilize complementary information, resulting in limited improvement in the accuracy and recall rate of sensitive word detection.
[0009] (3) In addition, the existing methods are difficult to accurately grasp the semantic role and function of words in the whole document, and cannot accurately judge the importance and sensitivity of words according to the overall context of the document. And it is difficult to comprehensively capture the rich and diverse semantic relationships between words, such as semantic similarity, opposition, implication, etc. The interaction relationship between features is not fully considered, resulting in the inability to fully consider the association of text elements when judging sensitive text.
[0010] Based on the above analysis, it can be concluded that the current sensitive text detection methods and the existing deep learning methods cannot achieve the ideal effect of accurately identifying and shielding sensitive information in the text to achieve privacy protection.
[0011] Therefore, there is an urgent need for a new method, device and storage medium for trip voice privacy protection to solve the above technical problems. Summary of the Invention
[0012] The present invention provides a method, device and storage medium for trip voice privacy protection, aiming to solve the problem of lagging detection in complex scenarios such as emerging hot meme variants and cross-language expressions, and meet the real-time privacy protection requirements in the travel scenario.
[0013] In the first aspect, the present invention provides a method for trip voice privacy protection, and the method for trip voice privacy protection includes the following steps:
[0014] S1. Obtain voice data, and perform speech-to-text processing on the voice data to convert it into text data;
[0015] S2. Collect historical itinerary text data, perform data annotation and data processing on the historical itinerary text data to obtain an itinerary text data set; perform model training based on the itinerary text data set to obtain a privacy protection model;
[0016] S3. Use the text data as the input of the privacy protection model, filter sensitive words in the text data to obtain privacy protection text data;
[0017] S4. Perform text-to-speech processing on the privacy protection text data to obtain privacy protection voice data.
[0018] Preferably, in step S3, the following sub-steps are included:
[0019] S31. Use the text data as the input of the privacy protection model, perform preprocessing on the text data to obtain a word segmentation result and a character segmentation result;
[0020] S32. Perform BERT embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain word-level vector features and character-level vector features;
[0021] S33. Construct a text graph, perform graph convolution processing on the word-level vector features and the character-level vector features according to the text graph to obtain word-level graph convolution features and character-level graph convolution features;
[0022] S34. Based on the cross-attention mechanism, fuse the word-level graph convolution features and the character-level graph convolution features to obtain fused features; perform residual connection and normalization optimization processing on the fused features, and filter out sensitive words in the text data based on the optimized fused features to obtain the privacy protection text data.
[0023] Preferably, in step S32, the following sub-steps are included:
[0024] S321. Respectively perform vector conversion on the word segmentation result and the character segmentation result, and add classification markers and sentence delimiters at the beginning and end of the sentences of the word segmentation result and the character segmentation result to obtain word segmentation vectors and character segmentation vectors; wherein, the classification marker is used to aggregate global semantics, and the sentence delimiter is used to define the text boundary;
[0025] S322. Respectively perform word embedding processing, segmentation embedding processing, and position embedding processing on the word segmentation vector and the character segmentation vector to obtain word-level embedding results and character-level embedding results;
[0026] S323. The word-level embedding result and the character-level embedding result are encoded through the Transformer neural network to obtain the word-level vector feature and the character-level vector feature.
[0027] Preferably, in step S33, the following sub-steps are included:
[0028] S331. Construct the text graph, which includes a document node and word nodes. The edges between the document node and the word nodes are constructed according to the number of times the corresponding word appears in the document, and the edges between two word nodes are constructed according to the probability of the corresponding word appearing in the corpus. Wherein, the corpus is composed of the data sets of the documents corresponding to all the document nodes;
[0029] S332. Perform graph convolution processing on the word-level vector feature and the character-level vector feature respectively as the node feature matrices of the text graph to obtain the word-level graph convolution feature and the character-level graph convolution feature.
[0030] Preferably, define the edge weight value in the text graph as A ij , and the edge weight value satisfies the following conditions:
[0031]
[0032] Wherein, TF-IDF ij represents the product of the word frequency and the inverse document frequency corresponding to nodes i and j, and PMI(i, j) represents the edge weight value when both nodes i and j are word nodes. PMI(i, j) satisfies the following conditions:
[0033]
[0034] Wherein, p(i, j) represents the probability that nodes i and j appear in the corpus at the same time, p(i) represents the probability that node i appears alone in the corpus, and p(j) represents the probability of the data set of node j appearing alone in the corpus.
[0035] Preferably, in step S34, the following sub-steps are included:
[0036] S341. Linearly transform the word-level graph convolution feature into a query vector, and transform the character-level graph convolution feature into a key vector and a value vector;
[0037] S342. Calculate the attention weights according to the cross-attention mechanism for the query vector, the key vector, and the value vector, and fuse the context information of the word-level graph convolution feature and the character-level graph convolution feature according to the attention weights to obtain the fusion feature;
[0038] S343. Optimize the fused feature through residual connection and normalization to obtain an optimized fused feature;
[0039] S344. Calculate the confidence of each word in the text data according to the optimized fused feature;
[0040] S345. Determine whether the confidence is higher than a preset threshold: if so, mark the corresponding word in the text data as a sensitive word and perform filtering processing.
[0041] In a second aspect, the present invention further provides a travel voice privacy protection device, including:
[0042] A voice conversion module, configured to obtain voice data and perform voice-to-text processing on the voice data to convert it into text data;
[0043] A model establishment module, configured to collect historical travel text data, perform data annotation and data processing on the historical travel text data to obtain a travel text data set; perform model training based on the travel text data set to obtain a privacy protection model;
[0044] A filtering module, configured to use the text data as an input of the privacy protection model to filter sensitive words in the text data to obtain privacy protection text data;
[0045] A text conversion module, configured to perform text-to-voice processing on the privacy protection text data to obtain privacy protection voice data.
[0046] Preferably, the filtering module includes the following sub-units:
[0047] A preprocessing unit, configured to use the text data as an input of the privacy protection model to preprocess the text data to obtain a word segmentation result and a character segmentation result;
[0048] A vector extraction unit, configured to perform BERT embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain a word-level vector feature and a character-level vector feature;
[0049] A text graph construction unit, constructs a text graph, and performs graph convolution processing on the word-level vector feature and the character-level vector feature according to the text graph to obtain a word-level graph convolution feature and a character-level graph convolution feature;
[0050] A sensitive word filtering unit, which is used to fuse the word-level graph convolution feature and the character-level graph convolution feature based on the cross-attention mechanism to obtain a fused feature; perform residual connection and normalization optimization processing on the fused feature, and filter out sensitive words in the text data based on the optimized fused feature to obtain the privacy protection text data.
[0051] In a third aspect, the present invention further provides a travel voice privacy protection device, including: a memory, a processor, and a travel voice privacy protection program stored on the memory and executable on the processor. When the processor executes the travel voice privacy protection program, the steps in any one of the above-mentioned embodiments of the travel voice privacy protection method are implemented.
[0052] In a fourth aspect, the present invention further provides a storage medium, on which a travel voice privacy protection program is stored. When the travel voice privacy protection program is executed by a processor, the steps in any one of the above-mentioned embodiments of the travel voice privacy protection method are implemented.
[0053] Compared with the prior art, the present invention integrates character and word-level multi-granularity feature representations, captures spelling variants and context semantics of sensitive words through BERT embedding and Transformer neural network encoding, and improves the recognition accuracy in complex contexts; the present invention constructs a text graph structure based on PMI and TF-IDF, combines graph convolutional networks to learn semantic associations between words, and enhances the ability to understand long-distance dependencies and semantic roles; the present invention realizes deep interaction between word-level and character-level features by adopting the cross-attention mechanism, accurately locates sensitive words through residual optimization and confidence threshold decision-making, and significantly improves the detection efficiency and recall rate; compared with traditional rule matching or single-granularity models, the present invention effectively solves the problem of detection lag in complex scenarios such as emerging hot meme variants and cross-lingual expressions through multi-dimensional feature fusion and graph structure modeling, and is particularly suitable for the real-time privacy protection requirements of travel scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The present invention will be described in detail below with reference to the drawings. Through the detailed description in combination with the following drawings, the above or other aspects of the present invention will become clearer and easier to understand. In the drawings:
[0055] Figure 1 is a flowchart of the travel voice privacy protection method provided in Embodiment 1 of the present invention;
[0056] Figure 2 is a functional schematic diagram of the privacy protection model of the travel voice privacy protection method provided in Embodiment 1 of the present invention;
[0057] Figure 3It is a flowchart of the privacy protection model of the travel voice privacy protection method provided in the first embodiment of the present invention;
[0058] Figure 4 It is a schematic diagram of the BERT embedding process of the travel voice privacy protection method provided in the first embodiment of the present invention;
[0059] Figure 5 It is a schematic diagram of the text graph structure of the travel voice privacy protection method provided in the first embodiment of the present invention;
[0060] Figure 6 It is a schematic diagram of the structure of the travel voice privacy protection device provided in the second embodiment of the present invention;
[0061] Figure 7 It is a schematic diagram of the structure of the travel voice privacy protection device provided in the third embodiment of the present invention. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] Embodiment 1
[0064] Please refer to Figures 1 - 5 , the present invention provides a travel voice privacy protection method, and the travel voice privacy protection method includes the following steps:
[0065] S1. Obtain voice data, and perform voice-to-text processing on the voice data to convert it into text data.
[0066] In the embodiment of the present invention, the language data is the travel voice data for communication between passengers and drivers during the travel in a online car-hailing or taxi, and the voice data is recognized by the voice recognition system Whisper.
[0067] S2. Collect historical travel text data, perform data annotation and data processing on the historical travel text data to obtain a travel text data set; perform model training based on the travel text data set to obtain a privacy protection model.
[0068] In the embodiment of the present invention, the historical travel text data is a text database composed of a large amount of text data obtained through step S1, and is used to train a privacy protection model that can accurately locate sensitive words according to the sensitive information in the text.
[0069] S3. Use the text data as the input of the privacy protection model, filter sensitive words in the text data to obtain privacy protection text data;
[0070] In an embodiment of the present invention, in step S3, the following sub-steps are included:
[0071] S31. Use the text data as the input of the privacy protection model, preprocess the text data to obtain a word segmentation result and a character segmentation result;
[0072] S32. Perform BERT Embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain word-level vector features and character-level vector features;
[0073] In an embodiment of the present invention, in step S32, the following sub-steps are included:
[0074] S321. Respectively perform vector transformation on the word segmentation result and the character segmentation result, and add a classification marker and a sentence separator at the beginning and end of the sentence of the word segmentation result and at the beginning and end of the sentence of the character segmentation result to obtain a word segmentation vector and a character segmentation vector; wherein, the classification marker is used to aggregate global semantics, and the sentence separator is used to define the text boundary;
[0075] S322. Respectively perform word embedding processing, segment embedding processing, and position embedding processing on the word segmentation vector and the character segmentation vector to obtain a word-level embedding result and a character-level embedding result;
[0076] S323. Perform encoding processing on the word-level embedding result and the character-level embedding result through the Transformer neural network to obtain the word-level vector features and the character-level vector features.
[0077] Specifically, please refer to Figure 4 , Figure 4 is a schematic diagram of the BERT embedding process of the travel voice privacy protection method provided in the first embodiment of the present invention. The BERT embedding processing includes TokenEmbedding for converting words in the text into fixed-dimensional word segment embeddings, PositionEmbedding for enabling the model to understand the word order by learning vectors at different positions, and Segment Embedding for distinguishing the belonging segments of different sentences.
[0078] Taking the character segmentation result as an example, the specific processing flow is as follows: First, use the Tokenizer model (token generator model) to perform vectorization conversion on the input text, add a classification marker (CLS) for aggregating global semantics and a sentence separator (SEP) for defining the text boundary at the beginning and end of the sentence, and then enter the multi-embedding processing.
[0079] Map the tokens to fixed-dimensional vectors through word embedding processing, and the output is expressed as Among them, E t represents the token embedding matrix, is the token vector of the i-th character (t represents token), l is the length of the text sequence, and d e is the embedding layer dimension, represents an l-row d-column matrix in the real number field e matrix.
[0080] Distinguish sentence or paragraph relationships between 0 and 1 through segmentation embedding, and output Among them, E s represents the segmentation embedding matrix, is the segmentation vector of the i-th character (s represents segment), and the meanings of the other symbols are the same as those in the token embedding layer.
[0081] Position embedding uses sine and cosine functions for position encoding, and the formula is where PE represents the position encoding matrix, pos is the absolute position of the character in the sequence (starting from 0), k is the dimension index (starting from 0), and d e is the embedding layer dimension, 2k and 2k + 1 correspond to even and odd dimensions respectively. This formula realizes the sine wave encoding of position information, and the output is expressed as Among them, E p represents the position embedding matrix, is the position vector of the i-th character (pos represents position), and the meanings of the other symbols are the same as before.
[0082] Add the results obtained after the three embeddings to get the sub-character result representation E c , and similarly obtain the word segmentation result E w , and then input them into the Transformer neural network encoding layer respectively. As the core module, this encoding layer first linearly maps the input vector into query Q, key K, and value V matrices, and through the multi-head attention mechanism MultiHead(Q, K, V) = Concat(h1, h2,..., h h )W O , where Q, K, and V are the query, key, and value matrices respectively, h i is the output of the i-th attention head, h is the number of attention heads, and W O is the output projection matrix, Concat represents the concatenation operation, and the calculation of each attention head is where are the query, key, and value mapping matrices of the i-th attention head respectively.
[0083] After fusing the input X and the attention output through residual connection, perform layer normalization Norm1 = LayerNorm(X + MultiHead(X, X, X)), where X is the input of the current layer, and LayerNorm represents the layer normalization operation. Then input it into the feed-forward neural network FFN(x) = max(0, xW1 + b1)W2 + b2, where x is the input vector, W1 and b1 are the weights and biases of the first fully connected layer, and W2 and b2 are the weights and biases of the second fully connected layer, and max(0, ·) is the ReLU activation function. After that, pass through the residual connection and layer normalization again Norm2 = LayerNorm(Norm1 + FFN(Norm1)) to output the high-level features that fuse semantic, structural, and positional information, and finally obtain the character-level vector feature T c Similarly, obtain the word-level vector feature T w Provide a representation for the follow-up
[0084] S33. Construct a text graph, and perform graph convolution processing on the word-level vector feature and the character-level vector feature according to the text graph to obtain the word-level graph convolution feature and the character-level graph convolution feature;
[0085] In the embodiment of the present invention, in step S33, the following sub-steps are included:
[0086] S331. Construct the text graph, the text graph includes document nodes and word nodes, the edges between the document nodes and the word nodes are constructed according to the number of times the corresponding words appear in the document, and the edges between two word nodes are constructed according to the probability that the corresponding words appear in the corpus; wherein, the corpus is composed of the data sets of all the documents corresponding to the document nodes.
[0087] Specifically, please refer to Figure 5 , Figure 5 which is a schematic diagram of the text graph structure of the travel voice privacy protection method provided in Embodiment 1 of the present invention. Let the text graph be G, the text graph G=(V, E), V={v1, v2,..., v n} is a set of n nodes, E is the set of edges between nodes, and its corresponding adjacency matrix is denoted as A∈{0, 1} n×n , and the node feature matrix is denoted as
[0088] S332. Perform graph convolution processing on the word-level vector feature and the character-level vector feature respectively as the node feature matrix of the text graph to obtain the word-level graph convolution feature and the character-level graph convolution feature.
[0089] Define the edge weight value in the text graph as A ij The edge weight value satisfies the following conditions:
[0090]
[0091] Among them, TF-IDF ij represents the product of the term frequency and the inverse document frequency corresponding to nodes i and j, and PMI(i, j) represents the edge weight value when both nodes i and j are the word nodes. PMI(i, j) satisfies the following conditions:
[0092]
[0093] Among them, p(i, j) represents the probability that nodes i and j appear in the corpus at the same time, p(i) represents the probability that node i appears alone in the corpus, and p(j) represents the data set probability that node j appears alone in the corpus.
[0094] TF-IDF is the result of multiplying the term frequency (TF, Term Frequency) and the inverse document frequency (IDF, Inverse Document Frequency). The calculation formula of the TF-IDF value is as follows:
[0095] TF-IDF(d, w) = tf(d, w) × idf(w);
[0096]
[0097] Among them, tf(d, w) represents the number of times the word w appears in the text d, that is, the term frequency TF. In idf(w), N represents the total number of texts, N(w) represents the number of texts in which the word w appears, and the larger the TF-IDF value of a certain word in the document, the higher the importance of this word in this document.
[0098] After constructing the text graph, graph convolution operation needs to be performed. Among them, the graph convolutional network model (Graph Convolutional Network, GCN) is defined as:
[0099]
[0100] Among them, H (l) and W (l) are respectively the node representation and the mapping weight matrix of the l-th layer, D ii = ∑ j A ij is the degree matrix, where I n is the identity matrix with a diagonal of 1, and the initial node representation matrix H (0) is X. Taking the word-level vector feature and the character-level vector feature as the feature matrix X of the graph node respectively, the character-level vector feature representation T obtained previously cAnd the word-level vector feature representation T w Obtain the output H (l) Is the graph convolution output, and finally obtain the word-level graph convolution feature H w And the character-level graph convolution feature H c .
[0101] S34. Based on the cross-attention mechanism, fuse the word-level graph convolution feature and the character-level graph convolution feature to obtain a fused feature; perform residual connection and normalization optimization processing on the fused feature, and filter out sensitive words in the text data based on the optimized fused feature to obtain the privacy-protected text data.
[0102] In the embodiment of the present invention, in step S34, the following sub-steps are included:
[0103] S341. Linearly transform the word-level graph convolution feature into a query vector, and transform the character-level graph convolution feature into a key vector and a value vector;
[0104] S342. Calculate the attention weights according to the query vector, the key vector, and the value vector according to the cross-attention mechanism, and fuse the context information of the word-level graph convolution feature and the character-level graph convolution feature according to the attention weights to obtain the fused feature;
[0105] S343. Perform residual connection and normalization optimization processing on the fused feature to obtain an optimized fused feature;
[0106] S344. Calculate the confidence of each word in the text data according to the optimized fused feature;
[0107] S345. Determine whether the confidence is higher than a preset threshold: if so, mark the corresponding word in the text data as a sensitive word and perform filtering processing.
[0108] Specifically, in the cross-attention fusion of the word-level graph convolution feature H w And the character-level graph convolution feature H c , first linearly transform the word-level graph convolution feature Into a query (W q Is a learnable weight matrix), and the character-level graph convolution feature Are respectively transformed into a key And a value (W k , W v Are learnable parameters).
[0109] The attention weights are calculated by the formula Where q iis the query vector, k j is the key vector. The Score function is a dot product operation Based on the weight α ij , the formula for the word vector to fuse character information is (v j is the value vector), the word-level graph convolutional feature is used as the query vector (Q), the character-level graph convolutional feature is used as the key vector (K) and the value vector (V). By calculating the attention weight, the word-level graph convolutional feature adaptively fuses the context information in the character-level graph convolutional feature. The fused feature can be obtained through the residual connection x res = FFN(x) + x layer normalization (γ, β are learnable, μ, σ 2 are the mean and variance) for optimization. Finally, the confidence is calculated through P = softmax(H final w + b). Set a preset threshold (such as 0.5). If P i > threshold, then mark the corresponding word as a sensitive word. Finally, remove the sensitive word from the text and output the filtered result. For example, process "sensitive word" as "*".
[0110] S4. Perform text-to-speech processing on the privacy-protected text data to obtain privacy-protected speech data.
[0111] In the embodiment of the present invention, the text-to-speech processing uses the third-party library ppytsx in the Python software. Through the above-mentioned hierarchical processing method, the privacy information in the travel speech can be protected efficiently and accurately, effectively preventing the leakage of sensitive content, providing a reliable guarantee for the privacy security of users, and is especially suitable for the processing of travel scenario data with high privacy requirements, such as intelligent travel assistants, in-vehicle voice systems, etc., meeting the real-time privacy protection requirements in the travel scenario.
[0112] Compared with the prior art, the present invention integrates character and word-level multi-granularity feature representations, captures the spelling variants and context semantics of sensitive words through BERT embedding and Transformer neural network encoding, and improves the recognition accuracy in complex contexts; the present invention constructs a text graph structure based on PMI and TF-IDF, combines graph convolutional networks to learn semantic associations between words, and enhances the understanding ability of long-distance dependencies and semantic roles; the present invention realizes the deep interaction between word-level and character-level features through the cross-attention mechanism, accurately locates sensitive words through residual optimization and confidence threshold decision-making, and significantly improves the detection efficiency and recall rate; compared with traditional rule matching or single-granularity models, the present invention effectively solves the problem of detection lag in complex scenarios such as emerging hot meme variants and cross-lingual expressions through multi-dimensional feature fusion and graph structure modeling, and is especially suitable for the real-time privacy protection requirements of travel scenarios.
[0113] Embodiment 2
[0114] An embodiment of the present invention further provides a travel voice privacy protection device. Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the travel voice privacy protection device 200 provided in the second embodiment of the present invention, and it includes:
[0115] 201. A voice conversion module, configured to obtain voice data and perform voice-to-text processing on the voice data to convert it into text data;
[0116] 202. A model establishment module, configured to collect historical travel text data, perform data annotation and data processing on the historical travel text data to obtain a travel text data set; and perform model training based on the travel text data set to obtain a privacy protection model;
[0117] 203. A filtering module, configured to use the text data as an input to the privacy protection model, filter sensitive words in the text data to obtain privacy protection text data;
[0118] 204. A text conversion module, configured to perform text-to-voice processing on the privacy protection text data to obtain privacy protection voice data.
[0119] In the embodiment of the present invention, the filtering module includes the following sub-units:
[0120] 2031. A preprocessing unit, configured to use the text data as an input to the privacy protection model, perform preprocessing on the text data to obtain a word segmentation result and a character segmentation result;
[0121] 2032. A vector extraction unit, configured to perform BERT embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain word-level vector features and character-level vector features;
[0122] 2033. A text graph construction unit, configured to construct a text graph, and perform graph convolution processing on the word-level vector features and the character-level vector features according to the text graph to obtain word-level graph convolution features and character-level graph convolution features;
[0123] 2034. A sensitive word filtering unit, configured to perform fusion processing on the word-level graph convolution features and the character-level graph convolution features based on a cross-attention mechanism to obtain a fusion feature; perform residual connection and normalization optimization processing on the fusion feature, and filter out sensitive words in the text data based on the optimized fusion feature to obtain the privacy protection text data.
[0124] The travel voice privacy protection device 200 can implement the steps in the travel voice privacy protection method in the above embodiments and can achieve the same technical effects. Refer to the description in the above embodiments, and details are not described herein again.
[0125] Embodiment III
[0126] The embodiment of the present invention further provides a travel voice privacy protection device. Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of the travel voice privacy protection device provided in Embodiment III of the present invention. The travel voice privacy protection device 300 includes: a memory 302, a processor 301, and a travel voice privacy protection program stored on the memory 302 and executable on the processor 301.
[0127] The processor 301 calls the travel voice privacy protection program stored in the memory 302 and executes the steps in the travel voice privacy protection method provided in the embodiment of the present invention. Please refer to Figures 1 - 5 , which specifically includes the following steps:
[0128] S1. Obtain voice data, and perform voice-to-text processing on the voice data to convert it into text data.
[0129] In the embodiment of the present invention, the language data is the travel voice data of the conversation between passengers and drivers during the travel in online car-hailing or taxis. The voice data is recognized by the voice recognition system Whisper.
[0130] S2. Collect historical travel text data, perform data annotation and data processing on the historical travel text data to obtain a travel text data set; perform model training based on the travel text data set to obtain a privacy protection model.
[0131] In the embodiment of the present invention, the historical travel text data is composed of a large amount of text data obtained through step S1, and is used to train a privacy protection model that can accurately locate sensitive words according to the sensitive information in the text.
[0132] S3. Use the text data as the input of the privacy protection model, filter out sensitive words in the text data to obtain privacy protection text data;
[0133] In the embodiment of the present invention, in step S3, the following sub-steps are included:
[0134] S31. Use the text data as the input of the privacy protection model, perform preprocessing on the text data to obtain a word segmentation result and a character segmentation result;
[0135] S32. Perform BERT Embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain word-level vector features and character-level vector features;
[0136] In the embodiment of the present invention, in step S32, the following sub-steps are included:
[0137] S321. Respectively perform vector transformation on the word segmentation result and the character segmentation result, and add classification markers and sentence delimiters at the beginning and end of the sentences of the word segmentation result and the character segmentation result respectively to obtain a word segmentation vector and a character segmentation vector; wherein, the classification marker is used to aggregate global semantics, and the sentence delimiter is used to define the text boundary;
[0138] S322. Respectively perform word embedding processing, segmentation embedding processing, and position embedding processing on the word segmentation vector and the character segmentation vector to obtain a word-level embedding result and a character-level embedding result;
[0139] S323. The word-level embedding result and the character-level embedding result are encoded through the Transformer neural network to obtain the word-level vector features and the character-level vector features.
[0140] Specifically, please refer to Figure 4 , Figure 4 is a schematic diagram of the BERT embedding process of the travel voice privacy protection method provided in Embodiment 1 of the present invention. The BERT embedding processing includes TokenEmbedding for converting words in the text into fixed-dimensional word segment embeddings, PositionEmbedding for enabling the model to understand the word order by learning vectors at different positions, and Segment Embedding for distinguishing the belonging segments of different sentences.
[0141] Taking the character segmentation result as an example, the specific processing flow is as follows: First, use the Tokenizer model (token generator model) to perform vectorization conversion on the input text, add a classification marker (CLS) for aggregating global semantics and a sentence delimiter (SEP) for defining the text boundary at the beginning and end of the sentence, and then enter the multi-embedding processing.
[0142] Map the tokens to fixed-dimensional vectors through word embedding processing, and the output is expressed as where E t represents the token embedding matrix, is the token vector of the i-th character (t represents token), l is the length of the text sequence, d e is the dimension of the embedding layer, represents an l-row d e column matrix over the real number field.
[0143] Distinguish sentence or paragraph relationships between 0 and 1 through split embedding, and output where E s represents the split embedding matrix, is the split vector of the i-th character (s represents segment), and the meanings of the remaining symbols are the same as those in the token embedding layer.
[0144] Position embedding uses sine and cosine functions for position encoding, and the formula is where PE represents the position encoding matrix, pos is the absolute position of the character in the sequence (starting from 0), k is the dimension index (starting from 0), and d e is the embedding layer dimension. 2k and 2k + 1 correspond to even and odd dimensions respectively. This formula realizes the sine wave encoding of position information, and the output is expressed as where E p represents the position embedding matrix, is the position vector of the i-th character (pos represents position), and the meanings of the remaining symbols are the same as before.
[0145] Add the results obtained from the three embedding processes to get the sub-character result representation E c Similarly, obtain the word segmentation result E w and then input them into the Transformer neural network encoding layer respectively. As the core module, this encoding layer first linearly maps the input vector into query Q, key K, and value V matrices, and through the multi-head attention mechanism MultiHead(Q, K, V) = Concat(h1, h2,..., h h )W O where Q, K, and V are the query, key, and value matrices respectively, h i is the output of the i-th attention head, h is the number of attention heads, and W O is the output projection matrix. Concat represents the concatenation operation, and the calculation of each attention head is where are the query, key, and value mapping matrices of the i-th attention head respectively.
[0146] After fusing the input X and the attention output through a residual connection, layer normalization Norm1 = LayerNorm(X + MultiHead(X, X, X)) is performed, where X is the input of the current layer and LayerNorm represents the layer normalization operation. Then, it is input into the feed-forward neural network FFN(x) = max(0, xW1 + b1)W2 + b2, where x is the input vector, W1 and b1 are the weights and biases of the first fully connected layer, and W2 and b2 are the weights and biases of the second fully connected layer, and max(0, ·) is the ReLU activation function. After that, it passes through the residual connection and layer normalization Norm2 = LayerNorm(Norm1 + FFN(Norm1)) again to output high-level features that fuse semantic, structural, and positional information, and finally obtain the character-level vector feature T c Similarly, the word-level vector feature T is obtained w to provide a representation for the follow-up
[0147] S33. Construct a text graph, and perform graph convolution processing on the word-level vector feature and the character-level vector feature according to the text graph to obtain the word-level graph convolution feature and the character-level graph convolution feature
[0148] In the embodiment of the present invention, in step S33, the following sub-steps are included
[0149] S331. Construct the text graph, where the text graph includes document nodes and word nodes. The edges between the document nodes and the word nodes are constructed according to the number of times the corresponding words appear in the document, and the edges between the two word nodes are constructed according to the probability of the corresponding words appearing in the corpus; wherein, the corpus is composed of the data sets of the documents corresponding to all the document nodes
[0150] Specifically, please refer to Figure 5 , Figure 5 which is a schematic diagram of the text graph structure of the travel voice privacy protection method provided in Embodiment 1 of the present invention. Let the text graph be G, and the text graph G = (V, E), where V = {v1, v2,..., v n} is a set of n nodes, and E is the set of edges between the nodes, and its corresponding adjacency matrix is denoted as A ∈ {0, 1} n×n , and the node feature matrix is denoted as
[0151] S332. Perform graph convolution processing on the word-level vector feature and the character-level vector feature respectively as the node feature matrix of the text graph to obtain the word-level graph convolution feature and the character-level graph convolution feature
[0152] Define the edge weight value in the text graph as A ij , and the edge weight value satisfies the following conditions
[0153]
[0154] Among them, TF-IDF ij represents the product of the term frequency and the inverse document frequency corresponding to nodes i and j, and PMI(i, j) represents the edge weight value when both nodes i and j are the word nodes. PMI(i, j) satisfies the following conditions:
[0155]
[0156] Among them, p(i, j) represents the probability that nodes i and j appear in the corpus at the same time, p(i) represents the probability that node i appears alone in the corpus, and p(j) represents the probability of the data set that node j appears alone in the corpus.
[0157] TF-IDF is the result of multiplying the term frequency (TF, Term Frequency) and the inverse document frequency (IDF, Inverse Document Frequency). The calculation formula of the TF-IDF value is as follows:
[0158] TF-IDF(d, w) = tf(d, w) × idf(w);
[0159]
[0160] Among them, tf(d, w) represents the number of times the word w appears in the text d, that is, the term frequency TF. In idf(w), N represents the total number of texts, N(w) represents the number of texts in which the word w appears, and the larger the TF-IDF value of a certain word in the document, the higher the importance of this word in this document.
[0161] After constructing the text graph, graph convolution operation needs to be performed. Among them, the graph convolutional network model (Graph Convolutional Network, GCN) is defined as:
[0162]
[0163] Among them, H (l) and W (l) are respectively the node representation and the mapping weight matrix of the l-th layer, D ii = ∑ j A ij is the degree matrix, Among them, I n is the identity matrix with a diagonal of 1, and the initial node representation matrix H (0) is X. Taking the word-level vector feature and the character-level vector feature as the feature matrix X of the graph node respectively, the character-level vector feature representation T obtained previously cAnd the word-level vector feature representation T w Obtain the output H (l) As the graph convolution output, finally obtain the word-level graph convolution feature H w And the character-level graph convolution feature H c .
[0164] S34. Based on the cross-attention mechanism, fuse the word-level graph convolution feature and the character-level graph convolution feature to obtain a fused feature; perform residual connection and normalization optimization processing on the fused feature, and filter out sensitive words in the text data based on the optimized fused feature to obtain the privacy-protected text data.
[0165] In the embodiment of the present invention, in step S34, the following sub-steps are included:
[0166] S341. Linearly transform the word-level graph convolution feature into a query vector, and transform the character-level graph convolution feature into a key vector and a value vector;
[0167] S342. Calculate the attention weights according to the cross-attention mechanism for the query vector, the key vector, and the value vector, and fuse the context information of the word-level graph convolution feature and the character-level graph convolution feature according to the attention weights to obtain the fused feature;
[0168] S343. Perform residual connection and normalization optimization processing on the fused feature to obtain an optimized fused feature;
[0169] S344. Calculate the confidence of each word in the text data according to the optimized fused feature;
[0170] S345. Determine whether the confidence is higher than a preset threshold: if so, mark the corresponding word in the text data as a sensitive word and perform filtering processing.
[0171] Specifically, in the cross-attention fusion of the word-level graph convolution feature H w and the character-level graph convolution feature H c , first linearly transform the word-level graph convolution feature into a query (W q is a learnable weight matrix), and transform the character-level graph convolution feature into a key and a value (W k , W v are learnable parameters).
[0172] The attention weights are calculated by the formula where q iis the query vector, k j is the key vector. The Score function is a dot product operation Based on the weight α ij , the formula for the word vector to fuse character information is (v j is the value vector), the word-level graph convolution feature is used as the query vector (Q), the character-level graph convolution feature is used as the key vector (K) and the value vector (V). By calculating the attention weight, the word-level graph convolution feature adaptively fuses the context information in the character-level graph convolution feature. The fused feature can be obtained through the residual connection x res = FFN(x) + x layer normalization (γ, β are learnable, μ, σ 2 are the mean and variance) for optimization. Finally, the confidence is calculated through P = softmax(H final w + b). Set a preset threshold (such as 0.5). If P i > threshold, then mark the corresponding word as a sensitive word. Finally, remove the sensitive word from the text and output the filtered result. For example, process "sensitive word" as "*".
[0173] S4. Perform text-to-speech processing on the privacy-protected text data to obtain privacy-protected speech data.
[0174] In the embodiment of the present invention, the text-to-speech processing uses the third-party library ppytsx in the Python software. Through the above-described hierarchical processing method, the privacy information in the travel speech can be protected efficiently and accurately, effectively preventing the leakage of sensitive content, providing a reliable guarantee for the user's privacy security, and is particularly suitable for processing travel scenario data with high privacy requirements, such as intelligent travel assistants, in-vehicle voice systems, etc., meeting the real-time privacy protection requirements in the travel scenario.
[0175] The travel speech privacy protection device 300 provided by the embodiment of the present invention can implement the steps in the travel speech privacy protection method in the above embodiment, and can achieve the same technical effects. Refer to the description in the above embodiment, and details are not described here again.
[0176] Embodiment Four
[0177] The embodiment of the present invention further provides a storage medium. The storage medium is a computer-readable storage medium. A travel speech privacy protection program is stored on the storage medium. When the travel speech privacy protection program is executed by a processor, it implements each process and step in the travel speech privacy protection method provided by the embodiment of the present invention, and can achieve the same technical effects. To avoid repetition, details are not described here again.
[0178] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0179] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element.
[0180] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0181] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. What is disclosed is only the preferred embodiments of the present invention. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can make many equivalent changes in form without departing from the purpose of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.
Claims
1. A method for protecting travel voice privacy, characterized in that, The described travel voice privacy protection method includes the following steps: S1. Obtain voice data, and perform voice-to-text processing on the voice data to convert it into text data; S2. Collect historical travel text data, perform data annotation and data processing on the historical travel text data to obtain a travel text data set; based on the travel text data set, perform model training to obtain a privacy protection model; S3. Use the text data as the input of the privacy protection model, filter out sensitive words in the text data to obtain privacy protection text data; S4. Perform text-to-voice processing on the privacy protection text data to obtain privacy protection voice data.
2. The itinerary voice privacy protection method according to claim 1, wherein In step S3, the following sub-steps are included: S31. Use the text data as the input of the privacy protection model, perform preprocessing on the text data to obtain a word segmentation result and a character segmentation result; S32. Perform BERT embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain word-level vector features and character-level vector features; S33. Construct a text graph, perform graph convolution processing on the word-level vector features and the character-level vector features according to the text graph to obtain word-level graph convolution features and character-level graph convolution features; S34. Based on the cross-attention mechanism, fuse the word-level graph convolution features and the character-level graph convolution features to obtain fused features; perform residual connection and normalization optimization processing on the fused features, and filter out sensitive words in the text data based on the optimized fused features to obtain the privacy protection text data.
3. The travel voice privacy protection method according to claim 2, wherein In step S32, the following sub-steps are included: S321. Respectively perform vector conversion on the word segmentation result and the character segmentation result, and add classification markers and sentence delimiters at the beginning and end of the sentences of the word segmentation result and the character segmentation result respectively to obtain word segmentation vectors and character segmentation vectors; wherein, the classification marker is used to aggregate global semantics, and the sentence separator is used to define the text boundary; S322. Respectively perform word embedding processing, segmentation embedding processing, and position embedding processing on the word segmentation vector and the character segmentation vector to obtain word-level embedding results and character-level embedding results; S323. The word-level embedding result and the character-level embedding result are encoded through the Transformer neural network to obtain the word-level vector features and the character-level vector features.
4. The travel voice privacy protection method according to claim 2, wherein In step S33, the following sub-steps are included: S331. Construct the text graph, the text graph includes document nodes and word nodes, the edges between the document nodes and the word nodes are constructed according to the number of times the corresponding words appear in the document, and the edges between two word nodes are constructed according to the probability of the corresponding words appearing in the corpus; wherein, the corpus is composed of the data sets of the documents corresponding to all the document nodes; S332. Respectively use the word-level vector features and the character-level vector features as the node feature matrices of the text graph for graph convolution processing to obtain the word-level graph convolution features and the character-level graph convolution features.
5. The travel voice privacy protection method according to claim 4, wherein Define the edge weight value in the text graph as A ij , and the edge weight value satisfies the following conditions: Among them, TF-IDF ij represents the product of the word frequencies corresponding to node i and node j and the inverse document frequency. PMI(i, j) represents the edge weight value when both node i and node j are the word nodes. PMI(i, j) satisfies the following conditions: Among them, p(i, j) represents the probability that nodes i and j appear in the corpus simultaneously, p(i) represents the probability that node i appears alone in the corpus, and p(j) represents the probability that node j appears alone in the dataset of the corpus.
6. The trip voice privacy protection method according to claim 2, characterized in that, In step S34, the following sub-steps are included: S341. Linearly transform the word-level graph convolution features into query vectors, and transform the character-level graph convolution features into key vectors and value vectors; S342. Calculate attention weights based on the query vectors, the key vectors, and the value vectors according to the cross-attention mechanism, and fuse the context information of the word-level graph convolution features and the character-level graph convolution features according to the attention weights to obtain the fused features; S343. Optimize the fused features through residual connection and normalization to obtain optimized fused features; S344. Calculate the confidence of each word in the text data based on the optimized fused features; S345. Determine whether the confidence is higher than a preset threshold: if so, mark the corresponding word in the text data as a sensitive word and perform filtering processing.
7. A travel voice privacy protection device, characterized in that, Including: A speech conversion module, configured to obtain speech data and perform speech-to-text processing on the speech data to convert it into text data; A model establishment module, configured to collect historical trip text data, perform data annotation and data processing on the historical trip text data to obtain a trip text dataset; perform model training based on the trip text dataset to obtain a privacy protection model; A filtering module, configured to use the text data as an input to the privacy protection model to filter sensitive words in the text data to obtain privacy protection text data; A text conversion module, configured to perform text-to-speech processing on the privacy protection text data to obtain privacy protection speech data.
8. The travel voice privacy protection device according to claim 7, wherein, The filtering module includes the following sub-units: A preprocessing unit, configured to use the text data as an input to the privacy protection model and perform preprocessing on the text data to obtain a word segmentation result and a character segmentation result; A vector extraction unit, configured to perform BERT embedding processing and Transformer neural network encoding processing on the word segmentation result and the character segmentation result to obtain word-level vector features and character-level vector features; A text graph construction unit, configured to construct a text graph, and perform graph convolution processing on the word-level vector features and the character-level vector features according to the text graph to obtain word-level graph convolution features and character-level graph convolution features; A sensitive word filtering unit, configured to fuse the word-level graph convolution features and the character-level graph convolution features based on the cross-attention mechanism to obtain fused features; perform residual connection and normalization optimization processing on the fused features, and filter out sensitive words in the text data based on the optimized fused features to obtain the privacy protection text data.
9. A travel voice privacy protection device, characterized in that, Including: A memory, a processor, and a trip voice privacy protection program stored on the memory and executable on the processor, where when the processor executes the trip voice privacy protection program, the steps in the trip voice privacy protection method described in any one of claims 1-6 are implemented.
10. A storage medium, characterized in that, A privacy protection program is stored on the storage medium, and when the travel voice privacy protection program is executed by a processor, the steps in the travel voice privacy protection method described in any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Robot auditory privacy information monitoring processing method
CN111597580A
Text classification method and system, computer equipment and storage medium
CN112529071A
Sensitive text detection method for social platform
CN116561318A
Unstructured data processing method, apparatus, and device, and medium
WO2021212968A1