Text recognition method based on language model
Through the text recognition method based on language model, asymmetric word segmentation algorithm and bidirectional attention mechanism are used to generate multidimensional semantic vector sequences, combined with dynamic weight allocation and multi-level condition combination analysis, the problems of insufficient understanding of complex semantics and lack of effective protection mechanisms in the existing technology are solved, efficient risk detection and multi-layer protection are achieved, and the security of the application is guaranteed.
Patent Information
- Application Number
- CN202510495466.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing text recognition technologies lack the ability to understand complex semantics, it is difficult to accurately identify potential risks, and lack dynamic weight adjustment and multi-level conditional analysis mechanisms, resulting in poor accuracy and flexibility in risk determination. After risk determination, there is a lack of an effective multi-layer protection mechanism, and it is impossible to deal with high-risk risks in a timely manner and ensure the security of the application.
The text recognition method based on language model is adopted to generate a multidimensional semantic vector sequence through asymmetric word segmentation algorithm and bidirectional attention mechanism. The dynamic weight allocation module modulates the features. The logical judgment engine performs multi-level conditional combination analysis to generate risk probability distributions, and generates interpretable risk feature vectors through demodulation processing. Based on the cross-validation of these feature vectors and preset thresholds, the final risk determination result is generated. When the determination result is a high-risk risk, a multi-layer protection mechanism is triggered, including content replacement, session interruption and security alerts.
It effectively improves the risk detection ability of text data, improves the ability to analyze and feature extraction of complex semantics, enhances the accuracy and reliability of risk analysis, and takes timely protective measures when high-risk risks are detected to ensure the safety of the application.
Smart Images

Figure CN120012079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing and information security technology, and more specifically, to a text recognition method based on a language model. Background Art
[0002] In today's digital age, with the rapid development of information technology, text data plays a vital role in various applications. Whether it is user input, text displayed on the interface, or text information in network transmission, it may contain sensitive or risky information. Existing text recognition technology mainly relies on traditional keyword matching and simple semantic analysis. Although these methods can identify risks in text to a certain extent, they have many limitations. For example, the keyword matching method cannot understand the contextual semantics and is prone to misjudgment; and the simple semantic analysis method lacks accuracy and reliability when faced with complex sentence structures and diverse semantic expressions. In addition, the existing text recognition technology lacks an effective protection mechanism after risk determination, and cannot respond to high-risk risks in a timely manner, resulting in security vulnerabilities and potential threats.
[0003] In the process of implementing the embodiments of the present invention, there are at least the following problems or defects in the prior art: the existing text recognition technology is insufficient in its ability to understand complex semantics and has difficulty in accurately identifying potential risks; it lacks dynamic weight adjustment and multi-level condition analysis mechanisms, resulting in poor accuracy and flexibility in risk determination; after risk determination, there is a lack of effective multi-layer protection mechanisms, making it impossible to respond to high-risk risks in a timely manner and ensure the security of applications. Summary of the invention
[0004] The present invention provides a text recognition method based on a language model, comprising: S1 obtains the target application to be detected text data stream, the text data stream includes user input content, interface display text and network transmission text; S2. Preprocessing the text data stream to generate a multidimensional semantic vector sequence; S3. Using a dynamic weight allocation module to perform feature modulation on a multi-dimensional semantic vector sequence to generate a modulation feature matrix; S4. Input the modulation feature matrix into the logic judgment engine, perform multi-level condition combination analysis, and generate risk probability distribution; S5. Demodulate the risk probability distribution to generate an interpretable risk feature vector; S6. Perform cross validation based on the risk feature vector and the preset threshold set to generate the final risk determination result; S7. When the judgment result is a high risk, a multi-layer protection mechanism is triggered, including content replacement, session interruption and security alert.
[0005] Furthermore, the step S2 specifically includes: S21. Using an asymmetric word segmentation algorithm to perform word segmentation semantic analysis on the text data stream, the word segmentation algorithm sets a dynamic window size W, W is adaptively adjusted according to the complexity of the sentence; S22. Perform context-related encoding on the word segmentation result to generate an initial semantic vector with a dimension of D, where D is a positive integer; S23. Use the bidirectional attention mechanism to spatially project the initial semantic vector to form a multi-dimensional semantic vector sequence , is the feature vector of the i-th semantic unit.
[0006] Furthermore, the characteristic modulation in step S3 satisfies: The calculation formula of the modulation characteristic matrix H is:
[0007] in, is the dynamic weight coefficient matrix, is the feature enhancement factor, is the space compression coefficient, V is the multi-dimensional semantic vector sequence, is the transposed matrix of V, is the hyperbolic tangent function, is the sigmoid function, is the Hadamard product, n is the total number of semantic units, and D is the dimension of the semantic vector.
[0008] Furthermore, the multi-level condition combination analysis in step S4 includes: S41. Extract the singular value decomposition components of the modulation feature matrix H and calculate the energy distribution entropy E;
[0009] in, is the i-th singular value, S is the sum of singular values, is the first entropy threshold, is the second entropy threshold, and is a positive real number and > ; S42. When E> When , the first-level risk analysis module is activated to detect semantic coherence anomalies; S43. When ≤E≤ When the second-level risk analysis module is activated to detect sensitive pattern matching; S44. When E< When the third-level risk analysis module is activated, the context logic contradiction is detected.
[0010] Furthermore, the demodulation process of step S5 includes: The demodulation function is:
[0011] in, is the demodulation coefficient matrix, is the trainable demodulation weight matrix, C is the risk category mapping matrix, K is the total number of risk categories, F is the interpretable risk feature vector, H is the modulation feature matrix, softmax is the normalization function, and ReLU is the linear rectification function.
[0012] Furthermore, the cross validation in step S6 includes: S61. Calculate the risk feature vector F and the preset typical risk pattern set The cosine similarity set ; in, is the jth typical risk pattern vector, m is the total number of patterns, is the jth cosine similarity, is the similarity threshold, is the confidence threshold, is the composite risk determination threshold, is the risk weight coefficient of the i-th category, and N is the number of consecutive determinations; S62. When there is > and > If the risk level of the deterministic risk reaches the preset high-risk standard, it is determined to be a high-risk risk. S63. When It is determined to be a compound risk when the risk level of the compound risk reaches the preset high-risk standard, and it is determined to be a high-risk risk. S64. When the results of N consecutive determinations meet the risk increasing rule, it is determined to be a high risk and the enhanced protection mechanism is triggered.
[0013] Furthermore, the multi-layer protection mechanism in step S7 includes: S71. The content replacement phase generates semantically preserved replacement text ,in, is the original text, M is the secure corpus; S72. Generate a progressive interruption instruction sequence during the session interruption phase ,in, is the i-th interrupt instruction, t is the total number of terminal steps, It is the final interrupt instruction; S73. Security alarm stage generates multi-dimensional alarm signals ,in, For local logging, Notify the remote server. It is the user interface warning layer.
[0014] Furthermore, the dynamic weight coefficient matrix The calculation method is shown in the following formula:
[0015] Among them, U is the trainable parameter matrix, b is the bias vector, is an improved activation function, and V is a multi-dimensional semantic vector sequence.
[0016] Furthermore, the method for updating the risk category mapping matrix C includes:
[0017] in, is the risk category mapping matrix at the current moment, is the learning rate, is the true risk label vector, To predict the risk feature vector, is the outer product operation, and H is the modulation feature matrix.
[0018] Furthermore, the generation of the semantically-preserving replacement text satisfies:
[0019] in, is the balance coefficient, is the cosine distance, E is the semantic encoding function, is the edit distance, is the scaling factor, R is the current risk level, is the sigmoid function, T is the original text, To replace the text.
[0020] According to the above-mentioned embodiments of the present invention, at least the following beneficial effects are achieved: the text recognition method of the present invention can effectively improve the risk detection capability of text data. By adopting an asymmetric word segmentation algorithm and a bidirectional attention mechanism, it is possible to accurately parse and extract features of complex semantics, generate a multidimensional semantic vector sequence, and thus provide richer semantic information for subsequent risk analysis. At the same time, the dynamic weight allocation module can adaptively adjust the weights according to the importance of the semantic units, further enhancing the expressive power of the features. In addition, the multi-level conditional combination analysis can flexibly activate risk analysis modules of different levels according to the energy distribution entropy of the feature matrix, and respectively detect semantic coherence anomalies, sensitive pattern matching, and contextual logic contradictions, thereby achieving a comprehensive and multi-level analysis of text risks and improving the accuracy and reliability of risk detection.
[0021] The present invention can also take effective measures in a timely manner through a multi-layer protection mechanism when high-risk risks are detected. The content replacement stage can generate semantically preserved replacement text to ensure that the core semantics of the original text is retained while eliminating risks; the session interruption stage can generate a progressive interruption instruction sequence to avoid an overly abrupt experience for users; the security alert stage can promptly notify relevant personnel to take further measures through local logging, remote server notifications, and user interface warnings. This multi-level, multi-dimensional protection mechanism can not only effectively reduce security risks, but also ensure the normal operation of applications and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which: Figure 1 A flowchart of a language model-based text recognition method provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0023] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0024] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, apparatus, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0025] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0026] Reference below Figure 1 , Figure 1 The following is a flow chart of a text recognition method based on a language model provided by an embodiment of the present invention. Figure 1 As shown, a text recognition method 100 based on a language model includes: S1 obtains the target application to be detected text data stream, the text data stream includes user input content, interface display text and network transmission text; S2. Preprocessing the text data stream to generate a multidimensional semantic vector sequence; S3. Using a dynamic weight allocation module to perform feature modulation on a multi-dimensional semantic vector sequence to generate a modulation feature matrix; S4. Input the modulation feature matrix into the logic judgment engine, perform multi-level condition combination analysis, and generate risk probability distribution; S5. Demodulate the risk probability distribution to generate an interpretable risk feature vector; S6. Perform cross validation based on the risk feature vector and the preset threshold set to generate the final risk determination result; S7. When the judgment result is a high risk, a multi-layer protection mechanism is triggered, including content replacement, session interruption and security alert.
[0027] It should be noted that in the text recognition method of the present invention, the first step involves obtaining the text data stream to be detected within the target application. The purpose of this step is to comprehensively cover possible risk sources. The text data stream here refers to various text forms generated during the operation of the application, including user input content, interface display text, and network transmission text. User input content refers to the text information entered by the user through the keyboard or other input devices in the application; interface display text refers to the text information displayed to the user on the application interface, such as button labels, prompt information, etc.; network transmission text refers to the text data transmitted when the application interacts with other devices or servers through the network. The acquisition of these text data streams is the basis for subsequent risk detection, ensuring that the text in the application can be fully monitored.
[0028] Specifically, the acquisition of the text data stream to be detected can be achieved through the interface or middleware of the application. For example, in an instant messaging application, the text input by the user can be captured through the input event of the chat box, the interface display text can be obtained through the output data of the interface rendering engine, and the network transmission text can be captured through the monitoring mechanism of the network communication module. The acquisition method of these data streams can be adapted according to the specific application architecture to ensure the integrity and real-time performance of the data. In addition, the preprocessing step of the text data stream is a key link in generating a multidimensional semantic vector sequence, and its purpose is to convert the original text into a semantic feature form that can be processed by the model. The asymmetric word segmentation algorithm is an important link in the preprocessing. It sets a dynamic window size and adaptively adjusts the granularity of word segmentation according to the complexity of the sentence. For example, for simple sentences, the window size can be small to improve the efficiency of word segmentation; for complex sentences, the window size can be appropriately increased to ensure the integrity of semantics. After the word segmentation result is context-associated encoded, an initial semantic vector with a dimension of D is generated. Context-associated encoding here refers to encoding each word or phrase using context information to capture its semantics in context. For example, a pre-trained language model such as BERT can be used for encoding, and the generated initial semantic vector will contain rich contextual information.
[0029] Preferably, some optimization measures can be introduced in the preprocessing stage of the text data stream. For example, for the adjustment of the dynamic window size, a threshold can be set based on the complexity of the syntactic structure of the text. When a sentence contains multiple clauses or complex modifiers, the dynamic window size can be automatically adjusted to a larger value to ensure that the word segmentation result can fully reflect the semantic structure of the sentence.
[0030] Furthermore, when generating the initial semantic vector, a variety of context encoding techniques can be combined, such as the multi-head attention mechanism in the Transformer architecture, to further improve the quality of the semantic vector. For the generation of a multi-dimensional semantic vector sequence, a bidirectional attention mechanism can be used to spatially project the initial semantic vector to form a multi-dimensional semantic vector sequence, thereby better capturing the contextual association information in the text.
[0031] In some embodiments, step S2 specifically includes: S21. Using an asymmetric word segmentation algorithm to perform word segmentation semantic analysis on the text data stream, the word segmentation algorithm sets a dynamic window size W, W is adaptively adjusted according to the complexity of the sentence; S22. Perform context-related encoding on the word segmentation result to generate an initial semantic vector with a dimension of D, where D is a positive integer; S23. Use the bidirectional attention mechanism to spatially project the initial semantic vector to form a multi-dimensional semantic vector sequence , is the feature vector of the i-th semantic unit.
[0032] It should be noted that the step of preprocessing the text data stream in the present invention includes using an asymmetric word segmentation algorithm to perform word segmentation semantic analysis on the text, and setting a dynamic window size W, which can be adaptively adjusted according to the complexity of the sentence. The asymmetric word segmentation algorithm here is an improved word segmentation method, which takes into account the importance and contextual information of different words in the text. Compared with the traditional symmetric word segmentation method, it can handle complex sentences more flexibly. The word segmentation result will further perform context-related coding to generate an initial semantic vector with a dimension of D. The context-related coding here refers to combining words with their contextual information through an encoding method to generate a vector representation that can reflect semantics. Finally, the initial semantic vector is spatially projected through a bidirectional attention mechanism to form a multidimensional semantic vector sequence. This process can effectively extract semantic features in the text and provide a basis for subsequent risk analysis.
[0033] Specifically, the core of the asymmetric word segmentation algorithm lies in the setting of the dynamic window size W. The window size can be adjusted according to the complexity of the sentence, for example, by analyzing the number of clauses in the sentence, the length of modifiers, or the distribution of keywords to determine the window size. When the sentence structure is relatively simple, the window size can be smaller to improve the efficiency of word segmentation; when the sentence structure is complex, the window size can be appropriately increased to ensure the integrity of the semantics. The dimension D of the initial semantic vector generated by the context-related encoding is a positive integer, which can be set according to the specific application scenario and model requirements. For example, when processing shorter texts, D can be set to a smaller value, such as 128 or 256; when processing longer or more complex texts, D can be set to 512 or higher. The bidirectional attention mechanism spatially projects the initial semantic vector by considering the bidirectional flow of context information, thereby generating a multi-dimensional semantic vector sequence. This process can be implemented using the multi-head attention mechanism in the Transformer architecture, which can effectively capture long-distance dependencies in the text.
[0034] Preferably, in the asymmetric word segmentation algorithm, the adjustment of the dynamic window size W can be further refined. For example, a scoring system based on sentence complexity can be introduced to calculate the complexity score based on factors such as the number of clauses, the length of modifiers, and the distribution of keywords in the sentence, and then the window size can be dynamically adjusted based on the score. In addition, context-dependent encoding can be combined with pre-trained language models (such as BERT or GPT) to generate initial semantic vectors. These models have been trained on large-scale corpora and can provide richer semantic information. For the bidirectional attention mechanism, an improved Transformer architecture can be adopted, such as introducing relative position encoding or adaptive attention weights, to further enhance the model's ability to capture text semantics.
[0035] In some embodiments, the characteristic modulation in step S3 satisfies: The calculation formula of the modulation characteristic matrix H is:
[0036] in, is the dynamic weight coefficient matrix, is the feature enhancement factor, is the space compression coefficient, V is the multi-dimensional semantic vector sequence, is the transposed matrix of V, is the hyperbolic tangent function, is the sigmoid function, is the Hadamard product, n is the total number of semantic units, and D is the dimension of the semantic vector.
[0037] It should be noted that the step of feature modulation of the multidimensional semantic vector sequence in the present invention is implemented by a dynamic weight allocation module, the core of which is to generate a modulation feature matrix. This process optimizes the multidimensional semantic vector sequence through a series of mathematical formulas and parameter operations to enhance the expressiveness and discrimination of the features. Among them, the calculation formula of the modulation feature matrix involves the dynamic weight coefficient matrix ( ), feature enhancement factor ( )、Space Compression Factor( ) and other key parameters, which are calculated through specific mathematical operations (such as the hyperbolic tangent function tanh, the sigmoid function σ, and the Hadamard product ) modulates the semantic vector. This process aims to dynamically adjust weights and enhance features so that the feature matrix can better reflect the risk characteristics of text data and provide more accurate input for subsequent risk analysis.
[0038] Specifically, the dynamic weight coefficient matrix ( ) is one of the core parameters of the modulation feature matrix, and its role is to dynamically assign weights according to the importance of semantic units. For example, higher weights can be assigned to keywords or sensitive words in the text, while lower weights can be assigned to common words. Feature enhancement factor ( ) is used to enhance the significance of feature vectors. Its value can be adjusted according to the specific application scenario and is usually between 0 and 1. Spatial compression coefficient ( ) is used to adjust the spatial distribution of the eigenvector and reduce the redundancy of the feature dimension. Its value is usually between 0 and 1. When calculating the modulation feature matrix, the hyperbolic tangent function (tanh) is used to map the eigenvalues to the range of -1 to 1, while the Sigmoid function ( ) maps the eigenvalues to the range of 0 to 1. The combination of the two can effectively handle the nonlinear relationship of the features. ) is an element-by-element multiplication operation, which is used to combine the weights with the feature vectors point by point, thereby realizing dynamic modulation of the features.
[0039] Preferably, the dynamic weight coefficient matrix ( ) can be further optimized by introducing contextual information. For example, the weight assignment can be dynamically adjusted in combination with the similarity or relevance of contextual semantics, so that the weight depends not only on the vocabulary itself but also on its role in the context. ) and the spatial compression factor ( ), can be adaptively adjusted according to different text types or risk levels. For example, when processing high-risk text, you can appropriately increase to enhance the saliency of the feature and reduce to retain more feature details.
[0040] Furthermore, the calculation of the modulation feature matrix can also introduce a regularization mechanism to prevent overfitting, such as by adding an L2 regularization term to constrain the size of the weight coefficient. This optimization method can further improve the robustness and adaptability of feature modulation, so that it can perform well in different types of text data.
[0041] In some embodiments, the multi-level condition combination analysis in step S4 includes: S41. Extract the singular value decomposition components of the modulation feature matrix H and calculate the energy distribution entropy E;
[0042] in, is the i-th singular value, S is the sum of singular values, is the first entropy threshold, is the second entropy threshold, and is a positive real number and > ; S42. When E> When , the first-level risk analysis module is activated to detect semantic coherence anomalies; S43. When ≤E≤ When the second-level risk analysis module is activated to detect sensitive pattern matching; S44. When E< When the third-level risk analysis module is activated, the context logic contradiction is detected.
[0043] It should be noted that the step of performing multi-level conditional combination analysis on the modulated feature matrix in the present invention is intended to calculate the energy distribution entropy by extracting the singular value decomposition components of the feature matrix, and activate risk analysis modules at different levels according to different ranges of entropy values. The singular value decomposition here is a mathematical method for decomposing a matrix into singular values and singular vectors, while the energy distribution entropy is the entropy value obtained by singular value calculation, which is used to measure the energy distribution of the feature matrix. According to the different ranges of energy distribution entropy, different levels of risk analysis modules are activated respectively, thereby realizing refined detection of text data risks. This process can effectively identify potential risks in texts and improve the accuracy and flexibility of risk detection through multi-level risk analysis.
[0044] Specifically, the energy distribution entropy (E) of the modulation feature matrix is obtained by calculating its singular value decomposition components. Singular value decomposition decomposes the matrix into singular values and singular vectors, where the magnitude of the singular values reflects the energy distribution in different directions in the matrix.
[0045] Preferably, in order to further optimize the performance of multi-level conditional combination analysis, the setting of entropy thresholds can be refined. For example, the thresholds can be dynamically adjusted according to the type of text data (such as news, social media, technical documents, etc.) or field (such as finance, medical care, education, etc.). and For texts in the financial field, stricter semantic coherence detection may be required, so the For social media text, we may pay more attention to sensitive pattern matching, so we can appropriately increase The value of .
[0046] Furthermore, an adaptive learning mechanism can be introduced to automatically adjust the threshold through a machine learning algorithm to adapt to the changing characteristics of text data. During the activation of the risk analysis module, contextual information and historical data can be combined to further enhance the accuracy and reliability of risk detection. For example, when detecting semantic coherence anomalies, a contextual semantic consistency check mechanism can be introduced; when detecting sensitive pattern matching, a sensitive word library that is updated in real time can be combined; when detecting contextual logical contradictions, a logical reasoning engine can be introduced. These optimization measures can effectively improve the performance of multi-level conditional combination analysis, making it more robust and adaptable in complex and changing text data.
[0047] In some embodiments, the demodulation process of step S5 includes: The demodulation function is:
[0048] in, is the demodulation coefficient matrix, is the trainable demodulation weight matrix, C is the risk category mapping matrix, K is the total number of risk categories, F is the interpretable risk feature vector, H is the modulation feature matrix, softmax is the normalization function, and ReLU is the linear rectification function.
[0049] It should be noted that the step of demodulating the risk probability distribution in the present invention is to convert the modulated feature matrix into an interpretable risk feature vector through a specific demodulation function. The demodulation process here refers to extracting the feature information related to the risk through mathematical transformation of the matrix after feature modulation so that it can be understood and used by the subsequent risk determination module. The key parameters involved in the demodulation function include the demodulation coefficient matrix ( ), trainable demodulation weight matrix ( ), risk category mapping matrix (C), etc. These parameters generate the final interpretable risk feature vector (F) through specific mathematical operations (such as matrix multiplication, softmax normalization function and ReLU linear rectification function). This process aims to transform the complex feature matrix into a feature vector with clear risk indication meaning, providing a clear basis for subsequent risk judgment.
[0050] During the demodulation process, the softmax function is used to normalize the feature vector into a probability distribution so that the score of each risk category is between 0 and 1 and the sum is 1; the ReLU function is used to introduce nonlinear characteristics to enhance the expressiveness of the model. The settings of these parameters and functions can be adjusted according to specific risk detection requirements to improve the accuracy and efficiency of demodulation.
[0051] Preferably, in order to further optimize the effect of the demodulation process, the demodulation coefficient matrix ( ) and the risk category mapping matrix ( ) is dynamically adjusted. For example, the demodulation coefficient matrix ( ) can be achieved by introducing an adaptive learning mechanism based on the input modulation feature matrix ( ) dynamically adjusts its value to better balance the relationship between the feature matrix and the risk category mapping matrix. Risk category mapping matrix ( ) can be updated in real time through online learning combined with real risk annotation data to improve the accuracy and timeliness of risk identification.
[0052] Furthermore, a multi-scale demodulation strategy can be introduced to extract richer and more detailed risk feature information by demodulating the modulated feature matrix at different levels. For example, demodulation can be performed separately at the local feature layer and the global feature layer, and then the results can be fused to improve the interpretability and distinguishability of the risk feature vector. These optimization measures can effectively improve the performance of the demodulation process, making it more adaptable and accurate in complex risk detection scenarios.
[0053] In some embodiments, the cross validation in step S6 includes: S61. Calculate the risk feature vector F and the preset typical risk pattern set The cosine similarity set ; in, is the jth typical risk pattern vector, m is the total number of patterns, is the jth cosine similarity, is the similarity threshold, is the confidence threshold, is the composite risk determination threshold, is the risk weight coefficient of the i-th category, and N is the number of consecutive determinations; S62. When there is > and > If the risk level of the deterministic risk reaches the preset high-risk standard, it is determined to be a high-risk risk. S63. When It is determined to be a compound risk when the risk level of the compound risk reaches the preset high-risk standard, and it is determined to be a high-risk risk. S64. When the results of N consecutive determinations meet the risk increasing rule, it is determined to be a high risk and the enhanced protection mechanism is triggered.
[0054] It should be noted that the step of cross-validating the risk feature vector in the present invention is to generate the final risk determination result by calculating the similarity between the risk feature vector and the preset typical risk pattern set. Cross-validation here refers to a comprehensive evaluation of the risk feature vector by combining a plurality of validation methods with a preset threshold value to ensure the accuracy and reliability of the risk determination. The preset typical risk pattern set ( ) is a set of feature vectors predefined according to known risk types and used to compare with the current risk feature vector ( ) for comparative analysis. By calculating the cosine similarity ( ) and confidence levels, the system can determine whether the current text has deterministic risks or compound risks, and trigger enhanced protection mechanisms based on the law of increasing risks.
[0055] High-risk risks are those that may pose a serious threat to the security and user experience of the target application. High-risk risks include: Risks based on specific pattern matching: The similarity between the risk feature vector and the preset typical risk pattern and the strength of the key risk features meet the standards at the same time, and the corresponding risk level is high. That is, when the similarity between the risk feature vector and a typical risk pattern exceeds the set standard, and the strength of the risk features of a specific key dimension in the vector also exceeds the set confidence standard, and the typical risk pattern belongs to the high-risk level in the risk level classification, it is judged as a high-risk risk.
[0056] Risks caused by compound factors: The result of the calculation of multiple risk feature dimensions and their weights exceeds the set standard, and the corresponding risk level is high risk. That is, the value of each dimension of the risk feature vector is multiplied by the corresponding weight and then added. If the comprehensive value exceeds the composite risk determination standard, and this composite risk belongs to the high risk level in the pre-set risk level system, it is determined to be a high risk.
[0057] Risks caused by increasing risks: When the results of N consecutive risk assessments show a trend of gradually increasing risk levels and meet the law of increasing risks, they are judged as high-risk risks.
[0058] The risk feature vector obtained after demodulation is compared with a pre-set typical risk pattern set to calculate the similarity between them. When the similarity between the risk feature vector and a typical risk pattern in the typical risk pattern set exceeds the pre-set similarity standard, and the risk feature intensity of a specific key dimension in the risk feature vector also exceeds the pre-set confidence standard, it is determined that there is a deterministic risk. If this deterministic risk belongs to a high-risk level risk according to the pre-set risk level classification standard, then the current text is determined to have a high-risk risk.
[0059] Comprehensively consider the risk characteristics of each dimension of the risk feature vector, and assign corresponding weights according to the importance of each dimension risk. Multiply the value of each dimension of the risk feature vector by the corresponding weight and add them together. If the result of this comprehensive calculation exceeds the pre-set composite risk determination standard, it is determined that there is a composite risk. If this composite risk belongs to a high-risk level risk in the pre-set risk level system, then the current text is also determined to have a high-risk risk.
[0060] In the process of continuous risk monitoring and judgment of the text, if N consecutive risk judgments are made, and the risk level obtained in each judgment shows a gradually increasing trend, that is, the risk level of the latter judgment is higher than that of the previous one, this situation is called satisfying the risk increasing law. When this risk increasing law occurs, regardless of whether the current risk feature vector meets the high-risk judgment conditions of the above-mentioned deterministic risk or compound risk, it is directly judged as a high-risk risk.
[0061] Specifically, the cross-validation process first calculates the risk feature vector ( ) and each pattern vector in the preset typical risk pattern set ( )'s cosine similarity ( ), the calculation formula of cosine similarity is: The cosine similarity value ranges from -1 to 1. The closer the value is to 1, the higher the similarity. The system sets the similarity threshold ( ) and confidence threshold ( ), when the cosine similarity of a pattern ( ) is greater than the similarity threshold ( ) and the corresponding risk characteristic value ( ) is greater than the confidence threshold ( ), it is determined to be a deterministic risk. In addition, the system will also calculate the composite risk determination index by accumulating the weighted risk values of different risk categories ( ), and the composite risk determination threshold ( ) are compared. When the accumulated value is greater than the threshold, it is determined to be a compound risk. These thresholds can be adjusted according to specific application scenarios and risk tolerance to meet different risk detection needs.
[0062] Preferably, in order to further improve the accuracy and flexibility of cross-validation, the similarity threshold ( ), confidence threshold ( ) and composite risk determination threshold ( ) for dynamic adjustment. For example, these thresholds can be dynamically set according to the field of text data (such as finance, medical care, social networking, etc.) and risk type (such as fraud, leakage, attack, etc.) to adapt to the risk characteristics in different scenarios. At the same time, the number of consecutive judgments ( ) mechanism. When the results of multiple consecutive judgments meet the law of increasing risk, the enhanced protection mechanism is triggered, thereby improving the system's response speed and protection capabilities. In addition, the typical risk pattern set can be dynamically updated in combination with machine learning algorithms. By continuously learning new risk samples and optimizing the risk pattern set, the accuracy and adaptability of cross-validation can be further improved.
[0063] In some embodiments, the multi-layer protection mechanism in step S7 includes: S71. The content replacement phase generates semantically preserved replacement text ,in, is the original text, M is the secure corpus; S72. Generate a progressive interruption instruction sequence during the session interruption phase ,in, is the i-th interrupt instruction, t is the total number of terminal steps, It is the final interrupt instruction; S73. Security alarm stage generates multi-dimensional alarm signals ,in, For local logging, Notify the remote server. It is the user interface warning layer.
[0064] It should be noted that in the present invention, when the judgment result is a high risk, a multi-layer protection mechanism will be triggered. This mechanism includes three stages: content replacement, session interruption and security alert. The multi-layer protection mechanism here refers to the coordinated action of multiple protection means to deal with different types of high-risk risks and ensure the security and stability of the system. Content replacement refers to replacing the detected high-risk text with safe text while retaining the original semantics as much as possible; session interruption refers to interrupting the current session when a risk is detected to prevent the risk from spreading further; security alerts notify relevant users or system administrators in various ways so that timely measures can be taken. The design of this mechanism aims to effectively reduce the impact of high-risk risks on the system through multi-level protection measures.
[0065] Specifically, in the multi-layer protection mechanism, the content replacement phase generates semantically preserved replacement text ( ), through the function Implementation, where is the original text, The process replaces high-risk content while preserving the semantics of the original text through an optimization algorithm. During the session interruption phase, a progressive interruption instruction sequence is generated ( ),in For the An interrupt instruction, is a time delay parameter used to control the rhythm of interruption to avoid excessive impact on user experience. In the security alert stage, a multi-dimensional alarm signal is generated ( ), including local logging ( )、Remote Server Notification( ) and user interface alerts ( ), to ensure that risk information can be communicated to relevant personnel in a timely manner. These parameters and mechanisms can be set and adjusted according to specific application scenarios to meet different risk protection needs.
[0066] Preferably, in order to further optimize the effect of the multi-layer protection mechanism, a more sophisticated semantic replacement strategy can be introduced in the content replacement stage. For example, by combining natural language generation technology, a more natural replacement text that is closer to the original semantics is generated, and a semantic similarity evaluation mechanism is introduced to ensure that the replaced text is highly semantically consistent with the original text. In the session interruption stage, the time delay parameter can be dynamically adjusted according to the risk level ( ), quickly interrupting high-risk scenarios and gradually interrupting low-risk scenarios to balance security and user experience. In the security alert stage, an intelligent alert system can be introduced to automatically select the appropriate alert method based on the risk type and severity, such as notifying relevant personnel in a timely manner through email, SMS or system notification. In addition, the security corpus can be dynamically updated in combination with machine learning algorithms to cope with changing risk characteristics and further improve the adaptability and effectiveness of the multi-layer protection mechanism.
[0067] In some embodiments, the dynamic weight coefficient matrix The calculation method is shown in the following formula:
[0068] Among them, U is the trainable parameter matrix, b is the bias vector, is an improved activation function, and V is a multi-dimensional semantic vector sequence.
[0069] It should be noted that the calculation method of the dynamic weight coefficient matrix in the present invention is implemented by a specific formula, which combines the trainable parameter matrix ( )、Bias vector( ) and the improved activation function ( ). Here the dynamic weight coefficient matrix ( ) is a key parameter for adjusting the feature vector weight, which can be adjusted according to the input multi-dimensional semantic vector sequence ( ) is generated dynamically, thereby enhancing the model's attention to important features. The trainable parameter matrix in the formula ( ) and the bias vector ( ) is a parameter learned during model training and is used to adjust the distribution of weights; the improved activation function ( ) is used to introduce nonlinear characteristics, making weight allocation more flexible. Through this calculation method, the model can dynamically adjust the weight according to the semantic features of the input text, improving the adaptability and accuracy of feature modulation.
[0070] For example, if the dimension of the semantic vector is D, then The dimension of can be set to D×D to ensure the legality of matrix multiplication. The bias vector ( ) has the same dimension as the semantic vector and is used to adjust the weight offset. The improved activation function ( ) can use variants of common activation functions such as ReLU, Sigmoid or Tanh to enhance the nonlinear expression ability of the model. During the training process, and The back propagation algorithm is used to update the model to minimize the loss function of the model, thereby optimizing the effect of weight distribution. In this way, the dynamic weight coefficient matrix can dynamically adjust the weights according to the semantic features of the input text, so that the model can better focus on important features and improve the accuracy of risk detection.
[0071] Preferably, in order to further optimize the calculation effect of the dynamic weight coefficient matrix, the activation function ( ) can be improved. For example, adaptive activation functions such as Swish or Mish are introduced. These activation functions can dynamically adjust their shapes according to the input, so as to better adapt to different semantic feature distributions. In addition, the parameter matrix ( ) is introduced into the training process of the model to prevent overfitting and improve the generalization ability of the model. ) settings, initial adjustments can be made according to specific application scenarios. For example, when processing high-risk texts, the bias value can be appropriately increased to improve the initial sensitivity of the weight.
[0072] Furthermore, we can combine multi-task learning to jointly train weight calculation with other tasks (such as semantic classification or sentiment analysis) to further improve the rationality and accuracy of weight distribution. These optimization measures can effectively improve the calculation effect of the dynamic weight coefficient matrix, making it more adaptable and robust in complex and changeable text data.
[0073] In some embodiments, the method for updating the risk category mapping matrix C includes:
[0074] in, is the risk category mapping matrix at the current moment, is the learning rate, is the real risk label vector, To predict the risk feature vector, is the outer product operation, and H is the modulation feature matrix.
[0075] It should be noted that the updating method of the risk category mapping matrix in the present invention is to dynamically adjust the real risk label vector and the predicted risk feature vector to optimize the risk identification ability of the model. ) is a matrix used to map feature vectors to specific risk categories, and its update process is achieved by learning the true risk annotations ( ) and the model prediction results ( ) is achieved by the difference between . The learning rate in the update formula is ( ) is a parameter that controls the update step size, and the outer product operation ( ) is used to feed back the error information into the matrix. Through this updating mechanism, the model can dynamically adjust the risk category mapping relationship according to the new labeled data, thereby improving the accuracy and adaptability of risk identification.
[0076] Specifically, the learning rate ( ) is usually set between 0 and 1, such as 0.01 or 0.001. The specific value can be adjusted according to the convergence speed and stability of the model. ) will error vector ( - ) and the feature matrix ( ) to generate a The update matrix of the same dimension can realize the dynamic adjustment of the risk category mapping matrix. This update mechanism enables the model to continuously learn new risk features during the training process and gradually optimize the accuracy of risk identification.
[0077] Preferably, in order to further optimize the update effect of the risk category mapping matrix, an adaptive learning rate adjustment mechanism can be introduced. For example, the learning rate is dynamically adjusted according to the error change of the model during the training process ( ), when the error is large, increase the learning rate to speed up the convergence, and when the error is small, reduce the learning rate to improve the accuracy.
[0078] Furthermore, regularization terms, such as L1 or L2 regularization, can be introduced to prevent overfitting of the mapping matrix during the update process. During the update process, batch normalization technology can be combined to normalize the error vector, thereby improving the stability and consistency of the update.
[0079] Furthermore, multi-source annotated data can be introduced to further enrich the model's learning content and improve the model's ability to identify different risk types by fusing real risk annotated vectors from different sources. These optimization measures can effectively improve the update effect of the risk category mapping matrix, making it more adaptable and accurate in a dynamically changing risk environment.
[0080] In some embodiments, the generation of the semantically-preserving replacement text satisfies:
[0081] in, is the balance coefficient, is the cosine distance, E is the semantic encoding function, is the edit distance, is the scaling factor, R is the current risk level, is the sigmoid function, T is the original text, To replace the text.
[0082] It should be noted that the method for generating semantically preserved replacement text in the present invention aims to retain the semantic information of the original text to the greatest extent while replacing the high-risk content through an optimization algorithm. ) refers to the text that meets security requirements and retains the core semantics of the original text after processing. The generation process minimizes the semantic distance ( ) and edit distance ( ), where the semantic distance measures the semantic similarity through cosine similarity and the edit distance measures the degree of text modification. This method combines the two dimensions of semantic consistency and text similarity to ensure that the replacement text strikes a balance between security and semantic coherence.
[0083] Specifically, and Represent the semantic encoding of the original text and the replacement text respectively, is the cosine distance, which is used to measure semantic similarity; is the edit distance, which is used to measure the degree of modification of the text; Is the balance coefficient, which is used to adjust the weights of semantic distance and edit distance. Balance coefficient The value of is usually between 0 and 1, for example, it can be set to 0.7 or 0.8, depending on the priority of semantic preservation and text modification. This can be achieved by using a pre-trained language model (such as BERT or GPT) to generate high-quality semantic vectors. By minimizing the above objective function, a replacement text with a high degree of semantic consistency with the original text can be generated while meeting security requirements.
[0084] Preferably, in order to further optimize the generation effect of semantically preserved replacement text, the balance coefficient For example, according to the current risk level ( ) Dynamic settings When the risk level is high, increase The value of should pay more attention to semantic preservation and ensure the semantic coherence of the replacement text; when the risk level is low, appropriately reduce to allow for more text modification.
[0085] Furthermore, a multi-candidate generation mechanism can be introduced to further improve the quality of the replacement text by generating multiple candidate replacement texts and screening them in combination with context information. ), a weight factor can be introduced , assigning different weights according to the importance of different parts of the text, thereby achieving more refined text modification control. These optimization measures can effectively improve the generation effect of semantically preserved replacement text, so that it can better preserve the semantic information of the original text while ensuring security, thereby improving user experience.
[0086] The above-mentioned embodiments of the present invention have the following beneficial effects: the present invention can generate a multidimensional semantic vector sequence by acquiring multiple text data streams in the target application and performing preprocessing, thereby providing a rich semantic basis for subsequent risk analysis. By adopting an asymmetric word segmentation algorithm and a dynamic window size, the word segmentation effect can be adaptively adjusted according to the complexity of the sentence, thereby further improving the accuracy of semantic parsing. The initial semantic vector is spatially projected by a bidirectional attention mechanism to form a multidimensional semantic vector sequence, which can better capture the contextual information in the text. The dynamic weight allocation module can perform feature modulation on the multidimensional semantic vector sequence, generate a modulation feature matrix, and enhance the expressiveness of the feature. The logic judgment engine can perform multi-level conditional combination analysis on the modulation feature matrix to generate a risk probability distribution, thereby realizing accurate identification of text risks. The demodulation process can convert the risk probability distribution into an interpretable risk feature vector, which is convenient for subsequent risk determination and analysis. The final risk determination result can be generated based on the cross-validation of the risk feature vector and the preset threshold set, thereby improving the accuracy and reliability of the risk determination. When the determination result is a high risk, the multi-layer protection mechanism can trigger content replacement, session interruption and security alarm to respond to the risk in a timely manner and ensure the security of the application. The content replacement stage can generate semantically preserved replacement text to ensure that the core semantics of the original text is retained while eliminating risks; the session interruption stage can generate a progressive interruption instruction sequence to avoid an overly abrupt experience for users; the security alert stage can promptly notify relevant personnel to take further measures through multi-dimensional alarm signals.
[0087] In addition, the calculation method of the dynamic weight coefficient matrix can further optimize the feature modulation process and improve the adaptability and flexibility of the system. The update method of the risk category mapping matrix can be dynamically adjusted based on the real risk label vector and the predicted risk feature vector to improve the self-learning ability and accuracy of the system. The method of generating semantically preserved replacement text can maintain the semantic consistency of the original text as much as possible while meeting security requirements, reduce the semantic loss caused by replacement, and improve user experience.
[0088] Furthermore, the storage medium of the embodiment of the present application stores program instructions that can implement all the above methods, wherein the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0089] The above descriptions are only some preferred embodiments of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A text recognition method based on a language model, characterized in that: The following steps are involved: S1 obtains the target application to be detected text data stream, the text data stream includes user input content, interface display text and network transmission text; S2. Preprocessing the text data stream to generate a multidimensional semantic vector sequence; S3. Using a dynamic weight allocation module to perform feature modulation on a multi-dimensional semantic vector sequence to generate a modulation feature matrix; S4. Input the modulation feature matrix into the logic judgment engine, perform multi-level condition combination analysis, and generate risk probability distribution; S5. Demodulate the risk probability distribution to generate an interpretable risk feature vector; S6. Perform cross validation based on the risk feature vector and the preset threshold set to generate the final risk determination result; S7. When the judgment result is a high risk, a multi-layer protection mechanism is triggered, including content replacement, session interruption and security alert.
2. The method according to claim 1, characterized in that: The step S2 specifically includes: S21. Using an asymmetric word segmentation algorithm to perform word segmentation semantic analysis on the text data stream, the word segmentation algorithm sets a dynamic window size W, and the window size W is adaptively adjusted according to the complexity of the sentence; S22. Perform context-related encoding on the word segmentation result to generate an initial semantic vector with a dimension of D, where the dimension D is a positive integer; S23. Use the bidirectional attention mechanism to spatially project the initial semantic vector to form a multi-dimensional semantic vector sequence , is the feature vector of the i-th semantic unit.
3. The method according to claim 2, characterized in that The characteristic modulation in step S3, the calculation formula of the modulation characteristic matrix H is: ; in, is the dynamic weight coefficient matrix, is the feature enhancement factor, is the space compression coefficient, V is the multi-dimensional semantic vector sequence, is the transposed matrix of V, is the hyperbolic tangent function, is the sigmoid function, is the Hadamard product, n is the total number of semantic units, and D is the dimension of the semantic vector.
4. The method according to claim 3, characterized in that: The multi-level condition combination analysis in step S4 includes: S41. Extract the singular value decomposition components of the modulation feature matrix H and calculate the energy distribution entropy E; ; in, is the i-th singular value, S is the sum of singular values, is the first entropy threshold, is the second entropy threshold, and is a positive real number and > ; S42. When E> When , the first-level risk analysis module is activated to detect semantic coherence anomalies; S43. When ≤E≤ When the second-level risk analysis module is activated to detect sensitive pattern matching; S44. When E< When the third-level risk analysis module is activated, the context logic contradiction is detected.
5. The method according to claim 4, characterized in that The demodulation function of the demodulation process in step S5 is: ; in, is the demodulation coefficient matrix, is the trainable demodulation weight matrix, C is the risk category mapping matrix, K is the total number of risk categories, F is the interpretable risk feature vector, H is the modulation feature matrix, softmax is the normalization function, and ReLU is the linear rectification function.
6. The method according to claim 5, characterized in that The cross validation in step S6 includes: S61. Calculate the risk feature vector F and the preset typical risk pattern set The cosine similarity set ; S62. When there is > and > If the risk level of the deterministic risk reaches the preset high-risk standard, it is determined to be a high-risk risk. S63. When It is determined to be a compound risk when the risk level of the compound risk reaches the preset high-risk standard, and it is determined to be a high-risk risk. S64. When the results of N consecutive determinations meet the risk increasing rule, it is determined to be a high risk and the enhanced protection mechanism is triggered; in, is the jth typical risk pattern vector, m is the total number of patterns, is the jth cosine similarity, is the kth dimension value of the risk feature vector F, is the i-th dimension value of the risk feature vector F, is the similarity threshold, is the confidence threshold, is the composite risk determination threshold, is the risk weight coefficient of the i-th category, and N is the number of consecutive judgments.
7. The method according to claim 6, characterized in that The multi-layer protection mechanism in step S7 includes: S71. The content replacement phase generates semantically preserved replacement text ,in, is the original text, M is the secure corpus; S72. Generate a progressive interruption instruction sequence during the session interruption phase ,in, is the i-th interrupt instruction, t is the total number of terminal steps, It is the final interrupt instruction; S73. Security alarm stage generates multi-dimensional alarm signals ,in, For local logging, Notify the remote server. It is the user interface warning layer.
8. The method according to claim 3, characterized in that The dynamic weight coefficient matrix The calculation method is shown in the following formula: ; Among them, U is the trainable parameter matrix, b is the bias vector, is an improved activation function, and V is a multi-dimensional semantic vector sequence.
9. The method according to claim 5, characterized in that The updating method of the risk category mapping matrix C includes: ; in, is the risk category mapping matrix at the current moment, is the learning rate, is the true risk label vector, To predict the risk feature vector, is the outer product operation, and H is the modulation feature matrix.
10. The method according to claim 7, characterized in that The generation of the semantically-preserving replacement text satisfies: ; in, is the balance coefficient, is the cosine distance, E is the semantic encoding function, is the edit distance, is the scaling factor, R is the current risk level, is the sigmoid function, T is the original text, To replace the text.
Citation Information
Cited By
Road standard relation construction and query method, equipment, medium and program product
CN120086356A
Business candidate person and member information management and storage method and system
CN120372669A
A method and system for managing and storing business candidate information
CN120372669B
Fraud phone real-time identification method and device based on AI semantic understanding
CN120639897A
Risk text accurate semantic recognition method based on improved deep learning model
CN121212158A