Document semantic risk identification method and device, computer equipment and storage medium
By segmenting documents and converting them using pre-trained language models, the problem of low efficiency in traditional document review is solved, enabling accurate identification and deep understanding of semantic risks in documents and generating structured risk reports.
Patent Information
- Application Number
- CN202511002990.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional document review methods are inefficient and fail to provide a deep understanding of semantic risks within documents, making it difficult to efficiently and accurately identify potential risks.
By dividing a document into multiple logically independent and semantically complete parts, we obtain the feature vectors of each part, construct the target semantic operator expression, and use a pre-trained language model for transformation and matching to accurately identify potential semantic risks.
It enables accurate risk identification of documents, enhances semantic clarity and comprehension, and allows for refined processing to avoid ambiguity in overall analysis, generating structured risk reports.
Smart Images

Figure CN120996045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of semantic recognition technology, and in particular to a method, apparatus, computer device, and storage medium for identifying semantic risks in documents. Background Technology
[0002] In today's information-saturated world, the number of documents is growing exponentially, making risk identification crucial. Traditional document review methods are inefficient and error-prone, failing to meet the demands of large-scale document processing. With technological advancements, a series of automated risk identification technologies have emerged, aiming to improve the efficiency and accuracy of document review to meet the stringent risk management requirements of various industries, thereby ensuring the safe and stable operation of businesses across all sectors.
[0003] In related technologies, risk identification methods based on regular expressions have been widely used in document review systems. These methods involve experts pre-defining a series of rules and regular expressions, and then automatically analyzing documents based on these rules to identify potentially risky content.
[0004] However, the relevant technologies lack deep semantic understanding capabilities, and the way rules are expressed differs significantly from human natural thinking, making it difficult to efficiently and accurately identify semantic risks in documents. Summary of the Invention
[0005] Therefore, it is necessary to provide a document semantic risk identification method, device, computer equipment, computer-readable storage medium, and computer program product that can achieve deep semantic understanding and accurate risk identification of document content, in response to the above-mentioned technical problems.
[0006] Firstly, this application provides a method for identifying document semantic risks. The method includes:
[0007] Obtain the block feature vectors of multiple document blocks of the document to be identified;
[0008] Expressions are constructed for each document block to obtain the target semantic operator expression;
[0009] The target semantic operator expression is transformed according to the language transformation model to obtain the target text; the language transformation model is used to convert abstract concepts into concrete text.
[0010] The target document segment is determined from multiple document segments by matching the pre-trained language model, the feature vectors of each segment, and the target text.
[0011] The target document is segmented and analyzed to obtain the document semantic risk analysis results.
[0012] In one embodiment, expressions are constructed for each document block to obtain target semantic operator expressions, including:
[0013] Determine target risk rules based on the type of document to be processed;
[0014] The preprocessing blocks are analyzed based on the target risk rules to determine the key target content; the key target content includes the vocabulary, phrases and abstract concepts of the preprocessing blocks.
[0015] Based on the target key content and the operator mapping relationship, the target operator is determined; the operator mapping relationship is the correspondence between the key content and the operator.
[0016] Determine the intermediate semantic operator expression based on the operator combination rules of the target operator and the target operator;
[0017] The intermediate semantic operator expression is optimized based on the operator parsing model to determine the target semantic operator expression.
[0018] In one embodiment, the target operator is determined based on the target key content and the operator mapping relationship, including:
[0019] The key content of the target is processed using a pre-trained language model to obtain the first block vector;
[0020] The key content in the operator mapping relationship is processed using a pre-trained language model to obtain the second block vector;
[0021] The similarity between the first block vector and the second block vector is calculated to obtain the similarity calculation result;
[0022] The target operator is determined based on the calculation results and the similarity threshold.
[0023] In one embodiment, matching is performed based on a pre-trained language model, feature vectors of each segment, and the target representation text to determine the target document segment from multiple document segments, including:
[0024] The retrieval vector is obtained by processing the pre-trained language model and the target text.
[0025] The vector database is searched based on the retrieval vector to determine the target block feature vector; the vector database is constructed from multiple block feature vectors.
[0026] The target document blocks are determined based on the target block feature vector.
[0027] In one embodiment, obtaining the block feature vectors of multiple document blocks of the document to be identified includes:
[0028] For each document block, a pre-trained language model is used to process the document blocks to obtain the block feature vectors.
[0029] In one embodiment, the method further includes:
[0030] Obtain the document to be recognized and segment it to obtain multiple blocks to be processed;
[0031] The preprocessing scheme is used to process each block to be processed, resulting in multiple document blocks.
[0032] Secondly, this application also provides a document semantic risk identification device. The device includes:
[0033] The vector extraction module is used to obtain the block feature vectors of multiple document blocks of the document to be identified;
[0034] The expression construction module is used to construct expressions for each document block to obtain the target semantic operator expression;
[0035] The text processing module is used to transform the target semantic operator expression based on the pre-trained language model to obtain the target expression text; the pre-trained language model is used to convert abstract concepts into concrete text;
[0036] The text matching module is used to match the pre-trained language model, the feature vectors of each block, and the target text to determine the target document block from multiple document blocks;
[0037] The risk analysis module is used to analyze the target document in segments to obtain the semantic risk analysis results.
[0038] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0039] Obtain the block feature vectors of multiple document blocks of the document to be identified;
[0040] Expressions are constructed for each document block to obtain the target semantic operator expression;
[0041] The target semantic operator expression is transformed based on the pre-trained language model to obtain the target text; the pre-trained language model is used to convert abstract concepts into concrete text.
[0042] The target document segment is determined from multiple document segments by matching the pre-trained language model, the feature vectors of each segment, and the target text.
[0043] The target document is segmented and analyzed to obtain the document semantic risk analysis results.
[0044] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0045] Obtain the block feature vectors of multiple document blocks of the document to be identified;
[0046] Expressions are constructed for each document block to obtain the target semantic operator expression;
[0047] The target semantic operator expression is transformed based on the pre-trained language model to obtain the target text; the pre-trained language model is used to convert abstract concepts into concrete text.
[0048] The target document segment is determined from multiple document segments by matching the pre-trained language model, the feature vectors of each segment, and the target text.
[0049] The target document is segmented and analyzed to obtain the document semantic risk analysis results.
[0050] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0051] Obtain the block feature vectors of multiple document blocks of the document to be identified;
[0052] Expressions are constructed for each document block to obtain the target semantic operator expression;
[0053] The target semantic operator expression is transformed based on the pre-trained language model to obtain the target text; the pre-trained language model is used to convert abstract concepts into concrete text.
[0054] The target document segment is determined from multiple document segments by matching the pre-trained language model, the feature vectors of each segment, and the target text.
[0055] The target document is segmented and analyzed to obtain the document semantic risk analysis results.
[0056] The aforementioned document semantic risk identification methods, devices, computer equipment, storage media, and computer program products, by segmenting the document to be identified into blocks and obtaining block feature vectors, can refine the processing and avoid ambiguity in overall analysis; then, by constructing expressions to obtain target semantic operator expressions, the semantic information of the document is presented in mathematical form, enhancing semantic clarity; further, with the help of a language conversion model, the target semantic operator expressions are converted into target descriptive text, thereby concretizing abstract concepts and further strengthening the understanding of the document; finally, by combining a pre-trained language model, block feature vectors, and target descriptive text, key content can be accurately extracted from multiple document blocks, thus enabling more accurate identification of potential semantic risks. Attached Figure Description
[0057] Figure 1 This is a flowchart illustrating a document semantic risk identification method in one embodiment;
[0058] Figure 2 This is a flowchart illustrating the steps for obtaining the target semantic operator expression in one embodiment;
[0059] Figure 3 This is a flowchart illustrating the steps for determining target document chunks in one embodiment;
[0060] Figure 4 This is a structural block diagram of a document semantic risk identification device in one embodiment;
[0061] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0063] In one embodiment, such as Figure 1 As shown, a document semantic risk identification method is provided. This embodiment illustrates the method applied to a terminal, but it is understood that the method can also be applied to a server, or to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0064] Step 102: Obtain the block feature vectors of multiple document blocks of the document to be identified.
[0065] For example, after obtaining the document to be identified, it needs to be segmented according to the principle of semantic integrity. This means considering the coherence and logic of the text's semantics and fully identifying the semantic relationships between different parts of the document during the segmentation process. Through detailed analysis, the document is divided into multiple logically independent and semantically complete text blocks for more efficient subsequent processing.
[0066] The text blocks are then processed to remove invalid characters, special symbols, formatting tags, etc., and then further segmented to obtain multiple document blocks. These multiple document blocks are then input into a pre-trained language model to obtain the block feature vectors corresponding to each document block.
[0067] Step 104: Construct expressions for each document block to obtain the target semantic operator expression.
[0068] For example, after obtaining multiple document chunks, it is necessary to construct operator expressions for each document chunk. The requirements need to be clearly defined, and the required operators should be determined based on the risk identification objectives. For text matching requirements, basic text matching operators such as "contains," "does not contain," and "exact match" are used. For instance, to find document chunks containing "default," "contains('default')" can be used. If semantic judgment is required, semantic judgment operators such as "semantic similarity" are used. For complex rules, composite semantic operators are used to combine basic operators, such as AND(contains('penalty'), contains('payment')). For domain-specific requirements, domain-specific semantic operators are used, such as the expression for "find default clauses." Finally, an operator expression parser is used to parse and optimize the constructed expressions to obtain the target semantic operator expression, ensuring its executableness.
[0069] Step 106: Transform the target semantic operator expression according to the language conversion model to obtain the target expression text; the language conversion model is used to convert abstract concepts into concrete text.
[0070] For example, the obtained target semantic operator expression refers to a kind of abstract, structured semantic representation, which may contain knowledge or rules of a certain domain. Therefore, a language conversion model is needed to generate the target text based on the target semantic operator expression, convert it into a natural language expression of the target domain, and map the abstract concepts to concrete text output. The language conversion model can be, but is not limited to, the Instruct language model and large open-source models such as Qwen.
[0071] For example, when dealing with the abstract concept of "breach of contract clauses," using explicit prompts (such as "Please expand the description of 'breach of contract clauses' into a specific form in the text"), the large language model will generate multiple different text descriptions based on the prompts, such as "if delivery is not made within the agreed time" or "if there is a breach of contract," etc. These expanded expressions can better adapt to different scenarios, providing more operable and understandable text.
[0072] Step 108: Match the pre-trained language model, the feature vectors of each segment, and the target text to determine the target document segment from multiple document segments.
[0073] For example, after obtaining the target representation text, the target representation text is processed using a pre-trained language model to convert it into a text vector, and the text vector is compared with the feature vectors of each block. For example, the similarity between the text vector and the feature vectors of each block is calculated, thereby determining the target document block corresponding to the feature vector of the block that is semantically similar to the text vector.
[0074] Step 110: Analyze the target document into blocks to obtain the document semantic risk analysis results.
[0075] For example, after determining the target document segments, they are analyzed in depth to extract their core semantics. Then, the core semantics are combined with the original semantic requirements of the semantic operator expressions corresponding to the text vectors and compared with preset risk standards to determine whether they constitute potential risks.
[0076] Following this, the identified risks will be categorized and organized to generate a structured risk report, which includes information such as the location of the risk, the type of risk, the severity of the risk, and improvement suggestions.
[0077] In the aforementioned document semantic risk identification method, by segmenting the document to be identified into blocks and obtaining block feature vectors, the ambiguity of the overall analysis can be avoided through refined processing. Then, by constructing expressions, the target semantic operator expression is obtained, which presents the semantic information of the document in mathematical form, enhancing the clarity of the semantics. With the help of a language conversion model, the target semantic operator expression is converted into the target descriptive text, thereby making abstract concepts concrete and further enhancing the understanding of the document. Finally, by combining the pre-trained language model, block feature vectors, and target descriptive text, key content can be accurately extracted from multiple document blocks, thereby enabling more accurate identification of potential semantic risks.
[0078] In one embodiment, such as Figure 2 As shown, expressions are constructed for each document block to obtain the target semantic operator expression, including:
[0079] Step 202: Determine the target risk rules based on the type of the document to be processed.
[0080] For example, after obtaining the documents to be processed, it is necessary to identify the document type and formulate specific risk rules, i.e., target risk rules, based on their nature. For instance, legal documents should focus on compliance, the accuracy of contract terms, and potential legal liabilities; technical documents need to focus on technical accuracy, data security, and intellectual property protection; and marketing materials should avoid misleading advertising, false promises, etc. By accurately identifying the document type, risk rules can be formulated in a targeted manner.
[0081] Step 204: Analyze the preprocessing blocks according to the target risk rules to determine the target key content; the target key content includes the vocabulary, phrases and abstract concepts of the preprocessing blocks.
[0082] For example, the content in the preprocessing blocks is analyzed according to target risk rules. These rules are formulated based on the risk characteristics and business needs of a specific domain to determine whether the text content is related to risk. In the specific analysis process, risk-related elements, i.e., target key content, are carefully screened from the preprocessing blocks.
[0083] The key content of the objectives covers multiple levels. Vocabulary is the most basic; words like "crash" and "fraud," which clearly point to risk, can directly raise awareness of risk. Phrases further enrich the information; for example, "broken capital chain" more clearly depicts a risk scenario. Abstract concepts, while not concrete, can reflect risk at a macro level; for example, "systemic risk" provides a perspective for a comprehensive understanding of potential crises.
[0084] Step 206: Determine the target operator based on the target key content and the operator mapping relationship; the operator mapping relationship is the correspondence between the key content and the operator.
[0085] For example, after determining the key content of the target, the operators to be used, i.e., the target operators, are determined based on a pre-set operator mapping relationship. These operators include basic semantic operators, composite semantic operators, and domain-specific semantic operators.
[0086] For basic semantic operators, there are mainly two types: text matching operators and semantic judgment operators, specifically:
[0087] Text matching operators include, but are not limited to, "contains", "does not contain", "exact match", or "equals" operators. The "contains" operator uses regular expressions (such as `match`, `search`, etc.) or string matching functions (such as `contains`, `startswith`, `endswith`) to directly match text segments. For example, "contains 'default'" can be expressed as `match(".*default.*")`. The "does not contain" operator is achieved by logically negating the result of the "contains" operator; for example, "text does not contain 'fine'" can be expressed as `NOT(match(".*fine.*"))`. The "exact match" or "equals" operator uses strict text equality comparisons, such as `equals("contract validity period")`.
[0088] Semantic judgment operators perform semantic matching through vector similarity calculation or embedding representations of language models. For example, the "semantic similarity" operator can be implemented by generating text block vector representations using a language model and then calculating cosine similarity, such as similarity(embedding(text A), embedding(text B)) > threshold.
[0089] For composite semantic operators, logical combination operators such as "AND", "OR", and "NOT" are defined, and nested combinations of multiple basic semantic operators are supported. For example:
[0090] The "AND" operator is used to express that multiple conditions are met simultaneously. For example, AND("penalty"), "payment")) represents a text block that contains both "penalty" and "payment".
[0091] The "OR" operator is used to express any of the following conditions being met: OR(contains("compensation"), contains("reimbursement")) represents a text block containing either "compensation" or "reimbursement".
[0092] For domain-specific semantic operators, expression is achieved through user-defined or expert-defined specific business terms or risk scenarios, combined with rule logic and semantic judgment. For example:
[0093] It can be defined as OR("breach of contract liability"), similarity(embedding(text chunks), embedding("breach of contract consequences"))>0.85), which combines direct text matching and semantic similarity judgment to achieve accurate and flexible expression.
[0094] Step 208: Determine the intermediate semantic operator expression based on the operator combination rules of the target operator and the target operator.
[0095] For example, first, the basic semantic operators, composite semantic operators, or domain-specific semantic operators involved in the target operator are identified. Then, according to operator combination rules, such as "AND" requiring multiple conditions to be satisfied simultaneously, and "OR" requiring only one condition to be satisfied, these operators are nested and combined. For instance, if the goal is to find content that contains specific words and is semantically similar, "AND" can be used to connect the "contains" and "semantically similar" operators. In this way, the target operator and its combination rules are transformed into specific intermediate semantic operator expressions.
[0096] Step 210: Optimize the intermediate semantic operator expression based on the operator parsing model to determine the target semantic operator expression.
[0097] For example, the operator parsing model is responsible for converting user-defined semantic operator expressions into executable code or structured queries, enabling subsequent steps to identify semantic risks in the document text. This allows for multifaceted optimization of the intermediate expressions, such as checking for syntax errors to ensure compliance with system specifications and prevent execution errors, as well as optimizing the expression structure by removing redundant parts and improving execution efficiency. Therefore, after constructing the intermediate semantic operator expression, the operator parsing model performs nested expression parsing, syntax error checking, and operator expression optimization to obtain the optimized semantic operator expression, i.e., the target semantic operator expression.
[0098] In one embodiment, determining the target operator based on the target key content and the operator mapping relationship includes:
[0099] The first block vector is obtained by processing the key content of the target using a pre-trained language model; the second block vector is obtained by processing the key content in the operator mapping relationship using the pre-trained language model; the similarity between the first block vector and the second block vector is calculated to obtain the similarity calculation result; the target operator is determined based on the calculation result and the similarity threshold.
[0100] For example, a pre-trained language model is used to process the target key content. This processing scheme, which converts text into vectors with semantic information, transforms the target key content into a first block vector, i.e., a numerical semantic representation of the target key content. Similarly, the key content in the operator mapping relationship is also processed using a pre-trained language model to obtain a second block vector.
[0101] Then, similarity calculation can be performed using cosine similarity. The pre-calculated similarity result is then compared with a similarity threshold. If the similarity exceeds the threshold, such as similarity(embedding(textA), embedding(textB)) > threshold, the operator corresponding to the key content in the operator mapping relationship can be set as the target operator.
[0102] In one embodiment, such as Figure 3 As shown, the target document segment is determined from multiple document segments by matching the pre-trained language model, the feature vectors of each segment, and the target text representation, including:
[0103] Step 302: Process the pre-trained language model and the target text to obtain the retrieval vector.
[0104] For example, the obtained target text is input into a pre-trained language model, which transforms the target text into a high-dimensional vector representation of the semantic representation, i.e., a retrieval vector.
[0105] Step 304: Search the vector database based on the retrieval vector to determine the target block feature vector; the vector database is constructed from multiple block feature vectors.
[0106] For example, when processing the document to be recognized, the content of the document blocks is encoded into high-dimensional vectors by a pre-trained language model to build a unified vector space database.
[0107] After obtaining the retrieval vector, its similarity is calculated with each vector in the vector space data. For example, a cosine similarity calculation algorithm can be used. Specifically:
[0108]
[0109] Where A and B are the retrieval vector and the document block vector, respectively. The cosine similarity value ranges from [-1, 1], with values closer to 1 indicating greater semantic similarity.
[0110] Then, the target block feature vectors that are semantically similar to the retrieval vector are determined.
[0111] Step 306: Determine the target document blocks based on the target block feature vector.
[0112] For example, after determining the target block feature vector through similarity calculation, the document block corresponding to the target block feature vector is set as the target document block.
[0113] In one embodiment, obtaining the block feature vectors of multiple document blocks of the document to be identified includes:
[0114] For each document block, a pre-trained language model is used to process the document blocks to obtain the block feature vectors.
[0115] For example, after processing the document to be identified to obtain multiple document blocks, for each document block:
[0116] The preprocessed text chunks are input into a pre-trained language model. The model converts the text into word vectors through a word embedding layer, adds positional information to the word vectors through a positional embedding layer, and then encodes the input sequence using the Transformer attention mechanism to extract semantic features, obtaining the hidden representation of the text chunks in the semantic space. Next, high-dimensional vector feature extraction is performed, extracting feature vectors from the hidden representation output by the pre-trained language model, thus obtaining the chunk feature vectors of the document chunks.
[0117] In one embodiment, the method further includes:
[0118] The document to be identified is obtained and segmented into multiple blocks to be processed. Each block is then processed based on a preprocessing scheme to obtain multiple document blocks.
[0119] For example, after obtaining the document to be identified, a segmentation operation is carried out based on the principle of semantic integrity. The semantic relationships between the various parts of the document are carefully identified, and through in-depth and detailed analysis, the document is divided into multiple logically independent and semantically complete text blocks.
[0120] Then, these text blocks undergo preprocessing, primarily removing invalid characters, special symbols, and irrelevant formatting tags. Afterward, the blocks are further segmented to obtain multiple document blocks.
[0121] In one exemplary embodiment, a document semantic risk identification method is provided, the method comprising the following steps:
[0122] Obtain the document to be recognized and segment it to obtain multiple blocks to be processed.
[0123] The preprocessing scheme is used to process each block to be processed, resulting in multiple document blocks.
[0124] For each document block, a pre-trained language model is used to process the document blocks to obtain the block feature vectors.
[0125] Determine target risk rules based on the type of document to be processed.
[0126] The preprocessing blocks are analyzed according to the target risk rules to determine the key target content; the key target content includes the vocabulary, phrases and abstract concepts of the preprocessing blocks.
[0127] Based on the target key content and the operator mapping relationship, the target operator is determined; the operator mapping relationship is the correspondence between the key content and the operator.
[0128] The intermediate semantic operator expression is determined based on the operator combination rules of the target operator and the target operator.
[0129] The intermediate semantic operator expression is optimized based on the operator parsing model to determine the target semantic operator expression.
[0130] The target semantic operator expression is transformed according to the language transformation model to obtain the target text; the language transformation model is used to convert abstract concepts into concrete text.
[0131] The retrieval vector is obtained by processing the pre-trained language model and the target text.
[0132] The vector database is searched based on the retrieval vector to determine the target block feature vector; the vector database is constructed from multiple block feature vectors.
[0133] The target document blocks are determined based on the target block feature vector.
[0134] The target document is segmented and analyzed to obtain the document semantic risk analysis results.
[0135] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0136] Based on the same inventive concept, this application also provides a document semantic risk identification device for implementing the document semantic risk identification method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more document semantic risk identification device embodiments provided below can be found in the limitations of the document semantic risk identification method described above, and will not be repeated here.
[0137] In one embodiment, such as Figure 4As shown, a document semantic risk identification device is provided, including: a vector extraction module 402, an expression construction module 404, a text processing module 406, a text matching module 408, and a risk analysis module 410, wherein:
[0138] The vector extraction module 402 is used to obtain the block feature vectors of multiple document blocks of the document to be identified.
[0139] The expression construction module 404 is used to construct expressions for each document block to obtain the target semantic operator expression.
[0140] The text processing module 406 is used to transform the target semantic operator expression according to the pre-trained language model to obtain the target expression text; the pre-trained language model is used to convert abstract concepts into concrete text.
[0141] The text matching module 408 is used to match the pre-trained language model, the feature vectors of each block, and the target text to determine the target document block from multiple document blocks.
[0142] The risk analysis module 410 is used to analyze the target document in blocks to obtain the document semantic risk analysis results.
[0143] In one embodiment, the expression construction module 404 is further configured to: determine target risk rules based on the type of the document to be processed; analyze the preprocessing blocks according to the target risk rules to determine target key content; the target key content includes the vocabulary, phrases, and abstract concepts of the preprocessing blocks; determine target operators according to the target key content and operator mapping relationship; the operator mapping relationship is the correspondence between key content and operators; determine intermediate semantic operator expressions according to the operator combination rules of the target operators and the target operators; and optimize the intermediate semantic operator expressions based on the operator parsing model to determine the target semantic operator expressions.
[0144] In one embodiment, the expression construction module 404 is further configured to process the target key content using a pre-trained language model to obtain a first block vector; process the key content in the operator mapping relationship using a pre-trained language model to obtain a second block vector; perform similarity calculation between the first block vector and the second block vector to obtain a similarity calculation result; and determine the target operator based on the calculation result and a similarity threshold.
[0145] In one embodiment, the text matching module 408 is used to process the pre-trained language model and the target representation text to obtain a retrieval vector; to search the vector database based on the retrieval vector to determine the target block feature vector; the vector database is constructed from multiple block feature vectors; and to determine the target document block based on the target block feature vector.
[0146] In one embodiment, the vector extraction module 402 is used to process the document blocks using a pre-trained language model to obtain the block feature vectors of the document blocks.
[0147] In one embodiment, the document semantic risk identification device is used to acquire a document to be identified and to segment the document to be identified to obtain multiple blocks to be processed; and to process each block to be processed based on a preprocessing scheme to obtain multiple document blocks.
[0148] Each module in the aforementioned document semantic risk identification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0149] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores document chunks and chunk feature vector data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a document semantic risk identification method.
[0150] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0151] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0152] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0153] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0155] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0156] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0157] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for identifying semantic risks in documents, characterized in that, The method includes: Obtain the block feature vectors of multiple document blocks of the document to be identified; Expressions are constructed for each document block to obtain the target semantic operator expression; The target semantic operator expression is transformed according to the language conversion model to obtain the target expression text; the language conversion model is used to convert abstract concepts into concrete text; The target document block is determined from the multiple document blocks by matching the pre-trained language model, the feature vectors of each block, and the target text. The target document is segmented and analyzed to obtain the document semantic risk analysis results.
2. The method according to claim 1, characterized in that, The step of constructing expressions for each document block to obtain target semantic operator expressions includes: Determine target risk rules based on the type of document to be processed; The preprocessing blocks are analyzed based on the target risk rules to determine the key target content; the key target content includes the vocabulary, phrases and abstract concepts of the preprocessing blocks. Based on the target key content and the operator mapping relationship, the target operator is determined; the operator mapping relationship is the correspondence between the key content and the operator. Determine the intermediate semantic operator expression based on the operator combination rules of the target operator and the target operator; The intermediate semantic operator expression is optimized based on the operator parsing model to determine the target semantic operator expression.
3. The method according to claim 2, characterized in that, The step of determining the target operator based on the target key content and the operator mapping relationship includes: The target key content is processed using a pre-trained language model to obtain the first block vector; The key content in the operator mapping relationship is processed using a pre-trained language model to obtain the second block vector; The similarity between the first block vector and the second block vector is calculated to obtain the similarity calculation result; The target operator is determined based on the calculation results and the similarity threshold.
4. The method according to claim 1, characterized in that, The step of matching the pre-trained language model, the feature vectors of each segment, and the target text representation to determine the target document segment from the multiple document segments includes: The retrieval vector is obtained by processing the pre-trained language model and the target text. The vector database is searched based on the retrieval vector to determine the target block feature vector; the vector database is constructed from multiple block feature vectors. The target document blocks are determined based on the target block feature vector.
5. The method according to claim 1, characterized in that, The step of obtaining the block feature vectors of multiple document blocks of the document to be identified includes: For each document segment, the pre-trained language model is used to process the document segment to obtain the segment feature vector of the document segment.
6. The method according to claim 1, characterized in that, The method further includes: The document to be identified is obtained and segmented to obtain multiple blocks to be processed; The preprocessing scheme is used to process each of the blocks to be processed, resulting in multiple document blocks.
7. A document semantic risk identification device, characterized in that, The device includes: The vector extraction module is used to obtain the block feature vectors of multiple document blocks of the document to be identified; An expression construction module is used to construct expressions for each of the document blocks to obtain target semantic operator expressions; The text processing module is used to transform the target semantic operator expression according to the pre-trained language model to obtain the target expression text; the pre-trained language model is used to convert abstract concepts into concrete text; The text matching module is used to match the pre-trained language model, each of the block feature vectors and the target representation text to determine the target document block from the multiple document blocks; The risk analysis module is used to analyze the target document segments to obtain the document semantic risk analysis results.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.