Information technology project compliance inspection methods, devices, equipment, media and products

By vectorized processing and entity relationship completion of information project document materials, combined with knowledge graphs and embedding learning, compliance inspection rules are generated, and the accuracy and automation of compliance inspection of information project are solved, and efficient and accurate compliance evaluation and dynamic rule updates are achieved.

CN119623453BActive Publication Date: 2025-05-06GUANGDONG SCI & TECH INFRASTRUCTURE CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510163193.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-06
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

It is difficult for the prior art to comprehensively and accurately evaluate the compliance of information projects, and there are problems such as inaccurate or omission of inspection results.

Method used

By obtaining the document materials of the information project, the entity is determined, and the preset seed rules are used for iterative learning to complete the entity relationship. Map entities and entity relationships to a common vector space, calculate similarity and align them, define symbolic rules based on compliance check files, and perform embedding learning to generate compliance check rules.

Benefits of technology

It has realized the automation of compliance inspections for information projects, reduced manual intervention, improved inspection efficiency and accuracy, and can reveal complex relationships between data, discover potential risk points, and support dynamic updates of rules to ensure that compliance inspections comply with the latest regulatory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623453B_ABST
    Figure CN119623453B_ABST
Patent Text Reader

Abstract

The present invention discloses a compliance inspection method, device, equipment, medium and product for an information technology project. The method includes vectorizing compliance inspection materials and identifying entities in the materials according to the contextual relationship of embedded vectors; iteratively learning and training entities using preset seed rules until the entity relationship of the entity is completed; mapping entities and entity relationships to a common vector space, calculating the similarity between entities and entity relationships, and aligning entities and entity relationships according to the similarity; defining symbol rules according to relevant compliance inspection files, iteratively learning aligned entities and entity relationships according to the symbol rules and embedded learning, obtaining compliance inspection rules, thereby performing compliance inspection on information technology projects and obtaining compliance inspection results of information technology projects. A comprehensive, accurate and reliable evaluation of the compliance inspection process of information technology projects is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of evaluation and detection technology, and in particular to an information technology project compliance inspection method, device, equipment, medium and product. Background Art

[0002] The compliance inspection of information technology projects refers to the process of comprehensively reviewing the project approval opinions, bidding documents, contract documents, construction plans, requirements specifications, change instructions, etc. according to the information technology project construction and acceptance standards and relevant policy documents formulated by the state, industry and local governments. This inspection process aims to ensure that the project is implemented in accordance with the established requirements, specifications and standards to achieve the expected results of the project. This inspection process not only helps to promote the standardization of project management and provide transparency and traceability of project implementation, but also ensures the effective use of project funds, avoids waste and losses, and ensures the maximization of investment benefits. However, in the actual implementation process, it may be impossible to conduct a comprehensive evaluation of the compliance of information technology projects due to reasons such as construction cycle, changes in project leaders, product iterations, missing documents, and policy document updates. Summary of the invention

[0003] The present invention provides an information technology project compliance inspection method, device, equipment, medium and product, which utilizes a knowledge graph to review information technology project compliance inspection materials, reduce inaccurate inspection results or omissions caused by overly simple or subjective inspection methods, and solve defects in information technology projects.

[0004] In order to achieve the above object, an embodiment of the present invention provides an information technology project compliance inspection method, including:

[0005] Acquire document materials of an information technology project, perform vectorization processing on the document materials to obtain embedding vectors of the document materials, and determine entities of the document materials according to contextual relationships of the embedding vectors;

[0006] Iteratively learning and training the entity using a preset seed rule until the entity relationship of the entity is completed;

[0007] Mapping the entities and entity relationships to a common vector space, calculating the similarity between the entities and entity relationships, and finding entities and entity relationships with similar spatial distances in the vector space according to the similarity and aligning them;

[0008] Defining symbolic rules according to relevant compliance check files, iteratively learning the aligned entities and entity relationships according to the symbolic rules and embedding learning to obtain compliance check rules;

[0009] The compliance check rule is used to perform a compliance check on the information project to obtain a compliance check result of the information project.

[0010] As an improvement of the above solution, the step of acquiring document materials of an information technology project, performing vectorization processing on the document materials to obtain an embedding vector of the document materials, and determining an entity of the document materials according to a contextual relationship of the embedding vector includes:

[0011] Acquire document materials of an information technology project, perform vectorization processing on the document materials, and obtain an embedding vector of the document materials;

[0012] Generating a contextual representation of the embedded vector by linear transformation according to the embedded vector;

[0013] According to the context representation, the entity of the document material is obtained through an agent classifier; wherein the agent classifier includes a BiLSTM layer, a CRF layer and an agent layer.

[0014] As an improvement of the above solution, obtaining the entity of the document material through an agent classifier according to the context representation includes:

[0015] Inputting the context representation into the BiLSTM layer to obtain context information of the embedding vector;

[0016] Input the context information into the CRF layer to obtain all label sequences and corresponding probabilities of the context information;

[0017] According to the tag sequence with the highest probability, the context information and the corresponding embedding vector, the entity of the document material is obtained through intelligent agent recognition.

[0018] As an improvement of the above solution, the entity is iteratively learned and trained using a preset seed rule until the entity relationship of the entity is completed, including:

[0019] According to the entity, a preset seed rule is set, and the preset seed rule is used to perform relationship matching on the entity to obtain the entity relationship of the entity; wherein the preset seed rule is a predefined entity relationship;

[0020] The preset seed rules and the entity relationship are used as a seed rule collection, and the trained relationship classifier is used to iteratively learn the entity according to the seed rule collection to obtain the entity relationship of the current iteration of the entity;

[0021] If the entity relationship of the entity is not complete, the entity relationship of the current iteration is included in the seed rule collection to obtain a new seed rule collection;

[0022] Then, the trained relationship classifier is used to iteratively learn the entity according to the new seed rule set to obtain the entity relationship of the current iteration of the entity until the entity relationship of the entity is completed.

[0023] As an improvement of the above solution, mapping the entities and entity relationships to a common vector space, calculating the similarity between the entities and entity relationships, and finding entities and entity relationships with similar spatial distances in the vector space according to the similarity and aligning them includes:

[0024] Normalizing the differences between the entities and entity relationships to obtain processed entities and entity relationships; performing representation learning on the processed entities and entity relationships using an embedding model to obtain triples of the processed entities and entity relationships;

[0025] Mapping the triples into a common low-dimensional vector space, and obtaining vector representations of processed entities and entity relationships through linear transformation;

[0026] The similarity between the vector representations is calculated, and entities and entity relationships with similar spatial distances in the low-dimensional vector space are found and aligned according to the similarity.

[0027] As an improvement of the above solution, the symbol rules are defined according to the relevant compliance check files, and the aligned entities and entity relationships are iteratively learned according to the symbol rules and embedding learning to obtain compliance check rules, including:

[0028] S41, according to the triples corresponding to the aligned entities and entity relationships, using embedding learning to obtain the embedding vectors corresponding to the aligned entities and entity relationships;

[0029] S42, based on the OWL2 object property axioms, and according to the relevant compliance check files, define the symbol rules to obtain a symbol rule collection;

[0030] S43, deriving a new symbol rule according to the symbol rule in the symbol rule collection and the embedding vector, and adding the new symbol rule to the symbol rule collection;

[0031] S44, using the deductive ability of the symbolic rule collection to infer new triples corresponding to the aligned entities and entity relationships;

[0032] S45, repeatedly executing S41-S44 until no new symbol rules can be derived, and obtaining compliance check rules according to the symbol rules in the symbol rule collection.

[0033] In order to achieve the above object, an embodiment of the present invention provides an information technology project compliance inspection device, including:

[0034] A document entity acquisition module, used to acquire document materials of an information project, perform vectorization processing on the document materials to obtain an embedding vector of the document materials, and determine the entity of the document materials according to the contextual relationship of the embedding vector;

[0035] An entity relationship completion module, used for iteratively learning and training the entity using a preset seed rule until the entity relationship of the entity is completed;

[0036] An entity relationship alignment module, used to map the entities and entity relationships to a common vector space, calculate the similarity between the entities and entity relationships, and find entities and entity relationships with similar spatial distances in the vector space according to the similarity and align them;

[0037] A check rule generation module is used to define symbol rules according to relevant compliance check files, and iteratively learn the aligned entities and entity relationships according to the symbol rules and embedding learning to obtain compliance check rules;

[0038] The inspection result obtaining module is used to perform a compliance inspection on the informationization project using the compliance inspection rules to obtain the compliance inspection result of the informationization project.

[0039] In order to achieve the above-mentioned purpose, an embodiment of the present invention provides an information technology project compliance checking device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the above-mentioned information technology project compliance checking method when executing the computer program.

[0040] In order to achieve the above-mentioned purpose, an embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned information project compliance checking method.

[0041] To achieve the above objectives, an embodiment of the present invention further provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the above-mentioned information technology project compliance checking method.

[0042] Compared with the prior art, the embodiments of the present invention disclose an information technology project compliance check method, device, equipment, medium and product. The method obtains document materials of the information technology project, vectorizes the document materials, obtains embedding vectors of the document materials, and determines entities of the document materials according to the contextual relationship of the embedding vectors; it iteratively learns and trains the entities using preset seed rules until the entity relationships of the entities are completed; maps the entities and entity relationships to a common vector space, calculates the similarities between the entities and entity relationships, and finds entities and entity relationships with similar spatial distances in the vector space according to the similarities and aligns them; defines symbol rules according to relevant compliance check files, iteratively learns the aligned entities and entity relationships according to the symbol rules and embedding learning, and obtains compliance check rules; uses the compliance check rules to perform compliance check on the information technology project to obtain compliance check results of the information technology project. It can automate compliance checks, reduce manual intervention, and improve inspection efficiency and accuracy. Through preset algorithms and models, the system can automatically analyze and compare data to quickly identify problems in compliance checks. Through entity and relationship modeling, it can reveal complex relationships between data, including direct and indirect relationships. This relationship mining capability helps to discover potential risk points in compliance checks, such as inconsistencies, contradictions, or potential violations. With the continuous changes in national, industry, and local policy documents, compliance check rules also need to be continuously updated. It supports dynamic update of rules to ensure that compliance checks always meet the latest regulatory requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flowchart of an information technology project compliance inspection method provided by an embodiment of the present invention;

[0044] Figure 2 It is a schematic diagram of a learning framework of deductive reasoning and inductive reasoning provided by an embodiment of the present invention;

[0045] Figure 3 It is a structural schematic diagram of an information project compliance inspection device provided by an embodiment of the present invention;

[0046] Figure 4 It is a structural block diagram of an information project compliance checking device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] It should be noted that the terms "comprises" and "specifically" and any variations of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0049] See also Figure 1 , Figure 1 The present invention provides a flow chart of a method for checking compliance of an information technology project, which includes:

[0050] S1, obtaining document materials of an information technology project, performing vectorization processing on the document materials to obtain an embedding vector of the document materials, and determining an entity of the document materials according to a contextual relationship of the embedding vector;

[0051] S2, iteratively learning and training the entity using a preset seed rule until the entity relationship of the entity is completed;

[0052] S3, mapping the entities and entity relationships to a common vector space, calculating the similarity between the entities and entity relationships, and finding entities and entity relationships with similar spatial distances in the vector space according to the similarity and aligning them;

[0053] S4, defining symbol rules according to relevant compliance check files, and iteratively learning the aligned entities and entity relationships according to the symbol rules and embedding learning to obtain compliance check rules;

[0054] S5, using the compliance check rule to perform a compliance check on the informationization project to obtain a compliance check result of the informationization project.

[0055] It is worth noting that the knowledge graph is a structured knowledge base used to represent entities, entity attributes, and relationships between entities. It presents this information in the form of nodes and edges, where nodes represent entities or concepts and edges represent relationships between them. Through the connection of entities and relationships, the knowledge graph can capture and represent rich semantic information, thereby achieving a deeper level of semantic understanding, while supporting complex queries and reasoning, information integration across data sources, dynamic updates and scalability, and improving data quality and accuracy.

[0056] Specifically, the step S1 includes:

[0057] S11, obtaining document materials of an information technology project, performing vectorization processing on the document materials, and obtaining an embedding vector of the document materials;

[0058] S12, generating a context representation of the embedded vector through a linear transformation according to the embedded vector;

[0059] S13, obtaining the entity of the document material through an agent classifier according to the context representation; wherein the agent classifier includes a BiLSTM layer, a CRF layer and an agent layer.

[0060] Exemplarily, in step S11, the information project document materials to be checked are collected, the compliance detection materials are vectorized, and the contextual relationship of different vectors is learned through the agent-based classifier to identify the entities in the materials. The document materials include the project approval opinions, bidding documents, contract documents, construction plans, requirements specifications, change instructions, etc. All materials are subjected to text preprocessing such as word segmentation, font unification, stop word removal, and stem extraction; then BERT is used to search for the corresponding embedding vector in the embedder's vocabulary for each word or vocabulary in the text, and the embedding vectors of long sentences are averaged, summed, or concatenated to obtain the embedding vector of the sentence.

[0061] In step S12, a query vector (Q), a key vector (K), and a value vector (V) are generated according to the embedding vector through linear transformation (i.e., multiplying by a weight matrix and adding a bias term), and a dot product operation is performed between the query vector and the key vectors of all words in the sequence, and then divided by a scaling factor (usually the square root of the key vector dimension) to obtain an attention score and convert the attention score into a probability distribution (i.e., attention weights). These weights represent the contribution of each word in the sequence to the current embedding vector, and finally the value vectors of all words in the sequence are weighted and summed using the attention weights to generate a contextual representation of the current embedding vector;

[0062] More specifically, the step S13 includes:

[0063] S131, inputting the context representation into a BiLSTM layer to obtain context information of the embedded vector;

[0064] S132, inputting the context information into a CRF layer to obtain all label sequences and corresponding probabilities of the context information;

[0065] S133, obtaining the entity of the document material through agent recognition according to the tag sequence with the highest probability, the context information and the corresponding embedding vector.

[0066] It is worth noting that the agent classifier learns behavioral strategies through interaction with the environment to maximize a certain cumulative reward. It can be regarded as treating the entity recognition process as a sequential decision problem, and optimizing the labeling strategy through the agent to improve the accuracy of entity recognition. The agent classifier structure includes the BiLSTM layer, the CRF layer, and the agent layer. Among them, the forward LSTM of the BiLSTM layer processes the input sequence according to the natural order of the text (from left to right) to generate a forward hidden state sequence. The reverse LSTM processes the input sequence according to the reverse order of the text (from right to left) to generate a reverse hidden state sequence. The hidden states of the forward LSTM and the reverse LSTM are concatenated as the output of the BiLSTM layer. In this way, the output of each time step contains the contextual information of the word; the CRF layer takes the output of the BiLSTM layer as input, uses the bidirectional processing capability of the BiLSTM to capture the contextual information of the text, and obtains the hidden state of each time step. At the same time, a transition matrix is ​​maintained inside the CRF layer to represent the transition probability between different labels. Through the transfer matrix and the output features of the BiLSTM layer, the CRF layer can calculate the probabilities of all possible label sequences and select the label sequence with the highest probability. The agent layer consists of three parts: state representation, action space, and reward function. The state representation can be represented as the word at the current position and its context information, which is provided by the BiLSTM layer; the action space is a set of all possible labels, that is, the agent needs to select the most appropriate label according to the current state; the reward function can be defined according to the accuracy of the labeling. For example, if the label selected by the agent is consistent with the true label, a positive reward is given; otherwise, a negative reward is given.

[0067] Pre-train the BiLSTM layer and the CRF layer to enable them to have basic entity recognition capabilities. Then, the agent layer is jointly trained with the BiLSTM layer and the CRF layer. During the training process, the agent layer selects actions (i.e., labels) based on the current state and obtains rewards based on the reward function. Through continuous iterative training, the agent can learn a better labeling strategy. Among them, the parameters of the agent layer can be updated through a policy gradient method (such as the REINFORCE algorithm). In each iteration, the gradient is calculated based on the actions selected by the agent and the rewards obtained, and the parameters of the agent are updated. At the same time, a threshold can be set. When the uncertainty of the labeling result of a certain position is greater than the threshold, the agent will correct the label of the position. Through the above step S13, entities such as the project amount, development content, operation and maintenance time, changes, equipment models, etc. in the document materials can be identified.

[0068] Specifically, the step S2 includes:

[0069] S21, according to the entity, set a preset seed rule, use the preset seed rule to perform relationship matching on the entity, and obtain the entity relationship of the entity; wherein the preset seed rule is a predefined entity relationship;

[0070] S22, using the preset seed rule and the entity relationship as a seed rule collection, using a trained relationship classifier, and iteratively learning the entity according to the seed rule collection to obtain the entity relationship of the current iteration of the entity;

[0071] S23, if the entity relationship of the entity is not complete, the entity relationship of the current iteration is included in the seed rule collection to obtain a new seed rule collection;

[0072] S24, then using the trained relationship classifier, iteratively learn the entity according to the new seed rule set to obtain the entity relationship of the current iteration of the entity, until the entity relationship of the entity is completed.

[0073] Exemplarily, based on the existing entities as the initial seed set, a new rule base is learned, and then new rules and relationships are extracted based on the new and old rule bases and the seed rule set is expanded. Through continuous iteration, new potential rules and relationships are found, discovered, and supplemented from unstructured data. First, the preset seed rules to be learned are defined (seed rules are predefined entity relationships that can be changed, but the identified entity relationships cannot be changed), such as "infrastructure services-public infrastructure", "software development services-public support platform", etc., and a certain number of seed instances are prepared for each relationship type, such as "public infrastructure-government cloud", "public support platform-data middle platform", etc.; then, the pre-trained model is used to match the defined seed rules and relationships with the entities and text corpus in the text library. If two entities are successfully matched in a sentence at the same time, it is assumed that the sentence describes the r relationship, and the relationship classifier is trained as a sample of the r relationship; the relationship classifier is jointly trained using the relationship twin network for learning, and then the obtained new rules and relationships are iteratively learned to obtain the entity relationship of the current iteration of the entity, until the entity relationship of the entity is completed. The identified entities are input into the twin network and the relationship classifier and iterated continuously. While learning new rules and new relationships, the learned rules and relationships are expanded to the collection of existing rule bases and relationship bases, thereby realizing the search and discovery of new potential rules and relationships in unstructured data.

[0074] It is worth noting that the twin network is a special neural network architecture, which consists of two neural networks with the same structure and shared weights, and is used to compare the similarity of two inputs. The twin network determines the similarity of two inputs by calculating the distance between them in the embedding space (such as the Euclidean distance). Two neural networks with similar structures and shared weights are designed as the two branches of the twin network (each branch accepts an entity as input and outputs the representation of the entity pair in the embedding space), and the branch output is used as the input of the relation classifier.

[0075] Furthermore, after completing the entity relationship of the entity, the step S2 further includes:

[0076] The similarity or distance between the entity and the corresponding entity attribute in the embedding space is calculated to complete the entity attribute of the entity.

[0077] For example, attribute completion can fill in the missing entity attributes in the knowledge graph, and can be used as a post-processing step to correct and improve the extraction results. After a certain number of iterations, the two branches of the twin network are adjusted to process entities and entity attributes (or attribute values) respectively, and the similarity or distance between entities and entity attributes in the embedding space is calculated to determine the correlation between them and thus achieve entity attribute completion.

[0078] Specifically, the step S3 includes:

[0079] S31, normalizing the differences between the entities and entity relationships to obtain processed entities and entity relationships; performing representation learning on the processed entities and entity relationships using an embedding model to obtain triples of the processed entities and entity relationships;

[0080] S32, mapping the triples into a common low-dimensional vector space, and obtaining vector representations of processed entities and entity relationships through linear transformation;

[0081] S33, calculating the similarity between the vector representations, finding entities and entity relationships with similar spatial distances in the low-dimensional vector space according to the similarity, and aligning them.

[0082] Exemplarily, in step S31, after the differences between entities and entity relationships are processed, representation learning is performed using an embedding model to linearly transform heterogeneous entities and heterogeneous entity relationships into aligned network representations. It can be understood that the heterogeneity of entities and entity relationships refers to the phenomenon that entities and relationships between them differ in structure and semantics in different data sources, different systems or different fields. The heterogeneity in this embodiment is mainly reflected in naming differences, representation differences and semantic differences. Therefore, different modules are designed to process the main differences to reduce the heterogeneity in different data sources, different systems or different fields. For example, the naming module performs preliminary sorting of naming, removes redundant, invalid or inconsistently formatted naming, uses string matching algorithms (such as exact matching, fuzzy matching, etc.) to identify possible naming differences in different ontologies, and identifies variant forms such as synonyms, near synonyms, abbreviations, and full names in naming. According to the identified naming differences, the corresponding normalization strategy is used for processing: using a synonym dictionary or semantic similarity calculation method, synonyms / near synonyms are merged into a standard name; according to the correspondence between abbreviations and full names, abbreviations are converted to full names, or vice versa; the naming is formatted in a unified manner, such as removing special characters, font formats, etc.; the representation module extracts information such as the attribute name, data type, and value range of entities and entity relationships from the input data; using the semantic similarity calculation method, the representation differences of the same or similar entities and entity relationships in different ontologies are identified. According to the identified representation differences in attribute names, data types, value ranges, etc., corresponding processing is performed, such as using synonym dictionaries, semantic similarity calculation methods, etc. for alignment of differences in attribute names; for differences in data types, conversion is performed according to data type conversion rules, such as converting string types to integer types; for differences in value ranges, mapping is performed according to value range mapping rules, and values ​​in different ranges are normalized. The semantic module inputs the sentence into the structure matcher, identifies the sentence components using the entities and entity relationships processed by the naming module and the representation module, and completes the attributes of the entities and entity relationships in the sentence according to the entity attributes in the attribute completion in step S223. Then, the context information of the sentence is integrated, and the semantic similarity score and the long-distance attention score are calculated using the Transformer model, and the total score threshold is set. When the total score is higher than the threshold, it means that the sentences are semantically similar, otherwise they are not similar. The processed entities and entity relationships are embedded in the TransE (Translating Embedding) model for representation learning, that is, the triple (h, r, t), where h represents the head entity, r represents the relationship, and t represents the tail entity (or heterogeneous entity), optimizes the vector representation of the entity and relationship, so that for the correct triple (h, r, t), the head entity vector h plus the relationship vector r can be close to or equal to the tail entity vector t, that is, h + r≈t.

[0083] In step S32, a vector space is constructed through a pre-trained language model (such as BERT), in which each entity and entity relationship is represented by a high-dimensional vector, where the distance in the vector space reflects the semantic similarity between entities and entity relationships, and the triples are mapped into a low-dimensional vector space, and translation invariance is used to optimize the vector representation of entities and relationships.

[0084] In step S33, in the vector space, for the entities and entity relationships in each domain, the vector most similar to the entities and entity relationships in the target domain (i.e., the nearest vector) is searched by calculating the cosine similarity between the vectors. Based on the search results of the nearest vector, a similarity threshold is set to align the entities and entity relationships in the source domain with the most similar entities and entity relationships in the target domain.

[0085] Specifically, the step S4 includes:

[0086] S41, according to the triples corresponding to the aligned entities and entity relationships, using embedding learning to obtain the embedding vectors corresponding to the aligned entities and entity relationships;

[0087] S42, based on the OWL2 object property axioms, and according to the relevant compliance check files, define the symbol rules to obtain a symbol rule collection;

[0088] S43, deriving a new symbol rule according to the symbol rule in the symbol rule collection and the embedding vector, and adding the new symbol rule to the symbol rule collection;

[0089] S44, using the deductive ability of the symbolic rule collection to infer new triples corresponding to the aligned entities and entity relationships;

[0090] S45, repeatedly executing S41-S44 until no new symbol rules can be derived, and obtaining compliance check rules according to the symbol rules in the symbol rule collection.

[0091] For example, Figure 2 As shown, Figure 2 It is a schematic diagram of a learning framework for deductive reasoning and inductive reasoning provided by an embodiment of the present invention; in step S41, the corresponding triples of aligned entities and entity relationships are obtained as input for embedding learning; it can be understood that the input of embedding learning includes forward input and reverse input, wherein the forward input includes fused entity triples and entity relationship triples, and the reverse input is randomly generated entity triples and entity relationship triples, and then the cross entropy loss function is used to calculate the average value of all input triples, find the embedding vector that can minimize the average loss, and obtain the embedding vector (vector rule) corresponding to the aligned entities and entity relationships.

[0092] In step S42, based on the OWL2 object property axioms, symbolic rules are defined according to the relevant compliance check files, and scores are assigned to each symbolic rule through effective pruning strategies and relationship embedding calculations to obtain a symbolic rule collection. It can be understood that OWL2 is the second version of Web Ontology Language, a formal ontology language for the Semantic Web, which aims to provide a richer and more flexible way to express information in the network and the relationships between them to support automated reasoning processes.

[0093] In step S43, a new symbolic rule is derived according to the symbolic rules in the symbolic rule collection and the embedding vector, and the new symbolic rule is included in the symbolic rule collection. For each derived symbolic rule, axiomatic induction will calculate the confidence of the symbolic rule through the applicability and consistency of the symbolic rule on the existing data (embedded vector). At the same time, in order to improve the calculation efficiency of the confidence, a set strategy for generating symbolic rules based on pattern matching is also designed, through patterns such as transitivity, equivalence, symmetry, and reflexivity, so as to effectively reduce the number of symbolic rules that need to be calculated.

[0094] In steps S44-S45, the deductive ability of axioms is used to infer new triples of entities and entity relationships, and a learning framework of deductive reasoning and inductive reasoning is constructed, and rule representation and entity representation are integrated to enhance the effect of embedded learning. For a given embedding vector and axiom set, axiom injection uses deductive reasoning to generate new triples and injects them into the next embedding learning. Among them, only triples whose tuple attributes are related to the embedding vector are retained in the generated new triples, and irrelevant new triples are filtered. After the above steps, this embodiment can use rule reasoning to complete general compliance checks on the content of document materials. For example, "check whether the project construction goals and construction contents are completed in accordance with the requirements of the project approval, filing plan, procurement documents and contracts", "whether the actual investment amount of the project main implementation fee and supporting fees is consistent with the filing amount", "capital investment should not exceed the filing amount or the project approval amount", etc. At the same time, if the compliance management method has detailed updates, the rules or axioms are defined using OWL2 in step S42, and after multiple iterations, until no new symbolic rules can be derived, compliance check rules are obtained according to the symbolic rules in the symbolic rule collection, so that the new management method details can be obtained.

[0095] The embodiment of the present invention discloses a compliance checking method for an information technology project. The method comprises the following steps: obtaining document materials of an information technology project, vectorizing the document materials, obtaining embedding vectors of the document materials, and determining entities of the document materials according to contextual relationships of the embedding vectors; iteratively learning and training the entities according to preset seed rules until entity relationships of the entities are completed; mapping the entities and entity relationships to a common vector space, calculating similarities between the entities and entity relationships, and finding and aligning entities and entity relationships with similar spatial distances in the vector space according to the similarities; defining symbol rules according to relevant compliance checking files, iteratively learning the aligned entities and entity relationships according to the symbol rules and embedding learning, and obtaining compliance checking rules; and using the compliance checking rules to perform compliance checking on the information technology project to obtain compliance checking results of the information technology project. It can automate compliance checks, reduce manual intervention, and improve inspection efficiency and accuracy. Through preset algorithms and models, the system can automatically analyze and compare data to quickly identify problems in compliance checks. Through the modeling of entities and relationships, it can reveal complex relationships between data, including direct and indirect relationships. This relationship mining capability helps to discover potential risk points in compliance checks, such as inconsistencies, contradictions, or potential violations. With the continuous changes in national, industrial and local policy documents, compliance inspection rules also need to be continuously updated. It supports dynamic update of rules to ensure that compliance checks always meet the latest regulatory requirements, thereby achieving a comprehensive, accurate and reliable evaluation of the compliance inspection process of information technology projects.

[0096] See also Figure 3 , Figure 3 1 is a schematic diagram of the structure of an information technology project compliance checking device 10 provided in an embodiment of the present invention. The information technology project compliance checking device 10 includes:

[0097] The document entity acquisition module 11 is used to acquire document materials of the information project, perform vectorization processing on the document materials to obtain embedding vectors of the document materials, and determine the entities of the document materials according to the contextual relationship of the embedding vectors;

[0098] An entity relationship completion module 12, configured to iteratively learn and train the entity using a preset seed rule until the entity relationship of the entity is completed;

[0099] An entity relationship alignment module 13, used to map the entities and entity relationships to a common vector space, calculate the similarity between the entities and entity relationships, and find entities and entity relationships with similar spatial distances in the vector space according to the similarity and align them;

[0100] A check rule generation module 14 is used to define symbol rules according to relevant compliance check files, and iteratively learn the aligned entities and entity relationships according to the symbol rules and embedding learning to obtain compliance check rules;

[0101] The inspection result obtaining module 15 is used to perform a compliance inspection on the informationization project using the compliance inspection rule to obtain a compliance inspection result of the informationization project.

[0102] An information technology project compliance checking device 10 provided in an embodiment of the present invention can implement all processes of the information technology project compliance checking method of the above-mentioned embodiment. The functions of each module in the device and the technical effects achieved are respectively the same as the functions and technical effects achieved by the information technology project compliance checking method of the above-mentioned embodiment, and will not be repeated here.

[0103] See also Figure 4 , Figure 4 1 is a schematic diagram of the structure of an information project compliance check device 20 provided in an embodiment of the present invention. The information project compliance check device 20 of this embodiment includes: a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above-mentioned information project compliance check method embodiment are implemented. Alternatively, when the processor 21 executes the computer program, the functions of each module in the above-mentioned information project compliance check device embodiment are implemented.

[0104] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory 22 and executed by the processor 21 to implement the present invention. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the information project compliance check device 20.

[0105] The information project compliance check device 20 may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The information project compliance check device 20 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art may understand that the schematic diagram is merely an example of the information project compliance check device 20 and does not constitute a limitation on the information project compliance check device 20. The information project compliance check device 20 may include more or fewer components than shown in the figure, or may combine certain components, or different components. For example, the information project compliance check device 20 may also include input and output devices, network access devices, buses, etc.

[0106] The processor 21 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor 21 is the control center of the information project compliance check device 20, and uses various interfaces and lines to connect various parts of the entire information project compliance check device 20.

[0107] The memory 22 can be used to store the computer program and / or module. The processor 21 realizes various functions of the information project compliance inspection device 20 by running or executing the computer program and / or module stored in the memory 22 and calling the data stored in the memory 22. The memory 22 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0108] Wherein, if the module integrated in the information project compliance check device 20 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-mentioned method embodiments can be implemented. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electrical carrier signals and telecommunication signals.

[0109] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0110] An embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the information technology project compliance checking method as described in the above embodiment.

[0111] In addition, an embodiment of the present invention also provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the information technology project compliance checking method of the above embodiment.

[0112] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for checking compliance of an information technology project, characterized in that: include: Acquire document materials of an information technology project, perform vectorization processing on the document materials to obtain embedding vectors of the document materials, and determine entities of the document materials according to contextual relationships of the embedding vectors; Iteratively learning and training the entity using a preset seed rule until the entity relationship of the entity is completed; Mapping the entities and entity relationships to a common vector space, calculating the similarity between the entities and entity relationships, and finding entities and entity relationships with similar spatial distances in the vector space according to the similarity and aligning them; Defining symbolic rules according to relevant compliance check files, iteratively learning the aligned entities and entity relationships according to the symbolic rules and embedding learning to obtain compliance check rules; Performing a compliance check on the information technology project using the compliance check rule to obtain a compliance check result of the information technology project; The step of acquiring document materials of an information technology project, performing vectorization processing on the document materials to obtain an embedding vector of the document materials, and determining an entity of the document materials according to a contextual relationship of the embedding vector includes: Acquire document materials of an information technology project, perform vectorization processing on the document materials, and obtain an embedding vector of the document materials; Generating a contextual representation of the embedded vector by linear transformation according to the embedded vector; According to the context representation, the entity of the document material is obtained through an agent classifier; wherein the agent classifier includes a BiLSTM layer, a CRF layer and an agent layer; The step of obtaining the entity of the document material by using an agent classifier according to the context representation includes: Inputting the context representation into the BiLSTM layer to obtain context information of the embedding vector; Input the context information into the CRF layer to obtain all label sequences and corresponding probabilities of the context information; The entity of the document material is obtained through intelligent agent recognition according to the tag sequence with the highest probability, the context information and the corresponding embedding vector.

2. The information project compliance inspection method according to claim 1, characterized in that: The adopting of preset seed rules to iteratively learn and train the entity until the entity relationship of the entity is completed includes: According to the entity, a preset seed rule is set, and the preset seed rule is used to perform relationship matching on the entity to obtain the entity relationship of the entity; wherein the preset seed rule is a predefined entity relationship; The preset seed rules and the entity relationship are used as a seed rule collection, and the trained relationship classifier is used to iteratively learn the entity according to the seed rule collection to obtain the entity relationship of the current iteration of the entity; If the entity relationship of the entity is not complete, the entity relationship of the current iteration is included in the seed rule collection to obtain a new seed rule collection; Then, the trained relationship classifier is used to iteratively learn the entity according to the new seed rule set to obtain the entity relationship of the current iteration of the entity until the entity relationship of the entity is completed.

3. The information project compliance inspection method according to claim 1, characterized in that: The mapping of the entities and entity relationships to a common vector space, calculating the similarity between the entities and entity relationships, and finding entities and entity relationships with similar spatial distances in the vector space according to the similarity and aligning them, comprises: Normalizing the differences between the entities and entity relationships to obtain processed entities and entity relationships; performing representation learning on the processed entities and entity relationships using an embedding model to obtain triples of the processed entities and entity relationships; Mapping the triples into a common low-dimensional vector space, and obtaining vector representations of processed entities and entity relationships through linear transformation; The similarity between the vector representations is calculated, and entities and entity relationships with similar spatial distances in the low-dimensional vector space are found and aligned according to the similarity.

4. The information project compliance inspection method according to claim 1, characterized in that: The method of defining a symbol rule according to the relevant compliance check file, iteratively learning the aligned entities and entity relationships according to the symbol rule and embedding learning, and obtaining the compliance check rule includes: S41, according to the triples corresponding to the aligned entities and entity relationships, using embedding learning to obtain the embedding vectors corresponding to the aligned entities and entity relationships; S42, based on the OWL2 object property axioms, and according to the relevant compliance check files, define the symbol rules to obtain a symbol rule collection; S43, deriving a new symbol rule according to the symbol rule in the symbol rule collection and the embedding vector, and classifying the new symbol rule into the symbol rule collection; S44, using the deductive ability of the symbolic rule collection to infer new triples corresponding to the aligned entities and entity relationships; S45, repeatedly executing S41-S44 until no new symbol rules can be derived, and obtaining compliance check rules according to the symbol rules in the symbol rule collection.

5. An information project compliance inspection device, characterized in that: include: A document entity acquisition module is used to acquire document materials of an information project, perform vectorization processing on the document materials to obtain an embedding vector of the document materials, and determine the entity of the document materials according to the contextual relationship of the embedding vector; An entity relationship completion module, used for iteratively learning and training the entity using a preset seed rule until the entity relationship of the entity is completed; An entity relationship alignment module, used to map the entities and entity relationships to a common vector space, calculate the similarity between the entities and entity relationships, and find entities and entity relationships with similar spatial distances in the vector space according to the similarity and align them; A check rule generation module is used to define symbol rules according to relevant compliance check files, and iteratively learn the aligned entities and entity relationships according to the symbol rules and embedding learning to obtain compliance check rules; An inspection result obtaining module, used to perform a compliance inspection on the informationization project using the compliance inspection rule to obtain a compliance inspection result of the informationization project; Wherein, the document entity acquisition module includes: Acquire document materials of an information technology project, perform vectorization processing on the document materials, and obtain an embedding vector of the document materials; Generating a contextual representation of the embedded vector by linear transformation according to the embedded vector; According to the context representation, the entity of the document material is obtained through an agent classifier; wherein the agent classifier includes a BiLSTM layer, a CRF layer and an agent layer; The step of obtaining the entity of the document material by using an agent classifier according to the context representation includes: Inputting the context representation into the BiLSTM layer to obtain context information of the embedding vector; Input the context information into the CRF layer to obtain all label sequences and corresponding probabilities of the context information; The entity of the document material is obtained through intelligent agent recognition according to the tag sequence with the highest probability, the context information and the corresponding embedding vector.

6. An information technology project compliance inspection device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the information project compliance checking method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the information technology project compliance checking method as described in any one of claims 1 to 4.

8. A computer program product, characterized in that The computer program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the information project compliance checking method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data processing method and device

    CN111966716A

  • Bridge field construction scheme examination method based on large model and knowledge graph

    CN118411016A