Person and social knowledge graph adaptive optimization method and system based on uncertain information

By combining multi-domain weighted membership degree and fuzzy rough set methods, along with multi-stage Markov processes and multi-hop attention layers, the knowledge graph structure is dynamically adjusted. This solves the accuracy and timeliness issues of traditional knowledge graphs under cross-domain and policy changes, achieving highly stable and real-time updated knowledge graph optimization.

CN120806085APending Publication Date: 2025-10-17DAREWAY SOFTWARE
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510880139.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional knowledge graphs struggle to accurately depict the fuzzy relationships between entities and attributes in cross-domain, complex logic, and multimodal scenarios, and they cannot respond promptly to policy changes and business needs, resulting in insufficient accuracy and timeliness.

Method used

A method combining multi-domain weighted membership degree and fuzzy rough set is adopted. Through multi-stage Markov process and multi-hop attention layer, the knowledge graph structure is dynamically adjusted to identify high-risk objects and perform adaptive updates. Combined with policy index and cross-departmental interaction factors, it can achieve accurate identification and adaptive optimization of uncertain information.

Benefits of technology

It improves the stability and real-time update capability of cross-domain data fusion, reduces boundary bias, ensures that the knowledge graph maintains the reliability of entities and relationships in the time-series dimension, and enhances the accuracy and responsiveness of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806085A_ABST
    Figure CN120806085A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of knowledge graph updating, and provides a human social knowledge graph adaptive optimization method and system based on uncertain information, and the technical scheme is as follows: associating each piece of data to a corresponding policy label according to a policy index, and constructing an initial multi-dimensional graph structure according to service connection and shared fields; establishing a multi-field weighted membership function and initializing a fuzzy set, and performing comprehensive credibility combination on cross-department entities and relationships; performing cross-department difference degree calculation on the suspicious entities and relationships, and correcting a state transition probability based on a multi-stage Markov process in combination with a cross-department interaction factor and a policy aging factor; carrying out aggregation on the high-confidence nodes and the edges after state transition by applying a multi-hop attention layer; performing versioning management on the optimized graph structure and establishing an incremental updating queue; according to the method, the credible evolution of the knowledge graph tracking nodes and edges in the time sequence dimension can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of knowledge graph updating, and particularly relates to a human resources knowledge graph adaptive optimization method and system based on uncertain information. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] In the application scenario of multi-department human resources business, the cross-domain data is large in scale and different in structure, often involving multiple association relationships and policy dimensions. The traditional knowledge graph focuses more on the representation of static entities and relationships, and cannot fully solve the problems of dynamic fusion and conflict correction between heterogeneous data. Especially when the data sources come from different departments such as health and education, field differences, record duplication, and uncertain attributes often bring significant difficulty in fusion.

[0004] If the existing method relies too much on fixed rules or a single threshold, it will be difficult to accurately determine the validity of the data at the "boundary". Once the parameters are not properly selected, it is easy to cause false filtering or delayed updating.

[0005] The traditional human resources knowledge graph construction method has the following problems: 1. When rules involving cross-domain, complex logic and multi-modal are involved, the knowledge graph cannot alleviate the differences in attribute values and entity boundaries in different dimensions. The traditional knowledge graph construction method often relies on clear and consistent data input, and cannot accurately depict the fuzzy relationship between different entities and attributes in the graph in the scene where the data involves cross-domain, complex logic and dependent relationships, and includes multi-modal. It is difficult to ensure the accuracy and integrity of the knowledge graph; 2. Due to the frequent updates of policy-driven human resources field and the variability of business scenarios, under the demand of high-frequency updates and real-time decision-making, the traditional mechanism of periodically updating knowledge base cannot timely capture the rule changes brought by new policies, resulting in that the graph content is difficult to keep up with the pace of business evolution, causing the knowledge graph to be unable to respond to knowledge support with accuracy and timeliness.

[0006] To solve the above technical problems, the fuzzy set and rough set theory shows certain value in the description of uncertainty: the fuzzy membership degree can gradually represent the cross-department data in different confidence intervals, and the upper and lower approximations of rough set can provide flexible division for high-risk or boundary objects; by introducing the multi-field weighted membership degree into the dynamic upper and lower approximations of rough set, the adaptive expansion and contraction classification of suspicious objects can be realized in the actual environment where the policy frequency and data update rate are not the same; compared with simple filtering or hard-coded rules, this method can more accurately capture the difference and uncertainty of cross-department data; however, only fuzzy-rough processing of uncertain parts is not enough to ensure the continuous evolution and real-time decision of the graph. SUMMARY

[0007] To solve at least one technical problem in the above background art, the present application provides an adaptive optimization method and system for human social knowledge graph based on uncertain information, which realizes accurate identification and adaptive update of uncertain information through multi-department data fusion, upper and lower approximation determination and state transition analysis, thereby further improving the tracking of the credibility evolution of each node and edge in the time dimension of the knowledge graph.

[0008] To achieve the above purpose, the present application adopts the following technical solutions: The first aspect of the present application provides an adaptive optimization method for human social knowledge graph based on uncertain information, comprising: Preprocessing the obtained multi-department business data, associating each data to the corresponding policy label according to the policy index, and constructing an initial multi-dimensional graph structure according to the business connection and shared fields; Based on the initial multi-dimensional graph structure, a fuzzy set is established, the fuzzy set is adjusted, and a multi-field weighted membership degree function of the adjusted fuzzy set is established; the dynamic fuzzy rough boundary region of the fuzzy set is calculated combined with the multi-field weighted membership degree function, the suspicious entities and relationships in the dynamic fuzzy rough boundary region are identified, and a high-risk object set is formed; Calculate the cross-department similarity, and clean the high-risk object set by counting all cross-department similarity calculation results, and extract the knowledge graph set of observable state; Introducing a multi-stage Markov process, updating the knowledge graph set of observable state, integrating the nodes and edges with high confidence after state transition into the updated knowledge graph structure, and combining the multi-hop attention layer to weight and aggregate the adjacent information and change frequency of different hop numbers, to obtain the optimized knowledge graph structure; Taking the optimized knowledge graph structure as the benchmark version, adaptively correcting the membership degree of the new nodes or relationships, and again measuring the difference between the new and old versions through the multi-stage Markov process, when the difference is greater than a set value, backtrack the data source and adjust the graph structure.

[0009] Further, the policy index is used to associate each piece of data with the corresponding policy label, and a preliminary multi-dimensional graph structure is constructed according to the business connection and shared fields, including: Establish a set of human resources business policy indexes; Combine the set of human resources business policy indexes, define a policy mapping function for each data record in the multi-department business data, determine the set of policy labels required to be associated with each piece of data, and perform labeling; The data records after labeling are regarded as a set of nodes, and a set of edges is defined based on the business connection and shared fields between data, and the initial multi-dimensional graph structure is constructed by combining the association relationship of different nodes under each policy index; Wherein, the policy mapping function is expressed as: , Wherein, represents the association degree evaluation of record and policy , is the association degree threshold, when exceeds , is included in the policy label set of the record.

[0010] Further, the calculation formula of the dynamic fuzzy rough boundary region of the fuzzy set is: , Wherein, represents the universal set covered by the current knowledge graph, and represent the upper approximation and lower approximation of the adjusted fuzzy set , and represent the comprehensive membership representation of the upper approximation and the lower approximation respectively, is a dynamic threshold reflecting the policy update frequency and data change rate; Or, the cross-department similarity calculation formula is: , , Wherein, is the cross-department similarity, is the cumulative weight value calculated on the matrix , , represents the difference measure of department and department on object , represents the total number of departments, Indicates that the department The global weight coefficient of Representation object local weight in its sector, Represents the exponential amplification factor for the difference measure.

[0011] Furthermore, the transfer probability correction formula for each stage is: , in, Indicates that From the state Transfer to state The basic probability of represents the corrected transition probability, represents the cross-sector interaction factor, represents the policy time factor; if multiple departments activate interactions on the same entity, then The value increases; if the policy update frequency increases, then Increase.

[0012] Furthermore, the multi-stage Markov process is introduced to update the knowledge graph set of observable states to obtain an updated knowledge graph structure, including: Based on the knowledge graph set of observable states, the spatial state is defined. Combined with the multi-stage Markov process, the cross-departmental interaction factors and policy timeliness factors are incorporated into the state transition matrix of each stage to correct the state transition probability of each stage. The knowledge graph structure is updated by combining the spatial state and the multi-stage state transition probability to obtain the updated knowledge graph structure.

[0013] Furthermore, in the initial reconstructed graph, each node and edge is attached based on Timestamp of the stage , record cross-departmental interaction information.

[0014] Furthermore, the multi-hop attention layer is combined to perform weighted aggregation on the adjacency information of different hops and the change frequency as follows: , , in, is the multi-hop fusion matrix, Indicates Adjacency extension within hop range, Indicates the maximum number of multi-hop layers, Indicates the Jump attention coefficient, represents the weighting function for cross-sector interactions, Indicates that a node or relationship is in the frequency of change appearing in the jump range, represents the equilibrium constant; Or, the adaptive membership correction function for the newly added node or relationship is: , wherein, represents the number of departments; represents any newly introduced relationship or node, represents the original membership in the benchmark version, represents the department the given membership of the newly added relationship , the global weight coefficient of the department , represents the control factor.

[0015] The second aspect of the present application provides a human social knowledge graph adaptive optimization system based on uncertain information, comprising: An initial multi-dimensional graph structure construction module is used for preprocessing the obtained multi-department business data, associating each data to the corresponding policy label according to the policy index, and constructing an initial multi-dimensional graph structure according to the business connection and sharing field; An observable knowledge graph extraction module is used for establishing a fuzzy set based on the initial multi-dimensional graph structure, adjusting the fuzzy set, establishing a multi-field weighted membership function of the adjusted fuzzy set; combining the multi-field weighted membership function to calculate the dynamic fuzzy rough boundary region of the fuzzy set, identifying the suspicious entities and relationships in the dynamic fuzzy rough boundary region to form a high-risk object set; calculating the cross-department similarity, cleaning the high-risk object set by counting all cross-department similarity calculation results, and extracting the knowledge graph set in the observable state; A knowledge graph optimization module is used for introducing a multi-stage Markov process, updating the knowledge graph set in the observable state, integrating the nodes and edges with high confidence after state transition into the updated knowledge graph structure, combining a multi-hop attention layer to weight and aggregate the adjacency information and the change frequency of different hop numbers, and obtaining the optimized knowledge graph structure; An adaptive correction module is used for taking the optimized knowledge graph structure as a benchmark version, performing adaptive membership correction on the newly added node or relationship, and again measuring the difference between the new and old versions through the multi-stage Markov process. When the difference is greater than a set value, backtrack the data source and adjust the graph structure.

[0016] The third aspect of the present application provides a computer readable storage medium.

[0017] A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method for adaptive optimization of a human social knowledge graph based on uncertain information as described above.

[0018] A fourth aspect of the present application provides a computer device.

[0019] A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method for adaptive optimization of a human social knowledge graph based on uncertain information as described above when executing the program.

[0020] Compared with the prior art, the beneficial effects of the present application are: 1. Compared with the traditional knowledge graph which focuses more on the processing of static entities and relationships, the present application uses a combination of multi-field weighted membership and fuzzy rough set to improve the description of uncertain data and the division of dynamic boundaries from the mathematical level. By adjusting the weighted membership and dynamic threshold, hierarchical screening of potential conflicts and suspected redundant elements is achieved, which not only ensures the accuracy of multi-source data, but also reduces the boundary deviation caused by large-scale heterogeneous input. Thus, it reflects the ability to make finer granularity division in high-risk areas, thereby improving the overall stability of cross-domain data fusion.

[0021] 2. The present application integrates cross-department interaction factors and policy time effectiveness factors into the state transition equation by constructing a multi-stage Markov process, so that the knowledge graph can maintain continuous estimation of entity and relationship reliability in the time sequence dimension. Unlike single or fixed window static reasoning, multi-stage iteration can refer to the data update rate of each department in each state transition and correct the transition probability according to the change frequency of each policy dimension. This combination of hierarchical iteration and multi-hop attention can more effectively capture the iterative stabilization process of cross-domain data and avoid global judgment deviation due to local anomalies or fragmented repeated records. By iteratively correcting the confidence distribution of nodes and edges in the multi-stage Markov chain, the present application ultimately realizes adaptive fusion and real-time update of large-scale heterogeneous data, thereby exhibiting better convergence and robustness in cross-department human social business scenarios.

[0022] The advantages of the additional aspects of the present application will be partially given in the following description, partially will become apparent from the following description, or will be understood by the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings, which form a part of the present application, are included to provide a further understanding of the application, and are incorporated in and constitute a part of this specification, illustrate embodiments of the present application and together with the description serve to explain the present application. Embodiments of the present application and its

[0024] Figure 1 is a flowchart of a human social knowledge graph adaptive optimization method based on uncertain information provided by an embodiment of the present application; Figure 2 is a weighted membership calculation schematic diagram of a human social knowledge graph update optimization method provided by an embodiment of the present application; Figure 3 is a knowledge graph evolution schematic diagram of a human social knowledge graph update optimization method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] The present application will be further described below in conjunction with the accompanying drawings and embodiments.

[0026] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0027] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and furthermore, it should be understood that when the terms "comprise" and / or "include" are used in the specification, there is a feature, step, operation, device, component and / or combination thereof.

[0028] The traditional knowledge graph construction method has the following problems: First, when rules involving cross-domain, complex logic and multi-modal are involved, the knowledge graph cannot alleviate the difference between attribute value and entity boundary in different dimensions; Specifically: 1. The human social field involves multiple sub-fields such as employment, social security, labor relations, etc. Each sub-field has its unique concepts and rules, and the terms and concepts between different fields may differ, making it difficult to represent and map uniformly in the knowledge graph; 2. The human social field knowledge graph usually involves a large number of entities and relationships, and the business rules often involve multiple conditions, multiple steps and multiple participants, and there may be complex dependency relationships between rules. These complex rules are difficult to represent and reason directly through simple triple relationships in the knowledge graph; 3. The data sources of the human social field are extensive, which may include government documents, enterprise internal data, network public information and other modalities. The data of different modalities may be heterogeneous in underlying representation. For example, the text description in the policy document and the accompanying structured data such as tables and charts, the knowledge graph faces difficulties in processing such multi-modal rules.

[0029] 4. Business rules in the human resources and social security field often involve multiple conditions, which may have complex dependencies. For example, when calculating social security benefits, it may be necessary to consider multiple factors such as an individual's years of contribution, contribution base, and retirement age, and these factors may affect each other.

[0030] For multi-source data with inconsistent quality and unclear relevance, traditional knowledge graph construction methods often rely on clear and consistent data input. They are unable to accurately depict the fuzzy relationships between different entities and attributes in the graph in scenarios where the data involves cross-domain, complex logic and dependencies, and includes multimodality, making it difficult to ensure the accuracy and completeness of the knowledge graph.

[0031] Second, when based on dynamically changing rules, knowledge graphs cannot handle the accurate mapping of surging business rules and policy iterations; 1. Frequent policy-driven updates: Policies in the human resources and social security field are constantly adjusted as the socioeconomic environment changes. For example, social security contribution rates, minimum wage standards, retirement policies, etc. may all be adjusted based on national or local policy changes. These policy adjustments will directly affect the entities, attributes, and relationships in the knowledge graph. Traditional knowledge graphs that are updated monthly cannot reflect the latest policy changes.

[0032] 2. Diversity of business scenarios: With the development of new technologies and new business models, the business scenarios in the human resources and social security sector are constantly changing, and the integration with other fields is also deepening. For example, the emergence of new employment models such as remote work, flexible employment, and the sharing economy has put forward new requirements for labor relations management and social security contributions.

[0033] Under the demand for high-frequency updates and real-time decision-making, the traditional mechanism of regularly updating the knowledge base cannot timely capture the rule changes brought about by new policies, resulting in the graph content being unable to keep up with the pace of business evolution, and the knowledge graph being unable to respond to accurate and timely knowledge support.

[0034] To address these challenges, fuzzy set and rough set theories have demonstrated their value in characterizing uncertainty. Fuzzy membership allows for a progressive representation of cross-departmental data within varying credibility intervals, while rough set upper and lower approximations provide a flexible classification method for high-risk or borderline objects. By incorporating multi-domain weighted membership into the dynamic upper and lower approximations of rough sets, adaptive expansion and contraction of suspicious object classification can be achieved in real-world environments where policy frequencies and data update rates vary. Compared to simple filtering or hard-coded rules, this approach can more accurately capture the differences and uncertainties in cross-departmental data. However, simply fuzzy-roughening the uncertain parts is not enough to ensure the continuous evolution of the graph and real-time decision-making. To this end, the present invention constructs a multi-stage Markov process on this basis, integrating cross-departmental interaction factors and policy timeliness factors into the process model of state transfer, so that the knowledge graph can "track" the credible evolution of each node and edge in the time dimension. When the state distribution of certain nodes or relationships reaches a high confidence level, the graph will be updated in time. If it is still in the fuzzy range, it will enter the subsequent stage for re-evaluation.

[0035] Example 1 like Figure 1 As shown, this embodiment provides a method for adaptive optimization of a human and social knowledge graph based on uncertain information, including the following steps: Step 1: Preprocess the acquired multi-department business data, associate each piece of data with the corresponding policy label based on the policy index, and build a preliminary multi-dimensional graph structure based on business connections and shared fields; The specific steps include: Step 101: Fusing the acquired heterogeneous data from multiple departments to obtain a multi-department dataset; Acquire heterogeneous data from multiple departments and construct original data sets. Collect business data from the health, education, and public security departments, analyze the data fields and their storage formats, and extract them into a unified structured style to form a health department data set. , education department data collection , Public Security Department Data Collection ; Merge the above three data sets to form a multi-department original data set , using the set representation as follows: ; Step 102: preprocess the acquired original data set to obtain a preprocessed multi-department data set; In this embodiment, the original data set is cleaned and formatted. Check each non-compliant record and the degree of field missing, eliminate entries that do not meet the unified analysis standards, and remove obvious outliers and duplicate records to obtain a preliminary cleansed data set; Perform weighted smoothing filtering on the retained entries, defining each data record as is the number of fields), when the window length is When , let the weighted smoothing function Expressed as: , in, represents the adjacent record vector, express The weighting coefficient at the relative position, denotes the result of smoothing, denotes the smoothing window radius; This step retains the main data characteristics while suppressing extreme value interference, obtaining a denoised multi-department data set ; Step 103, based on the pre-established policy index, label each data record in the preprocessed multi-department data set, associate each labeled data with its corresponding human social policy label, and define edges according to business connection and shared fields to construct a preliminary multi-dimensional graph structure; Specifically, it includes: Step 1031, establish a human social business policy index set ; Step 1032, in combination with the human social business policy index set, define a policy mapping function for each data record in the multi-department data to determine its required associated policy label set, complete data labeling, and the expression is as follows: , where denotes the association degree evaluation of the record and the policy , is the association degree threshold; when exceeds , it will be included in the policy label set of the record; through this labeling process, the matching relationship between each data and human social policy can be clearly defined when constructing the knowledge graph later; This step completes data labeling by taking the association degree between multi-department data and each policy as the key labeling basis; Step 1033, consider the data records after completing labeling as a node set, define an edge set based on the business connection and shared fields between data, and construct an initial multi-dimensional graph structure based on the association relationship of different nodes under each policy index; Specifically, consider the data records after completing labeling as a node set , define an edge set based on the business connection and shared fields between data , and form a graph ; To simultaneously present entity association under different policy dimensions, introduce the policy into the adjacency matrix , define the association relationship of node and under the policy dimension ​, the expression is as follows: , in, express and Under this policy dimension, there is a direct business relationship. If there is no relationship, .

[0036] Step 2: Establish a fuzzy set based on the initial multi-dimensional graph structure, adjust the fuzzy set, and establish a multi-domain weighted membership function for the adjusted fuzzy set; calculate the dynamic fuzzy rough boundary region of the fuzzy set based on the multi-domain weighted membership function, identify suspicious entities and relationships in the dynamic fuzzy rough boundary region, and form a high-risk object set; like Figure 2 As shown, the specific steps include: Step 201: establishing a fuzzy set based on the initial multi-dimensional graph structure, adjusting the fuzzy set, and establishing a multi-domain weighted membership function of the adjusted fuzzy set; Specifically, based on the established initial multi-dimensional graph structure, entities from different departments and their attributes are considered as original fuzzy sets , Indicates the department index; Specifically, the adjustment of fuzzy sets includes: Building cross-sectoral domain weights Indicates its credibility adjustment coefficient in a comprehensive business scenario; through cross-departmental field weights Determine the weighting method for adjusting the original fuzzy set. When adjusting the fuzzy set, the cross-departmental field weights are involved, indicating the impact of these weights on the credibility in the comprehensive business scenario; combined with the cross-departmental field weights The fuzzy membership of each department is weighted and combined to form a comprehensive fuzzy set , to more accurately reflect the differences in multi-department data and their credibility in business scenarios.

[0037] Specifically, the multi-domain weighted membership function of the adjusted fuzzy set is established including: set up is any entity or relationship to be evaluated, Represents the entity or relationship In the department The original membership in Represents the entity or relationship In the department The local credibility factor under Indicates department The weighted coefficient under the global situation, then the multi-domain weighted membership function is recorded as : , in, Indicates the number of departments; Through this weighting method, the comprehensive fuzzy set Can better reflect cross-departmental differences; Step 202: Calculate the dynamic fuzzy rough boundary region of the fuzzy set by combining the multi-domain weighted membership function of the adjusted fuzzy set; Specifically, the calculation formula for the motion blur rough boundary area is: , in, Represents the full set covered by the current knowledge graph, is the adjusted fuzzy set The upper approximation of for Lower approximation, and Respectively represent the comprehensive membership representation of upper approximation and lower approximation, A dynamic threshold that reflects the frequency of policy updates and the rate of data changes, used to control the extent of the boundary area. When the value is large, the boundary area shrinks, and when the value is small, the boundary area expands. Step 203: Mark high-risk items by combining the dynamic fuzzy rough boundary area of ​​the fuzzy set; For entities and relationships, the confidence performance of multiple departments in the dynamic fuzzy rough boundary area is comprehensively considered. When the confidence is lower than the set threshold, the comprehensive fuzzy set is recalculated. The comprehensive attributes or data part is in The comprehensive membership in is used for secondary division; For those who remain in The part is marked as high risk or high uncertainty items and recorded in the subsequent scheduling queue. All nodes and edges marked as high risk or high uncertainty items constitute a high risk object set , The objects in the ,are subject to high uncertainty due to the large differences in the ,affiliation between departments; In confirmation Finally, the edges and nodes involving multiple policy dimensions are preliminarily marked, including the source department information and the core attributes provided by each department, to facilitate the precise distinction of data information from different departments in the subsequent weighted comparison link.

[0038] Step 3: Calculate cross-department similarity, count all cross-department similarity calculation results, clean the high-risk object set, and extract the knowledge graph set of observable states; The specific steps include: Step 301, based on the high-risk object set, combined with the constructed multi-department weighted comparison strategy, calculate the cross-department similarity; Specifically, according to the object attribute distribution in the high-risk object set , define the cross-department weighted comparison matrix ; For any high-risk object , let denote the key attribute difference measure provided by department and department to the object; Quantify the difference degree for all department pairs and organize them into a matrix ; To comprehensively evaluate the difference degree in multiple dimensions of attributes, propose a cross-department similarity function , which is realized based on the weighted sum of ; let denote the global weight coefficient of department pair , let denote the local weight of object in its own department, and define the cross-department similarity ; Calculate the cumulative weight value on the matrix : , Combine the exponential decay factor to enhance the proportion from reliable departments: , where denotes the total number of departments, denotes the difference measure of department and department on object , denotes the exponential amplification coefficient of the difference measure; By this strategy, global and local weights are added to the multi-department data difference, so that the overall similarity of the object in is accurately measured; Step 302, count all cross-department similarity calculation results to clean the high-risk object set, and extract the knowledge graph set of observable state; Specifically, when the cross-department similarity of object is lower than the threshold , it is determined that there is serious inconsistent information in some departments, marked as invalid data and executed for degradation or deletion operation; the degradation operation retains the basic attributes but reduces its priority in the knowledge graph reasoning link; The cleaned or degraded object set is sent to the subsequent process, avoiding invalid data interference with dynamic updating model; the remaining high-risk objects If is higher than , it is retained and monitored in the subsequent step to form the overall set of cleaned knowledge graph , and the entities or relationships in the high-risk area that have not been deleted are marked as , represent the abnormal state set that can be observed in the subsequent Markov modeling; The above steps compare the suspicious entities and relationships in the boundary area across departments, calculate the core attribute difference between different departments, and calculate the cross-department similarity based on global and local weights; when the similarity is lower than the set threshold, the object is marked as invalid or degraded, which not only ensures the accuracy of multi-source data, but also reduces the boundary deviation caused by large-scale heterogeneous input.

[0039] Step 4: Define the spatial state based on the knowledge graph set of observable states, combine the multi-stage Markov process, and modify the state transition probability of each stage by incorporating cross-department interaction factors and policy time effectiveness factors in the state transition matrix of each stage, update the knowledge graph structure based on the spatial state and multi-stage state transition probability, and obtain the updated knowledge graph structure; Specifically, the following steps are included: Step 401, define the spatial state based on the knowledge graph set of observable states; Specifically, the discrete state of each node or edge in is regarded as the possible value range of the process, and the state space is defined, where represents the number of states; for the elements in , an additional one-dimensional identifier is added to distinguish whether to continue to retain, upgrade or transfer to a higher level of trust in the subsequent stage, so as to accurately measure the state evolution process when modeling.

[0040] Step 402, map the observable state to input, and perform state modeling and transition analysis on the nodes and edges of the knowledge graph set of observable states, and introduce cross-department interaction factors and policy time effectiveness factors to modify the transition probability of each stage; Specifically, for the multi-stage characteristics of human resources business, the transition matrix is defined, where represents the state transition matrix of the stage, ​Indicates the total number of stages; each Mapping Status Towards state The stage transition probability of In order to more accurately reflect the timeliness impact of cross-departmental interactions and policy updates, Incorporating cross-departmental interaction factors and policy timeliness factor ; make Indicates in From the state Transfer to state The basic probability of the final transition probability is corrected using the following equations: , in represents the corrected transition probability, represents the cross-sector interaction factor, represents the policy time factor; if multiple departments activate interactions on the same entity, then The value increases; if the policy update frequency increases, then Increase; both lead to a corresponding increase in the intensity of the transition probability, which is used to characterize the dynamic evolution brought about by cross-sector data coupling and policy effects; Step 403: State prediction and graph structure update: the graph structure is updated immediately for parts with high confidence transfer, while parts below the threshold remain in their original state. Based on the state space S constructed in step 401 and the multi-stage transfer matrix in step 402, for any moment The state distribution vector is recorded as , then in the stage The state iteration process can be recorded as: , in, It means that the multi-stage transfer matrices are sequentially multiplied to achieve the cumulative effect of multiple iterations from the initial state to the final moment; Represents the distribution results of each state at the next moment; The iteratively obtained state distribution is mapped back to the knowledge graph entities and relationships. If the state distribution of some nodes or edges exceeds the specified threshold, it is marked as a high-confidence conversion and the graph structure is updated immediately. If the distribution value does not reach the threshold, it is retained in the original state and recorded for re-evaluation in the subsequent stage, forming a cyclic iterative update of the graph structure to obtain the updated knowledge graph structure.

[0041] Step 5: Integrate the high-confidence state transition nodes and edges into the updated knowledge graph structure, and combine the multi-hop attention layer to weight and aggregate the adjacent information and frequency of changes of different hops to obtain an optimized knowledge graph structure, highlighting high-frequency cross-department change relationships and suppressing scattered noise. As shown in Figure 3 , the method specifically comprises the following steps: Step 501: Generate an initial reconstructed graph based on the marked high-confidence migration path set and the knowledge graph set of observable states. Specifically, select nodes and edges with transition probabilities higher than a threshold value from the final state distribution of step 4, mark them as a high-confidence migration path set , and record their significant jump information in the transition matrix of each stage ; Integrate the entities and relationships determined as valid and completed in step 3 into a set , and cross-reference them with ; if an entity or relationship appears in both and , it indicates that it has high credibility and transition stability, and is suitable for inclusion in the next knowledge graph reconstruction process. After obtaining and , generate an initial reconstructed graph , where contains all nodes in a high-confidence state or that have passed effective verification, contains high-confidence migration edges from ; New business rules and policy update nodes can be dynamically added.

[0042] Step 502: Attach a time stamp based on the stage to each node and edge in the initial reconstructed graph to record cross-department interaction information; the purpose is to enable the subsequent multi-hop attention layer to determine the frequency of changes and the level of activity under different policy cycles; Step 503: Combine the multi-hop attention layer to weight and aggregate the adjacent information and frequency of changes of different hops to obtain an optimized knowledge graph structure. Specifically, let represent the adjacency matrix of , let represent the extension matrix of -hop adjacent information in , and let represent the number of hops; to aggregate multi-hop information, introduce attention weights , and then aggregate all The multi-hop adjacency information is weighted and superimposed to obtain a multi-hop fusion matrix : , wherein represents an adjacency expansion within a multi-hop range, represents a maximum multi-hop layer number, represents an attention coefficient of the i-th hop, and the above formula is used for accumulation of the multi-hop adjacency information; To determine , an adaptive normalization mechanism is performed based on the variable frequency of the node at different hop numbers and the cross-department interaction features obtained in step S4 to highlight the hop numbers with more frequent and more significant cross-department influence: , , wherein, represents a weighted function of cross-department interaction, represents a variable frequency of the node or relationship within a multi-hop range, represents a balance constant for regulating the contribution of different hop numbers in the overall attention allocation; Through this multi-hop attention mechanism, the high-frequency variable relationship can be strengthened while the error accumulation caused by low-confidence or sporadic variable edges can be suppressed. Step 6: The optimized knowledge graph structure is taken as a benchmark version, and an incremental update queue is established; for newly entered nodes or relationships, the membership degree thereof is adaptively corrected based on the fuzzy rough set method, and the difference between the new and old versions is measured again through the multi-stage Markov process; if the difference is too large, the data source is analyzed in retrospect and the graph structure is adjusted in time;

[0043] Specifically, the following steps are included: Step 601, establishing a benchmark version and initializing an incremental update queue, taking the final graph structure generated in step S5 as a benchmark version , recording the complete topological information of the node set and the edge set as well as the existing business rules, policy references and time stamps; An incremental update queue is established , stores real-time business data and policy change records from various departments since the benchmark version was released; for each data or update request, the department and related authority attributes thereof are retrieved for subsequent fine processing; Step 602, introducing a fuzzy rough set to adaptively correct the membership degree of the newly added relationship to the incremental update queue ​Each new relationship or node in the incremental update is based on the weighted method of cross-department credibility to correct the initial membership degree. Let denote the historical membership degree of the existing object in the graph, the historical membership degree of the existing object in the graph, denote the membership degree given by the department to the new relationship, the membership degree given by the department to the new relationship, the global weight coefficient of the department , and let denote the regulatory factor, which balances the contribution of existing data and new information, and define the new comprehensive membership degree function , wherein denotes the number of departments; denotes any newly introduced relationship or node, denotes the original membership degree in the baseline version, and when there is no historical record, the initial value is a certain constant. In this way, the new relationship in the incremental update can be adaptively iteratively corrected according to the cross-department specific data credibility evaluation. After the membership degree of all new relationships or nodes is corrected, the upper and lower approximation operations of fuzzy rough set are used to preliminarily divide the boundaries of the objects newly entered into the graph. If the membership degree of an object and the threshold interval have a large intersection, the object is marked as a high-uncertainty candidate and waits for dynamic judgment in the subsequent modeling; Step 603, based on the multi-stage Markov process to measure the change rate between versions; For each graph structure version obtained after the completion of each incremental update , let denote the probability distribution vector of the state in version , which is composed of the nodes and edges combined in the previous steps; to measure the difference between version and version , define the version difference function as: , wherein denotes the set of all possible node or edge states, is the steady-state probability of state in version . weak If the value is greater than the set value, it indicates that the new version has a significant difference with the old version in the state distribution level, and then the added data source is traced back and the impact on the whole graph structure is analyzed; Embodiment two The embodiment provides a human social knowledge graph adaptive optimization system based on uncertain information, which comprises: An initial multi-dimensional graph structure construction module is configured to preprocess the obtained multi-department business data, associate each piece of data to a corresponding policy label according to a policy index, and construct an initial multi-dimensional graph structure according to business connection and sharing fields; An observable knowledge graph extraction module is configured to establish a fuzzy set based on the initial multi-dimensional graph structure, adjust the fuzzy set, establish a multi-field weighted membership function of the adjusted fuzzy set, calculate a dynamic fuzzy rough boundary region of the fuzzy set in combination with the multi-field weighted membership function, identify suspicious entities and relationships in the dynamic fuzzy rough boundary region to form a high-risk object set, calculate cross-department similarity, and clean the high-risk object set by counting all cross-department similarity calculation results to extract a knowledge graph set in an observable state; A knowledge graph optimization module is configured to introduce a multi-stage Markov process, update the knowledge graph set in the observable state, integrate nodes and edges with high confidence after state transition into an updated knowledge graph structure, and combine a multi-hop attention layer to weight and aggregate different hop adjacency information and change frequencies to obtain an optimized knowledge graph structure; An adaptive correction module is configured to take the optimized knowledge graph structure as a benchmark version, adaptively correct the membership of added nodes or relationships, and again measure the difference between new and old versions by the multi-stage Markov process, and when the difference is greater than a set value, trace back the data source and adjust the graph structure.

[0044] It should be noted that the specific implementation manner of the human social knowledge graph adaptive optimization system based on uncertain information in the embodiment of the present application is similar to that of the human social knowledge graph adaptive optimization method based on uncertain information in the embodiment of the present application, and specific details are described in the method part. In order to reduce redundancy, this part is not described here.

[0045] Embodiment three The embodiment provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to realize the steps in the human social knowledge graph adaptive optimization method based on uncertain information.

[0046] Embodiment four The embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps in the human social knowledge graph adaptive optimization method based on uncertain information as described above when executing the program.

[0047] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage, etc.) containing computer-usable program code.

[0048] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the functions specified in one or more flows and / or blocks.

[0049] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the functions specified in one or more flows and / or blocks.

[0050] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the functions specified in one or more flows and / or blocks.

[0051] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.

[0052] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. The adaptive optimization method of human and social knowledge graph based on uncertain information is characterized by: include: Pre-process the acquired multi-department business data, associate each piece of data with the corresponding policy label based on the policy index, and build an initial multi-dimensional graph structure based on business connections and shared fields; Based on the initial multi-dimensional graph structure, a fuzzy set is established, the fuzzy set is adjusted, and a multi-domain weighted membership function of the adjusted fuzzy set is established; the dynamic fuzzy rough boundary region of the fuzzy set is calculated by combining the multi-domain weighted membership function, and suspicious entities and relationships in the dynamic fuzzy rough boundary region are identified to form a high-risk object set; Calculate cross-department similarity, count all cross-department similarity calculation results, clean the high-risk object set, and extract the knowledge graph set of observable states; A multi-stage Markov process is introduced to update the knowledge graph set of observable states. Nodes and edges with high confidence after state transfer are integrated into the updated knowledge graph structure. Combined with a multi-hop attention layer, adjacency information of different hop counts and change frequencies are weightedly aggregated to obtain an optimized knowledge graph structure. The optimized knowledge graph structure is used as the baseline version, and the membership of newly added nodes or relationships is adaptively corrected. The difference between the new and old versions is measured again through a multi-stage Markov process. When the difference is greater than the set value, the data source is traced back and the graph structure is adjusted.

2. The method for adaptive optimization of human and social knowledge graph based on uncertain information according to claim 1 is characterized in that: The policy index is used to associate each piece of data with the corresponding policy label, and a preliminary multi-dimensional graph structure is constructed based on business connections and shared fields, including: Establish a human resources and social security business policy index collection; Combined with the human resources and social security business policy index set, a policy mapping function is defined for each data record in the multi-department business data, and the policy label set required to be associated with each data record is determined and labeled; The annotated data records are considered as a set of nodes. Based on the business connection and shared fields between the data, an edge set is defined. The initial multi-dimensional graph structure is constructed by combining the association between different nodes under each policy index. Among them, the policy mapping function Expressed as: , in, Represents a record and policy The correlation evaluation of is the correlation threshold, when Exceed When The collection of policy labels to assign to this record.

3. The method for adaptive optimization of human and social knowledge graph based on uncertain information according to claim 1 is characterized in that: The calculation formula of the dynamic fuzzy rough boundary area of ​​the fuzzy set is: , in, Represents the full set covered by the current knowledge graph, and Represents the adjusted fuzzy set The upper and lower approximations of and Respectively represent the comprehensive membership representation of upper approximation and lower approximation, A dynamic threshold that reflects the frequency of policy updates and the rate of data changes; Or, the formula for calculating cross-department similarity is: , , in, is the cross-department similarity, For the matrix Calculate the cumulative weighted value, , Indicates department With the department In the object The difference measure on Indicates the total number of departments, Indicates that the department The global weight coefficient of Representation object local weight in its sector, Represents the exponential amplification factor for the difference measure.

4. The method for adaptive optimization of human and social knowledge graph based on uncertain information according to claim 1 is characterized in that: The transfer probability correction formula for each stage is: , in, Indicates that From the state Transfer to state The basic probability of represents the corrected transition probability, represents the cross-sector interaction factor, represents the policy time factor; if multiple departments activate interactions on the same entity, then The value increases; if the policy update frequency increases, then Increase.

5. The method for adaptive optimization of human and social knowledge graph based on uncertain information according to claim 1 is characterized in that: The multi-stage Markov process is introduced to update the knowledge graph set of observable states to obtain an updated knowledge graph structure, including: Based on the knowledge graph set of observable states, the spatial state is defined. Combined with the multi-stage Markov process, the cross-departmental interaction factors and policy timeliness factors are incorporated into the state transition matrix of each stage to correct the state transition probability of each stage. The knowledge graph structure is updated by combining the spatial state and the multi-stage state transition probability to obtain the updated knowledge graph structure.

6. The method for adaptive optimization of human and social knowledge graph based on uncertain information according to claim 1 is characterized in that: In the initial reconstructed graph, each node and edge is attached based on Timestamp of the stage , record cross-departmental interaction information.

7. The method for adaptive optimization of human and social knowledge graph based on uncertain information according to claim 1 is characterized in that: The weighted aggregation of the adjacency information of different hops and the change frequency by combining the multi-hop attention layer is: , , in, is the multi-hop fusion matrix, Indicates Adjacency extension within hop range, Indicates the maximum number of multi-hop layers, Indicates the Jump attention coefficient, represents the weighting function for cross-sector interactions, Indicates that a node or relationship is in The frequency of changes within the jump range, represents the equilibrium constant; Or, the adaptive membership correction function for newly added nodes or relationships is: , in, Indicates the number of departments; Represents any newly introduced relationship or node, express The original membership in the baseline version, Indicates department For new relationships Given the membership degree, For the department The global weight coefficient of Represents a regulatory factor.

8. The adaptive optimization system of human and social knowledge graph based on uncertain information is characterized by: include: The initial multi-dimensional graph structure construction module is used to pre-process the acquired multi-department business data, associate each piece of data with the corresponding policy label based on the policy index, and construct the initial multi-dimensional graph structure based on business connections and shared fields; The observable knowledge graph extraction module is used to establish a fuzzy set based on the initial multi-dimensional graph structure, adjust the fuzzy set, and establish a multi-domain weighted membership function for the adjusted fuzzy set; combine the multi-domain weighted membership function to calculate the dynamic fuzzy rough boundary area of ​​the fuzzy set, identify suspicious entities and relationships in the dynamic fuzzy rough boundary area, and form a high-risk object set; Calculate cross-department similarity, count all cross-department similarity calculation results, clean the high-risk object set, and extract the knowledge graph set of observable states; The knowledge graph optimization module is used to introduce a multi-stage Markov process to update the knowledge graph set of observable states, integrate the nodes and edges with high confidence after state transfer into the updated knowledge graph structure, and combine the multi-hop attention layer to perform weighted aggregation of adjacency information of different hop counts and change frequencies to obtain the optimized knowledge graph structure; The adaptive correction module is used to use the optimized knowledge graph structure as the baseline version, perform adaptive membership correction on newly added nodes or relationships, and measure the difference between the new and old versions again through a multi-stage Markov process. When the difference is greater than the set value, the data source is traced back and the graph structure is adjusted.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the adaptive optimization method of human and social knowledge graph based on uncertain information are implemented as described in any one of claims 1 to 7.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the human and social knowledge graph adaptive optimization method based on uncertain information are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Medical consumable abnormal data type-2 fuzzy screening method based on cross-domain collaboration and isomerism map

    CN120974390A

  • Intelligent tax operation and maintenance management system based on digital intelligent base

    CN122155875A

  • A tax intelligent operation and maintenance management system based on a digital base

    CN122155875B