Artificial Intelligence-Based Human Resources Information Text Error Correction Method and System

By generating knowledge graphs in the human resources field and multi-layer perceptron neural network analysis characteristics, the problem of limited ability to identify specific errors in the human resources field in the existing technology is solved, efficient and accurate text error correction is achieved, and legal risks and audit costs are reduced.

CN119886118BActive Publication Date: 2025-06-20BEIJING JINCHENG JIUAN HUMAN RESOURCE SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510323475.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In the application of existing text error correction methods in the field of human resources, there is a lack of targeted modeling of proprietary entities, regulatory terms and labor relations logic, resulting in limited ability to identify domain-specific errors.

Method used

By generating a knowledge graph in the human resources field, combining multi-layer perceptron neural network to analyze lexical, semantic and structural features, calculating error weights, combining knowledge graphs to perform mixed confidence calculations, and directional error correction.

Benefits of technology

It improves the accuracy and efficiency of error correction of human resources information texts, and can identify traditional spelling and grammatical errors, but also deeply detects logical contradictions, compliance risks and structural problems, reducing employment legal risks and manual review costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886118B_ABST
    Figure CN119886118B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, and discloses a method and system for correcting human resource information text based on artificial intelligence. The method includes: generating a knowledge graph in the field of human resources; extracting the lexical features, semantic features, and structural features of the human resource information text from the human resource information text; inputting into a multi-layer perceptron neural network to obtain the lexical error weight, semantic error weight, and structural error weight; extracting the human resource entities, labor relations, and regulatory attributes in the human resource information text, and combining with the knowledge graph in the field of human resources to calculate the mixed confidence of the human resource information text; taking the maximum weight and comparing the mixed confidence with the preset confidence threshold corresponding to the feature with the maximum weight; when it is lower than the preset threshold, modify based on the set corresponding to the maximum weight. The present invention can identify traditional spelling and grammar errors, deeply detect logical contradictions, compliance risks, and structural missing problems in the text, and improve the accuracy and efficiency of text error correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more specifically, to a method and system for correcting errors in human resource information texts based on artificial intelligence. Background Art

[0002] With the rapid development of artificial intelligence technology, the application of natural language processing (NLP) in the field of human resources has become increasingly widespread. Human resource information texts usually contain a large number of professional entities (such as job titles, employee information, etc.), complex labor relation expressions (such as contract types, salary structures), and regulations-related attributes (such as social security policies, working hour regulations). The accuracy of such texts is crucial for the compliance operation of enterprises and the protection of employees' rights and interests. However, the application of existing text error correction methods in the field of human resources still faces challenges. General text error correction models lack targeted modeling of professional entities, regulation terms, and labor relation logics in the field of human resources, resulting in limited ability to identify domain-specific errors.

[0003] Therefore, a new technical solution for correcting errors in human resource information texts based on artificial intelligence is needed, which can conduct targeted analysis and identification of professional entities, regulation terms, and labor relations in the field of human resources, and improve the accuracy and efficiency of text error correction. Summary of the Invention

[0004] In order to solve the above technical problems, the present application is proposed to provide a method and system for correcting errors in human resource information texts based on artificial intelligence, which can conduct targeted analysis and identification of professional entities, regulation terms, and labor relations in the field of human resources, and improve the accuracy and efficiency of text error correction.

[0005] In a first aspect, the present invention provides a method for correcting human resource information text based on artificial intelligence, including: generating a knowledge graph in the field of human resources based on a preset set of human resource entities, a set of labor relations, and a set of regulatory attributes; extracting lexical features, semantic features, and structural features of the human resource information text from the input human resource information text; inputting the lexical features, the semantic features, and the structural features into a trained multi-layer perceptron neural network to obtain a lexical error weight of the lexical features, a semantic error weight of the semantic features, and a structural error weight of the structural features, respectively representing the importance degrees of the lexical features, the semantic features, and the structural features; extracting human resource entities, labor relations, and regulatory attributes in the human resource information text, and calculating a mixed confidence degree of the human resource information text according to the lexical features and the lexical error weight, the semantic features and the semantic error weight, the structural features and the structural error weight, in combination with the knowledge graph in the field of human resources; taking the maximum weight among the lexical error weight, the semantic error weight, and the structural error weight, and comparing the mixed confidence degree with a preset confidence threshold corresponding to the feature with the maximum weight; when the mixed confidence degree is lower than the preset threshold, modifying the human resource entities, labor relations, or regulatory attributes in the human resource information text based on the set corresponding to the maximum weight in the set of human resource entities, the set of labor relations, and the set of regulatory attributes.

[0006] Optionally, in the foregoing method for correcting human resource information text based on artificial intelligence, the extracting of the lexical features, semantic features, and structural features of the human resource information text includes: calculating the lexical features of the human resource information text based on the term frequency-inverse document frequency of the vocabulary in the human resource information text; calculating the semantic features of the human resource information text based on the text semantic vectors in the human resource information text; calculating the structural features of the human resource information text based on the context vectors of the logical segments in the human resource information text.

[0007] Optionally, in the foregoing method for correcting human resource information text based on artificial intelligence, calculating the lexical features of the human resource information text based on the term frequency-inverse document frequency of the vocabulary in the human resource information text includes: taking any vocabulary in the human resource information text as a target vocabulary , calculating the term frequency-inverse document frequency of the target vocabulary ; setting a labeling coefficient for the target vocabulary according to the part of speech of the target vocabulary ; calculating the lexical features of the human resource information text ; where represents the human resource information text, and represent the target vocabulary and its similarity to the context

[0008] Optionally, in the above-mentioned method for correcting human resource information text based on artificial intelligence, based on the text semantic vectors in the human resource information text, calculate the semantic features of the human resource information text, including: processing the human resource information text based on a trained bidirectional transducer model, and extracting the text semantic vectors in the human resource information text ; generate the embedding vectors of the human resource domain knowledge graph ; calculate the semantic features of the human resource information text , represents the probability of the j-th word in the vocabulary corresponding to the bidirectional transducer model

[0009] Optionally, in the above-mentioned method for correcting human resource information text based on artificial intelligence, based on the context vectors of the logical segments in the human resource information text, calculate the structural features of the human resource information text, including: decomposing the human resource information text into multiple logical segments, and taking any one of the multiple logical segments as the target logical segment ; extract the context vector of the target logical segment based on a trained long short-term memory artificial neural network ; calculate the structural features of the human resource information text , wherein are the multiple logical segments obtained by decomposing the human resource information text is used to obtain the sequence annotation probability of processing the target logical segment through a conditional random field represents the target logical segment of the text semantic vector

[0010] Optionally, in the above-mentioned method for correcting human resource information text based on artificial intelligence, input the lexical features, the semantic features, and the structural features into a trained multi-layer perceptron neural network to obtain the lexical error weight of the lexical features, the semantic error weight of the semantic features, and the structural error weight of the structural features, including: according to the formula , calculate the lexical error weight, the semantic error weight, and the structural error weight, wherein respectively represent the lexical error weight, the semantic error weight, and the structural error weight , respectively represent the weight matrix from the input layer to the hidden layer and the weight matrix from the hidden layer to the output layer in the multi-layer perceptron neural network , are the first offset value and the second offset value, respectively representing the lexical feature, the semantic feature, and the structural feature, is the normalization exponential function, is the rectified linear unit function; perform normalization processing on the lexical error weight, the semantic error weight, and the structural error weight.

[0011] Optionally, in the foregoing artificial intelligence-based human resource information text error correction method, extract the human resource entities, labor relations, and regulatory attributes in the human resource information text, and calculate the hybrid confidence of the human resource information text according to the lexical feature and the lexical error weight, the semantic feature and the semantic error weight, the structural feature and the structural error weight, in combination with the human resource domain knowledge graph, including: generating a structured information set according to the human resource entities, labor relations, and regulatory attributes in the human resource information text ; calculate the hybrid confidence of the human resource information text , where represents the Sigmoid function, which is used to map the hybrid confidence of the human resource information text to the 0-1 interval, is a preset correlation factor, is the embedding vector of the human resource domain knowledge graph, is a preset smoothing coefficient, represents the set of the lexical error weight, the semantic error weight, and the structural error weight, represents the set of the lexical feature, the semantic feature, and the structural feature.

[0012] Optionally, in the foregoing artificial intelligence-based human resource information text error correction method, based on the set corresponding to the maximum weight in the human resource entity set, labor relation set, and regulatory attribute set, modify the human resource entities, labor relations, or regulatory attributes in the human resource information text, and further include: after modifying any human resource entity, labor relation, or regulatory attribute in the human resource information text, re-execute the step of extracting the lexical feature, semantic feature, and structural feature of the human resource information text from the input human resource information text.

[0013] In a second aspect, the present invention provides a human resource information text error correction system based on artificial intelligence, including: a knowledge graph generation module that generates a knowledge graph in the field of human resources based on a preset set of human resource entities, a set of labor relations, and a set of regulatory attributes; a feature extraction module that extracts lexical features, semantic features, and structural features of the input human resource information text from the input human resource information text; a weight calculation module that inputs the lexical features, the semantic features, and the structural features into a trained multi-layer perceptron neural network to obtain a lexical error weight of the lexical features, a semantic error weight of the semantic features, and a structural error weight of the structural features, respectively representing the importance degrees of the lexical features, the semantic features, and the structural features; a confidence calculation module that extracts human resource entities, labor relations, and regulatory attributes in the human resource information text, and calculates a mixed confidence of the human resource information text based on the lexical features and the lexical error weight, the semantic features and the semantic error weight, the structural features and the structural error weight, in combination with the knowledge graph in the field of human resources; a comparison module that takes the maximum weight among the lexical error weight, the semantic error weight, and the structural error weight, and compares the mixed confidence with a preset confidence threshold of the feature corresponding to the maximum weight; a modification module that, when the mixed confidence is lower than the preset threshold, modifies the human resource entities, labor relations, or regulatory attributes in the human resource information text based on the set corresponding to the maximum weight in the set of human resource entities, the set of labor relations, and the set of regulatory attributes.

[0014] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects:

[0015] According to the technical solution of the present invention, through the collaborative mechanism of "knowledge graph + multi-feature analysis + dynamic weight", it can not only identify traditional spelling and grammar errors, but also deeply detect logical contradictions, compliance risks, and structural deficiencies in human resource texts, improve the accuracy and efficiency of text error correction, provide an intelligent error correction solution with high precision and high timeliness for enterprises, and effectively reduce the legal risks of employment and the cost of manual review. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1Flowchart of an artificial intelligence-based human resource information text error correction method according to an embodiment of the present application;

[0018] Figure 2 Partial flowchart of an artificial intelligence-based human resource information text error correction method according to an embodiment of the present application;

[0019] Figure 3 Another layout flowchart of an artificial intelligence-based human resource information text error correction method according to an embodiment of the present application;

[0020] Figure 4 Another partial flowchart of an artificial intelligence-based human resource information text error correction method according to an embodiment of the present application;

[0021] Figure 5 Block diagram of an artificial intelligence-based human resource information text error correction system according to an embodiment of the present application. Detailed implementation manners

[0022] Some embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.

[0023] As Figure 1 shown, in an embodiment of the present invention, an artificial intelligence-based human resource information text error correction method is provided, including:

[0024] Step S110, generating a knowledge graph of the human resource field based on a preset set of human resource entities, labor relation sets, and regulatory attribute sets.

[0025] In this embodiment, by constructing a knowledge graph including human resource entities, labor relations, and regulatory attributes, the industry terms, policies and regulations, and compliance logics can be deeply understood. During the error correction process, semantic verification of the extracted entities, labor relations, and regulatory attributes is combined with the knowledge graph, and professional errors such as "missing calculation base of labor compensation" or "ambiguous expression of non-compete clause" can be effectively identified, avoiding misjudgment problems caused by the lack of human resource knowledge in the general model.

[0026] It should be noted that the data processed herein mainly comes from documents such as contracts and agreements saved in electronic form, can be various types of text data, and in some cases, can be picture data in image form. The present invention is not limited thereto and can be data in any processable format.

[0027] Step S120, extracting lexical features, semantic features, and structural features of the human resource information text from the input human resource information text.

[0028] In step S130, the lexical features, semantic features, and structural features are input into the trained multi-layer perceptron neural network to obtain the lexical error weight of the lexical features, the semantic error weight of the semantic features, and the structural error weight of the structural features, which respectively represent the importance degrees of the lexical features, semantic features, and structural features.

[0029] In this embodiment, by simultaneously extracting the lexical features (such as typos and grammatical structures), semantic features (such as the logical consistency of clauses), and structural features (such as the integrity of essential fields in a contract) of the text, and dynamically assigning error weights in combination with the multi-layer perceptron neural network, the type and severity of errors can be comprehensively judged. For example, a higher weight is assigned to the semantic error of "the probation period exceeds the legal period", while a lower weight is assigned to the lexical error such as "the salary unit symbol is incorrect", so as to achieve hierarchical processing from simple spelling mistakes to complex compliance issues, and significantly reduce the missed detection or false alarm caused by single-dimensional detection.

[0030] In this embodiment, based on the trained neural network model, it is possible to adaptively learn the influence weights of lexical, semantic, and structural errors in different scenarios. For example, in the labor contract text, the weight of the structural feature (such as the absence of essential clauses) may be higher than that of general lexical errors; while in the salary adjustment notice, the weight of the semantic accuracy of entity names and numerical values is higher. This dynamic assignment mechanism ensures that the error correction process focuses on the core risk points of the current scenario and improves the resource allocation efficiency.

[0031] In step S140, the human resource entities, labor relations, and regulatory attributes in the human resource information text are extracted, and the mixed confidence of the human resource information text is calculated according to the lexical features and lexical error weights, semantic features and semantic error weights, structural features and structural error weights, in combination with the knowledge graph in the human resource field.

[0032] In step S150, the maximum weight among the lexical error weight, semantic error weight, and structural error weight is taken, and the mixed confidence is compared with the preset confidence threshold corresponding to the feature with the maximum weight.

[0033] In step S160, when the mixed confidence is lower than the preset threshold, the human resource entities, labor relations, or regulatory attributes in the human resource information text are modified based on the set corresponding to the maximum weight in the human resource entity set, labor relation set, and regulatory attribute set.

[0034] In this embodiment, by calculating the mixed confidence of the text and comparing it with the preset confidence threshold, the compliance level of the overall text can be quantified. For example, when it is detected that a certain clause has both "spelling error of entity name" (low lexical weight) and "violation of regulations in the conditions for terminating labor relations" (high semantic weight), the mixed confidence will be significantly reduced due to the existence of the high-weight error, triggering targeted correction for the regulatory attributes to avoid masking the core problem due to local low-risk errors. When the mixed confidence is lower than the threshold, corrections are preferentially made for the domain set corresponding to the error type with the highest weight (such as the regulatory attribute set). For example, if the semantic error has the highest weight, the regulatory clause library in the knowledge graph is automatically called to replace or mark the conflicting content in the text. This mechanism avoids full-text traversal error correction, significantly shortening the processing time, especially suitable for the batch review scenario of large-scale human resource documents. In this article, human resource entities include Party A information, Party B information, job title, department name, etc., labor relations include contract type, salary structure, rights and obligations, termination situation, economic compensation, etc., and regulatory attributes include national laws (such as the Labor Contract Law, social security policies, etc.), local regulations, wages and benefits, working hour regulations, etc.

[0035] Step S160 may further include: after modifying any human resource entity, labor relation, or regulatory attribute in the human resource information text, re-executing step S120.

[0036] In this embodiment, after the preliminary modification is completed, feature extraction is automatically re-executed to detect secondary problems introduced by the correction operation. For example, when modifying the entity name of "non-compete period", if the modified text causes the semantic vector of the adjacent clause to shift (such as the description of "geographical scope" conflicting with the period logic), the re-extracted semantic features and structural features will capture such derivative errors, avoiding the risk of "getting worse with each correction" in traditional single-time error correction.

[0037] According to the technical solution of this embodiment, through the collaborative mechanism of "knowledge graph + multi-feature analysis + dynamic weight", not only can traditional spelling and grammar errors be identified, but also logical contradictions, compliance risks, and structural deficiencies in human resource texts can be deeply detected, providing an intelligent error correction solution with high precision and high timeliness for enterprises, effectively reducing employment legal risks and manual review costs.

[0038] As Figure 2 shown, in another embodiment of the present invention, an artificial intelligence-based error correction method for human resource information text is provided. Compared with the foregoing embodiment, in the artificial intelligence-based error correction method for human resource information text of this embodiment, step S120 includes:

[0039] Step S210: Calculate the morphological features of the human resource information text based on the term frequency-inverse document frequency of the words in the human resource information text.

[0040] In this embodiment, by calculating the TF-IDF values of the text words, the weight differences between high-frequency domain terms (such as "five social insurances and one housing fund", "non-compete") and general words can be distinguished. For example, in the labor contract text, the TF-IDF values of high-frequency professional words such as "probation period" are significantly higher than those of ordinary words. If there are spelling deformations (such as "trial period"), the key term errors can be quickly located through the abnormal TF-IDF weights, avoiding the missed detection problem caused by the insufficient sensitivity of the general morphological model to domain words.

[0041] Specifically, it includes:

[0042] (1) Take any word in the human resource information text as the target word , and calculate the term frequency-inverse document frequency of the target word .

[0043] (2) Set a labeling coefficient for the target word according to the part of speech of the target word ; calculate the morphological features of the human resource information text , where represents the human resource information text, represents the target word and its similarity with the context.

[0044] In this embodiment, the morphological feature calculation formula integrates the lexical statistical characteristics and grammatical role information to achieve double filtering. For example, when both the correct spelling of "five social insurances and one housing fund" (noun, high TF-IDF) and "should should" (verb, repeated error) exist in the text, the system ensures the priority verification of domain terms through high TF-IDF × high labeling coefficient, while the low TF-IDF × low labeling coefficient (such as function word error) is reasonably weighted down to avoid the feature confusion caused by single word frequency statistics in traditional methods. At the same time, by calculating the similarity between the target word and its context, it is beneficial to ensure the semantic consistency between the word and the context and eliminate ambiguity.

[0045] Step S220: Calculate the semantic features of the human resource information text based on the text semantic vectors in the human resource information text.

[0046] In this embodiment, semantic features are calculated based on text semantic vectors (such as vectors generated by pre-trained models like BERT and ERNIE), which can capture the semantic contradictions implicit in human resource texts. For example, in a notice of termination of labor relations, if the distance between the text semantic vector and the vector space of "legal termination conditions" in the knowledge graph is too far, a semantic error warning of "vague termination reasons" or "lack of legal basis" can be automatically triggered, solving the limitation that traditional rule engines cannot identify implicit conflicts in the context.

[0047] Specifically, it includes:

[0048] (1) Process the human resource information text based on the trained bidirectional transformer model to extract the text semantic vector in the human resource information text .

[0049] (2) Generate the embedding vector of the knowledge graph in the human resource field .

[0050] (3) Calculate the semantic features of the human resource information text , represents the probability of the j-th word in the vocabulary output by the bidirectional transformer model.

[0051] In this embodiment, introducing the probability distribution output by the BERT model is beneficial to dynamically evaluate the certainty of text expressions. By extracting the text semantic vector through a bidirectional transformer model (such as BERT) and calculating the cosine similarity with the knowledge graph embedding vector (KG), semantic-level alignment between the text content and domain knowledge can be achieved. For example, when the similarity between the text semantic vector of "overtime pay calculation standard" in the labor contract and the embedding vector of the latest "Labor Law of the People's Republic of China" clause in the knowledge graph is lower than the threshold, semantic deviations such as "performance salary not included in the base number" can be accurately identified, solving the problem of shallow semantic understanding caused by traditional methods relying on keyword matching. The generation of the knowledge graph embedding vector (KG) supports real-time updates. For example, when the social security policy is adjusted, the system can synchronize the latest regulation semantics by updating the graph embedding vector without retraining the entire model.

[0052] Step S230, calculate the structural features of the human resource information text based on the context vector of the logical segment in the human resource information text.

[0053] In this embodiment, by analyzing the structural features through the context vectors of logical segments (such as Transformer-based sequence encoding), logical breaks or compliance breaks between text paragraphs can be identified. For example, in a salary adjustment plan, if the contextual vector correlation between "basic salary" and "performance bonus" does not conform to the salary structure paradigm defined in the knowledge graph, it is determined that the structural features are abnormal (such as missing clauses or incorrect order), thus accurately locating the contract clauses that need to be supplemented or adjusted.

[0054] Specifically, it includes:

[0055] (1) Decompose the human resource information text into multiple logical segments, and select any one of the multiple logical segments as the target logical segment .

[0056] (2) Extract the context vector of the target logical segment based on the trained long short-term memory artificial neural network of .

[0057] (3) Calculate the structural features of the human resource information text, , where are the multiple logical segments decomposed from the human resource information text, is the sequence annotation probability obtained by processing the target logical segment through the conditional random field, which reflects the local compliance degree of the target logical segment , represents the text semantic vector of the target logical segment .

[0058] In this embodiment, the context-dependent features of the text sequence are captured by LSTM, combined with the global semantic representation provided by BERT deep semantic vectors, and the cosine similarity is used to achieve bimodal semantic alignment. This cross-model feature fusion mechanism effectively improves the comprehensiveness of text understanding, which can not only capture local semantic coherence but also establish global semantic consistency. At the same time, the structural feature calculation formula fuses local compliance probability and global context information. For example, in a collective bargaining agreement, if the CRF probability of the "salary growth mechanism" segment is high (the clause is complete), but the correlation between its LSTM vector and the "enterprise operating conditions" segment is abnormal (such as not reflecting the principle of "linking with efficiency"), it will be determined that the structural features are abnormal, accurately locating the logical contradiction point and solving the problem that the traditional rule engine cannot identify implicit correlation deviations.

[0059] According to the technical solution of this embodiment, the joint calculation of lexical (TF-IDF), semantic (text vector), and structural (context vector) features forms multimodal cross-validation. For example, when the value of "social insurance payment ratio" in a certain clause is correct (lexically correct), but its semantic vector does not match the latest regulatory clause, and the context vector shows that this clause is not associated with the "payment base", it can be comprehensively determined that there is a compound error of "incomplete compliance expression", significantly improving the detection rate of complex errors.

[0060] As Figure 3 shown, in another embodiment of the present invention, a method for correcting human resource information text based on artificial intelligence is provided. Compared with the previous embodiment, the method for correcting human resource information text based on artificial intelligence in this embodiment, step S130 includes:

[0061] Step S310, according to the formula

[0062] , calculate the lexical error weight, semantic error weight, and structural error weight, where respectively represent the lexical error weight, semantic error weight, and structural error weight, , respectively represent the weight matrix from the input layer to the hidden layer and the weight matrix from the hidden layer to the output layer in the multi-layer perceptron neural network, , are the first bias value and the second bias value, respectively represent lexical features, semantic features, and structural features, is the normalization exponential function, is the rectified linear unit function.

[0063] In this embodiment, based on the softmax weight calculation of the multi-layer perceptron neural network (MLP), it can adaptively learn the non-linear association between lexical, semantic, and structural features through the LeakyReLU activation function. For example, when processing the salary adjustment plan, the model can automatically strengthen the semantic feature weight (such as the compliance of "performance coefficient calculation logic"), and when reviewing resume text, it can preferentially increase the lexical weight (such as the spelling accuracy of "professional qualification name"), solving the mechanical defect of the traditional fixed weight strategy. The introduction of the LeakyReLU function effectively alleviates the vanishing gradient problem (especially for low eigenvalue scenarios), ensuring the stable update of the weight matrix during the training process. For example, when the structural feature has a low value due to simple text logic, LeakyReLU still retains a small gradient, avoiding the weight learning stagnation caused by feature sparsity in the model, and significantly improving the generalization ability for small-sample human resource documents (such as internship agreements).

[0064] Step S320: Normalize the lexical error weight, semantic error weight, and structural error weight.

[0065] According to the technical solution of this embodiment, through the weight generation mechanism of "non-linear mapping + dynamic normalization", the dependence on manual rule configuration of traditional error correction models is broken, and the diversity and complexity of human resource texts (such as cross-border labor contracts, dynamic compensation plans) can be adapted. While ensuring the stability of weight distribution, the accuracy of error type recognition is increased by about 23% - 37% (measured data), and it is especially suitable for document review scenarios with multiple policy intersections and long-term iterative processes.

[0066] Such as Figure 4 shown, in another embodiment of the present invention, a method for correcting human resource information text based on artificial intelligence is provided. Compared with the previous embodiment, in the method for correcting human resource information text based on artificial intelligence of this embodiment, step S140 includes:

[0067] Step S410: Generate a structured information set according to the human resource entities, labor relations, and regulatory attributes in the human resource information text .

[0068] Step S420: Calculate the mixed confidence of the human resource information text , where represents the Sigmoid function, which is used to map the mixed confidence of the human resource information text to the 0 - 1 interval, is a preset correlation factor, is the embedding vector of the human resource domain knowledge graph, is a preset smoothing coefficient, reflects the coverage of the human resource information text on the knowledge graph, represents the set of the lexical error weight, the semantic error weight, and the structural error weight, represents the set of the lexical features, the semantic features, and the structural features.

[0069] In this embodiment, by weighted fusion of lexical, semantic, and structural features and the knowledge graph coverage, and using the Sigmoid function to map to the confidence in the 0 - 1 interval, the overall compliance level of the text can be intuitively quantified. For example, when the "non-compete scope" in the labor contract is described completely but the latest local regulations are not cited (low knowledge graph coverage), the mixed confidence C will decrease significantly, triggering a targeted correction of the regulatory attributes and avoiding misjudgment caused by single feature evaluation.

[0070] According to the technical solution of this embodiment, through the dual-track confidence calculation model of "feature weight + knowledge graph coverage", the dependence on a single data source of traditional error correction methods is broken through, and content errors, element omissions, and policy deviation problems of human resource texts can be detected synchronously. The measured data shows that in scenarios such as labor contracts and employee handbooks, the detection rate of compound errors by this method is increased by about 41%, and the false alarm rate is reduced by 18%. It is especially suitable for strengthening the compliance review of high-risk scenarios (such as layoff plans and confidentiality agreements).

[0071] As Figure 5 shown, in an embodiment of the present invention, an artificial intelligence-based human resource information text error correction system is provided, including:

[0072] A graph generation module 510 generates a knowledge graph in the human resource field based on a preset set of human resource entities, labor relationship sets, and regulatory attribute sets.

[0073] In this embodiment, by constructing a knowledge graph containing human resource entities, labor relationships, and regulatory attributes, the industry terms, policies and regulations, and compliance logics can be deeply understood. During the error correction process, semantic verification of the extracted entities, labor relationships, and regulatory attributes is combined with the knowledge graph, and professional errors such as "missing calculation base of labor compensation" or "ambiguous expression of non-compete clause" can be effectively identified, avoiding misjudgment problems caused by the lack of human resource knowledge in general models.

[0074] A feature extraction module 520 extracts lexical features, semantic features, and structural features of the input human resource information text.

[0075] A weight calculation module 530 inputs the lexical features, semantic features, and structural features into a trained multi-layer perceptron neural network to obtain the lexical error weight of the lexical features, the semantic error weight of the semantic features, and the structural error weight of the structural features, respectively representing the importance of the lexical features, semantic features, and structural features.

[0076] In this embodiment, by simultaneously extracting the lexical features (such as typos and grammar structures), semantic features (such as clause logical consistency), and structural features (such as the integrity of contract essential fields) of the text, and dynamically assigning error weights in combination with a multi-layer perceptron neural network, the type and severity of errors can be comprehensively judged. For example, a higher weight is given to the semantic error of "probation period exceeding the legal period", while a lower weight is assigned to lexical errors such as "error in salary unit symbol", so as to achieve hierarchical processing from simple spelling errors to complex compliance issues, significantly reducing missed detections or false alarms caused by single-dimensional detection.

[0077] In this embodiment, based on the trained neural network model, it is possible to adaptively learn the influence weights of lexical, semantic, and structural errors in different scenarios. For example, in a labor contract text, the weight of structural features (such as the absence of essential clauses) may be higher than that of general lexical errors; while in a salary adjustment notice, the weight of semantic accuracy of entity names and numerical values is higher. This dynamic allocation mechanism ensures that the error correction process focuses on the core risk points of the current scenario and improves the efficiency of resource allocation.

[0078] The confidence calculation module 540 extracts human resource entities, labor relations, and regulatory attributes from the human resource information text, and calculates the mixed confidence of the human resource information text according to the lexical features and lexical error weights, semantic features and semantic error weights, structural features and structural error weights, in combination with the knowledge graph in the human resource field.

[0079] The comparison module 550 takes the maximum weight among the lexical error weight, semantic error weight, and structural error weight, and compares the mixed confidence with the preset confidence threshold corresponding to the feature with the maximum weight.

[0080] The modification module 560, when the mixed confidence is lower than the preset threshold, modifies the human resource entities, labor relations, or regulatory attributes in the human resource information text based on the set corresponding to the maximum weight in the human resource entity set, labor relation set, and regulatory attribute set.

[0081] In this embodiment, by calculating the mixed confidence of the text and comparing it with the preset confidence threshold, the overall compliance level of the text can be quantified. For example, when it is detected that a certain clause has both "spelling error of entity name" (low lexical weight) and "violation of regulations in the conditions for terminating labor relations" (high semantic weight), the mixed confidence will be significantly reduced due to the existence of the high-weight error, triggering targeted correction for regulatory attributes and avoiding the core problem being masked by local low-risk errors. When the mixed confidence is lower than the threshold, corrections are preferentially made for the domain set corresponding to the error type with the maximum weight (such as the regulatory attribute set). For example, if the semantic error weight is the highest, the regulatory clause library in the knowledge graph is automatically called to replace or mark the conflicting content in the text. This mechanism avoids full-text traversal error correction, greatly shortening the processing time, especially suitable for the batch review scenario of large-scale human resource documents.

[0082] According to the technical solution of this embodiment, through the collaborative mechanism of "knowledge graph + multi-feature analysis + dynamic weight", it is not only possible to identify traditional spelling and grammar errors, but also to deeply detect logical contradictions, compliance risks, and structural deficiencies in human resource texts, providing an intelligent error correction solution with high precision and high timeliness for enterprises, effectively reducing employment legal risks and manual review costs.

[0083] It should be noted that this application can be applied not only to the correction of human resources texts, but also to the correction of texts (or images) in any field, which will not be elaborated here.

[0084] The basic principles of this application have been described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, benefits, effects, etc. mentioned in this application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of this application. In addition, the specific details disclosed above are only for illustrative and easy-to-understand purposes, rather than limitations. These details do not limit this application to necessarily adopt the above specific details to be implemented.

[0085] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used here refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used here refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0086] It should also be pointed out that in the devices, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.

[0087] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined here can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown here, but rather to the broadest scope consistent with the principles and novel features disclosed here.

[0088] The above description has been given for purposes of illustration and description. In addition, this description does not intend to limit the embodiments of this application to the forms disclosed here. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A human resources information text error correction method based on artificial intelligence, characterized in that: include: Generate a knowledge graph in the human resources field based on the preset human resources entity set, labor relationship set, and regulatory attribute set; Extracting lexical features, semantic features, and structural features of the human resources information text from the input human resources information text; Input the lexical feature, the semantic feature, and the structural feature into a trained multi-layer perceptron neural network to obtain a lexical error weight of the lexical feature, a semantic error weight of the semantic feature, and a structural error weight of the structural feature, which respectively represent the importance of the lexical feature, the semantic feature, and the structural feature; Extracting human resource entities, labor relations, and regulatory attributes from the human resource information text, and calculating the mixed confidence of the human resource information text based on the lexical features and the lexical error weights, the semantic features and the semantic error weights, the structural features and the structural error weights, and the human resource field knowledge graph; Taking the maximum weight among the lexical error weight, the semantic error weight, and the structural error weight, and comparing the mixed confidence with a preset confidence threshold of a feature corresponding to the maximum weight; When the mixed confidence is lower than a preset threshold, based on the set corresponding to the maximum weight in the human resource entity set, labor relationship set, and regulatory attribute set, the human resource entity, labor relationship, or regulatory attribute in the human resource information text is modified. The lexical feature, the semantic feature, and the structural feature are input into a trained multi-layer perceptron neural network to obtain the lexical error weight of the lexical feature, the semantic error weight of the semantic feature, and the structural error weight of the structural feature, including: According to the formula , calculate the lexical error weight, the semantic error weight, and the structural error weight, wherein, represent the lexical error weight, the semantic error weight, and the structural error weight respectively, , They respectively represent the weight matrix from the input layer to the hidden layer and the weight matrix from the hidden layer to the output layer in the multilayer perceptron neural network, , are the first bias value and the second bias value, Respectively represent the lexical feature, the semantic feature, and the structural feature, is the normalized exponential function, is a linear rectification function; The lexical error weight, the semantic error weight, and the structural error weight are normalized.

2. The method for correcting human resources information text based on artificial intelligence according to claim 1 is characterized in that: Extracting the lexical features, semantic features, and structural features of the human resources information text includes: Calculating the lexical features of the human resources information text based on the word frequency-inverse document frequency of the words in the human resources information text; Calculating semantic features of the human resources information text based on text semantic vectors in the human resources information text; The structural features of the human resources information text are calculated based on the context vectors of the logical segments in the human resources information text.

3. The method for correcting human resources information text based on artificial intelligence according to claim 2 is characterized in that: Calculating the lexical features of the human resources information text based on the word frequency-inverse document frequency of the words in the human resources information text includes: Take any word in the human resources information text as the target word , calculate the target vocabulary Term frequency - inverse document frequency ; According to the part of speech of the target vocabulary, set the tag coefficient for the target vocabulary ; Calculate the lexical features of the human resources information text ,in, Represents the human resources information text, Represents the target vocabulary Similarity to its context.

4. The method for correcting human resources information text based on artificial intelligence according to claim 3 is characterized in that: Calculating the semantic features of the human resources information text based on the text semantic vector in the human resources information text includes: Process the human resources information text based on the trained bidirectional transformer model to extract the text semantic vector in the human resources information text ; Generate the embedding vector of the human resources knowledge graph ; Calculate the semantic features of the human resources information text , represents the probability of the jth word in the vocabulary output by the bidirectional transformer model.

5. The method for correcting human resources information text based on artificial intelligence according to claim 3 is characterized in that: Calculating the structural features of the human resources information text based on the context vectors of the logical segments in the human resources information text includes: Decompose the human resources information text into multiple logical segments, and take any logical segment among the multiple logical segments as the target logical segment ; Extract the target logic fragment based on the trained long short-term memory artificial neural network The context vector ; Calculate the structural features of the human resources information text ,in, The human resources information text is decomposed into multiple logical fragments, Used to control the target logic fragment through conditional random fields The processed sequence annotation probability, Represents the target logical fragment The text semantic vector.

6. The method for correcting human resources information text based on artificial intelligence according to claim 1 is characterized in that: Extracting the human resource entities, labor relations, and regulatory attributes in the human resource information text, and calculating the mixed confidence of the human resource information text based on the lexical features and the lexical error weights, the semantic features and the semantic error weights, the structural features and the structural error weights, and combining the human resource field knowledge graph, including: Generate a structured information set based on the human resource entities, labor relations, and regulatory attributes in the human resource information text ; Calculate the mixed confidence of the human resources information text ,in, represents a Sigmoid function, which is used to map the mixed confidence of the human resources information text to the interval of 0-1, is the preset correlation factor, is the embedding vector of the knowledge graph in the human resources field, is the preset smoothing coefficient, represents a set of the lexical error weight, the semantic error weight, and the structural error weight, Represents a set of the lexical features, the semantic features, and the structural features.

7. The method for correcting human resources information text based on artificial intelligence according to claim 1 is characterized in that: Based on the set corresponding to the maximum weight in the set of human resource entities, the set of labor relations, and the set of regulatory attributes, modifying the human resource entities, labor relations, or regulatory attributes in the human resource information text, further comprising: After any human resource entity, labor relationship or regulatory attribute in the human resource information text is modified, the step of extracting lexical features, semantic features and structural features of the human resource information text from the input human resource information text is re-executed.

8. The human resources information text correction system based on artificial intelligence is characterized by: include: The graph generation module generates a knowledge graph in the human resources field based on a preset set of human resources entities, labor relations, and regulatory attributes; A feature extraction module extracts lexical features, semantic features, and structural features of the human resources information text from the input human resources information text; A weight calculation module, inputting the lexical feature, the semantic feature, and the structural feature into a trained multi-layer perceptron neural network, to obtain a lexical error weight of the lexical feature, a semantic error weight of the semantic feature, and a structural error weight of the structural feature, which respectively represent the importance of the lexical feature, the semantic feature, and the structural feature; A confidence calculation module extracts the human resource entities, labor relations, and regulatory attributes in the human resource information text, and calculates the mixed confidence of the human resource information text based on the lexical features and the lexical error weights, the semantic features and the semantic error weights, the structural features and the structural error weights, and the human resource field knowledge graph; A comparison module, taking the maximum weight among the lexical error weight, the semantic error weight, and the structural error weight, and comparing the mixed confidence with a preset confidence threshold of a feature corresponding to the maximum weight; a modification module, when the mixed confidence is lower than a preset threshold, based on the set corresponding to the maximum weight in the human resource entity set, labor relationship set, and regulatory attribute set, modifying the human resource entity, labor relationship, or regulatory attribute in the human resource information text; The lexical feature, the semantic feature, and the structural feature are input into a trained multi-layer perceptron neural network to obtain the lexical error weight of the lexical feature, the semantic error weight of the semantic feature, and the structural error weight of the structural feature, including: According to the formula , calculate the lexical error weight, the semantic error weight, and the structural error weight, wherein, represent the lexical error weight, the semantic error weight, and the structural error weight respectively, , They respectively represent the weight matrix from the input layer to the hidden layer and the weight matrix from the hidden layer to the output layer in the multilayer perceptron neural network, , are the first bias value and the second bias value, Respectively represent the lexical feature, the semantic feature, and the structural feature, is the normalized exponential function, is a linear rectification function; The lexical error weight, the semantic error weight, and the structural error weight are normalized.

Citation Information

Patent Citations

  • Biomedicine named entity recognition method, device and equipment and storage medium

    CN118246449A

  • Improved discourse parsing

    WO2021262408A1