Social security digital employee data processing and model building method based on cloud edge collaboration
Through the cloud-edge collaborative social security digital employee data processing method, an encoder-decoder neural network architecture is constructed to solve the problems of incomplete data collection, inaccurate semantic understanding and insufficient compliance with policy rules in social security data processing. It realizes the efficient, accurate and consistent answer generation of the social security question-and-answer model, and improves the intelligence level of social security services.
Patent Information
- Application Number
- CN202510914444.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing social security data processing technologies have problems with incomplete data collection, inaccurate semantic understanding, low efficiency in long text processing, and lack of strict compliance with policy rules, resulting in the inability of intelligent social security services to meet complex needs. Conventional methods find it difficult to model the long-term dependencies and hierarchical structures of social security policy texts, and the generated answers are inaccurate and do not meet policy requirements.
A social security digital employee data processing method based on cloud-edge collaboration is adopted. Data is collected collaboratively through the cloud platform and edge computing nodes, and an encoder-decoder neural network architecture is constructed. Combined with multi-granularity semantic fusion, relative position encoding, hierarchical sparse attention and differentiable logic reasoning units, a social security question-and-answer model is generated to ensure the accuracy of the answers and policy consistency.
It significantly improves the efficiency and accuracy of social security data processing, ensures the policy logic consistency of answers, reduces operating costs, optimizes resource allocation, and improves service quality and decision-making support capabilities.
Smart Images

Figure CN120764693A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital employee data processing, and in particular to a social security digital employee data processing and model building method based on cloud-edge collaboration. Background Art
[0002] With the development of society and changes in the demographic structure, social security services face the challenges of a growing number of insured persons, constantly updating policies, and increasingly complex business processes. Traditional social security business processing relies primarily on manual operations, which are inefficient, prone to human error, and fail to meet the public's demand for immediate and accurate social security information access and business processing. Against this backdrop, leveraging advanced information technology to enhance the intelligence of social security services has become a key development direction. However, existing technologies have numerous shortcomings in processing social security data and generating accurate answers, such as incomplete data collection, inaccurate semantic understanding, inefficient processing of long texts, and a lack of strict adherence to policy regulations. These shortcomings limit the development of intelligent social security services and prevent them from fully meeting the complex needs of social security services.
[0003] The problems with existing methods are as follows: Conventional word segmentation methods tend to chop up professional terms when processing them, resulting in semantic fragmentation, which affects the accuracy of understanding professional terms and long-tail expressions in the social security field. In addition, there is a lack of effective mechanisms to dynamically adjust the weights of embeddings of different granularities, making it difficult to fully integrate multi-granularity semantic information such as characters, words, and terms, limiting the model's ability to model complex semantics; social security policy texts have long-range dependencies across paragraphs and a hierarchical structure of chapters, sections, and clauses, but the conventional absolute position encoding in existing technologies is difficult to model long-range position correlations, and the fixed-pattern sparse attention strategy cannot adapt to the hierarchical logical chain of policies, making it easy for the model to lose important semantic information and logical structure when processing social security policy texts, resulting in inaccurate and incomplete understanding; when generating social security answers, the decoders of conventional technologies usually lack a hard constraint mechanism for policy rules and rely only on soft rule injection, which easily produces logical contradictions and answers that do not meet policy requirements, seriously affecting the practicality and reliability of the model and failing to meet the strict requirements of social security business for answer accuracy and policy consistency.
[0004] Therefore, the present invention proposes a social security digital employee data processing and model building method based on cloud-edge collaboration to solve the above problems. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention develops a social security digital employee data processing and model building method based on cloud-edge collaboration. The social security digital employee question-and-answer model constructed by the present invention can meet the strict requirements of social security business on answer accuracy and policy consistency, and is practical and reliable.
[0006] The technical scheme for solving the technical problem of the present application is a social security digital employee data processing and model building method based on cloud edge collaboration, comprising the following steps: S1, collect multi-source social security data through cloud platform and edge computing node collaboration, synchronize the collected data to the cloud structured data table based on API interface timing increment, fill in the missing insured years field with the mean value of the data of adjacent areas of the same risk type; real-time crawling of the updated policy interpretation files on the government website is realized through distributed crawler, and the policy text is parsed according to the "chapter-section-clause" level; S2, based on the collected multi-source social security data and policy text, construct training samples, the sample generation logic adopts the way of question and answer extraction, analyze the triplets of "policy clause-applicable condition-execution standard", generate standard question and answer templates, and divide the training samples into training set, validation set and test set according to the proportion; S3, a neural network architecture based on encoder and decoder is used to build a social security digital employee question and answer model, the model is of encoder-decoder structure, the encoder includes input layer, position coding layer, attention layer and output layer, the decoder includes input layer, logic enhancement layer, normalization layer and output layer, the training set is input into the model and processed by the encoder and decoder to generate answer sequence; S4, define the loss function to train the model and update the model parameters, and the model is trained for multiple iterations until the preset iteration stopping condition is met, then the model training is ended, and the model with the best performance is selected as the optimal model through the validation set; S5, input the test set into the optimal model to generate answer sequence for social security digital employee question and answer.
[0007] The specific process of the input layer of the encoder performing multi-granularity semantic fusion operation is as follows: The input text is subjected to character-level, word-level and term-level embedding fusion, to represent the input text, the input text contains questions and answers, and is represented as [CLS] question [SEP] answer [SEP], [CLS] represents the sequence start token, and [SEP] represents the separation token, and the specific steps are as follows: 1) Calculate the character-level embedding: Based on the input text, each character is mapped to a dense vector through a character embedding function to generate a character-level embedding matrix; 2) Calculate the word-level embedding: The input text is subjected to word segmentation processing to obtain a word sequence, and then each word sequence is mapped to a dense vector through a word embedding function to generate a word-level embedding matrix; 3) Calculate the term-level embedding: Based on the social security terminology database, the term search function is used to identify professional terms in the input text, and the corresponding term embedding vectors are directly retrieved from the terminology database to generate a term-level embedding matrix; 4) Calculate the gate vector: The character-level embedding matrix, word-level embedding matrix, and term-level embedding matrix corresponding to the same position are concatenated, and then the gate vector is calculated through linear transformation and Sigmoid activation function; 5) Compute fusion embedding: The word-level embedding matrix and the term-level embedding matrix are projected into the same dimensional space as the character-level embedding matrix. Then, the projected word-level embedding, projected term-level embedding, and character-level embedding corresponding to the same position are concatenated, and the concatenated vector is weighted using the gating vector of the position to obtain the fused embedding vector of the position. , Indicates the The fused embedding vector of each position.
[0008] The specific process of the encoder's position encoding layer performing relative position encoding operations is as follows: The semantic-aware relative position encoding mechanism is adopted. The specific steps are as follows: 1) Calculate the semantic decay coefficient: The fused embedding vectors of the two positions are concatenated, and then the semantic attenuation coefficient is calculated through linear transformation and activation function to reflect the degree of semantic association between the two positions. When the semantic association is high, the attenuation coefficient approaches zero. 2) Calculate relative position offset: Multiply the semantic attenuation coefficient between two positions by their relative position distance and take the negative value to obtain the relative position offset value, and then obtain the relative position offset matrix; 3) Generate semantic-aware attention scores: The query matrix and key matrix of the input sequence are obtained through linear transformation, and the standard attention score matrix is calculated. Then, the calculated relative position bias matrix is added to the standard attention score matrix to obtain the relative position attention score matrix based on semantic perception.
[0009] The specific process of the encoder's attention layer performing the layered coefficient attention enhancement operation is as follows: Based on the policy structure, a dynamic attention mask is constructed to reduce the computational cost and keep the policy logic chain intact. The specific steps are as follows: 1) Hierarchical boundary analysis: The hierarchical boundaries of the policy text are automatically identified by matching the label body of the policy text with a regular expression. The label system is represented as "{Chapter Number}Chapter{Section Number}Section{Clause Number}Article". The regular expression that matches the label system is "(\d+)Chapter(\d+)Section(\d+)Article". (\d+) is the regular expression pattern, indicating matching one or more numeric characters. (\d) matches the numbers 0-9, and (+) indicates one or more repetitions. 2) Structured Attention Mask Construction: Based on the hierarchical structure of the policy text, a binary mask matrix is generated for each level. The chapter-level mask stipulates that only positions belonging to the same chapter can establish attention connections, the section-level mask stipulates that only positions belonging to the same section can establish connections, and the clause-level mask stipulates that only positions belonging to the same clause can establish connections. Then, the three-level masks are added together to obtain a structured attention mask. In the attention calculation, only positions within the same level are allowed to establish connections. The calculation formula is as follows: , in, represents a structured attention mask; Represents chapter-level mask, the mask logic rule is , Indicates chapter-level mask Rank Elements of the column; Represents a section-level mask, and the mask logic rule is , Indicates the clause level mask Rank Elements of the column; Represents a clause-level mask, and the mask logic rule is , Indicates the clause level mask Rank Elements of the column; 3) Sparse attention calculation: The structured attention mask is added to the scaled standard attention score matrix, and then the sparse attention weight matrix is calculated using the Softmax function. The sparse attention weight matrix has non-zero weights only between the same-level position pairs allowed by the structured attention mask.
[0010] The specific process of the encoder's output layer performing feature output operations is as follows: Integrate the outputs of relative position attention and hierarchical sparse attention, aggregate position information and hierarchical structure information, and obtain the encoder output matrix. The specific steps are as follows: 1) Attention mechanism fusion: The relative position attention score matrix and the hierarchical sparse attention weight matrix are concatenated along the feature dimension, and the attention gating vector is calculated through linear transformation and Sigmoid activation function. The attention gating vector is then used to perform weighted summation on the relative position attention score matrix and the hierarchical sparse attention weight matrix to obtain the fused attention weight matrix. 2) Contextual representation generation: The attention value matrix is obtained by linear transformation of the fused embedding vector and the projection weight matrix. Then, the attention value matrix is weighted summed using the fused attention weight matrix and added to the fused embedding vector. Then, the addition result is layer normalized to obtain the context representation matrix output by the encoder.
[0011] The specific process of the decoder's input layer performing feature input operations is as follows: The context representation output by the encoder is fused with the target sequence embedding to semantically associate the input with the target sequence. The specific steps are as follows: 1) Target sequence embedding: Use the character embedding function to map each character of the target sequence into a dense vector to generate a character-level embedding matrix. Then, use the position encoding function to calculate the position embedding and superimpose it on the character-level embedding to construct the decoder input embedding matrix. 2) Masked self-attention calculation: The decoder input embedding matrix is processed using a masked multi-head self-attention mechanism, which restricts each position to only focus on itself and the previous position through the lower triangular mask matrix. 3) Encoder-Decoder Attention Fusion: The attention fusion operation is performed based on the standard multi-head attention mechanism. The standard multi-head attention mechanism uses the masked self-attention hidden state as the query matrix, the context representation of the encoder output as the key matrix and the value matrix, calculates the cross-attention weight, and fuses the encoder and decoder information through the weighted summation method to generate the attention hidden state matrix.
[0012] The specific process of the decoder's logic enhancement layer performing logic enhancement dynamic constraint operations is as follows: Hard constraints are injected through the differentiable logic reasoning unit to make the output conform to the policy logic. The specific steps are as follows: 1) Construction of logical rule base: A set of logical rules is predefined from the policy text to form a logical rule base. Each rule consists of a predicate logic expression about the input sequence and a consequent predicate logic expression about the output sequence. The two are in a logical implication relationship. The logical rule base is expressed as follows: , in, Represents a logic rule base, including rules; Indicates the Rule identifier; Represents the rule antecedent, which is about the input sequence Predicate logic expressions; Represents the rule postcondition, which is about the output sequence Predicate logic expressions; Represents logical implication relationship; The encoder is fed with an input sequence, i.e., policy text; Output sequence for the decoder, i.e. answer text; 2) Logical attention calculation: Based on the hidden state of the decoder at the current time step, the attention weight of each rule in the logical rule base at that time step is calculated through linear transformation, hyperbolic tangent function and dot product operation of the rule attention vector. This can be used to dynamically filter the subset of rules that are most relevant to the current generated context. 3) Differentiable logic constraint generation: The encoder's context representation is processed using the antecedent function to calculate the probability of satisfying the antecedent of each rule in the logical rule base. This probability is then element-wise multiplied by the corresponding rule's consequent embedding vector. The weighted sum of these results is then performed using the rule attention weight calculated at the current time step to obtain the logical constraint vector for that time step, thereby converting the symbolic logical rules into a differentiable constraint representation. 4) Gated residual enhancement: The hidden state of the decoder at the current time step is concatenated with the logic constraint vector, and the residual enhancement gating vector is calculated through linear transformation and Sigmoid activation function. Then, the logic constraint vector is projected into the latent space and weighted with the residual enhancement gating vector. The weighted constraint vector is then added to the original hidden state as the residual to obtain the logic enhancement hidden state matrix.
[0013] 8. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 7 is characterized in that the specific process of the layer normalization processing operation performed by the normalization layer of the encoder is as follows: Perform layer normalization on the logistically enhanced hidden state matrix to adjust the distribution of the hidden state, stabilize the training process, and improve the gradient flow, generating a layer-normalized hidden state matrix.
[0014] The specific process of the decoder's output layer performing the answer sequence output operation is as follows: The normalized hidden state matrix is mapped to the vocabulary space to generate an unnormalized score vector, which is then converted into a word probability distribution by a Softmax function, as follows: 1) Score vector calculation: The layer normalized hidden state vector is linearly transformed by the output projection weight matrix and bias vector to map the hidden state to the vocabulary space and generate an unnormalized score vector; 2) Probability distribution calculation: The unnormalized score vector is processed by the Softmax function to calculate the word probability distribution at the current time step, representing the probability of each candidate word being generated given the input sequence and the generated sequence, thereby guiding the sequence output. The calculation formula of the word probability distribution is as follows: , where, represents the probability distribution of the current word given the input sequence and the generated sequence ; represents the word to be generated at time step ; represents the word sequence generated before time step ; represents the encoder input sequence; 3) Sequence generation: a) If in the training phase: The cross-entropy loss is calculated based on the word probability distribution and the true label, and the model parameters are optimized through the cross-entropy loss function. The calculation formula of the cross-entropy loss function is as follows: , where, represents the true label at time step ; represents the target sequence length; represents the cross-entropy loss; represents the probability of the model predicting the true word given the input and the true historical sequence ; represents the true historical sequence before time step ; represents the logarithmic function, with the default base being 10; b) If in the inference phase: The sequence is generated step by step by selecting the word with the highest probability through the autoregressive method until the end token [SEP] is output, completing the generation of the answer sequence. The calculation formula is as follows: , in, This is the maximum index operation.
[0015] The model training process is as follows: S4.1. Define the loss function: Define a multi-task loss function, including weighted cross-entropy loss, logical rule loss, and boundary alignment loss. Use cross-entropy loss to optimize generation quality, use logical rule loss to ensure that the output meets the hard logical constraints of the policy rule base, and use boundary alignment loss to force key policy values to stay within a preset reasonable range. Combined with the multi-task loss function, improve the semantic accuracy and logical consistency of the answer. The logical rule loss is defined based on each rule in the logical rule base. The rule satisfaction score and rule violation score are calculated. The logical loss component of each rule is obtained based on the logarithmic function. The average of all rules is then taken to finally generate the logical rule loss value, which quantifies the degree to which the model output sequence violates the hard logical constraints in the policy logical rule base. Extract the values in the preset key value set from the model output sequence, determine whether they exceed the preset reasonable range, and then force the key policy values to be within the preset reasonable range; S4.2. Define the model parameter update method: An adaptive moment estimation optimizer is used to update model parameters. The initial learning rate is set to 0.0001 and decays to 90% of the original value every 10 training epochs. The gradients of the cross entropy loss, logistic rule loss, and boundary alignment loss in the multi-task loss function are calculated independently, and a gradient clipping strategy is implemented to scale the gradient when the L2 norm exceeds the threshold of 1.0; According to the model update method, the standard error back propagation algorithm is used to calculate the multi-task loss function value through forward propagation, and the gradient of each parameter is obtained through back propagation. After gradient clipping, the Adam optimizer parameter update rule is executed, and this cycle is repeated; S4.3. Define the conditions for stopping iterative training: The main stopping condition of the model is that early stopping is triggered when the multi-task loss function value of the validation set decreases by less than 0.1% for 8 consecutive epochs; The maximum training iteration epoch limit is set to 100 times. The top three model parameter snapshots with the highest validation set accuracy are saved in each iteration epoch. The final deployed model selects the version with the highest policy logic consistency score in the validation set as the optimal model.
[0016] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. The above technical solution has the following advantages or beneficial effects: At the encoder input layer, this paper adopts a three-level embedding fusion approach: character-level, word-level, and term-level. This can avoid the semantic fragmentation problem caused by the shredding of professional terms by conventional word segmentation methods, thereby improving semantic integrity and understanding accuracy. At the same time, the gating mechanism dynamically adjusts the weights of embeddings at different granularities, further enhancing the modeling capabilities of professional terms and long-tail expressions. In response to the long-range dependencies and hierarchical structure of social security policy texts, this paper adopts a semantically aware relative position encoding mechanism and a hierarchical sparse attention mechanism. The former allows highly correlated position pairs to maintain high attention weights even at large distances, while the latter constructs a dynamic attention mask based on the policy structure, reducing computational complexity while maintaining the integrity of the policy logic chain. This effectively improves the model's ability to understand and model social security policy texts. A differentiable logic reasoning unit is used in the decoder to build a logic rule library. Through operations such as logical attention calculation and gated residual enhancement, hard constraints are converted into differentiable constraint representations and injected into the decoding process. This enables the generated answers to strictly follow social security policy rules, avoiding the logical contradictions that are easily caused by the lack of hard constraint mechanisms in conventional decoders, thereby ensuring the policy logic consistency of the answers.
[0017] In summary, the social security digital employee data processing and model building method based on cloud-edge collaboration proposed in the present invention can significantly improve processing efficiency and speed, greatly enhance data accuracy and consistency, effectively reduce operating costs and optimize resource allocation, strengthen risk prevention and control and compliance assurance, improve service quality and decision-making support capabilities, and provide an intelligent, precise, efficient and secure social security digital employee question-and-answer model. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0019] Figure 1 Schematic diagram of the method of the present invention.
[0020] Figure 2 This is a comparison chart of the accuracy of the proposed model and the existing model on the social security question-answering test set.
[0021] Figure 3 A comparison chart of the multi-granularity embedding ablation experimental results.
[0022] Figure 4 This is a comparison chart of the logical consistency effects of the model of the present invention and other models on social security rules.
[0023] Figure 5 This is a comparison chart of the inference speed of the model of the present invention and the existing model. DETAILED DESCRIPTION
[0024] In order to clearly illustrate the technical features of this solution, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.
[0025] Example 1 A method for processing social security digital employee data and building a model based on cloud-edge collaboration includes the following steps: S1. Collaboratively collect multi-source social security data through cloud platforms and edge computing nodes. Synchronize the collected data to cloud-based structured data tables based on API interfaces at regular intervals. Fill missing insurance coverage years fields with the average value of data from adjacent regions with the same insurance type. Use distributed crawlers to capture updated policy interpretation documents from government websites in real time, and parse the policy texts at the "chapter-section-clause" level. The cloud data source is the social security government cloud platform's basic information database for insured persons, payment records, benefits verification database, and policy and regulations database. The edge data source is the on-site business processing records, self-service machine consultation logs, and remote video conversation transcripts collected by the terminal equipment in the social security service hall. The cloud-based structured data tables include details of insured persons’ payments, pension payment records, etc. Edge nodes can use voice recognition modules to convert consultation conversations into text streams, which are then uploaded to the edge data pool after desensitization. The data covers the entire life cycle of five major insurance types: pension, medical, work-related injury, unemployment, and maternity. It includes unstructured text such as the original policy terms, business processing rules, and a knowledge base of frequently asked questions. The data spans no less than 10 years to ensure the continuity of historical policies. S2. Construct training samples based on collected multi-source social security data and policy texts. The sample generation logic uses question-answer pair extraction to parse the "policy clause-applicable conditions-implementation standards" triples, generate standard question-answer templates, and divide the training samples into training, validation, and test sets in proportion. The training sample division strategy is shown in Table 1. For example, the clause "Pension = Basic Pension + Personal Account Pension" is parsed into the question-answer pair [How is the pension calculated? = Basic part plus personal account part].
[0026] Table 1 Training sample division strategy S3. A social security digital employee question-answering model is constructed using a neural network architecture based on an encoder and decoder. The model has an encoder-decoder structure. The encoder includes an input layer, a position encoding layer, an attention layer, and an output layer. The decoder includes an input layer, a logical enhancement layer, a normalization layer, and an output layer. The training set is input into the model and processed by the encoder and decoder to generate an answer sequence. S4. Define the loss function to train the model and update the model parameters. The model is trained for multiple iterations until the preset iteration stop condition is met. The model with the best performance is selected as the optimal model through the validation set. S5. Input the test set into the optimal model, conduct social security digital employee Q&A, and generate an answer sequence.
[0027] In a specific implementation, the specific process of the multi-granularity semantic fusion operation performed by the encoder input layer is as follows: The professional terms in the field of social security are chopped up by word segmentation tools, resulting in semantic fragmentation. Conventional word segmentation methods cannot retain the integrity of terms, affecting the accuracy of semantic understanding. The fusion of character-level, word-level, and term-level embedding can avoid the fragmentation of professional terms and improve semantic integrity. Represents the input text, which contains questions and answers, and is expressed as [CLS] question [SEP] answer [SEP], where [CLS] represents the sequence start marker and [SEP] represents the separation marker. The specific steps are as follows: 1) Compute character-level embeddings: Based on the input text, each character is mapped into a dense vector through the character embedding function to generate a character-level embedding matrix. The calculation formula is as follows: , in, represents the character-level embedding matrix, where each row corresponds to the embedding vector of a character; represents the character embedding function that maps each character into a dense vector; In natural language processing tasks, the sequence start marker [CLS] is used to aggregate global information, and the separation marker [SEP] is used to distinguish the question and answer parts. The question and answer are separated by position embedding and special markers. During model training, the model adopts an encoder-decoder architecture. The encoder processes the entire sequence to learn context representation, and the decoder uses [SEP] to identify the question part and use it as input to generate the answer sequence. 2) Calculate word-level embedding: The input text is segmented to obtain word sequences, and then each word sequence is mapped into a dense vector through the word embedding function to generate a word-level embedding matrix. The calculation formula is as follows: , in, Represents the word-level embedding matrix, where each row corresponds to the embedding vector of a word; Indicates the word segmentation algorithm operation based on the maximum matching method; represents the word embedding function that maps each word into a dense vector; Assume the input text It is the “Provident Fund Balance”; Set the character embedding dimension to 3. The character-level embedding matrix is expressed as , where each row corresponds to the embedding vector of a character, and “公” corresponds to , “Product” corresponds to ; After word segmentation, it is ["provident fund", "balance"]. If the word embedding dimension is set to 3, the word-level embedding matrix is expressed as , where each row corresponds to the embedding vector of a word, and “Provident Fund” corresponds to , "Balance" corresponds to .
[0028] 3) Compute term-level embeddings: Based on the social security terminology database, the term search function is used to identify professional terms in the input text, and the corresponding term embedding vectors are directly retrieved from the terminology database to generate a term-level embedding matrix. The calculation formula is as follows: , in, Represents a social security terminology database, such as the term "pension payment calculation months" and other professional terms, which are automatically extracted from the domain corpus using the TF-IDF (Term-Frequency-Inverse-Document-Frequency) algorithm; A term lookup function that identifies and retrieves term embeddings in text; represents the term-level embedding matrix, where each row corresponds to the embedding vector of a term; Assume that the input text contains the term "deemed payment years", and its embedding vector is directly obtained from the social security term library. Set the word embedding dimension to 3, then the term-level embedding matrix may be ; 4) Calculate the gate vector: The character-level embedding matrix, word-level embedding matrix, and term-level embedding matrix corresponding to the same position are concatenated, and then the gating vector is calculated through linear transformation and Sigmoid activation function. The gating vector can be used to dynamically adjust the contribution weight of embeddings of different granularities in subsequent fusion. The calculation formula is as follows: , in, Indicates in The character-level embedding vector of position , representing the Character embeddings; Indicates in The word-level embedding vector of the position represents the The embedding of each word needs to be aligned to the character position; Indicates the location The term-level embedding vector represents the first The embeddings of terms need to be aligned to character positions; represents the vector concatenation operation, which concatenates the three embedded vectors into a single vector with the concatenation dimension being the sum of the dimensions of the three embedded vectors; is the gating weight matrix, used for linear transformation; is the gate bias vector; is the Sigmoid activation function, and the output range is , represents the weights of embeddings of different granularities; For the The gating vector of the position; 5) Compute fusion embedding: The word-level embedding matrix and the term-level embedding matrix are projected into the same dimensional space as the character-level embedding matrix. Then, the projected word-level embedding, projected term-level embedding, and character-level embedding corresponding to the same position are concatenated, and the concatenated vector is weighted using the gating vector of the position to obtain the fused embedding vector of the position. , Indicates the The fused embedding vector of each position can effectively integrate the semantic information of three granularities: character, word, and term. The calculation formula is as follows: , in, Represents the word-level projection matrix, which projects the word-level embedding vector into a space with the same dimension as the character-level embedding vector, making it easier to perform subsequent fusion operations; Represents the term-level projection matrix, which projects the term-level embedding vector into a space with the same dimension as the character-level embedding vector, which facilitates subsequent fusion operations; is element-wise multiplication; It should be noted that the mapping method of the character embedding function and the word embedding function is implemented using a pre-trained model. For example, the character embedding function uses the pre-trained BERT model to map each character into a dense vector space of fixed dimension. For the character "人", it may be mapped to a 10-dimensional vector, which can capture the semantic characteristics of the character "人"; The terminology library injects domain prior knowledge and combines the gating mechanism to dynamically allocate the weights of word, phrase, and term level embeddings, which can effectively avoid semantic fragmentation caused by word segmentation and improve the modeling capabilities of professional terms and long-tail expressions. In addition, The word embedding vector after the item representation projection, The term embedding vector after the term representation projection, Item representation integrates the splicing vector of multi-granularity information, thereby realizing the dynamic fusion of three-level semantics of characters, words and terms, which can avoid the semantic fragmentation problem caused by the shredding of professional terms.
[0029] In a specific embodiment, the specific process of the relative position encoding operation performed by the position encoding layer of the encoder is as follows: Social security policy texts contain long-range dependencies across paragraphs, such as reimbursement conditions scattered across multiple chapters. Conventional absolute position encoding makes it difficult to model long-range position correlations. Therefore, this paper adopts a semantically aware relative position encoding mechanism, which can ensure that highly correlated position pairs maintain high attention weights even at long distances. The specific steps are as follows: 1) Calculate the semantic decay coefficient: The fused embedding vectors of the two positions are concatenated, and then the semantic attenuation coefficient is calculated through linear transformation and activation function to reflect the degree of semantic association between the two positions. When the semantic association is high, the attenuation coefficient approaches zero. The semantic attenuation coefficient can dynamically adjust the intensity of subsequent position attenuation. The calculation formula is as follows: , in, Indicates the The fused embedding vector of the position represents the Semantic features of each position; Indicates the The fused embedding vector of the position represents the Semantic features of each position; Represents a vector concatenation operation, which concatenates two vectors into a single vector; Represents the weighted semantic attenuation matrix, which is used to calculate the semantic relevance; represents the rectified linear unit activation function, The output range of the rectified linear unit activation function is , ensuring that the attenuation coefficient is non-negative; Indicates the Position and The semantic attenuation coefficient between positions can be used to dynamically adjust the position attenuation strength. Position and When the semantics of the two positions are highly correlated, such as the semantics of "payment years" and "retirement age" are highly correlated, Approaching 0; 2) Calculate relative position offset: Multiply the semantic attenuation coefficient between two positions by their relative position distance and take the negative value to obtain the relative position bias value, and then obtain the relative position bias matrix, which can be used to adjust the correlation between the two positions in the attention calculation. The calculation formula of the relative position bias value is as follows: , in, Indicates the Position and The relative position distance between two positions can be expressed, for example, by the absolute value of the index difference between the two positions; Indicates the Position to The relative position offset value of each position is a scalar; It should be noted that the relative position offset value During the calculation process, The negative sign before the term controls the direction of the correlation, indicating that the farther the distance or the lower the semantic correlation, the stronger the attenuation of the attention weight. That is, when the semantics are highly correlated, approaches 0, and Approaching 0 to reduce the position attenuation effect; 3) Generate semantic-aware attention scores: The query matrix and key matrix of the input sequence are obtained through linear transformation, and the standard attention score matrix is calculated. Then, the calculated relative position bias matrix is added to the standard attention score matrix to obtain the relative position attention score matrix based on semantic perception. The calculation formula is as follows: , in, is the query matrix, which is obtained by linear transformation of the input sequence; is the key matrix, which is obtained by linear transformation of the input sequence; is the bond matrix The transpose of is the relative position bias matrix, is the relative position bias matrix Rank Elements of the column; is the relative position attention score matrix; It should be noted that the fusion embedding vector is defined as , that is, The fused embedding vector of each position is the fused embedding vector No. elements, the fused embedding vector is used as the input sequence, and the query matrix is obtained by two different linear transformation layers. and bond matrix ,and The term is the standard attention score matrix, which is used to calculate the correlation between any two positions.
[0030] In a specific embodiment, the specific process of the encoder's attention layer performing the layered coefficient attention enhancement operation is as follows: Since social security policies have a chapter-section-clause hierarchical structure, user questions often require cross-level reasoning. For example, reimbursement of medical insurance in different places involves multiple chapters and policies. Conventional global attention calculations are expensive and prone to noise, while the fixed-mode sparse attention strategy of using local windows cannot adapt to the policy hierarchical logic chain. Therefore, this invention constructs a dynamic attention mask based on the policy structure to reduce the amount of calculation and keep the policy logic chain intact. The specific steps are as follows: 1) Hierarchical boundary analysis: The hierarchical boundaries of the policy text are automatically identified by matching the labeling of the policy text with regular expressions. The labeling system is expressed as "{chapter number} chapter {section number} section {clause number} article", and the regular expression that matches the labeling system is "(\d+) chapter (\d+) section (\d+) article". (\d+) is the regular expression pattern, which means matching one or more numeric characters, (\d) matches the numbers 0-9, and (+) indicates one or more repetitions. The start and end position indexes of chapters, sections, and clauses can be automatically identified without manual labeling.
[0031] It should be noted that the purpose of hierarchical boundary parsing is to automatically identify the boundary position indexes of chapters, sections, and clauses in the policy text, provide a hierarchical division basis for the construction of structured attention masks, and ensure that the structured attention masks can accurately reflect the hierarchical structure of the policy document.
[0032] 2) Structured attention mask construction: Based on the hierarchical structure of the policy text, a binary mask matrix is generated for each level. The chapter-level mask stipulates that only positions belonging to the same chapter can establish attention connections, the section-level mask stipulates that only positions belonging to the same section can establish connections, and the clause-level mask stipulates that only positions belonging to the same clause can establish connections. Then, the three-level masks are added together to obtain a structured attention mask. In the attention calculation, only positions within the same level are allowed to establish connections. The calculation formula is as follows: , in, for structured attention mask; Represents the chapter-level mask, defines the attention connectivity between chapters, and the mask logic rule is , Chapter-level mask Rank Elements of the column; Represents the section-level mask, defining the attention connectivity within the section, and the mask logic rule is , For clause level mask Rank Elements of the column; Represents the clause-level mask, defines the attention connectivity within the clause, and the mask logic rule is , For clause level mask Rank Elements of the column; 3) Sparse attention calculation: The structured attention mask is added to the scaled standard attention score matrix, and then the sparse attention weight matrix is calculated using the Softmax function. The sparse attention weight matrix has non-zero weights only between the same-level position pairs allowed by the structured attention mask. The calculation formula is as follows: , in, Indicates the dimension of the key vector, which determines the dimension of the vector in the attention mechanism. For example, when setting it, it is assigned according to the quotient of the hidden layer dimension of the model divided by the number of attention heads; is the Softmax function; is a sparse attention weight matrix that only retains non-zero weights within the same level.
[0033] It should be noted that structured attention masking uses the inherent chapter-section-clause hierarchical structure of policy documents to implement structural prior guidance and constrain the scope of attention, allowing the model to focus more on the content of the relevant hierarchy, improving the understanding and modeling capabilities of policy logic. For example, the "reimbursement ratio" only focuses on the "medical insurance benefits" section, thereby suppressing cross-section noise interference; When calculating the sparse attention weight matrix, the same level refers to the same chapter, section, or clause. During the calculation process, only the weights between positions belonging to the same level are retained, and the weights of other cross-level positions are set to zero to reduce the amount of calculation and avoid cross-level noise interference; When calculating the sparse attention weight matrix, the model avoids the fixed pattern of sliding windows to break up the logical chain of policy conditions and ensure semantic integrity. For example, the logical relationship of "continuous payment for ≥15 years" needs to be fully preserved. Through dynamic attention masks, the model can maintain the integrity of these conditional logical chains, thereby more accurately understanding the text that conforms to the policy logic. In order to retain the non-zero weights within the same level, when calculating the sparse attention weight matrix, Position and When the positions belong to the same level, ,otherwise, , so in the calculation process of the Softmax function, When used as the input of the Softmax function, the output weight of the corresponding position will approach 0, so that only the non-zero weights within the same level are retained, and the cross-level weights are suppressed.
[0034] In a specific implementation, the specific process of the encoder's output layer performing feature output operation is as follows: Since social security policy texts have complex and diverse semantic information and also contain rich hierarchical structures, conventional methods often focus on single feature extraction or simple feature combination when processing, and cannot effectively integrate position information and hierarchical structure information. Therefore, the present invention integrates the outputs of relative position attention and hierarchical sparse attention, aggregates position information and hierarchical structure information, and obtains the matrix output by the encoder. The specific steps are as follows: 1) Attention mechanism fusion: The relative position attention score matrix and the hierarchical sparse attention weight matrix are concatenated along the feature dimension, and the attention gating vector is calculated by linear transformation and Sigmoid activation function. Then, the attention gating vector is used to perform weighted summation on the relative position attention score matrix and the hierarchical sparse attention weight matrix to obtain the fused attention weight matrix. The calculation formula is as follows: , , in, Represents the attention mechanism fusion weight matrix for linear transformation; Represents the attention mechanism fusion bias vector; Represents the relative position attention score matrix and the sparse attention weight matrix Splicing along feature dimensions; Represents the fusion attention weight matrix, through the relative position attention score and hierarchical sparse attention weights Get, dimension and relative position attention score same; is the attention gate vector, which is output by the Sigmoid function and has a value range of , dynamically adjust the relative position attention score and hierarchical sparse attention weights Contribution weight; 2) Contextual representation generation: The attention value matrix is obtained by linear transformation of the fused embedding vector and the projection weight matrix. Then, the attention value matrix is weighted summed using the fused attention weight matrix and added to the fused embedding vector. Then, the addition result is layer-normalized to obtain the context representation matrix output by the encoder, which can effectively integrate relative position information and hierarchical structure information. The calculation formula is as follows: , in, represents the attention value matrix, , Represents the attention value projection weight matrix, which is used for linear transformation; The context representation matrix representing the encoder output integrates the fusion attention and position information; represents the fused embedding vector; Representation layer normalization operation.
[0035] In a specific embodiment, the specific process of the feature input operation performed by the input layer of the decoder is as follows: The social security question-answering task requires the model to accurately understand the input question and generate answers that conform to policy logic. Conventional methods often use simple concatenation or weighted summation to fuse the encoder output with the target sequence, which cannot fully capture the complex semantic associations between the input and target sequences. Therefore, this paper embeds and fuses the contextual representation of the encoder output with the target sequence to further semantically associate the input and target sequences. The specific steps are as follows: 1) Target sequence embedding: The character embedding function is used to map each character of the target sequence into a dense vector to generate a character-level embedding matrix. Then, the position embedding is calculated by the position encoding function and superimposed on the character-level embedding to construct the decoder input embedding matrix. The calculation formula is as follows: , in, represents the target sequence; represents the decoder input embedding matrix; represents the character embedding function that maps each character into a dense vector; Represents the absolute position encoding function, which is implemented by using sine and cosine functions to calculate the position embedding; 2) Masked self-attention calculation: The decoder input embedding matrix is processed using a masked multi-head self-attention mechanism. The lower triangular mask matrix is used to limit each position to only focus on itself and the previous position, which can prevent information leakage and effectively capture the internal dependencies of the target sequence. The calculation formula is as follows: , in, Represents the masked multi-head self-attention mechanism operation; Represents the masked self-attention hidden state matrix, which can capture the internal dependencies of the target sequence; It should be noted that when implementing the masked multi-head self-attention mechanism, a lower triangular mask matrix is added to the standard multi-head attention. , and the lower triangular mask matrix No. Rank The elements of the column are , when the The row index value of the row is greater than or equal to the The column index value of the column, ,otherwise , which ensures that each decoder position can only focus on itself and the previous position, preventing information leakage; 3) Encoder-Decoder Attention Fusion: The attention fusion operation is performed based on the standard multi-head attention mechanism. The standard multi-head attention mechanism uses the masked self-attention hidden state as the query matrix, the context representation output by the encoder as the key matrix and the value matrix, calculates the cross-attention weight, and fuses the encoder and decoder information through the weighted summation method to generate the attention hidden state matrix, which can realize the semantic association between the input sequence and the target sequence. The calculation formula is as follows: , in, represents the standard multi-head attention mechanism; represents the attention hidden state matrix.
[0036] In a specific embodiment, the specific process of the logic enhancement layer of the decoder performing the logic enhancement dynamic constraint operation is as follows: Social security answers must strictly follow policy rules. Since conventional decoders lack a hard constraint mechanism, soft rule injection easily leads to logical contradictions. Therefore, this invention injects hard constraints through a differentiable logic reasoning unit to make the output conform to policy logic. The specific steps are as follows: 1) Construction of logical rule base: A set of logical rules is predefined from the policy text to form a logical rule base. Each rule consists of a predicate logic expression about the input sequence and a consequent predicate logic expression about the output sequence. The two are in a logical implication relationship. The logical rule base is expressed as follows: , in, Represents a logic rule base, including rules; Indicates the Rule identifier; Represents the rule antecedent, which is about the input sequence Predicate logic expressions; Represents the rule postcondition, which is about the output sequence Predicate logic expressions; Represents logical implication relationship; The encoder is fed with an input sequence, i.e., policy text; Output sequence for the decoder, i.e. answer text; 2) Logical attention calculation: Based on the hidden state of the decoder at the current time step, the attention weight of each rule in the logical rule base at this time step is calculated through linear transformation, hyperbolic tangent function and dot product operation of the rule attention vector. It can be used to dynamically filter the rule subset most relevant to the current generation context. The calculation formula is as follows: , in, Represents the time step Rules The attention weight of Denotes the decoder at time step The hidden state of is the attention hidden state matrix No. OK; represents the logical attention weight matrix; Represents the implicit dimension of the rule, represents the dimension of the implicit representation in the rule attention mechanism, and sets ; Indicates the Rule identifier The attention vector of express The transpose of represents the natural exponential function; represents the hyperbolic tangent function; Indicates the Rule identifier The attention vector of Indicates the Rule identifier; express The transpose of 3) Differentiable logic constraint generation: The antecedent function is used to process the context representation of the encoder, and the probability of satisfying the antecedent of each rule in the logical rule base is calculated. Then, the probability of satisfying the antecedent is multiplied element-by-element by the corresponding rule's consequent embedding vector. The above results are weighted and summed using the rule attention weight calculated at the current time step to obtain the logical constraint vector of the time step. The symbolic logical rule is then converted into a differentiable constraint representation. The calculation formula is as follows: , in, is the time step The constraint vector of It is the antecedent function, which converts the symbolic logic expression of the rule antecedent into a differentiable operation. It is implemented by a multi-layer perceptron, and the input is the context representation matrix output by the encoder. , the output is a scalar value, which represents the probability that the antecedent is satisfied, and the range is ; Embed the vector for the consequent and map the rule consequent into a vector of fixed dimension; 4) Gated residual enhancement: The hidden state of the decoder at the current time step is concatenated with the logic constraint vector, and the residual enhancement gating vector is calculated through linear transformation and Sigmoid activation function. Then, the logic constraint vector is projected into the hidden space and weighted with the residual enhancement gating vector. The weighted constraint vector is then added to the original hidden state as the residual to obtain the logic enhancement hidden state matrix. This can achieve the effect of dynamically adjusting the injection strength of the constraint information. The calculation formula is as follows: , , in, Represents the residual enhancement gating weight matrix for linear transformation; represents the residual enhancement gate bias vector; Represents the time step The logic enhances the hidden state matrix to integrate the constraint information; Represents the time step The residual enhancement gate vector is output through the Sigmoid activation function, and the value range is ,dynamically adjust the constraint strength; Represents the constraint projection weight matrix, constraining the vector Projection into hidden space; It should be noted that in the process of generating differentiable logic constraints, symbolic rules are converted into differentiable operations through the antecedent function to achieve soft injection of hard constraints. The antecedent function uses neural network to approximate logical operations and outputs the probability of satisfying the rule antecedent to ensure that the gradient can be propagated during training, thereby enforcing policy rules when generating answers; dynamic attention screening is achieved through attention weights. Calculated based on the current hidden state, according to the hidden state Select relevant rule subsets to avoid irrelevant constraints interfering with high-weight rules and suppress low-weight rules, thereby dynamically adapting to the context of the generation step; the gated residual enhancement operation enhances the gate vector through the residual Adaptively adjust the constraint strength to ensure compatibility with context semantics, and enhance the gate vector with residuals As the Sigmoid gate vector, the weighted constraint vector When the constraints conflict with the semantics, the residual enhances the gate vector Approaching 0 weakens the injection, and vice versa, strengthening it to achieve flexible integration of hard constraints; the logically enhanced hidden state vector is obtained by stacking the logically enhanced hidden states of all time steps , which is then used as the input of the normalization layer to maintain the integrity of the sequence structure.
[0037] In a specific implementation, the specific process of the layer normalization operation performed by the normalization layer of the encoder is as follows: Perform layer normalization on the logistically enhanced hidden state matrix to adjust the distribution of the hidden state, stabilize the training process and improve the gradient flow, and generate a layer-normalized hidden state matrix. The calculation formula is as follows: , in, The representation layer normalizes the hidden state matrix, which stabilizes the training process and improves the gradient flow through normalization; Representation layer normalization operation.
[0038] In a specific implementation, the specific process of the decoder's output layer performing the answer sequence output operation is as follows: The normalized hidden state matrix is mapped to the word table space to generate an unnormalized score vector, which is then converted into a word probability distribution through the Softmax function. The specific steps are as follows: 1) Score vector calculation: By outputting the projection weight matrix and bias vector, the layer normalized hidden state vector is linearly transformed, the hidden state is mapped to the vocabulary space, and an unnormalized score vector is generated. The calculation formula is as follows: , in, represents the output projection weight matrix, mapping the hidden state to the vocabulary space; Represents the output projection bias vector; Represents the time step The unnormalized score vector of ; Represents the time step The layer normalized hidden state vector of ; 2) Probability distribution calculation: The Softmax function is used to process the unnormalized score vector and calculate the word probability distribution at the current time step, which represents the probability of each candidate word being generated under the given input sequence and the generated sequence, thereby guiding the sequence output. The calculation formula for the word probability distribution is as follows: , where, denotes the probability distribution of the current word given the input sequence and the generated sequence ; denotes the word to be generated at time step ; denotes the sequence of words generated before time step ; denotes the encoder input sequence; 3) Sequence generation: a) If in the training phase: Calculate the cross-entropy loss based on the word probability distribution and the true label, and optimize the model parameters through the cross-entropy loss function, the calculation formula of the cross-entropy loss function is as follows: , where, denotes the true label, i.e. the target word, at time step ; denotes the length of the target sequence; denotes the cross-entropy loss; denotes the probability of the model predicting the true word given the input and the true historical sequence ; denotes the true historical sequence before time step ; denotes the logarithmic function, with the default base being 10; b) If in the inference phase: Generate the sequence step by step by selecting the word with the maximum probability through autoregressive manner until the end-of-sequence token [SEP] is output, and complete the generation of the answer sequence, the calculation formula is as follows: , where, is the operation of taking the index of the maximum value, i.e. selecting the word with the maximum probability as the current output.
[0039] In the specific implementation, the training process of the model is as follows: S4.1, define the loss function: Define a multi-task loss function including cross-entropy loss, logical rule loss and boundary alignment loss weighted, optimize the generation quality through cross-entropy loss, ensure that the output meets the hard logical constraints of the policy rule library through logical rule loss, force the key policy numerical value to be in the pre-set reasonable interval through boundary alignment loss, and jointly improve the semantic accuracy and logical consistency of the answer through the multi-task loss function, the calculation formula is as follows: , in, is the multi-task loss function; is the weight coefficient of the logical rule loss; Loss of logic rules; is the weight coefficient of boundary alignment loss; is the boundary alignment loss; The definition of logical rule loss is based on each rule in the logical rule base. The rule satisfaction score and rule violation score are calculated. The logical loss component of each rule is obtained based on the logarithmic function. Then, the average of all rules is taken to finally generate the logical rule loss value, which quantifies the degree to which the model output sequence violates the hard logical constraints in the policy logical rule base. The calculation formula is as follows: , in, Loss of logic rules; is the total number of rules in the rule base; For rules The satisfied score represents the degree of satisfaction of differentiable logic; For rules The violated score represents the degree of violation of differentiable logic; Extract the values in the preset key value set from the model output sequence and determine whether they exceed the preset reasonable range, thereby forcing the key policy values to be within the preset reasonable range. This can avoid generating numerical results that are out of touch with the actual policy. The calculation formula is as follows: , in, is the boundary alignment loss; Represents a value extracted from the output sequence , which belongs to the preset key value set , such as, numerical is the payment period; Indicates when the value is extracted from the output sequence Less than the lower bound Penalty items when Indicates when the value is extracted from the output sequence Greater than the upper bound Penalty items when It should be noted that the logical rule loss calculates the satisfaction and violation of the rules based on the rule base to ensure that the answer meets the antecedent-consequent logical relationship of the social security policy; the boundary alignment loss determines whether it is out of bounds by extracting the key values of the output sequence. For each value, if it is lower than the preset lower bound, the lower bound difference is calculated; if it is higher than the preset upper bound, the upper bound difference is calculated. Finally, the differences of all out-of-bounds values are summed to constrain the values within the reasonable policy range.
[0040] S4.2. Define the model parameter update method: An adaptive moment estimation optimizer is used to update model parameters. The initial learning rate is set to 0.0001 and decays to 90% of the original value every 10 training epochs. The gradients of the cross entropy loss, logistic rule loss, and boundary alignment loss in the multi-task loss function are calculated independently, and a gradient clipping strategy is implemented to scale the gradient when the L2 norm exceeds the threshold of 1.0; Based on the model update method, a standard error backpropagation algorithm is used. 32 question-answer pairs are randomly sampled from the training set as a mini-batch. The multi-task loss function value is calculated through forward propagation. The gradients of each parameter are obtained through backpropagation. After gradient clipping, the Adam optimizer parameter update rule is executed, and this cycle is repeated. S4.3. Define the conditions for stopping iterative training: The main stopping condition of the model is that early stopping is triggered when the multi-task loss function value of the validation set decreases by less than 0.1% for 8 consecutive epochs; The maximum training iteration epoch limit is set to 100 times. The top three model parameter snapshots with the highest validation set accuracy are saved in each iteration epoch. The final deployed model selects the version with the highest policy logic consistency score in the validation set as the optimal model.
[0041] Example 2 The overall performance of different models on the social security question-answering task was compared, and the proposed model was compared with the traditional bidirectional long short-term memory network, the standard transformer model, and the basic version of the language representation model based on bidirectional transformer pre-training. The characteristics of each model are shown in Table 2. Table 2 Performance comparison between the model of the present invention and the existing model We use real user consultation questions collected at the edge as the test set. An example of the test set is shown in Table 3. Table 3 Test set example table The evaluation indicator is accuracy, which is calculated as follows: ; Depend on Figure 2 The experimental results show that the model of the present invention achieves significantly higher accuracy than existing models (standard neural network Transformer model, BiLSTM bidirectional long short-term memory model, BERT-base model), indicating that the multi-granularity embedding mechanism fully preserves the semantics of professional terms, enabling the model to answer practical questions more accurately.
[0042] Example 3 We use training curves to analyze the role of each component in the multi-granularity embedding architecture and conduct a systematic ablation study. The ablation comparison test configuration is shown in Table 4. Table 4 Ablation comparison experiment configuration table The validation set features, proportions, and validation starting point settings are shown in Table 5; Table 5 Validation set settings The experiments compare the performance of the full model with simplified versions such as missing term embeddings, missing gating mechanisms, and using only character embeddings. Figure 3 The experimental results shown in the figure show that the complete model significantly outperforms other configurations in both training efficiency and final accuracy. The domain knowledge directly injected into the terminology library and the three-level embedding weights dynamically adjusted by the gating mechanism effectively solve the problem of incorrect fragmentation of compound terms, enabling the model to grasp the deep semantics of policy texts more quickly.
[0043] Example 4 A heat map was used to evaluate the impact of different constraint mechanisms on policy rule compliance. The experiment focused on five core social security rules, including eligibility determination, amount calculation, and time limit requirements. The performance of the unconstrained model, the soft constraint model, and the hard constraint mechanism of the present invention were compared. The rule category comparison settings are shown in Table 6. Table 6 Rule category comparison setting table Logical consistency score evaluation method is: ; like Figure 4 The experimental results shown show that the model of the present invention maintains excellent logical consistency in all rule categories. For example, when dealing with pension calculation problems, the unit can force the model to output a policy formula that conforms to the "basic pension + personal account pension" and automatically focus on the relevant rule subsets through dynamic rule attention, avoiding interference from irrelevant rules such as the "maternity allowance application time limit".
[0044] Example 5 The model's efficiency advantage in processing long texts was verified by comparing inference time. The experiment measured the inference delay under social security policy texts of different lengths. The efficiency comparison dimensions are shown in Table 6. Table 6 Efficiency comparison of the model of the present invention and other models The test environment configuration is shown in Table 7; Table 7 Test environment configuration table Depend on Figure 5 The experimental results shown in the figure show that compared with existing models (standard neural network Transformer model, BiLSTM bidirectional long short-term memory model, BERT-base model), the model of the present invention maintains high accuracy while being significantly faster than traditional models, especially in processing long texts such as the 500-character "Medical Insurance Reimbursement Policy". This shows that the hierarchical sparse attention mechanism, based on the dynamic attention mask constructed based on the policy chapter structure, greatly reduces redundant calculations and achieves the best balance between efficiency and performance.
[0045] Although the above describes the specific implementation methods of the invention in conjunction with the accompanying drawings, it does not limit the scope of protection of the invention. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.
Claims
1. A method for processing and modeling social security digital employee data based on cloud-edge collaboration, characterized by: The following steps are involved: S1. Collaboratively collect multi-source social security data through cloud platforms and edge computing nodes. Synchronize the collected data to cloud-based structured data tables based on API interfaces at regular intervals. Fill missing insurance coverage years fields with the average value of data from adjacent regions with the same insurance type. Use distributed crawlers to capture updated policy interpretation documents from government websites in real time, parsing the policy text at the "chapter-section-clause" level. S2. Construct training samples based on collected multi-source social security data and policy texts. The sample generation logic uses question-answer pair extraction, parsing the "policy clause - applicable conditions - implementation standards" triples to generate standard question-answer templates. The training samples are then divided into training, validation, and test sets in proportion. S3. A social security digital employee question-answering model is constructed using a neural network architecture based on an encoder and decoder. The model has an encoder-decoder structure. The encoder includes an input layer, a position encoding layer, an attention layer, and an output layer. The decoder includes an input layer, a logical enhancement layer, a normalization layer, and an output layer. The training set is input into the model and processed by the encoder and decoder to generate an answer sequence. S4. Define the loss function to train the model and update the model parameters. The model is trained for multiple iterations until the preset iteration stop condition is met. The model with the best performance is selected as the optimal model through the validation set. S5. Input the test set into the optimal model, conduct social security digital employee Q&A, and generate an answer sequence.
2. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 1 is characterized in that: The specific process of the encoder's input layer performing multi-granularity semantic fusion operation is as follows: For input text Perform character-level, word-level, and term-level embedding fusion, Represents the input text, which contains questions and answers, and is expressed as [CLS] question [SEP] answer [SEP], where [CLS] represents the sequence start marker and [SEP] represents the separation marker. The specific steps are as follows: 1) Compute character-level embeddings: Based on the input text, each character is mapped into a dense vector through the character embedding function to generate a character-level embedding matrix; 2) Calculate word-level embedding: The input text is segmented to obtain word sequences, and then each word sequence is mapped into a dense vector through the word embedding function to generate a word-level embedding matrix; 3) Compute term-level embeddings: Based on the social security terminology database, the term search function is used to identify professional terms in the input text, and the corresponding term embedding vectors are directly retrieved from the terminology database to generate a term-level embedding matrix; 4) Calculate the gate vector: The character-level embedding matrix, word-level embedding matrix, and term-level embedding matrix corresponding to the same position are concatenated, and then the gate vector is calculated through linear transformation and Sigmoid activation function; 5) Compute fusion embedding: The word-level embedding matrix and the term-level embedding matrix are projected into the same dimensional space as the character-level embedding matrix. Then, the projected word-level embedding, projected term-level embedding, and character-level embedding corresponding to the same position are concatenated, and the concatenated vector is weighted using the gating vector of the position to obtain the fused embedding vector of the position. , Indicates the The fused embedding vector of each position.
3. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 2 is characterized in that: The specific process of the encoder's position encoding layer performing relative position encoding operations is as follows: The semantic-aware relative position encoding mechanism is adopted. The specific steps are as follows: 1) Calculate the semantic decay coefficient: The fused embedding vectors of the two positions are concatenated, and then the semantic attenuation coefficient is calculated through linear transformation and activation function to reflect the degree of semantic association between the two positions. When the semantic association is high, the attenuation coefficient approaches zero. 2) Calculate relative position offset: Multiply the semantic attenuation coefficient between two positions by their relative position distance and take the negative value to obtain the relative position offset value, and then obtain the relative position offset matrix; 3) Generate semantic-aware attention scores: The query matrix and key matrix of the input sequence are obtained through linear transformation, and the standard attention score matrix is calculated. Then, the calculated relative position bias matrix is added to the standard attention score matrix to obtain the relative position attention score matrix based on semantic perception.
4. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 3 is characterized in that: The specific process of the encoder's attention layer performing the layered coefficient attention enhancement operation is as follows: Based on the policy structure, a dynamic attention mask is constructed to reduce the computational complexity and keep the policy logic chain intact. The specific steps are as follows: 1) Hierarchical boundary analysis: The hierarchical boundaries of the policy text are automatically identified by matching the labeling of the policy text with regular expressions. The labeling system is represented as "{Chapter Number}Chapter{Section Number}Section{Clause Number}Article". The regular expression that matches the labeling system is "(\d+)Chapter(\d+)Section(\d+)Article". (\d+) is the regular expression pattern, indicating matching one or more numeric characters. (\d) matches the numbers 0-9, and (+) indicates one or more repetitions. 2) Structured Attention Mask Construction: Based on the hierarchical structure of the policy text, a binary mask matrix is generated for each level. The chapter-level mask stipulates that only positions belonging to the same chapter can establish attention connections, the section-level mask stipulates that only positions belonging to the same section can establish connections, and the clause-level mask stipulates that only positions belonging to the same clause can establish connections. Then, the three-level masks are added together to obtain a structured attention mask. In the attention calculation, only positions within the same level are allowed to establish connections. The calculation formula is as follows: , in, represents a structured attention mask; Represents chapter-level mask, the mask logic rule is , Indicates chapter-level mask Rank Elements of the column; Represents a section-level mask, and the mask logic rule is , Indicates the clause level mask Rank Elements of the column; Represents a clause-level mask, and the mask logic rule is , Indicates the clause level mask Rank Elements of the column; 3) Sparse attention calculation: The structured attention mask is added to the scaled standard attention score matrix, and then the sparse attention weight matrix is calculated using the Softmax function. The sparse attention weight matrix has non-zero weights only between the same-level position pairs allowed by the structured attention mask.
5. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 4 is characterized in that: The specific process of the encoder's output layer performing feature output operations is as follows: Integrate the outputs of relative position attention and hierarchical sparse attention, aggregate position information and hierarchical structure information, and obtain the encoder output matrix. The specific steps are as follows: 1) Attention mechanism fusion: The relative position attention score matrix and the hierarchical sparse attention weight matrix are concatenated along the feature dimension, and the attention gating vector is calculated through linear transformation and Sigmoid activation function. The attention gating vector is then used to perform weighted summation on the relative position attention score matrix and the hierarchical sparse attention weight matrix to obtain the fused attention weight matrix. 2) Contextual representation generation: The attention value matrix is obtained by linear transformation of the fused embedding vector and the projection weight matrix. Then, the attention value matrix is weighted summed using the fused attention weight matrix and added to the fused embedding vector. Then, the addition result is layer normalized to obtain the context representation matrix output by the encoder.
6. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 5 is characterized in that: The specific process of the decoder's input layer performing feature input operations is as follows: The context representation output by the encoder is fused with the target sequence embedding to semantically associate the input with the target sequence. The specific steps are as follows: 1) Target sequence embedding: Use the character embedding function to map each character of the target sequence into a dense vector to generate a character-level embedding matrix. Then, use the position encoding function to calculate the position embedding and superimpose it on the character-level embedding to construct the decoder input embedding matrix. 2) Masked self-attention calculation: The decoder input embedding matrix is processed using a masked multi-head self-attention mechanism, which restricts each position to only focus on itself and the previous position through the lower triangular mask matrix. 3) Encoder-Decoder Attention Fusion: The attention fusion operation is performed based on the standard multi-head attention mechanism. The standard multi-head attention mechanism uses the masked self-attention hidden state as the query matrix, the context representation of the encoder output as the key matrix and the value matrix, calculates the cross-attention weight, and fuses the encoder and decoder information through the weighted summation method to generate the attention hidden state matrix.
7. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 6 is characterized in that: The specific process of the decoder's logic enhancement layer performing logic enhancement dynamic constraint operations is as follows: Hard constraints are injected through the differentiable logic reasoning unit to make the output conform to the policy logic. The specific steps are as follows: 1) Construction of logical rule base: A set of logical rules is predefined from the policy text to form a logical rule base. Each rule consists of a predicate logic expression about the input sequence and a consequent predicate logic expression about the output sequence. The two are in a logical implication relationship. The logical rule base is expressed as follows: , in, Represents a logic rule base, including rules; Indicates the Rule identifier; Represents the rule antecedent, which is about the input sequence Predicate logic expressions; Represents the rule postcondition, which is about the output sequence Predicate logic expressions; Represents logical implication relationship; The encoder is fed with an input sequence, i.e., policy text; Output sequence for the decoder, i.e. answer text; 2) Logical attention calculation: Based on the hidden state of the decoder at the current time step, the attention weight of each rule in the logical rule base at that time step is calculated through linear transformation, hyperbolic tangent function and dot product operation of the rule attention vector. This can be used to dynamically filter the subset of rules that are most relevant to the current generated context. 3) Differentiable logic constraint generation: The encoder's context representation is processed using the antecedent function to calculate the probability of satisfying the antecedent of each rule in the logical rule base. This probability is then element-wise multiplied by the corresponding rule's consequent embedding vector. The weighted sum of these results is then performed using the rule attention weight calculated at the current time step to obtain the logical constraint vector for that time step, thereby converting the symbolic logical rules into a differentiable constraint representation. 4) Gated residual enhancement: The hidden state of the decoder at the current time step is concatenated with the logic constraint vector, and the residual enhancement gating vector is calculated through linear transformation and Sigmoid activation function. Then, the logic constraint vector is projected into the latent space and weighted with the residual enhancement gating vector. The weighted constraint vector is then added to the original hidden state as the residual to obtain the logic enhancement hidden state matrix.
8. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 7 is characterized in that: The specific process of the layer normalization operation performed by the encoder's normalization layer is as follows: Perform layer normalization on the logistically enhanced hidden state matrix to adjust the distribution of the hidden state, stabilize the training process, and improve the gradient flow, generating a layer-normalized hidden state matrix.
9. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 8 is characterized in that: The specific process of the decoder's output layer performing the answer sequence output operation is as follows: The normalized hidden state matrix is mapped to the word table space to generate an unnormalized score vector, which is then converted into a word probability distribution through the Softmax function. The specific steps are as follows: 1) Score vector calculation: By outputting the projection weight matrix and bias vector, the layer normalized hidden state vector is linearly transformed, the hidden state is mapped to the vocabulary space, and an unnormalized score vector is generated; 2) Probability distribution calculation: The Softmax function is used to process the unnormalized score vector and calculate the word probability distribution at the current time step, which represents the probability of each candidate word being generated under the given input sequence and the generated sequence, thereby guiding the sequence output. The calculation formula for the word probability distribution is as follows: , Where, Represents a given input sequence and the generated sequence Conditional current word The probability distribution of Represents the time step The word to be generated; Represents the time step Previously generated word sequences; represents the encoder input sequence; 3) Sequence generation: a) If you are in the training phase: The cross entropy loss is calculated based on the word probability distribution and the true label, and the model parameters are optimized through the cross entropy loss function. The calculation formula of the cross entropy loss function is as follows: , in, Represents the time step The true label of Indicates the target sequence length; represents the cross entropy loss; Indicates that the model is and real historical sequences Predicting real words probability; Represents the time step The real historical sequence before; Represents a logarithmic function, with a default base of 10; b) If in the inference stage: The word with the highest probability is selected through autoregression to gradually generate a sequence until the end mark [SEP] is output, completing the generation of the answer sequence. The calculation formula is as follows: , in, This is the maximum index operation.
10. The method for processing and modeling social security digital employee data based on cloud-edge collaboration according to claim 9 is characterized in that: The model training process is as follows: S4.
1. Define the loss function: Define a multi-task loss function, including weighted cross-entropy loss, logical rule loss, and boundary alignment loss. Use cross-entropy loss to optimize generation quality, use logical rule loss to ensure that the output meets the hard logical constraints of the policy rule base, and use boundary alignment loss to force key policy values to stay within a preset reasonable range. Combined with the multi-task loss function, improve the semantic accuracy and logical consistency of the answer. The logical rule loss is defined based on each rule in the logical rule base. The rule satisfaction score and rule violation score are calculated. The logical loss component of each rule is obtained based on the logarithmic function. The average of all rules is then taken to finally generate the logical rule loss value, which quantifies the degree to which the model output sequence violates the hard logical constraints in the policy logical rule base. Extract the values in the preset key value set from the model output sequence, determine whether they exceed the preset reasonable range, and then force the key policy values to be within the preset reasonable range; S4.
2. Define the model parameter update method: An adaptive moment estimation optimizer is used to update model parameters. The initial learning rate is set to 0.0001 and decays to 90% of the original value every 10 training epochs. The gradients of the cross entropy loss, logistic rule loss, and boundary alignment loss in the multi-task loss function are calculated independently, and a gradient clipping strategy is implemented to scale the gradient when the L2 norm exceeds the threshold of 1.0; According to the model update method, the standard error back propagation algorithm is used to calculate the multi-task loss function value through forward propagation, and the gradient of each parameter is obtained through back propagation. After gradient clipping, the Adam optimizer parameter update rule is executed, and this cycle is repeated; S4.
3. Define the conditions for stopping iterative training: The main stopping condition of the model is that early stopping is triggered when the multi-task loss function value of the validation set decreases by less than 0.1% for 8 consecutive epochs; The maximum training iteration epoch limit is set to 100 times. The top three model parameter snapshots with the highest validation set accuracy are saved in each iteration epoch. The final deployed model selects the version with the highest policy logic consistency score in the validation set as the optimal model.
Citation Information
Patent Citations
Relational triple combined extraction method and automatic question-answering system construction method
CN114691848A
Banking business question and answer matching method and device and customer service robot
CN116932721A
Cloud-edge collaborative big language model intelligent customer service deployment optimization method
CN117808481A
Public accumulation fund policy question and answer data processing method and system
CN118427329A
Metacosmic-based intelligent customer service interaction system and digital human service method
CN118551016A