An intelligent customer service dialogue method and system based on AIGC content generation

CN122840243APending Publication Date: 2026-09-29ANHUI JIUGUANG PANORAMIC INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610988942.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0010]鉴于现有技术的不足,本发明实施例提供一种基于AIGC内容生成的智能客服对话方法及系统,用以解决现有基于大语言模型的智能客服系统在自回归解码过程中缺乏逐词元级别的多维度协同控制能力,导致生成质量在专业性、合规性、个性化和安全性方面难以满足企业级客服场景要求的技术问题

Benefits of technology

[0020](1)通过分别计算领域专家模型和通用基线模型对各候选词元的对数概率并以两者之差作为对比分数调整概率分布,在逐词元生成过程中放大领域专家模型相对于通用基线模型的专业知识优势,提升客服回复的领域专业性和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840243A_ABST
    Figure CN122840243A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent customer service dialogue method and system based on AIGC content generation. Step S1 concatenates a security prefix vector sequence to the user input embedding front end, inputs it to a domain expert model for autoregressive decoding, and executes steps S2 to S5 sequentially. Step S2 adjusts the candidate lexical probability using the difference between the logarithmic probability of the domain expert model and the general baseline model as a comparison score, obtaining a first adjusted probability distribution. Step S3 verifies legality using a temporal logic automaton and a finite state machine, normalizing the probability of violating lexical terms to zero, obtaining a second adjusted probability distribution. Step S4 weights the distribution using the similarity between the user's implicit preference vector and the candidate lexical embedding representation as weights, obtaining a third adjusted probability distribution. Step S5 predicts the semantic risk probability based on the hidden state; if the probability does not exceed a threshold, it is written to the output buffer; if the probability exceeds the threshold, it backtracks to the security checkpoint and re-executes steps S2 to S4. Thus, professionalism, compliance, personalization, and security are cascaded and coordinated word-by-word, making it suitable for enterprise-level intelligent customer service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence content generation technology, and in particular to an intelligent customer service dialogue method and system based on AIGC content generation. Background Technology

[0002] With the rapid development of large language model technology, AIGC-based intelligent customer service systems are gradually replacing traditional template matching and retrieval-based customer service, becoming the mainstream technical solution for enterprise customer service. These systems, centered on large language models, generate customer service responses word-by-word through autoregressive decoding, enabling more flexible and natural conversational interactions.

[0003] However, existing intelligent customer service systems based on large language models lack multi-dimensional collaborative control capabilities at the word-by-word meta-level during autoregressive decoding, resulting in generation quality that fails to meet the requirements of enterprise-level customer service scenarios in terms of professionalism, compliance, personalization, and security. Specifically:

[0004] (1) In terms of professionalism, the large language model after domain fine-tuning still retains the tendency to generate general corpus. Existing technology does not distinguish and differentiate between domain professional content and general content in the decoding stage, resulting in unstable accuracy of professional terminology and process description in customer service replies.

[0005] (2) In terms of compliance, existing technologies mainly rely on data cleaning and instruction alignment during the training phase to enable the model to implicitly learn compliant behavior. This approach does not have deterministic guarantees during the inference phase. The model may generate content that violates business process rules or compliance rules at any word position. Existing technologies lack a mechanism to perform rigid constraint verification on candidate words in each word generation step.

[0006] (3) In terms of personalization, existing technologies mainly use prompting engineering to concatenate user profile text in the input to guide generation. This method operates at the input level rather than the probability distribution level, and cannot make fine-grained probability adjustments to candidate words based on user preferences during word-by-word generation.

[0007] (4) In terms of security, on the one hand, the security detection adopts a post-filtering scheme, and the security detection is only carried out after the response is fully generated. Once non-compliant content is detected, it needs to be discarded or regenerated, which wastes computing resources and increases response delay. On the other hand, there is a lack of proactive defense mechanism against malicious input, and users can bypass security alignment by prompting injection attacks.

[0008] The fundamental reason for the above shortcomings is that existing technologies treat professional enhancement, compliance constraints, personalized adjustments and security control as independent links, and fail to establish a unified cascaded control pipeline in each word generation step of autoregressive decoding, so that the control of each dimension can be coordinated in an orderly manner at the probability distribution level.

[0009] Therefore, there is an urgent need for an intelligent customer service dialogue method and system that can achieve multi-dimensional word-by-word meta-cascade control during autoregressive decoding. Summary of the Invention

[0010] In view of the shortcomings of the prior art, this invention provides an intelligent customer service dialogue method and system based on AIGC content generation, in order to solve the technical problem that the existing intelligent customer service system based on large language models lacks multi-dimensional collaborative control capabilities at the word-by-word meta-level during autoregressive decoding, resulting in the generation quality failing to meet the requirements of enterprise-level customer service scenarios in terms of professionalism, compliance, personalization and security.

[0011] In a first aspect, the present invention provides an intelligent customer service dialogue method based on AIGC content generation, comprising:

[0012] S1: Concatenate the secure prefix vector sequence to the front end of the user input embedding representation, and input the concatenated embedding representation into the domain expert model for autoregressive decoding to generate customer service responses. Steps S2 to S5 are executed sequentially in each word generation step of the autoregressive decoding. The secure prefix vector sequence is a fixed-length continuous vector sequence obtained through adversarial training. The domain expert model is a large language model fine-tuned with domain corpus.

[0013] S2: Calculate the log probability of each candidate word in the domain expert model and the general baseline model respectively. Use the difference between the log probabilities of the domain expert model and the general baseline model as the comparison score. Adjust the probability distribution of the candidate words output by the domain expert model based on the comparison score to obtain the first adjusted probability distribution. The general baseline model is a general language model with fewer parameters than the domain expert model.

[0014] S3: Based on the current states of the sequential logic automaton and the finite state machine, set the probability of the candidate word that leads to an illegal state transition in the first adjusted probability distribution to zero and normalize the remaining probabilities to obtain the second adjusted probability distribution; the sequential logic automaton is a deterministic finite automaton obtained by encoding business process rules, and the finite state machine is one or more deterministic finite automata obtained by encoding compliance rules. An illegal state transition means that the automaton does not have a legal transition target for the candidate word in the current state;

[0015] S4: Within the candidate word range where the probability is non-zero in the second adjusted probability distribution, the probability is weighted and adjusted using the similarity between the user's implicit preference vector and the embedding representation of each candidate word as the weight to obtain the third adjusted probability distribution; the user's implicit preference vector is a dense vector representing user preferences extracted from the user's historical interaction data through contrastive learning.

[0016] S5: Based on the hidden state output by the domain expert model in the current lexical generation step, perform look-ahead verification to predict the semantic risk probability of generating non-compliant content in the future. When the semantic risk probability does not exceed the first preset threshold, write the lexical sampled based on the third adjusted probability distribution into the output buffer. When the semantic risk probability exceeds the first preset threshold, backtrack to the safety checkpoint, and revert the autoregressive decoding state of the temporal logic automaton, finite state machine, and domain expert model to the historical state corresponding to the safety checkpoint. After applying probability penalties to the lexical sampled before this backtracking, re-execute steps S2 to S4 and then execute step S5. The safety checkpoint is the lexical position corresponding to the most recent update where the semantic risk probability was lower than the second preset threshold.

[0017] Secondly, the present invention provides an intelligent customer service dialogue system based on AIGC content generation, comprising: a processor and a memory;

[0018] The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the intelligent customer service dialogue method based on AIGC content generated according to the first aspect.

[0019] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0020] (1) By calculating the log probability of each candidate word element by the domain expert model and the general baseline model respectively, and using the difference between the two as the comparison score to adjust the probability distribution, the professional knowledge advantage of the domain expert model over the general baseline model is amplified in the word element generation process, thereby improving the domain professionalism and accuracy of customer service responses.

[0021] (2) By encoding business process rules into a sequential logic automaton, encoding compliance rules into a finite state machine, and verifying whether candidate words lead to illegal state transitions in each word generation step, words that violate the encoded rules are prevented from being sampled in the decoding stage by setting the probability to zero, so that the generated content meets the preset compliance requirements.

[0022] (3) By calculating the similarity between the user's implicit preference vector and the embedding representation of each candidate word within the scope of compliant words and adjusting the probability distribution accordingly, fine-grained personalized generation is achieved at the word-by-word level.

[0023] (4) By concatenating the security prefix vector sequence to the front end of the user input embedding representation and obtaining the vector sequence through adversarial training, an active defense mechanism against malicious input is established at the embedding level, enhancing the robustness of the domain expert model against prompt injection attacks. At the same time, by performing look-ahead verification based on the hidden state of the domain expert model to predict the probability of semantic risks, detection and backtracking are performed before non-compliant content is actually generated, avoiding the waste of computing resources and increased response latency caused by performing security detection only after the complete response is generated. In addition, by imposing a probability penalty on high-risk words that trigger backtracking during backtracking, the probability of repeatedly sampling the high-risk words is further reduced beyond the randomness of probability sampling, improving the effectiveness of the look-ahead verification mechanism and the overall generation efficiency of autoregressive decoding. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating an intelligent customer service dialogue method based on AIGC content generation, provided as an embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram of the architecture of an intelligent customer service dialogue system based on AIGC content generation, provided as an embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in further detail below with reference to the accompanying drawings.

[0027] This invention provides an intelligent customer service dialogue method based on AIGC content generation. Its core idea is to establish a cascaded control pipeline that includes contrastive decoding enhancement, symbol constraint filtering, personalized preference adjustment, and forward security verification in each word generation step of the autoregressive decoding of a large language model. This enables the generation process to be subject to word-level collaborative control in four dimensions: professionalism, compliance, personalization, and security.

[0028] Reference manual attached Figure 1 The diagram illustrates a flowchart of an intelligent customer service dialogue method based on AIGC content generation, provided by an embodiment of the present invention. Figure 1 As shown, the method includes steps S1 to S5. The specific implementation of each step is described in detail below.

[0029] Step S1, Secure Prefix Concatenation and Autoregressive Decoding: Concatenate the secure prefix vector sequence Embedded representation concatenated to user input The front end forms the embedded representation after splicing. ,in This indicates a concatenation operation along the sequence dimension. Input Domain Expert Model Autoregressive decoding is performed to generate customer service responses. Domain expert model. For a large language model fine-tuned with domain corpus, the user-input text is segmented by a word segmenter and then mapped to an embedding representation through the embedding layer of a domain expert model. Secure prefix vector sequence The vector is a fixed-length continuous sequence obtained through adversarial training. Its length is a preset value, which is an integer between 8 and 32 in this embodiment of the invention. The dimension of each vector is the same as the dimension of the embedding layer of the domain expert model. The specific length can be adjusted according to security requirements.

[0030] In each lexical generation step of autoregressive decoding, the domain expert model... Output the hidden state of the current step. The probability distribution of candidate words in the vocabulary is then calculated, and steps S2 to S5 are executed sequentially to adjust this probability distribution. Finally, the words are sampled and output. The autoregressive decoding process continues until the end-of-sequence marker is sampled or the number of generated words reaches the preset maximum sequence length. In this embodiment of the invention, the maximum sequence length is an integer between 512 and 4096, but this invention does not limit it. At this point, all words temporarily stored in the output buffer are concatenated to form the final customer service response.

[0031] In one possible implementation, a secure prefix vector sequence Obtained through minimax adversarial training, including alternating maximization and minimization phases. Secure prefix vector sequence. Before training, the embedding vectors of several safe semantic units from the real vocabulary are concatenated according to their prefix lengths as initial values. Alternatively, Xavier initialization can be used; this invention does not limit this approach. During the maximization phase, the projective gradient descent algorithm is used to search the embedding space for a domain expert model. Adversarial input perturbations that produce unsafe responses The initial value of the adversarial input perturbation Let it be the zero vector, and solve it using the following iterative update formula:

[0032]

[0033] in, Indicates the first The adversarial input perturbation in the next iteration and , This is the number of iterations of the projective gradient descent, and in this embodiment of the invention, it is an integer between 10 and 20. Representing the norm up to infinity, not exceeding The projection operator of the constraint set is used to constrain the disturbance within a preset range, superscript It is the iteration number, not the power. This represents the upper bound against disturbances and, in this embodiment of the invention, is a real number between 0.01 and 0.1. This represents the step size of the projected gradient descent and is taken in the embodiments of the present invention. , This indicates that the element-wise sign function is used to extract the gradient direction. Indicates to gradient operator, The safety loss term is used to measure the deviation of the model output from the safe response. To what extent, This indicates a pre-labeled safety response label. In equation (1) Pick It can be guaranteed that After the next iteration, the magnitude of the perturbation approaches the upper bound. This ensures that the detected adversarial perturbations have sufficient attack strength.

[0034] In the minimization phase, the adversarial input perturbation obtained during the maximization phase is fixed. Optimize the secure prefix vector sequence This ensures that the domain expert model can still generate a safe response under this perturbation. The training loss is:

[0035]

[0036] in, This represents the total training loss during the minimization phase. This represents the optimal adversarial input perturbation obtained during the maximization phase search. The generation quality loss term measures the deviation of the model's generation quality from the standard response under normal input conditions; that is, it measures the model's response quality under normal input conditions without adversarial perturbations. To what extent, This indicates a standard response tag used in normal conversation. Represents the weighting coefficients of the generated quality loss term and In this embodiment of the invention, the value is a real number between 0.5 and 2.0. The training loss described above incorporates... The reason for this is: if only security losses are optimized... The safety prefix may cause a decrease in the quality of model generation in normal dialogue scenarios, affecting the weight coefficients. This is used to balance the relationship between security and generation quality. After training, the secure prefix vector sequence... It is fixed as a constant tensor and is no longer updated during the inference phase.

[0037] In one possible implementation, a domain expert model Employing a sparse hybrid expert architecture, including Subnetwork of experts in various fields ,in It is a positive integer. Only the route with the highest weight is activated in each inference iteration. Subnetworks of domain experts, among which The preset number of activations and In this embodiment of the invention, the value is 2. Routing weight and the output of the current step The calculation method is as follows:

[0038]

[0039] Based on routing weight The output of the current step For activated The weighted sum of the outputs of the domain expert subnetworks:

[0040]

[0041] in, This represents an exponential function with the natural constant as its base, indicated by the superscript. Represents the transpose of a vector or matrix. Indicates the first The routing weights of the subnetworks of domain experts and , To sum the indices and iterate from 1 to All domain expert subnetworks, Equation (3) denominator for all Summation of subnetworks of experts in various fields This represents the weight matrix of the routing network. This indicates the hidden state of the current lexical generation step. The first term represents the result of a matrix-vector product. One portion, Represents the routing weight vector. Indicates the first For each domain expert subnetwork, equation (4) only applies to the route with the largest weight. The active domain expert subnetworks are summed. The routing network re-evaluates routing decisions at preset lexical intervals to adapt to topic switching during the dialogue process; in this embodiment of the invention, this interval is an integer between 8 and 64. Equation (4) only applies to the active domain expert subnetworks. Each domain expert subnetwork performs forward computation, reducing the computational cost relative to the activation... The number of expert subnetworks in each field is proportional to, rather than proportional to, the total number of networks. It is proportional to the total number of model parameters while controlling inference overhead.

[0042] Among them, the total number of domain expert subnetworks In this embodiment of the invention, the integer is taken between 8 and 64. The specific value can be adjusted according to the coverage width of the business domain; this invention does not limit this. (Hidden state) The dimension is denoted as Weight matrix of routing network The dimension is OK The parameters of the routing network and the subnetworks of each domain expert are obtained together with the pre-training and domain fine-tuning of the domain expert model.

[0043] In this embodiment of the invention, step S1 establishes an active security defense at the embedding level by searching for adversarial perturbations in the embedding space using the projective gradient descent algorithm and optimizing the security prefix vector sequence. Unlike existing technologies that perform post-filtering on the complete response at the output level, adversarial training enables the security prefix to generalize to attack patterns not covered by the training set, while post-filtering can only detect known violation patterns similar to the training samples.

[0044] During adversarial training, the secure prefix vector sequence learns a continuous transformation within the embedding space that anchors the initial hidden state of the domain expert model to the secure semantic region. During the inference phase, even when faced with user input not covered during the training phase, the initial state anchoring provided by the secure prefix can suppress the domain expert model from generating potentially misleading content in the input, enabling autoregressive decoding to start from the secure semantic region and providing robust extrapolation capabilities against cue injection attacks.

[0045] Step S2, Comparison and Decoding Probability Adjustment: In each word generation step of the autoregressive decoding initiated in step S1, steps S201 to S203 are executed sequentially. S201: Calculate the domain expert model respectively. And the general baseline model for each candidate lexical unit log probability and ,in , Given the vocabulary size, the general baseline model is a general language model with fewer parameters than the domain expert model. In this embodiment of the invention, the number of parameters in the general baseline model is one-tenth to one-half that of the domain expert model, and its log probability... Reflecting the tendency of general language generation; the difference in log probabilities between the domain expert model and the general baseline model is defined as the contrast score. .

[0046] In this embodiment of the invention, the general baseline model is a smaller pre-trained version of the domain expert model that shares the same model family and the same word segmenter. For example, it could be a small-sized base model of the same family as the domain expert model without domain fine-tuning, or a small model derived from it through knowledge distillation while maintaining the same vocabulary. Since domain fine-tuning does not change the word segmenter and vocabulary, the general baseline model and the domain expert model share the same word segmenter and the same vocabulary, enabling them to target the same candidate lexical units. The output log probabilities are comparable at the same vocabulary position, compared by scores. The calculations are clearly comparable at the word-by-word meta-level.

[0047] S202: Comparison of scores The candidate lexical units are superimposed with a contrast enhancement term on the logarithmic probability of the domain expert model. The contrast enhancement term is the contrast score compared to the preset contrast enhancement coefficient. The product of, where In this embodiment of the invention, a real number between 0.5 and 2.0 is used; for comparison fractions The candidate word units are selected while maintaining the original log probabilities of the domain expert model, thus avoiding amplification of non-professional content favored by the general baseline model. S203: Normalize the log probabilities of each candidate word unit adjusted in step S202 to obtain the first adjusted probability distribution. :

[0048]

[0049] in, The natural logarithm is the logarithm with the natural constant as its base. Indicates candidate word elements The first adjusted probability distribution, Indicates candidate word elements The comparison scores To sum the index and traverse the vocabulary from 1 to All candidate lexical units and the summation of the denominator of equation (5) over all candidate lexical units, This indicates taking the larger value between the parameter and zero. Equation (5) is obtained through... The selective enhancement in step S202 is implemented at the log probability level, that is, the log probability of candidate words with confidence levels higher than the general baseline model is amplified only, while candidate words with confidence levels lower than the general baseline model are not enhanced.

[0050] Among them, vocabulary size In this embodiment of the invention, the integer value is between 30,000 and 150,000. The specific value is determined by the word segmenter used by the domain expert model, and this invention does not limit it.

[0051] In this embodiment of the invention, step S2 compares the confidence differences between the domain expert model and the general baseline model on each candidate lexical unit by quantifying the scores, and selectively enhances the probability of lexical units in which the domain expert model has a professional knowledge advantage. Unlike existing technologies that directly use the original output probability of the domain fine-tuning model, contrastive decoding can identify and amplify domain-specific professional knowledge signals at each lexical unit generation step, while suppressing non-professional content that is indistinguishable from the general model. The first adjusted probability distribution output by step S2 serves as the input to step S3, enabling the symbol constraint verification in step S3 to be performed on the domain-enhanced probability distribution. Candidate lexical units with strong domain expertise obtain higher probabilities after enhancement in step S2, and are less likely to be diluted during normalization in the automata verification in step S3 due to the increased probability weight.

[0052] Step S3, Symbolic Constraint Filtering: Based on the current states of the sequential logic automaton and the finite state machine, verify the first adjustment probability distribution. Whether each candidate term leads to an illegal state transition. A temporal logic automaton is a deterministic finite automaton obtained by compiling the business process rules as finite-trace linear temporal logic formulas before executing step S1. Finite-trace linear temporal logic refers to linear temporal logic applicable to finite-length sequences, whose formulas can be transformed into equivalent deterministic finite automata using standard compilation algorithms. Its transition function is denoted as... ,in This represents the current state of the sequential logic automaton. Indicates candidate word elements The business actions are obtained through action recognition mapping. In one possible implementation, action recognition mapping is implemented through a predefined keyword mapping table, matching lexical units or sequences of lexical units to predefined business action categories. Lexical units that do not belong to any business action are mapped to no operation, and no operation does not trigger the state transition of the sequential logic automaton. The finite state machine is the automaton obtained by encoding the compliance rules before executing step S1, and its transition function is denoted as... ,in This represents the current state of the finite state machine.

[0053] In the token-by-token generation step, the action recognition mapping maintains a cache to be recognized, and each time a new token is sampled, it is appended to the cache to be recognized; if any suffix of the cache to be recognized has the longest match with an entry in the keyword mapping table, the corresponding service action is output and the cache to be recognized is cleared; unmatched tokens remain in the cache to be recognized until the matching succeeds or the length of the cache to be recognized reaches a preset upper limit. In the embodiment of the present invention, the preset upper limit is an integer between 16 and 64. When the preset upper limit is exceeded, the token that entered the cache to be recognized the earliest is removed from the cache to be recognized. Thereby, the token-level verification granularity and the service action-level state transition granularity are bridged by the cache to be recognized, and cross-token service action recognition is completed without increasing the design complexity of the temporal logic automaton.

[0054] The token sequence generated by the domain expert model through token-by-token sampling is successively the character "check", the character "query", the character "bal" and the character "ance". After the longest matching by the cache to be recognized, the token sequence is mapped to the service action named "query balance", which triggers the temporal logic automaton to perform state transition from the standby state to the balance query state; for characters that are modal particles, such as the character "of" and the character "le", there is no corresponding service action in the keyword mapping table, the character is mapped to a no-op, and the state of the temporal logic automaton remains unchanged.

[0055] In a possible implementation, the compliance rules include at least one of prohibition pattern rules, mandatory inclusion rules and format constraint rules. The prohibition pattern rule is encoded as a prefix matching deterministic finite automaton, which is used to block token sequences containing forbidden words or forbidden phrases; the mandatory inclusion rule is encoded as an endpoint checking deterministic finite automaton, which is used to verify whether necessary elements have been output before the end of a dialogue; the format constraint rule is encoded as a deterministic finite automaton corresponding to a regular expression, which is used to constrain the output format. After the service flow rules and compliance rules are delivered online through an interface, they are compiled online into corresponding automata, loaded immediately and replace the original automata, so that rule update does not require retraining the domain expert model.

[0056] Illegal state transition means that there is no legal transition target for the candidate token when the automaton is in the current state, that is or . The set of candidate tokens that have legal transition targets in both the temporal logic automaton and the finite state machine at the current state is recorded as the compliant token set , the candidate tokens in the compliant token set are re-normalized according to the first adjusted probability distribution, and the probabilities of candidate tokens outside the compliant token set are set to zero, so as to obtain the second adjusted probability distribution :

[0057]

[0058] wherein, represents the candidate token The second adjusted probability distribution, Represents the set of compliant terms, and , This indicates the current state of the sequential logic automaton. Next business action There is a legitimate transfer target. This indicates the current state of the finite state machine. Lower pair of words There is a legitimate transfer target. An empty set indicates that there is no valid transfer target. For summation index and the denominator of equation (6) for the compliant word set The summation of all candidate lexical units. The denominator of equation (6) is the sum of the probabilities of all candidate lexical units that pass the verification of the two automata. For the set of compliant lexical units... For candidate words other than those specified, their second adjustment probability distribution is zero. When the compliance rule is encoded as multiple deterministic finite automata, the candidate words must be determined as compliant by the check of all deterministic finite automata.

[0059] In the case where all candidate terms lead to illegal state transitions in the current state of a sequential logic automaton or finite state machine, the denominator is zero, and the set of compliant terms involved in step S4 is... At the same time, it is an empty set; at this time, the probability zeroing and normalization operation in step S3 and the personalized adjustment operation in step S4 in the current word generation step are skipped, and the backtracking in step S5 is directly triggered, and step S2 is re-executed after returning to the nearest safety checkpoint; if the current word position is the initial safety checkpoint, the preset safety prompt is used as the final output of this autoregressive decoding and the generation of this round ends, so as to avoid numerical anomalies caused by normalization operation on zero vector and ensure the robustness of the autoregressive decoding process.

[0060] In one possible implementation, during the training phase, the transfer function of the deterministic finite automaton is relaxed into a continuous function by performing differentiable relaxation on the transfer function using differentiable logical operators. :

[0061]

[0062] in, The transfer function after differentiability and relaxation is expressed in the state. and candidate lexical elements The output value and the range of values ​​are as follows. , Indicates the state of the automaton Embedded vector, The weight matrix represents the differentiable transition, used to map candidate lexical embeddings to the automaton state embedding space. Indicates candidate word elements The embedding representation is taken from the row vector corresponding to the candidate word in the projection layer of the domain expert model output. Its dimension is the same as the hidden state dimension of the domain expert model, and therefore it is in the same vector space as the user implicit preference vector obtained by sentence-level mean pooling. Represents the state embedding vector of the automaton Dimension and In this embodiment of the invention, the integer value is taken as an integer between 256 and 1024. This invention does not limit the value of the integer value. This means that the Sigmoid activation function maps the output to... Interval. The above differentiable relaxation function is divided by the automaton state embedding dimension. The square root is used to prevent the inner product from becoming too large as the dimension increases, causing the output of the Sigmoid function to approach the saturation region of 0 or 1, thus ensuring that the gradient is effectively propagated during training.

[0063] Among them, the state of the automaton Embedded vector It is obtained by looking up a learnable state embedding table by state index, and its dimension is Differentiable transferable weight matrix The number of rows equals the automaton state embedding dimension. The number of columns equals the candidate word embedding representation. The dimension of the state embedding vector. and weight matrix Xavier initialization is used before training, and the parameters are jointly optimized with those of the domain expert model during backpropagation of the auxiliary loss term. The automaton state during the training phase... Based on the labeled lexical sequence of the training samples, the transition function is updated deterministically. The candidate lexical probability distribution output by the domain expert model is used to calculate the symbol consistency loss, rather than directly driving the automaton state transition. In the case of the temporal logic automaton, the candidate lexical in Equation (7) is mapped to the business action through action recognition. The differentiable and relaxed transition function is a differentiable proxy for the legal transition of the deterministic finite automaton at the lexical level. It only provides gradient signals during the training phase and does not change the discrete state transition of the deterministic finite automaton based on the business action during the inference phase.

[0064] The sign consistency between the output of the domain expert model and the legal transitions of the deterministic finite automaton is used as an auxiliary loss term in backpropagation training to reduce the frequency of illegal state transitions during the inference phase.

[0065]

[0066] in, This represents the auxiliary loss term for sign consistency. This indicates the sequence length of the current training sample, i.e., the number of words in that training sample. , Indicates the index of the lexical generation step and , The domain expert model represents the first Step-by-step candidate word elements The output probability, Indicates the automaton in the th... The state of the step, This represents a small positive constant used for numerical stability and is taken as follows in the embodiments of the present invention. Add to the total training loss Item, of which The weighting coefficients of the symbol consistency loss are... In this embodiment of the invention, a real number between 0.1 and 1.0 is taken. The mechanism of the above symbol consistency loss is as follows: Equation (8) outer layer for all sequences The inner layer sums the positions of each word element, and the inner layer sums all candidate words in the vocabulary. The proportion of the probability quality of a computational domain expert model falling on a legal transition target; the closer this value is to 1, the more consistent the model's output distribution is with the legal transitions of the automaton; by minimizing During the training phase, the model learns to avoid generating lexical units that lead to illegal state transitions, thereby reducing the extent to which the probability distribution is modified by zeroing and normalization during the inference phase. It should be noted that Equation (8) uses the original output probability of the domain expert model instead of the probability distribution adjusted by step S2 during the inference phase for calculation, thereby internalizing the legal transition constraint into the output of the domain expert model itself, so that the constraint does not depend on the post-processing during the inference phase.

[0067] Wherein, the sequence length of the current training sample In this embodiment of the invention, the integer is taken between 128 and 2048, but the invention does not limit this.

[0068] In this embodiment of the invention, step S3 achieves deterministic rigid constraints by encoding business process rules and compliance rules into a formal automaton and verifying the legality of candidate lexical units in each lexical unit generation step, with a probability of zero. Unlike existing technologies that rely on implicit learning of compliance behavior during the training phase, symbolic constraints provide verifiable deterministic guarantees during the inference phase. That is, any lexical unit that violates the encoded rules cannot be sampled and output, and rule updates take effect immediately through online compilation, without requiring model retraining.

[0069] Step S4, Personalized Preference Adjustment: In the second adjustment probability distribution Within the range of candidate lexical terms with non-zero probabilities, calculate the user's implicit preference vector. Embedded representations of each candidate lexical unit The similarity is used as a weight to adjust the probability, resulting in a third adjusted probability distribution. :

[0070]

[0071] in, Indicates candidate word elements The third adjusted probability distribution, For the set of compliant terms as defined in equation (6), Indicates personalized temperature coefficient and This value is used to control the intensity of personalized adjustments, and in this embodiment of the invention, it is a real number between 0.1 and 1.0. User implicit preference vector With candidate lexical embedding representation The cosine similarity between them, where and All are vectors with the same dimension as the hidden states of the domain expert model. Represents the L2 norm. This represents a small positive constant used to avoid fractions with a denominator of zero. To sum the index and the denominator of equation (9) for the compliant word set The summation is performed on all candidate word units. In the third adjusted probability distribution described above, the cosine similarity is exponentially calculated and then summed with the second adjusted probability distribution. The effect of multiplication is that the exponential transformation expands the cosine similarity from a finite closed interval. Mapped to an interval of positive real numbers This ensures that the weighting factors are strictly positive and preserves the order of cosine similarity; when When the exponent term is always 1, the third adjustment probability distribution degenerates into the second adjustment probability distribution, meaning no personalized adjustment is applied.

[0072] Among them, personalized temperature coefficient In this embodiment of the invention, a real number between 0.1 and 1.0 is used; however, this invention does not limit this value. This is used to avoid small positive constants with a denominator of zero. In the embodiments of the present invention, take Its function is to maintain the denominator strictly greater than zero when the L2 norm of the user's implicit preference vector or candidate word embedding representation is extremely small, thus avoiding division by zero anomalies in floating-point operations.

[0073] User implicit preference vector This refers to a dense vector representing user preferences extracted from user history interaction data through contrastive learning before performing step S1. In one possible implementation, positive sample pairs are constructed using sentence-level hidden layer representations extracted by a domain expert model from the same user's high-satisfaction historical responses in different sessions. Negative samples are constructed using sentence-level hidden layer representations of responses from different users or historical responses from the same user with low satisfaction levels. Training is performed using the InfoNCE contrastive loss function:

[0074] In this embodiment of the invention, the sentence-level hidden layer representation refers to: after forward propagation of the corresponding historical response input domain expert model, the dense vector obtained by mean pooling the hidden states of each word position in the last layer; the sentence-level hidden layer representation is related to the word-by-word hidden states in steps S1 and S5. The difference lies in the granularity of the objects: the former is an aggregate representation of the entire historical response, while the latter is a step-by-step representation of the current word generation step.

[0075]

[0076] in, Indicates InfoNCE contrast loss, and This represents a positive sample pair consisting of sentence-level hidden layer representations of two highly satisfactory historical replies from the same user in different sessions. Indicates the first The sentence-level hidden layer representation of each negative sample and , The number of negative samples is an integer between 64 and 256 in this embodiment of the invention. Represents the cosine similarity function. The temperature coefficient representing the comparative learning and In this embodiment of the invention, a real number between 0.05 and 0.1 is used. With equation (9) For different parameters with different value ranges: Controlling the discrimination between positive and negative samples during contrastive learning training. The intensity of personalized adjustments during the control reasoning phase.

[0077] After training, the sentence-level hidden layer representations of the target user's high-satisfaction historical responses are taken, and a weighted average is calculated using the user's satisfaction score as the weight to obtain the user's implicit preference vector:

[0078]

[0079] in, This indicates the number of historical responses from the target user with high satisfaction levels. , Indicates the index of historical replies and , Indicates the first Sentence-level hidden layer representation of historical responses with high satisfaction rates. This indicates the first user feedback. The satisfaction score of each reply and .

[0080] Among them, the number of historical responses with high satisfaction among target users In this embodiment of the invention, the value is an integer between 5 and 50; however, this invention does not impose a limitation on this value. (Satisfaction score) The range of values ​​is The real number is used in this embodiment of the invention after the original user feedback is normalized to this interval by a linear mapping. For example, the five-star rating of 1 to 5 stars is mapped to 0.2, 0.4, 0.6, 0.8 and 1.0 respectively. In this embodiment of the invention, historical responses with a satisfaction score of not less than 0.8 are counted as high satisfaction and used as positive samples, and historical responses with a satisfaction score of not more than 0.4 are counted as low satisfaction and used as negative samples. Other intermediate responses are not included in the comparative training. This invention does not limit the specific dividing threshold.

[0081] For new users with no historical interaction data, extract the input style vector from the new user's input text. Input style vector To obtain the vector from the mean pooling of the last hidden state after forward propagation of the new user's input text through the domain expert model, the input style vector is... Nearest neighbor matching is performed with the user prototype cluster centers obtained through K-means clustering during the training phase, and the nearest neighbors are weighted and fused using the reciprocal of the distance as the weight. The preference vector corresponding to the center of each user prototype cluster:

[0082] In this embodiment of the invention, the user prototype cluster center is constructed in the following way: For all existing users whose implicit preference vectors have been collected during the training phase, the K-means clustering algorithm is used to cluster all user implicit preference vectors, and the number of clusters is denoted as . In this embodiment of the invention, the value is an integer between 32 and 512. This invention does not limit the value of this value. The number of clusters can be selected according to the scale of existing users using known methods such as the elbow rule or the silhouette coefficient. After clustering, the representative preference vector of the cluster is obtained by taking the average of the implicit preference vectors of all users in each cluster. The cluster center vector is obtained by averaging the input style vectors of all users within the cluster. The cluster center vector and the representative preference vector are then compared. Together they form the core of the user prototype cluster, used for constructing preference vectors for new users without historical interaction data.

[0083]

[0084] in, This represents the implicit preference vector of a new user. Indicates the first The preference vector corresponding to the center of each user prototype cluster and , This represents the number of nearest neighbor clusters participating in the fusion, and in this embodiment of the invention, it is an integer between 3 and 5. Represents the input style vector With the Cluster center L2 distance between them It is a small positive constant as defined in equation (9). In equation (12), it is the reciprocal of the distance. The weights are set so that cluster centers that are closer to each other contribute more, thus achieving preference transfer based on style similarity.

[0085] In this embodiment of the invention, step S4 applies personalized adjustments within the range of compliant lexical units that have passed the symbol constraint verification in step S3, due to the summation range of step S4. Only candidate words with non-zero probabilities output in step S3 are included. Personalization adjustments do not restore the probabilities of violating words that were set to zero in step S3, thus achieving personalization while ensuring compliance. Unlike existing technologies that concatenate user profile text at the input level, step S4 calculates the semantic similarity between user preferences and each candidate word at the probability distribution level, achieving fine-grained personalization control. Furthermore, since the operation object of step S4 is the probability distribution after domain enhancement in step S2 and compliance filtering in step S3, personalization adjustments are selected from a range of domain-specific and compliant candidate words, with the constraints of the three dimensions tightening layer by layer in the same probability space.

[0086] Step S5, Look-ahead Verification and Backtracking: In each lexical generation step of autoregressive decoding, steps S501 to S504 are executed sequentially. S501: Based on the hidden state output by the domain expert model in the current lexical generation step. Conduct forward-looking verification to predict the semantic risk probability of generating non-compliant content in the future. In one possible implementation, a separate multilayer perceptron and linear classification layer are set up outside the domain expert model, based on the hidden state output by the domain expert model at the current lexical generation step. Using the following formula as input, the semantic risk probability is calculated. :

[0087]

[0088] in, Indicate steps semantic risk probability and , This represents the hidden state output by the domain expert model at the current lexical generation step. and The weight matrix and bias vector of the multilayer perceptron are trainable parameters independent of the domain expert model, respectively. They are obtained during the training phase using manually labeled safe and unsafe response pairs as training samples and trained using the binary cross-entropy loss function. The linear rectified activation function is... , and These represent the weight vector and bias scalar of the linear classification layer, respectively. This means that the Sigmoid activation function maps the output to... The interval is used as the semantic risk probability. The parameters of the multilayer perceptron and linear classification layer in equation (13) are... Independent of the parameters of the domain expert model, it only reuses the hidden states generated by the domain expert model during the autoregressive decoding process. As input, this allows look-ahead validation to proceed without modifying the forward propagation process of the domain expert model, thus avoiding structural reasoning overhead on the domain expert model itself.

[0089] in, for real matrix, for 3D real vector, for 3D real vector, Let be a real scalar, representing the weight matrix of the multilayer perceptron and the linear classification layer. and weight vector Xavier initialization and bias vectors are used before training. With bias scalar Initialize to zero; Hidden state The dimension is taken as an integer between 2048 and 8192 in this embodiment of the invention, and is determined by the architecture of the domain expert model. This invention does not limit it. In the embodiments of the present invention, the hidden layer dimension of the multilayer perceptron is... The integer to be used is between 128 and 512; this invention does not limit this value. The training samples for the multilayer perceptron and linear classification layer are constructed as follows: The hidden state sequence output by the domain expert model at each word generation step during the generation of manually labeled responses is collected. ,in The number of terms in the response and In this embodiment of the invention, the integer is taken between 16 and 512, but this invention does not limit it; for responses marked as insecure, each of the hidden state sequences is... Set the tag to 1; for responses marked as safe, set each The labels are set to 0; the obtained hidden states and label pairs are used as training samples to minimize the binary cross-entropy loss function. The above training sample construction method passes the security annotations at the response level to the hidden states at the word level, enabling the multilayer perceptron to predict the probability of generating non-compliant content based on the current hidden state in the intermediate step of autoregressive decoding.

[0090] S502: When semantic risk probability Not exceeding the first preset threshold At that time, the third adjusted probability distribution will be used. The sampled lexical units are written to the output buffer, which is a temporary storage area for the sampled lexical units; when the semantic risk probability... Below the second preset threshold When the current lexical position is updated to a safety checkpoint, the states of the sequential logic automaton and finite state machine at this time are recorded. The safety checkpoint is when the semantic risk probability is lower than a certain value. The most recent update corresponds to the lexical position. In this embodiment of the invention, the security checkpoint update operation is performed before the sampled lexical is written to the output buffer.

[0091] S503: When semantic risk probability Exceeding the first preset threshold If the number of backtracking iterations at the current lexical position has not reached the backtracking limit, backtrack to the safety checkpoint and re-execute steps S2 to S4, then execute step S5. Since the output buffer and automaton state have been synchronously rolled back, and the domain expert model recalculates the hidden state based on the rolled-back context during re-execution, This leads to the recalculation of the first, second, and third adjusted probability distributions. Due to the inherent randomness of probability sampling, the re-execution may result in sampling terms different from those before the current backtracking. In this embodiment of the invention, Take a real number between 0.5 and 0.8. Take a real number between 0.1 and 0.3, and The criteria for determining security checkpoints are stricter than the criteria for triggering backtracking, ensuring sufficient security margins for the target locations being backtracked.

[0092] During backtracking, the autoregressive decoding states of the sequential logic automaton, finite state machine, and domain expert model are synchronously rolled back to the historical states corresponding to the safety checkpoints, and tokens after the safety checkpoints in the output buffer are discarded. An upper limit is set for the number of backtracking iterations for the same token position. In the embodiments of the present invention Take an integer between 2 and 5; Indicates position The number of times backtracked, the number of times backtracked The initial value is 0, and the number of backtracks is incremented by 1 each time a backtrack is executed. The backtrack count is reset to zero when a new word is written to the output buffer or the security checkpoint position is updated.

[0093] The autoregressive decoding state of the domain expert model includes the key-value cache maintained by the model during autoregressive decoding and the hidden state sequence output by each lexical generation step. The rollback operation truncates the key-value cache and hidden state sequence to the position corresponding to the safety checkpoint. This ensures that when step S2 is re-executed, the domain expert model recalculates the log probabilities of each candidate lexical based on the context prefix consistent with the safety checkpoint, thus avoiding the continued participation of hidden states obtained based on discarded lexicals in subsequent calculations. The key-value cache stores the key and value vectors of generated lexicals at each attention layer during the autoregressive decoding process. Reusing the key and value vectors in each new lexical generation step avoids repeatedly calculating attention for already generated lexicals, a well-known technique in large language model inference.

[0094] To avoid resampling the same candidate lexical that triggered a high risk before the current backtracking after re-executing steps S2 to S4, a backtracking penalty set is introduced in step S2 after the backtracking. The backtracking penalty set is initially empty, and each time a backtracking is triggered, the lexical sampled before the backtracking is added to the backtracking penalty set. When re-executing step S2, for candidate lexicals in the backtracking penalty set, a negative penalty term is added to the log probability of the domain expert model before normalization. The negative penalty term is denoted as... and Its absolute value In this embodiment of the invention, the dynamic range of the logarithmic probability distribution of the domain expert model is taken as a real number between 3 and 10. This invention does not limit this range and the value can be selected by parameter tuning on the validation set, using the criterion that the frequency of repeated backtracking at the same security checkpoint no longer decreases. Whenever a new term is written to the output buffer or the security checkpoint position is updated, the backtracking penalty set is cleared. By introducing the backtracking penalty set, the probability of resampling to the same high-risk term is further reduced beyond the randomness of probability sampling, avoiding a cycle of triggering the same backtracking in the same context.

[0095] S504: When semantic risk probability Exceeding the first preset threshold And the number of backtracking attempts has been completed. Reaching the maximum number of backtracking attempts At this point, backtracking is no longer performed, and the lexical units obtained from sampling based on the third adjusted probability distribution are written to the output buffer.

[0096] In this embodiment of the invention, step S5 performs look-ahead verification by reusing the hidden states of the domain expert model, predicting the semantic risks of subsequent generation before the actual output of the lexical units, thus shifting security detection from a post-processing to a pre-processing approach. Unlike existing technologies that perform security detection only after the complete generation of the response, look-ahead verification assesses risks in real-time at each lexical unit generation step. Once a risk is detected, it can immediately backtrack to the security checkpoint and regenerate, avoiding the waste of computational resources and increased response latency caused by discarding the entire generated content.

[0097] Reference manual attached Figure 2 The diagram shows a structural schematic of an intelligent customer service dialogue system based on AIGC content generation provided by an embodiment of the present invention.

[0098] This invention also provides an intelligent customer service dialogue system 20 based on AIGC content generation, including: a processor 201 and a memory 202;

[0099] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the above-mentioned intelligent customer service dialogue method based on AIGC content generation and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.

[0100] It should be understood that the processor 201 in this embodiment of the invention can be a central processing unit, or it can be other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, or other programmable logic devices. The general-purpose processor can be a microprocessor or any conventional processor.

[0101] It should also be understood that the memory 202 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, or flash memory. The volatile memory can be random access memory, which is used as an external cache.

[0102] The technical solutions in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or other arbitrary combinations. When implemented using software, the above technical solutions can be implemented, in whole or in part, in the form of a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. A computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more sets of available media. The available medium can be a magnetic medium, an optical medium, or a semiconductor medium.

[0103] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0104] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0105] If the functionality is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A method for intelligent customer service dialogue based on AIGC content generation, characterized in that, include: S1: Concatenate the secure prefix vector sequence to the front end of the user-input embedded representation, and input the concatenated embedded representation into the domain expert model for autoregressive decoding to generate customer service responses. Steps S2 to S5 are executed sequentially in each word generation step of the autoregressive decoding. The secure prefix vector sequence is a fixed-length continuous vector sequence obtained through adversarial training. The domain expert model is a large language model fine-tuned with domain corpus. S2: Calculate the log probability of each candidate word for the domain expert model and the general baseline model respectively, and use the difference between the log probability of the domain expert model and the general baseline model as the comparison score. Adjust the probability distribution of the candidate words output by the domain expert model based on the comparison score to obtain the first adjusted probability distribution; the general baseline model is a general language model with fewer parameters than the domain expert model. S3: Based on the current states of the temporal logic automaton and the finite state machine, set the probability of the candidate word element that leads to an illegal state transition in the first adjusted probability distribution to zero and normalize the remaining probabilities to obtain the second adjusted probability distribution; the temporal logic automaton is a deterministic finite automaton obtained by encoding business process rules, and the finite state machine is one or more deterministic finite automata obtained by encoding compliance rules; the illegal state transition refers to the automaton not having a legal transition target for the candidate word element in the current state. S4: Within the range of candidate lexical units with non-zero probabilities in the second adjusted probability distribution, the probabilities are weighted and adjusted using the similarity between the user's implicit preference vector and the embedding representation of each candidate lexical unit as weights to obtain a third adjusted probability distribution; the user's implicit preference vector is a dense vector representing user preferences extracted from user's historical interaction data through contrastive learning. S5: Based on the hidden state output by the domain expert model in the current lexical generation step, perform look-ahead verification to predict the semantic risk probability of generating non-compliant content in the future. When the semantic risk probability does not exceed the first preset threshold, write the lexical sampled based on the third adjusted probability distribution into the output buffer. When the semantic risk probability exceeds the first preset threshold, backtrack to the safety checkpoint, and revert the autoregressive decoding states of the temporal logic automaton, the finite state machine, and the domain expert model to the historical state corresponding to the safety checkpoint. After applying a probability penalty to the lexical sampled before this backtracking, re-execute steps S2 to S4 and then execute step S5. The safety checkpoint is the lexical position corresponding to the most recent update where the semantic risk probability is lower than the second preset threshold.

2. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, The secure prefix vector sequence is obtained by training in the following way: during the maximization phase, the projective gradient descent algorithm is used to search for adversarial input perturbations in the embedding space that cause the domain expert model to produce an insecure response; During the minimization phase, the secure prefix vector sequence is optimized so that the domain expert model still generates a secure response under the adversarial input perturbation; The training loss includes a generation quality loss term and a security loss term, wherein the generation quality loss term is used to ensure that the security prefix vector sequence does not degrade the generation quality of normal dialogue; After training, the sequence of secure prefix vectors is fixed into a constant tensor.

3. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, The domain expert model adopts a sparse hybrid expert architecture, which includes multiple domain expert sub-networks and a routing network. In each inference, only a preset number of the domain expert sub-networks with the highest routing weight are activated. The domain expert model calculates the routing weight of each domain expert sub-network based on the semantic embedding representation of the current dialogue context through the routing network. The routing network re-evaluates the routing decision at preset lexical intervals to adapt to topic switching during the dialogue process.

4. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, Step S2 specifically includes: S201: Calculate the log probability of each candidate word for the domain expert model and the general baseline model respectively, and subtract the log probability of the general baseline model from the log probability of the domain expert model to obtain the comparison score of each candidate word. S202: For candidate words with positive contrast scores, a contrast enhancement term is superimposed on the logarithmic probability of the domain expert model, wherein the contrast enhancement term is the product of the contrast score and a preset contrast enhancement coefficient; for candidate words with zero or negative contrast scores, the original logarithmic probability of the domain expert model remains unchanged. S203: Perform a normalization operation on the logarithmic probabilities of each candidate word after processing in step S202 to obtain the first adjusted probability distribution.

5. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, The temporal logic automaton is a deterministic finite automaton compiled by expressing the business process rules as finite-trace linear temporal logic formulas. In the autoregressive decoding process, words are mapped to business actions through action recognition, and the deterministic finite automaton performs state transitions according to the business actions. During the training phase, the transition function of the deterministic finite automaton is differentially relaxed using differentiable logic operators. The sign consistency between the candidate word probability distribution output by the domain expert model at each training step and the legal transition probability obtained through the differentially relaxed transition function is used as an auxiliary loss term for backpropagation training. The sign consistency is defined as the probability quality of the candidate word probability distribution output by the domain expert model falling on the legal transition target. The auxiliary loss term aims to maximize the probability quality to reduce the triggering frequency of illegal state transitions during the inference phase.

6. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, The compliance rules include at least one of prohibited pattern rules, forced inclusion rules, and format constraint rules. The prohibited pattern rules are encoded as prefix matching deterministic finite automata, the forced inclusion rules are encoded as endpoint checking deterministic finite automata, and the format constraint rules are encoded as deterministic finite automata corresponding to regular expressions. After the business process rules and the compliance rules are distributed online through the interface, they are compiled online into the corresponding automata and loaded and replaced with the original automata in real time.

7. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, The user implicit preference vector is constructed as follows: positive sample pairs are formed by the sentence-level hidden layer representations of the same user's high-satisfaction historical replies in different sessions, and negative sample pairs are formed by the sentence-level hidden layer representations of replies from different users or the same user's low-satisfaction historical replies. The InfoNCE contrastive loss function is used for training. After training, the sentence-level hidden layer representations of the target user's high-satisfaction historical replies are weighted and averaged using the user's feedback satisfaction score as the weight to obtain the user implicit preference vector. For new users without historical interaction data, an input style vector is extracted from the new user's input text. The input style vector is then matched with the nearest neighbor of the user prototype cluster centers obtained through clustering during the training phase. The preference vectors corresponding to the nearest multiple user prototype cluster centers are weighted and fused using the reciprocal of the distance as the weight to obtain the user implicit preference vector.

8. The intelligent customer service dialogue method based on AIGC content generation according to claim 1, characterized in that, Step S5 specifically includes: S501: Based on the hidden state output by the domain expert model in the current lexical generation step, perform look-ahead verification to predict the semantic risk probability of generating non-compliant content in the future. S502: When the semantic risk probability does not exceed the first preset threshold, the word character sampled based on the third adjusted probability distribution is written into the output buffer, and when the semantic risk probability is lower than the second preset threshold, the current word character position is updated to the security checkpoint. S503: When the semantic risk probability exceeds the first preset threshold and the number of backtrackings at the current word position has not reached the preset backtracking limit, backtrack to the safety checkpoint, and backtrack the key-value cache and hidden state sequence of the temporal logic automaton, the finite state machine and the domain expert model to the historical state corresponding to the safety checkpoint. Discard the words in the output buffer that are located after the safety checkpoint, apply a probability penalty to the words sampled before this backtracking to avoid resampling the words, and then re-execute steps S2 to S4, and then execute step S5. S504: When the semantic risk probability exceeds the first preset threshold and the number of backtracking attempts reaches the upper limit of the number of backtracking attempts, the lexical units sampled based on the third adjusted probability distribution will be written into the output buffer.

9. The intelligent customer service dialogue method based on AIGC content generation according to claim 8, characterized in that, The look-ahead verification in step S501 is achieved through a multilayer perceptron and a linear classification layer independent of the domain expert model. The hidden state output by the domain expert model in the current word generation step is used as input to predict the semantic risk probability. The parameters of the multilayer perceptron and the linear classification layer are independent of the parameters of the domain expert model, and are trained using manually labeled safe and unsafe response pairs as training samples through a binary cross-entropy loss function during the training phase.

10. An intelligent customer service dialogue system based on AIGC content generation, characterized in that, include: Processor and memory; The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the intelligent customer service dialogue method based on AIGC content generation according to any one of claims 1 to 9.