Scientific and technological achievement intelligent pushing and public service question and answer method, device and equipment

By using intelligent agents based on large language models and Bayesian engines, the problems of low efficiency and poor accuracy in manual processing modes have been solved, enabling efficient, accurate, and consistent intelligent push of scientific and technological achievements and public service Q&A.

CN121882243APending Publication Date: 2026-04-17BEIJING SCI & TECH PATENT OFFICE +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SCI & TECH PATENT OFFICE
Filing Date
2025-12-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, manual processing is inefficient, inaccurate, and incomplete in scenarios such as customer service inquiries, government services, and enterprise after-sales service. It is difficult to cope with sudden surges in inquiries, and the problems of delayed and inconsistent responses are serious.

Method used

An intelligent agent based on a large language model and a Bayesian engine is used to receive raw text, determine common posterior beliefs, and push scientific and technological achievements or public service response texts under certain conditions. By combining optimal action decisions and optimal stopping strategies, intelligent push can be achieved.

Benefits of technology

It significantly improves the efficiency, accuracy, and comprehensiveness of intelligent technology achievement delivery and public service Q&A, reduces reliance on manual processing, and ensures the timeliness and consistency of responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121882243A_ABST
    Figure CN121882243A_ABST
Patent Text Reader

Abstract

The invention provides a scientific and technological achievement intelligent pushing and public service question and answer method, device and equipment, and relates to the technical field of intelligent agents, and the method comprises the steps: receiving an original text sent by a terminal; through an intelligent agent based on a large language model and a Bayesian engine, a public posterior belief is determined based on a public prior belief observed at the current time, an original text and corresponding low-dimensional observation, the public posterior belief is used for determining and executing an optimal action corresponding to the original text, and under the condition that it is determined that search and question seeking are stopped based on the public posterior belief, the optimal action corresponding to the original text is executed. And pushing target push content corresponding to the original text to the terminal, wherein the target push content comprises the scientific and technological achievement items or the public service reply text. According to the method, the efficiency, accuracy and comprehensiveness of scientific and technological achievement intelligent pushing and public service question answering can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent technology, and in particular to a method, apparatus and equipment for intelligent push of scientific and technological achievements and public service Q&A. Background Technology

[0002] With the rapid development of internet technology, various information exchange platforms are proliferating, and the number of user inquiries and feedback is growing exponentially. In scenarios such as customer service inquiries, government services, and enterprise after-sales service, the traditional manual processing model requires customer service personnel to manually search knowledge bases, review relevant clauses one by one, and then write responses to user questions. This approach not only requires customer service personnel to possess specialized knowledge but also consumes a significant amount of time in information retrieval and text organization.

[0003] In existing technologies, manual processing typically involves the following steps: customer service personnel receive user inquiries, manually search a knowledge base system using keywords, sift through massive amounts of information to find relevant clauses, paraphrase or integrate the clause content based on their personal understanding, and finally send a response to the user. The entire process heavily relies on the professional skills and experience of customer service personnel, and different personnel may handle the same issue significantly differently. Furthermore, due to the limited speed of manual processing, delays and waiting times are common when facing sudden surges in inquiries. This traditional approach often struggles to guarantee the accuracy and comprehensiveness of responses to complex inquiries or situations requiring the citation of multiple clauses. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, apparatus and equipment for intelligent push of scientific and technological achievements and public service Q&A, which can significantly improve the efficiency, accuracy and comprehensiveness of intelligent push of scientific and technological achievements and public service Q&A.

[0005] In a first aspect, the present invention provides a method for intelligent delivery of scientific and technological achievements and public service Q&A, including: The raw text sent by the receiving terminal; An intelligent agent based on a large language model and a Bayesian engine determines a common posterior belief based on the common prior belief of the current observation, the original text and its corresponding low-dimensional observation. This belief is used to determine and execute the optimal action corresponding to the original text. When the agent determines to stop searching for evidence and asking follow-up questions based on the common posterior belief, the agent pushes the target content corresponding to the original text to the terminal. The target content includes scientific and technological achievements or public service response texts.

[0006] Secondly, the present invention also provides a device for intelligent push of scientific and technological achievements and public service Q&A, comprising: The text receiving module is used to receive raw text sent by the terminal; The content push module is used to determine the public posterior belief based on the public prior belief of the current observation, the original text and its corresponding low-dimensional observation by an intelligent agent based on a large language model and a Bayesian engine. It is used to determine and execute the optimal action corresponding to the original text, and push the target push content corresponding to the original text to the terminal when the search for evidence and further questioning are stopped based on the public posterior belief. The target push content includes scientific and technological achievement items or public service response text.

[0007] Thirdly, the present invention also provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement any of the methods provided in the first aspect.

[0008] Fourthly, the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement any of the methods provided in the first aspect.

[0009] This invention provides a method, apparatus, and device for intelligent technology achievement push and public service Q&A. Using an intelligent agent based on a large language model and a Bayesian engine, it determines a common posterior belief based on the common prior belief of the current observation, the original text, and its corresponding low-dimensional observation. This belief is used to determine and execute the optimal action corresponding to the original text. Upon determining to stop evidence searching and follow-up questioning based on the common posterior belief, it pushes the target content corresponding to the original text to the terminal. The target content includes technology achievement entries or public service response texts, significantly improving the efficiency, accuracy, and comprehensiveness of intelligent technology achievement push and public service Q&A.

[0010] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating a method for intelligent push of scientific and technological achievements and public service Q&A provided in an embodiment of the present invention; Figure 2 A technical framework diagram of a method for intelligent push of scientific and technological achievements and public service Q&A provided in an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a device for intelligent push of scientific and technological achievements and public service Q&A provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0015] Currently, existing technologies suffer from drawbacks such as low efficiency, unstable quality, poor interpretability, and difficulty in quantifying governance fairness due to manual retrieval and line-by-line responses. Based on this, the present invention provides a method, apparatus, and equipment for intelligent push of scientific and technological achievements and public service Q&A, which can significantly improve the efficiency, accuracy, and comprehensiveness of intelligent push of scientific and technological achievements and public service Q&A.

[0016] To facilitate understanding of this embodiment, a method for intelligent technology achievement push and public service Q&A disclosed in this embodiment of the invention will first be described in detail. See [link to relevant documentation]. Figure 1 The diagram shows a method for intelligent technology achievement push and public service Q&A, which mainly includes the following steps S102 to S104: Step S102: Receive the raw text sent by the terminal; Step S104: An agent based on a large language model and a Bayesian engine determines a common posterior belief based on the common prior belief of the current observation, the original text and its corresponding low-dimensional observation. This belief is used to determine and execute the optimal action corresponding to the original text. If the agent determines to stop searching for evidence and asking follow-up questions based on the common posterior belief, the agent pushes the target content corresponding to the original text to the terminal. The target content includes scientific and technological achievement items or public service response text.

[0017] For ease of understanding, this embodiment of the invention provides a specific implementation of a method for intelligent technology achievement push and public service Q&A, deployed in a community setting, allowing residents and community workers to interact through a client. See [link to relevant documentation]. Figure 2 The diagram illustrates a technical framework for a method of intelligent technology achievement recommendation and public service question answering, including: S0, data preparation and indexing; S1, parsing and single-time posterior; S2, batch fusion posterior; S3, action observation filtering; S4, expected cost and optimal action; S5, threshold determination; S6, compliance and risk mitigation; S7, delivery and logging; S8, strategy learning; S9, behavior auditing; and S10, attention cost and weight closed-loop correction. The specific process is shown below: S0, Data Preparation and Indexing: Acquire national / provincial / municipal / district / county-level scientific and technological achievements, public demands texts, etc. Scientific and technological achievements include policies, guidelines, projects, patents, and paper interpretations. The above data is processed as follows: cleaning, slicing, terminology alignment, graph anchoring, and vector / inverted indexing (RAG) to obtain a searchable evidence base and metadata index. The searchable evidence base refers to the original unstructured or semi-structured scientific and technological achievements and government information, such as policy documents and public demands texts; the metadata index is an indexing system built upon the structured descriptive information attached to each knowledge entry in the searchable evidence base.

[0018] S1, analytic and one-way posterior, includes: S1.1, using a large language model to process the original text Low-dimensional observation data were obtained through analysis. Low-dimensional observation data That is, low-dimensional interpretable observations.

[0019] S1.2, based on the original text and low-dimensional observation data Construct a joint observation kernel Specifically, the input will be the original text. Analysis into low-dimensional observation data Low-dimensional observation data Includes slots for topic, region, target audience, timeframe, and event stage, along with their confidence levels, based on the original text. Low-dimensional observation data Original text potential state Constructing the original text Analysis into low-dimensional observation data Analytical probability kernel and by potential state To the original text Generation likelihood The two form a joint observation core : ; in, Represents the potential state To the original text The generation likelihood, as shown below, refers to the likelihood of the system being in its true potential state. (For example, when "Policy Subject = Low-income assistance application", "Target = Resident", "Region = Fengtai District"), generate the original text. The probability of (such as policy provisions or questions from the public) This modeling approach depicts the generative relationship between real-world needs and policy conditions that produces specific textual content. In practical implementation, this can be estimated through manual annotation or weakly supervised methods (such as mapping legal clauses to topics). and in In the model, it is part of the forward model or observation matrix, providing a generative likelihood term for subsequent Bayesian inference.

[0020] The parsing probability kernel of an LLM parser represents the parsing probability of the LLM parser when reading the original text. Then, it was analyzed into structured low-dimensional observation data. The probability, This reflects the LLM parser's understanding and classification confidence of the original text, and is a probability distribution of the mapping from high-dimensional original text to low-dimensional observation data. In one implementation, It originates from the output distribution or calibration probability of the LLM parser, such as the confidence level of the parsing result, temperature scaling, discriminator score, etc., and is obtained after calibration through prompt word chain, template and discriminator model; in the system, it represents the parsing accuracy of the LLM parser, that is, the probability response function of the perception stage.

[0021] and The combination yields a joint observation kernel. This is a joint observation kernel consisting of the latent state, the original text, and the low-dimensional observation data. Original text Analysis into low-dimensional observation data Analytical probability kernel and by potential state To the original text Generation likelihood This combination transforms the LLM parser from a black-box responder into an interpretable sensor, and in Bayesian posterior computation, it integrates the observation kernel. Allows for probabilistic tracing of "why the label or push decision was made".

[0022] S1.3, a single Bayesian update is performed based on the joint observation kernel and the common prior belief of the current observation to obtain the common posterior belief of the current observation. Specifically, in the common prior belief... Below, for single low-dimensional observation data Perform a post-hoc update, as shown in the formula below: ; in, This represents diagonal matrix operations. The joint observation kernel is in diagonal matrix form. A normalized vector consisting entirely of 1s. For public a priori beliefs, Refers to given low-dimensional observation data After the potential state The Bayesian posterior probability, which is also the common posterior belief. In essence, it is an updated estimate of implicit states such as "policy theme / popular profile / demand type" for subsequent action decisions.

[0023] S2, Batch Fusion Posterior: Under the condition of satisfying the preset batch posterior, the batch posterior is performed based on the common prior belief of the current observation and the analytical probability kernels corresponding to multiple historical low-dimensional observation data within the preset batch posterior window to obtain the updated common posterior belief.

[0024] In one implementation, if there is a recent caching of second-low dimensional observation data Then, by using analytical probability kernels for batch fusion, a more robust batch posterior can be obtained. If there is no cache, skip this step. (This refers to batch post-verification.) This refers to the updated common posterior beliefs obtained through batch posterior analysis. The expression for batch posterior analysis is shown below: ; ; During batch posterior ablation, sharing of the most recent L low-dimensional observations is permitted to balance privacy and accuracy. Specifically: It refers to the first The posterior distribution of the state obtained after several observations, i.e., the batch posterior. This indicates that the system obtains the first... Sub-low dimensional observation data After the potential state Probability estimates. It is the Bayesian engine based on the belief of the previous moment. The posterior result obtained by updating the probability kernel; It refers to the first The second-lowest dimensional observation data is processed by the LLM parser from the first... The original text input next time Structured tags or slot combinations extracted from the data, such as theme, region, target audience, timeframe, and process. This represents the system in the [section / stage]. The "interpretable observations" seen by the wheel; This refers to the current observation time or batch index, indicating the current calculation or update time step, which is a low-dimensional observation data sequence. The indexes are typically arranged by time or interaction order; The batch fusion window length (number of sharing times) represents the nearest shared window that the system allows during batch a posteriori fusion. Low-dimensional observation data. The value of balancing privacy and accuracy: The larger the size, the more information is integrated, but the greater the risk of privacy leaks. The smaller the value, the stronger the protection, but the estimated variance increases; This is the lag or backtracking step size index, indicating the number of observation steps to look back during the current batch computation. ).For example Indicates the first The low-dimensional observation data is used to control the weighting or truncation range of observations in the sliding window; For the first The shared prior beliefs before the first observation indicate that the system receives the first... The distribution of beliefs prior to the second observation, i.e., the distribution of beliefs about the potential state. Prior estimates, Combined with the public posterior beliefs from the last update The possible propagation models are formed, which serve as the input for the Bayesian update; For the first The analytical probability kernel corresponding to the low-dimensional observation data of the th order represents the LLM parser in the th order. During the second observation, the original text was parsed into low-dimensional observation data. The probability kernel. It describes the analytical confidence level or calibration probability, and is the kernel term for the analytical probability when calculating the batch posterior.

[0025] In the expression for batch a posteriori, it represents the system at the th... Always combine with the latest Sub-low dimensional observation For potential states Perform fusion inference. Each Provide an analytical probability kernel. The system provides public prior beliefs and ultimately derives new public posterior beliefs, aiming to balance the dual goals of privacy protection and estimation accuracy. The system achieves this by... Based on the integration of recent Subanalysis probability kernel Enable batch post-hoc updates to improve estimation accuracy while maintaining privacy boundaries.

[0026] S3, Action Observation Filtering: Determine whether there is a historical action corresponding to the original text, and if the determination result is yes, perform social learning filtering on the public posterior belief of the current observation based on the historical action to obtain the updated public posterior belief.

[0027] In one implementation, in multi-agent sequential interactions, subsequent agents often only observe preceding actions rather than their private observations. In this case, public beliefs are filtered using action likelihood. ; ; ; ; The above formula transforms "seeing others' actions" back into a Bayesian increment of the state. Specifically, the main symbols are explained below: The updated public posterior belief refers to the distribution of group beliefs updated by Bayesian filtering after the system observes the actions (or batch actions) of others. Let be the belief transfer operator, representing the common posterior belief. In performing the action The subsequent transition function. Within the Markov Decision Process (MDP) framework, Used to update current public posterior beliefs For new public a posteriori beliefs . Let be the expected reward function, representing the current public posterior belief. Next action The expected benefits or costs obtained. For the similarity of the action. This represents the distribution of action strategies under the observed conditions. For low-dimensional observation data The probability kernel is analyzed. These parameters together constitute a Bayesian inference loop for social learning filtering and public belief updating, realizing the mechanism of "inferring state changes through the actions of others".

[0028] S4, Expected Cost and Optimal Action: Based on updated common posterior beliefs and low-dimensional observation data, determine the expected value of actions corresponding to candidate actions. With the objective of minimizing the expected value of actions, determine the optimal action corresponding to the original text from the candidate actions. That is, given LLM parsing of low-dimensional observation data... With public beliefs Then, through the posterior distribution Weighted cost function Find the expected value and select the action that minimizes the expected cost. .

[0029] First, a rational inattention mechanism is introduced for individual agents to determine the optimal observation kernel. Specifically, the expression for the joint optimality of attention and utility is: , , , , Parameter meaning explanation: In the scene The selected observation kernel (attention policy / analysis strategy) represents the kernel selected at the 1st... In various application scenarios (such as different government affairs or different service types), the parsing strategy selected by the intelligent agent to balance information cost and accuracy. It defines low-dimensional observation data. In potential state The probability of it occurring.

[0030] In order to establish public a priori beliefs and low-dimensional observation data The expected utility function under the condition that the system's common prior beliefs are... And obtain low-dimensional observation data At that time, the corresponding state-action pair The weighted utility (or negative cost). Representing a scene Immediate utility function .

[0031] Public a priori beliefs refer to the beliefs that an agent holds about a potential state before it has received any observations. The initial belief distribution. Information cost The reference distribution is also the starting point for Bayesian updates.

[0032] The cost of information acquisition represents the initial belief. The following observation kernel is used The required attention or computational overhead. Mutual information or entropy-based regularization is often used.

[0033] For the scene The joint utility objective under the scenario indicates that in the scenario Given action strategy The overall expected return at that time. It incorporates the task's rewards. With attention cost The trade-off between these factors is the objective function for optimizing the RIBUM model.

[0034] For the scene The conditional parsing kernel under the condition represents the conditional parsing kernel under the condition of the first condition. In a specific environment, for specific low-dimensional observation data The analytical probability distribution.

[0035] The optimal action is defined as the action taken given low-dimensional observation data. and public a priori beliefs Then, the action choice that maximizes expected utility.

[0036] Given low-dimensional observation data With updated public posterior beliefs The optimal action selection at that time satisfies: ; It is used to make action decisions within a framework that prioritizes on-demand response and supplements it with proactive push notifications.

[0037] Given low-dimensional observation data and beliefs The optimal action that minimizes the expected cost. Given low-dimensional observation data After the potential state The posterior. In the potential state Next action The cost. : A set of actions. It is the expected cost, and It is the action that minimizes the expected cost. .

[0038] S5, Threshold determination: Optimal stopping control is adopted to determine whether the updated common posterior belief exceeds the stopping threshold curve; if the determination result is negative, it is determined to continue the evidence search and questioning to obtain new original text, and the agent continues to determine the optimal action corresponding to the new original text; if the determination result is positive, it is determined to stop the evidence search and questioning.

[0039] To suppress information cascading / conformity (actions eventually becoming consistent regardless of new evidence) in sequential interactions, suppose the agent switches between two modes at each step: Sharing private observations (continue to gather evidence and ask follow-up questions); Stop following the group (stop collecting evidence and asking follow-up questions before pushing / publishing).

[0040] in, Indicates the first The first intelligent agent (or the first In group interaction processes, the action choice variable (step decision-making) takes values ​​determined by the prevailing public beliefs. With the system's set stop threshold Decide: ; In other words, The value is determined by whether the current belief exceeds the stopping threshold curve, reflecting the system's optimal switching between the two modes of "continue to search for evidence / interrogate" and "publish / push".

[0041] Specifically, let: : Expected value of continuing the search for evidence; : Expected value for immediate stop and push; discount factor. Then the optimal strategy satisfies: ; Under the monotonicity and TP2 condition (Total Positivity of Order 2), there exists a threshold belief π. This results in a "threshold structure" in the strategy: when the belief strength exceeds a certain threshold, the process immediately stops and the push is initiated; otherwise, the evidence search continues. Therefore: ; Threshold beliefs are obtained through offline training or online learning based on historical data. , Taking into account privacy costs, budget constraints, and the risk of false pushes, this function can be approximated by policy gradients or reinforcement learning algorithms (referred to in the paper as the "policy gradient search threshold"). Thus, The actual decision is determined by the model's adaptive approach, achieving a balanced strategy that prioritizes on-demand response and supplements it with proactive push notifications.

[0042] The parameters have the following meanings: The stopping value function represents the expected social benefit (or negative cost) of immediately stopping and continuing push / publishing when the current public belief is π. It includes factors such as the risk of false pushes, the cost of privacy breaches, and budget consumption.

[0043] The continuing value function represents the expected discounted return at future moments, given the current belief π, if one chooses to continue collecting private observations or seeking clarification. Its calculation includes the expected value of future belief updates.

[0044] / Let be the threshold intervals, representing the two decision regions in the belief space: "stop pushing" and "continue searching for evidence." The boundaries are defined by the threshold function. Confirmed. Satisfied. = The point is the threshold.

[0045] The optimal stopping threshold is defined as making = The point of belief. When If the expected benefit of the push action exceeds the benefit of continued observation, the push should be executed immediately; otherwise, evidence collection should continue. It can be learned through policy gradient or dynamic programming algorithms.

[0046] Let be the action variable, representing the first . Action decisions for an agent at any given moment: The representative continued to collect evidence. This represents a halt and push. It is based on current public beliefs. Does it exceed the threshold? Decide.

[0047] , , Representing privacy / budget / latency cost coefficients, these coefficients quantify the system's costs in information sharing, resource consumption, and service timeliness, respectively; these costs are related to... , Linkage is used to construct overall social welfare goals.

[0048] The threshold switching curve in this invention is not an independent parameter, but rather determined by the stopping value function. With continuing value function The belief space boundary curve is derived from the equality condition. This curve is used to distinguish between the two decision regions of "continue to search for evidence / interrogate" and "stop and push," and is the geometric representation of the optimal stopping strategy in the belief space.

[0049] Furthermore, the above discounts The construction process is as follows: Constructing group social welfare goals (discount) ), and with public beliefs Design a threshold strategy for sufficient statistics: ; in, Let be the objective function. As a strategy, Regarding strategy Expectations For discrete time steps, For system status, For the first The information set of the step, The expected condition is indicated by... The given information pertains to the state-action cost at step k. Expectations Let be the conditional probability, representing the probability based on ,state equal The probability, The state is not in The probability, The discount factor is a coefficient used to discount future returns, typically ranging from (0,1), where d is the value relative to the state. The relevant penalty coefficient, Let the state-action cost function represent the state-action cost. Take action The immediate costs incurred at that time For state non The relevant penalty coefficient.

[0050] This part belongs to the "Optimal Stopping" stochastic control model. During the sequential interaction process, at each time step k, the system needs to make a decision between "continuing to collect evidence" and "stopping and publishing the push notification". To achieve the optimal balance between privacy, budget, and push notification accuracy, this embodiment of the invention defines a social welfare objective function, using the discount factor ρ and public belief π as core variables, and designs an optimal strategy based on a threshold.

[0051] Based on the previous embodiments, the embodiments of the present invention will have the following branches: Branch 1: If the public posterior belief determines to continue the search for evidence and further questioning, generate clarifying questions / trigger deeper searches, obtain new evidence, and return to steps S1 to S5 for iteration until the threshold is crossed.

[0052] Branch Two: If, based on shared posterior beliefs, evidence gathering and further questioning cease, the target content corresponding to the original text is determined by matching structured tags from a pre-built searchable evidence base using the joint observation kernel corresponding to the original text. Specifically, under high-confidence π, the most suitable "policy interpretation / process card / outcome item" is selected based on joint observation kernel / slot matching, ready for external output.

[0053] S6, Compliance and Risk Mitigation; S7, Delivery and Logs: Push the target content corresponding to the original text to the terminal when the target push content meets the risk and security constraints and engineering mapping constraints.

[0054] (1) Risk and safety constraints (taking CVaR as an example): To constrain tail events such as "false push / illusion risk" or "group unfairness," a uniform risk metric or CVaR constraint is used, as shown in the following expression: , in For loss variables; Indicates system loss Exceeding the threshold That portion of the risk; The expected excess loss represents the value of the system when it loses money. Exceed The expected excess at that time is used to quantify the degree of tail risk. Confidence level; To optimize variables; This represents the upper limit of risk tolerance.

[0055] (2) Engineering mapping constraints: Latency: Constraints on single response time , This is the actual time taken for a single response. This is the threshold for the time taken for a single response.

[0056] Budget / Token: Constraints on the number of tokens called / generated , This represents the actual amount of tokens consumed. The token budget threshold.

[0057] Hallucination risk: For non-fact confidence / discriminator scores; use probability constraints or CVaR.

[0058] Fairness: A measure of the difference in rewards / errors between groups Set group constraints , For the first The actual measure of group fairness For the first Group fairness threshold.

[0059] Compliance: Penalties for triggering privacy / confidential content .

[0060] The core idea of ​​step S6 is to maximize revenue while introducing a "Unified RiskMeasure" or CVaR (Conditional Value at Risk) constraint to control potential tail risks, including extreme events such as false push notifications, illusory output (non-factual content), or group fairness bias. After passing the S6 verification, the system distributes the product across multiple platforms (agent / app / SMS, etc.), collecting action logs and constraint indicators (latency, token, risk score) for clicks, dwell time, and conversions, providing data for subsequent learning and auditing.

[0061] S8, Policy Learning: Construct a Lagrange regularized objective based on the log information corresponding to the target push content, and use the Lagrange regularized objective to perform original-dual updates on the agent.

[0062] Specifically, based on the Lagrange regularization objective For strategy parameters Perform gradient ascent, adjust multipliers Perform gradient ascent (project to non-negativity) and use (Entropy / norm) steady-state regularization; single timescale update, satisfying last-iterate convergence and zero dual gap (Slater).

[0063] The principle is as follows: The system's "push / follow-up question / manual transfer / drill-down retrieval" strategy is modeled as a Constrained Markov Decision Process (CMDP). The goal is to maximize revenue while satisfying multiple constraints such as service latency, budget, illusion risk, and fairness. These constraints include service latency, budget (token or call volume), illusion risk, fairness, and compliance. CMDP achieves strategy optimization and constraint balance through a constrained policy gradient method.

[0064] CMDP is defined as: ; ; ; in, For expectation operators; Discount factor; For immediate reward function, it represents the state variable. Execution action variable The immediate rewards obtained, such as push accuracy and user satisfaction; For the first The constraint cost function represents the constraint cost of the system under this state and action, such as latency cost, computing power consumption, illusion risk, unfairness or privacy penalty; For the first Class constraint upper limit; For the first The state-action cost function of class constraints represents the state-action cost at time step [number]. Executing action variables under state variables The immediate constraint cost obtained.

[0065] This invention embodiment can simultaneously perform strategy updates ( θL) and constraint weight update ( λL), thereby obtaining the optimal strategy under multiple constraints. This CMDP modeling form enables the system to maximize the benefits of government services while meeting the time, budget, and security requirements under multiple constraints, and achieves dynamic optimization and global convergence through subsequent (D2) primitive-dual policy gradients.

[0066] Lagrange regularization objective and primitive-dual update: , , . in, For Lagrange functions, For strategy parameters, Let Lagrange multiplier vector be the vector. Let the objective function (reward) of the strategy be denoted as . For the first Cost function with constraints For the first The threshold of each constraint, The regularization coefficient is . For regularization terms, , For the first Step and the first The strategy parameters for each step, For projection operators, The learning rate is a policy parameter. For Lagrange functions gradient, , For the first Step and the first Step 1 One Lagrange multiplier, The learning rate is the Lagrange multiplier.

[0067] The original-dual method has zero duality gap under the Slater condition, and compared to the dual-timescale method, the embodiments of this invention can also achieve last-iterate guarantee under a single timescale (compared to the results of only average iteration or dimension dependence). In the above formula, Let be the Lagrange regularization objective function, used to control the cost of multiple constraints while maximizing service benefits; Indicates the strategy parameters, Let Lagrange multiplier vectors be used. For regularization terms, The discount occupancy metric is used. This design ensures that the government affairs push system achieves stable and optimal operation while meeting service latency, budget, risk, and fairness constraints.

[0068] S9, Behavior Audit; S10 Attention Cost and Weight Closed-Loop Correction: Determine whether the log information corresponding to the target push content meets the consistency conditions of NIAS and NIAC, and if the determination result is negative, perform inverse reconstruction of utility and attention cost for the agent.

[0069] Specifically, S9: Use the "observation → action" log to verify whether NIAS / NIAC holds true; if inconsistent, perform inverse reconstruction of utility / attention cost to pinpoint the source of the bias. S10: Based on the conclusion of S9, adjust the attention cost weight λ downwards / upwards or adjust accordingly. The (information cost) and threshold are fed back to the parsing depth and stopping strategy in S1 to S5, forming an adaptive closed loop. The principle is as follows: Behavioral consistency can be verified (NIAS / NIAC): if only black-box observable agent behavior data is available. Whether a system meets RIBUM can be determined using No Improvement Action Transition (NIAS) and No Improvement Cycle (NIAC), and utility can be reconstructed accordingly, facilitating compliance audits and deviation diagnosis. ; ; Explanation: The two sets of inequalities are necessary and sufficient conditions, and the basis for the utility reconstruction algorithm is given. This indicates that low-dimensional observations can be explained. Indicates the action of the intelligent agent. For potential real state; It is a utility function; For observation of the nucleus; For prior beliefs; In order to observe The probability of the action at that time; Pay attention to cost weighting; For information acquisition costs, For the scene The immediate effect of the following For the scene The probability of the following behavior, For the scene Attention costs, For the scene The cost of attention.

[0070] The parameters of the two sets of inequalities, NIAS (No Improved Action Transition) and NIAC (No Improved Cycle), , This is used to determine whether there is an improvement action or a cyclical improvement. If all inequalities are true, then there exists a set of... A Bayesian utility maximization model that enables agents to satisfy rational attention proves that their behavior is consistent and auditable; otherwise, deviation diagnosis and utility reconstruction can be performed based on the deviation terms.

[0071] In the "Behavioral Consistency Testable (NIAS / NIAC)" section, this invention introduces two sets of inequalities: No Improving Action Switch (NIAS) and No Improving Action Cycle (NIAC). These are used to verify whether a system that can only observe input-output behavior (i.e., a black-box agent) can be mathematically interpreted as a Bayesian utility-maximizing decision-maker under rational attention. If these inequalities are satisfied, it indicates the existence of a latent set of utility functions. Note the cost function ;Observation kernel Cost weight This makes the agent's decision-making behavior consistent with RIBUM theory, which can be used to reconstruct its potential utility and attentional preferences for fairness and compliance auditing.

[0072] In summary, this invention uses "Single Intelligent Agent = LLM Sensor + Bayesian Engine + Rational Attention", "Group Interaction = Bayesian Social Learning + Optimal Stopping Control", and "Policy Optimization = Constrained Policy Gradient and Primitive-Dual Update" as the three-layer kernel of the system. This system is designed for community scenarios (residents + grid workers / community workers). Its functions include: Question Answering: Residents or staff ask questions (e.g., "How do I apply for minimum living allowance?"), the system first understands the question, and then provides actionable application guidance; Push Notifications: The system can accurately push matching scientific and technological achievements / policies / process cards to appropriate groups "primarily on demand, secondarily proactively". Its key idea is to treat the Large Language Model (LLM) as an interpretable social sensor, rather than a black box: text is parsed from high-dimensional original text w (policy provisions, public demands) into low-dimensional interpretable features y (topic, region, target, timeliness, process, etc.). The system maintains a probabilistic belief π (prior / public belief) about real needs / policy topics. Upon encountering new y, π is updated to π+ using a Bayesian method. When making decisions, the system compares the benefits and costs (privacy, budget, latency, error risk, fairness, etc.) of "continuing to search for evidence / follow up" versus "directly pushing the request." It only pushes the request when the probability of success is high enough; otherwise, it first follows up or drills down to investigate further. In long-term operation, the system treats decisions such as "pushing / following up / transferring to human intervention / deepening the investigation" as a "constrained reinforcement learning problem" for continuous optimization, ensuring service quality while simultaneously meeting the goals of latency, budget, compliance, and fairness.

[0073] This invention, aimed at modernizing grassroots government governance, proposes a large-scale intelligent system integrating precise delivery of scientific and technological achievements to grassroots governance. The system primarily uses on-demand response, supplemented by proactive delivery, to accurately match and deliver scientific and technological achievements (policies, guidelines, projects, patents, and standards, etc.) to grassroots staff and residents. To address the pain points of existing manual retrieval and item-by-item reply methods, such as low efficiency, inconsistent quality, poor interpretability, and difficulty in quantifying governance fairness, this invention achieves the following technical implementation: Explanatoryness and consistency: The Large Language Model (LLM) is modeled as an interpretable social sensor, outputting low-dimensional interpretable features and feeding them into a Bayesian engine for posterior inference and action selection, ensuring that the decision-making process is transparent and auditable.

[0074] Anti-information cascade (herding) and robust group interaction: Introduce stochastic control with optimal stopping and incentive / privacy-cost trade-off in multi-agent sequential interaction to delay cascade and suppress herding error.

[0075] Policy learning with guaranteed service quality: Introducing constrained policy gradient (CMDP) and original-dual update, and providing a policy optimization process with global convergence at the final iteration point under multiple constraints such as "latency / budget / fairness / illusion risk".

[0076] Based on the foregoing embodiments, this invention provides a device for intelligent technology achievement push and public service Q&A, see [link to related document]. Figure 3 The schematic diagram shown illustrates the structure of a device for intelligent technology achievement push and public service Q&A, which includes the following parts: The text receiving module 302 is used to receive raw text sent by the terminal; The content push module 304 is used to determine the public posterior belief based on the public prior belief of the current observation, the original text and its corresponding low-dimensional observation by an intelligent agent based on a large language model and a Bayesian engine. It is used to determine and execute the optimal action corresponding to the original text, and push the target push content corresponding to the original text to the terminal when the search for evidence and further questioning are stopped based on the public posterior belief. The target push content includes scientific and technological achievement items or public service response text.

[0077] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0078] This invention provides an electronic device, specifically, the electronic device includes a processor and a memory; the memory stores a computer program, which, when run by the processor, executes the method described in any of the above embodiments.

[0079] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 40, a memory 41, a bus 42, and a communication interface 43. The processor 40, the communication interface 43, and the memory 41 are connected through the bus 42. The processor 40 is used to execute executable modules, such as computer programs, stored in the memory 41.

[0080] The memory 41 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 43 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0081] Bus 42 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0082] The memory 41 is used to store programs. After receiving an execution instruction, the processor 40 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 40 or implemented by the processor 40.

[0083] Processor 40 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 40 or by instructions in software form. Processor 40 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 41. The processor 40 reads the information in memory 41 and, in conjunction with its hardware, completes the steps of the above method.

[0084] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.

[0085] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0086] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for intelligent delivery of scientific and technological achievements and Q&A in public services, characterized in that, include: The raw text sent by the receiving terminal; An intelligent agent based on a large language model and a Bayesian engine determines a common posterior belief based on the common prior belief of the current observation, the original text and its corresponding low-dimensional observation. This belief is used to determine and execute the optimal action corresponding to the original text. If the agent determines to stop searching for evidence and asking follow-up questions based on the common posterior belief, it pushes the target content corresponding to the original text to the terminal. The target content includes scientific and technological achievements or public service response text.

2. The method for intelligent push of scientific and technological achievements and public service Q&A according to claim 1, characterized in that, Based on the common prior beliefs of the current observation, the original text, and its corresponding low-dimensional observations, common posterior beliefs are determined, including: The original text is parsed using a large language model to obtain low-dimensional observation data, and a joint observation kernel is constructed based on the original text and the low-dimensional observation data. A single Bayesian update is performed based on the joint observation kernel and the common prior belief of the current observation to obtain the common posterior belief of the current observation. Determine whether there is a historical action corresponding to the original text, and if the determination result is yes, perform social learning filtering on the public posterior belief of the current observation based on the historical action to obtain the updated public posterior belief.

3. The method for intelligent push of scientific and technological achievements and public service Q&A according to claim 1, characterized in that, The method further includes: Under the condition of satisfying the preset batch posterior, the updated common posterior belief is obtained by performing batch posterior based on the common prior belief of the current observation and the analytical probability kernels corresponding to multiple historical low-dimensional observation data within the preset batch posterior window.

4. The method for intelligent push of scientific and technological achievements and public service Q&A according to claim 1, characterized in that, Determine and execute the optimal action corresponding to the original text, including: Based on the updated public posterior beliefs and the low-dimensional observation data, the expected values ​​of actions corresponding to candidate actions are determined. With the goal of minimizing the expected values ​​of actions, the optimal action corresponding to the original text is determined from the candidate actions.

5. The method for intelligent push of scientific and technological achievements and public service Q&A according to claim 1, characterized in that, The method further includes: Optimal stopping control is employed to determine whether the updated common posterior belief exceeds the stopping threshold curve; If the result is negative, the investigation and questioning will continue to obtain new original text, and the agent will continue to determine the optimal action corresponding to the new original text. If the determination result is yes, then stop the search for evidence and further questioning.

6. The method for intelligent push of scientific and technological achievements and public service Q&A according to claim 5, characterized in that, If, based on the aforementioned public posterior belief, the evidence gathering and questioning are halted, the target content corresponding to the original text is pushed to the terminal, including: Based on the joint observation kernel corresponding to the original text, the structured tags in the pre-built searchable evidence base are matched to determine the target push content corresponding to the original text; If the target content satisfies the risk and security constraints and the engineering mapping constraints, the target content corresponding to the original text is pushed to the terminal.

7. The method for intelligent push of scientific and technological achievements and public service Q&A according to claim 1, characterized in that, The method further includes: Based on the log information corresponding to the target push content, a Lagrange regularized target is constructed, and the agent is updated using the Lagrange regularized target. And / or, determine whether the log information corresponding to the target push content meets the NIAS and NIAC consistency conditions, and if the determination result is negative, perform inverse reconstruction of the utility and attention cost of the agent.

8. A device for intelligent technology achievement push and public service Q&A, characterized in that, include: The text receiving module is used to receive raw text sent by the terminal; The content push module is used to determine the public posterior belief based on the public prior belief of the current observation, the original text and its corresponding low-dimensional observation by an intelligent agent based on a large language model and a Bayesian engine. It is used to determine and execute the optimal action corresponding to the original text, and push the target push content corresponding to the original text to the terminal when it is determined to stop the search for evidence and further questioning based on the public posterior belief. The target push content includes scientific and technological achievement items or public service response text.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent customer service question and answer method based on large language model technology

    CN118364084A

  • Data processing method and device, information recommendation system, electronic equipment and medium

    CN120011550A

  • Intelligent agent method and device based on intelligent language model, equipment and medium

    CN120763305A