Information processing method and device

Through online reinforcement learning and interaction model optimization, the intelligent customer service system can adapt to user needs in real time, solving the problem of insufficient flexibility in existing systems and achieving more efficient user response and improved service quality.

CN121833790APending Publication Date: 2026-04-10ZHEJIANG CAINIAO SUPPLY CHAIN MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511631869.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing intelligent customer service robot systems lack flexibility, cannot adapt to user needs in real time, and have weak real-time performance and learning capabilities, resulting in an inability to generate satisfactory responses.

Method used

By employing online reinforcement learning methods, user interaction data is recorded in real time, external tools are called to obtain information, and reinforcement learning algorithms such as GRPO are used to optimize the interaction model, thereby achieving self-updating and improved adaptability of the model.

Benefits of technology

This has enabled the intelligent customer service system to learn and adapt, improving response accuracy and user satisfaction, and ensuring continuous improvement in service quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833790A_ABST
    Figure CN121833790A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an information processing method and device, and the method comprises the steps: obtaining a logistics information processing request, and determining a processing type corresponding to the logistics information processing request; calling a target interface corresponding to the processing type, obtaining logistics information associated with the logistics information processing request, inputting the logistics information and the logistics information processing request into an interaction model for processing, and obtaining multiple pieces of candidate feedback information; according to user feedback information associated with the logistics information processing request, target feedback information is determined in the multiple pieces of candidate feedback information, and the interaction model is optimized into a target interaction model according to the target feedback information. The external information and the user query are input into the model together, so that the model can be combined with the external information for model prediction, the model prediction accuracy is improved, the model is optimized online through the feedback information, the model has the self-updating capability, and the model adaptability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of artificial intelligence technology, and in particular to information processing methods and apparatus. Background Technology

[0002] With the continuous development of artificial intelligence technology, intelligent customer service robots have been widely applied in e-commerce, logistics, finance, and technical support, becoming a key tool for enterprises to improve service efficiency and reduce operating costs. Traditional intelligent customer service systems are typically built based on predefined rule bases or static knowledge bases, generating responses to user queries through pattern matching or simple decision trees. In addition, some systems employ machine learning models trained offline on historical data (such as classification models or early dialogue generation models) in an attempt to achieve more natural interactions. However, current systems still suffer from insufficient flexibility, failing to generate or return satisfactory results to users. Therefore, providing users with a more intelligent and flexible information generation method is a pressing issue that needs to be addressed. Summary of the Invention

[0003] In view of this, embodiments of this specification provide information processing methods. One or more embodiments of this specification also relate to information processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address technical deficiencies in the prior art.

[0004] According to a first aspect of the embodiments of this specification, an information processing method is provided, comprising: Obtain a logistics information processing request and determine the processing type corresponding to the logistics information processing request; Call the target interface corresponding to the processing type to obtain the logistics information associated with the logistics information processing request; The logistics information and the logistics information processing request are input into the interaction model for processing to obtain multiple candidate feedback information; Based on the user feedback information associated with the logistics information processing request, the target feedback information is determined from the plurality of candidate feedback information, and the interaction model is optimized into the target interaction model based on the target feedback information.

[0005] According to a second aspect of the embodiments of this specification, another information processing method is provided, including: Receive a target logistics information processing request submitted by a target user, and determine the target processing type corresponding to the target logistics information processing request; Call the target interface corresponding to the target processing type to obtain the target logistics information associated with the target logistics information processing request; The target logistics information and the target logistics information processing request are input into the target interaction model for processing to obtain target feedback information, and the target feedback information is fed back to the target user. The target interaction model is obtained through the above method.

[0006] According to a third aspect of the embodiments of this specification, an information processing apparatus is provided, comprising: The acquisition module is configured to acquire logistics information processing requests and determine the processing type corresponding to the logistics information processing requests; The calling module is configured to call the target interface corresponding to the processing type to obtain the logistics information associated with the logistics information processing request; The input module is configured to input the logistics information and the logistics information processing request into the interaction model for processing, and obtain multiple candidate feedback information. The determination module is configured to determine the target feedback information from among the multiple candidate feedback information based on the user feedback information associated with the logistics information processing request, and optimize the interaction model into the target interaction model based on the target feedback information.

[0007] According to a fourth aspect of the embodiments of this specification, another information processing apparatus is provided, comprising: The receiving module is configured to receive a target logistics information processing request submitted by a target user and determine the target processing type corresponding to the target logistics information processing request. The acquisition module is configured to call the target interface corresponding to the target processing type to obtain the target logistics information associated with the target logistics information processing request. The processing module is configured to input the target logistics information and the target logistics information processing request into the target interaction model for processing, obtain target feedback information, and provide the target feedback information to the target user, wherein the target interaction model is obtained through the above method.

[0008] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described information processing method.

[0009] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the information processing method described above.

[0010] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the information processing method described above.

[0011] This specification provides an information processing method that, based on the processing type of a logistics information processing request, calls the target interface corresponding to the processing type to obtain the associated logistics information for the request. The logistics information and the logistics information processing request are input into an interactive model for processing, enabling the model to combine external information with user queries for prediction, thus improving prediction accuracy. Based on user feedback associated with the logistics information processing request, the target feedback information is determined from multiple candidate feedback information output by the model. This target feedback information is then used to optimize the interactive model into the target interactive model, achieving dynamic model optimization based on user feedback information. This gives the model self-updating capabilities, improves its adaptability, and further enhances the user experience of the request feedback function. Attached Figure Description

[0012] Figure 1A A schematic diagram of the system structure of an information processing method according to an embodiment of this specification is shown; Figure 1B A flowchart of an information processing method according to an embodiment of this specification is shown; Figure 2 A flowchart of another information processing method provided according to one embodiment of this specification is shown; Figure 3 A flowchart illustrating the processing procedure of an information processing method according to an embodiment of this specification is shown. Figure 4 This specification shows a schematic diagram of the structure of an information processing apparatus according to one embodiment. Figure 5 This specification shows a schematic diagram of the structure of another information processing apparatus provided in one embodiment; Figure 6 A structural block diagram of a computing device provided according to one embodiment of this specification is shown. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0017] The technical solutions provided in this application can employ deep learning models with relatively large parameter scales. However, this large model is merely an example; this application does not limit the number of model parameters supported by the deep learning model used, aiming to meet actual needs. The deep learning models involved in this application can be artificial intelligence-based language models (LM) or multimodal models (MM).

[0018] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0019] Online reinforcement learning is a machine learning method that allows a system to learn and optimize its policies in real time during actual operation. It obtains feedback through interaction with the environment and continuously adjusts and improves its decisions.

[0020] Tool Invocation: The process by which a customer service robot requests data or services from an external system, the results of which are used to assist in decision-making or generate responses.

[0021] User feedback: The evaluation provided by users to the robot's responses, which is used to adjust the robot's learning strategy and improve response quality.

[0022] Conversation Log: Records detailed information about user interactions with the bot, including user queries, bot responses, and user feedback.

[0023] Intelligent Customer Service: A system that uses artificial intelligence technology to improve the efficiency and intelligence of customer service.

[0024] Real-Time Update: The ability to dynamically adjust system models or strategies based on the latest data and feedback.

[0025] With the rapid development of intelligent customer service technology, its application in the logistics fulfillment chain is becoming increasingly widespread. However, many existing customer service robots still rely on static knowledge and predefined rules when processing user queries, making them unable to adapt to rapidly changing user needs, and exhibiting the following three shortcomings: Poor adaptability: Because the robot's response relies on pre-set rules and answers, it is unable to flexibly handle unexpected situations or users' personalized needs.

[0026] Insufficient real-time capability: Many robots cannot access external information in real time and cannot answer user questions promptly.

[0027] Weak learning ability: In traditional methods, the process of updating the knowledge base of robots is often time-consuming and laborious, and they cannot learn from user feedback in a timely manner and optimize their own answering strategies.

[0028] Based on this, an information processing method is provided in this specification. This specification also relates to an information processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0029] See Figure 1A , Figure 1A A schematic diagram of the system architecture of an information processing method according to an embodiment of this specification is shown. The system architecture of this information processing method includes three parts: data feedback / admission, online reinforcement learning, and deployment, evaluation, and online execution. The following explanation uses a customer service scenario as an example to illustrate this system architecture.

[0030] The data feedback / admission section includes the current human online service logs. These logs are the system's data source, recording the complete dialogue history between real online users and the customer service robot (or human customer service representative), including user queries, model responses (LLM responses), and crucial user feedback (e.g., satisfaction / dissatisfaction) or signals of human customer service takeover (delegation of management services). Interaction data streamed from the online logs is synchronized and temporarily stored in real-time via Redis / OSS (a high-speed caching / storage system), serving as a data buffer. The system periodically polls data from the buffer and performs data feedback / admission checks. Only data that meets certain criteria (e.g., clear feedback, low overlap with existing data) is added to the Training Data set.

[0031] The Tools section in the online reinforcement learning component can be considered external tools for the model. These tools can provide external knowledge to the model by leveraging log data through API interfaces or RAG retrieval enhancement systems. The model invokes these external Tools in real-time to retrieve information based on user queries. The online-model-ref (the current online benchmark model) is trained on the training set. During training, the RL-Policy (reward function) can be used to score new responses based on user feedback. For example, a high score is given if the semantics of a new response are similar to historically positive responses; otherwise, a low score is given. The trained model is then passed to the save-Best-checkpoint.

[0032] In the deployment and evaluation phase, the model checkpoint compares the performance of the newly trained model (latest model) with the benchmark model (online-model-ref) on a unified test set to determine if the new model is superior. If the new model wins the evaluation, the system will automatically deploy and complete the model switch, allowing the better-performing new model to begin serving real users.

[0033] In summary, the above system enables real-time recording of interactions between the online service model and users in the service logs. Log data is instantly synchronized to Redis / OSS. The system periodically checks this data, admitting high-quality interaction data (such as those with clear feedback) into the Training Data. The online reinforcement learning module interacts with the environment using the current online model. It receives user queries, obtains real-time information through tools, and generates candidate responses. Based on feedback from historical Training Data, these candidate responses are scored. The GRPO algorithm uses these scores to update and optimize the model's policy parameters, generating a better new model. The trained new model (latest model) is automatically compared and evaluated with the current online benchmark model (online-model-ref). If the evaluation results show that the new model performs better, the system automatically triggers the service deployment process, replacing the old online model with the new model. After the new model goes live, new interaction logs are generated again. This new data is collected, filtered, and used for the next round of training and optimization, thus forming a continuous self-iterative and self-improving closed loop.

[0034] Based on this, by seamlessly integrating online interaction, tool calls, reinforcement learning, and automated operation and maintenance, an intelligent customer service system with a lifecycle has been built, which can continuously learn and evolve during use, ultimately achieving continuous improvement in service quality and efficiency.

[0035] See Figure 1B , Figure 1B A flowchart of an information processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0036] Step S102: Obtain the logistics information processing request and determine the processing type corresponding to the logistics information processing request.

[0037] In this context, a logistics information processing request can be understood as any question or instruction from a user to the intelligent customer service robot related to the logistics fulfillment process. The logistics information processing request is the trigger signal for the entire process. It can include user queries and is an important part of the conversation log. The content of logistics information processing requests is usually in natural language, such as: "Where is my package?", "Could you urge the courier for me?", "I want to complain about a damaged package". The processing type can be understood as the specific operation category categorized by the system after semantic understanding and intent recognition of the user request. It determines which external tools (APIs) the robot needs to call subsequently and what response strategy to adopt. The processing type corresponding to the logistics information processing request directly maps to specific "tool calls". For example, a processing type of "logistics status query" directly maps to the "logistics details API", and subsequently, the "logistics details API" can be directly called to respond to the logistics information processing request.

[0038] In practical applications, intelligent customer service robots receive user queries (such as logistics status inquiries) and record the process in real time, including the user's query content, the robot's initial response, and user feedback. After receiving a user query, the intelligent customer service robot determines the processing type of the query and selects the appropriate API to call to process the query. For example, it might call the logistics details API to retrieve the current location of the package, the work order API to help the user submit a work order, the delivery expediting API to help the user expedite the delivery, or the package damage detection API to check for damage in the signed package photos. At this point, the user query, the model's response, and any user feedback need to be stored in a database (such as Redis or OSS). Data that meets certain conditions (e.g., within a controllable range of overlap with existing database data, clear user feedback (satisfaction, dissatisfaction, whether the problem is resolved, or whether gratitude is expressed)) is used as training samples for the next step of the model (data that does not meet these conditions is not used as training samples).

[0039] In a specific embodiment of this specification, suppose a user sends the following message to the customer service robot: "The corners of the box I received yesterday are dented, and the contents seem to be damaged as well. What should I do?" The system receives this text message from the user and identifies it as a valid "logistics information processing request." The system performs Natural Language Processing (NLP) and intent recognition on this message, analyzing the keywords "dented box" and "damaged contents," thus determining that the user's core intent is to report damaged goods. Therefore, the system classifies this request as a "package damage report." Subsequently, an after-sales service ticket can be created for the user by calling the ticket API, thereby providing a corresponding response and solution to the user's message.

[0040] Therefore, converting user needs into machine-executable instructions through intent recognition is the cornerstone for the entire system to achieve high efficiency, accuracy, and adaptability.

[0041] Step S104: Call the target interface corresponding to the processing type to obtain the logistics information associated with the logistics information processing request.

[0042] Since the processing type corresponding to the logistics information processing request has been determined, the corresponding target interface can be called based on the processing type. The target interface can be understood as a backend service API that corresponds one-to-one with the processing type and can complete that type of task. The target interface can also be understood as the API interface of the tool to be called based on the processing type, enabling the system or model to interact with the external world to obtain real-time data and perform operations. By calling the target interface, logistics information related to the logistics information processing request can be obtained. Logistics information can be understood as all structured data results or operational feedback related to the user request returned by calling the target interface. It is not only the narrow definition of "logistics trajectory" but also includes status reports after the operation is performed. Logistics information may include query results such as the current location of the package and the estimated delivery time; and / or operational feedback such as the work order has been successfully created and a work order number has been returned, and an order reminder instruction has been issued.

[0043] In a specific embodiment of this specification, referring to the example above, the processing type is determined to be "Reporting a Damaged Package". Based on the processing type "Reporting a Damaged Package", the system determines that the target interface to be called is the "Work Order API". Subsequently, the system constructs a request (possibly automatically filled with user identity, order number, etc.) and initiates a call to this API. The "Work Order API" processes the request in the backend, successfully creating an after-sales work order and returning a structured response, i.e., logistics information. Its content may be: {"status":"success","ticket_id":"123","message":"Work order created; customer service will contact you within 24 hours."}. The logistics information and the logistics information processing request can then be input into the interactive model to obtain multiple candidate feedback messages, which are used to determine the training data for the training model.

[0044] Based on this, by calling the target interface, the logistics information of the associated request can be obtained, so that the input of the subsequent model can include new external knowledge, enabling the model to output more accurate and reliable results.

[0045] Step S106: Input the logistics information and the logistics information processing request into the interaction model for processing to obtain multiple candidate feedback information.

[0046] The interaction model can be understood as a core intelligent agent, typically a large language model. It possesses the ability to perform tasks, understand language, and interact. The interaction model can integrate tool invocation functions and online reinforcement learning strategies, specifically as an intelligent customer service robot. Through processing by the interaction model, multiple candidate feedback messages can be output. These candidate feedback messages can be understood as multiple potential responses generated by the interaction model for the same set of inputs (request + information), with different wording, styles, or emphases. This forms the basis for the subsequent "optimization" in the reinforcement learning optimization stage.

[0047] In practical applications, after acquiring real-time information, the system can input the real-time information queried and retrieved by the user into the interaction model, which then generates multiple candidate responses. Specifically, the interaction model can employ reinforcement learning algorithms to optimize the generation of candidate responses. For example, the optimization objective of the interaction model can be expressed as:

[0048] Where Rt is the cumulative reward at time t, rt+k is the immediate reward obtained at time t+k (e.g., based on the quality of user feedback), and γ is the discount factor.

[0049] In a specific embodiment of this specification, referring to the example above, the system inputs the logistics information processing request ("The corners of the box I received yesterday were dented...") and the logistics information (the JSON result of a successful work order creation) into the interaction model. The interaction model understands the user's original request and, based on the fact that a work order was successfully created, generates creative text. The model generates multiple candidate feedback messages, such as: Candidate 1: "Work order 123 has been created for you, and customer service will contact you within 24 hours to handle the damage issue." Candidate 2: "We are very sorry that you encountered this problem! We have entered a work order (number 123) for you, and a specialist will contact you as soon as possible to verify and handle it. Please rest assured." Subsequently, appropriate candidate feedback can be selected as training data based on user feedback to optimize the interaction model, improve its accuracy, and meet the user's needs.

[0050] Based on this, generating multiple candidate feedbacks is equivalent to providing the learning algorithm with a "set of actions to be evaluated." The system can then select the best solution from these candidates using a subsequent reward function, thereby achieving iterative optimization of the model strategy. Since all candidate feedbacks are based on the same, real logistics information, the accuracy of the content is guaranteed.

[0051] Step S108: Based on the user feedback information associated with the logistics information processing request, determine the target feedback information from the plurality of candidate feedback information, and optimize the interaction model into the target interaction model based on the target feedback information.

[0052] User feedback can be understood as the most crucial reward signal in the optimization process. User feedback refers to users' evaluations of historical responses, directly reflecting the quality of the response. User feedback includes explicit and implicit feedback. Explicit feedback includes "Satisfied / Dissatisfied" buttons and five-star ratings. Implicit feedback includes whether the user expresses gratitude in subsequent conversations, whether the problem is resolved, and whether the user immediately requests a transfer to a human operator (management by a customer service representative). Based on user feedback, target feedback can be determined from multiple candidate feedback options. The target feedback is the superior response selected from these candidates. Selection criteria can include the target feedback's potential to elicit positive user feedback and its potential to be the final answer returned to the user. Therefore, target feedback can be used to optimize the interaction model, resulting in a target interaction model that directly outputs results that elicit positive user feedback. The target interaction model is the optimized and updated model, which tends to generate high-quality responses like the target feedback.

[0053] In practical applications, user feedback can be used as reward information to optimize reinforcement learning, and Generalized Regression Policy Optimization (GRPO) can be used for policy updates. The principle of GRPO lies in improving the policy based on the value function, constructing an enhanced policy objective at each update. Using the policy gradient method, it can be expressed as:

[0054] Where πθ represents the policy, θ is the parameter, and τ is the execution trajectory. By maximizing the expected reward J(θ), we can obtain a more optimized action selection policy. The reward function includes the following policies: (1) Referencing the training data, calling APIs (such as GPT5, Qwen-MAX, and other SOTA closed-source models) to score candidate responses. Based on the user response feedback in the training data, if the feedback is good, a higher score is given; otherwise, a lower score is given. If the feedback is poor, the candidate response is close to the robot's semantics and expression in the training data, and a lower score is given. (2) Other format constraints and tool call specification constraints on the reward.

[0055] In practical implementation, in addition to GRPO, we can also integrate other derived reinforcement learning algorithms into this scheme to enhance the robot's learning ability in complex environments, for example: (1) TRPO (Trust Region Policy Optimization): TRPO ensures that the new policy does not deviate too far from the old policy when updating the policy, reducing instability. By introducing the constraint of the trust region, the optimization objective is:

[0056] The constraint (subject to) is:

[0057] Where δ is the tolerance KL divergence.

[0058] (2) PPO (Proximal Policy Optimization): PPO balances exploration and exploitation by limiting the magnitude of policy updates and introducing pruning during update execution. The optimization objectives are as follows:

[0059] Where rt(θ) is the probability ratio of the new and old strategies, At^ is the advantage function, and ϵ is the pruning range.

[0060] (3) DDPG (Deep Deterministic Policy Gradient): Suitable for problems involving continuous action spaces. DDPG uses both a policy network and a value network, learning through an actor-critic architecture. The update steps typically include the following:

[0061] Where J(θ) is the objective function of the policy.

[0062] (4) A3C (Asynchronous Actor-Critic): A3C uses multi-threaded asynchronous training, which can reduce training variance and speed up convergence. This method uses multiple independent agents to share information through an undirected graph.

[0063] In a specific embodiment of this specification, the system queries historical user feedback information related to the current logistics information processing request. It is assumed that historical data shows that in this "problem reporting" scenario, responses containing empathetic tone and explicit reassuring promises (such as "Please rest assured") have a significantly higher probability of receiving a "satisfactory" rating than other styles. A reward function scores three candidates: Candidate 2 receives the highest score because it includes elements such as "I'm very sorry" and "Please rest assured." The system uses the target feedback information (Candidate 2) as a successful example and initiates a reinforcement learning algorithm such as the GRPO algorithm. The goal of the GRPO algorithm is to adjust the parameters of the interaction model so that the probability of the model generating a high-scoring response like Candidate 2 when encountering similar situations in the future is greatly increased. Through policy gradient updates, the model's internal parameters are fine-tuned. After optimization, the original interaction model evolves into a more powerful target interaction model.

[0064] Based on this, by learning directly from real user feedback, user preferences are directly transformed into model parameters, enabling data-driven autonomous performance evolution of the model. The reward function directly quantifies the business objective of "user satisfaction" as a metric for technical optimization, ensuring that the model's optimization direction is highly consistent with the goal of improving user experience.

[0065] The information processing methods provided in this specification will be further described in detail below through various embodiments.

[0066] Furthermore, determining the target feedback information from the plurality of candidate feedback information based on the user feedback information associated with the logistics information processing request includes: determining a sample data sequence containing the logistics information processing request, and reading the user feedback information from the sample data sequence; evaluating the plurality of candidate feedback information based on the user feedback information, and determining the target feedback information from the plurality of candidate feedback information based on the evaluation result.

[0067] The sample data sequence can be understood as a structured unit of historical interaction records. It can be a complete entry in a session log or a context-dependent sequence of multiple log entries. The sample data sequence may include user queries (logistics information processing requests), historical responses (feedback from the interaction model), user feedback (user evaluations of the historical response such as "satisfied" or "dissatisfied"), or implicit feedback such as subsequent expressions of gratitude.

[0068] In practical applications, user feedback information can be retrieved from a sample database. When the system receives a new logistics information processing request, it searches the historical database for "sample data sequences" that are highly similar to the current request in semantics, scenario, or processing type, and extracts key "user feedback information" from these sequences. This yields a quantifiable "empirical value" based on real-world history, used to predict the user satisfaction that each candidate response might elicit. Subsequently, multiple candidate responses can be evaluated based on the user feedback information; this is the core calculation logic of the reward function. Evaluation is not performed directly but through a learned or pre-defined mapping relationship. If historical sample data sequences show that responses containing specific elements (such as reassuring language or clear follow-up steps) received positive feedback, then candidates containing similar elements in the current response will receive high scores. Conversely, if historical data shows that a certain type of response (such as vague replies or no solution provided) is always accompanied by negative feedback, then similar candidates in the current response will be penalized. Finally, based on the evaluation results, the target feedback information can be determined from multiple candidate responses to select the most suitable response to the logistics information processing request.

[0069] In a specific embodiment of this specification, the system searches for samples related to "delayed delivery" in historical training data. It finds a historical sequence with high similarity. Key information is extracted from this sample sequence: users provided positive feedback to responses containing "empathy" and "providing specific reasons and expectations." For the current query, the model generates candidate A and candidate B. Based on the retrieved historical feedback information and an evaluation using a reward function, candidate A is determined to have a low score, and candidate B to have a high score. Therefore, based on the evaluation results, candidate B is selected as the target feedback information.

[0070] Based on this, by reading user feedback information from the sample data sequence containing physical information processing requests, historical information is introduced to select the target feedback information that should be hit at the moment. Subsequently, the target feedback information is used to optimize the model, thereby improving the training efficiency and capability of the model.

[0071] Furthermore, optimizing the interaction model into a target interaction model based on the target feedback information includes: optimizing the interaction model based on the target feedback information until an intermediate interaction model that meets the training stopping condition is obtained; evaluating the intermediate interaction model and the interaction model using an evaluation set, and determining the target interaction model based on the evaluation results.

[0072] After optimizing the interaction model using the target feedback information, multiple rounds of optimization can be repeated until an intermediate interaction model that meets the training stopping condition is obtained. The training stopping condition can be reaching a preset number of rounds or the model parameters reaching preset values. At this point, the model in the current optimization round can be considered the intermediate interaction model, and the intermediate interaction model is evaluated. The intermediate interaction model is the model version obtained during the optimization process when the training stopping condition is met, which has not yet undergone final validation. The intermediate interaction model may be better than the original model, or it may deteriorate due to improper training. Therefore, it needs to be compared and evaluated with the original interaction model. Based on the evaluation results, it is determined whether the original interaction model needs to be replaced, thus ultimately determining the target interaction model. During evaluation, an evaluation set is used to evaluate both the intermediate interaction model and the original interaction model. The evaluation set is a representative, high-quality benchmark dataset containing various typical user queries and their expected ideal responses, serving as a "standard test" for evaluating model performance. Evaluation results can include quantitative metrics for assessing the model's performance on the evaluation set, such as accuracy and recall, which are core indicators for measuring the performance of information retrieval and classification systems. Specifically, evaluation results can also include metrics such as model response relevance, fluency, and user satisfaction prediction values, thus achieving a multi-faceted and multi-dimensional evaluation.

[0073] In practical applications, after training meets certain conditions, such as simultaneously satisfying 1000 training samples and two iterations of training, the latest model training file is saved. An intermediate interactive model is obtained based on the saved training file, and using prepared benchmark samples, the performance of the new model and the online service model is automatically scheduled and evaluated. If the new model achieves better precision and recall performance on the benchmark compared to the online service model, it is automatically deployed and switched to become the target interactive model for serving users.

[0074] In a specific embodiment of this specification, the system uses new Training Data (e.g., 1000 accumulated interaction logs with feedback) to train the v1.0 model using reinforcement learning. Model parameters are continuously updated. Training pauses when the system detects that the training samples have been iterated twice and the total data volume has reached the target. The resulting model is saved as an intermediate interaction model. The system automatically starts an evaluation task. The intermediate interaction model and the current interaction model are run on the same fixed evaluation set (e.g., containing 500 test questions covering various logistics scenarios). The evaluation results of the two models on various metrics are automatically calculated, for example: the interaction model has an accuracy of 85% and a recall of 80%; the intermediate interaction model has an accuracy of 88% and a recall of 82%. Based on the evaluation results, the intermediate interaction model is selected as the final target interaction model for online deployment and use.

[0075] Based on this, by evaluating and comparing the intermediate interaction model with the original interaction model after training, the risk of model performance degradation due to unstable reinforcement learning training or data noise is eliminated, ensuring that each online model update is a definite performance improvement, thereby guaranteeing the stability and reliability of service quality.

[0076] Furthermore, determining the sample data sequence includes: acquiring at least one interactive structured data recorded by the interaction model during the deployment phase, wherein the interactive structured data includes candidate processing requests, candidate user feedback information, and candidate model feedback information; verifying the at least one interactive structured data according to a preset data verification strategy, and selecting the interactive structured data that passes the verification as the sample data sequence.

[0077] The structured interaction data recorded during the deployment phase can be understood as the raw interaction records generated in real time in the online production environment, with a uniform format. Unlike temporary logs, this data is specifically structured for subsequent learning. The structured interaction data can include candidate processing requests, candidate model feedback information, and candidate user feedback information. Candidate processing requests can be understood as the user's original query, such as a logistics processing request. Candidate model feedback information can be understood as the response generated by the model for that request and ultimately returned to the user. Candidate user feedback information can be understood as the user's evaluation of the response (e.g., satisfaction / dissatisfaction, or whether the problem was solved). The pre-set data validation strategy can be understood as a set of automated data cleaning and quality control rules. Its purpose is to filter out the beneficial and reliable parts for model training from massive amounts of raw interaction data that may contain noise.

[0078] In practical applications, the core criteria for data validation strategies include feedback clarity, data redundancy / novelty, and compliance with transaction rules. Feedback clarity means that user feedback must be explicit and interpretable (e.g., clearly labeled "satisfied" or "dissatisfied," or the user expressing gratitude / dissatisfaction). Data redundancy / novelty means that new data should not highly overlap with existing training data; it must be within a "controllable range" to ensure data diversity and information content, avoiding model overfitting. Transaction rule compliance means that response content must conform to format constraints, tool usage guidelines, etc. Validated interactive structured data can be understood as "high-quality data units" that have successfully passed all the above quality checks. They are deemed suitable for training models and are thus formally accepted as "sample data sequences," flowing into the Training Data pool.

[0079] In a specific embodiment of this specification, the system pulls a batch of new raw data from online logs as interactive structured data. For example, data A includes a candidate processing request: "Will my package arrive tomorrow?"; candidate model feedback: "Your package has been found to be at the delivery station and is expected to arrive tomorrow."; and candidate user feedback: Positive (the user clicked "Satisfied"). The system applies validation rules to each piece of data according to a preset data validation strategy and selects the validated data as a sample data sequence.

[0080] Based on this, through a rigorous data validation strategy, the system automatically filters out invalid, ambiguous, and redundant data. This ensures that the dataset used for reinforcement learning is "high-purity," enabling the model to learn from high-quality signals, greatly accelerating the convergence speed, and avoiding learning incorrect patterns from noisy data.

[0081] Furthermore, optimizing the interaction model into a target interaction model based on the target feedback information includes: determining a preset generalization regression strategy, trust region strategy, proximal strategy, depth determination strategy, or asynchronous training strategy; optimizing the interaction model based on the generalization regression strategy, the trust region strategy, the proximal strategy, the depth determination strategy, or the asynchronous training strategy and the target feedback information to obtain the target interaction model.

[0082] In the embodiments of this specification, different optimization strategies can be selected when optimizing the interaction model, such as generalized regression strategy, trust region strategy, proximal strategy, deep deterministic strategy, or asynchronous training strategy. The generalized regression strategy, also known as the GRPO algorithm, is a policy optimization method based on value function regression. It guides policy updates through generalized value estimation, aiming to improve model performance more efficiently and stably. The trust region strategy, also known as the TRPO algorithm, sets a trust region with each policy update to ensure that the new policy does not deviate too far from the old policy. This is achieved by constraining the KL divergence between the new and old policies, effectively preventing training collapse and making it very suitable for scenarios with extremely high stability requirements, such as online learning. The proximal strategy, also known as the PPO algorithm, directly limits the magnitude of policy updates by introducing a pruning function, also aiming to stabilize training. PPO has good convergence and robustness in practice and is one of the currently popular reinforcement learning algorithms. The deep deterministic strategy, also known as the DDPG algorithm, is an algorithm specifically designed for continuous action spaces. If the response generation process of intelligent customer service is viewed as selecting actions in a continuous semantic space (i.e., generating vector representations of each word or sentence), then DDPG is very suitable. It employs an Actor-Critic architecture, enabling it to learn deterministic policies. The asynchronous training strategy, known as the A3C algorithm, utilizes multiple threads to interact with the environment asynchronously and in parallel, updating a global model. This significantly improves data collection and training efficiency, accelerates convergence, and reduces variance during training through multiple parallel explorations.

[0083] In practical applications, in order to improve training efficiency or training quality, one or more algorithms can be used in combination. The specific number of algorithms to be selected can be determined according to the actual situation.

[0084] In a specific embodiment of this specification, the system architect or algorithm engineer will pre-define one or more core algorithms based on specific transaction requirements, the nature of the response action space (discrete selection vs. continuous generation), and the trade-off between stability and efficiency. Optimization is performed based on the strategy and target feedback information. After iterative optimization using one or more of the above strategies and meeting the training stopping condition, a target interaction model with improved performance is born. The target interaction model internalizes the successful dialogue strategies learned from the target feedback information.

[0085] Based on this, by pre-setting a variety of different cutting-edge algorithms, the training needs of different training scenarios are met, ensuring that the training process of the model tends more stably toward the first optimal solution, and avoiding the drastic performance fluctuations that are prone to occur in traditional policy gradient methods.

[0086] Furthermore, the step of inputting the logistics information and the logistics information processing request into the interaction model for processing to obtain multiple candidate feedback information includes: determining a prompt word template that matches the target interface, adding the logistics information and the logistics information processing request to the prompt word template to obtain input context information; and inputting the input context information into the interaction model for processing to obtain multiple candidate feedback information.

[0087] The prompt word template matching the target interface can be understood as a predefined, structured text framework—a complete context template containing role settings, task descriptions, output format requirements, and data placeholders. The purpose of the prompt word template is to organize the raw, unstructured user queries and API-returned data into a format that the model can easily understand and execute accurately. By adding logistics information and logistics information processing requests to the prompt word template, input context information that can be directly input into the interaction model is obtained. This input context information is the final, complete input text formed by filling the placeholders in the prompt word template with the specific logistics information processing requests and logistics information (API return results). It is the sole basis for the interaction model to execute its tasks.

[0088] In a specific embodiment of this specification, the system selects a preset template that matches the processing type "Logistics Status Inquiry and Problem Reporting". The template content may be as follows: [Role] You are a professional logistics customer service assistant.

[0089] [Task] Based on the user's query and system information, generate a reply that reassures the user and informs them of the processing result.

[0090] [System Information] Logistics Status: {logistics_status} Order Number: {ticket_id} [User query] {user_query}

Requirements

[0091] The system fills in the specific data into the template, generating a completed input context ready to be sent to the model. This context is then input into the interactive model, which, based on this highly structured and comprehensive prompt, generates multiple candidate feedback messages that meet the requirements.

[0092] Based on this, by forcing the model to respond based on the provided logistics information (system information) through templates, the possibility of the model "fabricating" information is fundamentally eliminated, ensuring the accuracy and consistency of information, achieving standardized responses, and avoiding the generation of random and unprofessional content.

[0093] This specification provides an information processing method, comprising: acquiring a logistics information processing request and determining the processing type corresponding to the logistics information processing request; calling a target interface corresponding to the processing type to acquire logistics information associated with the logistics information processing request; inputting the logistics information and the logistics information processing request into an interaction model for processing to obtain multiple candidate feedback information; determining a target feedback information from the multiple candidate feedback information based on the user feedback information associated with the logistics information processing request; and optimizing the interaction model into a target interaction model based on the target feedback information. This method achieves the goal of acquiring logistics information associated with the logistics information processing request by calling the target interface corresponding to the processing type. Inputting the logistics information and the logistics information processing request into the interaction model for processing allows external information and user queries to be input into the model simultaneously, enabling the model to combine external information for prediction and improving the model's prediction accuracy. Based on user feedback information associated with logistics information processing requests, the target feedback information is determined from multiple candidate feedback information output by the model. The interaction model is then optimized into the target interaction model using the target feedback information. This achieves the function of dynamically optimizing the model based on user feedback information, enabling the model to have self-updating capabilities, improving the model's adaptability, and further enhancing the user experience of the request feedback function.

[0094] See Figure 2 , Figure 2 A flowchart of another information processing method provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0095] Step S202: Receive the target logistics information processing request submitted by the target user, and determine the target processing type corresponding to the target logistics information processing request.

[0096] Step S204: Call the target interface corresponding to the target processing type to obtain the target logistics information associated with the target logistics information processing request.

[0097] Step S206: Input the target logistics information and the target logistics information processing request into the target interaction model for processing, obtain target feedback information, and feed back the target feedback information to the target user, wherein the target interaction model is obtained by the above method.

[0098] In this context, the target user can be understood as a real user requesting customer service, and the target logistics information processing request is the real-time request submitted by that user. After determining the target processing type of the target logistics information processing request, the corresponding interface can be called to retrieve the target logistics information associated with the request. Inputting the target logistics information and the target logistics information processing request into a trained target interaction model yields the model's output target feedback information, which is then sent back to the target user, thereby providing intelligent customer service.

[0099] This specification provides an information processing method that, based on a trained target interaction model, predicts target logistics information and the target logistics information processing request in relation to the target logistics information processing request, thereby obtaining target feedback information. This allows the model to combine user queries and external knowledge to accurately generate the user's expected response, improving the accuracy of the model's output and further increasing the resolution rate of user-submitted questions. Simultaneously, it effectively reduces the need for repeated follow-up inquiries and transfers to human customer service. This directly reduces the workload of human customer service representatives and the company's operating costs, achieving cost reduction and efficiency improvement.

[0100] The following is in conjunction with the appendix Figure 3 Taking the application of the information processing method provided in this specification in a courier information query scenario as an example, the information processing method will be further explained. Among other things, Figure 3 A flowchart illustrating the processing steps of an information processing method according to an embodiment of this specification is shown, specifically including the following steps.

[0101] Step S302: Obtain the logistics information processing request and determine the processing type corresponding to the logistics information processing request.

[0102] Step S304: Call the target interface corresponding to the processing type to obtain the logistics information of the associated logistics information processing request.

[0103] Step S306: Determine the prompt word template that matches the target interface, add the logistics information and logistics information processing request to the prompt word template, and obtain the input context information.

[0104] Step S308: Input the input context information into the interaction model for processing to obtain multiple candidate feedback information.

[0105] Step S310: Determine the sample data sequence containing the logistics information processing request, and read user feedback information from the sample data sequence.

[0106] Step S312: Evaluate multiple candidate feedback information based on user feedback information, and determine the target feedback information from the multiple candidate feedback information based on the evaluation results.

[0107] Step S314: Optimize the interaction model based on the target feedback information until an intermediate interaction model that meets the training stopping condition is obtained.

[0108] Step S316: Use the evaluation set to evaluate the intermediate interaction model and the interaction model respectively, and determine the target interaction model based on the evaluation results.

[0109] Step S318: Upon receiving a target logistics information processing request submitted by a target user, determine the target processing type corresponding to the target logistics information processing request.

[0110] Step S320: Call the target interface corresponding to the target processing type to obtain the target logistics information associated with the target logistics information processing request.

[0111] Step S322: Input the target logistics information and the target logistics information processing request into the target interaction model for processing, obtain the target feedback information, and then send the target feedback information back to the target user.

[0112] This specification provides an information processing method that, based on the processing type of a logistics information processing request, calls the target interface corresponding to the processing type to obtain the associated logistics information for the request. The logistics information and the logistics information processing request are input into an interactive model for processing, enabling the model to combine external information with user queries for prediction, thus improving prediction accuracy. Based on user feedback associated with the logistics information processing request, the target feedback information is determined from multiple candidate feedback information output by the model. This target feedback information is then used to optimize the interactive model into the target interactive model, achieving dynamic model optimization based on user feedback information. This gives the model self-updating capabilities, improves its adaptability, and further enhances the user experience of the request feedback function.

[0113] Corresponding to the above method embodiments, this specification also provides embodiments of an information processing apparatus. Figure 4 A schematic diagram of the structure of an information processing apparatus according to one embodiment of this specification is shown. Figure 4 As shown, the device includes: The acquisition module 402 is configured to acquire a logistics information processing request and determine the processing type corresponding to the logistics information processing request; Module 404 is configured to invoke the target interface corresponding to the processing type to obtain the logistics information associated with the logistics information processing request. Input module 406 is configured to input the logistics information and the logistics information processing request into the interaction model for processing to obtain multiple candidate feedback information; The determination module 408 is configured to determine the target feedback information from the plurality of candidate feedback information based on the user feedback information associated with the logistics information processing request, and optimize the interaction model into the target interaction model based on the target feedback information.

[0114] In an optional embodiment, determining the target feedback information from the plurality of candidate feedback information based on the user feedback information associated with the logistics information processing request includes: A sample data sequence containing the logistics information processing request is determined, and the user feedback information is read from the sample data sequence; the multiple candidate feedback information is evaluated based on the user feedback information, and the target feedback information is determined from the multiple candidate feedback information based on the evaluation results.

[0115] In an optional embodiment, optimizing the interaction model into a target interaction model based on the target feedback information includes: The interaction model is optimized based on the target feedback information until an intermediate interaction model that meets the training stopping condition is obtained; the intermediate interaction model and the interaction model are evaluated using an evaluation set, and the target interaction model is determined based on the evaluation results.

[0116] In an optional embodiment, determining the sample data sequence includes: At least one interactive structured data recorded by the interaction model during the deployment phase is obtained, wherein the interactive structured data includes candidate processing requests, candidate user feedback information and candidate model feedback information; the at least one interactive structured data is verified according to a preset data verification strategy, and the interactive structured data that passes the verification is selected as the sample data sequence.

[0117] In an optional embodiment, optimizing the interaction model into a target interaction model based on the target feedback information includes: Determine a preset generalization regression strategy, trust region strategy, proximal strategy, depth determination strategy, or asynchronous training strategy; optimize the interaction model based on the generalization regression strategy, trust region strategy, proximal strategy, depth determination strategy, or asynchronous training strategy, and the target feedback information to obtain the target interaction model.

[0118] In one optional embodiment, the step of inputting the logistics information and the logistics information processing request into the interaction model for processing to obtain multiple candidate feedback information includes: A prompt word template matching the target interface is determined, and the logistics information and the logistics information processing request are added to the prompt word template to obtain input context information; the input context information is input into the interaction model for processing to obtain multiple candidate feedback information.

[0119] The above is an illustrative scheme of an information processing device according to this embodiment. It should be noted that the technical solution of this information processing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the information processing device, please refer to the description of the technical solution of the information processing method described above.

[0120] Corresponding to the above method embodiments, this specification also provides another embodiment of an information processing apparatus. Figure 5 A schematic diagram of another information processing apparatus provided in one embodiment of this specification is shown. Figure 5 As shown, the device includes: The receiving module 502 is configured to receive a target logistics information processing request submitted by a target user and determine the target processing type corresponding to the target logistics information processing request. The acquisition module 504 is configured to call the target interface corresponding to the target processing type to acquire the target logistics information associated with the target logistics information processing request; The processing module 508 is configured to input the target logistics information and the target logistics information processing request into the target interaction model for processing, obtain target feedback information, and provide the target feedback information to the target user, wherein the target interaction model is obtained by the above method.

[0121] The above is an illustrative scheme of another information processing device according to this embodiment. It should be noted that the technical solution of this information processing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the information processing device, please refer to the description of the technical solution of the information processing method described above.

[0122] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0123] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0124] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0125] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.

[0126] The processor 620 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described information processing method.

[0127] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the information processing method described above.

[0128] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described information processing method.

[0129] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the information processing method described above.

[0130] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described information processing method.

[0131] The above is an illustrative example of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the information processing method described above belong to the same concept. Details not described in detail in the technical solution of the computer program can be found in the description of the technical solution of the information processing method described above.

[0132] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described information processing method.

[0133] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the information processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the information processing method described above.

[0134] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0135] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0136] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0137] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0138] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An information processing method, comprising: Obtain a logistics information processing request and determine the processing type corresponding to the logistics information processing request; Call the target interface corresponding to the processing type to obtain the logistics information associated with the logistics information processing request; The logistics information and the logistics information processing request are input into the interaction model for processing to obtain multiple candidate feedback information; Based on the user feedback information associated with the logistics information processing request, the target feedback information is determined from the plurality of candidate feedback information, and the interaction model is optimized into the target interaction model based on the target feedback information.

2. The information processing method according to claim 1, wherein determining the target feedback information from the plurality of candidate feedback information based on the user feedback information associated with the logistics information processing request includes: Determine a sample data sequence containing the logistics information processing request, and read the user feedback information from the sample data sequence; The multiple candidate feedback messages are evaluated based on the user feedback information, and the target feedback message is determined from the multiple candidate feedback messages based on the evaluation results.

3. The information processing method according to claim 1, wherein optimizing the interaction model into a target interaction model based on the target feedback information comprises: The interaction model is optimized based on the target feedback information until an intermediate interaction model that meets the training stopping condition is obtained. The intermediate interaction model and the interaction model are evaluated using an evaluation set, and the target interaction model is determined based on the evaluation results.

4. The information processing method according to claim 2, wherein determining the sample data sequence includes: Obtain at least one structured interaction data recorded by the interaction model during the deployment phase, wherein the structured interaction data includes candidate processing requests, candidate user feedback information, and candidate model feedback information; The at least one interactive structured data is verified according to a preset data verification strategy, and the interactive structured data that passes the verification is selected as the sample data sequence.

5. The information processing method according to any one of claims 1 to 4, wherein optimizing the interaction model into a target interaction model based on the target feedback information comprises: Determine the preset generalization regression strategy, trust region strategy, proximal strategy, depth determination strategy, or asynchronous training strategy; The interaction model is optimized based on the generalized regression strategy, the trust region strategy, the proximal strategy, the depth determination strategy, or the asynchronous training strategy, as well as the target feedback information, to obtain the target interaction model.

6. The information processing method according to claim 1, wherein inputting the logistics information and the logistics information processing request into an interaction model for processing to obtain multiple candidate feedback information includes: Determine a prompt word template that matches the target interface, add the logistics information and the logistics information processing request to the prompt word template, and obtain input context information; The input context information is input into the interaction model for processing to obtain multiple candidate feedback information.

7. An information processing method, comprising: Receive a target logistics information processing request submitted by a target user, and determine the target processing type corresponding to the target logistics information processing request; Call the target interface corresponding to the target processing type to obtain the target logistics information associated with the target logistics information processing request; The target logistics information and the target logistics information processing request are input into the target interaction model for processing to obtain target feedback information, and the target feedback information is fed back to the target user, wherein the target interaction model is obtained by the method of any one of claims 1 to 6.

8. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program or instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.