Multi-agent recommendation method and device combined with preference learning, electronic equipment, storage medium and program product

By combining the multi-agent recommendation method of preference learning, using the user's fine preferences and context information, the problem of insufficient recommendation relevance and accuracy in the existing recommendation system is solved, and more efficient recommendation results are achieved.

CN120045697APending Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202411904063.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing multi-agent recommendation system fails to fully utilize the user's fine preferences and context information, resulting in insufficient relevance and accuracy of recommendations.

Method used

A multi-agent recommendation method combining preference learning is proposed. By determining the user's task query information, translating it into task demand information, inference based on user information, and evaluating and optimizing using pre-trained preference learning alignment model, we finally generate optimization recommendation answers.

Benefits of technology

By accurately analyzing user queries and converting them into specific task requirements, using personalized reasoning to generate initial answers, and then evaluating and optimizing through preference learning alignment models, we will ultimately provide optimized recommendation answers that are more in line with user's fine preferences and context information, thereby improving the performance of the recommendation system and improving the relevance and accuracy of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045697A_ABST
    Figure CN120045697A_ABST
Patent Text Reader

Abstract

The invention provides a preference learning combined multi-agent recommendation method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: determining task query information of a query user, translating the task query information, and obtaining task demand information; reasoning based on the query user, the task query information and the task demand information to obtain an initial answer; evaluating the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback; and adjusting the initial answer based on the optimization feedback to obtain an optimization recommendation answer. According to the method, the fine preference and the context information of the user can be fully utilized, so that the relevance and the accuracy of recommendation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of natural language processing, and in particular, to a multi-agent recommendation method, apparatus, electronic device, storage medium, and program product that combines preference learning. Background Art

[0002] This section aims to provide background or context for the embodiments of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely by virtue of being included in this section.

[0003] A multi-agent recommendation system is a technology that uses multiple collaborative agents to provide personalized recommendations. These agents can simulate different roles or strategies and analyze users' behaviors, preferences, and context information through interaction and collaborative work to generate recommendation results.

[0004] However, in the related art, the fine preferences and context information of users are not fully utilized, resulting in insufficient relevance and accuracy of recommendations. Summary of the Invention

[0005] In view of this, an object of the present disclosure is to provide a multi-agent recommendation method, apparatus, electronic device, storage medium, and program product that combines preference learning, which can at least solve one of the technical problems in the related art to a certain extent.

[0006] Based on the above object, in the first aspect of the exemplary embodiments of the present disclosure, a multi-agent recommendation method that combines preference learning is provided, which is applied to a server. The method includes:

[0007] Determine the task query information of the query user, and translate the task query information to obtain task requirement information;

[0008] Infer based on the query user, the task query information, and the task requirement information to obtain an initial answer;

[0009] Evaluate the initial answer through a pre-trained preference learning alignment model to obtain an optimization feedback;

[0010] Adjust the initial answer based on the optimization feedback to obtain an optimized recommendation answer.

[0011] Based on the same inventive concept, in the second aspect of the exemplary embodiments of the present disclosure, a multi-agent recommendation apparatus that combines preference learning is provided, including:

[0012] A requirement information determination module, configured to determine the task query information of the query user, and translate the task query information to obtain task requirement information;

[0013] An initial answer determination module, configured to perform reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer;

[0014] An optimization feedback determination module, configured to evaluate the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback;

[0015] A recommended answer determination module, configured to adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer.

[0016] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0017] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect.

[0018] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product including computer program instructions that, when run on a computer, cause the computer to execute the method described in the first aspect.

[0019] As can be seen from the above, the multi-agent recommendation method, apparatus, electronic device, storage medium, and program product provided by the embodiments of the present disclosure in combination with preference learning include:

[0020] Determine the task query information of the query user, translate the task query information to obtain task requirement information; perform reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer; evaluate the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback; adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer. The present disclosure can make full use of the user's fine-grained preferences and context information, thereby improving the relevance and accuracy of recommendations. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 Schematic diagram of the application scenario of the multi-agent recommendation method combined with preference learning provided by an exemplary embodiment of the present disclosure;

[0023] Figure 2 Flow chart of a multi-agent recommendation method combined with preference learning provided by an exemplary embodiment of the present disclosure;

[0024] Figure 3 Schematic diagram of the structure of a multi-agent system of the multi-agent recommendation method combined with preference learning provided by an exemplary embodiment of the present disclosure;

[0025] Figure 4 Schematic diagram of the structure of a multi-agent recommendation device combined with preference learning provided by an exemplary embodiment of the present disclosure;

[0026] Figure 5 Schematic diagram of the hardware structure of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0027] It can be understood that, before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0028] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the technical solutions of the present application according to the prompt message.

[0029] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0030] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present application. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present application.

[0031] It can be understood that the data involved in the technical solution of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related regulations.

[0032] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to convey the scope of the present disclosure fully to those skilled in the art.

[0033] In this document, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0034] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by those of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms "comprising" or "including" and similar terms mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. The article "a" or "an" before an element does not exclude the existence of multiple such elements.

[0035] The principles and spirit of the present disclosure will be explained in detail below with reference to several representative embodiments of the present disclosure.

[0036] As described in the background art, in the related art, the fine-grained preferences and context information of users have not been fully utilized, resulting in insufficient relevance and accuracy of recommendations. Specifically, in the related art, the SciAgents framework is an advanced technical system that combines knowledge graph construction and reasoning, as well as the application of large language models and multi-agent systems. In terms of knowledge graph construction and reasoning, this framework constructs an ontology knowledge graph through the GraphReasoning library, extracts key knowledge points from scientific literature, and determines the relationships between these concepts through random path or shortest path algorithms. At the same time, the framework also utilizes large language models, such as GPT-4, to generate, expand, and optimize research hypotheses. In addition, the system is divided into multiple role-based agents, including "planner", "ontologist", "scientist No. 1", "scientist No. 2", and "critic", etc. Each agent executes tasks according to its specific instruction set, jointly promoting the progress of research.

[0037] In terms of the construction and reasoning of the knowledge graph in the SciAgents framework, although it can extract information from scientific literature and construct an ontology knowledge graph through the GraphReasoning library, this process relies on a large amount of domain knowledge and expert input, so the cost is relatively high. At the same time, problems such as data noise and inconsistency may be encountered during the extraction process, and using random path or shortest path algorithms may sometimes not ensure the generation of the most accurate relationships between concepts. In the application of large language models and multi-agent systems, although it can generate, expand, and optimize research hypotheses, the operation of the system requires a complex coordination mechanism to manage numerous agents. Communication between agents may cause delays and efficiency problems, and when dealing with unstructured tasks, agents may have difficulty proposing precise hypotheses.

[0038] In the related art, the content-based Agent+ recommendation scheme uses large language models (LLMs) to improve the performance of the recommendation system. This scheme can generate personalized recommendation descriptions by analyzing the historical preferences and interaction data of users, making the recommended content more in line with the unique personality of each user. In addition, LLMs are also used to understand users' search queries or recommendation requests, including those complex expressions and nuances contained in natural language.

[0039] However, when generating personalized descriptions, the content-based Agent recommendation scheme relies on a large amount of user data to ensure the accuracy of the descriptions, which may raise concerns about privacy. At the same time, since users' preferences change over time, the system must be continuously updated to maintain the relevance of its recommendations. In terms of query understanding, in the face of complex or ambiguous queries, even advanced LLMs may have difficulty accurately grasping users' intentions.

[0040] In the related art, the AGENTiGraph platform is a comprehensive technical framework that optimizes the interaction between users and knowledge graphs and reduces its complexity by converting natural language queries into structured graph operations through semantic parsing capabilities. In addition, the platform's adaptive multi-agent system can integrate multimodal inputs from LLM agents to generate coherent action plans consistent with the user's intent. Figure 1 At the same time, AGENTiGraph also has the ability of dynamic knowledge integration, supports continuous knowledge extraction and update, ensures that the information in the knowledge graph remains up-to-date, and provides dynamic visualization functions.

[0041] However, in terms of semantic parsing of the AGENTiGraph platform, although the platform can convert natural language queries into structured graph operations, this process may be ambiguous and inaccurate. Especially for non-standard queries, the system may have difficulty providing accurate parsing. In terms of the adaptive multi-agent system, processing multimodal inputs may require a large amount of computing resources, and the adaptability of the agents may sometimes lead to unpredictable behaviors, increasing the complexity of the system. As for dynamic knowledge integration, continuous knowledge extraction and integration may cause system overload, and dynamic updates may introduce new errors or inconsistencies, all of which may affect the performance and stability of the platform.

[0042] In the related art, the MetaGPT framework is a system specifically designed for multi-agent collaboration. It ensures a structured approach to problem-solving by encoding standardized operation (SOP) procedures as prompts. In this framework, each agent is required to participate in the collaboration in the form of an expert and must generate structured outputs according to the established requirements.

[0043] However, the dependence of the MetaGPT framework on SOPs may limit the flexibility and creativity of the agents.

[0044] In the related art, the Agent4Rec system is an advanced movie recommendation system simulator composed of 1,000 intelligent agents initialized based on real user data. Powered by ChatGPT-3.5, the system can make personalized responses to different recommendation algorithms and the movies they recommend according to each user's unique preferences and characteristics.

[0045] As a movie recommendation system simulator, the Agent4Rec system can simulate the behavior of 1,000 intelligent agents, which are initialized by real users and driven by ChatGPT-3.5 to achieve personalized movie recommendations. However, it may not be able to fully capture all the complexities of real user behavior. In addition, the performance of the system may depend on the initial user data, which may lead to biased recommendation results. For new or unpopular movies, the recommendations provided by the Agent4Rec system may not be accurate enough because it may not have enough data to understand and predict user preferences for these contents.

[0046] In the study of alignment methods, the task is usually regarded as a binary classification problem, and the BT model is applied to calculate the negative log-likelihood loss. To simplify the analysis, two functions are defined: r φ (x,y w ) is expressed as logx 1 ,r φ (x,y l ) is expressed as logx 2 . Based on this, the objective function of the BT model is expressed as:

[0047]

[0048] In the formula, represents the objective function of the BT model, that is, the loss function; r φ represents the parameters of the model; represents the data set; x represents the input of the model; y w represents the correct answer or positive sample; y l represents the wrong answer or negative sample; x 1 represents the score or probability of the model for the correct answer; x 2 represents the score or probability of the model for the wrong answer; β represents a hyperparameter used to adjust the sensitivity of the score; σ represents the sigmoid function used to convert the input into a probability between 0 and 1.

[0049] Furthermore, this objective function is simplified to By studying the gradient vector field of the BT model, two main limitations are identified:

[0050] The influence on x 2 is greater than that on x 1 , which indicates that the gradient changes more significantly on x 2 . Therefore, the model is more inclined to reduce the probability of generating non-preferred data rather than increasing the probability of generating preferred responses.

[0051] The initial state in the gradient vector field has a significant impact on the final optimization result. Specifically, the initial position of the large language model after SFT may be at a position where both x 1 and x 2 are relatively small, indicating a lower probability of generating human preferences and non-preferred responses, and the gradient direction is not completely prioritized for enhancing the preferred response. If the initial position is at a position where both x 1 and x 2 are relatively large, the very small gradient present will lead to slow convergence and difficulty in escaping local minima, which may result in insufficient learning from human preference data.

[0052] To solve the above problems, the present disclosure provides a multi-agent recommendation method, device, electronic device, storage medium, and program product solution that combines preference learning, specifically including:

[0053] Determine the task query information of the query user, translate the task query information to obtain task requirement information; perform reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer; evaluate the initial answer through a pre-trained preference learning alignment model to obtain an optimization feedback; adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer. The present disclosure precisely analyzes the user query and converts it into specific task requirements, generates an initial answer using personalized reasoning, and then evaluates and optimizes it through a preference learning alignment model, ultimately providing an optimized recommended answer that better conforms to the user's fine preferences and context information, thereby improving the performance of the recommendation system and enhancing the relevance and accuracy of the recommendation.

[0054] After introducing the basic principle of the present disclosure, the various non-limiting embodiments of the present disclosure will be specifically introduced below.

[0055] Refer to Figure 1 which is a schematic diagram of an application scenario of the multi-agent recommendation method that combines preference learning provided by an exemplary embodiment of the present disclosure.

[0056] In this application scenario, it includes a terminal device 101 and a server 102. Among them, the terminal device 101 and the server 102 can be connected through a wired or wireless communication network to achieve data interaction.

[0057] The terminal device 101 can be an electronic device near the user side with data transmission, multimedia input / output functions, including but not limited to desktop computers, mobile phones, mobile computers, tablets, media players, smart wearable devices, personal digital assistants (PDAs), or other electronic devices capable of implementing the above functions. The electronic device may include a processor and a display screen with touch input function, the display screen is used to present a graphical user interface, the graphical user interface can display an application interface, and the processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.

[0058] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0059] In some exemplary embodiments, the multi-agent recommendation method combined with preference learning can run on the terminal device 101 or the server 102.

[0060] When the multi-agent recommendation method combined with preference learning runs on the server 102, the server 102 is used to provide the multi-agent recommendation service combined with preference learning to the user of the terminal device 101.

[0061] The server 102 receives the task query information of the query user from the terminal device 101, and the server 102 translates the task query information to obtain task requirement information.

[0062] The server 102 performs reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer.

[0063] The server 102 evaluates the initial answer through a pre-trained preference learning alignment model to obtain an optimization feedback.

[0064] After the server 102 adjusts the initial answer based on the optimization feedback to obtain an optimized recommended answer; the server 102 transmits the optimized recommended answer to the terminal device 101.

[0065] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0066] Reference Figure 2 , a multi-agent recommendation method combined with preference learning, the method comprising the following steps:

[0067] Step S210, determine the task query information of the query user, translate the task query information to obtain task requirement information.

[0068] When specifically implemented, the method for determining the task query information of the query user:

[0069] Reference Figure 3 , in this exemplary embodiment, the multi-agent system includes five agents, including: a task interpreter, a manager, a searcher, a user / item analyzer, and a reflector, where the multi-agent system can receive the task query information input by the query user.

[0070] In the above exemplary embodiment, the method for determining the task query information of the query user is introduced. Next, the method for obtaining the task requirement information is introduced:

[0071] When specifically implemented, the method for translating the task query information to obtain task requirement information:

[0072] Reference Figure 3 , after the multi-agent system receives the task query information input by the query user, the multi-agent system translates the task query information into a specific recommendation task through the task interpreter and generates a task requirement description. Specifically, the task interpreter translates the conversation about the task query information into an executable recommendation task. The task interpreter will obtain the session history when starting to run. Since the history of the session may be very long, the task interpreter can only obtain the last part of the history. The task interpreter can also call a text summarization tool to obtain a more concise overview of the history. Finally, the task interpreter will give a specific description of the task requirement to guide the subsequent operation of the manager.

[0073] Step S220, perform reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer.

[0074] In this exemplary embodiment, performing reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer includes:

[0075] Performing keyword search based on the task requirement information to obtain relevant information of the task requirement information; analyzing the query user and the task query information to obtain the preference information of the query user; performing information aggregation based on the relevant information and the preference information to obtain the initial answer.

[0076] During specific implementation, keyword search is performed based on the task requirement information to obtain relevant information about the task requirement information; the query user and the task query information are analyzed to obtain the preference information of the query user; information aggregation is performed based on the relevant information and the preference information to obtain the initial answer in the following manner:

[0077] Reference Figure 3 , after the manager gets the task interpreter to translate the task query information into a specific recommended task, the manager starts to call other agents to obtain a detailed analysis of the user and the project. These agents, including the searcher and the user / project analyst, support the invocation of some tools. For example, the searcher can access a search engine, and the user / project analyst can access detailed information about the user and the project. After receiving the responses from the searcher and the user / project analyst, the manager will try to provide an answer, that is, give a ranking of the candidate set. Among them, the specific role of the manager is: for a given task, the manager assigns subtasks to other agents to complete the main execution process. It supervises the collaboration among all other agents. The manager always alternately executes three steps: "think", "act", and "observe". In the thinking stage, the manager analyzes the reasons for the current situation of the task (such as whether the analysis is sufficient, whether additional information is needed, etc.). In the action stage, the manager can choose to give an answer to end the task or seek help from other agents (in a specific interface format). The responses of other agents will be given in the observation stage of the manager.

[0078] In the above exemplary embodiment, the manner of obtaining the initial answer is introduced. Next, the manner of obtaining relevant information about the task requirement information will be specifically introduced:

[0079] In this exemplary embodiment, keyword search is performed based on the task requirement information to obtain relevant information about the task requirement information, including:

[0080] Keyword extraction is performed on the task requirement information to obtain requirement information keywords; based on a search engine, the requirement information keywords are searched to obtain requirement information reference data; the requirement information reference data is screened and summarized to obtain the relevant information.

[0081] During specific implementation, the manner of performing keyword extraction on the task requirement information to obtain requirement information keywords, searching the requirement information keywords based on a search engine to obtain requirement information reference data, and screening and summarizing the requirement information reference data to obtain the relevant information is as follows:

[0082] Reference Figure 3, the searcher first extracts keywords from the task requirement information, uses a search engine to search for the task requirement information, and further retrieves paragraphs containing the keywords in specific entries based on the keywords. Then, the searcher will screen and summarize the relevant data retrieved to respond to the manager's query. Specifically, the role of the searcher is to be responsible for searching using a search tool according to the requirements given by the manager, and finally summarizing the text and replying to the manager. Taking Wikipedia as an example of a search tool. The searcher can provide a search query to obtain the most relevant entries in Wikipedia. The searcher can further retrieve paragraphs containing the given keywords in specific entries. Finally, the searcher is required to summarize the paragraphs to respond to the manager's query.

[0083] In the above exemplary embodiment, the method for specifically introducing the relevant information for obtaining the task requirement information is described below. Specifically, the method for analyzing the query user and the task query information to obtain the preference information of the query user is described:

[0084] In this exemplary embodiment, analyzing the query user and the task query information to obtain the preference information of the query user includes:

[0085] Constructing a user portrait for the query user based on the information database; performing project analysis on the task query information based on the information database to obtain project attributes; retrieving the interaction history between the query user and the task query information based on an interaction retriever; and analyzing to obtain the preference information based on the user portrait, the project attributes, and the interaction history.

[0086] Specifically in implementation, the user portrait refers to:

[0087] A virtual image constructed based on information such as the user's behavior, preferences, personal characteristics, and historical data; it is used to represent the typical characteristics of the user group.

[0088] Specifically in implementation, the project attributes refer to:

[0089] A series of characteristics and information related to a specific project or content, and these attributes help to comprehensively describe and understand the characteristics of the project.

[0090] Specifically in implementation, the method for constructing a user portrait for the query user based on the information database; performing project analysis on the task query information based on the information database to obtain project attributes; retrieving the interaction history between the query user and the task query information based on an interaction retriever; and analyzing to obtain the preference information based on the user portrait, the project attributes, and the interaction history:

[0091] Reference Figure 3 ,The user / project analyst focuses on examining and understanding the characteristics and preferences of the user, as well as the attributes of the project. The analyst will obtain two tools to assist in the analysis, including an information database and an interactive retriever. The analyst can obtain the user profile of each query user and the attributes of each project through the information database. Through the interactive retriever, the analyst can obtain the user / project interaction history before the current moment. By combining these two tools, the analyst can conduct in-depth analysis of the user or the project.

[0092] Step S230: Evaluate the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback.

[0093] Specifically, when implementing, the method of evaluating the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback is as follows:

[0094] Reference Figure 3 ,In a multi-agent system, a pre-trained preference learning alignment model is used as a reflector to improve the reasoning ability. The reflector will review the answers given by the manager and provide improvement suggestions for their content and format. Among them, the role of the reflector is: responsible for judging the correctness of the answers given by the manager; if the reflector judges that the answer is correct, it will give further reflections. When the manager is about to execute the second or more runs for the same task input, the reflector will intervene. If the reflector judges that there is no room for improvement in the answers given by the manager, the manager will no longer execute the current run. Otherwise, the reflector will further summarize the areas where the manager can improve, such as not considering a few highly rated items / movies in the user's historical interactions. In this exemplary embodiment, the pre-trained preference learning alignment model is guided by ASFT to make LLMs simultaneously focus on improving responses that conform to human preferences and reducing those that do not conform to human preferences; among them, ASFT (Aligned Supervised Fine-Tuning through Absolute Likelihood) is a technique for optimizing large language models (LLMs), aiming to make the output of the model closer to human preferences and expectations. The benefits of applying ASFT to large language models (LLMs) are shown in Table 1:

[0095] Table 1 Overview of ASFT Performance Improvement

[0096]

[0097]

[0098] In the above exemplary embodiment, the method of obtaining optimization feedback is introduced. Next, the method of training the preference learning alignment model is introduced:

[0099] In this exemplary embodiment, the preference learning alignment model is trained by the following method:

[0100] Construct a sample set including a number of samples; wherein, the samples include: sample data and label data; the sample data includes an initial answer for training; the label data includes an expected feedback for training; input the sample data into the preference learning alignment model to obtain predicted data output by the model, wherein the predicted data includes an optimized feedback for training output by the model; determine the similarity and difference between the predicted data and the label data; construct a composite objective function based on the similarity and the difference; update the model parameters through the composite objective function to obtain the pre-trained preference learning alignment model.

[0101] Specifically, when implementing, construct a sample set including a number of samples; wherein, the samples include: sample data and label data; the sample data includes an initial answer for training; the label data includes an expected feedback for training; input the sample data into the preference learning alignment model to obtain predicted data output by the model, wherein the predicted data includes an optimized feedback for training output by the model; determine the similarity and difference between the predicted data and the label data; construct a composite objective function based on the similarity and the difference; update the model parameters through the composite objective function to obtain the pre-trained preference learning alignment model in the following manner:

[0102] By constructing a sample set containing sample data (initial answer for training) and label data (expected feedback for training), a basis for model learning and optimization is provided; input the sample data into the model, the model outputs predicted data (optimized feedback for training), and then the similarity and difference are determined by comparing the predicted data and the label data; design a loss function to quantify the difference between the model output and human preferences, and adjust the model parameters by optimizing this loss function; based on the gradient flow analysis framework, the model updates the parameters according to the gradient of the composite objective function in each iteration, gradually increasing the similarity between the output and human preferences while reducing the difference between the output and human preferences.

[0103] In the above exemplary embodiment, the method for training the preference learning alignment model is introduced. Next, specifically introduce the method for constructing a composite objective function based on the similarity and the difference; updating the model parameters through the composite objective function to obtain the pre-trained preference learning alignment model:

[0104] In this exemplary embodiment, constructing a composite objective function based on the similarity and the difference; updating the model parameters through the composite objective function to obtain the pre-trained preference learning alignment model includes:

[0105] Construct a supervised fine-tuning loss function and an alignment loss function based on the similarity and the difference; combine the supervised fine-tuning loss function and the alignment loss function to obtain the composite objective function; update the model parameters through the composite objective function until the similarity between the predicted data and the label data is maximized while the difference is minimized, so as to obtain the pre-trained preference learning alignment model.

[0106] When specifically implementing, the method for constructing a supervised fine-tuning loss function and an alignment loss function based on the similarity and the difference; combining the supervised fine-tuning loss function and the alignment loss function to obtain the composite objective function; updating the model parameters through the composite objective function until the similarity between the predicted data and the label data is maximized while the difference is minimized, so as to obtain the pre-trained preference learning alignment model is as follows:

[0107] First, transform the preference response score, that is, define the score function S(x, y ω ), which is based on the conditional probability π θ (y ω |x), and is converted into an optimization objective through the sigmoid function, as shown in the following formula:

[0108]

[0109] In the formula, S(x, y ω ) represents the score or scoring function given the input x and the preference response y ω ; π θ (y ω |x) represents the conditional probability that the preference response y ω occurs given the input x under the model parameters θ; log represents the logarithmic function, which is used to transform the probability value to make it suitable for optimization algorithms such as gradient descent; θ represents the model parameters, including weights and biases, etc., which will be optimized during the training process; 1 - π θ (y ω |x) represents the probability that the non-preference response occurs given the input x.

[0110] This objective aims to maximize the score logσ(f θ (x, y w )) of the preference response and minimize the score 1 - σ(f θ (x, y l )) of the non-preference response, that is, let The following optimization objective is obtained:

[0111]

[0112] In the formula, πθ represents the probability distribution under the model parameter θ, usually used to represent conditional probability; E represents the expectation, which is used to calculate the expectation of a certain random variable under the probability distribution; ∼ represents "obeys" or "comes from", and is used to represent that the random variable (x, y ω , y l ) obeys or comes from the dataset D; f θ represents the score function under the model parameter θ, which is used to calculate the scores of the preference response and the non-preference response.

[0113] Based on this goal, the alignment loss function of the ASFT method is proposed. This function combines the negative log-likelihood of the preference and non-preference responses. That is, based on this goal, the alignment loss function of the ASFT method is proposed:

[0114]

[0115] By performing a gradient flow analysis on , it is found that the alignment loss in ASFT has a balanced impact on the parameters x 1 and x 2 , enabling the model to learn to generate responses preferred by humans while avoiding generating disliked responses. After performing a transformation on , we get That is, a similar transformation is performed on :

[0116]

[0117] For each pairwise preference data (x, y w , y l ) ∈ D, the update rate of x 1 with respect to x 2 is Since x 1 tends to increase and x 2 tends to decrease during the optimization process, we have This indicates that updates x 1 and x 2 in a balanced manner. Therefore, ASFT effectively identifies and processes the key factors in different scenarios during the optimization process. It prioritizes solving the main problems without showing a fixed bias towards x 1 or x 2 . This feature enables ASFT to avoid unnecessary actions, thereby improving the optimization efficiency. In other words, ASFT guides LLMs to simultaneously focus on improving responses that conform to human preferences and reducing those that do not.

[0118] In addition, based on Regarding the characteristics of the gradient field, it is found that ASFT is not sensitive to the initial model and can maintain robustness at different starting points. ASFT can increase the probability of generating human-preferred responses while reducing the probability of generating responses that humans dislike. Finally, by combining the supervised fine-tuning (SFT) loss and the alignment loss, a composite objective function is proposed. That is, by combining the supervised fine-tuning (SFT) loss and the alignment loss, the following composite objective function is proposed to optimize the model's language generation ability and alignment with human preferences simultaneously:

[0119]

[0120] Among them, Following the standard method of minimizing the negative log-likelihood ratio loss, which is derived from causal language modeling. The goal is to maximize the probability of generating the target token given the input context.

[0121] Step S240: Adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer.

[0122] Specifically in implementation, the way to adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer:

[0123] Referring to Figure 3 , the reflector will be responsible for analyzing and reflecting on the manager's answer in the last trial and giving suggestions, such as modifying the answer format to follow the task requirements. Finally, the manager will solve the task according to the reflector's feedback and provide a more reasonable answer, for example, adding the missing item ID.

[0124] Based on the above exemplary embodiments, in this solution, by integrating the preference learning alignment model into the multi-agent system as a reflector, the reasoning and decision-making capabilities of the system are improved, thus achieving better performance in the recommendation task. Specifically:

[0125] The performance of this solution for the rating prediction task on the ml-100k and Beauty datasets is shown in Table 2 as follows:

[0126] Table 2 Analysis Table of Results for ml-100k and Beauty Datasets

[0127]

[0128] RMSE and MAE respectively represent the root mean square error and the mean absolute error. The results show that the proposed solution in this invention is basically on par with the baseline method on the Beauty dataset and performs slightly better on the ml-100k dataset.

[0129] Sequential recommendation systems predict a user's next possible interests by analyzing the sequence of items the user has interacted with. In sequential recommendation tasks, modeling the user's long-term and short-term interests is very important. Therefore, the role of user analysts is self-evident. The number of relevant items in the sequence is significantly higher than in the rating prediction task. It is difficult to require item analysts to analyze every item that appears in the historical or candidate sets. In addition, considering that the answer to the sequential recommendation task is more complex (i.e., a ranking order of the candidate set), reflectors can help prevent managers from getting into formatting troubles. Single-round behavior analysis may miss considering the user's long-term behavior.

[0130] The performance of this solution for sequential recommendation tasks on the ml-100k and Beauty datasets is shown in Table 3 as follows:

[0131] Table 3 Performance Comparison Table for Sequential Recommendation Tasks

[0132]

[0133] Among them, HR@K, that is, Hits@K, represents the proportion that at least one of the top K items recommended by the system hits the actual user's selection. The higher HR@K is, the greater the probability that the recommendation system accurately recommends the items the user likes among the top K recommendations; NDCG@K is a ranking metric that considers not only whether the recommendation system hits the items the user likes, but also whether the order of recommendation is reasonable. The higher NDCG is, the more it means that the system can not only hit the items the user prefers, but also rank these items in a relatively front position. The results show that the metrics of this solution on the Beauty dataset are all better than other baseline methods, and it also performs well on the ml-100k dataset.

[0134] It should be noted that the method of this embodiment of the present disclosure can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In such a distributed scenario, one of these multiple devices can only execute one or more steps of the method of this embodiment of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0135] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0136] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a multi-agent recommendation device combined with preference learning.

[0137] Reference Figure 4 , the multi-agent recommendation device combined with preference learning includes:

[0138] A demand information determination module 410, configured to determine the task query information of the query user, and translate the task query information to obtain task demand information;

[0139] An initial answer determination module 420, configured to reason based on the task demand information to obtain an initial answer;

[0140] An optimization feedback determination module 430, configured to evaluate the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback;

[0141] A recommended answer determination module 440, configured to adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer.

[0142] In this exemplary embodiment, the demand information determination module 410 is specifically configured to:

[0143] Determine the task query information of the query user, and translate the task query information to obtain task demand information.

[0144] In this exemplary embodiment, the initial answer determination module 420 is specifically configured to:

[0145] Extract keywords from the task demand information to obtain demand information keywords; search for demand information reference data based on a search engine for the demand information keywords; screen and summarize the demand information reference data to obtain the relevant information; construct a user profile for the query user based on an information database; perform project analysis on the task query information based on the information database to obtain project attributes; retrieve the query user and the task query information based on an interactive retriever to obtain the interaction history between the query user and the task query information; analyze based on the user profile, the project attributes, and the interaction history to obtain the preference information; summarize the relevant information and the preference information to obtain the initial answer.

[0146] In this exemplary embodiment, the optimization feedback determination module 430 is specifically configured to:

[0147] Evaluate the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback, where the preference learning alignment model is trained by the following method: construct a sample set including a number of samples; wherein, the samples include: sample data and label data; the sample data includes the initial answer for training; the label data includes the expected feedback for training; input the sample data into the preference learning alignment model to obtain predicted data output by the model, wherein the predicted data includes the optimized feedback for training output by the model; determine the similarity and difference between the predicted data and the label data; construct a supervised fine-tuning loss function and an alignment loss function based on the similarity and the difference; combine the supervised fine-tuning loss function and the alignment loss function to obtain the composite objective function; update the model parameters through the composite objective function until the similarity between the predicted data and the label data is maximized while the difference is minimized, to obtain the pre-trained preference learning alignment model.

[0148] In this exemplary embodiment, the recommended answer determination module 440 is specifically configured to:

[0149] Adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer.

[0150] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0151] The device of the above embodiment is used to implement the corresponding multi-agent recommendation method combined with preference learning in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0152] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the multi-agent recommendation method combined with preference learning in any of the above embodiments.

[0153] Figure 5 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0154] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0155] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0156] The input / output interface 1030 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0157] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0158] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0159] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification and does not necessarily include all the components shown in the figure.

[0160] The electronic device of the above embodiment is used to implement the corresponding multi-agent recommendation method for combined preference learning in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0161] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the multi-agent recommendation method for combined preference learning as described in any of the foregoing embodiments.

[0162] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0163] The above non-transitory computer-readable storage medium can be any available medium or data storage device accessible by a computer, including but not limited to magnetic memory (such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc.), optical memory (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid state drives (SSD)), etc.

[0164] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the multi-agent recommendation method for combined preference learning as described in any of the foregoing exemplary method embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0165] Based on the same inventive concept, corresponding to the multi-agent recommendation method with combined preference learning described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to execute the multi-agent recommendation method with combined preference learning. Corresponding to the execution subjects of the respective steps in the respective embodiments of the multi-agent recommendation method with combined preference learning, the processors that execute the corresponding steps can belong to the corresponding execution subjects.

[0166] The computer program product of the above embodiments is used to cause the computer and / or the processor to execute the multi-agent recommendation method with combined preference learning described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0167] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to as "circuit", "module", or "system" in this document. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.

[0168] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive examples) of the computer-readable storage medium can include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in combination with an instruction execution system, apparatus, or device.

[0169] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0170] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0171] The computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0172] It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine. These computer program instructions, when executed by a computer or other programmable data processing device, produce an apparatus that implements the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0173] These computer program instructions can also be stored in a computer-readable medium that can cause a computer or other programmable data processing device to operate in a specific manner. In this way, the instructions stored in the computer-readable medium produce a product that includes an instruction apparatus for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0174] Computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operation steps are executed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, thereby enabling the instructions executed on the computer or other programmable apparatus to provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0175] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the illustrated operations must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart may be changed in the order of execution. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0177] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0178] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; Under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and for the sake of brevity, they are not provided in detail.

[0179] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0180] Although the present application has been described in connection with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0181] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.

[0182] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit, and this division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.

Claims

1. A multi-agent recommendation method combined with preference learning, characterized in that: include: Determine the task query information of the querying user, translate the task query information, and obtain task requirement information; Perform reasoning based on the query user, the task query information, and the task requirement information to obtain an initial answer; The initial answer is evaluated by a pre-trained preference learning alignment model to obtain optimization feedback; The initial answer is adjusted based on the optimization feedback to obtain an optimized recommended answer.

2. The method according to claim 1, characterized in that The reasoning based on the query user, the task query information and the task requirement information to obtain an initial answer includes: Perform a keyword search based on the task requirement information to obtain relevant information of the task requirement information; Analyzing the query user and the task query information to obtain the preference information of the query user; The initial answer is obtained by aggregating the relevant information and the preference information.

3. The method according to claim 2, characterized in that The keyword search based on the task requirement information is performed to obtain relevant information of the task requirement information, including: Extracting keywords from the task requirement information to obtain keywords of the requirement information; Searching the demand information keywords based on a search engine to obtain demand information reference data; The demand information reference data is screened and summarized to obtain the relevant information.

4. The method according to claim 2, characterized in that: The analyzing the query user and the task query information to obtain the preference information of the query user includes: Constructing the query user based on the information database to obtain a user portrait; Performing project analysis on the task query information based on the information database to obtain project attributes; Retrieving the query user and the task query information based on an interactive retriever to obtain an interaction history between the query user and the task query information; The preference information is obtained by analyzing the user portrait, the project attributes and the interaction history.

5. The method according to claim 1, characterized in that The method further comprises training the preference learning alignment model by: Constructing a sample set including several samples; wherein the samples include: sample data and label data; the sample data includes initial answers for training; the label data includes expected feedback for training; Inputting the sample data into the preference learning alignment model to obtain prediction data output by the model, wherein the prediction data includes training optimization feedback output by the model; Determining similarities and differences between the predicted data and the label data; Based on the similarity and the difference, a composite objective function is constructed; the model parameters are updated by the composite objective function to obtain the pre-trained preference learning alignment model.

6. The method according to claim 5, characterized in that Based on the similarity and the difference, a composite objective function is constructed; and the model parameters are updated by the composite objective function to obtain the pre-trained preference learning alignment model, including: Constructing a supervised fine-tuning loss function and an alignment loss function based on the similarity and the difference; Combining the supervised fine-tuning loss function with the alignment loss function to obtain the composite objective function; The model parameters are updated through the composite objective function until the similarity between the predicted data and the label data is maximized while the difference is minimized, thereby obtaining the pre-trained preference learning alignment model.

7. A multi-agent recommendation device combined with preference learning, characterized in that: include: A requirement information determination module is configured to determine the task query information of the querying user, translate the task query information, and obtain task requirement information; An initial answer determination module is configured to perform reasoning based on the query user, the task query information and the task requirement information to obtain an initial answer; An optimization feedback determination module is configured to evaluate the initial answer through a pre-trained preference learning alignment model to obtain optimization feedback; The recommended answer determination module is configured to adjust the initial answer based on the optimization feedback to obtain an optimized recommended answer.

8. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Multi-agent collaborative data query method and device

    CN120470019A

  • Large model reasoning optimization method and device based on background information fusion, equipment and medium

    CN120764692A

  • Artificial intelligence-based product data analysis methods, devices, equipment, and media

    CN122570554A

  • A scoring prediction model based on multi-agent debate enhancement

    CN122713444A