Data processing method and device, information recommendation system, electronic equipment and medium

By using Bayesian project distribution and large language models in the personalized dialogue recommendation system, natural language query statements are generated and user feedback is received, which solves the problem that the system is difficult to grasp user preferences in the early stage of startup, and achieves the effect of quickly identifying user core preferences.

CN120011550AActive Publication Date: 2025-05-16BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510093299.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

It is difficult for a personalized dialogue recommendation system to quickly grasp the user's natural language preferences in the early stages of launching, and large language models lack decision-making theoretical reasoning when generating dialogues, which makes it difficult to balance exploration and development.

Method used

Using the current probability belief state based on Bayesian project distribution, the item description text is selected from the description text set, and a natural language query statement is generated based on the pre-trained large language model, and user feedback information is received to extract user preference value.

Benefits of technology

By receiving user feedback information and extracting user preference values, it can better reflect user preferences. Combining Bayesian optimization methods and large language models, quickly identify user core preferences and improve the accuracy and efficiency of preference extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011550A_ABST
    Figure CN120011550A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, and relates to the technical field of data processing, in particular to the technical fields of natural language processing, statistics, machine learning, large language models and the like. According to the specific implementation scheme, the method comprises the steps of selecting a project description text from a description text set based on a current probability belief state of Bayesian project distribution; obtaining and sending a natural language query statement based on the project description text and a pre-trained large language model; receiving feedback information of the user on the natural language query statement; and based on the feedback information, extracting a preference degree value of the user to the item description text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to technical fields such as natural language processing, statistics, machine learning, and large language models, and more particularly to a data processing method and device, an information recommendation system, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the early stages of the personalized conversational recommendation system, an effective natural language preference guidance strategy is urgently needed. This strategy should be able to quickly understand the user's favorites without the user knowing anything about the system. Ideally, the user's preferences can be understood simply through some random item descriptions. The rise of large language models provides technical support for this natural language preference guidance.

[0003] However, these models lack the necessary decision-theoretic reasoning when generating dialogues and cannot effectively balance exploration and exploitation, which may lead to excessive attention to known preferences or wasting too much time on low-value items. In addition, a single large language model also faces challenges in processing a large number of unknown item descriptions, as well as in terms of system behavior control and interpretability. Summary of the invention

[0004] The present disclosure provides a data processing method and device, an information recommendation system, an electronic device, and a computer-readable storage medium.

[0005] According to a first aspect, a data processing method is provided, the method comprising: selecting a project description text from a description text set based on a current probabilistic belief state of a Bayesian project distribution; obtaining and sending a natural language query statement based on the project description text and a pre-trained large language model; receiving user feedback information on the natural language query statement; and extracting a user preference value for the project description text based on the feedback information.

[0006] According to a second aspect, a data processing device is provided, which includes: a selection unit, configured to select project description text from a description text set based on the current probabilistic belief state of the Bayesian project distribution; an acquisition unit, configured to obtain and send a natural language query statement based on the project description text and a pre-trained large language model; a query unit, configured to receive user feedback information on the natural language query statement; and an extraction unit, configured to extract the user's preference value for the project description text based on the feedback information.

[0007] According to a third aspect, an information recommendation system is provided, the system comprising: a user terminal and a server; the user terminal sends an information recommendation request to the server; the server executes the data processing method described in any implementation method of the first aspect based on the information recommendation request, and sends a natural language query statement to the user terminal; the user terminal is used to receive the natural language query statement sent by the server, and feed back the user's feedback information on the natural language query statement to the server; the server extracts the user's preference value for the item description text in the description text set based on the feedback information, generates an information recommendation result based on the preference degree, and sends the information recommendation result to the user terminal.

[0008] According to a fourth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect.

[0009] According to a fifth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in any implementation of the first aspect.

[0010] The data processing method and device provided by the embodiment of the present disclosure first selects a project description text from a description text set based on the current probability belief state of the Bayesian project distribution; secondly, obtains and sends a natural language query statement based on the project description text and a pre-trained large language model; then, receives user feedback information on the natural language query statement; finally, extracts the user's preference value for the project description text based on the feedback information. Thus, by receiving user feedback information on the natural language query statement and extracting the user's preference value, the user's preference can be better reflected; combining the Bayesian optimization method with the large language model can quickly identify the user's core preference within a small number of dialogue rounds, thereby improving the accuracy and efficiency of preference extraction.

[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0013] Figure 1 is a flow chart of an embodiment of a data processing method according to the present disclosure;

[0014] Figure 2 It is a structural schematic diagram of the data processing flow in the present disclosure;

[0015] Figure 3 It is a schematic diagram of the structure of the probability belief state of the present disclosure after multiple rounds of natural language dialogue;

[0016] Figure 4 is a schematic structural diagram of an embodiment of a data processing device according to the present disclosure;

[0017] Figure 5 is a structural diagram of an embodiment of the information recommendation system according to the present disclosure;

[0018] Figure 6 It is a block diagram of an electronic device used to implement the data processing method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Unless explicitly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising”, etc., will be understood to include the stated elements or components but not to exclude other elements or components.

[0020] The technical solution of the present disclosure is described below by means of specific embodiments. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.

[0021] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0022] In the field of education, personalized learning resource recommendation systems need to quickly understand students' learning preferences and needs at the initial stage of launch in order to provide the most suitable learning materials and resources. Traditional methods often rely on students' direct ratings or choices, but when students are unfamiliar with a large number of learning resources, this method is inefficient and has limited effects.

[0023] In the prior art, Bayesian optimization methods are mainly used in question answering and recommendation systems based on fixed templates. These methods usually rely on predefined query templates to guide user feedback and update the user preference model. For example, the system may ask users about their preferences for certain specific features, and then adjust the recommendation strategy based on the user's answer. This method is much less effective when users are unfamiliar with the items or the item features are diverse, because the fixed template cannot flexibly adapt to the complexity of various item descriptions and user preferences.

[0024] With the development of Large Language Models (LLMs), some recommendation systems have begun to use these models to generate more natural conversations and recommendations. These systems can understand and generate complex natural language, thereby interacting with users in a more humane way. These systems often lack effective decision-making theory support when generating conversations, and cannot strike a good balance between exploring new items and using known information, resulting in low recommendation efficiency.

[0025] The single large language model also has the limitation of prompts. Specifically, an intuitive natural language preference elicitation (NL-PE) method is to provide a single LLM with the description x of all products and the conversation history H in each round of the conversation. t and instructions for generating new queries, wherein the instruction is a hint provided to the LLM to guide it to generate a new query to further explore the user's preferences. The instruction may include some instructions or examples to help the LLM understand what kind of query should be generated to effectively guide the user to express his or her preferences.

[0026] However, for most sets of products, it is not possible to describe all products [x1, ..., x N ] into the context window of the LLM is computationally expensive. Although large language models can be made to internalize product knowledge through fine-tuning, each update of product information means the need to retrain the entire system. More importantly, the behavior of the LLM in terms of preference elicitation cannot be controlled except through hint engineering or further fine-tuning, neither of which guarantees that the model will provide predictable and explainable behavior when developing and exploring user preferences.

[0027] In recommender systems, exploration vs. exploitation is a core issue. Existing algorithms, such as Thompson sampling and UCB, aim to address this issue, but these algorithms usually require users to directly provide ratings or comparisons, which is impractical when users are unfamiliar with most items. These algorithms do not work well when applied to item recommendations based on natural language descriptions because they cannot effectively handle the diversity and ambiguity of natural language.

[0028] In view of the defects in traditional technologies, the present disclosure proposes a data processing method, which can quickly and effectively extract the subject's learning preferences through the Bayesian optimization acquisition method of a large language model. Figure 1 A process 100 according to an embodiment of a data processing method of the present disclosure is shown. The data processing method comprises the following steps:

[0029] Step 101 : selecting a project description text from a description text set based on the current probability belief state of the Bayesian project distribution.

[0030] In this embodiment, the Bayesian item distribution refers to the prior or posterior distribution of the user's utility for multiple items, which can be updated using the Bayesian optimization strategy. The union of the distribution parameters of all current items in the Bayesian item distribution is the current probabilistic belief state. Among them, the item is related to the user's preference for a certain object, and is a category or specific object that the user may be interested in, which is described in natural language. For example, the item can be "children's books" or "science fiction movies", and the user's utility value represents the user's preference for these items. In the present disclosure, the item can be a natural language description text of various aspects of the Bayesian item distribution, such as the item is a product name, the item is a book name, etc. The item is the carrier of the user's preference, and the utility is a quantitative representation of the user's preference for these items. Specifically, the utility can be a value in the range of [0, 1]. The Bayesian optimization strategy can be used to select the item description text with the optimal utility from the Bayesian item distribution, that is, to obtain the optimal item with the best utility.

[0031] In this embodiment, the Bayesian item distribution can adopt the Beta distribution. Although the Beta distribution is a natural and convenient choice, other distributions can also be considered to represent the probabilistic belief state according to the specific application scenario and system design requirements. For example, the normal distribution (Gaussian distribution): the utility can be constrained to the interval [0, 1] by transformation (for example, by logistic transformation). However, the normal distribution itself is not naturally suitable for probability representation in the interval [0, 1]. For example, the logarithmic probability distribution (Logit-Normal Distribution): This is a distribution obtained by applying the normal distribution to the logarithmic probability transformed variable, which can be used to represent the probability in the interval [0, 1]. Another example is the Dirichlet distribution: if multiple categories of preferences are considered instead of binary preferences, the Dirichlet distribution can be used as a conjugate prior for the multinomial distribution.

[0032] In this embodiment, when selecting an alternative distribution to represent the probabilistic belief state, multiple factors need to be considered comprehensively. First, the Beta distribution is an ideal choice for modeling probabilities due to its definition in the interval [0, 1] and its conjugation with the Bernoulli distribution, especially for cold start scenarios of user preferences. However, if a more complex model is needed to capture the relationship between preferences, other distributions can be considered. In this embodiment, the description text set includes at least one item description text, such as Figure 2 As shown, a description text set x, each item description text in the description text set can be a text of a user-related object of interest (such as a product, a course). At the beginning of a conversation with the user, it is assumed that the user's preference utility for each item description text in the description text set obeys a uniform distribution, and the uniform distribution has corresponding distribution parameters (i.e., an initial probability belief state). In each subsequent round of conversation, the execution subject on which the data processing method runs obtains the user's preference degree value, and based on the preference degree value, updates the distribution parameters representing the preference distribution to obtain the current probability belief state.

[0033] In this embodiment, for the commodity preference scenario, the description text set is a commodity description text set, and the item description text can be a text selected from a predefined commodity description text set. These item descriptions are known item descriptions in the system and can be any natural language descriptions, such as product descriptions, book introductions, etc. These descriptions constitute the "knowledge base" of the system's optional items and are the basis for the system to guide preferences. These natural language description texts can come from product databases, knowledge graphs, etc.

[0034] In this embodiment, the current probability belief state is the state value of the distribution parameter of the user's preference utility following the distribution, wherein the content of the distribution parameter is different when the principle of following the distribution is different; for example, the preference utility follows the Beta distribution, and its parameter is αi and β i , then the current probability belief state is α i and β i Status value.

[0035] In this embodiment, project description text can be selected from the description text set according to the current probability belief state of the Bayesian project distribution in a variety of ways, such as inputting the current probability belief state and the description text set into a pre-trained selection model to obtain the project description text output by the selection model.

[0036] Step 102: obtain and send a natural language query statement based on the project description text and the pre-trained large language model.

[0037] In this embodiment, a pre-trained large language model is used to generate natural language queries. This pre-trained large language model has been pre-trained on a large amount of public text data and has strong natural language understanding and generation capabilities. Through prompt engineering, this pre-trained large language model can be guided to generate natural language queries that meet the requirements.

[0038] In this embodiment, the natural language query sentence is a sentence sent to the user to inquire about the user's preference information. The above step 102 includes: inputting the project description text and the query prompt word into the pre-trained large language model to obtain the natural language query sentence output by the pre-trained large language model. The query prompt word is used to prompt the pre-trained large language model to generate a natural language query sentence.

[0039] In this embodiment, the execution subject on which the data processing method runs is to ask the user in real time whether he likes a specific aspect in each round of dialogue and wait for the user's reply. This real-time interaction is a key step in the entire preference extraction process, because the user's reply will be used to accurately infer the user's preference.

[0040] Step 103: Receive user feedback on the natural language query statement.

[0041] In this embodiment, the natural language query statement is an inquiry statement. After obtaining the natural language query statement, the user feeds back feedback information to the execution entity on which the data processing method runs, wherein the feedback information is the user's reply to the natural language query statement. The feedback information can reflect the user's positive or negative feedback on various aspects of the natural language query statement.

[0042] In this embodiment, the execution subject on which the data processing method runs will select a project description text that best balances exploration and development according to the current probabilistic belief state in each round of dialogue, and generate a natural language query statement (such as "Are you interested in books for children to read?") for asking the user. The user's response to the natural query statement (such as "yes" or "no") will be captured by the execution subject, and the user's preference value for each project description text will be inferred through the natural language inference (NLI, Natural Language Interface) model. This process can be iterative, and each round of replies will be used to update the current probabilistic belief state, thereby guiding the next round of query generation. Therefore, the whole process is a real-time, interactive dialogue process, and the system will adjust its behavior in each round according to the user's immediate feedback to more effectively extract the user's natural language preferences.

[0043] Step 104: extracting the user's preference value for the item description text based on the feedback information.

[0044] In this embodiment, the preference value is used to reflect the user's preference for the item description text. The preference value can be expressed as a probability value w i t (i>1, used to represent items, t>1, used to represent conversation turns) is used to represent the probability value w i t It can be obtained through a natural language inference model, representing the project description text x i Contains the probability of describing the user's preference. This probability value can be used to update the user's utility belief about the item, that is, the current probability belief state.

[0045] Specifically, Figure 2 As shown, suppose that in the tth (t>1) round of dialogue, the system generates a natural language query sentence q t , and got feedback from users t Through the NLI model, the system can infer the item description text x i Whether it contains the user's preference description text and gives a probability value w i t (i∈1..N, N>1). This probability value reflects the user's evaluation of the item description text x i degree of preference.

[0046] The data processing method provided in this embodiment realizes natural language preference extraction by formalizing Bayesian optimization of any natural language project description text. At the same time, the present disclosure also introduces and evaluates an algorithm called preference elicitation using Bayesian optimization and large language models, which enhances the ability of large language models in natural language preference extraction through Bayesian optimization.

[0047] In this embodiment, the above data processing method can be applied to the field of education, and the description text set can be a random course introduction or learning material description of multiple learning subjects. The execution subject on which the above data processing method runs can interact with the learning subject through natural language, quickly extract the learning preference of the learning subject, and thus infer the learning subject's interest in different learning resources. Applying the data processing method to the field of education can significantly improve the efficiency and accuracy of personalized learning resource recommendations. Through natural language interaction, the system can communicate with students more naturally and have a deep understanding of their needs. At the same time, it uses decision theory to optimize the recommendation strategy to ensure the best use of educational resources. This not only improves students' learning experience, but also promotes the development of educational technology, and provides solid technical support for the realization of truly personalized education.

[0048] The data processing method provided by the embodiment of the present disclosure first selects a project description text from a description text set based on the current probability belief state of the Bayesian project distribution; secondly, obtains and sends a natural language query statement based on the project description text and a pre-trained large language model; then, receives user feedback information on the natural language query statement; finally, extracts the user's preference value for the project description text based on the feedback information. Thus, by receiving user feedback information on the natural language query statement and extracting the user's preference value, the user's preference can be better reflected; combining the Bayesian optimization method with the large language model can quickly identify the user's core preference within a small number of dialogue rounds, thereby improving the accuracy and efficiency of preference extraction.

[0049] In some embodiments of the present disclosure, the above-mentioned data processing also includes: updating the current probability belief state based on the preference degree value to obtain an updated probability belief state; detecting whether the preference extraction stop condition is met based on the updated probability belief state and the current probability belief state; in response to detecting that the preference extraction stop condition is not met, using the updated concept belief state as the current concept belief state.

[0050] In this embodiment, based on the preference degree value, the current probability belief state is updated to obtain the updated probability belief state, which includes: selecting a parameter value of a probability distribution parameter that matches the preference degree value from a preset preference probability table; and replacing the current probability belief state with the parameter value to obtain the updated probability belief state.

[0051] In this embodiment, the stop condition refers to the condition for deciding when to stop the next round of preference value extraction process. That is, this stop condition is used to determine whether to continue to execute the preference value extraction method combining Bayesian optimization and a large language model (the data processing method of the present disclosure) to avoid continuing preference exploration in unnecessary dialogue rounds. Specifically, the stop condition can be when a preset number of dialogue rounds is reached, or when the update amplitude of the probability belief state is lower than a preset amplitude threshold, wherein the preset amplitude threshold can be set based on development requirements.

[0052] In this embodiment, the stop condition is used to limit the continuation of the preference extraction process to avoid wasting resources in unnecessary dialogue rounds, or to stop redundant exploration when the user has clearly expressed the preference or the system has sufficiently understood the user's preference.

[0053] In this embodiment, the updated concept belief state is used as the current concept belief state, and the next round of dialogue can be conducted based on the current concept belief state, thereby updating the user's preference value.

[0054] exist Figure 2 In the process, the system continuously accumulates information about user preferences through multiple rounds of dialogue and uses this information to update the current probability belief state. In each round of dialogue, the system generates a natural language query statement to obtain the user's feedback information and obtains the probability value w through the NLI model. i t , and then use this probability value w i t Update the parameters of the Beta distribution to gradually approach the user's true preferences.

[0055] Figure 3 The data processing method of the present disclosure is used to perform a preference heuristic algorithm through three rounds (such as Figure 3 The natural language dialogue corresponding to the images at t=1, t=2, and t=3 in the figure gradually updates the cold start (the image corresponding to t=0) user's utility u for the item. i The current probability belief state exist Figure 3In the figure, during the cold start, the natural language query generated by the execution subject is: "Are you looking for a children's book?", and the user's feedback in the t = 1 round of dialogue is "Yes"; the natural language query generated by the execution subject in the t = 1 round of dialogue is: "Are you interested in magic?", and the user's feedback in the t = 2 round of dialogue is "Yes"; the natural language query generated by the execution subject in the t = 2 round of dialogue is: "Do you like reading?", and the user's feedback in the t = 3 round of dialogue is "Yes". The images corresponding to the moments t = 0, t = 1, t = 2, and t = 3 are the horizontal coordinates of the item utility u i , the vertical axis is the probability distribution diagram of the current probability belief state of each round of dialogue. The items in the image corresponding to t=0, t=1, t=2, and t=3 include s1, s2, s3, and s4, which are used to represent different children's books respectively.

[0056] The following is a detailed description of the process of obtaining the current probability belief state:

[0057] Before starting any conversation, an initial uninformative belief p(u) about the utility relationship between users and items is set.

[0058] Assuming that the utility of each item is independent of each other, we get the utility formula shown in formula (1).

[0059]

[0060] And each utility u i The prior distribution of is a Beta distribution

[0061]

[0062] This paper discloses a completely cold start setting, assuming a uniform Beta prior distribution with parameters Beta distribution lies in the interval [0, 1] - this is the normalization interval of the finite ratings in the classic recommendation system. Therefore, the utility value u i =1 or u i = 0 is interpreted as a complete like or dislike for item i, while u i Values ​​of ∈(0, 1) provide preference strengths between these two extremes.

[0063] In order to t To update the utility belief a posteriori, an observation model is needed to represent the likelihood p(r t |x,u,q t ). Modeling t The likelihood of is a challenging task, so some simplifying assumptions are needed. First, assume that the single response r tThe likelihood of any previous dialogue history H t-1 is independent, so:

[0064]

[0065] Having determined the decomposed distribution of item utility and observation history probability, we need to provide a specific observation model that conditions the response probability based on the query, item description, and potential utility: p(r t |x,u,q t ).

[0066] Since the prior distribution is for conditionally independent u i To decompose, we can introduce a separate binary response for each item Here It can be 0 (for No) or 1 (for Yes) to express the relevance of individual preferences for each item i in round t. Importantly, we do not really need to have a separate response for each item - this will be handled by the natural language inference model, but for simplicity, we start with a separate binary response model for each item.

[0067]

[0068] After determining the response probabilities, this leads us to our first attempt to update the full posterior utility for the observed binary rating feedback. Specifically, when a series of binary ratings is observed When t = 1, the Beta prior distribution (Equation (2)) is combined with the Bernoulli likelihood function (Equation (4)) to construct a standard Beta-Bernoulli conjugate model, and the posterior utility belief is calculated based on it.

[0069]

[0070] In formula (6), The subsequent incremental update follows equation (4) and uses the same conjugacy property to give

[0071]

[0072] In formula (7),

[0073] In the data processing method provided by the present disclosure, the current probabilistic belief state is adopted, which not only effectively promotes personalized recommendations, but also allows the Bayesian optimization strategy to guide the query generation of large language models. This avoids over-exploration, that is, unnecessary inquiries about items that are obviously not of high value, and over-exploitation, that is, excessive concentration on the user's known preferences.

[0074] The data processing method provided in this embodiment updates the current probability belief state based on the preference degree value to obtain an updated probability belief state; detects whether the preference extraction stop condition is met based on the updated probability belief state and the current probability belief state; in response to detecting that the preference extraction stop condition is not met, uses the updated concept belief state as the current concept belief state; through the preference extraction stop condition, when the user has clearly expressed the preference or the system has sufficiently understood the user's preference, no redundant exploration is performed, thereby improving the efficiency of preference degree value extraction, enhancing the user experience, and ensuring that the preference extraction task is completed within a reasonable time and resource range; it can effectively adjust the estimate of the user's preference based on the user's natural language feedback, thereby improving the accuracy and efficiency of the recommendation.

[0075] In some optional implementations of the present disclosure, the above-mentioned updating of the current probability belief state based on the preference degree value to obtain the updated probability belief state includes: determining the beta distribution parameters of each item in the current probability belief state; updating the beta distribution parameters based on the preference degree value to obtain the updated beta distribution parameters; obtaining the updated probability belief state based on the updated beta distribution parameters.

[0076] In this optional implementation, the Beta distribution is used in the Bayesian item distribution to represent the prior and posterior distribution of utility, and the parameters (α i and β i ).

[0077] In this optional implementation, the preference value is used to reflect the user's preference for the item description text. The preference value can be expressed as a probability value w i t In the Bayesian update process, this probability value w i t Used to update the parameter α of the Beta distribution i and β i If w i t A higher value indicates that the user is more likely to prefer the item, so α can be increased. i The value of i t If it is lower, you can increase β iThis updating method is similar to the process of updating the prior distribution based on the observation results in the Beta-Bernoulli conjugate model.

[0078] In this optional implementation, the process of updating the probability belief state is based on the Bayesian update rule, assuming that the utility u of each item i It follows the Beta distribution with parameter α i and β i In each round of dialogue, according to the user's feedback information r t And the result of natural language reasoning, update α i and β i The steps for updating the parameters of the Beta distribution are as follows:

[0079] Initially, the utility u of each item is i The prior distribution of is Beta distribution with parameters α0=1 and β0=1, indicating uniform prior.

[0080] Assume that the user’s feedback information r t is binary (e.g., “yes” or “no”, as in Figure 2 N or Y in , and regard it as the result of the Bernoulli test.

[0081] For each item i, based on the output of the NLI model, we infer whether the item i meets the user’s preferences. If item i is considered to meet the user’s preferences, then r i t =1, otherwise r i t =0.

[0082] For each item i, according to the user's feedback information r i t , update the parameters of the Beta distribution:

[0083] α i t =α i t-1 +r i t ;

[0084] β i t =β i t-1 +(1-r i t );

[0085] Thus, the utility u of each item is i The posterior distribution of is:

[0086] p(u i |xi, H t )=Beta(α i t ,β i t )

[0087] The probabilistic belief state is the joint distribution of all item utility distributions, namely:

[0088] p(u|x,H t )=∏ i=1 N Beta(α i t ,β i t )

[0089] This joint distribution reflects the current belief state about the utility of all items. The expression of joint distribution represents the given condition x and historical information H. t The probability distribution of variable u is a product of N Beta distributions. Each Beta distribution has parameters α i t and β i t Determine, where i ranges from 1 to N (N is a natural number > 1).

[0090] The stopping condition can be reaching a preset number of dialogue rounds, or stopping when the update amplitude of the probability belief state is lower than a certain threshold. If the stopping condition is not met, continue to the next round of dialogue and repeat the parameter update step.

[0091] This optional implementation provides a method for updating the current probability belief state, determining the beta distribution parameters of each item in the current probability belief state; based on the preference degree value, updating the beta distribution parameters to obtain updated beta distribution parameters; based on the updated beta distribution parameters, obtaining the updated probability belief state, providing a reliable implementation method for updating the current probability belief state, and improving the reliability of updating the current probability belief state.

[0092] In some optional implementations of the present disclosure, the above-mentioned selection of project description text from the description text set based on the current probabilistic belief state of the Bayesian project distribution includes: determining a context acquisition function from any one of the Thompson sampling algorithm, the upper confidence limit algorithm, the entropy reduction strategy and the greedy strategy; controlling the context acquisition function to select project description text from the description text set under the current probabilistic belief state of the Bayesian project distribution.

[0093] In this optional implementation, the project description text selected from the description text set refers to selecting projects close to the optimal value, which can strike a balance between development (i.e., utilizing known high-value projects) and exploration (i.e., exploring unknown projects that may have high value); when selecting projects, the execution subject on which the data processing method runs will select projects close to the optimal value according to different Bayesian optimization strategies. The above-mentioned determination of the context acquisition function from any one of the Thompson sampling algorithm, the upper confidence limit algorithm, the entropy reduction strategy, and the greedy strategy includes: using any one of the Thompson sampling algorithm, the upper confidence limit algorithm, the entropy reduction strategy, and the greedy strategy as the context acquisition function.

[0094] In this optional implementation, the Thompson Sampling algorithm extracts a sample from the posterior distribution of each item and selects the item with the highest utility, so that the expected utility of the selected item is close to the optimal value.

[0095] In this optional implementation, the upper confidence bound algorithm (UCB) selects the items with the highest confidence upper bound, which are likely to have higher expected utility and thus closer to the optimal value. For example, the UCB strategy will tend to select items with higher uncertainty, which are likely to have higher expected utility and thus closer to the optimal value.

[0096] In this optional implementation, the entropy reduction strategy (ER) refers to optimizing the performance and stability of the system by reducing the disorder of the system and increasing the order of the system.

[0097] In this optional implementation, the greedy strategy (Greedy Algorithm) is an algorithm that takes the best or optimal choice in the current state in each step of selection, hoping to lead to the best or optimal result in the world. The core idea of ​​the greedy algorithm is to make a local optimal decision in each step of selection, that is, to select the best or optimal solution in the current state, and then continue to build a solution to the problem based on this choice until a complete solution is obtained or no more choices can be made.

[0098] In this optional implementation, γ C It stands for context acquisition function, which is used to select an item description text to generate a natural language query statement. The role of this context acquisition function is to decide which item description text to ask in each round of dialogue based on the current probabilistic belief state (i.e., the probability distribution of user preferences) in order to strike a balance between exploration and exploitation.

[0099] In this optional implementation, the context acquisition function γ CAlternative strategies, such as the Thompson sampling algorithm, the upper confidence bound algorithm, the entropy reduction strategy, and the greedy strategy, are all methods used to guide how to select item description text. The relationship between these strategies and the concept belief state is that they all make choices based on the current concept belief state. In other words, the above strategies all select the most appropriate item description to generate queries under the guidance of the belief state, in order to obtain the user's core preferences within the shortest number of dialogue turns.

[0100] This optional implementation provides a method for selecting project description text from a text collection, determining a context acquisition function from any one of the Thompson sampling algorithm, the upper confidence limit algorithm, the entropy reduction strategy, and the greedy strategy; controlling the context acquisition function to select project description text from the description text collection under the current probabilistic belief state of the Bayesian project distribution, and improving the reliability of obtaining the project description text through the Bayesian optimization algorithm.

[0101] In some optional implementations of the present disclosure, the above-mentioned obtaining and sending a natural language query statement based on the project description text and a pre-trained large language model includes: obtaining a conversation history text; based on the conversation history text, detecting whether the project description text is a repeated query text; in response to detecting that the project description text is not a repeated query text, inputting the project description text and the query prompt word into a pre-trained large language model to obtain a natural language query statement output by the large language model; and sending the natural language query statement.

[0102] In this optional implementation, the conversation history text is a natural language query statement in one or more rounds of conversations generated during the historical time period. The execution subject on which the data processing method runs will record all previously raised queries in each round of conversation, and when generating new queries, ensure that the new queries do not repeat the historical queries. Specifically, the conversation history text is obtained through the following steps: the execution subject on which the data processing method runs will maintain a query record list, which stores the queries generated and raised in all previous rounds. Each time a new query is generated, the execution subject on which the data processing method runs will check whether the newly generated query already exists in the list to obtain the conversation history text.

[0103] In this optional implementation, the repeated query text is repeated text that has appeared in one or more rounds of conversations generated in a historical time period. The project description text and the conversation history text are matched for similarity (such as cosine similarity, Jaccard similarity, etc.). If the similarity between the project description text and the conversation history text is greater than a similarity threshold, it is determined that the project description text is a repeated query text; if the similarity between the project description text and the conversation history text is less than or equal to the similarity threshold, it is determined that the project description text is not a repeated query text.

[0104] Optionally, the execution subject on which the data processing method runs can convert the conversation history text and the item description text into vector representations (such as using a pre-trained language model such as BERT to generate an embedding vector for the query), and then calculate the similarity between them. If the similarity between the conversation history text and the item description text is higher than the similarity threshold, it is considered to be a repeated query text.

[0105] Optionally, the execution entity on which the data processing method runs can also set some simple rules, such as checking whether the project description text includes specific keywords or phrases in the conversation history text. If these keywords or phrases have appeared in the historical query, it is considered to be a repeated query text.

[0106] In this optional implementation, the query prompt word is used to prompt the pre-trained large language model to generate a natural language query statement. The large language model pre-trained with the query prompt word can generate a natural language query statement.

[0107] In this optional implementation, the execution subject on which the data processing method runs can directly send a natural language query statement to the user's terminal and obtain user feedback information from the user's terminal. The natural language query statement can be sent to the user terminal in a variety of forms, such as text messages or emails.

[0108] The method for obtaining and sending natural language query statements provided in this optional implementation method obtains conversation history text; based on the conversation history text, detects whether the project description text is a repeated query text; in response to detecting that the project description text is not a repeated query text, inputs the project description text and the query prompt word into a pre-trained large language model to obtain a natural language query statement output by the large language model; sends the natural language query statement, thereby improving the accuracy of project description text detection, and can prevent the pre-trained large language model from repeatedly generating natural language query statements; uses the pre-trained large language model to process the project description text, thereby improving the accuracy and reliability of obtaining the natural language query statement.

[0109] In some optional implementations of the present disclosure, the above-mentioned extraction of the user's preference value for the project description text based on the feedback information includes: determining the preference description text corresponding to the feedback information based on a pre-constructed preference description text set; and obtaining the preference value based on the preference description text and the project description text.

[0110] In this optional implementation, the preference description text is a pre-set text related to the user's preference, and the preference description text set is a text set of the user's preference setting. The above-mentioned determination of the preference description text corresponding to the feedback information based on the pre-built preference description text set includes: in response to the feedback information being characterized as a positive reply, the corresponding deviation description text is selected from the preference description text set; in response to the feedback information being characterized as a negative reply, the preference description text selected from the preference description text set is multiplied by a preset coefficient. The preference degree value indicates the degree to which the item description text contains the preference description text or preference, and is usually a probability value (such as the probability of implied).

[0111] In this optional implementation, since the large language model is guided to generate a natural language query sentence q of the “yes” type t , asking the user whether they like a particular aspect a t , the user's feedback information will be "yes" or "no". The information content of "yes" or "no" is relatively simple, based on this, the user's preference description text ρ is constructed t : If the user answers “yes”, then ρ t Directly equal to a t ; If the user answers “no”, then ρ t for "not" and a t For example, if the query asks whether the user prefers the "children's books" element in the project, when the user answers "yes", the preference description text ρ t If the answer is “no”, it is described as “not a children’s book”. This processing method generates concise and universal preference description text, which is suitable for natural language inference models.

[0112] In this optional implementation, obtaining the preference degree value based on the preference description text and the project description text includes: inputting the preference description text and the project description text into a pre-trained natural language inference model to obtain the preference degree value output by the natural language inference model.

[0113] In this optional implementation, the pre-trained natural language inference model is responsible for analyzing the relationship between preferences obtained through conversations and product descriptions, and inferring the user's preference for a specific product. To achieve this function, the natural language inference model needs to have strong natural language understanding capabilities and be able to determine whether a product description contains a user's preference description.

[0114] Typically, a natural language inference model is a model that is further fine-tuned based on a pre-trained language model (such as BERT, GPT, etc.). The pre-training phase enables the model to learn rich semantic and grammatical knowledge, while the fine-tuning phase adapts it to specific task requirements, such as judging the implicit relationship between product descriptions and user preferences. When the natural language inference model is a BERT model, the BERT model usually accepts two pieces of text (preference description text and item description text) as input, and the two pieces of text are separated by special tags (such as [CLS] and [SEP]). In the natural language inference task, these two pieces of text are the premise and hypothesis, respectively. The BERT model will output three probability values ​​in the natural language inference task, corresponding to the three categories of "contradiction", "neutral" and "entailment", among which the probability value corresponding to the entailment is the preference degree value.

[0115] In this optional implementation, the natural language inference model can be a pre-trained model or a model fine-tuned according to a specific task. In either case, the natural language inference model needs to be trained before use to ensure that it can accurately perform natural language inference.

[0116] The method for extracting the user's preference value for the project description text provided by this optional implementation method determines the preference description text corresponding to the feedback information based on a pre-constructed preference description text set; obtains the preference value based on the preference description text and the project description text, and generates the preference value through the intermediate preference description text, thereby improving the reliability of obtaining the preference value.

[0117] In some optional implementations of the present disclosure, obtaining the preference degree value based on the preference description text and the project description text includes: inputting the preference description text and the project description text into a pre-trained natural language reasoning model to obtain the preference degree value output by the natural language reasoning model.

[0118] In this optional implementation, the preference value is a probability value representing a rating vector, which is a vector of probabilities that the item description text belongs to a preference. Natural language inference technology is used to predict the item description text x i Is there an implicit preference ρ t , in order to obtain the preference value As expected, the natural language inference model can determine that the project description text of "Narnia" contains a preference for "children", while "Les Miserables" does not contain this preference, and thus infer that the user who expresses this preference will prefer project i rather than project j. ω (x i , ρt ) Predict x i Contains ρ t Probability And return the rounded score

[0119] The method for obtaining the preference degree value provided by this optional implementation inputs the preference description text and the project description text into a pre-trained natural language inference model to obtain the preference degree value output by the natural language inference model, thereby improving the accuracy and reliability of obtaining the preference degree value.

[0120] Optionally, the above-mentioned obtaining the preference degree value based on the preference description text and the project description text includes: concatenating the preference description text and the project description text to obtain a concatenated text; inputting the concatenated text into a pre-trained language representation model (such as BERT) to obtain a preference degree value of whether the project description conforms to the user's preference description predicted by the language representation model. The language representation model can effectively process text, and performs well in processing text similarity and semantic matching tasks, and can be directly used to predict project scores without explicitly judging the implied relationship.

[0121] Optionally, the above-mentioned method of obtaining the preference degree value based on the preference description text and the project description text includes: judging whether the project description text conforms to the preference description text by predefined rules or keywords. Specifically, detecting whether the preference description text includes preset keywords, in response to detecting that the preference description text includes the preset keywords, detecting whether the project description text includes the preset keywords, in response to detecting that the project description text includes the preset keywords, determining that the project description text conforms to the preference description text, and determining the preference degree value based on the number of preset keywords included in the project description text and the preference description text. The method of obtaining the preference degree value is simple to implement, easy to understand and debug.

[0122] In some optional implementations of the present disclosure, the above-mentioned extraction of the user's preference value for the project description text based on the feedback information includes: performing sentiment classification on the feedback information to obtain a sentiment classification result; performing sentiment classification on the project description text to obtain a project classification result; and extracting the user's preference value for the project description text based on the sentiment classification result and the project classification result.

[0123] In this optional implementation, the sentiment classification result can be a result represented by different sentiment values, specifically, the sentiment value is the category of sentiment classification and the corresponding numerical representation. For example, the sentiment classification can be "positive", "neutral", and "negative", corresponding to the values ​​1, 0.5, and 0, respectively. The above-mentioned sentiment classification of feedback information includes: inputting the feedback information into a pre-trained sentiment classification model to obtain the sentiment classification result output by the sentiment classification model. Assuming that the user's answer to the query is "yes", after being processed by the sentiment classification model, the sentiment classification result is "positive" (value 1).

[0124] In this optional implementation, the project classification result can be represented by the above-mentioned sentiment values, which reflect the intensity of the sentiment. The above-mentioned sentiment classification of the project description text to obtain the project classification result includes: inputting the project description text into the above-mentioned trained sentiment classification model to obtain the project classification result output by the sentiment classification model. For example, the project description text is "This is a very suitable book for children to read", which may be classified as "positive" (value 1) through the sentiment classification model.

[0125] In this optional implementation, the above-mentioned extraction of the user's preference value for the project description text based on the sentiment classification results and the project classification results includes: calculating the initial degree value based on the sentiment classification results and the project classification results; and obtaining the preference degree value based on the initial degree value.

[0126] In this optional implementation, the initial degree value is initial data reflecting the preference, and the initial degree value can be determined in any of the following three ways:

[0127] 1. Initial degree value = user feedback sentiment value × item description sentiment value, for example, preference degree value = 1 (positive) × 1 (positive) = 1.

[0128] 2. Initial degree value = (user feedback sentiment value × weight 1 + item description sentiment value × weight 2) / (weight 1 + weight 2), where weight 1 and weight 2 can be adjusted according to actual needs, for example, weight 1 = 0.6, weight 2 = 0.4. For example, preference degree value = (1 × 0.6 + 1 × 0.4) / (0.6 + 0.4) = 1.

[0129] 3. Initial degree value = |user feedback sentiment value - item description sentiment value| For example, preference degree value = |1-1| = 0. Here, the smaller the value, the higher the preference degree.

[0130] In this optional implementation, the above-mentioned obtaining of the preference value based on the initial degree value includes: comparing the initial degree value with different level thresholds to determine the preference value; in order to further simplify the representation of the preference value, a threshold (0.5) can be set to binarize the initial degree value, for example: if the initial degree value ≥ 0.5, the preference value = 1 (indicating preference; otherwise, the preference value = 0 (indicating non-preference).

[0131] The method for obtaining the preference degree value provided by this optional implementation method performs sentiment classification on the feedback information to obtain the sentiment classification result; performs sentiment classification on the project description text to obtain the project classification result; based on the sentiment classification result and the project classification result, the user's preference degree value for the project description text is extracted, thereby improving the reliability of obtaining the preference degree value.

[0132] Optionally, obtaining the preference degree value based on the preference description text and the project description text includes: using NLI, BERT and the above sentiment classification method to obtain multiple initial degree values; calculating the average value of the multiple initial degree values, or using a voting mechanism to vote on the multiple initial degree values ​​to obtain the user's preference degree value for the project description text. The method for obtaining the preference degree value disclosed in the present invention can comprehensively utilize the advantages of multiple technologies to improve the accuracy and robustness of inference.

[0133] Further references Figure 4 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data processing device, which is similar to Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0134] like Figure 4 As shown, the data processing device 400 provided in this embodiment includes: a selection unit 401, a obtaining unit 402, a query unit 403, and an extraction unit 404. Among them, the above-mentioned selection unit 401 can be configured to select project description text from a description text set based on the current probability belief state of the Bayesian project distribution. The above-mentioned obtaining unit 402 can be configured to obtain and send a natural language query statement based on the project description text and a pre-trained large language model. The above-mentioned query unit 403 can be configured to receive user feedback information on the natural language query statement. The above-mentioned extraction unit 404 can be configured to extract the user's preference value for the project description text based on the feedback information.

[0135] In this embodiment, the specific processing of the selection unit 401, the obtaining unit 402, the query unit 403, and the extraction unit 404 and the technical effects thereof can be referred to in the respective Figure 1The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0136] In some embodiments of the present disclosure, the above-mentioned device also includes: an updating unit (not shown in the figure), and the above-mentioned updating unit is configured to: update the current probability belief state based on the preference degree value to obtain an updated probability belief state; based on the updated probability belief state and the current probability belief state, detect whether the preference extraction stop condition is met; in response to detecting that the preference extraction stop condition is not met, use the updated concept belief state as the current concept belief state.

[0137] In some embodiments of the present disclosure, the above-mentioned updating unit is further configured to: determine the beta distribution parameters of each item in the current probabilistic belief state; based on the preference degree value, update the beta distribution parameters to obtain updated beta distribution parameters; based on the updated beta distribution parameters, obtain an updated probabilistic belief state.

[0138] In some embodiments of the present disclosure, the selection unit 401 is further configured to: determine a context acquisition function from any one of the Thompson sampling algorithm, the upper confidence limit algorithm, the entropy reduction strategy, and the greedy strategy; and control the context acquisition function to select project description text from a description text set under the current probability belief state of the Bayesian project distribution.

[0139] In some embodiments of the present disclosure, the obtaining unit 402 is further configured to: obtain the conversation history text; based on the conversation history text, detect whether the project description text is a repeated query text; in response to detecting that the project description text is not a repeated query text, input the project description text and the query prompt word into a pre-trained large language model to obtain a natural language query statement output by the large language model; and send the natural language query statement.

[0140] In some embodiments of the present disclosure, the extraction unit 404 is further configured to: determine the preference description text corresponding to the feedback information based on a pre-constructed preference description text set; and obtain a preference degree value based on the preference description text and the item description text.

[0141] In some embodiments of the present disclosure, the extraction unit 404 is further configured to: input the preference description text and the item description text into a pre-trained natural language inference model to obtain a preference degree value output by the natural language inference model.

[0142] In some embodiments of the present disclosure, the extraction unit 404 is further configured to: perform sentiment classification on the feedback information to obtain a sentiment classification result; perform sentiment classification on the project description text to obtain a project classification result; and extract the user's preference value for the project description text based on the sentiment classification result and the project classification result.

[0143] The data processing device provided by the embodiment of the present disclosure firstly selects the project description text from the description text set based on the current probability belief state of the Bayesian project distribution by the selection unit 401; secondly, the obtaining unit 402 obtains and sends the natural language query statement based on the project description text and the pre-trained large language model; then, the query unit 403 receives the user's feedback information on the natural language query statement; finally, the extraction unit 404 extracts the user's preference value for the project description text based on the feedback information. Thus, by receiving the user's feedback information on the natural language query statement and extracting the user's preference value, the user's preference can be better reflected; combining the Bayesian optimization method with the large language model can quickly identify the user's core preference within a small number of dialogue rounds, thereby improving the accuracy and efficiency of preference extraction.

[0144] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an information recommendation system. Figure 1 The method embodiments shown correspond.

[0145] like Figure 5 As shown, the information recommendation system 500 provided in this embodiment includes: a server 501 and a user terminal 502; the user terminal 502 sends an information recommendation request to the server 501; the server 501 executes the data processing method provided in the above embodiment based on the information recommendation request, and sends a natural language query statement to the user terminal 502; the user terminal 502 is used to receive the natural language query statement sent by the server, and feed back the user's feedback information on the natural language query statement to the server 501; the server 501 extracts the user's preference value for the item description text in the description text set based on the feedback information, generates an information recommendation result based on the preference value, and sends the information recommendation result to the user terminal 502.

[0146] In this embodiment, the information recommendation request is a request for recommended information sent by the user to the server through the user terminal, and the information recommendation result is a result generated by the server based on the user's information recommendation request and the user's preference value for the item description text. The above feedback information is the reply content to the natural language query statement, and the natural language query statement can be a query statement that asks the user once. Specifically, the server sends this query statement to the user terminal based on the information recommendation request, determines the user's preference value based on the user's feedback information on this query statement, generates information recommendation results based on the preference value, and sends the information recommendation result to the user terminal. Optionally, since the user's preference value cannot be directly determined by asking once, the above natural language query statement can also be the sum of query statements generated multiple times, and the above feedback information is the reply content of multiple query statements. Specifically, the server sends multiple query statements to the user terminal based on the information recommendation request, and determines the user's preference value based on the reply content of multiple query statements.

[0147] Specifically, in a specific example, after receiving the user's information recommendation request, the server 501 can also select a project description text from a description text set based on the current probabilistic belief state of the Bayesian project distribution; obtain and send the current natural language query statement to the user terminal 502 based on the project description text and a pre-trained large language model; receive the user's feedback information on the current natural language query statement through the user terminal 502; based on the current feedback information, extract the user's preference value for the project description text, generate an information recommendation result based on the preference value, and send the information recommendation result to the user terminal.

[0148] In an actual example, the user's query and focus is to find children's books that can be purchased. To this end, the user sends an information recommendation request for recommended books to the server 501 through the user terminal 502. The server 501 generates a natural language query sentence "Are you interested in books for children to read?" based on the information recommendation request. The user feedbacks "Yes" through the user terminal 502. The server 501 determines the user's preference value for children's books based on the feedback information, and based on the preference value, determines and feeds back information recommendation results of multiple children's books to the user terminal 502.

[0149] In this embodiment, in the information recommendation system 500, the specific processing of the server 501 and the technical effects thereof can be referred to in Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0150] In the information recommendation system provided in this embodiment, the user terminal 502 sends an information recommendation request to the server 501; based on the information recommendation request, the server 501 executes the data processing method provided in the above embodiment and sends a natural language query statement to the user terminal 502; the user terminal 502 is used to receive the natural language query statement sent by the server and feed back the user's feedback information on the natural language query statement to the server 501; based on the feedback information, the server 501 extracts the user's preference value for the item description text in the description text set, generates an information recommendation result based on the preference value, and sends the information recommendation result to the user terminal 502. In this way, the information recommendation result can be generated based on the user's preference value, which can effectively take into account the user's personalized needs, improve the accuracy and reliability of the information recommendation result, and improve the user experience.

[0151] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0152] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0153] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0154] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0155] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0156] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).

[0157] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, causes the modes / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0159] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0161] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0162] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0163] The foregoing description of specific exemplary embodiments of the present disclosure is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present disclosure to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present disclosure and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present disclosure and various different selections and changes. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

Claims

1. A data processing method, the method comprising: Selecting an item description text from the description text set based on the current probability belief state of the Bayesian item distribution; Based on the project description text and the pre-trained large language model, obtain and send a natural language query statement; Receiving user feedback information on the natural language query statement; Based on the feedback information, the user's preference value for the item description text is extracted.

2. The method according to claim 1, further comprising: Based on the preference degree value, updating the current probability belief state to obtain an updated probability belief state; Based on the updated probability belief state and the current probability belief state, detect whether the preference extraction stop condition is met; In response to detecting that the preference extraction stop condition is not satisfied, the updated concept belief state is used as the current concept belief state.

3. The method according to claim 2, wherein: The updating of the current probability belief state based on the preference degree value to obtain the updated probability belief state includes: determining beta distribution parameters for each item in the current probabilistic belief state; Based on the preference degree value, updating the beta distribution parameter to obtain an updated beta distribution parameter; Based on the updated beta distribution parameters, an updated probability belief state is obtained.

4. The method according to any one of claims 1 to 3, wherein: The selecting of the project description text from the description text set based on the current probability belief state of the Bayesian project distribution includes: Determine a context acquisition function from any one of a Thompson sampling algorithm, an upper confidence bound algorithm, an entropy reduction strategy, and a greedy strategy; The context acquisition function is controlled to select item description text from the description text set under the current probability belief state of the Bayesian item distribution.

5. The method according to any one of claims 1 to 3, wherein: The obtaining and sending of a natural language query statement based on the project description text and the pre-trained large language model includes: Get the conversation history text; Based on the conversation history text, detecting whether the item description text is a repeated query text; In response to detecting that the project description text is not a repeated query text, inputting the project description text and the query prompt word into a pre-trained large language model to obtain a natural language query sentence output by the large language model; Send the natural language query statement.

6. The method according to any one of claims 1 to 3, wherein: The extracting the user's preference value for the item description text based on the feedback information includes: Determining the preference description text corresponding to the feedback information based on a pre-built preference description text set; The preference degree value is obtained based on the preference description text and the item description text.

7. The method according to claim 6, wherein: The obtaining of the preference degree value based on the preference description text and the item description text comprises: The preference description text and the item description text are input into a pre-trained natural language inference model to obtain a preference degree value output by the natural language inference model.

8. The method according to any one of claims 1 to 3, wherein: The extracting the user's preference value for the item description text based on the feedback information includes: Performing sentiment classification on the feedback information to obtain a sentiment classification result; Performing sentiment classification on the project description text to obtain a project classification result; Based on the sentiment classification result and the item classification result, the user's preference value for the item description text is extracted.

9. A data processing device, comprising: A selection unit is configured to select an item description text from a description text set based on a current probability belief state of a Bayesian item distribution; An obtaining unit is configured to obtain and send a natural language query sentence based on the project description text and a pre-trained large language model; A query unit, configured to receive user feedback information on the natural language query statement; The extraction unit is configured to extract the user's preference value for the item description text based on the feedback information.

10. An information recommendation system, comprising: Servers and user terminals; The user terminal sends an information recommendation request to the server; The server executes the data processing method according to any one of claims 1 to 8 above based on the information recommendation request, and sends a natural language query statement to the user terminal; The user terminal is used to receive the natural language query statement sent by the server, and feed back user feedback information on the natural language query statement to the server; The server extracts the user's preference value for the item description text in the description text set based on the feedback information, generates an information recommendation result based on the preference value, and sends the information recommendation result to the user terminal.

11. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Explanatable recommendation system, method and equipment based on large language model and medium

    CN119046529A

  • Question and answer information processing method and device, model training method and device, electronic equipment and medium

    CN119168093A

  • Bayesian modeling for risk assessment based on integrating information from dynamic data sources

    US20230153662A1

Cited By

  • Sound effect enhancement processing method and system and storage medium

    CN121438852A

  • Scientific and technological achievement intelligent pushing and public service question and answer method, device and equipment

    CN121882243A