Information recommendation method and device based on closed loop and storage medium
By deploying information recommendation and discrimination models in edge devices, and combining user personal and physical examination information, the model parameters are updated to avoid conflicts. This solves the shortcomings of existing technologies that cannot provide personalized recommendations and ignore health issues, and enables healthy meal recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-03-31
AI Technical Summary
Existing food service systems cannot deeply integrate information recommendations with users' personal circumstances and ignore the impact on users' health, which may lead to the recommendation of food that conflicts with their health conditions.
Pre-trained information recommendation and discrimination models are deployed on the user's edge devices. By acquiring the user's personal and physical examination information, it is determined whether the recommended information conflicts with the physical examination information. The model parameters are then updated through reinforcement learning to ensure that the recommended information is consistent with the user's health status.
It enables personalized meal recommendations, avoids conflicts with users' health conditions, improves the healthiness of the recommendations, and aligns with users' physical conditions.
Smart Images

Figure CN121765142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent recommendation technology for catering, and in particular to a closed-loop information recommendation method, device and storage medium. Background Technology
[0002] Currently, in canteens or other dining scenarios, food service providers can recommend dishes to users, allowing them to make informed decisions when ordering.
[0003] In existing technologies, food service providers can use recommendation algorithms to recommend meals to users. For example, they can combine historical ordering behavior with recommendation algorithms to make recommendations. This approach typically involves training a unified information recommendation model for a large number of users, which then outputs corresponding recipes to be recommended to those users.
[0004] This approach has the following problems: 1. The information recommendation model is trained uniformly, and it is impossible to make information recommendations in a deep way by combining the user's personal situation.
[0005] 2. Existing recommendation methods neglect the impact on users' health. Users may continue to eat according to dietary habits that conflict with their health conditions.
[0006] There are currently no effective solutions to the technical problems mentioned above, namely, the inability of existing technologies to recommend information based on users' personal circumstances and the neglect of the impact on users' physical health. Summary of the Invention
[0007] The embodiments of this disclosure provide a closed-loop information recommendation method, apparatus, and storage medium to at least solve the technical problems in the prior art that it cannot combine the user's personal situation for information recommendation and ignores the impact on the user's physical health.
[0008] According to one aspect of the present disclosure, an information recommendation method is provided. The method is used in a user-held edge device, where an information recommendation model and a discrimination model pre-trained by a server device are deployed. The method includes: obtaining the user's personal information and obtaining the user's physical examination information, wherein the personal information includes first health information for a first preset time period; determining first recommendation information, which is recipe information recommended to the user, based on the information recommendation model deployed in the user's edge device and the personal information, wherein the first recommendation information is recipe information recommended to the user, and the discrimination model is used to determine whether there is a conflict between the physical examination information and the first recommendation information; determining whether to recommend the first recommendation information to the user based on the first recommendation information and the physical examination information, based on the discrimination model; if the first recommendation information is not recommended to the user, determining first reward information, and updating the parameters of at least a portion of the network structure in the information recommendation model through reinforcement learning based on the first reward information; and if the first recommendation information is recommended to the user and the user performs business based on the first recommendation information, obtaining the user's updated health information, determining second reward information based on the updated health information, and updating the parameters of at least a portion of the network structure in the information recommendation model through reinforcement learning based on the second reward information, wherein the updated information recommendation model is used to continue to recommend information to the user.
[0009] According to another aspect of the present disclosure, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is executed, a processor performs any of the methods described above.
[0010] According to another aspect of the present disclosure, an information recommendation device is also provided, comprising: an acquisition module, configured to acquire a user's personal information and acquire the user's physical examination information, wherein the personal information includes first health information for a first preset time period; a recommendation information determination module, configured to determine first recommendation information based on the personal information and an information recommendation model deployed on the user's edge device, wherein the first recommendation information is recipe information recommended to the user, and the information recommendation model and the discrimination model are pre-trained by a server device; and a discrimination module, configured to determine, based on the first recommendation information and the physical examination information and the discrimination model, whether to recommend the first recommendation information to the user, wherein the discrimination model is used to determine whether there is a correlation between the physical examination information and the first recommendation information. Whether there is a conflict; the first update module is used to determine the first reward information when the first recommendation information is not recommended to the user, and update the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the first reward information; the second update module is used to obtain the user's updated health information when the first recommendation information is recommended to the user and the user performs business based on the first recommendation information, and determine the second reward information based on the updated health information, and update the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the second reward information, wherein the updated information recommendation model is used to continue to recommend information to the user.
[0011] According to another aspect of the present disclosure, an information recommendation device is also provided, comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions for processing the following steps: acquiring a user's personal information and acquiring the user's physical examination information, the personal information including first health information for a first preset time period; determining first recommendation information based on the personal information and an information recommendation model deployed on the user's edge device, the first recommendation information being recipe information recommended to the user, the information recommendation model and the discrimination model being pre-trained by a server device; and determining whether to recommend the first recommendation information to the user based on the first recommendation information and the physical examination information, using the discrimination model, the discrimination model being used for... The system determines whether there is a conflict between the physical examination information and the first recommendation information; if the first recommendation information is not recommended to the user, it determines the first reward information and updates the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the first reward information; if the first recommendation information is recommended to the user and the user performs business operations based on the first recommendation information, it obtains the user's updated health information and determines the second reward information based on the updated health information; based on the second reward information, it updates the parameters of at least part of the network structure in the information recommendation model through reinforcement learning, wherein the updated information recommendation model is used to continue to recommend information to the user.
[0012] In this embodiment, an initial information recommendation model and a discriminant model are first trained using training samples from each user. Then, the information recommendation model and the discriminant model are deployed on each user's edge device. The user's edge device determines meal recommendations using the information recommendation model. On one hand, the output of the discriminant model is used to evaluate whether the recommended information output by the information recommendation model conflicts with the health examination information, and the parameters of at least a portion of the network structure of the information recommendation model are updated based on the output of the discriminant model. On the other hand, if the recommended information is recommended to the user based on the discriminant model, the user's updated health status after eating the meal recommended by the information recommendation model (updated health information) is obtained, and the parameters of at least a portion of the network structure in the information recommendation model on the edge device are updated based on the updated health status. Therefore, this method not only updates the information recommendation model in a personalized way based on the user's real-time health status, making the information recommendation model recommend meals with the aim of making the user's personal health status more favorable, but also updates the information recommendation model with the aim of not conflicting with the user's physical examination information, so as to minimize the conflict between the information recommendation model output and the user's physical examination information (for example, it will not recommend high-sugar diets to patients with high blood sugar). Thus, compared with the existing technology, it can make meal recommendations based on the user's personal situation and health status. Attached Figure Description
[0013] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings: Figure 1 This is a hardware structure block diagram of a computing device for implementing the method described in Embodiment 1 of this disclosure; Figure 2 This is a flowchart illustrating the information recommendation method according to the first aspect of Embodiment 1 of this disclosure; Figure 3 This is a schematic diagram illustrating a process for updating the information recommendation model in the edge devices of each user, as provided in Embodiment 1 of this disclosure. Figure 4 This is a schematic diagram illustrating the relationship between an information recommendation model and a discrimination model provided in Embodiment 1 of this disclosure; Figure 5 This is a schematic diagram of a process for evaluating the value information output by a network, provided in Embodiment 1 of this disclosure. Figure 6 This is a schematic diagram of an information recommendation device according to the first aspect of Embodiment 2 of this disclosure; and Figure 7This is a schematic diagram of an information recommendation device according to the first aspect of Embodiment 3 of this disclosure. Detailed Implementation
[0014] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this disclosure.
[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0016] Example 1 According to this embodiment, an information recommendation method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0017] The method embodiments provided in this example can be executed in a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computing device for implementing an information recommendation method is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory for storing data, a transmission device for communication functions, and an input / output interface. The memory, transmission device, and input / output interface are connected to the processor via a bus. In addition, it may also include a display, keyboard, and cursor control device connected to the input / output interface. Those skilled in the art will understand that... Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0018] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element in a computing device. As involved in the embodiments of this disclosure, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).
[0019] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the information recommendation method in this embodiment of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the information recommendation method of the aforementioned application. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0020] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0021] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.
[0022] It should be noted here that, in some optional embodiments, the above... Figure 1The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.
[0023] According to a first aspect of this embodiment, an information recommendation method is provided. This method can be implemented by a user's edge device, and some technical features of the method can be implemented by a server device. The hardware architecture of both the edge device and the server device can refer to the hardware architecture of the computing devices mentioned above. The edge device is deployed with an information recommendation model pre-trained by the server device. Figure 2 A flowchart illustrating the method is shown below. (Refer to...) Figure 2 As shown, the method includes: S202: Obtaining the user's personal information and physical examination information, the personal information including the first health information for a first preset time period; S204: Based on personal information and an information recommendation model deployed on the user's edge device, determine the first recommendation information, which is recipe information recommended to the user; S206: Based on the first recommendation information and the physical examination information, and using a discriminant model, determine whether to recommend the first recommendation information to the user. The discriminant model is used to determine whether there is a conflict between the physical examination information and the first recommendation information. S208: If the first recommendation information is not recommended to the user, determine the first reward information, and update the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the first reward information; and S210: When the first recommendation information is recommended to the user and the user performs business operations based on the first recommendation information, the updated health information of the user is obtained, and the second reward information is determined based on the updated health information. Based on the second reward information, the parameters of at least a portion of the network structure in the information recommendation model are updated through reinforcement learning. The updated information recommendation model is used to continue to recommend information to the user.
[0024] refer to Figure 3As shown, the server device can pre-train an initial information recommendation model and a discrimination model. The initial information recommendation model has a certain recommendation capability and can be directly used to recommend meal information to users. The discrimination model is used to determine whether the recommended information determined by the information recommendation model conflicts with the user's physical examination information. After the initial information recommendation model is trained by the server device, it can be deployed on each user's edge device. This specification does not limit the form of the edge device; for example, the edge device can be a mobile phone, tablet computer, smart wearable device, etc.
[0025] The edge device can obtain the user's personal information and physical examination information. The personal information includes first-level health information within a first preset time period (S202), and the physical examination information refers to the user's most recent physical examination. The physical examination information is in text format, and may be, for example, "thyroid nodules; sinus bradycardia with arrhythmia" or "fatty liver (moderate); dyslipidemia (high total cholesterol and low-density lipoprotein cholesterol)." Then, the edge device can determine the first recommended information based on personal information and an information recommendation model (S204). The first recommended information is meal information recommended to the user. The edge device can then recommend the first recommended information to the user.
[0026] The specific duration of the first preset time period can be set in advance according to business needs; for example, the duration of the first preset time period can be set to 7 days, 14 days, or 21 days. In this embodiment, health information (including first health information and updated health information and second health information, which will be mentioned later) can be used to represent the user's physical health status.
[0027] Then, based on the first recommendation information and the physical examination information, the edge device determines, using a discrimination model, whether to recommend the first recommendation information to the user (S206). (See reference) Figure 4 As shown, after the edge device determines the recommended information (such as the first recommended information) through the information recommendation model, it uses a discriminant model to determine whether there is a conflict between the first recommended information and the physical examination information. If it is determined that there is no conflict between the first recommended information and the physical examination information, the edge device can directly recommend the first recommended information to the user. If it is determined that there is a conflict between the first recommended information and the physical examination information, the edge device will not recommend the first recommended information to the user.
[0028] If the first recommendation information is not recommended to the user, the edge device determines the first reward information and updates the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the first reward information (S208).
[0029] When the edge device recommends the first recommendation information to the user and the user performs business based on the first recommendation information, the edge device obtains the user's updated health information, determines the second reward information based on the updated health information, and updates the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the second reward information. The updated information recommendation model is used to continue to recommend information to the user (S210).
[0030] In other words, the parameters of the personalized network in the information recommendation model are updated using two different rewards (first reward information and second reward information) depending on whether the discriminant model determines that the first recommendation information will not be recommended to the user or that it will be recommended. The difference between the first reward information and the second reward information is that the first reward information is mainly used to penalize the information recommendation model for outputting the first recommendation information (because the first recommendation information conflicts with the health examination information), while the second reward information is determined based on changes in the user's health status after the first recommendation information has been recommended to the user, and for the information recommendation model, it can be either a penalty or a reward.
[0031] The aforementioned user's execution of services based on the first recommendation information refers to the user choosing to dine on the meal recommended in the first recommendation information. In other words, when a user chooses a meal recommended in the first recommendation information, the edge device can obtain the user's updated health information. This updated health information is used to update the information recommendation model through reinforcement learning. Specifically, the edge device determines the second reward information corresponding to the first recommendation information based on the updated health information, and updates at least some parameters of the network structure in the information recommendation model using reinforcement learning based on this second reward information. However, if the user does not choose the meal corresponding to the first recommendation information but instead dines on another meal according to their own preference, then this sample (and therefore the corresponding updated health information) will not be obtained to update the parameters of at least some network structures in the information recommendation model.
[0032] Continue to refer to Figure 3 As shown, with Figure 3 Taking user i as an example, after the edge device i held by user i deploys the discrimination model and the initial information recommendation model, the edge device i will recommend information (recommendation of meal information) to user i through its own deployed information recommendation model and discrimination model, and update the information recommendation model according to the feedback of user i's real-time health information or the output of the discrimination model.
[0033] The following example illustrates the process of updating the information recommendation model and the process of recommending information using the information recommendation model in this embodiment.
[0034] For example, suppose the first health information above was obtained when making information recommendations to the user for the sth time, and this first health information is called H. s H s Includes health information within a first preset time period (e.g., from day l to day k) [h l , ..., h k ]. h l This is the health information for day l.
[0035] Edge devices, based on an information recommendation model, determine the first health information H. s In addition to other information contained in the personal information, the corresponding primary recommendation information R is determined. s Other information may include the user's user ID, location, user-inputted dining preferences, and historical dining records. Of course, edge devices can also input the user's health checkup information into the recommendation model to determine the appropriate first recommendation, R. s Then, the edge device determines whether to use the first recommendation information R based on the discrimination model. s Recommended to users, and then determined by a discriminative model not to use the first recommendation information R s When recommended to users, determine the first reward information r. s_1 And based on the first reward information r s_1 The parameters of at least a portion of the network structure of the information recommendation model are updated using reinforcement learning, wherein the first reward information r s_1 It is a negative value.
[0036] After determining the first recommendation information R through the discriminant model s When recommended to users, the edge device will use the first recommendation information R s The information is displayed to the user, and it can then be determined whether the user dined according to the first recommended information. If the user did so, updated health information can be obtained. The updated health information corresponds to a time period following the first preset time period. For example, the updated health information H... up It can be [h] k+1 , ..., h k+t (For example, health information from day k+1 to day k+t), the specific value of t can be preset, for example, the shortest health information after the update can be the user's health information within h days. k+1 The day's health information. Then, the edge device can update the health information based on the current day's data. up Determine the corresponding second reward information r s_2 And according to the second reward information r s_2The parameters of at least a portion of the network structure in the information recommendation model are updated. The determination of the second reward information and the method for updating the parameters of at least a portion of the network structure in the information recommendation model based on the second reward information will be explained in detail later.
[0037] As described in the background section, existing technologies have the following problems: 1. Information recommendation models are uniformly trained, making it impossible to deeply integrate user's personal circumstances for information recommendations. 2. Existing recommendation methods neglect the impact on user's health. Users may continue to eat according to dietary habits that conflict with their health conditions.
[0038] Therefore, in this method, an initial information recommendation model and a discriminant model are first trained uniformly using training samples from each user. Then, the information recommendation model and the discriminant model are deployed on each user's edge device. The user's edge device determines the recommended information for meal recommendations using the information recommendation model. On one hand, the output of the discriminant model is used to evaluate whether there is a conflict between the recommended information output by the information recommendation model and the health examination information, and the parameters of at least a portion of the network structure of the information recommendation model are updated based on the output of the discriminant model. On the other hand, if the recommended information is recommended to the user based on the discriminant model, the user's updated health status after eating the meal recommended by the information recommendation model (updated health information) is obtained, and the parameters of at least a portion of the network structure of the information recommendation model in the edge device are updated based on the updated health status. Therefore, this method not only updates the information recommendation model in a personalized way based on the user's real-time health status, making the information recommendation model recommend meals with the aim of making the user's personal health status more favorable, but also updates the information recommendation model with the aim of not conflicting with the user's physical examination information, so as to minimize the conflict between the information recommendation model output and the user's physical examination information (for example, it will not recommend high-sugar diets to patients with high blood sugar). Thus, compared with the existing technology, it can make meal recommendations based on the user's personal situation and health status.
[0039] Optionally, the first health information includes at least one of the following: the user's emotional information, heart rate information, blood pressure information, blood sugar information, height information, and weight information. Personal information also includes the user's user ID, age, gender, and location.
[0040] In other words, the dimensions of health information (including but not limited to first health information, second health information, and updated health information) in this embodiment can include emotional information, heart rate information, blood pressure information, blood sugar information, height information, and weight information. Information other than health information in personal information, such as user ID, age, gender, and location, can be obtained by the user themselves. However, most health information needs to be updated in real time. The edge device can obtain health information each time it acquires it, either through user input or through real-time measurement. For example, the edge device can identify the user's facial image itself or through a server device to determine the user's current emotional information, and can also obtain the user's heart rate information through a smart wearable device. As another example, this embodiment can recommend information to users in a cafeteria setting. This allows for the setup of measurement devices for heart rate, blood pressure, blood sugar, weight, and height, and the edge device can obtain the measured information through the Internet of Things (IoT). For adults, height usually does not change, so when the user is an adult, height information only needs to be obtained once initially. For minors, edge devices can obtain height information once every certain period, instead of obtaining height information every time health information is confirmed.
[0041] Optionally, the information recommendation model includes a first feature extraction network, a second feature extraction network, and a personalization network; based on personal information and physical examination information, the operation of determining the first recommended information based on the information recommendation model deployed on the user's edge device includes: inputting personal information into the first feature extraction network to obtain a first semantic feature corresponding to the personal information, and inputting physical examination information into the second feature extraction network to obtain a second semantic feature corresponding to the physical examination information; inputting the first semantic feature and the second semantic feature into the personalization network to determine the first recommended information.
[0042] For details, please refer to [link / reference]. Figure 4 As shown, the information recommendation model may include a first feature extraction network, a second feature extraction network, and a personalization network. In this embodiment, the edge device mainly updates the parameters of the personalization network structure. (Reference) Figure 4 Edge devices can input personal information into a first semantic feature extraction network to obtain first semantic features corresponding to the personal information, and input physical examination information into a second semantic feature extraction network to obtain second semantic features corresponding to the physical examination information. Then, the first and second semantic features can be input into a personalization network to determine the first recommendation information.
[0043] The personalized network can be a neural network such as a feedforward neural network or a multilayer perceptron. The information recommendation model can output the probability of each preset food combination. Then, the edge device generates the corresponding recipe text based on the food combination with the highest probability, which serves as the first recommendation information. For example, the first recommendation information can be in the form of a text such as "rice, cola chicken wings, garlic broccoli, and tomato and egg soup".
[0044] It should be noted that, Figure 4 The first, second, and third feature extraction networks in the model can all use the BERT model.
[0045] Optionally, the discrimination model includes: a third feature extraction network and a discrimination network; based on the first recommendation information and the physical examination information, the operation of determining whether to recommend the first recommendation information to the user based on the discrimination model includes: determining a third semantic feature corresponding to the first recommendation information based on the third feature extraction network; inputting the first semantic feature, the second semantic feature, and the third semantic feature into the discrimination network to determine discrimination information, which is used to indicate whether there is a conflict between the physical examination information and the first recommendation information; determining whether to recommend the first recommendation information to the user based on the discrimination information; and wherein, the operation of updating the parameters of at least some network structures in the information recommendation model through reinforcement learning includes: updating the parameters of the personalization network in the information recommendation model through reinforcement learning, while freezing the parameters of other network structures in the information recommendation model other than the personalization network.
[0046] Further reference Figure 4 As shown, for the discriminative network branch, the first recommendation information is input into the third feature extraction network to obtain the third semantic feature corresponding to the first recommendation information. Then, the first semantic feature, the second semantic feature, and the third semantic feature are input into the discriminative network to obtain discriminative information, which is binary classification information (0 or 1). The discriminative information is used to indicate whether the first recommendation information conflicts with the user's health examination information. The discriminative network can be, for example, a feedforward neural network or a multilayer perceptron neural network. When the discriminative information indicates that the first recommendation information conflicts with the user's health examination information, the first recommendation information will not be recommended to the user; when the discriminative information indicates that the first recommendation information does not conflict with the user's health examination information, the first recommendation information can be recommended to the user. For example, if the health examination information indicates that the user's blood sugar is high, then if the first recommendation information contains high-sugar foods, the discriminative model is likely to determine that the first recommendation information conflicts with the corresponding user's health examination information.
[0047] Then, we can divide the process into two scenarios, each with different reward determination methods, to update the personalized network parameters in the information recommendation model. It's important to note that the parameters of other network structures in the information recommendation model, besides the personalized network, are not updated; for example, the parameters of the first and second feature extraction networks will be frozen. Scenario 1: If the first recommendation information is not recommended to the user, a first reward is determined, and the parameters of the personalized network in the information recommendation model are updated based on this first reward information. Scenario 2: If the first recommendation information is recommended to the user, a second reward is determined based on the subsequently obtained updated health information, and the parameters of the personalized network in the information recommendation model are updated based on this second reward information.
[0048] Specifically, the first reward information can be a fixed negative value, such as -50. This first reward information is given when the discriminative model determines that the first recommended information conflicts with the corresponding user's physical examination information. Therefore, assigning a negative value to the first reward information directly aims to prevent the information recommendation model from providing recommendations that conflict with the user's physical examination information.
[0049] Optionally, the operation of determining the second reward information corresponding to the first recommendation information based on the updated health information includes: Based on the updated health information, a first reward value is determined, and based on the difference between the first health information and the updated health information, a second reward value is determined; based on the first reward value and the second reward value, second reward information corresponding to the first recommendation information is determined.
[0050] In other words, the determination of the second reward information in this embodiment can include two specific calculation methods: 1. If the updated health information indicates that the user's body is at a relatively healthy level, then the specific value of the second reward information is higher. 2. If the physical condition corresponding to the updated health information is better than the physical condition corresponding to the first health information, then the specific value of the second reward information is higher.
[0051] The following is an example of calculating the second reward information.
[0052] Continuing with the example above, the edge device can adjust its health information based on the updated health information H. up Determine the second reward information r corresponding to the first recommendation information. s_2 This involves determining the first reward value r1 and the second reward value r2 corresponding to the first recommendation information. Then, the first reward value r1 and the second reward value r2 can be summed, or weighted summed, to obtain the second reward information r. s_2 .
[0053] The first reward value can be determined solely based on the updated health information, while the second reward value can be determined using both the updated and first health information. Since both the first and updated health information may contain health data from multiple days, the average or mode of the first and updated health information can be used to determine the first and second reward values, respectively.
[0054] For the first reward value r1, the updated health information H can be determined based on the updated health information. up The corresponding time period determines whether the user is in a healthy or unhealthy state. When the user is in a healthy state, r1 = +10; when the user is in an unhealthy state, r1 = -10.
[0055] For the second reward value r2, if it is determined that the physical state corresponding to the updated health information is better than the physical state corresponding to the first health information, then r2 = +10. Conversely, if it is determined that the physical state corresponding to the updated health information is not better than the physical state corresponding to the first health information, then r2 = -10.
[0056] In this embodiment, the normal range for each dimension of health information can be predetermined. For example, heart rate between 60-100 beats / minute is considered normal; diastolic blood pressure between 91 mmHg and 139 mmHg is considered normal, and systolic blood pressure between 89 mmHg and 61 mmHg is considered normal. Blood glucose (in this example, only the user's fasting blood glucose is used) between 3.9 mmol / L and 6.1 mmol / L is considered normal. For height and weight information, the user's BMI can be determined based on these data. A BMI of 18.5 kg / m² is considered normal. 2 -24kg / m 2 The values between these ranges are within the normal range. For emotional information, when the emotional information falls within a preset emotion, the emotion is considered to be within the normal range. Assuming emotional information can be categorized into several types (such as joy, calmness, anger, and sadness), the preset emotions can be set to joy and calmness.
[0057] If the edge device determines that the data in the first preset quantity dimension of the updated health information is not within the normal range, it can determine that the updated health information indicates that the user is in an unhealthy state. If it determines that the data in the first preset quantity dimension of the updated health information is within the normal range, it can determine that the health information indicates that the user is in a healthy state.
[0058] Furthermore, when determining the second reward value r2, the edge device can consider situations where the updated health information indicates the user is in a healthy state, while the first health information indicates the user is in an unhealthy or healthy state, as cases where the updated health information indicates a better physical state than the first health information. Additionally, when both the updated and first health information indicate the user is in an unhealthy state, if the data in the updated health information exceeds a second preset number of dimensions that are closer to the normal range compared to the data in the first health information, then the edge device can also determine that the updated health information indicates a better physical state than the first health information. Both the first and second preset numbers can be pre-set.
[0059] Specifically, in this embodiment, the parameters of at least a portion of the network structure in the information recommendation model can be updated using the Actor-Critic approach in reinforcement learning. To achieve this update of the parameters of at least a portion of the network structure (i.e., the personalized network) in the information recommendation model via the Actor-Critic approach, an evaluation network for determining the value information Q also needs to be deployed in the edge device. (See reference...) Figure 5 As shown, this evaluation network outputs value information corresponding to the respective recommendations based on health information, physical examination information, and recommendation information. That is, health information and physical examination information are treated as states, and recommendation information is treated as actions.
[0060] As explained above, there are two scenarios in this embodiment, and the only difference between the two scenarios is the reward.
[0061] In the first scenario, the edge device does not recommend the primary recommendation information R to the user. s And directly determined the first reward information r corresponding to the first recommendation information. s_1 In this scenario, the edge device can replace the initial recommendation with pre-set standard recommendations, such as a healthy recipe suitable for various population groups.
[0062] In the second scenario, the edge device recommends the first recommendation information R to the user. s And determine the user based on the first recommendation information R s After the meal, updated health information H was obtained. up And based on the updated health information H up The second reward information r corresponding to the first recommendation information was determined. s_2 .
[0063] Regardless of which of the above scenarios occurs, the edge device subsequently determines the second health information H. y Second Health Information H yThis could refer to the user's health information obtained when the edge device makes its information recommendation for the yth time, where y > s, for example, y = s + 1. The second health information H... y It can contain health information within a second preset time period [h] p , ..., h q (Health information from day p to day q). Here, q can be, for example, k+t, meaning the last day of the second preset time period can coincide with the last day of the time period corresponding to the updated health information, allowing the determination of the second reward information corresponding to the first recommendation information on that day. The edge device will contain the second health information H. y Personal information and physical examination information are input into the information recommendation model to obtain the second recommendation information R. y Then, the edge device can use the first recommendation information R s Personal information and physical examination information, including primary health information, are input into the evaluation network to obtain the primary recommendation information R. s The corresponding first value information Q s And the second recommendation information R y The second health information and corresponding physical examination information are input into the evaluation network to obtain the second recommendation information R. y The corresponding second value information Q y It should be noted that the second recommendation information R y This is for training purposes only and should not be recommended for users.
[0064] Then, the edge device can determine the timing difference error corresponding to the first recommendation information. The formula is shown below: (1) In formula (1) above, γ is a preset discount factor, and Q s To match the first recommended information R s The corresponding first value information, Q y To match the second recommendation information R y The corresponding second value information, r, is related to the first recommendation information R. s The corresponding reward information is the first reward information in the first case and the second reward information in the second case. Edge devices can minimize this timing difference error. The evaluation network is trained to achieve the desired result. The specific formula for parameter updates is: (2) (3) in, This indicates the evaluation network. Let α represent the parameters of the evaluation network when making information recommendations for the sth time. α is the learning rate. In formula (2), S represents the state at the corresponding time (including personal information and physical examination information). Formula (2) is derived with respect to the parameters of the evaluation network. Formula (3) represents the updated parameters obtained by updating the parameters of the evaluation network using the gradient descent algorithm. .
[0065] Finally, the edge device bases its data on personal information and medical examination information (i.e., S in the formula) and first-value information Q. s And the first recommendation information R s The parameters of the personalized network in the information recommendation model are updated using the following formula: (4) (5) in, Let θ represent the information recommendation model, and let θ represent the parameters of the information recommendation model. s Let be the parameters of the information recommendation model when making information recommendations for the sth time. β is the learning rate. Formula (4) represents the derivative of the information recommendation model Re(). Formula (5) represents the updating of the parameters of the personalized network in the information recommendation model using the gradient ascent algorithm, resulting in the updated network parameters. Therefore, by continuously repeating the above process, the goal of continuously updating the information recommendation model can be achieved.
[0066] Optionally, the method further includes: obtaining a first training sample and a second training sample through a server device, wherein the first training sample includes: personal information sample, physical examination information sample, recommendation information sample, and discrimination result sample, and the second training sample includes: personal information sample, physical examination information sample, and recommendation information sample; training a first feature extraction network, a second feature extraction network, a third feature extraction network, and a discrimination network through the server device based on the first training sample; and updating the parameters of the personalized network in the information recommendation model through the server device based on the second training sample, wherein the network parameters of the first feature extraction network and the second feature extraction network are fixed during the parameter update process of the personalized network in the information recommendation model.
[0067] It should be noted that before deploying the discrimination model and information recommendation model on edge devices, both models need to be trained uniformly by the server device. The server device can obtain a first training sample and a second training sample. The first training sample is used to train the discrimination model, enabling it to determine whether recommended information conflicts with health examination information. The second training sample is used to train the information recommendation model, giving the information recommendation model initially deployed on the edge device a certain information recommendation capability. The first and second training samples can be obtained through manual annotation. A large number of user personal information samples and health examination information samples can be obtained. Through manual annotation, suitable meal information for the population corresponding to the corresponding personal information sample (this meal information is beneficial to the health of the population corresponding to the corresponding health information sample) can be determined, thereby annotating the corresponding recommendation information samples to obtain the second training sample. Then, based on the second training sample, the corresponding discrimination result samples are annotated through manual annotation to obtain the first training sample. Subsequently, the discrimination model can be trained in a supervised manner using the first training sample. Then, the information recommendation model can be trained in a supervised manner using the second training sample. Since the discriminative model and the information recommendation model share the first feature extraction network and the second feature extraction network, the parameters of the first feature extraction network and the second feature extraction network are fixed during the supervised training of the information recommendation model using the second training samples.
[0068] In addition, refer to Figure 1 As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.
[0069] Therefore, in this embodiment, an initial information recommendation model is first trained using training samples from each user, along with a discriminant model to determine whether the recommended information conflicts with the user's health checkup information. Then, the information recommendation model and the discriminant model are deployed on each user's edge device. The user's edge device determines the recommended information for meal recommendations using the information recommendation model. On one hand, it uses the output of the discriminant model to evaluate whether the recommended information output by the information recommendation model conflicts with the health checkup information, and updates at least a portion of the network structure parameters of the information recommendation model based on the output of the discriminant model. On the other hand, if the recommended information is recommended to the user based on the discriminant model, the user's updated health status after eating the meal recommended by the information recommendation model (updated health information) is obtained, and the parameters of at least a portion of the network structure in the information recommendation model on the edge device are updated based on the updated health status. Therefore, this method not only updates the information recommendation model in a personalized way based on the user's real-time health status, making the information recommendation model recommend meals with the aim of making the user's personal health status more favorable, but also updates the information recommendation model with the aim of not conflicting with the user's physical examination information, so as to minimize the conflict between the information recommendation model output and the user's physical examination information (for example, it will not recommend high-sugar diets to patients with high blood sugar). Thus, compared with the existing technology, it can make meal recommendations based on the user's personal situation and health status.
[0070] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0071] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0072] Example 2 Figure 6An information recommendation apparatus according to a first aspect of this embodiment is shown, which corresponds to the method described according to the first aspect of Embodiment 1. (See reference...) Figure 6 As shown, the information recommendation device includes: an acquisition module 601, used to acquire the user's personal information and physical examination information, the personal information including first health information for a first preset time period; a recommendation information determination module 602, used to determine first recommendation information based on the personal information and an information recommendation model deployed on the user's edge device, the first recommendation information being recipe information recommended to the user, the information recommendation model being pre-trained by a server device; a discrimination module 603, used to determine whether to recommend the first recommendation information to the user based on the first recommendation information and physical examination information, using a discrimination model, the discrimination model being used to determine whether there is a conflict between the physical examination information and the first recommendation information; and a first update module. Module 604 is used to determine first reward information when the first recommendation information is not recommended to the user, and update the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the first reward information; the second update module 605 is used to obtain the user's updated health information when the first recommendation information is recommended to the user and the user performs business based on the first recommendation information, and determine second reward information based on the updated health information, and update the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the second reward information, wherein the updated information recommendation model is used to continue to recommend information to the user.
[0073] Optionally, the information recommendation model includes a first feature extraction network, a second feature extraction network, and a personalization network; the recommendation information determination module 602 is used to input personal information into the first feature extraction network to obtain a first semantic feature corresponding to the personal information, and input physical examination information into the second feature extraction network to obtain a second semantic feature corresponding to the physical examination information; and input the first semantic feature and the second semantic feature into the personalization network to determine the first recommendation information.
[0074] Optionally, the discrimination model includes: a third feature extraction network and a discrimination network; a discrimination module 603, used to determine a third semantic feature corresponding to the first recommendation information based on the third feature extraction network; input the first semantic feature, the second semantic feature and the third semantic feature into the discrimination network to determine discrimination information; and determine whether to recommend the first recommendation information to the user based on the discrimination information, wherein the discrimination information is used to indicate whether there is a conflict between the physical examination information and the first recommendation information; and wherein, the operation of updating the parameters of at least some network structures in the information recommendation model through reinforcement learning includes: updating the parameters of the personalization network in the information recommendation model through reinforcement learning, while freezing the parameters of other network structures in the information recommendation model other than the personalization network.
[0075] Optionally, the first health information includes at least one of the following: the user's emotional information, heart rate information, blood pressure information, blood sugar information, height information, and weight information. Personal information also includes the user's user ID, age, gender, and location.
[0076] Optionally, the second update module 605 is used to determine a first reward value based on the updated health information, and to determine a second reward value based on the difference between the first health information and the updated health information; and to determine second reward information based on the first reward value and the second reward value.
[0077] Optionally, it further includes: a training module 606, used to acquire a first training sample and a second training sample through a server device, wherein the first training sample includes: personal information sample, physical examination information sample, recommendation information sample and discrimination result sample, and the second training sample includes: personal information sample, physical examination information sample and recommendation information sample; the server device trains a first feature extraction network, a second feature extraction network, a third feature extraction network and a discrimination network based on the first training sample; the server device updates the parameters of the personalized network in the information recommendation model based on the second training sample, wherein the network parameters of the first feature extraction network and the second feature extraction network are fixed during the parameter update process of the personalized network in the information recommendation model.
[0078] In this embodiment, an initial information recommendation model and a discriminant model are first trained using training samples from each user. Then, the information recommendation model and the discriminant model are deployed on each user's edge device. The user's edge device determines the recommended information for meal recommendations using the information recommendation model. On one hand, the output of the discriminant model is used to evaluate whether there is a conflict between the recommended information and the health examination information, and the parameters of at least a portion of the network structure of the information recommendation model are updated based on the output of the discriminant model. On the other hand, if the recommended information is recommended to the user based on the discriminant model, the user's updated health status after eating the recommended meal (updated health information) is obtained, and the parameters of at least a portion of the network structure of the information recommendation model in the edge device are updated based on the updated health status. This technology not only updates the information recommendation model in a personalized way based on the user's real-time health status, making meal recommendations more favorable to the user's individual health condition, but also updates the model to ensure that the recommended information does not conflict with the user's medical examination results. This is to minimize the conflict between the recommendations output by the information recommendation model and the user's medical examination results (for example, it will not recommend high-sugar diets to patients with high blood sugar). Therefore, compared with existing technologies, it can make meal recommendations by combining the user's individual situation and health status.
[0079] Example 3 Figure 7 An information recommendation device according to a first aspect of this embodiment is shown, which corresponds to the method described according to a first aspect of Embodiment 1. (Reference) Figure 7As shown, the information recommendation device includes: a processor 710; and a memory 720 connected to the processor 710, used to provide the processor 710 with instructions to process the following steps: obtaining the user's personal information and obtaining the user's physical examination information, the personal information including first health information for a first preset time period; determining first recommendation information based on the personal information and an information recommendation model deployed on the user's edge device, the first recommendation information being recipe information recommended to the user, the information recommendation model and the discrimination model being pre-trained by a server device; and determining whether to recommend the first recommendation information to the user based on the first recommendation information and the physical examination information, using the discrimination model, the discrimination model being used to determine the physical examination information. The system checks whether there is a conflict between the information and the first recommendation information; if the first recommendation information is not recommended to the user, it determines the first reward information and updates the parameters of at least part of the network structure in the information recommendation model through reinforcement learning based on the first reward information; if the first recommendation information is recommended to the user and the user performs business based on the first recommendation information, it obtains the user's updated health information and determines the second reward information based on the updated health information; based on the second reward information, it updates the parameters of at least part of the network structure in the information recommendation model through reinforcement learning, wherein the updated information recommendation model is used to continue to recommend information to the user.
[0080] Optionally, the information recommendation model includes a first feature extraction network, a second feature extraction network, and a personalization network; based on personal information, the operation of determining the first recommended information based on the information recommendation model deployed on the user's edge device includes: inputting personal information into the first feature extraction network to obtain a first semantic feature corresponding to the personal information, and inputting physical examination information into the second feature extraction network to obtain a second semantic feature corresponding to the physical examination information; inputting the first semantic feature and the second semantic feature into the personalization network to determine the first recommended information.
[0081] Optionally, the discrimination model includes: a third feature extraction network and a discrimination network; the operation of determining whether to recommend the first recommendation information to the user based on the first recommendation information and the physical examination information, according to the discrimination model, includes: determining the third semantic feature corresponding to the first recommendation information based on the third feature extraction network; inputting the first semantic feature, the second semantic feature, and the third semantic feature into the discrimination network to determine discrimination information; determining whether to recommend the first recommendation information to the user based on the discrimination information, wherein the discrimination information is used to indicate whether there is a conflict between the physical examination information and the first recommendation information; and wherein the operation of updating the parameters of at least some network structures in the information recommendation model through reinforcement learning includes: updating the parameters of the personalization network in the information recommendation model through reinforcement learning, while freezing the parameters of other network structures in the information recommendation model other than the personalization network.
[0082] Optionally, the first health information includes at least one of the following: the user's emotional information, heart rate information, blood pressure information, blood sugar information, height information, and weight information. Personal information also includes the user's user ID, age, gender, and location.
[0083] Optionally, the operation of determining the second reward information based on the updated health information includes: determining a first reward value based on the updated health information, and determining a second reward value based on the difference between the first health information and the updated health information; and determining the second reward information based on the first reward value and the second reward value.
[0084] Optionally, the memory 720 is further configured to provide the processor 710 with instructions to process the following steps: acquiring a first training sample and a second training sample via a server device, wherein the first training sample includes: a personal information sample, a physical examination information sample, a recommendation information sample, and a discrimination result sample, and the second training sample includes: a personal information sample, a physical examination information sample, and a recommendation information sample; training a first feature extraction network, a second feature extraction network, a third feature extraction network, and a discrimination network via the server device based on the first training sample; and training a personalized network in the information recommendation model via the server device based on the second training sample, wherein, during the training of the personalized network in the information recommendation model, the network parameters of the first feature extraction network and the second feature extraction network are fixed.
[0085] In this embodiment, an initial information recommendation model and a discriminant model are first trained using training samples from each user. Then, the information recommendation model and the discriminant model are deployed on each user's edge device. The user's edge device determines the recommended information for meal recommendations using the information recommendation model. On one hand, the output of the discriminant model is used to evaluate whether there is a conflict between the recommended information and the health examination information, and the parameters of at least a portion of the network structure of the information recommendation model are updated based on the output of the discriminant model. On the other hand, if the recommended information is recommended to the user based on the discriminant model, the user's updated health status after eating the recommended meal (updated health information) is obtained, and the parameters of at least a portion of the network structure of the information recommendation model in the edge device are updated based on the updated health status. This technology not only updates the information recommendation model in a personalized way based on the user's real-time health status, making meal recommendations more favorable to the user's individual health condition, but also updates the model to ensure that the recommended information does not conflict with the user's medical examination results. This is to minimize the conflict between the recommendations output by the information recommendation model and the user's medical examination results (for example, it will not recommend high-sugar diets to patients with high blood sugar). Therefore, compared with existing technologies, it can make meal recommendations by combining the user's individual situation and health status.
[0086] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0087] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0088] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0089] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0090] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0091] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0092] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An information recommendation method characterized by comprising: The method is used for an edge device held by a user, and a recommendation model and a discrimination model pre-trained by a server device are deployed in the edge device, and the method comprises: obtaining personal information of the user and obtaining physical examination information of the user, wherein the personal information comprises first health information of a first preset time period; determining first recommendation information according to the personal information and based on the recommendation model deployed in the edge device of the user, wherein the first recommendation information is recipe information recommended for the user; judging whether to recommend the first recommendation information to the user according to the first recommendation information and the physical examination information and based on the discrimination model, wherein the discrimination model is used to judge whether there is a conflict between the physical examination information and the first recommendation information; in a case where the first recommendation information is not recommended to the user, determining first reward information and updating parameters of at least part of network structures in the recommendation model through reinforcement learning according to the first reward information; and in a case where the first recommendation information is recommended to the user and the user performs a service according to the first recommendation information, obtaining updated health information of the user and determining second reward information according to the updated health information, and updating parameters of at least part of network structures in the recommendation model through reinforcement learning according to the second reward information, wherein the updated recommendation model is used to continue to recommend information for the user.
2. The method of claim 1, wherein, The recommendation model comprises a first feature extraction network, a second feature extraction network and a personalized network; The operation of determining the first recommendation information according to the personal information and based on the recommendation model deployed in the edge device of the user comprises: inputting the personal information into the first feature extraction network to obtain first semantic features corresponding to the personal information, and inputting the physical examination information into the second feature extraction network to obtain second semantic features corresponding to the physical examination information; inputting the first semantic features and the second semantic features into the personalized network to determine the first recommendation information.
3. The method of claim 2, wherein, The discrimination model comprises a third feature extraction network and a discrimination network; The operation of judging whether to recommend the first recommendation information to the user according to the first recommendation information and the physical examination information and based on the discrimination model comprises: determining third semantic features corresponding to the first recommendation information according to the third feature extraction network; inputting the first semantic features, the second semantic features and the third semantic features into the discrimination network to determine discrimination information, wherein the discrimination information is used to indicate whether there is a conflict between the physical examination information and the first recommendation information; judging whether to recommend the first recommendation information to the user according to the discrimination information; and wherein the operation of updating parameters of at least part of network structures in the recommendation model through reinforcement learning comprises: The parameters of the personalized network in the information recommendation model are updated by means of reinforcement learning while the parameters of other network structures in the information recommendation model except the personalized network are frozen.
4. The method of claim 1, wherein, The first health information includes at least one of emotional information, heart rate information, blood pressure information, blood glucose information, height information and weight information of the user, and the personal information further includes a user ID, age, gender and region of the user.
5. The method of claim 1, wherein, The operation of determining second reward information according to the updated health information includes: The first reward value is determined according to the updated health information, and the second reward value is determined according to the difference between the first health information and the updated health information; and the second reward information is determined according to the first reward value and the second reward value.
6. The method of claim 3, wherein, Further comprising: The server device acquires first training samples and second training samples, the first training samples including personal information samples, physical examination information samples, recommendation information samples and discrimination result samples, and the second training samples including personal information samples, physical examination information samples and recommendation information samples; The server device trains the first feature extraction network, the second feature extraction network, the third feature extraction network and the discrimination network according to the first training samples; The server device updates the parameters of the personalized network in the information recommendation model according to the second training samples, wherein the network parameters of the first feature extraction network and the second feature extraction network are fixed during the updating of the parameters of the personalized network in the information recommendation model.
7. A storage medium, characterized by The storage medium includes a stored program, wherein the program is executed by a processor to perform the method of any one of claims 1-6 when the program is running.
8. An information recommendation device characterized by comprising: Comprising: An acquisition module is configured to acquire personal information of a user and acquire physical examination information of the user, the personal information including first health information of a first preset time period; A recommendation information determination module is configured to determine first recommendation information based on an information recommendation model deployed in an edge device of the user according to the personal information, the first recommendation information being recipe information recommended for the user, the information recommendation model and a discrimination model being pre-trained by a server device; A discrimination module is configured to determine whether to recommend the first recommendation information to the user based on a discrimination model according to the first recommendation information and the physical examination information; A first update module is configured to determine first reward information in a case where the first recommendation information is not recommended to the user, and update parameters of at least part of network structures in the information recommendation model by means of reinforcement learning according to the first reward information. A second updating module is configured to, in a case where the first recommendation information is recommended to the user and the user performs a service according to the first recommendation information, acquire updated health information of the user, determine second reward information according to the updated health information, and update parameters of at least part of network structures in the information recommendation model according to the second reward information in a manner of reinforcement learning, wherein the updated information recommendation model is used to continue to recommend information for the user.
9. The apparatus of claim 8, wherein, The information recommendation model comprises a first feature extraction network, a second feature extraction network, and a personalized network; the recommendation information determination module is configured to input the personal information into the first feature extraction network to obtain first semantic features corresponding to the personal information, and input the physical examination information into the second feature extraction network to obtain second semantic features corresponding to the physical examination information; and input the first semantic features and the second semantic features into the personalized network to determine the first recommendation information.
10. An information recommendation device characterized by comprising: The method comprises: a processor; and a memory connected with the processor, configured to provide the processor with instructions for processing the following processing steps: acquiring personal information of a user and acquiring physical examination information of the user, wherein the personal information comprises first health information of a first preset time period; determining first recommendation information according to the personal information and the physical examination information based on an information recommendation model deployed in an edge device of the user, wherein the first recommendation information is recipe information recommended for the user, and the information recommendation model and a discrimination model are pre-trained by a server device; judging whether to recommend the first recommendation information to the user according to the discrimination model based on the first recommendation information; in a case where the first recommendation information is not recommended to the user, determining first reward information and updating parameters of at least part of network structures in the information recommendation model in a manner of reinforcement learning according to the first reward information; in a case where the first recommendation information is recommended to the user and the user performs a service according to the first recommendation information, acquiring updated health information of the user, and determining second reward information according to the updated health information, and updating parameters of at least part of network structures in the information recommendation model in a manner of reinforcement learning according to the second reward information, wherein an updated information recommendation model is used to continue to recommend information for the user.
Citation Information
Patent Citations
Drug recommendation method and device, electronic equipment and storage medium
CN111753543A
Healthy recipe adjusting method and device, equipment and storage medium
CN120183615A
Personalized recipe recommendation system, recommendation method and interaction system
CN120809081A
Model training method and device, computer equipment and storage medium
CN121306426A
Intelligent query decomposition, specialized model routing, and hierarchical aggregation with conflict resolution
US20250378099A1