Recommendation method and device based on large language model
By combining target images and user interest tendencies, using large language models to identify user intentions and generate recommended content, the problem that existing recommendation systems are difficult to meet the diverse query needs of vulnerable and technically sensitive groups is solved, and the accuracy of user experience and recommendations is improved.
Patent Information
- Application Number
- CN202510436153.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-08
AI Technical Summary
It is difficult for existing recommendation systems to provide reasonable services or content recommendations to vulnerable groups based on the real query needs of vulnerable groups. For groups with high technical sensitivity, typing input can provide limited information, which is difficult to meet the rich and diverse query needs.
By combining target images and user interest tendencies, a large language model is used to identify user intentions and generate recommended content to users. The method includes acquiring a target image of a user, analyzing the subject in the image through a large language model, combining the user's interest information, determining the target subject and generating recommended content.
It improves the user's query experience, ensures the fairness of the use of network services, and the generated recommended content is more in line with user interests and actual scenarios, and lowers the user's operation threshold.
Smart Images

Figure CN119988660A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and specifically, to a recommendation method and device based on a large language model. Background Art
[0002] With the continuous evolution of mobile Internet, people's lives are increasingly inseparable from the query and recommendation functions of the Internet. The current query and recommendation functions rely on users to search through typing input. For vulnerable groups with low technical sensitivity, the current "giant apps" or "super apps" usually deeply integrate multiple service functions on the same platform, resulting in complicated function entrances and too many levels, which significantly increases the threshold for these vulnerable groups to use functions such as query and recommendation. For example, they may need to spend a lot of time to find the input box and enter text, the recognition accuracy is not stable when using voice input, and the error rate is high when typing. Therefore, it is difficult for the current recommendation system to provide reasonable service or content recommendations for vulnerable groups based on their real query needs. In addition, even for groups with high technical sensitivity, the information that can be provided by typing input is limited, making it difficult for them to meet rich and diverse query needs.
[0003] Therefore, how to ensure fairness in the use of network services and provide a rich experience of fast access to information and services has become a real challenge that needs to be addressed urgently. Summary of the invention
[0005] The embodiments of this specification provide a recommendation method based on a large language model, which effectively improves the user's query experience by identifying the user's intention by combining the target image and the user's interest tendency.
[0006] In a first aspect, an embodiment of the present specification provides a recommendation method based on a large language model, comprising: obtaining at least one subject of a user query, wherein the at least one subject is identified from a target image; obtaining first information of the user, wherein the first information is used to indicate the interest tendency of the user; based on the first information, determining a target subject among the at least one subject; and generating at least one recommended content for the user based on the target subject through the large language model.
[0007] In some embodiments, the obtaining of the first information of the user includes: obtaining associated information of the user; and generating the first information of the user based on the associated information through the large language model.
[0008] In some embodiments, the user's associated information includes at least one of the following information: attribute information of the user, historical behavior data of the user, location information of the user, and current time information of the user.
[0009] In some embodiments, determining a target subject among the at least one subject based on the first information includes: determining a target subject among the at least one subject and a vocabulary associated with the target subject based on the first information; generating at least one recommended content for the user based on the target subject through the large language model includes: generating at least one recommended content for the user based on the target subject and a vocabulary associated with the target subject through the large language model.
[0010] In some embodiments, determining a target subject in the at least one subject and vocabulary associated with the target subject based on the first information includes: retrieving associated information and / or context information of the at least one subject based on a preset corpus, the context information including associated words expanded based on the subject and / or the associated information; generating vocabulary associated with the subject based on the associated information, the at least one subject and the context information through a first multimodal large model; determining the target subject and vocabulary associated with the target subject among the at least one subject and vocabulary associated with the subject based on the first information.
[0011] In some embodiments, the method of determining the target subject and the vocabulary associated with the target subject among the at least one subject and the vocabulary associated with the subject based on the first information includes: semantically encoding the first information to obtain a first semantic vector; semantically encoding each of the subject and the vocabulary associated with the subject to obtain a second semantic vector; and screening the second semantic vector based on the similarity between the first semantic vector and the second semantic vector to determine the target subject and the vocabulary associated with the target subject among the at least one subject.
[0012] In some embodiments, the vocabulary includes hypernyms of the target subject
[0013] In some embodiments, the large language model is a spatiotemporal large language model, and the spatiotemporal large language model is fine-tuned in the following manner: obtaining a spatiotemporal dataset and labels of sample users, the spatiotemporal dataset including: associated information of sample users, and the labels including services used by the sample users; using the spatiotemporal large language model, based on the spatiotemporal dataset, predicting the services used by the sample users; and adjusting the network parameters of the spatiotemporal large language model based on the prediction results and the labels.
[0014] In some embodiments, obtaining at least one subject of a user query includes: obtaining the target image input by the user, analyzing the structure in the target image based on the target image through a second multimodal large model, and generating a scene graph, wherein the scene graph includes the at least one subject and the relationship between the subjects.
[0015] In some embodiments, the target image is an image acquired by capturing the scene in which the user is currently located.
[0016] In some embodiments, generating at least one recommended content for the user based on the target subject through the large language model includes: generating at least one recommended content for the user based on the target subject and associated information through the large language model.
[0017] In a second aspect, an embodiment of the present specification provides a recommendation method based on a large language model, comprising: obtaining a target image of a user; displaying recommended content corresponding to the target image to the user, wherein the recommended content is generated by the large language model for the user based on a target subject, wherein the target subject is determined from at least one subject based on first information, wherein the first information is used to indicate the interest tendency of the user, and the at least one subject is identified from the target image.
[0018] In some embodiments, a content generation direction is obtained, where the content generation direction is used to indicate the information type of the recommended content, and the recommended content is generated for the user by the large language model based on the target subject and the content generation direction.
[0019] In some embodiments, after displaying the recommended content corresponding to the target image to the user, the method further includes: in response to a first operation of the user on the recommended content on an interactive interface, storing the recommended content in a preference record of the user, the first operation being used to indicate the user's interest in the recommended content, and the preference record being used to generate the first information.
[0020] In some embodiments, after displaying the recommended content corresponding to the target image to the user, the method further includes: storing the recommended content in a historical browsing record of the user, wherein the historical browsing record is used to generate the first information.
[0021] In a third aspect, an embodiment of the present specification provides a recommendation device based on a large language model, comprising: a subject query module, configured to obtain at least one subject queried by a user, wherein the at least one subject is identified from a target image; an information acquisition module, configured to obtain first information of the user, wherein the first information is used to indicate the interest tendency of the user; a target determination module, configured to determine a target subject among the at least one subject based on the first information; and a content generation module, configured to generate at least one recommended content for the user based on the target subject through the large language model.
[0022] In a fourth aspect, an embodiment of the present specification provides a computing device, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any implementation manner in the first aspect or the second aspect is implemented.
[0023] In the solution provided by the above-mentioned embodiments of the present specification, at least one subject in the identified target image is combined with the first information indicating the user's interest tendency, and a large language model is used to predict the search intent, thereby generating recommended content that the user may be interested in. The usage threshold is very low, ensuring the fairness of the use of network services. The generated recommended content is more in line with the user's interests and actual scenarios, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0025] Figure 1 It is a schematic diagram of an application scenario in which the embodiments of this specification can be applied; Figure 2 is a schematic diagram of a mobile phone interface in an embodiment of this specification; Figure 3 is a flow chart of a recommendation method based on a large language model in an embodiment of this specification; Figure 4 is a flow chart of the process of acquiring a spatiotemporal large language model in an embodiment of this specification; Figure 5 is a flowchart of a recommended process in the embodiments of this specification; Figure 6 is another flow chart of the recommendation method based on the large language model in the embodiment of this specification; Figure 7is a flowchart of another recommendation method based on a large language model in an embodiment of this specification; Figure 8 is a schematic diagram of an interactive interface in an embodiment of this specification; Fig. 9 is a schematic diagram of another interactive interface in an embodiment of this specification; Fig.10 is a schematic diagram of the structure of a recommendation device based on a large language model in an embodiment of this specification; Fig.11 It is a structural diagram of another recommendation device based on a large language model in the embodiments of this specification. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0027] As mentioned above, compared to querying by typing, it is easier for users to express their needs and provide more intention information through low-threshold methods such as taking photos. To this end, the embodiment of this specification proposes a recommendation scheme based on a large language model (LLM) to identify user intentions based on target images combined with user interest tendencies and generate recommended content for users. This scheme is based on a multimodal human-computer interaction idea: allowing users to use more natural interaction methods such as taking photos to truly, directly and conveniently reflect their query needs in different scenarios. Users only need to provide images as queries, and the system can explore and generate corresponding recommended content based on their visual content combined with user interest tendencies to meet the user's query needs.
[0028] Figure 1 FIG. 1 is a schematic diagram showing an application scenario in which the embodiments of this specification can be applied. Figure 1 As shown, in Figure 1 In the application scenario shown, the user is near a self-service terminal in a hospital. His or her possible needs are to register, print reports, and query information. The user's target image can be obtained through the client installed on the user's terminal device. The target image can be obtained by the user taking a photo of the self-service terminal in the hospital environment through the terminal device (such as a mobile phone), and then the recommended content corresponding to the target image is displayed to the user, such as Figure 2As shown, the recommended content displayed on the mobile phone interface includes: "How to register for medical insurance by yourself?" "How does the self-service terminal change the service experience?" The recommended content here is generated by the large language model for the user based on the self-service terminal as the target subject. The self-service terminal is determined from at least one subject in the target image based on the first information. In addition to the self-service terminal as the subject, the target image may also contain subjects such as green plants and water dispensers in the hospital. The first information is used to indicate the user's interest tendency. In this example, the user's interest tendency may be a tendency related to low blood pressure. At least one subject is identified from the target image taken by the user.
[0029] In the above process, the client may send the acquired target image to a remote server, which processes the target image and generates at least one recommended content for the user, and sends the recommended content to the client. Alternatively, the client may identify the target image to obtain at least one subject, and then send the at least one subject to the server, which processes the at least one subject and generates at least one recommended content for the user, and sends the recommended content to the client. It is understandable that when the computing power of the terminal device where the client is located is sufficient, the target image may also be processed on the client side to generate at least one recommended content for the user.
[0030] Continue to see Figure 3 , Figure 3 A flowchart of a recommendation method based on a large language model according to an embodiment is shown. The method can be executed by any device, platform or device cluster with computing and processing capabilities, and can be applied to the client side or the server side mentioned above, including steps S301-304 as shown below.
[0031] like Figure 3 As shown, in step S301, at least one subject of the user query is obtained.
[0032] Among them, at least one subject is identified from the target image, and the subject refers to an independent object or entity existing in the image or video. These objects or entities can be people, animals, objects or any other elements that can be identified and distinguished in visual content. The target image is an image that the user has a query requirement. In practice, the target image can be an image that the user takes, screenshots, selects or otherwise determines based on his query requirements and uploads to the client. The target image usually contains at least one subject queried by the user. For example, when the target image is an image obtained by the user taking a photo of the desktop, the teacup, paper, pen, laptop, tea leaves in the teacup and the user's exposed pants on the desktop in the image can all be considered as subjects.
[0033] This embodiment does not limit the method for obtaining at least one subject of the user query. For example, a specific implementation method may be to extract the subject from the target image input by the user through image analysis, entity recognition and other technologies.
[0034] In step S302, the first information of the user is obtained.
[0035] Among them, the first information is used to indicate the user's interest tendencies. In practice, the first information may be topics or services that the user is interested in, specifically topics or online services related to people's livelihood and government affairs, commodity information, cultural and technological fields, historical and geographical knowledge, and current affairs information. For example, the first information may be football matches that the user cares about, traditional culture that the user is interested in, provident fund withdrawal services, etc.
[0036] This embodiment does not limit the method of obtaining the user's first information. For example, the user's first information can be obtained by analyzing the user's historical behavior data, or by analyzing the attribute information uploaded by the user. For example, if the historical behavior data is the user's bill record, running shoes, knee pads, sports drinks and other items appear in the bill record many times, the user's first information can be analyzed to be a topic word related to sports and fitness, indicating the user's interest in sports and fitness.
[0037] As an implementation method, in order to accurately obtain the user's first information, this step may be to obtain the user's associated information, and generate the user's first information based on the associated information through a large language model.
[0038] The user's associated information is user personal information related to the user, such as the user's occupation, the user's crowd portrait, the user's historical behavior, etc. Exemplarily, the user's associated information may include at least one of the following information: the user's attribute information, the user's historical behavior data, the user's location information, and the user's current time information.
[0039] Specifically, the user's attribute information is used to describe the user's basic situation, interest preferences, behavioral habits, such as education level, social role, exercise habits, etc., in order to better model user portraits; the user's historical behavior data is used to describe the user's online behavior, for example, it may include the user's billing records, records of using mini-programs, search history, and historical records of visiting specific locations on the client at specific times, from which the user's continuous behavior in different time and space scenarios can be obtained to better analyze user behavior patterns; the user's location information can be the geographical location information of the user when taking or uploading the target image, or it can be the shooting location information carried by the target image itself, in order to better predict behavior patterns through location information; the user's current time information can be the time when the user takes the target image, or it can be the time when the user uploads the target image to the client, in order to better predict behavior patterns through time information.
[0040] The large language model used here is an artificial intelligence model based on deep learning, which is specifically used to understand and generate natural language. It can be obtained by fine-tuning the open source large language model, or it can be trained so that the large language model can learn the deep association between the user's associated information and the first information through supervised training. The acquisition process of the large language model used in the embodiments of this specification will be described in detail later, and no further description will be given here for the time being.
[0041] In practice, prompt information can be constructed based on the associated information, and the prompt information can be input into the large language model to obtain the first information of the user output by the large language model. Exemplarily, the constructed prompt information can be: based on the user's portrait information {p} and recent purchase history {h}, combined with the current time {t} and location information {l}, predict what kind of service {f} the user needs, thereby obtaining the service prediction {f} output by the model.
[0042] It is understandable that this embodiment does not limit the execution order of step S301 and step S302.
[0043] In one example, in order to reduce the amount of calculation when the solution is applied online, the process of generating the user's first information can be completed offline in advance and stored in an offline database, and the first information can be obtained from the offline database.
[0044] Next, in step S303, based on the first information, a target subject in at least one subject is determined.
[0045] Specifically, when there are multiple identified subjects, the target subject can be determined by the user's interest tendency indicated by the first information, that is, the target subject is the subject that the user is interested in, and this specification does not limit the number of target subjects.
[0046] In this embodiment, the target subject associated with the first information can be determined from multiple subjects through the association between the semantics of the first information and the semantics of the subject. Still taking the desktop object as an example, when there are multiple subjects on the desktop, such as a teacup, paper, pen, laptop, tea leaves in the teacup, and pants exposed by the user, when the user's first information indicates that the user is interested in traditional Chinese culture, because teacups and tea leaves are not only daily necessities, but also have semantic associations with traditional Chinese culture and carry rich cultural connotations and symbolic meanings, the determined target subject can be teacups and tea leaves. When the user's first information indicates that the user is interested in technology products, and there is a semantic association between laptops and technology products, the determined target subject can be a laptop.
[0047] Furthermore, when there is only one recognized subject, that subject may be directly determined as the target subject.
[0048] Finally, in step S304, at least one recommended content is generated for the user based on the target subject through the large language model.
[0049] In practice, a large language model with prompt words as input can be constructed based on the target subject and the user, and the large language model can output corresponding recommended content.
[0050] The prompt word contains information about the target subject and may also contain information related to the user, so that the large language model can generate recommended content that the user may be interested in, and the recommended content is semantically related to the target subject. The recommended content can be text information, goods, services or other resources that are personalized and recommended to the user.
[0051] For example, when the target subject is a computer host of brand A, the constructed prompt word may be "Based on the target subject, raise questions that the user may be interested in." The recommended content generated by the large language model based on the prompt word may be text information such as "How can a computer host of brand A improve work efficiency?" and "How is the cooling system of a computer host of brand A designed?" When the user clicks on the text information, the text information can be input into the large language model or other large language models to obtain more relevant text or picture content generated by the large language model to answer related questions and satisfy the user's curiosity.
[0052] In some embodiments, this step may generate at least one recommended content for the user based on the target subject and the associated information through a large language model.
[0053] For example, when the recommended content is text information, the constructed prompt words may be "comprehensively analyze the user's known background information and the target subject in the target image, and propose an interactive question for the user. The question content must be related to the content in the target image provided by the user, combined with the user's interests, background and scene characteristics." For example, when the target subject is tea and teacups, and the user is interested in traditional Chinese culture, the generated recommended content may be "What kind of tea is suitable for this teacup?", "What are the types and materials of traditional Chinese teacups?" and "What types of tea are there in China?", thereby arousing the user's interest.
[0054] Furthermore, the control where the recommended content is located can be associated with the online service interface provided by the client. When the user clicks on the recommended content, he can enter the corresponding online service page, thereby quickly mapping the user's time and space needs in the real scene to the functional entrance of the relevant application program, and obtaining timely feedback and guidance. For example, when the recommended content is "How to register for medical insurance online?", after the user clicks on the recommended content, he can jump to the mini program interface that provides the registration function; when the recommended content is "How to withdraw housing provident fund online?", he can jump to the mini program interface that provides housing provident fund withdrawal.
[0055] This multimodal interaction not only expands the application scenarios of large language models from pure text to visual fields, but also reflects the potential of large models in image and text understanding and multimodal search. Through multimodal search as an entry point, the user's personalized intention can be connected with the diversified functions in network services. Even for similar image queries, due to the differences in preferences and needs of different users, the solution of this embodiment can obtain personalized responses that vary from person to person. This intelligent and diversified answer is crucial to meeting user needs and improving user experience.
[0056] The current mainstream recommendation systems mainly rely on ID (Identity Document) recommendations and historical behavior analysis. They are good at mining user preferences from structured data and can provide relatively accurate targeted content recommendations for existing, patterned user behaviors. However, such systems often seem to be unable to cope with real-time multimodal requests, especially image-dominated query scenarios. For example, when a user takes a photo and uploads a scene in daily life for query, this unstructured visual information exceeds the modeling capabilities of traditional recommendation systems or retrieval systems, making it difficult to quickly generate reasonable service or information recommendations.
[0057] Taking the above problems into consideration, in response to the dynamic needs of users, the large language model used in the embodiments of this specification may be a spatial temporal-LLM (ST-LLM for short), which is obtained by training or fine-tuning the large language model using large-scale spatial temporal data. It has certain spatial and temporal prediction capabilities and can capture the specific needs of users in different scenarios. The generated recommended content can better meet the current needs of users.
[0058] Combine the following Figure 4 The example of the acquisition process of the spatiotemporal large language model used in the embodiments of this specification is described. Figure 4 The example shown may include steps S401-S403, specifically:
[0059] In step S401, a spatiotemporal dataset and labels of sample users are obtained.
[0060] First, we need to build a spatiotemporal dataset of a certain scale, which contains the association information of sample users. In spatiotemporal prediction tasks, fine-grained, multi-dimensional user behavior data is usually required, but existing public datasets often cannot meet this requirement. To this end, we need to collect user behavior information on online platforms and build a corpus as a spatiotemporal dataset.
[0061] When constructing a spatiotemporal dataset, relevant information about sample users can be obtained from a variety of data sources with the authorization and permission of the sample users. For example, the data source can be social networks, forums and communities, data provided by sample users by creating accounts or filling out forms, and the browsing behavior and interaction history of sample users on online platforms.
[0062] For example, the billing records and historical interaction information of sample users can be selected as data sources, where the historical interaction information mainly includes records of using mini programs, search history records, and records of visiting specific locations of the client at specific times. The reason for selecting these two types of data is that the continuous consumption behavior of users in different time and space scenarios can explicitly or implicitly reflect their goals, intentions, interests, and knowledge depth. Based on this long-term accumulated information, not only can the user's behavior patterns be accurately portrayed, but also key knowledge support can be provided for the generative recommendation system.
[0063] For example, for any unsupervised document d from the document set D, the corpus construction system F can be used to extract four types of key information from it: attribute information P of the sample user, time T, location L, and interaction history H. The corpus construction process can be formalized as the following formula: (1) Among them, the document set D refers to a set of unsupervised documents, that is, text data without annotation labels, such as the bill records and historical interaction information of sample users; the unsupervised document d is any document in the document set D, for example, the consumption record of a user on a certain day; the attribute information P of the sample user can be the sample user ID, occupation, etc.; the time T can be the consumption time or the interaction time; the location L can be the consumption location or the interaction location; the interaction history H can be the sample user's click, purchase, comment and other behavior records.
[0064] These four types of information are related to each other and together provide rich background context for spatiotemporal prediction and generative recommendation models. Based on this large-scale, high-quality corpus, the system can not only more accurately capture users' spatiotemporal behavior patterns, but also provide more precise and in-depth support for generative recommendation models.
[0065] Then, for each sample user in the spatiotemporal dataset, it can be annotated to obtain a label, which contains the services used by the sample user. For example, it can be manually labeled, or the services used by the sample user after the time period covered by the spatiotemporal dataset can be used as labels.
[0066] In step S402, the services used by the sample users are predicted based on the spatiotemporal data set using the spatiotemporal large language model.
[0067] In practice, Figure 5 As shown in the Spatial Temporal Guidance module in the figure, the spatial temporal guidance module depicts the fine-tuning process of the spatial-temporal large language model before its actual application. For each sample user, a prompt word can be constructed based on its corresponding data in the spatial-temporal dataset, and the prompt word can be input into the spatial-temporal large language model. The spatial-temporal large language model will output the predicted services that the sample user may use, such as which products to buy.
[0068] The predetermined format of the prompt words can be set by the staff as needed. Exemplarily, this embodiment designs a set of comprehensive input construction schemes, and the prompt words follow the following instruction template: given the attribute information of the sample user and the corresponding historical interaction data in the spatiotemporal dataset, combined with the current time and location information, predict the functional services that the sample user may use. For example, when the sample user is near a restaurant, it is predicted that the sample user may use online ordering or takeaway services, and when the sample user is in the gym, it is predicted that the sample user may need to buy healthy food online. By providing large-scale annotated examples in the supervised learning stage, the model can learn the deep connection between user characteristics and spatiotemporal factors, so as to more accurately understand the user context during the reasoning process, and then provide highly targeted and practical service recommendations.
[0069] In step S403, based on the prediction results and labels, the network parameters of the spatiotemporal large language model are adjusted.
[0070] In practice, the gap between the model's prediction results and the label can be measured by the loss function, and the prediction accuracy can be improved by adjusting the network parameters to minimize the loss function. This process can be fine-tuned by using supervised learning examples, so that the spatiotemporal large language model can obtain the ability to predict service behaviors based on the user's personal information and historical behavior information in a specific spatiotemporal scenario. This embodiment does not limit the specific method of adjusting the network parameters used. For example, some or all network parameters can be selected for fine-tuning based on the pre-trained model.
[0071] By providing large-scale annotated examples in the supervised learning stage, the model can learn the deep correlation between user characteristics and spatiotemporal factors, thereby more accurately understanding the user context during the reasoning process, and providing highly targeted and practical service recommendations.
[0072] Next, combine Figure 5 The flowchart of the recommendation process shown in FIG. 1 is a flowchart of the recommendation process shown in FIG. 1 , which introduces the recommendation process based on the large language model provided by the embodiment of this specification in more detail. Figure 5 It mainly includes: an input module (Input), which is used to display the input data required in the implementation of this solution; a spatiotemporal guidance module, which is used to fine-tune the spatiotemporal large language model; a preference discovery module, which consists of steps 1, 2, 3 and 4. The sequence numbers of the steps are only used to identify different steps, and the execution order of the steps is not restricted. In step 1: a scene graph is generated based on the target image through a recognizer (Recognizer); in step 2: a knowledge graph is generated based on the user's attribute information and the scene graph through a generator (Generator); in step 3: a first information is generated based on the user's attribute information and historical behavior data through a pre-trained spatiotemporal large language model; in step 4: a recaller (Retriever) is used to retrieve and sort the knowledge graph according to the first information to determine the target subject and the vocabulary associated with the target subject; a personalized recommendation module is used to finally generate recommended content using the spatiotemporal large language model.
[0073] Figure 6 This is a flowchart of another recommendation method based on a large language model in the embodiments of this specification. The method can be executed by any device, platform or device cluster with computing and processing capabilities, including steps S601-S604 as shown below.
[0074] like Figure 6As shown, in step S601, a target image input by a user is obtained, and a structure in the target image is analyzed based on the target image through a second multimodal large model to generate a scene graph.
[0075] The scene graph includes at least one subject and the relationship between the subjects, and the scene graph can be represented by a ternary set.
[0076] In some embodiments, the subject recognition of the target image can be performed directly through a multi-modal large language model (Multilingual Large Language Model, MLLM for short) to obtain information about at least one subject. However, considering that the multi-modal large model is used for visual content understanding, the model usually returns the recognition result in the form of text, which contains rich textual descriptive information. These textual description information are often embedded in complex grammar and sentence structures, making it difficult to efficiently retrieve structured knowledge directly from visual queries. It is necessary to process complex grammar and sentence structures to extract concise and effective information. In order to solve this problem, this embodiment proposes to use the form of Knowledge Graph (KG) triples to represent visual entities (i.e., subjects) and the relationships between them, so as to explicitly model the complex relationships between visual objects. This method provides greater flexibility for subsequent knowledge relationship screening and expansion by effectively constructing scene graphs for subsequent reasoning and recommendation.
[0077] Specifically, if Figure 5 The recognizer shown in step 1 in can introduce a second multimodal large model (such as an open source large model) to perform visual queries on the input target image and extract entity information therein to obtain the scene graph generated by MLLM, as shown below: (2) in, It consists of a series of triplets, such as , each triple represents a visual entity With another entity The relationship between , the triple set can be as follows: (3) Unlike natural language descriptions, knowledge graphs use triples to convert information into structured representations. This structured representation not only expresses complex knowledge concisely and clearly, but also has strong flexibility and scalability, facilitating subsequent information retrieval, reasoning, relationship discovery, knowledge integration and other tasks. By storing information in the form of standardized triples, knowledge graphs can more efficiently support a variety of subsequent tasks, including information retrieval, reasoning, relationship discovery, etc. In addition, the triple form makes it easier to add new knowledge and integrate existing knowledge, greatly enhancing the scalability of knowledge graphs and further improving their application value in complex tasks.
[0078] For example, Figure 5 Step 1 further illustrates a visualized scene graph, in which the subjects included are: a teacup, paper, a pen, a laptop, tea leaves in the teacup, and pants. An arrow indicates the relationship between two subjects, for example, tea leaves are in the teacup and the pen is on the paper.
[0079] It can be understood that the target image of the user query in this embodiment can be an image obtained by capturing the scene in which the user is currently located. Figure 5 As shown in the input module, the user uses the application on the mobile phone to take a picture of the desktop in front of him to obtain the target image and upload it, so as to make recommendations based on the real-time needs of the user in the real space-time scene.
[0080] In other embodiments, in order to reduce the amount of online calculations, the process of obtaining a scene graph may be to perform image recognition on a target image, obtain information of at least one subject in the target image, vectorize the information of at least one subject to obtain a subject vector, and then match it in a scene graph library. The scene graph library pre-stores a large number of scene graphs corresponding to different types of images. By matching the scene graph that is closest to the subject vector, the scene graph corresponding to the target image can be determined, thereby avoiding the use of large models online to generate scene graphs and speeding up the online response speed.
[0081] Next, in step S602, the user's associated information is obtained, and the user's first information is generated based on the associated information through a spatiotemporal large language model.
[0082] For example, Figure 5 As shown in step 3 of the preference discovery module, the associated information input into the spatiotemporal large language model may include the attribute information of the user and the historical behavior data of the user. After the associated information is input into the spatiotemporal large language model, the spatiotemporal large language model may model the user's interests and predict the first information indicating the user's interest tendency.
[0083] For a more detailed explanation of the associated information and the first information, please refer to the relevant instructions in the previous text, which will not be repeated here.
[0084] In step S603, based on the first information, a target subject in at least one subject and a vocabulary associated with the target subject are determined.
[0085] Among them, the vocabulary that has an association relationship with the target subject can be a vocabulary that has a semantic association with the target subject. The vocabulary is used to provide diversified information to expand the coverage of the recommended content. For example, when the target subject is tea, the vocabulary that has a semantic association with it can be "green tea", "tea art", "tea culture", "health preservation", etc.; the vocabulary that has an association relationship with the target subject can also be a vocabulary related to the user's interests, so that the recommended content is more targeted. For example, when the target subject is a laptop computer and the user is a bank clerk, the vocabulary related to the user can be "productivity tools", "data analysis", "efficient office", etc. It can be understood that the above-mentioned association relationship can be a direct association relationship or an indirect association relationship. In some embodiments, the vocabulary can include the hypernym of the target subject. For example, when the target subject is tea, its hypernym can be "plant", "drink", "tea", "leaves of plants", etc., to introduce a wider range of knowledge background information.
[0086] In some embodiments, each subject may be expanded and associated to obtain vocabulary, and then a target subject and corresponding vocabulary may be determined from multiple subjects and their corresponding vocabulary based on the first information; in other embodiments, the target subject may be first determined based on the first information, and then the target subject may be expanded to obtain vocabulary that has an associated relationship with the target subject.
[0087] As an implementation method, this step may specifically include the following steps a)-c):
[0088] Step a), based on a preset corpus, retrieve and obtain related information and / or context information of at least one subject.
[0089] Among them, the context information includes associated words expanded based on the subject and / or related information. Exemplarily, the associated words can be hypernyms or hyponyms. For example, the hypernym of bank employee can be bank, finance, securities, and the hyponym can be teller, account manager, etc.
[0090] The preset corpus can be various open knowledge bases or private corpora. For example, Figure 5The generator shown in step 2 of , in the scene graph containing entities, can use the OpenKnowledge Retriever to retrieve the contextual information most relevant to the current scene from multiple open corpora to provide semantically rich background knowledge and ensure that the information sources relied on by the scene expansion are extensive and diverse, thereby improving the coverage and accuracy of knowledge.
[0091] Step b) generates words that are associated with the subject based on the association information, at least one subject, and context information through the first multimodal large model.
[0092] In practice, a prompt word can be constructed based on the association information, at least one subject, and context information. The prompt word is used to prompt the big model to generate vocabulary related to both the user and the subject, and the prompt word is input into the first multimodal big model to obtain the vocabulary output by the big model. Exemplarily, the first multimodal big model can be an existing open source big model.
[0093] It can be understood that the vocabulary is essentially a knowledge enhancement based on the scene graph to better utilize large language models as user agents to simulate user preferences and interests.
[0094] Continue to see Figure 5 In step 2, the process of generating vocabulary in this step can be regarded as a contextual expansion of the original scene graph to generate a knowledge graph. The knowledge graph can also be represented by a set of triples. Each triple represents the association between the subject and the vocabulary, and multiple triples form a multi-hop chain relationship.
[0095] by Figure 5 Take the knowledge graph in as an example, where the vocabulary in the blue module is expanded from the subject, and the vocabulary in the yellow module is expanded from the associated information. It can be seen that there is a one-hop relationship between tea leaves and tea cups, and the relationship between the two subjects is represented by a triple. There is a two-hop relationship between tea leaves and tableware, and the relationship between the two subjects is represented by two triples, namely the triple composed of tea leaves and tea cups and the triple composed of tea cups and tableware. There is a three-hop relationship between tea leaves and idea extension, and the relationship between the two subjects is represented by a combination of multiple triples. This process not only retains the original structure of the scene, but also combines the user's attribute information (such as long-term interests, recent concerns, etc.) for personalized adjustments to ensure that the generated knowledge graph is not only rich but also controllable and targeted.
[0096] The predetermined format of the prompt word can be set by the staff as needed, and is generally set to a format with strong readability. For example, the constructed prompt word A can be as follows:
[0097] You are an advanced multimodal AI assistant that can analyze image content, structured data, and text descriptions, and combine multiple information to provide users with accurate and targeted output. You will think as a user, combine the user's interests and background, extract the knowledge related to their interests in the picture, and output it in a specific format. Your responsibility is to ensure the accuracy, logic, and high relevance of the output results.
[0098] **Mission Statement**: Given a picture taken by a user and multiple triplets describing the relationships between subjects in the picture, in the format of (head subject, relationship, tail subject), the encyclopedia information of these subjects and the basic information of the user (such as occupation, crowd portrait, region, and topics of interest) are provided.
[0099] 1. Multiple triplets describing the relationship between the subjects in the image: {content} 2. Encyclopedia information of each entity: {entities} 3. User Information: - Occupation: {work} - Crowd portrait: {crowd} - Region: {region_name}-{prov_name}-{city_name}-{city_degree} - Interesting topics: {interest} Your task is to: 1. Extract content related to user interests from the subject’s encyclopedia information.
[0100] 2. The extracted content must revolve around the subject relationship in the image and be highly relevant to the user’s interest background.
[0101] 3. The final output format is a triple: (subject, relation, other object or information). The "subject" in the output must be one of the head subject or tail subject of a given triple in the image.
[0102] **Tips and Rules**: 1. Filter information based on user interests: Based on the user's basic information (interests, occupation, etc.), extract the main encyclopedia content that the user may be interested in. You need to highlight the content that is highly relevant to the user's interests and ignore irrelevant or unnecessary details.
[0103] 2. Ensure the format and logic of the output: The final output is in the triple format: (subject, relationship, other objects or information). Among them: - Either the head body or the tail body must originate from a given triple.
[0104] - "Relationship" should briefly describe the related content.
[0105] - “Other objects or information” may include encyclopedia-related knowledge of the subject or other information that matches the user’s interests.
[0106] 3. Rigorousness and coherence: Ensure that the output triples are logically clear, the information is accurate, and highly consistent with the user's interests.
[0107] **Output requirements**: - Only output triples that are factually plausible.
[0108] - The triple results are given directly in strict accordance with the example output format, and no explanatory text is allowed.
[0109] **Example Format**: (XXX, Relationship A, YYY) (XXX, Relationship B, ZZZ) Then, after the prompt word A is input into the first multimodal large model, a set of triples consisting of multiple subjects and words associated with them, i.e., a knowledge graph, is output, thereby obtaining words associated with the subjects. It can be understood that this embodiment does not limit the number of words associated with each subject.
[0110] In other embodiments, in order to reduce the amount of online calculation, step a) and step b) can be replaced by an offline execution scheme. Specifically, a query vector can be constructed based on the associated information and at least one subject, and then matched in the knowledge graph library. The knowledge graph library pre-stores a large number of different types of images and knowledge graphs corresponding to users. By matching the knowledge graph closest to the query vector, the knowledge graph can be determined, avoiding the use of a large model online to generate the knowledge graph and speeding up the online response speed.
[0111] Step c) is to determine, based on the first information, a target subject and words associated with the target subject from among at least one subject and words associated with the subject.
[0112] In practice, the semantic similarity between the first information and each subject-vocabulary pair (i.e., the subject and the vocabulary associated with the subject) may be calculated, and then one or more subject-vocabulary pairs with the highest similarity may be determined as the target subject and the vocabulary associated with the target subject, so as to find the target subject and vocabulary that are closest to the user's interest tendency. This embodiment does not limit the method for calculating the similarity, for example, a neural network may be used for calculation.
[0113] As an implementation method, this step may specifically include the following steps 1)-3): Step 1) semantically encode the first information to obtain a first semantic vector.
[0114] Step 2) semantically encode each subject and the words associated with the subject to obtain a second semantic vector.
[0115] Still Figure 5 For example, see Figure 5 The recaller shown in step 4 of , in order to further build a personalized knowledge graph, can perform hop-by-hop semantic indexing on the knowledge graph to effectively retrieve relevant information and efficiently capture first-order neighbors with clear semantic associations as a new basis for knowledge retrieval, thereby enhancing the knowledge matching ability of user interests.
[0116] Exemplarily, a pre-trained language model (eg, a generative encoder) may be introduced to semantically encode the first information and the triples in the knowledge graph, respectively, to obtain a first semantic vector and a second semantic vector, thereby capturing their latent semantic information.
[0117] Step 3), based on the similarity between the first semantic vector and the second semantic vector, the second semantic vector is screened to determine a target subject in at least one subject and words associated with the target subject.
[0118] Then, the cosine similarity between the first semantic vector and the second semantic vector can be calculated. In other examples, the similarity can also be measured using methods such as Euclidean distance. In the knowledge graph, since the relationships between multiple subjects and words are connected through a chain relationship composed of multiple triples, the semantic similarity between the first semantic vector corresponding to the first information and the second semantic vector corresponding to the triple can be calculated hop by hop.
[0119] For example, Figure 5As shown in the retrieval module, when tea is taken as the starting point, for a chain relationship such as tea-pen-recording tool-expression medium, the first jump is the association relationship between tea and pen, and its corresponding second semantic vector can be obtained by semantically encoding the triple consisting of tea and pen, the second jump is the association relationship between tea and recording tool, and its corresponding second semantic vector can be obtained by semantically encoding the triple consisting of tea and pen and the triple consisting of pen and recording tool, and the third jump is the association relationship between tea and expression medium, and its corresponding second semantic vector can be obtained by semantically encoding the triple consisting of tea and pen, the triple consisting of pen and recording tool, and the triple consisting of recording tool and expression medium. For a chain relationship such as tea-beverage-cultural experience, the first jump is the association relationship between tea and beverage, and its corresponding second semantic vector can be obtained by semantically encoding the triple consisting of tea and beverage, and the second jump is the association relationship between tea and cultural experience, and its corresponding second semantic vector can be obtained by semantically encoding the triple consisting of tea and beverage and the triple consisting of beverage and cultural experience.
[0120] When calculating hop by hop, the similarity between the first semantic vector and the second semantic vector corresponding to the first hop can be calculated first. If the similarity is higher than the set threshold, the triplet is considered to be highly correlated and retained; otherwise, continue to propagate backward along the graph structure and recursively calculate the similarity of the adjacent triplet of the next hop to ensure the scalability of the retrieval range. For example, calculate the similarity between the first semantic vector and the second semantic vector corresponding to the second hop, and determine whether it is higher than the set threshold. Repeat this cycle until the entire knowledge graph is traversed.
[0121] Finally, after completing the hop-by-hop indexing, the retained associations are screened and sorted according to semantic similarity. For example, the top five associations are selected, and the subjects and words associated with the associations are used as the target subjects and the words that have an association relationship with the target subjects for subsequent personalized recommendations or knowledge enhancement tasks.
[0122] like Figure 5As shown in the re-ranking module, by comparing the similarity between the first semantic vector and the second semantic vector corresponding to the association between tea and beverages, the similarity score of the association between tea and beverages is 0.9. By comparing the similarity between the first semantic vector and the second semantic vector corresponding to the association between tea and cultural experience, the similarity score of the association between tea and cultural experience is 0.7. The higher the similarity score, the closer the two are semantically. It can be seen that the user has a high interest in beverages. Question generation based on this can meet the needs of users. If only the target image is analyzed without combining the user's interests, it is difficult to guide the concept of beverages. Through the above method, structured knowledge can be effectively extracted from complex visual content, and personalized knowledge graphs can be generated according to user preferences, thereby improving the accuracy and relevance of information retrieval and recommendation.
[0123] It can be understood that in this embodiment, when there is only one subject in the target image, the subject is the target subject, and there may be multiple words associated with the subject. In this case, the above processing also needs to be performed to determine a limited number of words according to needs.
[0124] In addition, in step S604, at least one recommended content for the user is generated based on the target subject and the vocabulary associated with the target subject through the large language model.
[0125] In practice, based on the target subject and the vocabulary associated with the target subject, the prompt words can be constructed to input into the big language model, and the big language model can output the corresponding recommended content. The prompt words can contain the user's attribute information and the user's historical behavior data to provide the big model with rich background information, and can also contain the user's location information and the user's current time information to assist the big model in capturing the user's complex and real dynamic needs that change with time and space according to the actual time and geographical environment.
[0126] Exemplarily, the constructed prompt word B is as follows: You are a multimodal AI assistant. Your main task is to generate an attractive question for the user based on the pictures uploaded by the user, known background information, and user behavior data, so as to guide the user to click and interact. The generated question must be highly relevant to the content of the picture, and combined with the user's interests, background, and current scene characteristics, to be attractive and personalized. The following are your task requirements: **Mission Statement**: 1. A user uploads a photo. We know the target subject of the photo and the words that are related to the target subject. We provide this information in the form of triplets: {triplets}.
[0127] 2. User background information includes: - Occupation: {work} - Crowd portrait: {crowd} - Region: {region_name}-{prov_name}-{city_name}-{city_degree} - Interesting topics: {interest} 3. User’s recent purchase history: {history}.
[0128] 4. The time and location where the user currently takes the photo.
[0129] - Time: {time} - Location: {location} 5. Based on the above multi-dimensional information, you need to comprehensively analyze the user’s background and current scenario (time, place) and ask an interactive question for the user.
[0130] **Tips and Rules**: 1. The question content must be relevant to the picture content provided by the user, taking into account the user's interests, background and scene characteristics.
[0131] 2. Prioritize areas that users may be interested in, such as combining their occupation, region, purchase history, and topics of interest.
[0132] 3. If the target subject of the image is a certain type of product or scenery, you can try to guide users to think from the perspective of practicality or emotional association.
[0133] 4. If there are multiple target subjects in the image, give priority to the angle that is more relevant to the user's interests or the shooting scene.
[0134] 5. The generated questions should be concise and clear, not exceeding 18 words.
[0135] 6. Questions must be engaging and encourage users to click or interact further.
[0136] **Output Format**: Directly output a final generated interactive question without explanation or supplementary description. The question must not exceed 18 words.
[0137] In this way, after the prompt word B is input into the spatiotemporal large language model, the output recommendation content is the interactive question , the generation process can be shown as follows: (4) in, Represents the user, represents the target image, Indicates the user's related information. Represents the five identified target entities and words that are associated with the target entities, for example: tea and beverages, tea and cultural experience, laptops and productivity tools, teacups and tableware, and teacups and idea expansion.
[0138] This interactive question can respond to users' multimodal requests in a timely and personalized manner, returning to Figure 1 For example, since the user has historical behavior data related to low blood pressure, the generated recommendation content includes "What are the daily care for low blood pressure?", which can meet the user's needs in a timely manner.
[0139] Figure 7 A flowchart of a recommendation method based on a large language model proposed in an embodiment of this specification is shown. The method can be executed by any computing device and is described below using execution on a client as an example, including the following steps S701-S702.
[0140] In step S701, a target image of a user is acquired.
[0141] In this embodiment, the client can obtain the target image of the user in a variety of ways, for example, it can be obtained from the Internet, it can also be obtained from a local database, and it can also be obtained by calling an image acquisition device (for example, a camera). For example, the user can call the camera to capture the target image of the surrounding environment through the client.
[0142] In step S702 , recommended content corresponding to the target image is displayed to the user.
[0143] The recommended content is generated by the large language model for the user based on the target subject, the target subject is determined from at least one subject based on the first information, the first information is used to indicate the user's interest tendency, and the at least one subject is identified from the target image. For the generation process of the recommended content, please refer to the description in the previous embodiment, and this embodiment will not be repeated.
[0144] In this step, the recommended content corresponding to the target image may be displayed on the interface of the client, and each recommended content may be displayed at a different position on the interface.
[0145] In some embodiments, in order to make recommendations more accurately based on user needs, before generating recommended content, a content generation direction may also be obtained. The content generation direction is used to indicate the information type of the recommended content, and the recommended content is generated for the user by the large language model based on the target subject and the content generation direction.
[0146] For example, the content generation direction can be "knowledge", to indicate that the information type of the recommended content is the knowledge type, and the prompt words input into the large language model are adjusted so that it generates recommended content from a popular science and rigorous perspective; the content generation direction can also be "inspiration", to indicate that the information type of the recommended content is the knowledge type, and the prompt words input into the large language model are adjusted so that it generates recommended content from an interesting and humorous perspective.
[0147] In one example, a preset content generation direction may be obtained, or a specified content generation direction may be obtained. For example, different content generation direction controls are displayed on the interface, and the user can select or switch the content generation direction by clicking the control. Figure 8 A schematic diagram of the interface when the content generation direction is knowledge is shown.
[0148] In other embodiments, after step S702, in response to a first operation of the user on the recommended content on the interactive interface, the recommended content may be stored in the user's preference record.
[0149] The first operation is used to indicate the user's interest in the recommended content. For example, the first operation may be a user clicking a favorite or like control on the interface. Fig. 9 As shown, the control can be located in the upper right corner of the interface and is heart-shaped. The preference record can be used as the user's historical behavior data to generate the first information in the subsequent generation, thereby improving the accuracy of the first information in indicating the user's interest tendency.
[0150] In other embodiments, after step S702, the recommended content may also be stored in the user's historical browsing record, and the historical browsing record is used to generate the first information.
[0151] In practice, each generated recommendation content can be recorded in the user's historical browsing history. The historical browsing history can be used as the user's historical behavior data to generate the first information later, thereby improving the accuracy of the first information in indicating the user's interest tendency.
[0152] In other embodiments, in response to a second operation of the user on the recommended content on the interactive interface, new recommended content corresponding to the target image may be displayed to the user. For example, the second operation may be a user clicking a control on the interface for regenerating recommended content, or a user swiping right on the interface. After receiving the instruction triggered by the second operation, the client may request the remote server to regenerate the recommended content, and then receive and display the regenerated recommended content.
[0153] Figure 7The solution provided by the corresponding embodiment obtains the user's target image to display recommended content that the user may be interested in. The usage threshold is very low, which ensures the fairness of the use of network services. The generated recommended content is more in line with the user's interests and actual scenarios, thereby improving the user experience.
[0154] Fig.10 Schematic diagram of the structure of the recommendation device based on the large language model in the embodiment of this specification. The device can be applied to any device, platform or device cluster with computing and processing capabilities. The device includes: The subject query module 101 is configured to obtain at least one subject queried by a user, wherein the at least one subject is identified from a target image; The information acquisition module 102 is configured to acquire first information of the user, where the first information is used to indicate the interest tendency of the user; A target determination module 103 is configured to determine a target subject in at least one subject based on the first information; The content generation module 104 is configured to generate at least one recommended content for the user based on the target subject by using a large language model.
[0155] In some embodiments, the subject query module 101 is specifically configured to: obtain the associated information of the user; and generate the first information of the user based on the associated information through a large language model.
[0156] In some embodiments, the user's associated information includes at least one of the following information: the user's attribute information, the user's historical behavior data, the user's location information, and the user's current time information.
[0157] In some embodiments, the target determination module 103 is specifically configured to: determine a target subject in at least one subject and a vocabulary associated with the target subject based on the first information; the content generation module 104 is specifically configured to: generate at least one recommended content for the user based on the target subject and the vocabulary associated with the target subject through a large language model.
[0158] In some embodiments, the target determination module 103, when configured to determine a target subject in at least one subject and vocabulary associated with the target subject based on the first information, is specifically configured to: retrieve associated information and / or context information of at least one subject based on a preset corpus, the context information including associated words expanded based on the subject and / or associated information; generate vocabulary associated with the subject based on the associated information, at least one subject and the context information through a first multimodal large model; determine the target subject and vocabulary associated with the target subject from at least one subject and vocabulary associated with the subject based on the first information.
[0159] In some embodiments, the target determination module 103, when configured to determine the target subject and the vocabulary associated with the target subject in at least one subject and the vocabulary associated with the subject based on the first information, is specifically configured to: semantically encode the first information to obtain a first semantic vector; semantically encode each subject and the vocabulary associated with the subject to obtain a second semantic vector; based on the similarity between the first semantic vector and the second semantic vector, screen the second semantic vector to determine the target subject and the vocabulary associated with the target subject in at least one subject.
[0160] In some embodiments, the vocabulary includes hypernyms of the target subject.
[0161] In some embodiments, the large language model is a spatiotemporal large language model, which is fine-tuned in the following manner: obtaining a spatiotemporal dataset and labels of sample users, the spatiotemporal dataset including: associated information of sample users, and labels including services used by sample users; using the spatiotemporal large language model, based on the spatiotemporal dataset, predicting the services used by sample users; and adjusting the network parameters of the spatiotemporal large language model based on the prediction results and labels.
[0162] In some embodiments, the subject query module 101 is specifically configured to: obtain a target image input by a user, analyze the structure in the target image based on the target image through a second multimodal large model, and generate a scene graph, wherein the scene graph includes at least one subject and the relationship between the subjects.
[0163] In some embodiments, the target image is an image acquired by capturing the scene in which the user is currently located.
[0164] In some embodiments, the content generation module 104 is specifically configured to generate at least one recommended content for the user based on the target subject and the associated information through a large language model.
[0165] Fig.11 This is a schematic diagram of the structure of another recommendation device based on a large language model in the embodiment of this specification. The device can be applied to any device, platform or device cluster with computing and processing capabilities. The device includes:
[0166] The image acquisition module 111 is configured to acquire a target image of a user;
[0167] The interface display module 112 is configured to display recommended content corresponding to the target image to the user. The recommended content is generated for the user by a large language model based on a target subject. The target subject is determined from at least one subject based on first information. The first information is used to indicate the user's interest tendency. The at least one subject is identified from the target image.
[0168] In some embodiments, the device also includes: a direction determination module (not shown in the figure), which is configured to obtain a content generation direction, the content generation direction is used to indicate the information type of the recommended content, and the recommended content is generated for the user by a large language model based on the target subject and the content generation direction.
[0169] In some embodiments, the device also includes: a preference recording module (not shown in the figure), which is configured to, after displaying the recommended content corresponding to the target image to the user, store the recommended content in the user's preference record in response to the user's first operation on the recommended content on the interactive interface, the first operation is used to indicate the user's interest in the recommended content, and the preference record is used to generate the first information.
[0170] In some embodiments, the device further includes: a history record module (not shown in the figure), which is configured to store the recommended content into the user's historical browsing record after displaying the recommended content corresponding to the target image to the user, and the historical browsing record is used to generate the first information.
[0171] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the following Figure 3 or Figure 7 Describe the method.
[0172] The embodiment of the present specification also provides a computing device, including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the following is implemented: Figure 3 or Figure 7 Describe the method.
[0173] The embodiments of the present specification also provide a computer program product, including a computer program / instruction, which is executed by a processor to implement the following Figure 3 or Figure 7 Describe the steps of the method.
[0174] It is understandable that before or when using the technical solutions of the various embodiments in this specification, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner in accordance with relevant laws and regulations, and the user's authorization will be obtained.
[0175] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to the electronic device, application, server, storage medium or other software or hardware that performs the operation of the technical solution of this specification according to the prompt message.
[0176] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0177] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation method of this specification. Other methods that meet relevant laws and regulations may also be applied to the implementation method of this specification.
[0178] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the multiple embodiments disclosed in this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0179] In some cases, the actions or steps described in the claims may be performed in a different order than in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0180] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not used to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. A recommendation method based on a large language model, the method comprising: Acquire at least one subject queried by a user, wherein the at least one subject is identified from a target image; Acquire first information of the user, where the first information is used to indicate the interest tendency of the user; Based on the first information, determining a target subject among the at least one subject; At least one recommended content is generated for the user based on the target subject through the large language model.
2. The method according to claim 1, wherein: The obtaining the first information of the user includes: Obtaining the associated information of the user; The first information of the user is generated based on the associated information through the large language model.
3. The method according to claim 2, wherein: The user's associated information includes at least one of the following information: attribute information of the user, historical behavior data of the user, location information of the user, and current time information of the user.
4. The method according to claim 1, wherein: The determining, based on the first information, a target subject in the at least one subject includes: Based on the first information, determining a target subject in the at least one subject and a vocabulary associated with the target subject; The generating, by using the large language model and based on the target subject, at least one recommended content for the user comprises: At least one recommended content for the user is generated through the large language model based on the target subject and the vocabulary associated with the target subject.
5. The method according to claim 4, wherein: The determining, based on the first information, a target subject in the at least one subject and a vocabulary associated with the target subject includes: Based on a preset corpus, retrieve the associated information and / or the context information of the at least one subject, wherein the context information includes associated words expanded based on the subject and / or the associated information; Generate, by means of a first multimodal macromodel, a vocabulary associated with the subject based on the association information, the at least one subject, and the context information; Based on the first information, the target subject and the vocabulary associated with the target subject are determined from among the at least one subject and the vocabulary associated with the subject.
6. The method according to claim 5, wherein: The determining, based on the first information, the target subject and the vocabulary associated with the target subject from among the at least one subject and the vocabulary associated with the subject, includes: Performing semantic encoding on the first information to obtain a first semantic vector; Performing semantic coding on each of the subjects and words associated with the subjects to obtain a second semantic vector; Based on the similarity between the first semantic vector and the second semantic vector, the second semantic vector is screened to determine a target subject in the at least one subject and words associated with the target subject.
7. The method according to claim 4, wherein: The vocabulary includes hypernyms of the target subject.
8. The method according to claim 2, wherein: The large language model is a spatiotemporal large language model, which is fine-tuned in the following way: Acquire a spatiotemporal data set and a label of a sample user, wherein the spatiotemporal data set includes: associated information of the sample user, and the label includes a service used by the sample user; Predicting the services used by the sample user based on the spatiotemporal data set by using the spatiotemporal large language model; Based on the prediction result and the label, a network parameter of the spatiotemporal large language model is adjusted.
9. The method according to claim 1, wherein: The obtaining of at least one subject of the user query includes: The target image input by the user is acquired, and a structure in the target image is analyzed based on the target image through a second multimodal large model to generate a scene graph, wherein the scene graph includes the at least one subject and the relationship between the subjects.
10. The method according to claim 1, wherein: The target image is an image obtained by capturing the scene in which the user is currently located.
11. The method according to claim 1, wherein: The generating, by using the large language model and based on the target subject, at least one recommended content for the user comprises: At least one recommended content for the user is generated based on the target subject and the associated information through the large language model.
12. A recommendation method based on a large language model, the method comprising: Get the user's target image; Recommended content corresponding to the target image is displayed to the user, the recommended content is generated by the large language model for the user based on a target subject, the target subject is determined from at least one subject based on first information, the first information is used to indicate the interest tendency of the user, and the at least one subject is identified from the target image.
13. The method according to claim 12, wherein: The method further comprises: A content generation direction is obtained, where the content generation direction is used to indicate an information type of the recommended content, and the recommended content is generated for the user by the large language model based on a target subject and the content generation direction.
14. The method according to claim 12, wherein: After displaying the recommended content corresponding to the target image to the user, the method further includes: In response to a first operation of the user on the recommended content on an interactive interface, the recommended content is stored in a preference record of the user, the first operation being used to indicate the user's interest in the recommended content, and the preference record being used to generate the first information.
15. The method according to claim 12, wherein: After displaying the recommended content corresponding to the target image to the user, the method further includes: The recommended content is stored in the user's historical browsing record, and the historical browsing record is used to generate the first information.
16. A recommendation device based on a large language model, the method comprising: A subject query module, configured to obtain at least one subject queried by a user, wherein the at least one subject is identified from a target image; An information acquisition module, configured to acquire first information of the user, where the first information is used to indicate the interest tendency of the user; a target determination module, configured to determine a target subject among the at least one subject based on the first information; The content generation module is configured to generate at least one recommended content for the user based on the target subject through the large language model.
17. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 15 is implemented.
Citation Information
Patent Citations
Information recommendation method and device
CN108334528A
Information recommendation method and device, front-end implementation method and device, equipment and storage medium
CN110020156A
Method and device for feeding back search result and storage medium
CN114969408A
Cross-domain news recommendation method based on knowledge base enhancement
CN115640462A
Commodity recommendation method and device, equipment and medium
CN116029793A
Cited By
Recommendation system-oriented spatio-temporal data enhancement method
CN120179926A
Image processing method and device based on artificial intelligence and electronic equipment
CN120783094A
Method, device and system for providing personalized content recommendation service using multi modal and multi context based on artificial intelligence model
KR102994451B1