Personalized recommendation method and device, computer equipment and storage medium
By fusing user facial features and eye-tracking interaction data with object database information to construct a joint feature vector, the problem of capturing users' real-time interests and preferences in existing technologies is solved, enabling personalized and instant recommendations and improving user experience.
Patent Information
- Application Number
- CN202511962577.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies struggle to capture users' real-time changing interests and preferences, and cannot fully mine the deep semantic information contained in multimodal data, resulting in limited recommendation performance in scenarios such as sparse historical interaction data, interest drift, or cold start.
By acquiring the user's facial feature data and eye-tracking interaction data at the current moment, and fusing them with multimodal data information in the object library, a joint feature vector is constructed, and personalized recommendations are made based on interest scores.
It enables accurate interpretation of users' immediate interests and potential preferences, enhances the personalization and immediacy of recommendations, strengthens user immersion and satisfaction, and improves information matching efficiency.
Smart Images

Figure CN121707682A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of cloud computing, and more particularly to a personalized recommendation method, apparatus, computer device, and storage medium. Background Technology
[0002] Personalized recommendation strategies can be customized for users by utilizing interaction data generated when users interact with multiple items or content.
[0003] In related technologies, collaborative filtering, matrix factorization, or deep learning models can be used to establish and analyze the relationship between users' historical interaction data and the attribute information of products or content, as well as to calculate the matching degree between users and objects, thereby recommending products or content similar to users' past interests. However, these methods mainly rely on historical interaction data, making it difficult to capture users' real-time changing interests and preferences, and also unable to fully mine the deep semantic information contained in multimodal data. This results in limited recommendation effectiveness in scenarios with sparse historical interaction data, interest drift, or cold start. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide a personalized recommendation method, apparatus, computer device, and storage medium to solve the problems existing in the related art.
[0005] A first aspect of this disclosure provides a personalized recommendation method, the method comprising: acquiring facial feature data and eye-tracking interaction data of a user at a current moment, as well as multimodal data information of each object in an object library; determining multiple candidate objects from the object library based on the user's browsing behavior at the current moment; for each candidate object, fusing the multimodal data information of the candidate object with the facial feature data and eye-tracking interaction data to obtain a joint feature vector of the candidate object; determining an interest score corresponding to each object based on the joint feature vector of each candidate object; and recommending objects to the user based on the interest scores of each object.
[0006] A second aspect of this disclosure provides a personalized recommendation apparatus applied to the personalized recommendation selection method of the first aspect. The apparatus includes: an acquisition module for acquiring facial feature data and eye-tracking interaction data of a user at a current moment, as well as multimodal data information of each object in an object library; a determination module for determining multiple candidate objects from the object library based on the user's browsing behavior at the current moment; a fusion module for fusing the multimodal data information of each candidate object with the facial feature data and eye-tracking interaction data to obtain a joint feature vector of the candidate object; and a recommendation module for determining an interest score corresponding to each object based on the joint feature vector of each candidate object, and recommending objects to the user based on the interest scores of each object.
[0007] A third aspect of this disclosure provides a computer device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the personalized recommendation method described above.
[0008] A fourth aspect of this disclosure provides a computer-readable storage medium having a computer program / operation instructions stored thereon, which, when executed by a processor, implement the steps of the personalized recommendation method described above.
[0009] According to a fifth aspect of this disclosure, a computer program product is provided that, when executed by a processor, implements the steps of the personalized recommendation method described above.
[0010] The at least one technical solution adopted in this disclosure can achieve the following beneficial effects: by acquiring the user's facial feature data and eye-tracking interaction data at the current moment, as well as the multimodal data information of each object in the object library; based on the user's browsing behavior at the current moment, multiple candidate objects are determined from the object library; for each candidate object, the multimodal data information of the candidate object is fused with the facial feature data and eye-tracking interaction data to obtain the joint feature vector of the candidate object; based on the joint feature vector corresponding to each candidate object, the interest score corresponding to each object is determined, and recommendations are made to the user based on the interest score of each object. In this way, this disclosure can capture the user's facial features and eye-tracking interaction data in real time and dynamically fuse them with the multimodal information in the object library to construct a joint feature vector that can simultaneously reflect the user's physiological response, visual attention, and the characteristics of the object itself. This fusion mechanism can more accurately interpret the user's immediate interests and potential preferences during the browsing process, and based on the interest score of the joint feature vector, realize the transformation from passive screening to active perception in the recommendation method, thereby improving the personalization and immediacy of the recommendation, enhancing the user's immersion and satisfaction, and ultimately effectively improving the information matching efficiency and user experience. Attached Figure Description
[0011] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 A schematic diagram of the architecture of a recommendation list generation method in a related art provided in an embodiment of this disclosure; Figure 2A flowchart illustrating a personalized recommendation method provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of a personalized recommendation system provided in an embodiment of the present disclosure; Figure 4 This is a schematic diagram of the structure of a personalized recommendation device provided in an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present disclosure; Figure 7 A schematic diagram of a computer program product provided according to an embodiment of this disclosure. Detailed Implementation
[0013] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0014] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0015] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0018] With the deep integration of artificial intelligence and computing networks, various websites and smart terminals have permeated all aspects of daily life and entertainment. In this process, users continuously generate massive amounts of diverse historical interaction data, widely stored across different platforms and devices. This historical interaction data not only includes rich media content such as text, images, audio, and video, but also covers behavioral sequences such as clicks, browsing, saving, and sharing, as well as contextual information such as social relationships and geographical location, collectively forming a multimodal, high-dimensional interaction graph. Based on this, personalized recommendation strategies can be customized for users.
[0019] In related technologies, when customizing personalized recommendation strategies for users based on their historical interaction data, the following steps can be taken: First, collect the user's historical interaction data and perform adaptive model training. Then, construct a comprehensive and three-dimensional user profile through multimodal fusion technology. Next, generate a personalized service strategy specific to the user based on this user profile to continuously serve the user. Furthermore, the user profile can be continuously optimized based on subsequent user interaction feedback. The core advantage of this technology compared to traditional methods lies in its ability to utilize a large multimodal model to uniformly represent and fuse various user data such as text, voice, images, and video, thereby significantly improving the completeness and accuracy of the user profile and providing a richer feature foundation for subsequent personalized services.
[0020] Figure 1 This is a schematic diagram illustrating the architecture of a recommendation list generation method in a related art, provided as an embodiment of this disclosure. For example... Figure 1 As shown below, the relevant technologies for products can be explained. These technologies can start from a massive collection of products, first using a candidate generation model to filter out hundreds of candidate products from millions of products based on user profiles, and then using a rating and ranking model to further filter and rank these hundreds of candidate products, finally outputting a recommendation list of dozens of products to the user.
[0021] Related technologies, such as collaborative filtering, matrix factorization, or deep learning models, combine users' historical interaction data with the attribute information of products or content to establish and analyze the relationship between the two, and calculate the matching degree between users and objects, thereby recommending products or content similar to users' past interests. However, these methods mainly rely on historical interaction data, making it difficult to capture users' real-time changing interests and preferences, and also unable to fully mine the deep semantic information contained in multimodal data. This results in limited recommendation effectiveness in scenarios with sparse historical interaction data, interest drift, or cold start.
[0022] Meanwhile, with sparse historical interaction data, related technologies struggle to find objects or content similar to users' historical behaviors, and the authenticity and reliability of the historical interaction data itself are also difficult to guarantee. Furthermore, in a multimodal data environment, relying solely on manually constructed features, simple feature crossings, or using only a single-modal network model often fails to fully extract the high-level semantic information hidden behind objects, and cannot construct a comprehensive and three-dimensional user profile. On the other hand, since network models are mostly trained or fine-tuned based on static historical interaction data or limited new samples, the system often cannot respond promptly when user interests or contextual information change dynamically, leading to a deviation between recommended content and the user's current expectations.
[0023] To address the aforementioned issues, this disclosure provides a personalized recommendation method. It acquires a user's current facial feature data and eye-tracking interaction data, as well as multimodal data information of objects in an object library. Based on the user's browsing behavior at the current moment, multiple candidate objects are identified from the object library. For each candidate object, the multimodal data information is fused with the facial feature data and eye-tracking interaction data to obtain a joint feature vector. An interest score is determined for each object based on its joint feature vector, and recommendations are made to the user based on each object's interest score. Thus, this disclosure can capture user facial features and eye-tracking interaction data in real time and dynamically fuse them with multimodal information from the object library to construct a joint feature vector that simultaneously reflects the user's physiological responses, visual attention, and the object's inherent characteristics. This fusion mechanism can more accurately interpret the user's immediate interests and potential preferences during browsing, and based on the interest score of the joint feature vector, it achieves a shift from passive filtering to proactive perception in recommendation, thereby improving the personalization and immediacy of recommendations, enhancing user immersion and satisfaction, and ultimately effectively improving information matching efficiency and user experience.
[0024] The personalized recommendation method provided in this disclosure can be executed by a terminal or by a chip applied to the terminal.
[0025] For example, the aforementioned terminals may include one or more of the following: mobile phones, tablets, wearable devices, in-vehicle devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, handheld computers (PDAs), and wearable devices based on augmented reality (AR) and / or virtual reality (VR) technologies. They may also include, but are not limited to, remote control devices, wearable devices, streetlights, home appliances, and other smart terminals. This disclosure does not impose specific limitations on these aspects.
[0026] Figure 2 This is a flowchart illustrating a personalized recommendation method provided in one embodiment of this disclosure. Figure 2 As shown, personalized recommendation methods include: S201: Obtain the user's current facial feature data and eye-tracking interaction data, as well as the multimodal data information of each object in the object library. Here, the objects displayed on the terminal front-end interface all come from the object library and are a visual representation of some objects in the object library. Their multimodal data information is consistent with the multimodal data information of the corresponding objects in the object library.
[0027] In some embodiments, a momentary facial image of the user can be acquired through an image acquisition device on the terminal (such as a camera installed on the terminal), and the Deepface computer vision library can be invoked to analyze and process the facial image to obtain facial feature data. The facial feature data may include at least one of the following: age, gender, ethnicity, and facial expression.
[0028] Next, the acquired facial feature data undergoes structured preprocessing to convert it into facial feature vectors that can be used as input to the model. The specific scheme is as follows: Age: Can be divided into common age groups and assigned corresponding code values. For example, code 1: age < 18 years old; code 18: 18 years old ≤ age ≤ 24 years old; code 25: 25 years old ≤ age ≤ 34 years old; code 35: 35 years old ≤ age ≤ 44 years old; code 45: 45 years old ≤ age ≤ 49 years old; code 50: 50 years old ≤ age ≤ 55 years old; code 56: age ≥ 56 years old.
[0029] Gender is categorized as either male or female.
[0030] Race is categorized into Asian, white, Middle Eastern, Indian, Latino, black, etc.
[0031] The emoji tags include at least one of anger (-2), fear (-1), neutral (0), sadness (-1), disgust (-2), happiness (2), and surprise (1).
[0032] In some embodiments, user eye-tracking interaction data can also be obtained through facial images. Specifically, this includes processing the eye video stream using the OpenFace computer vision library to achieve real-time detection and tracking of eye gaze. Specifically, by calculating and analyzing the geometric relationship between the gaze direction vector and the visual positions of objects on the current page, the user's gaze-object interaction behavior (such as fixation or saccade) can be determined, and the start time and gaze duration of each interaction behavior can be recorded, thereby constructing eye-tracking interaction data containing temporal information. Here, "object" refers to any interactive or fixational visual element contained in the terminal's front-end display interface, including but not limited to images, icons, buttons, menu items, text paragraphs, video playback windows, floating prompts, and animated elements; and the video stream of the eye region is a dynamic image sequence that is captured in real-time and focused on the eye region from a continuously acquired facial image sequence using face detection and cropping technology.
[0033] During the acquisition of eye-tracking interaction data, the user's facial feature data can be matched with the facial feature database of registered users to determine the user's identity information. Specifically: if no existing record can be matched, the user is identified as a new user, and the eye-tracking interaction data collected in real time can be directly initialized and used to construct the initial browsing behavior list for the new user; if a match is successful, the user is identified as an existing user, and the newly generated eye-tracking interaction data can be merged chronologically with the user's identity-associated, stored historical browsing behavior list to form an updated, chronologically ordered browsing behavior list for subsequent interest modeling and dynamic recommendations.
[0034] In some embodiments, the multimodal data information of each object in the object library is mainly obtained through the following methods: When an object is added to the library, the system automatically or the creator pre-extracts and structures its original visual, textual, and audio data—for example, by extracting image features through computer vision models, generating text summaries and tags using natural language processing technology, performing speech recognition and sentiment analysis on videos or audio, and integrating metadata such as source, author, and publication time. These multimodal features are encoded and uniformly stored in the object feature library, forming a complete digital profile for each object, which can be called and matched in real time during recommendations.
[0035] S202, Based on the user's browsing behavior at the current moment, identify multiple candidate objects from the object library.
[0036] In some embodiments, candidate objects can be determined from the object library based on the user's browsing behavior list: taking multimodal data information mined by multimodal large model as the core, combining user facial feature data and eye movement interaction data for similarity matching, referring to the user's interest tendencies in the browsing behavior list, prioritizing the recall of categories semantically related to objects of high interest in the user's history, excluding object types that are explicitly excluded, and filtering out hundreds of potential candidate objects.
[0037] Here, multiple candidate objects are identified at the current moment, primarily based on the user's real-time browsing context and immediate interaction cues. Specifically, the coordinates of the user's gaze point are first captured and mapped to a specific area on the screen or in space. All visible potential objects within that area, such as images, buttons, and product cards, are then used as the first-level candidate. Simultaneously, the current page content, application interface, or environmental context—such as an e-commerce listing page, video stream, or article being viewed—is used to create a basic candidate pool of the main objects within that scenario. Furthermore, the user's brief interaction history over the past few seconds is considered, such as quickly scanned items, mouse hover sequences, or gesture directions, to extract objects that may generate sustained interest. These elements together form a dynamic, real-time candidate object set. The entire process is completed within milliseconds, ensuring that subsequent analysis focuses on the most relevant targets.
[0038] S203, for each candidate object, the multimodal data information of the candidate object is fused with facial feature data and eye-tracking interaction data to obtain the joint feature vector of the candidate object.
[0039] In some embodiments, this disclosure can employ a Meta-Transformer architecture to construct a unified multimodal learning framework. This framework deeply fuses multimodal data information with the facial feature data and eye-tracking interaction data obtained in step S201 to generate a joint feature vector. The core design of this framework lies in learning a shared parameter space that contains the intersection of the independent parameter spaces of each data modality. Finally, a unified mapping function is used to achieve the joint representation of multimodal information, providing data support for subsequent recommendation tasks. And assume that the following conditions are met: Where θ1 represents the independent parameter space of the first mode; θ2 represents the independent parameter space of the second mode; θ n This represents the independent parameter space of the nth mode.
[0040] Specifically, within this framework, assuming there are n different data modes to process, for the i-th mode, the input mode data can be denoted as X. i(Where i = 1, 2, ..., j). For ease of explanation, this embodiment of the disclosure uses j = 4 as an example for related description. When j = 4, X1 can include X1, X2, X3, and X4. X1 can be image modal data (object side), such as the image pixel matrix of the object; X2 can be text modal data (object side), such as the text description sequence of the object; X3 can be facial feature modal data (user side), such as the encoded user facial feature vector; X4 can be eye movement temporal behavior modal data (user side), such as the user's eye movement interaction data.
[0041] During the supervised pre-training phase, each modality of data X... i It can correspond to a real label or target representation Y. i These are used to guide model learning. For example, Y1 corresponding to X1 could be the category label of the object; Y2 corresponding to X2 could be the sentiment polarity label of the description; Y3 corresponding to X3 could be a user demographic attribute classification based on features; and Y4 corresponding to X4 could be a user interest intensity regression target based on behavior. It should be understood that in practical applications, information such as the object's voice description can also be added.
[0042] Based on this, X1, X2, X3, X4 and their corresponding Y1, Y2, Y3, Y4 can be input into the mapping function F, and the shared parameter space can be jointly optimized using cross-entropy loss and mean squared error loss. This allows the mapping function F to gradually learn the intrinsic relationship between different modalities and user preferences. The formula for the mapping function F can be expressed as:
[0043] Where z represents the joint feature vector, and x represents the input data. The shared parameters of the mapping function are represented by , and y represents the true label of the input data. X represents the predicted value, which includes multimodal data information of the candidate object, facial feature data, and eye-tracking interaction data.
[0044] After optimization, F can directly receive four types of raw data: X1, X2, X3, and X4. After transformation by a unified labeler and processing by a shared Transformer encoder, it outputs a joint feature vector that fuses the semantic information of the object and the user's features. This joint feature vector not only contains a multimodal comprehensive representation of the object, but also strengthens the correlation between the user's facial feature data, eye-tracking interaction data, and the object through an intermodal cross-attention mechanism. This provides high-dimensional feature support for the dynamic improvement of user profiles and provides quantitative basis for interest score calculation and candidate object selection, realizing a closed-loop technical link from multimodal data input to recommendation decision-making.
[0045] S204. Determine the interest score for each candidate object based on the joint feature vector, and make recommendations to the user based on the interest score of each object.
[0046] In some embodiments, for any candidate object, the interaction state between the user and the candidate object is determined based on eye-tracking interaction data. If there is interaction between the user and the candidate object, the duration of the user's gaze on the candidate object and the total gaze duration of all candidate objects are determined based on the joint feature vector. The user's explicit interest score for the candidate object is then determined based on the duration of the user's gaze on the candidate object and the total gaze duration of all candidate objects. Here, the duration of gaze fixation is the duration the user's gaze lingers on the candidate object; the total gaze duration of all candidate objects is the total duration the user's gaze lingers on all candidate objects.
[0047] In the absence of interaction between the user and the candidate objects, the stability of the user's switching time and gaze direction changes between different candidate objects is determined based on the joint feature vector. Furthermore, the user's interest score for the candidate objects is determined based on the stability of the user's switching time and gaze direction changes between different candidate objects, along with facial feature data. Here, switching time refers to the time spent switching the user's gaze between different candidate objects, reflecting the fluency of their decision-making and the efficiency of their interest transfer; the stability of gaze direction changes refers to the synergy and smoothness of the user's gaze direction changes during interaction, reflecting the depth of their attention and the stability of their cognitive load.
[0048] In some embodiments, the interest coefficient corresponding to the browsing behavior can be calculated by capturing the user's eye gaze direction, the interaction record of the currently browsed page, and facial expression tags. The specific process may include: after the process starts, the user's facial image is first acquired through the terminal's image acquisition device. Simultaneously, the Deepface computer vision library is called to analyze the facial image to extract facial feature data, and data such as age and gender are preprocessed in a structured manner to convert them into facial feature vectors. At the same time, the OpenFace computer vision library is called to process the eye video stream extracted from the facial image sequence, detecting and tracking the gaze direction in real time. By analyzing the gaze direction vector and the geometric relationship between the gaze direction vector and the objects on the current page, interaction behaviors such as fixation and saccades are determined. The start time and gaze duration of each browsing behavior are recorded to construct eye-tracking interaction data containing temporal information.
[0049] During this process, the collected facial feature data can be matched with the facial feature database of registered users to determine the user's identity. For new users, the browsing behavior list is initialized directly with the current eye-tracking interaction data, while for existing users, the newly collected eye-tracking interaction data is merged with the historical browsing behavior list in chronological order to form a complete browsing behavior list.
[0050] Subsequently, based on eye-tracking interaction data, it is determined whether the user has a clear interactive behavior (such as clicking, long-pressing, etc.) with the currently viewed object. When the user's gaze has an interactive behavior record, the ratio of the gaze duration of the current candidate to the total gaze duration of all candidate objects is used as the explicit interest score (value range [0,1]). When the user's gaze has no interactive behavior record, the interaction time or dwell time between the gaze and the object is normalized by combining the consistency of the gaze direction and the dwell interval. Then, the user's facial expression label during the current browsing is associated with it. Disgust corresponds to -5, anger corresponds to -2, fear corresponds to -1, neutral corresponds to 0, surprise corresponds to 3, sadness corresponds to 4, and happiness corresponds to 5. Finally, the implicit interest coefficient (POI) for this scenario is generated.
[0051] The formula for calculating explicit interest rating is as follows:
[0052] Among them, P i t represents the explicit interest score of the i-th candidate; i This indicates the duration of the user's gaze on the candidate object. This represents the total gaze duration of all candidate objects, where n represents the total number of candidate objects. This represents a set of candidate objects, where i represents the identifier of a candidate object.
[0053] The formula for calculating implicit interest scores is as follows:
[0054] Where POI represents the interest coefficient; 0 represents no interest; and other represents excluding. Other factors that influence interest ratings besides the candidate's popularity among all candidates; Represents the current facial expression tag; 'a' represents the left eye's gaze direction vector; 'b' represents the right eye's gaze direction vector. The vector representing the gaze direction of the left eye in an image frame at interval t; s represents the gaze direction vector of the left eye in an image frame at interval t; i It is a normalized scalar value, typically ranging from [0,1]. It comprehensively reflects the consistency between the duration of the user's gaze on the candidate object and the direction of the gaze. The gaze duration is the length of time the user's gaze lingers on the candidate object.
[0055] The consistency of gaze direction can be obtained through the cosine similarity of the gaze direction vectors:
[0056] In some embodiments, after obtaining the corresponding interest score for each candidate object, the interest score can be adjusted based on the matching degree between the user's facial expressions and the object's semantics during interaction with the candidate object. Specifically, the matching degree between the user's facial expressions and the object's semantics during interaction with the candidate object can be multiplied by the interest score to obtain the comprehensive interest score for each candidate object. Finally, all candidate objects are sorted according to the comprehensive interest score and then recommended to the user. During the recommendation, all candidate objects can be recommended directly, or the candidate object with the highest comprehensive interest score can be recommended.
[0057] Figure 3 This is a schematic diagram of the structure of a personalized recommendation system provided in one embodiment of this disclosure. Figure 3 As shown, taking the personalized recommendation scenario of an e-commerce platform as an example, the complete process is as follows: First, the personalized recommendation system can collect users' facial feature data and eye-tracking interaction data in real time through the camera set on the terminal. Combined with the user's historical preferences stored in the basic model and the product information in the item database, millions of products are input into the candidate model. The candidate model quickly filters out hundreds of candidate products that match the user's basic profile, and then passes these hundreds of candidate products to the ranking model. The ranking model retrieves the interest scores of the current candidate products, performs preliminary ranking of the candidate products, and filters out dozens of highly interested products. Finally, the re-ranking model combines real-time interest priority, product diversity, and other rules to output the final recommendation result, completing a personalized recommendation based on the user's real-time behavior and expressions.
[0058] It should be noted that the interest score calculation in this embodiment does not directly rely on the original eye-tracking or facial data. Instead, it uses the joint feature vector z as the core input, and the joint feature vector z as a high-level semantic representation after deep fusion of a multimodal large model. A scoring model (such as a neural network) is then used to generate the final interest score. Specifically, t in the explicit interest score formula... i and The raw data, including transition time and gaze stability from implicit interest scoring, are first used as one of the input features to construct the joint feature vector z. Subsequently, the joint feature vector z is input into the scoring model, which, through its learned complex nonlinear relationships, outputs an interest score that accurately reflects the user-object correlation. This design ensures that the fused multimodal features directly dominate the scoring decision, thereby fully leveraging the fusion value of step S203 and avoiding the problem of the fusion step being rendered ineffective in the scoring process.
[0059] Figure 3The data acquisition module directly corresponds to step S201, responsible for collecting user facial features and eye-tracking interaction data in real time using sensors such as cameras, while simultaneously retrieving multimodal information of products from the object library. The candidate model corresponds to step S202, which quickly filters hundreds of candidate objects from a massive product library based on the user's real-time browsing behavior and historical preferences. The filtering process can be found in [reference needed]. Figure 1 The relevant technical architecture is shown below; the ranking model serves as the core of the system, corresponding to the scoring stages in steps S203 and S204. It uses fusion technologies such as Meta-Transformer to deeply fuse the multimodal data of candidate objects with the user's real-time physiological data, generating a joint feature vector, and calculating explicit or implicit interest scores based on this vector. The logical basis of this fusion and scoring process is as follows: Figure 2 The flowchart shown illustrates the method. The final re-ranking model corresponds to the recommendation action in step S204. Based on interest scoring, business rules are introduced for final adjustments, outputting an optimized recommendation list, thus completing the transformation from "passive filtering" to "active perception." The entire system achieves engineering implementation of the method steps through modular design, ensuring the personalization and immediacy of recommendations.
[0060] This method acquires the user's current facial feature data and eye-tracking interaction data, as well as multimodal data information of objects in the object library. Based on the user's browsing behavior at the current moment, multiple candidate objects are identified from the object library. For each candidate object, the multimodal data information of the candidate object is fused with the facial feature data and eye-tracking interaction data to obtain a joint feature vector of the candidate object. Based on the joint feature vector of each candidate object, an interest score is determined for each object, and recommendations are made to the user based on the interest score of each object. In this way, this embodiment of the disclosure can capture the user's facial features and eye-tracking interaction data in real time and dynamically fuse them with multimodal information in the object library to construct a joint feature vector that can simultaneously reflect the user's physiological response, visual attention, and the characteristics of the object itself. This fusion mechanism can more accurately interpret the user's immediate interests and potential preferences during the browsing process, and based on the interest score of the joint feature vector, it realizes the transformation from passive screening to active perception recommendation method, thereby improving the personalization and timeliness of recommendations, enhancing user immersion and satisfaction, and ultimately effectively improving information matching efficiency and user experience.
[0061] The foregoing primarily describes the solutions provided by the embodiments of this disclosure from the perspective of the server. It is understood that, in order to implement the above functions, the server includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0062] This disclosure embodiment can divide the server into functional units according to the above method example. For example, it can divide each function into separate functional modules, or it can integrate two or more functions into one management module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this disclosure embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0063] By dividing each functional module according to its corresponding function, an exemplary embodiment of this disclosure provides a personalized recommendation device, which can be a server or a chip applied to a server. Figure 4 This is a schematic diagram of a personalized recommendation device provided in one embodiment of the present disclosure. Figure 4 As shown, the personalized recommendation device 400 includes: The acquisition module 401 is used to acquire the user's facial feature data and eye-tracking interaction data at the current moment, as well as the multimodal data information of each object in the object library.
[0064] The determination module 402 is used to determine multiple candidate objects from the object library based on the user's browsing behavior at the current moment.
[0065] The fusion module 403 is used to fuse the multimodal data information of the candidate object with the facial feature data and the eye-tracking interaction data for each candidate object to obtain the joint feature vector of the candidate object; The recommendation module 404 is used to determine the interest score corresponding to each object based on the joint feature vector corresponding to each candidate object, and to make recommendations to the user based on the interest score of each object.
[0066] In one alternative approach, determining the interest score corresponding to each object based on the joint feature vector corresponding to each candidate object includes: for any candidate object, determining the interaction state between the user and the candidate object based on the joint feature vector; if there is an interaction between the user and the candidate object, determining the user's gaze duration on the candidate object and the total gaze duration of all candidate objects based on the eye-tracking interaction data; and determining the user's explicit interest score for the candidate object based on the user's gaze duration on the candidate object and the total gaze duration of all candidate objects.
[0067] In one alternative approach, the formula for calculating the explicit interest score is:
[0068] Among them, P i t represents the explicit interest score of the i-th candidate; i This indicates the duration of the current user's gaze on the candidate object; This represents the total gaze duration of all candidate objects, where n represents the total number of candidate objects. This represents a set of candidate objects, where i represents the identifier of a candidate object.
[0069] In one alternative approach, determining the interest score for each object based on the joint feature vector corresponding to each candidate object includes: In the absence of any interaction between the user and the candidate object, the stability of the user's switching time and binocular gaze direction changes between different candidate objects is determined based on the joint feature vector, and the user's interest score for the candidate object is determined based on the stability of the user's switching time and binocular gaze direction changes between different candidate objects and facial feature data.
[0070] In one alternative approach, fusing the multimodal data information of the candidate object with the facial feature data and the eye-tracking interaction data to obtain a joint feature vector of the candidate object includes: The multimodal data information of the candidate object, the facial feature data, and the eye-tracking interaction data are input into a preset mapping function to obtain the joint feature vector of the candidate object.
[0071] In an alternative approach, the formula for the mapping function is:
[0072] Where z represents the joint feature vector, and x represents the input data. The shared parameter of the mapping function is represented by y, which represents the true label of the input data, and X includes the multimodal data information of the candidate object, the facial feature data, and the eye-tracking interaction data.
[0073] This disclosure also provides an electronic device, including: at least one processor; a memory for storing at least one processor-executable instruction; wherein the at least one processor is used to execute the instruction to implement the steps of the method disclosed in this disclosure.
[0074] Figure 5 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present disclosure. Figure 5 As shown, the electronic device 500 includes at least one processor 501 and a memory 502 coupled to the processor 501, which can perform the corresponding steps in the methods disclosed in the embodiments of this disclosure.
[0075] The processor 501 described above can also be called a Central Processing Unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in this embodiment can be implemented by the integrated logic circuitry in the processor 501 or by software instructions. The processor 501 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 502, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 501 reads information from the memory 502 and, in conjunction with its hardware, completes the steps of the method described above.
[0076] Furthermore, various operations / processes according to this disclosure, implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, for example, Figure 6 The computer system 600 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above. Figure 6 This is a schematic diagram of the structure of a computer system provided in an embodiment of the present disclosure.
[0077] Computer system 600 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0078] like Figure 6 As shown, the computer system 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the computer system 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0079] Multiple components in the computer system 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device capable of inputting information into the computer system 600. The input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 608 may include, but is not limited to, a hard disk and an optical disk. The communication unit 609 allows the computer system 600 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, a modem, network card, infrared communication device, wireless communication transceiver, and / or chipset, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0080] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 602 and / or communication unit 609. In some embodiments, the computing unit 601 can be configured to perform the methods disclosed in this disclosure by any other suitable means (e.g., by means of firmware).
[0081] This disclosure also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the methods disclosed in this disclosure.
[0082] The computer-readable storage medium in this disclosure can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The aforementioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specifically, the aforementioned computer-readable storage medium may include electrical connections based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0083] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0084] Figure 7 This is a schematic diagram of a computer program product provided according to an embodiment of the present disclosure. Figure 7 As shown, the computer program product 700 includes a computer program 701, which, when executed by a processor, implements the methods disclosed in the embodiments of this disclosure.
[0085] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer.
[0086] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0087] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.
[0088] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0089] The above description is merely an illustration of some embodiments of this disclosure and the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0090] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A personalized recommendation method, characterized in that, include: Acquire the user's current facial feature data and eye-tracking interaction data, as well as multimodal data information of each object in the object library; Based on the user's browsing behavior at the current moment, multiple candidate objects are determined from the object library; For each candidate object, the multimodal data information of the candidate object is fused with the facial feature data and the eye-tracking interaction data to obtain the joint feature vector of the candidate object; Based on the joint feature vector corresponding to each candidate object, an interest score is determined for each object, and recommendations are made to the user based on the interest score of each object.
2. The method according to claim 1, characterized in that, The step of determining the interest score corresponding to each object based on the joint feature vector corresponding to each candidate object includes: For any candidate object, the interaction state between the user and the candidate object is determined based on the joint feature vector. If there is an interaction between the user and the candidate object, the duration of the user's gaze on the candidate object and the total duration of gaze on all candidate objects are determined based on the eye-tracking interaction data. The explicit interest score of the user on the candidate object is determined based on the duration of the user's gaze on the candidate object and the total duration of gaze on all candidate objects.
3. The method according to claim 2, characterized in that, The formula for calculating the explicit interest score is as follows: Among them, P i t represents the explicit interest score of the i-th candidate; i This indicates the duration of the current user's gaze on the candidate object; This represents the total gaze duration of all candidate objects, where n represents the total number of candidate objects. This represents a set of candidate objects, where i represents the identifier of a candidate object.
4. The method according to claim 1, characterized in that, The step of determining the interest score corresponding to each object based on the joint feature vector corresponding to each candidate object includes: In the absence of any interaction between the user and the candidate object, the stability of the user's switching time and binocular gaze direction changes between different candidate objects is determined based on the joint feature vector, and the user's interest score for the candidate object is determined based on the stability of the user's switching time and binocular gaze direction changes between different candidate objects and facial feature data.
5. The method according to claim 1, characterized in that, The step of fusing the multimodal data information of the candidate object with the facial feature data and the eye-tracking interaction data to obtain the joint feature vector of the candidate object includes: The multimodal data information of the candidate object, the facial feature data, and the eye-tracking interaction data are input into a preset mapping function to obtain the joint feature vector of the candidate object.
6. The method according to claim 5, characterized in that, The formula for the mapping function is: Where z represents the joint feature vector, and x represents the input data. The shared parameters of the mapping function are represented by , and y represents the true label of the input data. The predicted value X includes the multimodal data information of the candidate object, the facial feature data, and the eye-tracking interaction data.
7. A personalized recommendation device, characterized in that, include: The acquisition module is used to acquire the user's facial feature data and eye-tracking interaction data at the current moment, as well as the multimodal data information of each object in the object library; The determination module is used to determine multiple candidate objects from the object library based on the user's browsing behavior at the current moment; The fusion module is used to fuse the multimodal data information of each candidate object with the facial feature data and the eye-tracking interaction data to obtain the joint feature vector of the candidate object. The recommendation module is used to determine the interest score corresponding to each object based on the joint feature vector corresponding to each candidate object, and to make recommendations to the user based on the interest score of each object.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.