Method and device for recommending multimedia data based on in-vehicle face image

CN116563924BActive Publication Date: 2026-08-21CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310539434.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-12
Publication Date
2026-08-21
Estimated Expiration
2043-05-12

AI Technical Summary

Technical Problem

[0003]有鉴于此,本申请实施例提供了一种基于车内人脸图像推荐多媒体数据的方法、装置、电子设备及计算机可读存储介质,以解决相关技术中基于同一用户账号向不同乘客推荐个性化内容导致推荐准确度下降,影响不同乘客的观看体验的问题

Benefits of technology

[0008] The beneficial effects of this application embodiment compared with the prior art include at least the following: When the in-vehicle multimedia application is opened, this application embodiment can acquire the facial image of at least one object, and determine the gaze start point, gaze direction, and gaze length of each object based on the facial image of at least one object. Then, based on the gaze start point, gaze direction, and gaze length of each object, the target object for watching the in-vehicle multimedia application is determined. Based on the facial image of the target object, multimedia data recommended to the target object is determined and sent to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images. In this way, it can be determined whether a user is watching the in-vehicle multimedia application based on the user's facial image. When it is determined that the user is watching the in-vehicle multimedia application, multimedia data of similar users is recommended to the user based on the user's facial image. This can improve the accuracy of recommendations, allowing different passengers to watch their favorite content on the in-vehicle multimedia application and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116563924B_ABST
    Figure CN116563924B_ABST
Patent Text Reader

Abstract

The application provides a method and device for recommending multimedia data based on a face image in a vehicle. The method comprises: obtaining a face image of at least one object when a vehicle-mounted multimedia application is started; determining a visual line starting point, a visual line direction and a visual line length of each object according to the face image of the at least one object; determining a target object for watching the vehicle-mounted multimedia application according to the visual line starting point, the visual line direction and the visual line length of each object; determining multimedia data recommended to the target object according to the face image of the target object, and sending the multimedia data to the vehicle-mounted multimedia application. The technical solution of the application can determine whether a user is watching the vehicle-mounted multimedia application according to the face image of the user, and when it is determined that the user is watching the vehicle-mounted multimedia application, multimedia data of a similar user of the user is recommended to the user according to the face image of the user, so that the accuracy of the recommendation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of recommending data based on facial images, and more particularly to a method and apparatus for recommending multimedia data based on in-vehicle facial images. Background Technology

[0002] With the rapid development of internet technology, in-vehicle multimedia content has become increasingly abundant. Passengers in different locations within the vehicle can watch their preferred multimedia content on their respective in-vehicle screens. Different passengers may have different viewing preferences; for example, male passengers may prefer listening to financial and current affairs audio content in the morning, while children may prefer watching children's animated videos. Currently, personalized content recommendations for in-vehicle multimedia applications are mostly based on the same user account. Content viewing data is recorded in the user account, and personalized content recommendations are then made based on the user's browsing habits. However, in the in-vehicle scenario, since all passengers share the in-vehicle screen—for example, two children may be watching cartoons on the rear screen simultaneously, or the front passenger screen may be used by friends or family members at different times—and the user account remains unchanged in these scenarios, continuing to recommend personalized content to different passengers based on the same user account will lead to decreased recommendation accuracy and negatively impact the viewing experience for different passengers. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, apparatus, electronic device, and computer-readable storage medium for recommending multimedia data based on in-vehicle facial images, in order to solve the problem in related technologies where recommending personalized content to different passengers based on the same user account leads to a decrease in recommendation accuracy and affects the viewing experience of different passengers.

[0004] A first aspect of this application provides a method for recommending multimedia data based on in-vehicle facial images. The method includes: acquiring a facial image of at least one object when an in-vehicle multimedia application is activated; determining the gaze start point, gaze direction, and gaze length of each object based on the facial images of the at least one object; determining a target object for viewing the in-vehicle multimedia application based on the gaze start point, gaze direction, and gaze length of each object; determining multimedia data to recommend to the target object based on the facial image of the target object, and sending the multimedia data to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between historical multimedia data of the target object's facial image and historical multimedia data of other objects' facial images.

[0005] A second aspect of this application provides an apparatus for recommending multimedia data based on in-vehicle facial images. The apparatus includes: an acquisition module for acquiring a facial image of at least one object when an in-vehicle multimedia application is activated; a first determination module for determining the gaze start point, gaze direction, and gaze length of each object based on the facial images of the at least one object; a second determination module for determining a target object viewing the in-vehicle multimedia application based on the gaze start point, gaze direction, and gaze length of each object; and a third determination module for determining multimedia data to recommend to the target object based on the facial image of the target object, and sending the multimedia data to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between historical multimedia data of the target object's facial image and historical multimedia data of other objects' facial images.

[0006] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0007] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0008] The beneficial effects of this application embodiment compared with the prior art include at least the following: When the in-vehicle multimedia application is opened, this application embodiment can acquire the facial image of at least one object, and determine the gaze start point, gaze direction, and gaze length of each object based on the facial image of at least one object. Then, based on the gaze start point, gaze direction, and gaze length of each object, the target object for watching the in-vehicle multimedia application is determined. Based on the facial image of the target object, multimedia data recommended to the target object is determined and sent to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images. In this way, it can be determined whether a user is watching the in-vehicle multimedia application based on the user's facial image. When it is determined that the user is watching the in-vehicle multimedia application, multimedia data of similar users is recommended to the user based on the user's facial image. This can improve the accuracy of recommendations, allowing different passengers to watch their favorite content on the in-vehicle multimedia application and improving the user experience. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown;

[0011] Figure 2 This is a flowchart illustrating a method for recommending multimedia data based on in-vehicle facial images, according to an embodiment of this application.

[0012] Figure 3 This is a flowchart illustrating the steps for determining the target object to view the in-vehicle multimedia application according to an embodiment of this application;

[0013] Figure 4 This is a flowchart illustrating another step in determining the target object for viewing an in-vehicle multimedia application, as described in this application embodiment.

[0014] Figure 5 This is a flowchart of another method for recommending multimedia data based on in-vehicle facial images, according to an embodiment of this application.

[0015] Figure 6 This is a block diagram of a device for recommending multimedia data based on in-vehicle facial images, according to an embodiment of this application.

[0016] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0017] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0018] It should be noted that the user information (including but not limited to vehicle equipment information, user facial images, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0019] The method and apparatus for recommending multimedia data based on in-vehicle facial images according to embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0020] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.

[0021] like Figure 1 As shown, system architecture 100 may include vehicle devices 101, 102, and 103, network 104, and server 105. Network 104 is used as a medium to provide communication links between vehicle devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0022] It should be understood that Figure 1 The number of vehicles, devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of vehicles, devices, networks, and servers can be included. For example, server 105 could be a server cluster composed of multiple servers.

[0023] Users can use the multimedia applications on vehicle devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send multimedia data, etc. Vehicle devices 101, 102, and 103 can be various electronic devices with screens.

[0024] Server 105 can be a server providing various services. For example, when the in-vehicle multimedia application is turned on, server 105 can obtain facial images of at least one object from vehicle device 101 (or 102 or 103), and determine the gaze start point, gaze direction, and gaze length of each object based on the facial images of at least one object. Then, based on the gaze start point, gaze direction, and gaze length of each object, it can determine the target object watching the in-vehicle multimedia application, determine the multimedia data to recommend to the target object based on the facial image of the target object, and send the multimedia data to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images. In this way, it can determine whether a user is watching the in-vehicle multimedia application based on the user's facial image. When it is determined that the user is watching the in-vehicle multimedia application, it can recommend multimedia data of similar users to the user based on the user's facial image. This can improve the accuracy of recommendations, allowing different passengers to watch their favorite content on the in-vehicle multimedia application and improving the user experience.

[0025] In some embodiments, the method for recommending multimedia data based on in-vehicle facial images provided in this invention is generally executed by server 105, and correspondingly, the device for recommending multimedia data based on in-vehicle facial images is generally located in server 105. In other embodiments, certain vehicle devices may have functions similar to a server to execute this method. Therefore, the method for recommending multimedia data based on in-vehicle facial images provided in this invention is not limited to execution on the server side.

[0026] Figure 2 This is a flowchart illustrating a method for recommending multimedia data based on in-vehicle facial images, according to an embodiment of this application. The method provided in this embodiment can be executed by any electronic device with computer processing capabilities. The electronic device can be... Figure 1 The server shown.

[0027] like Figure 2 As shown, the method includes steps S210 to S240.

[0028] In step S210, when the in-vehicle multimedia application is turned on, a facial image of at least one object is acquired.

[0029] In step S220, the starting point of the gaze, the direction of the gaze, and the length of the gaze are determined based on the face image of at least one object.

[0030] In step S230, the target object for viewing the in-vehicle multimedia application is determined based on the starting point, direction, and length of the line of sight of each object.

[0031] In step S240, based on the facial image of the target object, multimedia data recommended to the target object is determined and sent to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of the facial images of other objects.

[0032] This method acquires facial images of at least one object when the in-vehicle multimedia application is activated. Based on these images, it determines the starting point, direction, and length of each object's gaze. Then, based on these parameters, it identifies the target object viewing the application. Using the target object's facial image, it determines recommended multimedia data and sends it to the application. The multimedia data is determined based on the similarity between the target object's historical multimedia data and the historical multimedia data of other objects' facial images. This allows the system to determine whether a user is viewing the application based on their facial image. When this is confirmed, it recommends multimedia data from similar users based on their facial image, improving recommendation accuracy and allowing different passengers to view content they prefer, thus enhancing the user experience.

[0033] In some embodiments, the aforementioned in-vehicle multimedia applications may be video applications, novel applications, audio applications, news applications, etc., but are not limited to these.

[0034] In some embodiments, when the in-vehicle multimedia application is activated, the in-vehicle camera can capture facial images of all passengers inside the vehicle and send these images to a server. For example, when the in-vehicle multimedia application starts playing a video, the vehicle control module activates the in-vehicle camera to capture the user's facial image. In this application, a local facial image database can be constructed, or facial images can be sent to a server and stored in that server's facial image database. The server can recognize the user's facial images; for example, it can store and learn facial data of the same user under different lighting conditions and different attire (such as wearing a hat or glasses) to achieve accurate facial image recognition.

[0035] In some embodiments, vehicle models equipped with in-vehicle cameras, such as those with passenger monitoring systems, generally meet the usage scenarios of this application. The multimedia data viewed by the user is recorded by the in-vehicle infotainment system and also serves as the carrier for the final recommended content. Under certain conditions, the in-vehicle camera is activated by the vehicle control module to acquire facial images, and then the user's browsing behavior is uploaded to the personalized recommendation system (server) via a remote vehicle service. Finally, the personalized recommendation system generates personalized recommended content using algorithms.

[0036] In some embodiments, the gaze origin, gaze direction, and gaze length of each object are determined based on the facial image of at least one object. For example, by recognizing the facial images of each object using a visual focusing algorithm, a three-dimensional vector of the gaze origin, a three-dimensional vector of the gaze direction, and the gaze length of each object can be obtained. In this embodiment, the gaze origin may refer to the position of the user's eyes, the gaze direction may refer to the direction the user is currently looking, and the gaze length may refer to the distance between the user's eyes and the in-vehicle screen.

[0037] In some embodiments, static line-of-sight parameters can be set according to the actual application scenario before executing the method of this application. For example, different static line-of-sight parameters can be set for the positions of different in-vehicle screens and vehicle seats. Specifically, the static line-of-sight parameters corresponding to the in-vehicle screen are obtained through actual detection based on the position information of the driver's seat and the position information of the in-vehicle screen corresponding to the driver's seat. The static line-of-sight parameters corresponding to the in-vehicle screen are obtained through actual detection based on the position information of the passenger seat and the position information of the in-vehicle screen corresponding to the passenger seat. The static line-of-sight parameters corresponding to the in-vehicle screen are obtained through actual detection based on the position information of the rear seats and the position information of the in-vehicle screen corresponding to the rear seats. In this embodiment, the static line-of-sight parameters include the static line-of-sight starting point, the static line-of-sight direction, and the static line-of-sight length. For example, the three-dimensional vector of the static line-of-sight starting point is (x0, y0, z0), the three-dimensional vector of the static line-of-sight direction is (x1, y1, z1), and the static line-of-sight direction is L.

[0038] In some embodiments, the target object for viewing the in-vehicle multimedia application is determined based on the starting point, direction, and length of each object's gaze. For example, the location information of each object within the vehicle can be determined based on its gaze starting point, direction, and length, thus identifying which seat each object is sitting in. Then, based on the location information, at least one in-vehicle screen corresponding to that location can be identified. At this point, the target object for viewing the in-vehicle multimedia application can be determined from among the objects based on the object's gaze starting point, direction, and length, and the static gaze parameters of the at least one in-vehicle screen. This effectively avoids recommending multimedia data based on the facial images of objects not currently viewing the application, which could negatively impact the user's multimedia viewing experience.

[0039] In some embodiments, multimedia data recommended to a target object is determined based on the target object's facial image, and the multimedia data is sent to the in-vehicle multimedia application. For example, multimedia data recommended to that FaceID (set for the target object's facial image) is retrieved from a server's database; that is, multimedia data recommended to the target object. In this embodiment, the multimedia data recommended to that FaceID can be determined based on the similarity between historical multimedia data bound to that FaceID and historical multimedia data bound to other FaceIDs. This approach can improve the accuracy of the recommendation.

[0040] In some embodiments, before sending the multimedia data recommended to the target object to the in-vehicle multimedia application, the method further includes: determining a score for each item in the multimedia data recommended to the target object based on historical multimedia data of the target object's facial image, historical multimedia data of other objects' facial images, and the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images; updating the ranking of each item in the multimedia data recommended to the target object based on the score of each item, to obtain updated multimedia data recommended to the target object. This approach can further improve the accuracy of the recommended multimedia data.

[0041] In some embodiments, after recommending multimedia data to the in-vehicle multimedia application, two types of user feedback behaviors regarding the recommended multimedia data can be acquired in real time. These include positive feedback behaviors (such as clicking, saving, and purchasing) and negative feedback behaviors (such as ignoring, skipping, reporting, and disliking). For positive feedback behaviors, the weight of content that has been clicked or saved is increased, making it more likely to be recommended. For negative feedback behaviors, they are treated as negative feedback in recommendation algorithms based on log data such as collaborative filtering, reducing the recommendation weight of these contents and lowering their recommendation frequency and priority.

[0042] In some embodiments, the historical multimedia data of the target object's facial image can be historical multimedia data bound to a FaceID set for the target user's facial image. This historical multimedia data can be multimedia objects viewed by the user in the past period, as well as behavioral data related to those multimedia objects. Similarly, the historical multimedia data of other objects' facial images can be historical multimedia data bound to a FaceID set for other users' facial images. This historical multimedia data can be multimedia objects viewed by those other users in the past period, as well as behavioral data related to those multimedia objects. The behavioral data can include actions such as liking, saving, forwarding, and taking screenshots of the multimedia objects.

[0043] Figure 3 This is a flowchart illustrating the steps for determining the target object to view the in-vehicle multimedia application according to an embodiment of this application, such as... Figure 3 As shown, step S230 may include steps S310 to S330.

[0044] In step S310, the position information of each object inside the vehicle is determined based on the starting point, direction, and length of the line of sight of each object.

[0045] In step S320, based on the location information of each object inside the vehicle, the static line-of-sight parameters of at least one in-vehicle screen corresponding to the location information are obtained.

[0046] In step S330, the target object for viewing the in-vehicle multimedia application is determined based on the starting point, direction, and length of the line of sight of each object and the static line of sight parameters of at least one in-vehicle screen.

[0047] This method can determine the position information of each object in the vehicle based on the starting point, direction, and length of its gaze. Based on the position information of each object in the vehicle, it obtains the static gaze parameters of at least one in-vehicle screen corresponding to the position information. Then, based on the starting point, direction, and length of each object's gaze and the static gaze parameters of at least one in-vehicle screen, it determines the target object for watching the in-vehicle multimedia application. This can effectively avoid recommending multimedia data based on the facial images of objects who are not watching the in-vehicle multimedia application, thereby avoiding affecting the user's experience of watching multimedia data.

[0048] Figure 4 This is a flowchart illustrating another step in determining the target object for viewing an in-vehicle multimedia application, as described in an embodiment of this application. Figure 4 As shown, step S330 may further include steps S410 and S420.

[0049] In step S410, for an object, a first deviation information is determined based on the object's line of sight starting point and static line of sight starting point, a second deviation information is determined based on the object's line of sight direction and static line of sight direction, and a third deviation information is determined based on the object's line of sight length and static line of sight length.

[0050] In step S420, based on the first deviation information, the second deviation information, and the third deviation information, it is determined whether the object is the target object for watching the in-vehicle multimedia application.

[0051] This method can determine first deviation information based on the object's gaze origin and static gaze origin, second deviation information based on the object's gaze direction and static gaze direction, and third deviation information based on the object's gaze length and static gaze length. Then, based on the first, second, and third deviation information, it can determine whether the object is the target object for watching the in-vehicle multimedia application. In this way, the target object for watching the in-vehicle multimedia application can be quickly and accurately identified from at least one object. This can effectively avoid recommending multimedia data based on the facial images of objects who are not watching the in-vehicle multimedia application, thereby avoiding affecting the user's experience of watching multimedia data.

[0052] In some embodiments, a first deviation information is determined based on the object's line-of-sight starting point and static line-of-sight starting point. This allows for the calculation of the specific deviation value between the object's line-of-sight starting point and the static line-of-sight starting point, and determines whether the deviation value meets the conditions set for the static line-of-sight starting point, such as whether the deviation value is within 10% of the static line-of-sight starting point. A second deviation information is determined based on the object's line-of-sight direction and static line-of-sight direction. This allows for the calculation of the specific deviation value between the object's line-of-sight direction and the static line-of-sight direction, and determines whether the deviation value meets the conditions set for the static line-of-sight direction, such as whether the deviation value is within 10% of the static line-of-sight direction. A third deviation information is determined based on the object's line-of-sight length and static line-of-sight length. This allows for the calculation of the specific deviation value between the object's line-of-sight direction and the static line-of-sight length, and determines whether the deviation value meets the conditions set for the static line-of-sight length, such as whether the deviation value is within 10% of the static line-of-sight length. When the first deviation information, the second deviation information, and the third deviation information meet the corresponding conditions, the object is determined to be the target object for viewing the in-vehicle multimedia application. Conversely, if the first deviation information, the second deviation information, and the third deviation information do not meet the corresponding conditions mentioned above, it is determined that the object is not the target object for watching the in-vehicle multimedia application.

[0053] Figure 5 This is a flowchart of another method for recommending multimedia data based on in-vehicle facial images, as described in this application embodiment. Figure 5 As shown, the above method may further include steps S510 to S540.

[0054] In step S510, a face image of at least one historical object is acquired.

[0055] In step S520, the starting point, direction, and length of the gaze of each historical object are determined based on the facial image of at least one historical object.

[0056] In step S530, the historical target object for viewing the in-vehicle multimedia application is determined based on the starting point, direction, and length of the line of sight of each historical object.

[0057] In step S540, the similarity between the historical multimedia data of the face image of the historical target object and the historical multimedia data of the face image of other historical objects is determined, and the historical multimedia data to be recommended to the historical target object is determined based on the similarity.

[0058] This method can determine the similarity between historical multimedia data of a target object's face image and historical multimedia data of other historical objects' face images, and determine the historical multimedia data to be recommended to the target object based on the similarity. This helps to establish a mapping relationship between the target object's face image and the multimedia data recommended to the target object based on the similarity, which facilitates the accurate recommendation of corresponding multimedia data to the target object based on the face image in the future.

[0059] In some embodiments, the historical target object for viewing the in-vehicle multimedia application is determined based on the starting point, direction, and length of the gaze of each historical object. This avoids subsequently calculating the similarity of historical multimedia data of facial images of objects not viewing the in-vehicle multimedia application with that of other historical objects. In this embodiment, a corresponding FaceID can be set for the facial image of the historical target object, and the historical multimedia data bound to that FaceID is obtained based on the FaceID. Similarly, a corresponding FaceID can be set for the facial images of other historical objects, and the historical multimedia data bound to that FaceID is obtained based on the FaceID. The aforementioned historical multimedia data may include multimedia objects viewed by the user in past time periods, as well as behavioral data related to those multimedia objects.

[0060] In some embodiments, after identifying the viewer through the object's facial image and a visual focusing algorithm, personalized content recommendation can be performed based on a collaborative filtering algorithm. Specifically, the similarity between the historical multimedia data of the historical target object's facial image and the historical multimedia data of other historical objects' facial images can be calculated using the following formula (1), where the historical target object is user j, and the other historical objects are a set of users v, as follows:

[0061]

[0062] Among them, u j This represents a list of content from user j's historical multimedia data that has generated actions related to the project content. v This represents a list of content from user v's historical multimedia data that has generated actions related to the project content, |u j| represents the number of instances in user j's historical multimedia data where user j has interacted with specific content. |U v | represents the number of instances in the historical multimedia data of user v that have generated actions related to the project content, sim(u j ,u v ) represents the similarity between the historical multimedia data of user j and the historical multimedia data of user v, sim(u j ,u v The smaller the value, the more similar the historical multimedia data of user j is to the historical multimedia data of user v.

[0063] In some embodiments, the above method further includes: determining the score of each item in the historical multimedia data recommended to the historical target object based on the historical multimedia data of the historical target object's face image, the historical multimedia data of other historical objects' face images, and the similarity between the historical multimedia data of the historical target object's face image and the historical multimedia data of other historical objects' face images; updating the ranking of each item in the historical multimedia data recommended to the historical target object based on the score of each item, to obtain the updated historical multimedia data recommended to the historical target object. In this embodiment, the score of each item in the historical multimedia data recommended to the historical target object is calculated using the following formula (2), where the historical target object is user j, and the other historical objects are a set of users v, as detailed below:

[0064]

[0065] Where S(j,k) represents the k users whose interests are closest to user j, N(i) represents the set of users who have interacted with item content i, and W jv r represents the similarity between the historical multimedia data of user j and the historical multimedia data of user v. vi p represents user v's rating of item content i. j i represents user j's rating of item content i, p j The larger the value of i, the higher the user j's interest in the project content i.

[0066] In some embodiments, k in formula (2) above can be obtained from the ranking results of multiple similarities calculated by formula (1) above. For example, the top 5 are taken from the ranking results, in which case k is 5.

[0067] In some embodiments, the ranking of each item in the historical multimedia data recommended to the historical target object is updated according to the score of each item calculated by the above formula (2), so as to obtain the updated historical multimedia data recommended to the historical target object. For example, according to the ranking of interest scores (i.e., the score of each item), several items with the highest scores are selected as the historical multimedia data recommended to the historical target object, and a mapping relationship between the FaceID set for the face image of the historical target object and the historical multimedia data is established. In this way, when the historical target object watches the in-vehicle multimedia application again, the historical multimedia data recommended to the historical target object can be obtained according to the FaceID.

[0068] In some embodiments, the method further includes, after determining that an object is viewing the in-vehicle multimedia application, recording in real time the multimedia objects viewed by the object on the in-vehicle multimedia application based on the object's facial image, and recording the object's behavioral data related to the multimedia objects; after a preset time, determining whether the object is viewing the in-vehicle multimedia application based on the object's current gaze start point, gaze direction, and gaze length; if yes, continuing to record the multimedia objects viewed by the object and the behavioral data related to the multimedia objects; or if not, stopping the recording of relevant data related to the object on the in-vehicle multimedia application; when it is determined that there are multiple objects viewing the in-vehicle multimedia application, determining the weight of each object for the multimedia objects of the in-vehicle multimedia application based on the number of objects viewing the in-vehicle multimedia application, storing the weight of each object for the multimedia objects of the in-vehicle multimedia application based on the facial images of each object, and using the weights to recommend multimedia data to each object. For example, when it is determined that an object is viewing the in-vehicle multimedia application, recording the multimedia objects viewed by the object and the behavioral data related to the multimedia objects, and binding this content as multimedia data with the FaceID set for the object's facial image. Then, every 5 minutes, based on the current starting point, direction, and length of the object's gaze, it is determined whether the object is watching the in-vehicle multimedia application to continuously confirm whether it is the same person watching, achieving accurate viewing data binding. That is, every 5 minutes, the object's facial image needs to be re-acquired, and the object's viewing status is determined based on this image. When multiple objects are detected simultaneously watching the screen (e.g., two children in the back seat watching cartoons), the multimedia data from the in-vehicle multimedia application is simultaneously recorded under the FaceID set for each object's facial image, and the records themselves are weighted accordingly. A single person watching has a weight of 1, two people watching has a weight of 0.5, and n people watching has a weight of 1 / n. Finally, when recommending content, these weight values ​​are incorporated into the recommendation logic.

[0069] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. The apparatus for recommending multimedia data based on in-vehicle facial images described below can be referred to in correspondence with the method for recommending multimedia data based on in-vehicle facial images described above. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of this application.

[0070] Figure 6 This is a block diagram of a device for recommending multimedia data based on in-vehicle facial images, according to an embodiment of this application.

[0071] like Figure 6 As shown, the device 600 for recommending multimedia data based on in-vehicle facial images includes an acquisition module 610, a first determination module 620, a second determination module 630, and a third determination module 640.

[0072] Specifically, the acquisition module 610 is used to acquire a facial image of at least one object when the in-vehicle multimedia application is turned on.

[0073] The first determining module 620 is used to determine the starting point of the gaze, the direction of the gaze, and the length of the gaze for each object based on the face image of at least one object.

[0074] The second determining module 630 is used to determine the target object for viewing the in-vehicle multimedia application based on the starting point, direction, and length of the line of sight of each object.

[0075] The third determining module 640 is used to determine the multimedia data recommended to the target object based on the target object's facial image, and send the multimedia data to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the target object's historical multimedia data and the historical multimedia data of other objects' facial images.

[0076] The device 600, which recommends multimedia data based on in-vehicle facial images, can acquire facial images of at least one object when the in-vehicle multimedia application is turned on. Based on these images, it determines the starting point, direction, and length of each object's gaze. Then, based on these parameters, it identifies the target object viewing the in-vehicle multimedia application. Based on the target object's facial image, it determines the multimedia data to recommend to that target object and sends the multimedia data to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the target object's historical multimedia data and the historical multimedia data of other objects' facial images. This method can determine whether a user is viewing the in-vehicle multimedia application based on their facial image. When it is determined that the user is viewing the application, it recommends multimedia data from similar users based on their facial image. This improves the accuracy of recommendations, allowing different passengers to view content they like on the in-vehicle multimedia application and enhancing the user experience.

[0077] In some embodiments, the second determining module 630 is configured to: determine the position information of each object in the vehicle based on the starting point, direction, and length of the line of sight of each object; obtain the static line of sight parameters of at least one in-vehicle screen corresponding to the position information of each object in the vehicle; and determine the target object for watching the in-vehicle multimedia application based on the starting point, direction, and length of the line of sight of each object and the static line of sight parameters of at least one in-vehicle screen.

[0078] In some embodiments, determining the target object for viewing the in-vehicle multimedia application based on the viewing start point, viewing direction, and viewing length of each object and the static viewing parameters of at least one in-vehicle screen includes: for an object, determining first deviation information based on the viewing start point and static viewing start point of the object, determining second deviation information based on the viewing direction and static viewing direction of the object, and determining third deviation information based on the viewing length and static viewing length of the object; and determining whether the object is the target object for viewing the in-vehicle multimedia application based on the first deviation information, the second deviation information, and the third deviation information.

[0079] In some embodiments, before determining the multimedia data to be recommended to the target object based on the face image of the target object, the apparatus 600 for recommending multimedia data based on the in-vehicle face image is further configured to: determine the score of each item in the historical multimedia data to be recommended to the historical target object based on the historical multimedia data of the historical target object's face image, the historical multimedia data of other historical objects' face images, and the similarity between the historical multimedia data of the historical target object's face image and the historical multimedia data of other historical objects' face images; and update the ranking of each item in the historical multimedia data to be recommended to the historical target object based on the score of each item, thereby obtaining the updated historical multimedia data to be recommended to the historical target object.

[0080] In some embodiments, the apparatus 600 for recommending multimedia data based on in-vehicle facial images is further configured to: determine a score for each item in the historical multimedia data recommended to the historical target object based on historical multimedia data of facial images of historical target objects, historical multimedia data of facial images of other historical objects, and the similarity between historical multimedia data of facial images of historical target objects and historical multimedia data of facial images of other historical objects; and update the ranking of each item in the historical multimedia data recommended to the historical target object based on the score of each item, thereby obtaining updated historical multimedia data recommended to the historical target object.

[0081] In some embodiments, before sending the multimedia data to the in-vehicle multimedia application, the apparatus 600 for recommending multimedia data based on in-vehicle facial images is further configured to: determine a score for each item in the multimedia data recommended to the target object based on historical multimedia data of the target object's facial images, historical multimedia data of other objects' facial images, and the similarity between the historical multimedia data of the target object's facial images and the historical multimedia data of other objects' facial images; update the ranking of each item in the multimedia data recommended to the target object based on the score of each item, to obtain updated multimedia data recommended to the target object; and send the updated multimedia data recommended to the target object to the in-vehicle multimedia application.

[0082] In some embodiments, the device 600 for recommending multimedia data based on in-vehicle facial images is further configured to: after determining an object watching an in-vehicle multimedia application, record in real time the multimedia objects viewed by the object on the in-vehicle multimedia application based on the object's facial image, and record the object's behavioral data related to the multimedia objects; after a preset time, determine whether the object is watching the in-vehicle multimedia application based on the object's current gaze starting point, gaze direction, and gaze length; if yes, continue recording the multimedia objects viewed by the object on the in-vehicle multimedia application and the behavioral data related to the multimedia objects; or if not, stop recording the relevant data related to the object on the in-vehicle multimedia application; when it is determined that there are multiple objects watching the in-vehicle multimedia application, determine the weight of each object for the multimedia objects of the in-vehicle multimedia application based on the number of objects watching the in-vehicle multimedia application, store the weight of each object for the multimedia objects of the in-vehicle multimedia application based on the facial images of each object, and use the weights to recommend multimedia data to each object.

[0083] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application.

[0084] like Figure 7 As shown, the electronic device 700 of this embodiment includes a processor 710, a memory 720, and a computer program 730 stored in the memory 720 and executable on the processor 710. When the processor 710 executes the computer program 730, it implements the steps in the various method embodiments described above. Alternatively, when the processor 710 executes the computer program 730, it implements the functions of each module in the various device embodiments described above.

[0085] Electronic device 700 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 700 may include, but is not limited to, processor 710 and memory 720. Those skilled in the art will understand that... Figure 7 This is merely an example of electronic device 700 and does not constitute a limitation on electronic device 700. It may include more or fewer parts than shown, or different parts.

[0086] The processor 710 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0087] The memory 720 can be an internal storage unit of the electronic device 700, such as a hard disk or RAM of the electronic device 700. The memory 720 can also be an external storage device of the electronic device 700, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 700. The memory 720 can also include both internal and external storage units of the electronic device 700. The memory 720 is used to store computer programs and other programs and data required by the electronic device.

[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0089] If an integrated module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in a computer-readable medium can be appropriately added to or subtracted according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0090] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for recommending multimedia data based on in-vehicle facial images, characterized in that, The method includes: When the in-vehicle multimedia application is turned on, acquire the facial image of at least one object; acquiring the facial image of at least one object when the in-vehicle multimedia application is turned on includes: when the in-vehicle multimedia application starts playing a video, activating the in-vehicle camera through the vehicle control module to acquire the facial image; Based on the facial image of at least one object, determine the gaze start point, gaze direction, and gaze length of each object; The target objects for viewing the in-vehicle multimedia application are determined based on the starting point, direction, and length of each object's line of sight. Based on the facial image of the target object, multimedia data recommended to the target object is determined, and the multimedia data is sent to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images. When it is determined that there are multiple objects watching the in-vehicle multimedia application, the weight of each object relative to the multimedia object of the in-vehicle multimedia application is determined according to the number of objects watching the in-vehicle multimedia application. The weight of each object relative to the multimedia object of the in-vehicle multimedia application is stored based on the face image of each object. The weight is used to recommend multimedia data to each object.

2. The method according to claim 1, characterized in that, Based on the starting point, direction, and length of each object's line of sight, the target objects for viewing the in-vehicle multimedia application are determined as follows: Based on the starting point, direction, and length of each object's line of sight, determine the position information of each object inside the vehicle; Based on the location information of each object inside the vehicle, obtain the static line-of-sight parameters of at least one in-vehicle screen corresponding to the location information; The target objects for viewing the in-vehicle multimedia application are determined based on the starting point, direction, and length of each object's line of sight and the static line of sight parameters of the at least one in-vehicle screen.

3. The method according to claim 2, characterized in that, Based on the starting point, direction, and length of each object's gaze, and the static gaze parameters of the at least one in-vehicle screen, the target objects for viewing the in-vehicle multimedia application are determined as follows: For an object, the first deviation information is determined based on the object's line of sight starting point and static line of sight starting point; the second deviation information is determined based on the object's line of sight direction and static line of sight direction; and the third deviation information is determined based on the object's line of sight length and static line of sight length. Based on the first deviation information, the second deviation information, and the third deviation information, it is determined whether the object is the target object for watching the in-vehicle multimedia application.

4. The method according to claim 1, characterized in that, Before determining the multimedia data to be recommended to the target object based on the target object's facial image, the method further includes: Obtain the face image of at least one historical object; Based on the facial image of at least one historical object, determine the gaze start point, gaze direction, and gaze length of each historical object; Based on the starting point, direction, and length of the line of sight of each historical object, the historical target object for viewing the in-vehicle multimedia application is determined; The similarity between the historical multimedia data of the face image of the historical target object and the historical multimedia data of the face image of other historical objects is determined, and historical multimedia data recommended to the historical target object is determined based on the similarity.

5. The method according to claim 4, characterized in that, The method further includes: Based on the historical multimedia data of the target object's facial image, the historical multimedia data of other historical objects' facial images, and the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other historical objects' facial images, a score is determined for each item in the historical multimedia data recommended to the target object. Based on the rating of each item, the ranking of each item in the historical multimedia data recommended to the historical target object is updated to obtain the updated historical multimedia data recommended to the historical target object.

6. The method according to claim 1, characterized in that, Before sending the multimedia data to the in-vehicle multimedia application, the method further includes: Based on the historical multimedia data of the target object's facial image, the historical multimedia data of other objects' facial images, and the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images, a score is determined for each item in the multimedia data recommended to the target object. Based on the rating of each item, the ranking of each item in the multimedia data recommended to the target object is updated to obtain the updated multimedia data recommended to the target object; The updated multimedia data recommended to the target object will be sent to the in-vehicle multimedia application.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: After identifying the object viewing the in-vehicle multimedia application, the system records in real time the multimedia objects viewed by the object on the in-vehicle multimedia application based on the object's facial image, and also records the object's behavioral data related to the multimedia objects. After a preset time, based on the object's current gaze starting point, gaze direction, and gaze length, it is determined whether the object is watching the in-vehicle multimedia application. If so, the recording of the multimedia objects viewed by the object on the in-vehicle multimedia application and the behavioral data performed on the multimedia objects continues; otherwise, the recording of relevant data on the object on the in-vehicle multimedia application is stopped.

8. A device for recommending multimedia data based on in-vehicle facial images, characterized in that, The device includes: The acquisition module is used to acquire a facial image of at least one object when the in-vehicle multimedia application is turned on; acquiring a facial image of at least one object when the in-vehicle multimedia application is turned on includes: when the in-vehicle multimedia application starts playing a video, the in-vehicle camera is activated through the vehicle control module to acquire the facial image; The first determining module is used to determine the starting point of the gaze, the direction of the gaze, and the length of the gaze for each object based on the face image of the at least one object; The second determining module is used to determine the target object for viewing the in-vehicle multimedia application based on the starting point, direction, and length of the line of sight of each object. The third determining module is used to determine the multimedia data recommended to the target object based on the target object's facial image, and send the multimedia data to the in-vehicle multimedia application. The multimedia data is determined based on the similarity between the historical multimedia data of the target object's facial image and the historical multimedia data of other objects' facial images. The device is further configured to, when it is determined that there are multiple objects watching the in-vehicle multimedia application, determine the weight of each object relative to the multimedia object of the in-vehicle multimedia application based on the number of objects watching the in-vehicle multimedia application, store the weight of each object relative to the multimedia object of the in-vehicle multimedia application based on the face image of each object, and use the weight to recommend multimedia data to each object.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video recommendation method and device

    CN112104914A

  • Media content recommendation method and device, vehicle-mounted terminal and storage medium

    CN114694199A

  • Vehicle-mounted screen control method and device and vehicle

    CN114882579A