Multimedia data recommendation method and vehicle
By acquiring occupant profile information and filtering multimedia data based on navigation system duration, the problem of poor adaptability of multimedia data recommendations in multi-passenger in-vehicle scenarios was solved, achieving a complete occupant experience and high group satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-07
AI Technical Summary
In multi-passenger in-vehicle scenarios, existing technologies struggle to meet the multimedia data recommendation needs of different passengers, resulting in poor adaptability and a reduced driving experience.
By acquiring profile information of multiple occupants, data recommendation features with common preferences are determined. Multimedia data is then filtered in conjunction with the vehicle's navigation system trip duration to ensure time-based adaptability. Multimodal sensing devices are used to acquire facial images and audio information of occupants, analyze personal information and physiological state, and adjust multimedia data recommendations accordingly.
It improves the adaptability of multimedia data to passengers, ensuring that passengers have a complete viewing/listening experience, and enhances the group satisfaction and driving experience of multiple passengers in the vehicle.
Smart Images

Figure CN121808073A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimedia display technology for vehicles, and more specifically, to a multimedia data recommendation method and vehicle in the field of multimedia display technology for vehicles. Background Technology
[0002] With the development of modern automotive technology, in-vehicle terminals can recommend multimedia data to passengers based on their preferences. However, in actual in-vehicle scenarios with multiple passengers, due to differences in the preferences of different passengers and the duration of the journey, in-vehicle terminals struggle to meet the entertainment needs of multiple passengers during the journey. This results in poor compatibility between the recommended multimedia data and the individual passengers in the vehicle, thus reducing the driving experience. Summary of the Invention
[0003] This application provides a multimedia data recommendation method and a vehicle, which can improve the group satisfaction of multiple occupants in the vehicle and ensure the occupants' driving experience.
[0004] Firstly, a multimedia data recommendation method is provided, which is applied to a vehicle. The method includes: acquiring first profile information of multiple occupants in the vehicle; determining first data recommendation features that conform to the common preferences of each occupant based on each first profile information; determining the first trip duration of the vehicle based on the vehicle's navigation system; filtering a first multimedia data set corresponding to the first data recommendation features based on the first trip duration to obtain a second multimedia data set; determining first multimedia recommendation data in the second multimedia data set; and playing the first multimedia recommendation data on the vehicle's display screen.
[0005] In the above technical solution, by acquiring data recommendation features based on the common preferences of multiple occupants, the adaptability of the recommended multimedia data to multiple occupants is improved. By filtering multimedia data that can be viewed in its entirety based on the trip duration, the adaptability of the recommended multimedia data to the vehicle driving scenario in the time dimension is improved, ensuring that occupants have a complete viewing / listening experience. This method improves the group satisfaction of multiple occupants in the vehicle and ensures the occupants' driving experience.
[0006] In conjunction with the first aspect, in some possible implementations, the vehicle is equipped with a multimodal perception device. The step of acquiring first profile information of multiple occupants in the vehicle and determining first data recommendation features that conform to the common preferences of each occupant based on each first profile information includes: acquiring first facial images and first audio information of multiple occupants in the vehicle based on the multimodal perception device; determining the personal information and first physiological state of each occupant based on each first facial image and each first audio information; determining the target preference label of each occupant based on the personal information and the first physiological state; and determining the first data recommendation features that conform to the common preferences of each occupant based on the target preference label of each occupant.
[0007] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of determining the target preference label of each occupant based on personal information and the first physiological state includes: if the personal information indicates that the first target occupant is a historical occupant of the vehicle, then the historical preference label of the first target occupant is obtained, and the first target occupant is any one of multiple occupants; and the target preference label of the first target occupant is determined based on the historical preference label and the first physiological state.
[0008] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of filtering the first multimedia data set corresponding to the first data recommendation feature based on the first trip duration to obtain the second multimedia data set includes: obtaining the first multimedia data set corresponding to the first data recommendation feature, wherein the first multimedia data set is composed of each first multimedia data; in the first multimedia data set, determining the first multimedia data whose playback duration is less than or equal to the first trip duration as the second multimedia data, and determining the second multimedia data set based on each second multimedia data.
[0009] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, the step of determining multimedia recommended data in the second multimedia data set and playing the multimedia recommended data on the vehicle's display screen includes: if the first trip duration is less than or equal to a first duration threshold, then in the second multimedia data set, the second multimedia data whose playback duration best matches the first trip duration is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen; if the first trip duration is greater than the first duration threshold, then in the second multimedia data set, the second multimedia data whose playback duration is less than or equal to a second duration threshold is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen, wherein the second duration threshold is less than or equal to the first duration threshold.
[0010] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the vehicle includes at least one display screen, and the step of playing the first multimedia recommendation data on the vehicle's display screen includes: determining, based on each first profile information, a first target display screen corresponding to the first multimedia recommendation data and a sound field playback strategy for the first target display screen in all display screens of the vehicle; displaying the first multimedia recommendation data on the first target display screen; and playing the multimedia recommendation data in a first sound zone corresponding to the first target display screen based on the sound field playback strategy.
[0011] In combination with the first aspect and the above implementation methods, in some possible implementation methods, after determining the first multimedia recommendation data in the second multimedia data set and playing the first multimedia recommendation data on the vehicle's display screen, the method further includes: receiving interactive operations from each passenger regarding the first multimedia recommendation data; determining the second target passenger corresponding to the interactive operation based on the multimodal perception device; and updating the target preference label of the second target passenger based on the interactive operation.
[0012] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, after determining the first multimedia recommendation data in the second multimedia data set and playing the first multimedia recommendation data on the vehicle's display screen, the method further includes: during the playback of the first multimedia recommendation data, acquiring second facial images and second audio information of each occupant through a multimodal sensing device; determining the second physiological state of each occupant based on each second facial image and each second audio information; and adjusting the first multimedia recommendation data based on each second physiological state.
[0013] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, the vehicle includes multiple displays, and the method further includes: acquiring second profile information of a third target occupant in the vehicle, wherein the third target occupant is any occupant in the vehicle; determining second data recommendation features corresponding to the third target occupant and a second target display screen corresponding to the seating position of the third target occupant based on the second profile information; determining a second travel duration of the vehicle based on the vehicle's navigation system; filtering a third multimedia data set corresponding to the second data recommendation features based on the second travel duration to obtain a fourth multimedia data set; determining second multimedia recommendation data in the fourth multimedia data set; displaying the second multimedia recommendation data on the second target display screen; and playing the second multimedia recommendation data in the second audio zone corresponding to the second target display screen.
[0014] Secondly, a multimedia data recommendation device is provided, the device comprising: The recommendation feature determination unit is used to obtain the first profile information of multiple occupants in the vehicle, and determine the first data recommendation features that conform to the common preferences of each occupant based on the first profile information. The multimedia data filtering unit is used to determine the first trip duration of the vehicle based on the vehicle's navigation system, and to filter the first multimedia data set corresponding to the first data recommendation features based on the first trip duration to obtain the second multimedia data set. The multimedia data playback unit is used to determine first multimedia recommendation data from a second multimedia data set and play the first multimedia recommendation data on a first target display screen, wherein the first target display screen is any display screen among all the display screens of the vehicle.
[0015] Thirdly, a vehicle is provided, the vehicle including: a memory for storing executable program code; A processor for calling and running executable program code from memory to perform the methods in the first aspect or any possible implementation of the first aspect described above.
[0016] Fourthly, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0017] Fifthly, a computer-readable storage medium is provided that stores computer program code, which, when executed on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof. Attached Figure Description
[0018] Figure 1 This is a system architecture diagram of a multimedia data recommendation method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a multimedia data recommendation method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a multimedia data recommendation method provided in an embodiment of this application; Figure 4 This is a system architecture diagram of a multimedia data recommendation method provided in an embodiment of this application; Figure 5 This is a flowchart illustrating a multimedia data recommendation method provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a multimedia data recommendation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0020] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0021] Please see Figure 1 , Figure 1 This is a system architecture diagram of a multimedia data recommendation method provided in an embodiment of this application. For example... Figure 1 As shown in the embodiments of this application, the multimedia data recommendation method is applied to a scenario where an in-vehicle terminal recommends multimedia data to occupants in a vehicle. The in-vehicle terminal is an electronic device in the vehicle with data acquisition and processing capabilities. It establishes communication connections with functional modules such as the navigation system, display screen, in-vehicle audio system, and multimodal sensing devices in the vehicle through communication protocols such as the Controller Area Network (CAN) bus. This allows it to acquire various information about the vehicle and its occupants, analyze their interests and preferences based on the acquired occupant information, and then recommend multimedia data accordingly. The multimodal sensing devices include camera equipment and audio recording equipment. The camera equipment is used to capture images inside and outside the vehicle, and the audio recording equipment is used to capture audio inside and outside the vehicle.
[0022] When there are multiple passengers in a vehicle, due to differences in their preferences and the duration of the journey, the in-vehicle terminal may not be able to meet the entertainment needs of multiple passengers during the journey. This results in poor compatibility between the recommended multimedia data and the individual passengers, thus reducing the driving experience.
[0023] To address the aforementioned issues, this application provides a multimedia data recommendation method. Based on the profiles of multiple occupants in a vehicle, it determines data recommendation features that align with the shared preferences of these occupants. Simultaneously, it determines the trip duration based on the vehicle's navigation system. From a first multimedia data set corresponding to the data recommendation features, it filters out a second multimedia data set that meets the trip duration. Multimedia recommendation data is then determined from the second multimedia data set and played on the vehicle's display screen. By acquiring data recommendation features shared by multiple occupants, the adaptability of the recommended multimedia data to multiple occupants is improved. By filtering multimedia data that can be viewed in its entirety based on the trip duration, the adaptability of the recommended multimedia data to the vehicle's driving scenario in the time dimension is improved, ensuring that occupants receive a complete viewing / listening experience. This method enhances the overall satisfaction of multiple occupants in the vehicle and guarantees a positive driving experience.
[0024] based on Figure 1 The system architecture diagram shown below will be used in conjunction with... Figures 2-5 This application provides a detailed description of the multimedia data recommendation method provided in its embodiments.
[0025] Please see Figure 2 This is a flowchart illustrating a multimedia data recommendation method provided in an embodiment of this application. Figure 2 As shown, the method in this application embodiment may include the following steps S101-S103.
[0026] S101, Obtain first profile information of multiple occupants in the vehicle, and determine first data recommendation features that conform to the common preferences of each occupant based on each first profile information; Specifically, the process involves acquiring the first profile information of multiple occupants in the vehicle, and determining the first data recommendation features that match the common preferences of all occupants based on the acquired first profile information of all occupants. In the application scenario of vehicles with multiple passengers, multiple occupants can be all occupants in the vehicle; or they can be multiple occupants viewing the same display screen. For example, when the vehicle includes multiple display screens such as the main control screen, the passenger screen, and the ceiling screen, the ceiling screen is used for rear-seat occupants to view together, thus acquiring the first profile information of multiple rear-seat occupants.
[0027] The first profile information includes the passenger's personal information and physiological state. Personal information includes, but is not limited to, the passenger's gender and age, representing their static identity attributes. Physiological state includes the passenger's facial expressions and sleep status, representing their dynamic physical and mental attributes. This profile information enables the in-vehicle terminal to perform preference analysis on the passengers and determine each passenger's target preference label. The data recommendation feature is a multimedia data feature that can satisfy the target preference labels of all passengers. For example, if passenger A is an 8-10 year old boy, not asleep, and his facial expression indicates he is excited, the in-vehicle terminal's preference analysis for passenger A yields a target preference label of "cartoons." If passenger B is a 30-35 year old woman, not asleep, and her facial expression indicates she is relaxed, the in-vehicle terminal's preference analysis for passenger B yields a target preference label of "light music." Therefore, the data recommendation feature could be music-themed cartoons.
[0028] S102, based on the vehicle's navigation system, determine the first trip duration of the vehicle, and based on the first trip duration, filter the first multimedia data set corresponding to the first data recommendation feature to obtain the second multimedia data set; Specifically, the vehicle includes a navigation system. During the vehicle's journey, the navigation system determines the first trip duration. Based on this duration, it filters and processes each piece of first multimedia data in the first multimedia data set corresponding to the first data recommendation features, resulting in a second multimedia data set. This second multimedia data set consists of at least one piece of second multimedia data. Multimedia data includes, but is not limited to, video, music, podcasts, and radio. After determining the first data recommendation features, the in-vehicle terminal identifies at least one piece of first multimedia data that matches these features and combines these pieces of first multimedia data into the first multimedia data set. To ensure that all passengers have a complete viewing experience of the multimedia data recommended by the in-vehicle terminal before the end of the journey, first multimedia data in the first multimedia data set with a playback duration less than or equal to the first trip duration is designated as second multimedia data.
[0029] S103, determine the first multimedia recommendation data from the second multimedia data set, and play the first multimedia recommendation data on the vehicle's display screen.
[0030] Specifically, first multimedia recommended data is determined from the second multimedia data set, and the first multimedia recommended data is played on the vehicle's display screen. Optionally, the in-vehicle terminal can display all second multimedia data from the second multimedia data set on the display screen, and receive the first multimedia recommended data selected by the occupant based on the display screen; alternatively, it can calculate a weighted score for each second multimedia data based on the first data recommendation features, and select the second multimedia data with the highest score as the first multimedia recommended data; or it can select the second multimedia data whose playback duration best matches the first trip duration as the first multimedia recommended data. This application embodiment does not limit this. When multiple occupants are all occupants in the vehicle, the display screen playing the first multimedia recommended data can be all display screens in the vehicle; it can also be any display screen other than the main control screen, to prevent the first multimedia recommended data from interfering with the driver's driving and causing a safety accident. When multiple occupants are multiple occupants viewing the same display screen, the first profile information can also include the seating position of each occupant, and the display screen playing the first multimedia recommended data is the target display screen corresponding to that seating position.
[0031] In this embodiment, data recommendation features that match the common preferences of multiple occupants are determined based on their profile information. Simultaneously, the trip duration is determined based on the vehicle's navigation system. A second multimedia data set that meets the trip duration is selected from the first multimedia data set corresponding to the data recommendation features. Multimedia recommendation data is then determined from the second multimedia data set and played on the vehicle's display screen. By acquiring data recommendation features that match the common preferences of multiple occupants, the adaptability of the recommended multimedia data to multiple occupants is improved. By selecting multimedia data that can be viewed in its entirety based on the trip duration, the adaptability of the recommended multimedia data to the vehicle's driving scenario in the time dimension is improved, ensuring that occupants receive a complete viewing / listening experience. This method improves the group satisfaction of multiple occupants in the vehicle and guarantees their driving experience.
[0032] Please see Figure 3 This is a flowchart illustrating a multimedia data recommendation method provided in an embodiment of this application. Figure 3 As shown, the method in this application embodiment may include the following steps S201-S209.
[0033] S201, Based on the multimodal perception device, acquire the first facial images and first audio information of multiple occupants in the vehicle; Specifically, the vehicle in this application embodiment is equipped with a multimodal sensing device, which includes a camera and a microphone. The camera acquires first facial images of multiple occupants in the vehicle, and the microphone acquires first audio information of the multiple occupants in the vehicle. In a multi-passenger vehicle application scenario, the multiple occupants can be all the occupants in the vehicle; or they can be multiple occupants viewing the same display screen. For example, when the vehicle includes multiple display screens such as a main control screen, a passenger-side screen, and a ceiling-mounted screen, the ceiling-mounted screen is used for rear-seat occupants to view together, thereby acquiring the first portrait information of the multiple rear-seat occupants.
[0034] S202, based on each first facial image and each first audio information, determine the personal information and first physiological state of each occupant; Specifically, the vehicle-mounted terminal is equipped with a multimodal large-scale model pre-trained with a large amount of sample data. This multimodal large-scale model can simultaneously understand and process image information, audio information, and text information. In this embodiment, the multimodal large-scale model extracts facial features such as skin texture and expression changes of each occupant from the first facial image, and extracts acoustic features such as voiceprint, speech rate, yawning, and breathing of each occupant from the first audio information. The facial features and acoustic features are fused across modally to determine the personal information and first physiological state of each occupant. The personal information includes, but is not limited to, information such as the occupant's gender and age, used to characterize the occupant's static identity attributes; the physiological state includes information such as the occupant's facial expressions and sleep status, used to characterize the occupant's dynamic physical and mental attributes.
[0035] S203, determine the target preference labels for each crew member based on personal information and primary physiological state; Specifically, after determining the personal information and primary physiological state of each occupant, the multimodal large model infers from these information to analyze each occupant's preferences and determine their target preference labels. The in-vehicle terminal also includes a memory system database, which maps and stores the occupant identifiers and characteristics of historical occupants. Occupant characteristics include facial features, acoustic features, and historical preference labels, which are preferences determined based on the occupant's historical interaction with multimedia data. The in-vehicle terminal can periodically clean the memory system database to prevent storage overload.
[0036] In one feasible implementation, if the personal information characterizes the first target occupant as a historical occupant of the vehicle, then the historical preference tag of the first target occupant is obtained. The first target occupant can be any one of multiple occupants. The target preference tag of the first target occupant is determined based on the historical preference tag and the first physiological state. When the multimodal large model acquires the occupant's facial and acoustic features, it matches these features with the occupant features stored in the memory system database to determine if the occupant is a historical occupant. If the occupant is not a historical occupant, an occupant identifier is assigned to the occupant in the memory system database, and the occupant identifier and occupant features are mapped and stored. The occupant's operational behavior regarding multimedia data in the vehicle is continuously acquired, and preference tags are generated and stored based on the operational behavior as historical preference tags for reference in the memory system database. If the occupant is a historical occupant, the current target preference tag can be determined by combining the occupant's historical preference tag and their current physiological state.
[0037] S204, Based on the target preference labels of each occupant, determine the first data recommendation feature that conforms to the common preferences of each occupant; Specifically, the target preference tags of each passenger are fused to determine the first data recommendation feature that matches the common preferences of all passengers. The data recommendation feature is a multimedia data feature that can satisfy the target preference tags of all passengers. For example, passenger A is an 8-10 year old boy who is not asleep and whose facial expression indicates that passenger A is in an excited state. The in-vehicle terminal performs preference analysis on passenger A and obtains that passenger A's target preference tag is "cartoons". Passenger B is a 30-35 year old woman who is not asleep and whose facial expression indicates that passenger B is in a relaxed state. The in-vehicle terminal performs preference analysis on passenger B and obtains that passenger B's target preference tag is "light music". Then the data recommendation feature can be music-themed cartoons.
[0038] In one feasible implementation, the intersection of the target preference labels of each passenger is determined as the first data recommendation feature that conforms to the common preferences of each passenger. For example, passengers C and D are both boys aged 8-10 years old and are both asleep. According to the radio equipment, passengers C and D are discussing anime E. Therefore, it is determined that the target preference label of passengers C and D is "anime E", and the data recommendation feature can be the name of anime E.
[0039] S205, determines the duration of the vehicle's first journey based on the vehicle's navigation system; Specifically, the in-vehicle terminal also includes an audio-visual management module, which obtains the vehicle's trip duration from the navigation system. To ensure that all passengers have a complete viewing experience of the multimedia data recommended by the in-vehicle terminal before the end of the trip, the trip duration indicated by the navigation system is used as an important factor in multimedia data recommendation. The audio-visual management module obtains the first data recommendation features returned by the multimodal big data model, generates prompt words based on the first data recommendation features and the first trip duration, and inputs the prompt words into the multimodal big data model.
[0040] S206, Obtain the first multimedia data set corresponding to the first data recommendation feature, wherein the first multimedia data set is composed of each first multimedia data; Specifically, the in-vehicle terminal also includes a multimedia module, which contains a multimedia resource database. After receiving prompts from the audio-visual management module, the multimodal large model first identifies at least one piece of first multimedia data that matches the first data recommendation characteristics through the multimedia module, and then combines these first pieces of multimedia data into a first multimedia data set. Multimedia data includes, but is not limited to, video, music, podcasts, radio, and other data.
[0041] S207, in the first multimedia data set, the first multimedia data whose playback duration is less than or equal to the first trip duration is determined as the second multimedia data, and the second multimedia data set is determined based on each second multimedia data; Specifically, after determining the first multimedia data set, the multimodal big data model filters the first multimedia data set according to the duration of the trip in the prompt words, and determines the first multimedia data whose playback duration is less than or equal to the duration of the first trip as the second multimedia data, and determines the second multimedia data set based on each second multimedia data.
[0042] S208, if the first trip duration is less than or equal to the first duration threshold, then in the second multimedia data set, the second multimedia data whose playback duration best matches the first trip duration is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen; Specifically, the prompts sent by the audio-visual management module also include constraints on a first duration threshold. After determining the second multimedia data set, the multimodal large model filters the final first multimedia recommendation data from the second multimedia data set based on a comparison between the first trip duration and the first duration threshold. If the first trip duration is less than or equal to the first duration threshold, the second multimedia data with the playback duration closest to the first trip duration is determined as the most suitable first multimedia recommendation data for the first trip duration and played on the vehicle's display screen. The first duration threshold is a pre-set time boundary point used to distinguish between short and long trips. For example, the first duration threshold can be 30 minutes. When the first trip duration is less than or equal to 30 minutes, the trip is determined to be a short trip; when the first trip duration is greater than 30 minutes, the trip is determined to be a long trip. In short trips, passengers tend to experience complete, continuous, and duration-matched multimedia data. Therefore, in the second multimedia data set, the second multimedia data with the playback duration closest to the first trip duration is determined as the first multimedia recommendation data.
[0043] In short trips, when there are many second multimedia data whose playback duration matches the duration of the first trip, the multimodal big data model can calculate the weighted score of each second multimedia data whose playback duration is close to the trip duration based on the recommendation features of the first data, and select the second multimedia data with the highest score as the first multimedia recommendation data.
[0044] S209, if the duration of the first trip is greater than the first duration threshold, then in the second multimedia data set, the second multimedia data with a playback duration less than or equal to the second duration threshold is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen.
[0045] Specifically, when the first trip duration exceeds a first duration threshold, the trip is defined as a long trip. In the multimedia data set, the second multimedia data with a playback duration less than or equal to a second duration threshold is selected as the final first multimedia recommended data, and this first multimedia recommended data is played on the vehicle's display screen. During long trips, passengers are prone to decreased attention or boredom when watching a single type of content for extended periods. Shorter content is easier to digest, and breaks or switching between content can be naturally introduced during playback intervals. The second duration threshold is less than or equal to the first duration threshold. For example, if the first duration threshold is 30 minutes and the trip duration is 2 hours, the second duration threshold is 20 minutes. The second multimedia data with a playback duration less than or equal to 20 minutes in the second multimedia data set is selected as the first multimedia recommended data, and the first multimedia recommended data is played sequentially. After each first multimedia recommended data is played, there can be a 5-minute pause before playing the next first multimedia recommended data. If the first multimedia recommended data is all in video format, passengers other than the driver can be prompted to relax their eyes during the 5-minute break, or light music or humorous audio can be played to provide a soothing and relaxing experience.
[0046] In long runs, when there is a large amount of first multimedia recommendation data, the multimodal large model can calculate the weighted score of each first multimedia recommendation data according to the first data recommendation features, and sort the first multimedia recommendation data from high to low according to the weighted score.
[0047] When multiple occupants refer to all occupants in the vehicle, the display screen playing the first multimedia recommendation data can be any display screen in the vehicle; or it can be any display screen other than the main control screen, to prevent the first multimedia recommendation data from interfering with the driver's driving and causing a safety accident. When multiple occupants refer to multiple occupants viewing the same display screen, the first profile information can also include the seating position of each occupant, and the display screen playing the first multimedia recommendation data is the target display screen corresponding to that seating position.
[0048] In one feasible implementation, when the vehicle includes at least one display screen, each display screen corresponds to its own audio zone. The vehicle includes an in-vehicle audio system, and the audio zone is an audio playback channel, such as a speaker or Bluetooth headset, separately allocated by the in-vehicle audio system for each display screen to ensure that the audio played on each display screen does not interfere with each other. Among all the display screens in the vehicle, a first target display screen corresponding to the first multimedia recommendation data and a sound field playback strategy for the first target display screen are determined based on each first profile information. The first multimedia recommendation data is displayed on the first target display screen, and the multimedia recommendation data is played in the first audio zone corresponding to the first target display screen based on the sound field playback strategy. The sound field playback strategy refers to the in-vehicle terminal sending the first physiological state and seating position of each occupant to the in-vehicle audio system, so that the in-vehicle audio system can control the sound field playback according to the current physiological state of the occupant in each seating position. The sound field playback control includes front-rear balance and left-right balance. The front-rear balance is used to adjust the intensity distribution of the multimedia recommendation data between the front and rear seats of the vehicle, and the left-right balance is used to adjust the intensity distribution of the sound between the left and right sides of the vehicle. For example, if the first image information indicates that multiple occupants are all rear-seat passengers, then the ceiling-mounted screen is identified as the first target display screen, and the vehicle audio system adjusts the balance backward. Assuming there are three rear-seat passengers, and the leftmost passenger's first physiological state indicates they are asleep, the vehicle audio system adjusts the balance to the right. It should be noted that the display screen playing the first multimedia recommendation data is determined by the audio-visual management module and sent to the multimedia module, so that the multimedia module displays the first multimedia recommendation data through the corresponding display screen and plays the first multimedia recommendation data through the corresponding audio zone of that display screen.
[0049] In one feasible implementation, during the playback of the first multimedia recommendation data, the in-vehicle terminal receives interactive operations from each passenger regarding the first multimedia recommendation data via a display screen or a terminal device connected to the display screen; based on a multimodal perception device, a second target passenger corresponding to the interactive operation is determined, and the target preference tag of the second target passenger is updated based on the interactive operation. After updating the target preference tag of the second target passenger, the historical preference tag of the second target passenger is updated or added in the memory system database according to the passenger identifier of the second target passenger.
[0050] In one feasible implementation, during the playback of the first multimedia recommendation data, the in-vehicle terminal can also acquire second facial images and second audio information of each occupant through a multimodal sensing device; determine the second physiological state of each occupant based on each second facial image and each second audio information; and adjust the first multimedia recommendation data based on each second physiological state. For example, if it is detected that three passengers in the back row have fallen asleep while watching the first multimedia recommendation data, the first multimedia recommendation data can be turned off, or switched to soothing soft music, or the playback volume of the first multimedia recommendation data can be reduced.
[0051] Please see Figure 4 , Figure 4 This is a system architecture diagram of a multimedia data recommendation method provided in an embodiment of this application. For example... Figure 4 As shown, the in-vehicle terminal includes a multimedia module, an audio-visual management module, a multimodal big data model, and a memory system database. The multimedia module provides multimedia data to the multimodal big data model, receives multimedia recommendation data returned by the model, and sends the playback status of the recommended multimedia data to the model. The multimedia module is connected to the display screen and the in-vehicle audio system to display multimedia recommendation devices on the display screen and play multimedia recommendation data through the in-vehicle audio system. The audio-visual management module obtains the trip duration from the navigation system, generates prompt words based on the trip duration and the first data recommendation features returned by the multimodal big data model, and sends the prompt words to the multimodal big data model. The multimodal big data model analyzes and processes the facial images and audio information of the occupants obtained by the camera and audio equipment to obtain the occupants' personal information and physiological state, thereby determining the target preference tags of each member, and fusing the target preference tags of all occupants to determine the first data recommendation features. The multimodal large model is also used to send the first data recommendation features to the audio-visual association module to receive prompts from the audio-visual management module based on the first data recommendation features and the trip duration. Based on the prompts, it filters and determines multimedia recommendation data in the multimedia module. The multimodal large model is also used to obtain the multimedia recommendation data sent by the multimedia module and the playback status of the multimedia recommendation data when playing recommended multimedia videos. It then combines the user's current facial and acoustic features to update or add to the user's historical preferences stored in the memory system database in real time. The memory system database is used to map and store occupants and their facial features, acoustic features, and historical preferences.
[0052] In this embodiment, facial images and audio information of occupants are acquired through multimodal sensing devices, improving the accuracy of distinguishing occupant personal information and physiological states. By fusing and reasoning the personal information and physiological states of multiple occupants using a multimodal large-scale model, the raw sensor data is transformed into understandable occupant intentions to determine each occupant's target preference tags. All target preference tags are then fused to determine data recommendation features shared by multiple occupants, improving the suitability of recommended multimedia data to multiple occupants. Multimedia data that can be viewed in its entirety is filtered by trip duration, improving the time-dimensional suitability of recommended multimedia data to the vehicle's driving scenario and ensuring occupants receive a complete viewing / listening experience. By combining data recommendation features shared by multiple occupants with the vehicle's trip duration, the overall satisfaction of multiple occupants in the vehicle is improved, guaranteeing a positive driving experience. When an occupant is a historical occupant, their current target preference tag is determined by combining their historical preference tags, enhancing the intelligence of multimedia data recommendation. The vehicle journey is divided into short and long trips based on a pre-set first duration threshold. Different multimedia recommendation data is determined for different trip types, improving the adaptability of the multimedia data recommendation method to various vehicle application scenarios and further enhancing the passenger experience. By dividing each display screen into sound zones and positioning each passenger in the optimal listening position according to the sound field playback strategy, the passenger listening experience is improved. While playing multimedia recommendation data, the system obtains real-time information on each passenger's interaction or physiological state, thereby adjusting user preference tags or multimedia recommendation data, further enhancing the intelligence of the multimedia data recommendation.
[0053] Please see Figure 5 This is a flowchart illustrating a multimedia data recommendation method provided in an embodiment of this application. Figure 5 As shown, the method in this application embodiment may include the following steps S301-S304.
[0054] S301, Obtain the second profile information of the third target occupant in the vehicle; Specifically, the multimedia data recommendation method provided in this application is applied to a multimedia data recommendation scenario for a single display screen among multiple display screens. The vehicle includes a multimodal perception device, which acquires a second profile of a third target occupant in the vehicle. The third target occupant is at least one occupant at a seating position corresponding to any display screen in the vehicle. The multimodal perception device includes a camera device and a sound recording device. The camera device acquires a third facial image and seating position of the third target occupant, and the sound recording device acquires third audio information of the third target occupant.
[0055] The vehicle provided in this application embodiment includes multiple displays, such as a main control screen, a passenger-side screen, and a ceiling-mounted screen. Different displays are used for viewing by occupants in different seating positions. Therefore, the second profile information includes the occupant's personal information, physiological state, and seating position. The personal information includes, but is not limited to, the occupant's gender, age, etc., which are used to characterize the occupant's static identity attributes. The physiological state includes, but is not limited to, the occupant's facial expressions, sleep state, etc., which are used to characterize the occupant's dynamic physical and mental attributes. The seating position is used to determine the second target display screen viewed by the third target occupant.
[0056] S302, based on the second profile information, determine the second data recommendation features corresponding to the third target occupant, and the second target display screen corresponding to the seating position of the third target occupant; Specifically, the personal information and physiological status in the second profile information are analyzed to determine the second data recommendation features corresponding to the third target passenger, as well as the second target display screen corresponding to the third target passenger's seating position. When there is only one third target passenger, the data recommendation features are multimedia data features that can satisfy the preference tags of the third target passenger; when there are at least two third target passengers, the data recommendation features are multimedia data features that conform to the common preferences of the third target passengers.
[0057] S303, based on the vehicle's navigation system, determine the second trip duration of the vehicle, and based on the second trip duration, filter the third multimedia data set corresponding to the second data recommendation features to obtain the fourth multimedia data set; Specifically, the vehicle includes a navigation system. During the vehicle's journey, the navigation system determines the second trip duration. Based on the trip duration, it filters and processes each third multimedia data in the third multimedia data set corresponding to the second data recommendation features to obtain a fourth multimedia data set. The fourth multimedia data set consists of at least one fourth multimedia data. Multimedia data includes, but is not limited to, video, music, podcasts, and radio. After determining the second data recommendation features, the in-vehicle terminal identifies at least one third multimedia data that matches the second data recommendation features from the local multimedia resource database and / or the cloud multimedia resource database, and combines these third multimedia data into a third multimedia data set. To ensure that all passengers have a complete viewing experience of the multimedia data recommended by the in-vehicle terminal before the end of the trip, third multimedia data with a playback duration less than or equal to the second trip duration is identified as fourth multimedia data in the third multimedia data set.
[0058] S304, determine the second multimedia recommendation data in the fourth multimedia data set, display the second multimedia recommendation data on the second target display screen, and play the second multimedia recommendation data in the second audio zone corresponding to the second target display screen.
[0059] Specifically, second multimedia recommendation data is determined from the fourth multimedia data set, displayed on the second target display screen, and played in the second audio zone corresponding to the second target display screen. When the vehicle includes at least one display screen, each display screen has its own audio zone. The vehicle includes an in-vehicle audio system, and the audio zone is an audio playback channel, such as a speaker or Bluetooth headset, separately allocated by the in-vehicle audio system for each display screen to ensure that the audio played on each display screen does not interfere with each other.
[0060] Optionally, the vehicle terminal can display all the fourth multimedia data in the fourth multimedia data set on the second target display screen, and receive the second multimedia recommendation data selected by the third target occupant based on the second target display screen; it can also calculate the weighted score of each fourth multimedia data according to the second data recommendation features, and select the fourth multimedia data with the highest score as the second multimedia recommendation data; it can also select the fourth multimedia data whose playback duration is most suitable for the second trip duration as the second multimedia recommendation data, and this application embodiment does not limit this.
[0061] In this embodiment, when the vehicle includes multiple displays, the second target display screen to be viewed by the third target occupant is determined based on their seating position. Data recommendation features matching the occupant's preferences are determined based on their profile information. Simultaneously, the trip duration is determined based on the vehicle's navigation system. A fourth multimedia data set that meets the trip duration is selected from the third multimedia data set corresponding to the data recommendation features. Multimedia recommendation data is then determined from the fourth multimedia data set and played on the second target display screen. By determining the multimedia recommendation data corresponding to a single display screen using the occupant's profile information and the trip duration from the navigation system, different multimedia recommendation data can be provided for different display screens, achieving personalized recommendations.
[0062] based on Figure 1 The system architecture diagram will be presented below, in conjunction with... Figure 6 This application provides a detailed description of the multimedia data recommendation device provided in its embodiments. It should be noted that... Figure 6 The multimedia data recommendation device in the present application is used to perform the following tasks. Figures 2-5 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 2-5 The example shown.
[0063] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a multimedia data recommendation device provided in an embodiment of this application. Figure 6As shown, the multimedia data recommendation device 1 of this application embodiment may include: a recommendation feature determination unit 11, a multimedia data filtering unit 12, and a multimedia data playback unit 13.
[0064] The recommendation feature determination unit 11 is used to obtain first profile information of multiple occupants in the vehicle and determine first data recommendation features that conform to the common preferences of each occupant based on each first profile information. The multimedia data filtering unit 12 is used to determine the first trip duration of the vehicle based on the vehicle's navigation system, and to filter the first multimedia data set corresponding to the first data recommendation feature based on the first trip duration to obtain the second multimedia data set. The multimedia data playback unit 13 is used to determine the first multimedia recommendation data in the second multimedia data set and play the first multimedia recommendation data on the first target display screen, which is any display screen among all the display screens of the vehicle.
[0065] Optionally, a multimodal perception device is installed in the vehicle, and the recommended feature determination unit 11 is specifically used to acquire first facial images and first audio information of multiple occupants in the vehicle based on the multimodal perception device; Based on each first facial image and each first audio information, determine the personal information and first physiological state of each occupant; Determine the target preference labels for each occupant based on personal information and primary physiological state; Based on the target preference labels of each passenger, the first data recommendation feature that matches the common preferences of each passenger is determined.
[0066] Optionally, the recommended feature determination unit 11 is specifically used to obtain the historical preference label of the first target occupant if the personal information characterization of the first target occupant is a historical occupant of the vehicle, and the first target occupant is any one of multiple occupants; The target preference label of the first target occupant is determined based on historical preference labels and the first physiological state.
[0067] Optionally, the multimedia data filtering unit 12 is specifically used to obtain the first multimedia data set corresponding to the first data recommendation feature, and the first multimedia data set is composed of each first multimedia data. In the first multimedia data set, the first multimedia data whose playback duration is less than or equal to the first trip duration is determined as the second multimedia data, and the second multimedia data set is determined based on each second multimedia data.
[0068] Optionally, the multimedia data playback unit 13 is specifically used to determine the second multimedia data whose playback duration best matches the first trip duration in the second multimedia data set as the first multimedia recommended data if the first trip duration is less than or equal to the first duration threshold, and play the first multimedia recommended data on the vehicle's display screen. If the duration of the first trip is greater than the first duration threshold, then in the second multimedia data set, the second multimedia data with a playback duration less than or equal to the second duration threshold is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen, where the second duration threshold is less than or equal to the first duration threshold.
[0069] Optionally, the vehicle includes at least one display screen, and the multimedia data playback unit 13 is specifically used to determine, based on each first profile information, the first target display screen corresponding to the first multimedia recommendation data and the sound field playback strategy of the first target display screen in all the display screens of the vehicle. First multimedia recommendation data is displayed on the first target display screen, and multimedia recommendation data is played in the first sound zone corresponding to the first target display screen based on the sound field playback strategy.
[0070] Optionally, the multimedia data recommendation device 1 is specifically used to receive interactive operations from each passenger regarding the first multimedia recommendation data; The second target occupant corresponding to the interaction operation is determined based on the multimodal sensing device, and the target preference label of the second target occupant is updated based on the interaction operation.
[0071] Optionally, the multimedia data recommendation device 1 is specifically used to acquire the second facial image and second audio information of each occupant through a multimodal sensing device during the playback of the first multimedia recommendation data. The second physiological state of each occupant is determined based on each second facial image and each second audio information; The first multimedia recommendation data is adjusted based on each second physiological state.
[0072] Optionally, the vehicle includes at least one display screen, and the multimedia data recommendation device 1 is specifically used to obtain the second profile information of a third target occupant in the vehicle, wherein the third target occupant is any occupant in the vehicle; Based on the second profile information, the second data recommendation features corresponding to the third target occupant are determined, as well as the second target display screen corresponding to the seating position of the third target occupant; The second trip duration of the vehicle is determined based on the vehicle's navigation system. Based on the second trip duration, the third multimedia data set corresponding to the second data recommendation features is filtered to obtain the fourth multimedia data set. The second multimedia recommendation data is determined from the fourth multimedia data set, displayed on the second target display screen, and played in the second audio zone corresponding to the second target display screen.
[0073] In this embodiment, by acquiring facial images and audio information of occupants, the accuracy of distinguishing occupant personal information and physiological states is improved. By fusing and reasoning about the personal information and physiological states of multiple occupants, the raw sensor data is transformed into understandable occupant intentions to determine the target preference tags of each occupant. All target preference tags are then fused to determine data recommendation features shared by multiple occupants, improving the suitability of the recommended multimedia data to multiple occupants. Multimedia data that can be viewed in its entirety is filtered by trip duration, improving the time-dimensional suitability of the recommended multimedia data to the vehicle's driving scenario, ensuring occupants receive a complete viewing / listening experience. By combining data recommendation features shared by multiple occupants with the vehicle's trip duration, the overall satisfaction of multiple occupants in the vehicle is improved, guaranteeing a positive driving experience. When an occupant is a historical occupant, their current target preference tag is determined by combining their historical preference tags, enhancing the intelligence of multimedia data recommendation. The vehicle trip is divided into short and long trips based on a pre-set first duration threshold. Different multimedia recommendation data is determined for different trip types, improving the adaptability of the multimedia data recommendation method to various vehicle application scenarios and further enhancing the passenger experience. By dividing each display screen into sound zones and positioning each passenger in the optimal listening position according to the sound field playback strategy, the passenger listening experience is improved. While playing multimedia recommendation data, the interactive operations or physiological states of each passenger are acquired in real time, thereby adjusting user preference tags or multimedia recommendation data, further enhancing the intelligence of multimedia data recommendation. The multimedia recommendation data corresponding to each display screen is determined by the passenger profile information in front of each individual display screen and the trip duration of the navigation system, allowing for different multimedia recommendation data for different display screens, achieving personalized recommendations.
[0074] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application.
[0075] For example, such as Figure 7 As shown, the vehicle 700 includes a processor 701 and a memory 702, wherein the processor 701 is electrically connected to the memory 702.
[0076] The processor 701 is the control center of the vehicle 700 and may include one or more processing cores. The processor 701 connects to various parts of the vehicle via various interfaces and lines, executing various vehicle functions and processing data by running or calling computer programs stored in the memory 702 and calling data stored in the memory 702, thereby providing overall control of the vehicle 700. Optionally, the processor 701 may be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 701 may integrate one or more of the following: CPU, Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user page, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 701 and may be implemented separately using a communication chip.
[0077] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the computer programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function, etc.; the data storage area may store data created based on the use of the vehicle 700, etc.
[0078] Furthermore, memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 702 may also include a memory controller to provide processor 701 with access to memory 702.
[0079] In this embodiment, the processor 701 in the vehicle 700 loads the instructions corresponding to the processes of one or more computer programs into the memory 702 according to the following steps, and the processor 701 runs the computer programs stored in the memory 702 to realize various functions, as follows: Obtain first profile information of multiple occupants in the vehicle, and determine first data recommendation features that conform to the common preferences of each occupant based on each first profile information; The first trip duration of the vehicle is determined based on the vehicle's navigation system. Based on the first trip duration, the first multimedia data set corresponding to the first data recommendation features is filtered to obtain the second multimedia data set. First multimedia recommendation data is determined from the second multimedia data set, and the first multimedia recommendation data is played on the vehicle's display screen.
[0080] Optionally, the vehicle is equipped with a multimodal perception device. When the processor 701 acquires first profile information of multiple occupants in the vehicle and determines first data recommendation features that conform to the common preferences of each occupant based on the first profile information, it specifically performs the following: First facial images and first audio information of multiple occupants in a vehicle are acquired using a multimodal sensing device; Based on each first facial image and each first audio information, determine the personal information and first physiological state of each occupant; Determine the target preference labels for each occupant based on personal information and primary physiological state; Based on the target preference labels of each passenger, the first data recommendation feature that matches the common preferences of each passenger is determined.
[0081] Optionally, when processor 701 determines the target preference labels for each occupant based on personal information and initial physiological state, it specifically performs the following: If the personal information indicates that the first target occupant is a historical occupant of the vehicle, then the historical preference tag of the first target occupant is obtained. The first target occupant can be any one of multiple occupants. The target preference label of the first target occupant is determined based on historical preference labels and the first physiological state.
[0082] Optionally, when processor 701 performs filtering processing on the first multimedia data set corresponding to the first data recommendation features based on the first trip duration to obtain the second multimedia data set, it specifically executes the following: Obtain the first multimedia data set corresponding to the first data recommendation feature, wherein the first multimedia data set is composed of each first multimedia data; In the first multimedia data set, the first multimedia data whose playback duration is less than or equal to the first trip duration is determined as the second multimedia data, and the second multimedia data set is determined based on each second multimedia data.
[0083] Optionally, when the processor 701 determines the recommended multimedia data from the second multimedia data set and plays the recommended multimedia data on the vehicle's display screen, it specifically performs the following: If the duration of the first trip is less than or equal to the first duration threshold, then the second multimedia data with the most suitable playback duration in the second multimedia data set is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen. If the duration of the first trip is greater than the first duration threshold, then in the second multimedia data set, the second multimedia data with a playback duration less than or equal to the second duration threshold is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the vehicle's display screen, where the second duration threshold is less than or equal to the first duration threshold.
[0084] Optionally, the vehicle includes at least one display screen, and when the processor 701 executes the action of playing first multimedia recommendation data on the vehicle's display screen, it specifically performs the following: Based on the first profile information, the first target display screen corresponding to the first multimedia recommendation data and the sound field playback strategy of the first target display screen are determined among all the display screens in the vehicle. First multimedia recommendation data is displayed on the first target display screen, and multimedia recommendation data is played in the first sound zone corresponding to the first target display screen based on the sound field playback strategy.
[0085] Optionally, after determining the first multimedia recommendation data from the second multimedia data set and playing the first multimedia recommendation data on the vehicle's display screen, the processor 701 further executes: Receive interactive operations from each passenger regarding the first multimedia recommendation data; The second target occupant corresponding to the interaction operation is determined based on the multimodal sensing device, and the target preference label of the second target occupant is updated based on the interaction operation.
[0086] Optionally, after determining the first multimedia recommendation data from the second multimedia data set and playing the first multimedia recommendation data on the vehicle's display screen, the processor 701 further executes: During the playback of the first multimedia recommendation data, the second facial images and second audio information of each occupant are acquired through a multimodal sensing device; The second physiological state of each occupant is determined based on each second facial image and each second audio information; The first multimedia recommendation data is adjusted based on each second physiological state.
[0087] Optionally, the vehicle includes at least one display screen, and the processor 701 also performs: Obtain the second profile information of a third target occupant in the vehicle, where the third target occupant is any occupant in the vehicle; Based on the second profile information, the second data recommendation features corresponding to the third target occupant are determined, as well as the second target display screen corresponding to the seating position of the third target occupant; The second trip duration of the vehicle is determined based on the vehicle's navigation system. Based on the second trip duration, the third multimedia data set corresponding to the second data recommendation features is filtered to obtain the fourth multimedia data set. The second multimedia recommendation data is determined from the fourth multimedia data set, displayed on the second target display screen, and played in the second audio zone corresponding to the second target display screen.
[0088] In this embodiment, by acquiring facial images and audio information of occupants, the accuracy of distinguishing occupant personal information and physiological states is improved. By fusing and reasoning about the personal information and physiological states of multiple occupants, the raw sensor data is transformed into understandable occupant intentions to determine the target preference tags of each occupant. All target preference tags are then fused to determine data recommendation features shared by multiple occupants, improving the suitability of the recommended multimedia data to multiple occupants. Multimedia data that can be viewed in its entirety is filtered by trip duration, improving the time-dimensional suitability of the recommended multimedia data to the vehicle's driving scenario, ensuring occupants receive a complete viewing / listening experience. By combining data recommendation features shared by multiple occupants with the vehicle's trip duration, the overall satisfaction of multiple occupants in the vehicle is improved, guaranteeing a positive driving experience. When an occupant is a historical occupant, their current target preference tag is determined by combining their historical preference tags, enhancing the intelligence of multimedia data recommendation. The vehicle trip is divided into short and long trips based on a pre-set first duration threshold. Different multimedia recommendation data is determined for different trip types, improving the adaptability of the multimedia data recommendation method to various vehicle application scenarios and further enhancing the passenger experience. By dividing each display screen into sound zones and positioning each passenger in the optimal listening position according to the sound field playback strategy, the passenger listening experience is improved. While playing multimedia recommendation data, the interactive operations or physiological states of each passenger are acquired in real time, thereby adjusting user preference tags or multimedia recommendation data, further enhancing the intelligence of multimedia data recommendation. The multimedia recommendation data corresponding to each display screen is determined by the passenger profile information in front of each individual display screen and the trip duration of the navigation system, allowing for different multimedia recommendation data for different display screens, achieving personalized recommendations.
[0089] It should be understood that the apparatus provided in this application embodiment is used to execute the above-described multimedia data recommendation method, and therefore can achieve the same effect as the above-described implementation method.
[0090] When using an integrated unit, the device may include a processing module and a storage module. When the device is applied to a vehicle, the processing module can be used to control and manage the vehicle's movements. The storage module can be used to support the vehicle in executing relevant program code.
[0091] The processing module may be a processor or a controller, which can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.
[0092] In addition, the device provided in this application embodiment may specifically be a chip, component or module. The chip may include a connected processor and a memory. The memory is used to store instructions. When the processor calls and executes the instructions, the chip can execute the multimedia data recommendation method provided in the above embodiment.
[0093] This application also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the aforementioned method steps to implement the multimedia data recommendation method provided in the above embodiments.
[0094] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the multimedia data recommendation method provided in the above embodiment.
[0095] In this embodiment, the device, computer-readable storage medium, computer program product, or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0096] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0097] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0098] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A multimedia data recommendation method, characterized in that, Applied to vehicles, the method includes: Obtain first profile information of multiple occupants in the vehicle, and determine first data recommendation features that conform to the common preferences of each of the occupants based on each first profile information; The first trip duration of the vehicle is determined based on the vehicle's navigation system. Based on the first trip duration, the first multimedia data set corresponding to the first data recommendation feature is filtered to obtain the second multimedia data set. First multimedia recommendation data is determined from the second multimedia data set, and the first multimedia recommendation data is played on the vehicle's display screen.
2. The method according to claim 1, characterized in that, The vehicle is equipped with a multimodal perception device. The step of acquiring first profile information of multiple occupants in the vehicle and determining first data recommendation features that conform to the common preferences of each of the occupants based on each of the first profile information includes: The first facial images and first audio information of multiple occupants in the vehicle are acquired based on the multimodal sensing device. Based on each of the first facial images and each of the first audio information, determine the personal information and first physiological state of each of the occupants; Based on the personal information and the first physiological state, the target preference tags of each of the occupants are determined; Based on the target preference labels of each of the occupants, a first data recommendation feature that conforms to the common preferences of each of the occupants is determined.
3. The method according to claim 2, characterized in that, The determination of target preference tags for each passenger based on the personal information and the first physiological state includes: If the personal information indicates that the first target occupant is a historical occupant of the vehicle, then the historical preference tag of the first target occupant is obtained, and the first target occupant is any one of the plurality of occupants; The target preference label of the first target occupant is determined based on the historical preference label and the first physiological state.
4. The method according to claim 1, characterized in that, The process of filtering the first multimedia data set corresponding to the first data recommendation feature based on the first trip duration to obtain the second multimedia data set includes: Obtain the first multimedia data set corresponding to the first data recommendation feature, wherein the first multimedia data set is composed of each first multimedia data; In the first multimedia data set, the first multimedia data whose playback duration is less than or equal to the first trip duration is determined as the second multimedia data, and the second multimedia data set is determined based on each of the second multimedia data.
5. The method according to claim 1, characterized in that, The step of determining multimedia recommendation data from the second multimedia data set and playing the multimedia recommendation data on the vehicle's display screen includes: If the duration of the first trip is less than or equal to the first duration threshold, then the second multimedia data whose playback duration best matches the duration of the first trip in the second multimedia data set is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the display screen of the vehicle. If the duration of the first trip is greater than the first duration threshold, then in the second multimedia data set, the second multimedia data with a playback duration less than or equal to the second duration threshold is determined as the first multimedia recommended data, and the first multimedia recommended data is played on the display screen of the vehicle, wherein the second duration threshold is less than or equal to the first duration threshold.
6. The method according to claim 1 or 5, characterized in that, The vehicle includes at least one display screen, and playing the first multimedia recommendation data on the display screen of the vehicle includes: In all displays of the vehicle, a first target display screen corresponding to the first multimedia recommendation data and a sound field playback strategy of the first target display screen are determined based on each first profile information. The first multimedia recommendation data is displayed on the first target display screen, and the multimedia recommendation data is played in the first sound zone corresponding to the first target display screen based on the sound field playback strategy.
7. The method according to claim 2, characterized in that, After determining the first multimedia recommendation data in the second multimedia data set and playing the first multimedia recommendation data on the vehicle's display screen, the method further includes: Receive interactive operations from each of the passengers regarding the first multimedia recommendation data; Based on the multimodal sensing device, the second target occupant corresponding to the interaction operation is determined, and the target preference label of the second target occupant is updated based on the interaction operation.
8. The method according to claim 2, characterized in that, After determining the first multimedia recommendation data in the second multimedia data set and playing the first multimedia recommendation data on the vehicle's display screen, the method further includes: During the playback of the first multimedia recommendation data, the second facial image and second audio information of each passenger are acquired through the multimodal sensing device; The second physiological state of each occupant is determined based on each of the second facial images and each of the second audio information; The first multimedia recommendation data is adjusted based on each of the second physiological states.
9. The method according to claim 1, characterized in that, The vehicle includes multiple displays, and the method further includes: Obtain second profile information of a third target occupant in the vehicle, wherein the third target occupant is at least one occupant at a seating position corresponding to any display screen in the vehicle; Based on the second profile information, the second data recommendation features corresponding to the third target occupant and the second target display screen corresponding to the seating position of the third target occupant are determined. The second trip duration of the vehicle is determined based on the vehicle's navigation system. Based on the second trip duration, the third multimedia data set corresponding to the second data recommendation feature is filtered to obtain a fourth multimedia data set. Second multimedia recommendation data is determined from the fourth multimedia data set, the second multimedia recommendation data is displayed on the second target display screen, and the second multimedia recommendation data is played in the second audio zone corresponding to the second target display screen.
10. A vehicle, characterized in that, The vehicles include: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, causing the vehicle to perform the method as described in any one of claims 1 to 9.