A vehicle-mounted video playing system and method
Through the modular design of the in-vehicle video playback system, the emotions of people in the car and the vehicle status are monitored in real time, and the video playback mode is automatically adjusted. This solves the problem that traditional systems cannot provide personalized services, and improves driving safety and user satisfaction.
Patent Information
- Application Number
- CN202510424986.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Traditional in-car video playback systems cannot provide personalized services based on users' real-time status and needs. The ways for users to interact with the system are limited and cannot meet consumers' personalized needs.
Through the in-vehicle scene perception module, occupant emotion monitoring module, video preference recognition module, playback mode switching module and multi-screen interaction module, the emotional state of the people in the car is monitored in real time, providing personalized in-vehicle video playback services.
It automatically adjusts the video playback mode according to the driver's attention and vehicle status, reduces visual interference, improves driving safety, enhances user experience, and meets the entertainment needs of different users.
Smart Images

Figure CN120343310B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle-mounted video playback, and in particular to a vehicle-mounted video playback system and method. Background Art
[0002] In-car video playback refers to the process of playing video content using electronic devices inside the car (such as the central control screen, rear entertainment screen, head-up display, etc.) and related multimedia playback software and hardware. The video playback content can be locally stored video files, online streaming videos, or auxiliary videos related to vehicle driving. With the development of the automotive industry and the improvement of people's consumption needs, the application of in-car video playback systems in modern cars is becoming more and more extensive. Therefore, automobile manufacturers and suppliers need to strengthen the optimization and upgrading of in-car video playback systems to meet consumer needs.
[0003] Traditional in-car video playback can usually only play content according to a fixed mode preset by the user. The user's interaction with the system is relatively limited. Usually, the user only selects the playback content and adjusts the volume through buttons or touch screens. It is unable to provide personalized services based on the user's real-time status and needs. Summary of the Invention
[0004] The present invention provides an in-vehicle video playback system and method, the main purpose of which is to monitor the emotional state of people in the vehicle in real time and provide personalized in-vehicle video playback services.
[0005] To achieve the above-mentioned object, the present invention provides an in-vehicle video playback system, comprising: an in-vehicle scene perception module, an occupant emotion monitoring module, a video preference recognition module, a playback mode switching module, a multi-screen interaction module, and a video playback execution module;
[0006] The vehicle scene perception module is used to obtain the occupants and vehicle scene of the target vehicle, wherein the vehicle scene includes the vehicle status, the interior scene and the exterior environment, and the scene perception unit of the target vehicle is set according to the vehicle scene;
[0007] The occupant emotion monitoring module is used to collect voice data and facial images of the occupants of the vehicle, and to create a real-time emotion monitoring network for the occupants by combining the voice data, facial images, and the vehicle scene;
[0008] The video preference identification module is used to read the historical video playlist of the target vehicle, extract the video playback content in the historical video playlist, and identify the user video playback preference of the target vehicle based on the historical video playlist and the real-time emotion monitoring network;
[0009] The playback mode switching module is configured to determine a current driving mode of the target vehicle based on the scene perception unit, monitor the driver's attention of the target vehicle based on the real-time emotion monitoring network, and switch the video playback mode of the video playback content to an audiobook mode based on the current driving mode and the driver's attention;
[0010] The multi-screen interaction module is configured to set automatic switching conditions between the video playback mode and the audiobook mode based on the real-time emotion monitoring network and the current driving mode; set a multimedia content control path for the vehicle occupants based on the automatic switching conditions, the current driving mode, and the real-time emotion monitoring network; and create a multi-screen entertainment interaction interface for the vehicle occupants based on the multimedia content control path and the user's video playback preferences;
[0011] The video playback execution module is used to combine the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface to execute the video playback processing of the target vehicle and obtain the in-vehicle video playback result.
[0012] Optionally, setting the scene perception unit of the target vehicle according to the vehicle-mounted scenario includes:
[0013] Collecting vehicle driving data, external environment data, and occupant behavior data of the target vehicle according to the vehicle-mounted scenario;
[0014] identifying a driving mode of the target vehicle based on the vehicle driving data, the external environment data, and the occupant behavior data;
[0015] Extracting facial features and hand movement features of the occupants of the target vehicle based on the occupant behavior data;
[0016] Identifying sound data of the target vehicle from the vehicle driving data, and extracting engine sound frequency and occupant voice features of the target vehicle based on the sound data;
[0017] Building a mapping relationship between the in-vehicle scenario and the driving mode by combining the facial features of the occupant, the hand movement features of the occupant, the engine sound frequency, and the voice features of the occupant;
[0018] Based on the mapping relationship, a scene perception unit of the target vehicle is set.
[0019] A method for playing an in-vehicle video, characterized in that the method comprises:
[0020] Acquire the occupants and vehicle scene of the target vehicle, wherein the vehicle scene includes the vehicle status, the vehicle scene, and the vehicle environment, and set the scene perception unit of the target vehicle according to the vehicle scene;
[0021] Collecting voice data and facial images of the occupants of the vehicle, and creating a real-time emotion monitoring network for the occupants of the vehicle by combining the voice data, the facial images, and the vehicle scene;
[0022] Reading a historical video playlist of the target vehicle, extracting video play content in the historical video playlist, and identifying a user video play preference of the target vehicle based on the historical video playlist and the real-time emotion monitoring network;
[0023] determining a current driving mode of the target vehicle based on the scene perception unit, monitoring the driver's attention of the target vehicle based on the real-time emotion monitoring network, and switching a video playback mode of the video content to an audiobook mode based on the current driving mode and the driver's attention;
[0024] setting an automatic switching condition between the video playback mode and the audiobook mode based on the real-time emotion monitoring network and the current driving mode; setting a multimedia content control path for the vehicle occupant based on the automatic switching condition, the current driving mode, and the real-time emotion monitoring network; and creating a multi-screen entertainment interaction interface for the vehicle occupant based on the multimedia content control path and the user's video playback preference;
[0025] In combination with the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface, the video playback processing of the target vehicle is executed to obtain the vehicle-mounted video playback result.
[0026] The embodiment of the present invention obtains the occupants and the in-vehicle scene of the target vehicle, and sets the scene perception unit of the target vehicle according to the in-vehicle scene, so as to judge the driver's attention state and the driving mode of the vehicle, avoid the video content from interfering with the driver, and ensure driving safety; further, the embodiment of the present invention creates a real-time emotion monitoring network of the occupants by collecting the voice data and facial images of the occupants, so as to timely detect the driver's fatigue state and concentration level, and monitor the emotional changes of the occupants, thereby helping to optimize the human-computer interaction function of the in-vehicle video playback system; the embodiment of the present invention uses the historical video playback list to play the video. and the real-time emotion monitoring network to identify the user video playback preferences of the target vehicle, and can give priority to updating the video content that the user often watches, thereby increasing the user's dependence on and satisfaction with the system; further, the embodiment of the present invention can reduce visual interference by converting the video playback mode of the video playback content into the audio book mode based on the current driving mode and the driver's attention, so as to enable the driver to focus on driving and help the driver make better use of driving time, and at the same time enrich the entertainment content in the car to meet the interests and hobbies of different users; the embodiment of the present invention sets the video playback mode and the audio book mode according to the real-time emotion monitoring network and the current driving mode. The automatic switching conditions of the vehicle can set differentiated in-vehicle multimedia content control paths to achieve personalized video playback experience for the occupants; further, the embodiment of the present invention sets the multimedia content control path for the occupants based on the automatic switching conditions, the current driving mode and the real-time emotion monitoring network, which can significantly distinguish the multimedia content control logic of the driver side and the rear screen passenger side to ensure that the operations of the two do not interfere with each other; the embodiment of the present invention creates a multi-screen entertainment interaction interface for the occupants based on the multimedia content control path and the user's video playback preferences, which can meet the entertainment or information acquisition needs without affecting driving safety, and adapt to different Driving scene, and at the same time, the playback content can be adjusted according to the needs of passengers in different seats to improve user satisfaction; finally, the embodiment of the present invention combines the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface to execute the video playback processing of the target vehicle, obtain the in-vehicle video playback result, and can obtain the user's status and emotional information in real time, and provide users with personalized video playback content and services. The combination of the multimedia content control path and the multi-screen entertainment interaction interface can avoid the video playback content from interfering with the driver, improve driving safety, and provide users with a smoother and more natural intelligent interactive experience. Therefore, the embodiment of the present invention provides an in-vehicle video playback system and method that can monitor the emotional state of people in the car in real time and provide personalized in-vehicle video playback services. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A functional module diagram of a vehicle-mounted video playback system provided by one embodiment of the present invention;
[0028] Figure 2 A schematic flow chart of a method for playing in-vehicle video according to an embodiment of the present invention;
[0029] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0031] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.
[0032] In fact, the server-side device deployed by the in-vehicle video playback system may be composed of one or more devices. The above-mentioned in-vehicle video playback system can be implemented as: a business instance, a virtual machine, and a hardware device. For example, the in-vehicle video playback system can be implemented as a business instance deployed on one or more devices in a cloud node. In simple terms, the in-vehicle video playback system can be understood as a software deployed on a cloud node, which is used to provide in-vehicle video playback services to each user terminal. Alternatively, the in-vehicle video playback system can also be implemented as a virtual machine deployed on one or more devices in a cloud node. The virtual machine is installed with application software for managing each user terminal. Alternatively, the in-vehicle video playback system can also be implemented as a server composed of many hardware devices of the same or different types, and one or more hardware devices are set to provide in-vehicle video playback services to each user terminal.
[0033] In terms of implementation, the in-car video playback system and the user end are mutually compatible. Specifically, if the in-car video playback system is an application installed on a cloud service platform, the user end is the client that establishes a communication connection with the application. Alternatively, if the in-car video playback system is implemented as a website, the user end is implemented as a webpage. Alternatively, if the in-car video playback system is implemented as a cloud service platform, the user end is implemented as a mini-program within an instant messaging application.
[0034] Reference Figure 1 FIG. 1 is a functional module diagram of a vehicle-mounted video playback system provided by an embodiment of the present invention.
[0035] The in-vehicle video playback system 100 described in the present invention can be installed in a cloud server. In terms of implementation, it can be implemented as one or more service devices, as an application installed in the cloud (e.g., a server or server cluster for in-vehicle video playback), or as a website. Depending on the functionality implemented, the in-vehicle video playback system 100 includes an in-vehicle scene perception module 101, an occupant emotion monitoring module 102, a video preference recognition module 103, a playback mode switching module 104, a multi-screen interaction module 105, and a video playback execution module 106.
[0036] In an embodiment of the present invention, in the tracking based on in-vehicle video playback, each of the above modules can be implemented independently and called with other modules. The call here can be understood as a module that can connect to multiple modules of another type and provide corresponding services to the multiple modules connected to it. In an in-vehicle video playback system provided by an embodiment of the present invention, the scope of application of the in-vehicle video playback architecture can be adjusted by adding modules and directly calling them without modifying the program code, thereby realizing cluster-type horizontal expansion, so as to achieve the purpose of quickly and flexibly expanding the in-vehicle video playback system. In actual applications, the above modules can be set in the same device or different devices, or they can be set in a virtual device, such as a service instance in a cloud server.
[0037] The following describes the various components and specific workflows of the vehicle-mounted video playback system in conjunction with specific embodiments.
[0038] The vehicle-mounted scene perception module 101 is used to obtain the occupants and vehicle-mounted scene of the target vehicle, wherein the vehicle-mounted scene includes the vehicle status, the interior scene and the exterior environment, and the scene perception unit of the target vehicle is set according to the vehicle-mounted scene.
[0039] The embodiment of the present invention can provide data support for subsequent in-vehicle video playback settings by obtaining the occupants and in-vehicle scenes of the target vehicle. The target vehicle refers to a vehicle used to transport people, goods or perform specific operations. The occupants refer to people located inside the target vehicle, including the driver and passengers. The in-vehicle scenes refer to various situations and status information inside and outside the target vehicle, such as vehicle speed, driving route, driving mode, etc.
[0040] Furthermore, the embodiment of the present invention can judge the driver's attention state and the vehicle's driving mode by setting the scene perception unit of the target vehicle according to the in-vehicle scenario, avoid interference of the video content on the driver, and ensure driving safety. The scene perception unit refers to an intelligent system that can comprehensively and accurately perceive and analyze the vehicle's environment and internal conditions.
[0041] As an embodiment of the present invention, the setting of the scene perception unit of the target vehicle according to the in-vehicle scenario includes: collecting vehicle driving data, external environment data and occupant behavior data of the target vehicle according to the in-vehicle scenario; identifying the driving mode of the target vehicle based on the vehicle driving data, the external environment data and the occupant behavior data; extracting the facial features and hand movement features of the occupants of the target vehicle according to the occupant behavior data; identifying the sound data of the target vehicle from the vehicle driving data, and extracting the engine sound frequency and occupant voice features of the target vehicle based on the sound data; constructing a mapping relationship between the in-vehicle scenario and the driving mode in combination with the facial features, hand movement features, engine sound frequency and voice features; and setting the scene perception unit of the target vehicle based on the mapping relationship.
[0042] Among them, the vehicle driving data refers to various data related to the driving status of the target vehicle itself, including but not limited to vehicle speed, acceleration, braking frequency, steering angle, gear information, engine speed, etc. The external environment data refers to the external environment information around the vehicle, such as weather conditions and road conditions. The occupant behavior data refers to the behavior and status information of the occupants (including the driver and passengers), such as the driver's attention state. The driving mode refers to the driving method adopted by the driver under different road conditions and driving requirements. The occupant facial features refer to the feature information of the occupant's face captured by the camera and other equipment in the car, such as the degree of eye opening, blinking frequency, head tilt angle, etc. The occupant hand movement features refer to the hand movement features of the occupant during driving or riding, such as the way of holding the steering wheel, gear shifting, passenger's grasping action, etc. Sound data refers to various sound information generated during vehicle driving, including engine sound, friction between tires and road surface, wind noise, etc., as well as voice communication and operating instructions of passengers in the vehicle. The engine sound frequency refers to the distribution characteristics of the sound signal generated when the engine is running in the frequency domain. For example, the engine fundamental frequency may be in the range of tens to hundreds of hertz when idling, while when accelerating at high speed, the fundamental frequency may rise to several thousand hertz. The passenger voice characteristics refer to the unique attributes and characteristics of the voice signals of vehicle occupants (including drivers and passengers) when speaking in the vehicle. For example, the speaking speed may increase when nervous or excited, while the speaking speed may slow down when relaxed or thinking. The mapping relationship refers to the correspondence established between the vehicle scenario and the driving mode, that is, in a specific vehicle scenario, which driving mode should the vehicle adopt to achieve the best driving effect and safety.
[0043] Optionally, according to the in-vehicle scenario, the vehicle driving data of the target vehicle can be collected through CAN bus data reading, and the occupant behavior data of the target vehicle can be collected using an in-vehicle camera. Based on the vehicle driving data, the external environment data and the occupant behavior data, the driving mode of the target vehicle can be identified by a machine learning algorithm, such as using a support vector machine algorithm to classify the collected data to identify different driving modes. Based on the sound data, the engine sound frequency of the target vehicle can be extracted using fast Fourier transform. In combination with the occupant facial features, the occupant hand movement features, the engine sound frequency and the occupant voice features, the mapping relationship between the in-vehicle scenario and the driving mode can be realized using a cluster analysis algorithm. For example, the K-means clustering algorithm can divide the in-vehicle scenario data and driving mode data into different clusters based on multi-dimensional features such as occupant facial features, occupant hand movement features, engine sound frequency and occupant voice features, each cluster corresponding to an in-vehicle scenario category or driving mode type.
[0044] The occupant emotion monitoring module 102 is used to collect voice data and facial images of the occupants of the vehicle, and to create a real-time emotion monitoring network for the occupants of the vehicle by combining the voice data, the facial images and the vehicle scene.
[0045] By collecting voice data and facial images of the people in the car, the embodiment of the present invention can identify the voice characteristics, language habits, facial expressions and identities of different people, thereby providing personalized video playback experience for different users. The voice data refers to the sound information of the people in the car speaking collected by a microphone or other audio equipment, and the facial image refers to the image information of the human face taken by a camera or other image acquisition device.
[0046] Furthermore, the embodiment of the present invention creates a real-time emotion monitoring network for the people in the car by combining the voice data, the facial image and the in-vehicle scenario, which can timely detect the driver's fatigue state and concentration level, and at the same time monitor the emotional changes of the people in the car, thereby helping to optimize the human-computer interaction function of the in-vehicle video playback system. The real-time emotion monitoring network refers to a system architecture that performs real-time perception, analysis and evaluation of the emotions of people in the car based on voice data and facial images collected by in-vehicle equipment.
[0047] As an embodiment of the present invention, the combination of the voice data, the facial image and the in-vehicle scenario to create a real-time emotion monitoring network for the people in the car includes: determining different seating areas of the people in the car based on the voice data; identifying the emotional states of the people in the different seating areas based on the voice data and the facial image; extracting facial feature points of the people in the car from the facial image, and generating an emotion heat map of the different seating areas by combining the emotional states of the people and the facial feature points; identifying the vehicle driving state in the in-vehicle scenario, and analyzing the scene sensitivity of the people in the car based on the vehicle driving state and the emotional states of the people; constructing cabin emotion zones of the people in the car based on the scene sensitivity and the emotion heat map; and creating a real-time emotion monitoring network for the people in the car based on the cabin emotion zones.
[0048] Among them, the different seating areas refer to different spatial areas in the vehicle divided according to the seating position, such as the driver's seat, the front passenger seat, the left rear seat, the middle rear seat, the right rear seat, etc. The emotional state of the person refers to the emotional state of the person in the vehicle at a certain moment, such as happiness or anger. The facial feature points refer to the key parts used for emotion recognition in facial images, such as the inner and outer corners of the eyes, and the end points and inflection points of the eyebrows. The vehicle driving state refers to the collection of various dynamic and static conditions in which the vehicle is in operation, including speed, position, road conditions, environmental conditions, vehicle system status, driving behavior, and safety status. The scene sensitivity refers to the sensitivity of the emotional state of the person in the vehicle to changes in the vehicle scene, such as the vehicle driving state. For example, the emotions of the person in the vehicle fluctuate greatly when the vehicle accelerates or brakes suddenly, but are relatively stable when the vehicle is driving smoothly. The cabin emotional partition refers to the seat partition obtained by dividing the cabin space into emotional areas based on scene sensitivity and emotional heat map.
[0049] Optionally, based on the voice data, the different seating areas of the occupants in the vehicle can be determined using an acoustic positioning method; based on the voice data and the facial image, the emotional states of the occupants in different seating areas can be identified by a deep learning model, such as a CNN model; based on the combination of the emotional states of the occupants and the facial feature points, the emotional heat maps of the different seating areas can be generated using Python's concat function; based on the scene sensitivity and the emotional heat map, the cabin emotional zoning of the occupants in the vehicle can be constructed using GIS spatial analysis technology, such as hotspot analysis technology.
[0050] In an optional embodiment of the present invention, based on the vehicle driving state and the emotional state of the occupant, the scene sensitivity of the occupant is analyzed using the following formula:
[0051]
[0052] Among them, A represents the scene sensitivity of the people in the car, w j The weight representing the sensitivity of the j-th vehicle’s driving state to the scene, such as weight w = 1, w j It can be 0.5, ΔE j Indicates the degree of change in the emotional state of the person in the j-th vehicle driving state, such as the degree of change ΔE includes 10, 20, 30, ΔE j It can be 20, where n represents the number of categories of vehicle driving status, j represents the category index of vehicle driving status, and m j represents the number of key characteristic parameters contained in the driving state of the j-th vehicle, q ij Indicates the influence weight of the i-th key characteristic parameter on the emotional state of the person in the j-th vehicle driving state, such as influence weight q = 1, q ij It can be 0.6, y ij Indicates the quantitative value of the impact of the i-th key characteristic parameter on the emotional state of the person in the j-th vehicle driving state, such as the impact quantitative value y = 10, y ij It can be 8, where i represents the serial number of the key feature parameter.
[0053] It should be noted that in this application, the above formula can comprehensively consider the vehicle driving state category, the emotional changes of people in each state, and the impact of key characteristic parameters on emotions, and accurately quantify the sensitivity of people in the car to different scenes. In particular, it should be noted that the formula It is used to perform weighted summation of each key feature parameter according to its importance in affecting the emotional state of the person, and obtain the quantitative impact value of the driving state of the j-th type of vehicle.
[0054] The video preference identification module 103 is used to read the historical video play list of the target vehicle, extract the video playback content in the historical video play list, and identify the user video playback preference of the target vehicle based on the historical video play list and the real-time emotion monitoring network.
[0055] The embodiment of the present invention can enhance personalized user experience and optimize content management of the video playback system by reading the historical video play list of the target vehicle and extracting the video playback content in the historical video play list. The historical video play list refers to a record list of video files or video resources that the user has played in the vehicle-mounted video playback system, and the video playback content refers to the specific content in the video file or video stream actually watched by the user. For example, the playback content of the movie "Avatar" played by the user in the vehicle-mounted system includes the movie's pictures, dialogues, background music, subtitles, etc.
[0056] Optionally, the historical video playlist of the target vehicle can be read through the OBD-II port.
[0057] Furthermore, the embodiment of the present invention can prioritize updating the video content that the user frequently watches by identifying the user's video playback preferences of the target vehicle based on the historical video playlist and the real-time emotion monitoring network, thereby increasing the user's dependence on and satisfaction with the system. The user's video playback preferences refer to the tendencies, preferences, and habits shown by the user when watching videos, such as concluding that the user likes to watch action movies based on the historical video playlist.
[0058] As an embodiment of the present invention, the identifying the user video playback preferences of the target vehicle based on the historical video playlist and the real-time emotion monitoring network includes: determining the occupant seating layout of the target vehicle based on the real-time emotion monitoring network; identifying the occupant combination in the target vehicle based on the occupant seating layout and the real-time emotion monitoring network; extracting the video ID and its corresponding timestamp data in the historical video playlist based on the occupant combination in the vehicle; identifying the high-frequency playback content of the occupant combination in the vehicle based on the video ID and the timestamp data; and identifying the user video playback preferences of the target vehicle based on the high-frequency playback content.
[0059] Among them, the passenger seat layout refers to the seat distribution and arrangement of passengers inside the vehicle, and the passenger combination in the vehicle refers to the composition of the passengers in the vehicle, including information such as the number of people, age, gender, and mutual relationships. For example, family travel can be understood as a combination of parents and children, and business travel can be understood as a combination of colleagues or customers. The video ID refers to the unique identification code for each video in the historical video playlist, and the timestamp data refers to the information recording the video playback time point, which marks the specific time when each video starts and ends. The high-frequency playback content refers to the video content that is repeatedly played by the passenger combination in the vehicle in the historical video playback record.
[0060] Optionally, according to the real-time emotion monitoring network, the occupant seating layout of the target vehicle can be determined by detecting the facial contours of the occupants. Based on the occupant seating layout and the real-time emotion monitoring network, the occupant combination in the target vehicle can be identified using voice recognition technology. For example, if a child is heard calling an adult "Dad" or "Mom", it can be determined that they are in a parent-child relationship; if members are heard discussing work-related content, it can be determined that they are colleagues. Based on the video ID and the timestamp data, the high-frequency playback content of the occupant combination in the vehicle can be identified by a sequence pattern mining algorithm, such as the PrefixSpan algorithm.
[0061] The playback mode switching module 104 is used to determine the current driving mode of the target vehicle based on the scene perception unit, monitor the driver's attention of the target vehicle according to the real-time emotion monitoring network, and convert the video playback mode of the video playback content into the audio book mode based on the current driving mode and the driver's attention.
[0062] The embodiment of the present invention determines the current driving mode of the target vehicle based on the scene perception unit, and can recommend video content that is more suitable for the current scene to passengers, thereby improving driving safety. The current driving mode refers to the driving state or operating mode of the vehicle at a certain moment.
[0063] As an embodiment of the present invention, determining the current driving mode of the target vehicle based on the scene perception unit includes: identifying the driving situation of the target vehicle based on the scene perception unit; analyzing the vehicle motion state and driving road type of the target vehicle according to the driving situation; collecting engine sound data and occupant sound data of the target vehicle based on the vehicle motion state; identifying the engine operating condition of the target vehicle based on the engine sound data; analyzing the voice characteristics of the occupants of the target vehicle based on the occupant sound data; extracting the driving behavior characteristics of the target vehicle in combination with the vehicle motion state, the driving road type and the engine operating condition; and determining the current driving mode of the target vehicle based on the driving behavior characteristics and the voice characteristics of the occupants.
[0064] The driving scenario refers to the specific scenario in which the vehicle is traveling, including the traffic environment, road conditions, weather conditions, traffic flow, etc. around the vehicle. The vehicle motion state refers to the dynamic situation of the vehicle during driving, such as speed, acceleration, steering angle, yaw angular velocity, braking status, etc. The road type refers to the type of road the vehicle is traveling on, such as expressways, urban roads, rural roads, mountain roads, etc. The engine sound data refers to various sound signals generated by the vehicle engine during operation, including engine speed changes, load changes, fault sounds, etc. The occupant sound data refers to audio signals such as voice conversations and emotional expressions of passengers or the driver in the vehicle. The engine operating condition refers to the real-time operating status of the engine, such as idling, acceleration, deceleration, high load, etc. The occupant voice characteristics refer to the voice characteristics of the occupant speaking in the vehicle, such as speaking rate, tone, volume, and voice content. The driving behavior characteristics refer to the driver's behavioral performance during driving, such as the depth of the accelerator pedal, the force of the brakes, the frequency of steering, etc.
[0065] Optionally, based on the driving scenario, the vehicle motion state of the target vehicle can be identified by vehicle sensors, based on the engine sound data, the engine operating condition of the target vehicle can be identified using engine operating parameters obtained by an engine control unit, based on the occupant sound data, the voice characteristics of the occupants of the target vehicle can be analyzed by NLP sentiment analysis, and in combination with the vehicle motion state, the driving road type and the engine operating condition, the driving behavior characteristics of the target vehicle can be extracted using a hidden Markov model.
[0066] Furthermore, the embodiment of the present invention can prevent the in-vehicle video playback system from interfering with the driver's concentration and prevent the driver from being in danger due to fatigue and video interference by monitoring the driver's attention of the target vehicle based on the real-time emotion monitoring network. The driver's attention refers to the driver's ability and state to concentrate his mental activities on driving-related tasks and the surrounding traffic environment during driving.
[0067] Optionally, based on the real-time emotion monitoring network, the target vehicle's driver attention monitoring can be achieved using eye tracking technology, such as using a camera device under the real-time emotion monitoring network to monitor the driver's line of sight and blinking frequency to determine the driver's concentration level.
[0068] The embodiment of the present invention converts the video playback mode of the video playback content into the audio book mode based on the current driving mode and the driver's attention, thereby reducing visual interference, allowing the driver to focus on driving, and helping the driver to make better use of driving time. At the same time, it can enrich the entertainment content in the car and meet the interests and hobbies of different users. The video playback mode refers to the playback mode of the in-vehicle entertainment system that mainly plays visual content, and the audio book mode refers to the playback mode of the in-vehicle entertainment system that mainly plays audio content.
[0069] As an embodiment of the present invention, the conversion of the video playback mode of the video playback content to the listening mode based on the current driving mode and the driver's attention includes: identifying the material type of the video playback content, and extracting the voice track of the video playback content according to the material type; performing format conversion processing of the video playback content based on the voice track to obtain an audio playback format; generating audio playback content corresponding to the video playback content according to the audio playback format; creating an audio control interface for the video playback content based on the audio playback content; setting audio output parameters of the audio playback content according to the current driving mode; setting mode conversion rules for the video playback mode according to the current driving mode and the driver's attention; and converting the video playback mode of the video playback content to the listening mode in combination with the mode conversion rules, the voice track, the audio output parameters and the audio control interface.
[0070] Among them, the material type refers to the content category of the video playback content, for example, the video material is a movie, TV series, short video, etc., the voice track refers to an independent audio channel in the video playback content for storing audio data, the format conversion processing refers to the process of extracting the voice track in the video file and converting it into a format suitable for audio playback, the audio playback content refers to a type of media content in the form of audio, the audio control interface refers to the interface for user interaction with the audio playback system, the audio output parameters refer to the setting parameters during audio playback, such as volume and sound effect mode, and the mode conversion rule refers to the logical rule by which the system automatically decides when to switch the video playback mode to the audiobook mode based on the current driving mode and the driver's attention status. For example, if the system detects that the driver's attention is distracted or the driving environment is complex, it may automatically switch the video playback mode to the audiobook mode.
[0071] Optionally, according to the material type, the voice track of the video playback content can be extracted by a video parsing algorithm, and based on the voice track, the format conversion processing of the video playback content is implemented using a text-to-speech engine, and based on the audio playback content, the audio control interface of the video playback content can be created by a graphics framework provided by the vehicle operating system, such as using the WindCreate function of PhotonMicroGUI to create an audio control interface window, and then hiding the video playback window and displaying the audio control window through the WindSetVisible function, and according to the current driving mode and the driver's attention, the mode conversion rule setting of the video playback mode can be set using a rule engine.
[0072] The multi-screen interaction module 105 is used to set the automatic switching conditions between the video playback mode and the audiobook mode according to the real-time emotion monitoring network and the current driving mode, set the multimedia content control path of the occupants based on the automatic switching conditions, the current driving mode and the real-time emotion monitoring network, and create a multi-screen entertainment interaction interface for the occupants based on the multimedia content control path and the user's video playback preference.
[0073] The embodiment of the present invention sets the automatic switching conditions between the video playback mode and the audiobook mode according to the real-time emotion monitoring network and the current driving mode, thereby setting differentiated in-vehicle multimedia content control paths and realizing a personalized video playback experience for the occupants of the vehicle. The automatic switching conditions refer to conditions that automatically trigger the conversion of the video playback mode to the audiobook mode. For example, if the vehicle is in a complex driving mode, such as high-speed driving or severe congestion, and the driver frequently operates the video or is not paying attention, the video playback mode will be triggered to switch to the audiobook mode.
[0074] As an embodiment of the present invention, the automatic switching conditions between the video playback mode and the audiobook mode are set according to the real-time emotion monitoring network and the current driving mode, including: identifying the driver and driving scene characteristics in the current driving mode; determining the scene complexity of the current driving mode based on the driving scene characteristics; detecting the driver's visual attention based on the real-time emotion monitoring network; analyzing the attention change trend of the visual attention according to the scene complexity, and identifying the driver's cognitive load; calculating the driver's continuous concentration time based on the attention change trend; defining the driver's attention level according to the continuous concentration time; setting the mode switching threshold corresponding to the attention level and the cognitive load based on the scene complexity; and setting the automatic switching conditions between the video playback mode and the audiobook mode according to the mode switching threshold.
[0075] Among them, the driver refers to the driver currently driving the vehicle, the driving scene characteristics refer to specific characteristics related to the driving task in the current driving environment, including road type, traffic flow, weather conditions, etc., the scene complexity refers to the difficulty of the driving task in the current driving scene. For example, the scene complexity of urban congested roads is usually higher, while the scene complexity of highway cruising is lower, the visual attention refers to the driver's attention to visual information during driving, including gaze direction, gaze time, etc., the attention change trend refers to the change of the driver's visual attention over time, the cognitive load refers to the amount of information and task difficulty that the driver needs to process during driving, the sustained attention time refers to the length of time the driver can maintain high visual attention within a period of time, the attention level refers to the classification of the driver's attention state according to the sustained attention time, and the mode switching threshold refers to the condition for switching the multimedia playback mode set according to the scene complexity, attention level and cognitive load.
[0076] Optionally, based on the driving scene characteristics, the scene complexity of the current driving mode can be determined by a multi-factor comprehensive evaluation method; based on the scene complexity, the attention change trend of the visual attention can be analyzed using a long short-term memory network algorithm; the driver's cognitive load can be identified by physiological signal monitoring technology; based on the scene complexity, the mode switching threshold corresponding to the attention level and the cognitive load can be set using a threshold judgment algorithm; for example, at a high attention level, the proportion of the time the gaze point stays in the road-related area should exceed 80%; at a low attention level, the proportion should be less than 60%; the standard deviation of the heart rate variability is greater than 60 milliseconds at low cognitive load, and should be less than 40 milliseconds at high cognitive load.
[0077] In an optional embodiment of the present invention, based on the attention change trend, the driver's continuous concentration time is calculated using the following formula:
[0078]
[0079] Among them, Time d represents the driver's continuous attention time, u represents the decay rate of the attention level under the attention change trend, θ represents the driver's attention threshold, f represents the lower limit of the attention level under the attention change trend, such as f>0, f can be 0.1, a represents the range of attention level under the attention change trend, Time start Indicates the time when the driving task starts.
[0080] It should be noted that u, a, and f in the formula can be calculated by a function of the attention level changing over time, such as A(t) = a×eut +f, where the a value is determined by the driver's initial state of attention at the beginning of the driving task and the subsequent range of attention changes, and the u value can be determined using a data fitting algorithm (such as the least squares method, etc.), e ut The value range of is between (0, 1]. When t=0, A(t) approaches f. θ can be understood as the preset driver's attention threshold, which ranges from 0 to 1. When the driver's attention level drops to θ, it is considered that the continuous concentration time has ended. For example, if θ is set to 0.7, when the attention level drops to 0.7 or below, it is determined that the driver is no longer in a continuous concentration state.
[0081] Furthermore, the embodiment of the present invention sets the multimedia content control path for the vehicle occupants based on the automatic switching conditions, the current driving mode, and the real-time emotion monitoring network, thereby significantly differentiating the multimedia content control logic between the driver side and the rear screen passenger side, ensuring that the operations of the two do not interfere with each other. The operation behavior control path refers to a path for classifying, operating, and managing various multimedia contents (video, audio, etc.) for different vehicle occupants (such as drivers and passengers) in the vehicle video playback system. For example, when in a complex driving mode and the driver's attention is distracted, the driver is only allowed to perform basic control over the audio in the listening mode, while the passengers in the rear screen seat can freely switch between various modes such as video playback, listening to books, and music playback according to their personal preferences.
[0082] As an embodiment of the present invention, the multimedia content control path for the occupants is set according to the automatic switching condition, the current driving mode and the real-time emotion monitoring network, including: identifying the driving scenario corresponding to the current driving mode and determining the driver's seat occupant and the fellow passengers among the occupants; analyzing the task complexity of the current driving mode based on the driving scenario; identifying the multimedia playback format under the current driving mode according to the automatic switching condition; extracting the emotion change trend of the fellow passengers under the multimedia playback format based on the real-time emotion monitoring network; determining the degree of enthusiasm of the fellow passengers for the multimedia playback format based on the emotion change trend; creating a free switching mechanism for the fellow passengers for the multimedia playback format based on the enthusiasm; setting the behavior restriction authority and restriction lifting rules for the driver's seat occupant for the multimedia playback format according to the driving scenario and the task complexity; and setting the multimedia content control path for the occupants in the vehicle in combination with the multimedia playback format, the free switching mechanism, the behavior restriction authority and the restriction lifting rules.
[0083] Among them, the driving scenario refers to the specific environment or working conditions in which the vehicle is located, and the task complexity refers to the difficulty and complexity of the driving task that the driver needs to complete in a specific driving scenario. For example, when driving on congested roads in the city, it is necessary to frequently start and stop the vehicle, pay attention to the dynamics of pedestrians and non-motor vehicles, and respond to complex traffic signals, etc., which can be understood as a high task complexity. The main driver's seat occupant refers to the person sitting in the driver's seat of the vehicle, that is, the person who is responsible for controlling the vehicle's driving and is primarily responsible for the vehicle's driving operation and safe driving. The co-passengers refer to the people sitting in other seats of the vehicle except the main driver's seat occupant, including the co-pilot seat and the passengers in the back seat. The multimedia playback format refers to the way of presenting multimedia content in the car, such as video type, audio type, and the emotional change trend refers to the passenger's emotional state obtained at any time through the emotional monitoring network. For example, a passenger's mood gradually becomes happy when watching a certain type of video, but becomes irritable when listening to certain audio. The enthusiasm refers to the degree of the passenger's liking or preference for a certain form of multimedia playback. The free switching mechanism refers to a system that allows passengers to freely switch the form of multimedia playback according to their preferences and emotional state. For example, a passenger can switch between different playback forms through voice commands or touch screen operations. The behavior restriction authority refers to the restriction on the driver's seat occupant's operation of the multimedia system in a specific driving scenario. For example, in a driving scenario with high task complexity, the driver is restricted from watching videos or performing complex operations. The restriction lifting rule refers to the rule for lifting the behavior restriction on the driver's seat occupant when certain conditions are met. For example, when the vehicle enters a driving scenario with low task complexity (such as highway cruising), some restrictions are lifted.
[0084] Optionally, the driver's seat occupant and fellow passengers in the vehicle can be determined by seat sensors, and based on the real-time emotion monitoring network, the emotion change trend of the fellow passengers under the multimedia playback format can be extracted using a time series analysis model, and based on the emotion change trend, the degree of enthusiasm of the fellow passengers for the multimedia playback format can be determined through an LSTM-Attention network, and based on the enthusiasm, the free switching mechanism of the fellow passengers for the multimedia playback format can be created using a direct connection channel between the rear independent control terminal and the domain controller, such as the rear cabin entertainment screen of Ideal L9 and Qualcomm SA8155P, and based on the driving scenario and the complexity of the task, the behavior restriction authority and restriction lifting rules of the driver's seat occupant for the multimedia playback format can be set through conditional trigger logic.
[0085] In an optional embodiment of the present invention, the task complexity of the current driving mode is analyzed using the following formula according to the driving scenario:
[0086]
[0087] Among them, C represents the task complexity of the current driving mode, b e represents the e-th driving action in the driving scene, p(b e ) represents the probability of taking the e-th driving action in the driving scenario, D obj It represents the number of dynamic obstacles detected per unit time in the driving scenario, g represents the number of driving actions in the driving scenario, e represents the action index of the driving action in the driving scenario, α and β represent weight coefficients, such as α and β can be 0.5, and T represents the time window.
[0088] It should be noted that, in this application, the above formula can be used to more comprehensively understand and evaluate the complexity of the driving task. In particular, it should be noted that the formula The entropy of the probability distribution of driving actions is calculated to reflect the uncertainty of driving behavior. The higher the entropy, the greater the uncertainty of driving actions and the higher the task complexity. At the same time, formula D obj / T is the number of dynamic obstacles D detected within a given time window T. obj , to reflect the complexity of the driving environment. The more obstacles there are, the higher the task complexity.
[0089] The embodiment of the present invention creates a multi-screen entertainment interaction interface for the occupants of the vehicle based on the multimedia content control path and the user's video playback preference. This can meet the entertainment or information acquisition needs without affecting driving safety, adapt to different driving scenarios, and adjust the playback content according to the needs of passengers in different seats to improve user satisfaction. The multi-screen entertainment interaction interface refers to a system inside the vehicle that provides an interactive entertainment experience for the occupants through multiple screens (or display devices).
[0090] As an embodiment of the present invention, the creation of a multi-screen entertainment interaction interface for the occupants based on the multimedia content control path and the user video playback preference includes: identifying the multi-screen media display device of the occupants based on the multimedia content control path, and defining the content display mode of the multi-screen media display device; setting personalized recommended content for the multi-screen media display device according to the user video playback preference and the content display mode; extracting the main driver's screen in the multi-screen media display device based on the multimedia content control path, and setting a priority control mechanism for the main driver's screen in the multi-screen media display device; analyzing the communication requirements between the multi-screen media display devices, and determining the communication protocol and signal connection channel of the multi-screen media display device based on the communication requirements; constructing an interactive communication unit of the multi-screen media display device according to the communication protocol and the signal connection channel; and creating a multi-screen entertainment interaction interface for the occupants by combining the multi-screen media display device, the priority control mechanism, the personalized recommended content and the interactive communication unit.
[0091] Among them, the multi-screen media display device refers to a display system installed inside the vehicle, consisting of multiple display screens. These display screens are distributed in different locations in the vehicle, such as the front center console, instrument panel, back of the rear seat or roof, etc. The content display mode refers to the presentation form of the content on the multi-screen media display device. For example, navigation information may be displayed in a larger proportion on the left side of the center console, while the music playback interface is displayed in a smaller window on the right. The personalized recommended content refers to multimedia content that meets the user's personal interests recommended to the user on the multi-screen media display device based on the user's video playback preferences (such as favorite movie types, actors, music styles, etc.). The main driver's screen refers to the screen in the multi-screen media display device that is mainly used by the driver. The priority control mechanism refers to the rules for determining the priority and control authority of the main driver's screen in the multi-screen media display device. For example, when the vehicle encounters an emergency or the driver needs to pay attention to important information, the multimedia content will automatically adjust the display to display key driving information. Information such as navigation prompts and vehicle fault warnings are displayed preferentially on the central control screen, and the playback parameters of other media display devices are adaptively adjusted (such as lowering the volume of video playback). The communication requirements refer to the conditions that need to be met for information interaction and collaborative work between the various screens in the multi-screen media display device, including data transmission rate, real-time requirements, data type, etc. The communication protocol refers to the rules and standards formulated to achieve communication between multi-screen media display devices. The signal connection channel refers to the physical or logical channel for signal transmission between the various screens in the multi-screen media display device, including wired connection (such as HDMI, USB, Ethernet, etc.) and wireless connection (such as Wi-Fi, Bluetooth, etc.). The interactive communication unit refers to a functional module constructed based on the communication protocol and signal connection channel for realizing interactive communication between the various screens of the multi-screen media display device. It is responsible for sending, receiving, processing and coordinating data, so that the various screens can cooperate with each other and realize the multi-screen interactive function.
[0092] Optionally, based on the multimedia content control path, the priority control mechanism of the main driver's screen in the multi-screen media display device can be set through C / C++ tools. For example, a permission management class is written in C++ to encapsulate the operation permissions and priority settings of the main driver's screen. Based on the communication requirements, the signal connection channel of the multi-screen media display device can be determined according to the electrical layout and communication protocol requirements of the vehicle. According to the communication protocol and the signal connection channel, the interactive communication unit of the multi-screen media display device can be constructed through Ethernet technology.
[0093] The video playback execution module 106 is used to combine the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface to execute the video playback processing of the target vehicle and obtain the in-vehicle video playback result.
[0094] The video playing process of the target vehicle is executed by combining the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interactive interface, and a vehicle video playing result is obtained. The state and emotion information of the user can be obtained in real time, and personalized video playing content and services are provided for the user. The combination of the multimedia content control path and the multi-screen entertainment interactive interface can avoid the interference of the video playing content on the driver, improve the driving safety, and provide the user with a smoother and more natural intelligent interaction experience. The video playing process refers to a series of processes for comprehensively managing and operating the vehicle video playing by combining multiple elements such as the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interactive interface. The vehicle video playing result refers to the actual effect presented to the user through the video playing process. For example, when the scene perception unit detects that the vehicle enters a tunnel and the light changes sharply, the system automatically converts the video playing mode to an audiobook mode to avoid the video picture from distracting the attention of the driver. The display navigation information and the audiobook information displayed on the main driver screen can be understood as the vehicle video playing result.
[0095] The embodiment of the present invention obtains the occupants and the in-vehicle scene of the target vehicle, and sets the scene perception unit of the target vehicle according to the in-vehicle scene, so as to judge the driver's attention state and the driving mode of the vehicle, avoid the video content from interfering with the driver, and ensure driving safety; further, the embodiment of the present invention creates a real-time emotion monitoring network of the occupants by collecting the voice data and facial images of the occupants, so as to timely detect the driver's fatigue state and concentration level, and monitor the emotional changes of the occupants, thereby helping to optimize the human-computer interaction function of the in-vehicle video playback system; the embodiment of the present invention uses the historical video playback list to play the video. and the real-time emotion monitoring network to identify the user video playback preferences of the target vehicle, and can give priority to updating the video content that the user often watches, thereby increasing the user's dependence on and satisfaction with the system; further, the embodiment of the present invention can reduce visual interference by converting the video playback mode of the video playback content into the audio book mode based on the current driving mode and the driver's attention, so as to enable the driver to focus on driving and help the driver make better use of driving time, and at the same time enrich the entertainment content in the car to meet the interests and hobbies of different users; the embodiment of the present invention sets the video playback mode and the audio book mode according to the real-time emotion monitoring network and the current driving mode. The automatic switching conditions of the vehicle can set differentiated in-vehicle multimedia content control paths to achieve personalized video playback experience for the occupants; further, the embodiment of the present invention sets the multimedia content control path for the occupants based on the automatic switching conditions, the current driving mode and the real-time emotion monitoring network, which can significantly distinguish the multimedia content control logic of the driver side and the rear screen passenger side to ensure that the operations of the two do not interfere with each other; the embodiment of the present invention creates a multi-screen entertainment interaction interface for the occupants based on the multimedia content control path and the user's video playback preferences, which can meet the entertainment or information acquisition needs without affecting driving safety, and adapt to different Driving scene, and at the same time, the playback content can be adjusted according to the needs of passengers in different seats to improve user satisfaction; finally, the embodiment of the present invention combines the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface to execute the video playback processing of the target vehicle, obtain the in-vehicle video playback result, and can obtain the user's status and emotional information in real time, and provide users with personalized video playback content and services. The combination of the multimedia content control path and the multi-screen entertainment interaction interface can avoid the video playback content from interfering with the driver, improve driving safety, and provide users with a smoother and more natural intelligent interactive experience. Therefore, the embodiment of the present invention provides an in-vehicle video playback system and method that can monitor the emotional state of people in the car in real time and provide personalized in-vehicle video playback services.
[0096] like Figure 2 FIG. 1 is a flow chart of a method for playing a video in a vehicle according to an embodiment of the present invention. In this embodiment, the method for playing a video in a vehicle includes:
[0097] Acquire the occupants and vehicle scene of the target vehicle, wherein the vehicle scene includes the vehicle status, the vehicle scene, and the vehicle environment, and set the scene perception unit of the target vehicle according to the vehicle scene;
[0098] Collecting voice data and facial images of the occupants of the vehicle, and creating a real-time emotion monitoring network for the occupants of the vehicle by combining the voice data, the facial images, and the vehicle scene;
[0099] Reading a historical video playlist of the target vehicle, extracting video play content in the historical video playlist, and identifying a user video play preference of the target vehicle based on the historical video playlist and the real-time emotion monitoring network;
[0100] determining a current driving mode of the target vehicle based on the scene perception unit, monitoring the driver's attention of the target vehicle based on the real-time emotion monitoring network, and switching a video playback mode of the video content to an audiobook mode based on the current driving mode and the driver's attention;
[0101] setting an automatic switching condition between the video playback mode and the audiobook mode based on the real-time emotion monitoring network and the current driving mode; setting a multimedia content control path for the vehicle occupant based on the automatic switching condition, the current driving mode, and the real-time emotion monitoring network; and creating a multi-screen entertainment interaction interface for the vehicle occupant based on the multimedia content control path and the user's video playback preference;
[0102] In combination with the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface, the video playback processing of the target vehicle is executed to obtain the vehicle-mounted video playback result.
[0103] In the several embodiments provided by the present invention, it should be understood that the provided systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and actual implementation may employ other division methods.
[0104] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A vehicle-mounted video playback system, characterized in that: The system includes: an in-vehicle scene perception module, an occupant emotion monitoring module, a video preference recognition module, a playback mode switching module, a multi-screen interaction module, and a video playback execution module; The vehicle scene perception module is used to obtain the occupants and vehicle scene of the target vehicle, wherein the vehicle scene includes the vehicle status, the interior scene and the exterior environment, and the scene perception unit of the target vehicle is set according to the vehicle scene; The occupant emotion monitoring module is used to collect voice data and facial images of the occupants of the vehicle, and to create a real-time emotion monitoring network for the occupants by combining the voice data, facial images, and the vehicle scene; The video preference identification module is used to read the historical video playlist of the target vehicle, extract the video playback content in the historical video playlist, and identify the user video playback preference of the target vehicle based on the historical video playlist and the real-time emotion monitoring network; The playback mode switching module is configured to determine a current driving mode of the target vehicle based on the scene perception unit, monitor the driver's attention of the target vehicle based on the real-time emotion monitoring network, and switch the video playback mode of the video playback content to an audiobook mode based on the current driving mode and the driver's attention; The multi-screen interaction module is configured to set automatic switching conditions between the video playback mode and the audiobook mode based on the real-time emotion monitoring network and the current driving mode; set a multimedia content control path for the vehicle occupants based on the automatic switching conditions, the current driving mode, and the real-time emotion monitoring network; and create a multi-screen entertainment interaction interface for the vehicle occupants based on the multimedia content control path and the user's video playback preferences; The video playback execution module is used to combine the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface to execute the video playback processing of the target vehicle and obtain the in-vehicle video playback result.
2. The vehicle-mounted video playback system according to claim 1, wherein: The step of setting the scene perception unit of the target vehicle according to the vehicle-mounted scenario includes: Collecting vehicle driving data, external environment data, and occupant behavior data of the target vehicle according to the vehicle-mounted scenario; identifying a driving mode of the target vehicle based on the vehicle driving data, the external environment data, and the occupant behavior data; Extracting facial features and hand movement features of the occupants of the target vehicle based on the occupant behavior data; Identifying sound data of the target vehicle from the vehicle driving data, and extracting engine sound frequency and occupant voice features of the target vehicle based on the sound data; Building a mapping relationship between the in-vehicle scenario and the driving mode by combining the facial features of the occupant, the hand movement features of the occupant, the engine sound frequency, and the voice features of the occupant; Based on the mapping relationship, a scene perception unit of the target vehicle is set.
3. The vehicle-mounted video playback system according to claim 1, wherein: The step of combining the voice data, the facial image, and the vehicle scene to create a real-time emotion monitoring network for the vehicle occupants includes: determining different seating areas for the occupants of the vehicle based on the voice data; identifying the emotional states of people in the different seating areas based on the voice data and the facial image; Extracting facial feature points of the occupant from the facial image, and combining the emotional state of the occupant with the facial feature points to generate an emotional heat map of the different seating areas; Identifying a vehicle driving state in the vehicle scenario, and analyzing the scene sensitivity of the occupant based on the vehicle driving state and the emotional state of the occupant; Constructing cabin emotion zones for the occupants of the vehicle based on the scene sensitivity and the emotion heat map; Based on the cabin emotion zoning, a real-time emotion monitoring network for the occupants of the vehicle is created.
4. The vehicle-mounted video playback system according to claim 1, wherein: The identifying the user video playback preference of the target vehicle based on the historical video playlist and the real-time emotion monitoring network includes: determining an occupant seating layout of the target vehicle based on the real-time emotion monitoring network; identifying the occupant composition of the target vehicle based on the occupant seating layout and the real-time emotion monitoring network; Extracting the video ID and the corresponding timestamp data from the historical video playlist according to the vehicle occupant combination; Based on the video ID and the timestamp data, identifying the frequently played content of the vehicle occupant combination; The user video playback preference of the target vehicle is identified based on the high-frequency playback content.
5. The vehicle-mounted video playback system according to claim 1, wherein: The determining, based on the scene perception unit, the current driving mode of the target vehicle includes: Based on the scene perception unit, identifying the driving situation of the target vehicle; analyzing a vehicle motion state and a driving road type of the target vehicle according to the driving scenario; Based on the vehicle motion state, collecting engine sound data and passenger sound data of the target vehicle; identifying an engine operating condition of the target vehicle based on the engine sound data; Analyzing the voice characteristics of the occupant of the target vehicle based on the occupant voice data; extracting driving behavior characteristics of the target vehicle based on the vehicle motion state, the driving road type, and the engine operating condition; The current driving mode of the target vehicle is determined based on the driving behavior characteristics and the voice characteristics of the occupant.
6. The vehicle-mounted video playback system according to claim 1, wherein: The converting the video playback mode of the video playback content to the audiobook mode based on the current driving mode and the driver's attention includes: Identifying the material type of the video playback content, and extracting the voice track of the video playback content based on the material type; Based on the voice track, performing format conversion processing on the video playback content to obtain an audio playback format; Generating audio playback content corresponding to the video playback content according to the audio playback format; Creating an audio control interface for the video playback content based on the audio playback content; Setting audio output parameters of the audio playback content according to the current driving mode; Setting a mode conversion rule for the video playback mode according to the current driving mode and the driver's attention; The video playback mode of the video playback content is converted into a book listening mode in combination with the mode conversion rule, the voice track, the audio output parameter and the audio control interface.
7. The vehicle-mounted video playback system according to claim 1, wherein: The automatic switching condition between the video playback mode and the audiobook mode is set according to the real-time emotion monitoring network and the current driving mode, including: Identifying the characteristics of the driver and driving scene in the current driving mode; Determining the scene complexity of the current driving mode according to the driving scene characteristics; detecting the driver's visual attention based on the real-time emotion monitoring network; analyzing a trend of changes in visual attention based on the complexity of the scene and identifying the driver's cognitive load; Calculating the driver's continuous concentration time based on the attention change trend; Defining the driver's attention level according to the continuous concentration time; Based on the complexity of the scene, setting a mode switching threshold corresponding to the attention level and the cognitive load; According to the mode switching threshold, an automatic switching condition between the video playback mode and the book listening mode is set.
8. The vehicle-mounted video playback system according to claim 1, wherein: The step of setting a multimedia content control path for the occupant based on the automatic switching condition, the current driving mode, and the real-time emotion monitoring network includes: Identifying a driving scenario corresponding to the current driving mode and determining the driver and fellow passengers in the vehicle; Analyzing the task complexity of the current driving mode based on the driving scenario; identifying a multimedia playback mode in the current driving mode according to the automatic switching condition; Extracting the emotion change trend of the fellow passenger under the multimedia playback mode based on the real-time emotion monitoring network; determining, based on the emotion change trend, the degree of the passenger's enthusiasm for the multimedia playback format; Based on the degree of enthusiasm, a mechanism is created for the passengers to freely switch the multimedia playback format; According to the driving scenario and the complexity of the task, setting the driver's seat occupant's behavioral restriction authority and restriction lifting rules for the multimedia playback format; In combination with the multimedia playback format, the free switching mechanism, the behavior restriction authority and the restriction lifting rules, a multimedia content control path for the occupants of the vehicle is set.
9. The vehicle-mounted video playback system according to claim 1, wherein: The step of creating a multi-screen entertainment interaction interface for the vehicle occupants based on the multimedia content control path and the user's video playback preference includes: Based on the multimedia content control path, identifying the multi-screen media display device of the occupant of the vehicle, and defining a content display mode of the multi-screen media display device; Setting personalized recommended content for the multi-screen media display device according to the user's video playback preference and the content display mode; extracting a main driver's screen in the multi-screen media display device based on the multimedia content control path, and setting a priority control mechanism for the main driver's screen in the multi-screen media display device; Analyzing communication requirements between the multi-screen media display devices, and determining communication protocols and signal connection channels of the multi-screen media display devices based on the communication requirements; constructing an interactive communication unit of the multi-screen media display device according to the communication protocol and the signal connection channel; By combining the multi-screen media display device, the priority control mechanism, the personalized recommended content and the interactive communication unit, a multi-screen entertainment interactive interface for the occupants of the vehicle is created.
10. A vehicle-mounted video playback method, using the vehicle-mounted video playback system according to any one of claims 1 to 9, characterized in that: The method comprises: Acquire the occupants and vehicle scene of the target vehicle, wherein the vehicle scene includes the vehicle status, the vehicle scene, and the vehicle environment, and set the scene perception unit of the target vehicle according to the vehicle scene; Collecting voice data and facial images of the occupants of the vehicle, and creating a real-time emotion monitoring network for the occupants of the vehicle by combining the voice data, the facial images, and the vehicle scene; Reading a historical video playlist of the target vehicle, extracting video play content in the historical video playlist, and identifying a user video play preference of the target vehicle based on the historical video playlist and the real-time emotion monitoring network; determining a current driving mode of the target vehicle based on the scene perception unit, monitoring the driver's attention of the target vehicle based on the real-time emotion monitoring network, and switching a video playback mode of the video content to an audiobook mode based on the current driving mode and the driver's attention; setting an automatic switching condition between the video playback mode and the audiobook mode based on the real-time emotion monitoring network and the current driving mode; setting a multimedia content control path for the vehicle occupant based on the automatic switching condition, the current driving mode, and the real-time emotion monitoring network; and creating a multi-screen entertainment interaction interface for the vehicle occupant based on the multimedia content control path and the user's video playback preference; In combination with the scene perception unit, the real-time emotion monitoring network, the multimedia content control path and the multi-screen entertainment interaction interface, the video playback processing of the target vehicle is executed to obtain the vehicle-mounted video playback result.
Citation Information
Patent Citations
Vehicle driving mode switching method and device and computer readable storage medium
CN116142197A
Scene mode determination method and device, vehicle-mounted terminal and vehicle
CN118799941A
Cited By
A video playing self-adaptive adjustment method and system based on multi-modal biometric feature fusion, an electronic device and a storage medium
CN122698805A