Vehicle-mounted music adaptive pushing method, vehicle and storage medium

By combining data on the emotional state and environment of in-vehicle users, target content is determined from the user's music preference library, and a push strategy is determined based on real-time data. This solves the problem of poor adaptability of in-vehicle music push, achieves accurate matching of music content with user emotions and scenarios, and improves user experience.

CN122196224APending Publication Date: 2026-06-12GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GREAT WALL MOTOR CO LTD
Filing Date
2026-03-12
Publication Date
2026-06-12

Smart Images

  • Figure CN122196224A_ABST
    Figure CN122196224A_ABST
Patent Text Reader

Abstract

The application provides a vehicle-mounted music adaptive pushing method, a vehicle and a storage medium, and relates to the technical field of intelligent vehicles.The method comprises the following steps: firstly, based on the emotional state of a user in a vehicle and scene environment data, target content matched with the emotional state and the scene environment data is determined from a pre-constructed user music preference library.The scene environment data is used to represent the corresponding natural environment, time dimension and space travel scene of the user during the driving process of the vehicle.Secondly, based on the music playing state in the vehicle, the change characteristics of the emotional state and the interaction data of the user and the in-vehicle playing device, a pushing strategy of the target content is determined.The change characteristics are used to indicate the dynamic change attribute, fluctuation level and state switching characteristics of the emotional state of the user over time.Finally, based on the pushing strategy, the target content is pushed to the user.The technical problem of poor adaptability of vehicle-mounted music pushing in the related art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent vehicle technology, and in particular to an adaptive in-vehicle music push method, a vehicle, and a storage medium. Background Technology

[0002] With the rapid development of automotive intelligence and connectivity technologies, in-vehicle entertainment systems have become one of the core configurations for enhancing the user's driving experience. As a core application module of in-vehicle entertainment systems, the development of in-vehicle music push function is showing a trend of transformation from manual operation to automation, and from generalization to personalization.

[0003] Most in-car music push solutions focus only on users' historical listening preferences, lacking awareness and adaptation to users' real-time emotional states. They cannot adjust the push content according to the dynamic changes in users' emotions, resulting in a disconnect between the push content and the user's current mood. The scene adaptation dimension is relatively simple, often only combining a single scene factor, resulting in low push accuracy.

[0004] In summary, the relevant technologies suffer from poor compatibility with in-vehicle music streaming. Summary of the Invention

[0005] In view of the above problems, this application provides a method, vehicle, and storage medium for adaptive in-vehicle music push that overcomes or at least partially solves the current problem of poor adaptability of in-vehicle music push. The technical solution is as follows: An adaptive in-vehicle music push method, the method comprising: Based on the emotional state and scene environment data of the users in the vehicle, target content that matches the emotional state and scene environment data is determined from a pre-built user music preference library. The scene environment data is used to characterize the natural environment, time dimension and spatial travel scenario of the user during the vehicle driving process. Based on the music playback status in the car, the change characteristics of the emotional state, and the interaction data between the user and the in-car playback device, a push strategy for pushing the target content is determined. The change characteristics are used to indicate the dynamic change attributes, fluctuation magnitude, and state switching characteristics of the user's emotional state over time. Based on the aforementioned push strategy, the target content is pushed to the user.

[0006] In this way, by combining the emotional state of the user in the car with the scene environment data to determine the matching target content, and then determining the push strategy based on the music playback status, emotional state change characteristics and user interaction data, the target content is finally pushed according to the push strategy. This enables the pushed music content to be adapted to both the user's emotions and the driving scenario, and the push method to fit the real-time status of the vehicle, realizing adaptive push of in-vehicle music and improving the accuracy and intelligence of music push.

[0007] Optionally, the method further includes: Acquire state perception data of the user inside the vehicle. The state perception data is used to characterize the user's emotion-related features and vehicle driving operation features that can reflect the user's emotions. The state perception data is classified to determine different types of state perception sub-data. The state perception sub-data is used to characterize the user's own emotion-related features and vehicle driving behavior association data related to the user's emotions. Feature extraction and emotion feature analysis are performed on the state-aware sub-data respectively to obtain the emotion analysis results corresponding to each type of state-aware sub-data. The emotion analysis results are fused to obtain the emotional state of the user in the vehicle.

[0008] In this way, by acquiring state perception data that characterizes user emotions and reflects driving operation characteristics, the data is classified into different types of state perception sub-data. Then, feature extraction and emotion analysis are performed on each sub-data separately, and finally, the analysis results are integrated to obtain the emotional state. This achieves multi-dimensional and refined processing of emotion-related data, avoids the limitations of single data processing, improves the accuracy and comprehensiveness of determining the emotional state of users in the vehicle, and provides reliable emotional data support for the matching of subsequent target content.

[0009] Optionally, determining target content matching the emotional state from a pre-built user music preference library based on the emotional state and scene environment data includes: Determine the first degree of match between the emotional state and multiple items in the user's music preference library; Based on the first matching degree, multiple candidate contents are determined; Determine the second matching degree between the scene environment data and multiple candidate contents; The target content is determined based on the second matching degree.

[0010] In this way, by first determining the initial match between the emotional state and multiple contents in the user's music preference library, filtering multiple candidate contents based on the initial match, then determining the second match between the scene environment data and each candidate content, and finally determining the target content based on the second match, the two-level matching method of "initial emotional screening + scene fine screening" is adopted to gradually narrow down the range of music content. This can filter out target content that is suitable for both the user's emotional state and scene environment data, effectively improving the accuracy of target content determination and reducing the probability of music content with low matching degree being pushed.

[0011] Optionally, the method further includes: The candidate content corresponding to the second matching degree is sorted according to the priority of the target matching degree to generate a target content list, which includes multiple candidate contents. Based on the push strategy, the candidate content is pushed to the user in the order of the target content list.

[0012] In this way, by prioritizing the candidate content corresponding to the second matching degree according to the target matching degree, a target content list containing multiple candidate contents is generated, and the candidate contents are pushed in the order of the list. The target matching degree can reflect the comprehensive adaptability of the content to the user's needs. Prioritizing the content can push the content with higher adaptability first, which can make music push more orderly and ensure that users get the music content that best matches their own emotions and needs first, further improving the adaptability and user experience of music push.

[0013] Optionally, the step of determining the push strategy for the target content based on the music playback status in the vehicle, the changing characteristics of the emotional state, and the interaction data between the user and the in-vehicle playback device includes: Based on the music playback status, the change characteristics of the emotional state, and the interaction data, the current playback scenario in the vehicle is determined. The playback scenario is used to characterize the scenarios related to the music push strategy formulation in the vehicle environment, as well as the music push intervention adaptation requirements corresponding to the change characteristics. Based on the playback scenario, the push strategy is determined. The push strategy includes at least one of the following: an automatic playback push strategy, a push strategy that displays the target content on the in-vehicle interactive interface, and a push strategy that seamlessly switches the music being played.

[0014] In this way, by first combining the in-car music playback status, emotional state change characteristics, and user interaction data to determine the current in-car playback scenario, and then determining the push strategy based on the playback scenario, the push strategy includes at least one of three selectable types. The playback scenario can characterize the intervention and adaptation requirements of music push in the in-vehicle environment. The strategy determined based on the scenario is more in line with the actual push needs, and multiple strategies are available, which can improve the adaptability of the push strategy to the real-time in-vehicle scenario, while increasing the flexibility of the push method and adapting to the music push needs in different in-vehicle scenarios.

[0015] Optionally, when there are multiple users in the vehicle, the method further includes: Based on the scene environment data and the emotional state of each user, the target content corresponding to each user is determined from the pre-built user music preference library; Based on the in-vehicle zone audio system, an independent audio playback channel is matched for each user's seating position; The target content for each user is pushed to the audio playback channel at the corresponding user's seat location.

[0016] In this way, when there are multiple users in the car, the target content for each user is determined based on scene environment data and each user's emotional state. Then, based on the in-vehicle zone audio system, an independent audio playback channel is matched for each user's seat position, and the target content for each user is pushed to the corresponding channel. The target content for each user is determined independently, and the playback channels are independent of each other. This can meet the personalized music needs of different users in the car, realize independent music push in multi-user scenarios, avoid mutual interference between music playback of different users, and improve the music experience when multiple users are driving.

[0017] Optionally, the method further includes: If a conflict is detected between the push strategies of different users, the target content or push strategy of each user will be adjusted based on the push priority rules and the interaction data between the user and the in-vehicle playback device.

[0018] In this way, when a conflict is detected between the push strategies of different users, the target content or push strategy of each user is adjusted based on the push priority rules and the interaction data between the user and the in-vehicle playback device. This provides a clear solution to the push conflict of multiple users, which can effectively solve the problem of music push conflict in multi-user scenarios, avoid push chaos, ensure the smoothness and rationality of in-vehicle music push in multi-user scenarios, and take into account the push needs of different users.

[0019] Optionally, the method further includes: When the emotional state of the user in the vehicle is identified as a preset anxiety, at least one music adjustment option is generated to adjust the style of the target content. Obtain the selection command input by the user based on the music adjustment options; Based on the selection instruction, the corresponding target adjustment music is determined, and the target adjustment music is used as the target content and pushed to the user according to the push strategy.

[0020] In this way, when the user's emotional state is identified as preset anxiety, at least one music adjustment option for adjusting the style of the target content is generated, the user's selection instruction is obtained, and the target adjustment music is determined and pushed based on the instruction. The system provides a music adjustment scheme that the user can choose independently for anxiety, and the adjustment direction focuses on the style and / or genre of the target content. This can achieve proactive adjustment of the user's anxiety, improve the targeting and practicality of in-car music push, and meet the user's personalized music adjustment needs when in an anxious state.

[0021] An adaptive music delivery device for in-vehicle use, the device comprising: The first determining module is used to determine target content that matches the emotional state and scene environment data of the user in the vehicle from a pre-built user music preference library. The scene environment data is used to characterize the natural environment, time dimension and spatial travel scenario of the user during the vehicle driving process. The second determining module is used to determine a push strategy for pushing the target content based on the music playback status in the vehicle, the change characteristics of the emotional state, and the interaction data between the user and the in-vehicle playback device. The change characteristics are used to indicate the dynamic change attributes, fluctuation magnitude, and state switching characteristics of the user's emotional state over time. The push module is used to push the target content to the user based on the push strategy.

[0022] Optionally, the in-vehicle music adaptive push device also includes an acquisition module: The acquisition module acquires the state perception data of the user inside the vehicle. The state perception data is used to characterize the user's emotion-related features and vehicle driving operation features that can reflect the user's emotions. The state perception data is classified to determine different types of state perception sub-data. The state perception sub-data is used to characterize the user's own emotion-related features and vehicle driving behavior association data related to the user's emotions. Feature extraction and emotion feature analysis are performed on the state-aware sub-data respectively to obtain the emotion analysis results corresponding to each type of state-aware sub-data. The emotion analysis results are fused to obtain the emotional state of the user in the vehicle.

[0023] Optionally, the first determining module is also used for: Determine the first degree of match between the emotional state and multiple items in the user's music preference library; Based on the first matching degree, multiple candidate contents are determined; Determine the second matching degree between the scene environment data and multiple candidate contents; The target content is determined based on the second matching degree.

[0024] Optionally, the in-vehicle music adaptive push device also includes a first generation module: The first generation module is used to sort the candidate content corresponding to the second matching degree according to the priority of the target matching degree and generate a target content list, wherein the target content list includes multiple candidate contents; Based on the push strategy, the candidate content is pushed to the user in the order of the target content list.

[0025] Optionally, the second determining module is also used for: Based on the music playback status, the change characteristics of the emotional state, and the interaction data, the current playback scenario in the vehicle is determined. The playback scenario is used to characterize the scenarios related to the music push strategy formulation in the vehicle environment, as well as the music push intervention adaptation requirements corresponding to the change characteristics. Based on the playback scenario, the push strategy is determined. The push strategy includes at least one of the following: an automatic playback push strategy, a push strategy that displays the target content on the in-vehicle interactive interface, and a push strategy that seamlessly switches the music being played.

[0026] Optionally, when multiple users are present in the vehicle, the in-vehicle music adaptive push device also includes a third determining module: The third determining module is used to determine the target content corresponding to each user from a pre-built user music preference library based on the scene environment data and the emotional state of each user. Based on the in-vehicle zone audio system, an independent audio playback channel is matched for each user's seating position; The target content for each user is pushed to the audio playback channel at the corresponding user's seat location.

[0027] Optionally, the adaptive in-vehicle music delivery system also includes an adjustment module: The adjustment module is used to adjust the target content or push strategy of each user based on the push priority rules and the interaction data between the user and the in-vehicle playback device if a conflict is detected between the push strategies of different users.

[0028] Optionally, the in-vehicle music adaptive push device also includes a second generation module: The second generation module is used to generate at least one music adjustment option for adjusting the target content when the emotional state of the user in the vehicle is identified as a preset anxiety emotion. The music adjustment option is used to adjust the style of the target content. Obtain the selection command input by the user based on the music adjustment options; Based on the selection instruction, the corresponding target adjustment music is determined, and the target adjustment music is used as the target content and pushed to the user according to the push strategy.

[0029] A vehicle comprising: the vehicle performing any of the optional in-vehicle music adaptive push methods described above.

[0030] A computer-readable storage medium includes: a computer program stored on the computer-readable storage medium, wherein the computer program, when executed by a processor, implements any of the optional in-vehicle music adaptive push methods described above.

[0031] By employing the aforementioned technical solution, this application provides an adaptive in-vehicle music push method. This method combines the emotional state of the in-vehicle user with scene environment data representing the vehicle's natural driving environment, time dimension, and spatial travel scenario. It then determines matching target content from the user's music preference library, ensuring that the pushed music content simultaneously matches the user's real-time emotions and actual driving scenario, significantly improving the accuracy of matching music content with the user's current needs. Furthermore, based on the real-time music playback status in the vehicle, emotional state change characteristics reflecting dynamic changes and fluctuations in the user's emotions, and interaction data between the user and the in-vehicle playback device, a corresponding push strategy is determined. This ensures that the push strategy aligns with the real-time playback status of the in-vehicle system, the dynamic changes in the user's emotions, and the user's interaction habits, guaranteeing the rationality and adaptability of the push method. Finally, by pushing the target content to the user according to this strategy, the entire process of adaptive intelligent push for in-vehicle music can be achieved without additional manual intervention from the user. This simplifies the operation of in-vehicle music use while effectively improving the intelligence level of in-vehicle music push and the overall driving music experience.

[0032] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0033] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This illustration shows one of the flowcharts of the adaptive in-vehicle music push method provided in an embodiment of this application; Figure 2 This is a second schematic flowchart of the adaptive in-vehicle music push method provided in an embodiment of this application; Figure 3 The third flowchart illustrates the adaptive in-vehicle music push method provided in this application embodiment; Figure 4 The fourth flowchart illustrates the adaptive in-vehicle music push method provided in this application embodiment; Figure 5The fifth flowchart illustrates the adaptive in-vehicle music push method provided in this application embodiment; Figure 6 A schematic diagram of the structure of an adaptive in-vehicle music push device provided in an embodiment of this application is shown. Detailed Implementation

[0034] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0035] To address the technical problem of poor adaptability of in-vehicle music push notifications in related technologies, this application provides an adaptive in-vehicle music push method, such as... Figure 1 As shown, Figure 1 This is a schematic flowchart illustrating an adaptive in-vehicle music push method provided in an embodiment of this application. The method includes: S11. Based on the emotional state and scene environment data of the users in the vehicle, determine the target content that matches the emotional state and scene environment data from the pre-built user music preference library.

[0036] Among them, scene environment data is used to characterize the natural environment, time dimension, and spatial travel scenario of the user during the vehicle driving process. Emotional state is the user's real-time mood (such as calm, anxiety, pleasure, etc.) obtained through multi-dimensional data identification. It is the core subjective basis for characterizing the user's current music needs. Scene environment data can include: natural environment data (such as weather, lighting, road conditions), time dimension data (such as time period, date, season), and spatial travel scenario data of the user driving (such as commuting, long-distance travel, leisure travel, destination type). It is the objective scene basis for characterizing the user's music needs.

[0037] Specifically, a pre-built user music preference library is used as the sole data source for music content filtering. This user music preference library is pre-generated based on the user's historical in-car music operation behavior (such as play, favorite, like, skip, delete). The library stores a massive amount of music content and corresponding feature tags (such as music style, rhythm, genre, mood matching type, scene matching type) to ensure that the filtered content conforms to the user's long-term music preferences.

[0038] Specifically, the process involves extracting emotional feature tags corresponding to the user's emotional state and matching them with the emotional matching tags of various music content in the user's music preference library to calculate the emotional matching degree. It also involves extracting scene feature tags (such as rainy day, soothing, commuting, upbeat, nighttime, and gentle) corresponding to the scene environment data and matching them with the scene matching tags of various music content in the library to calculate the scene matching degree. Finally, by combining the emotional matching degree and the scene matching degree, a reasonable matching threshold is set, and music content that simultaneously meets the threshold requirements for both emotional and scene matching degrees is selected as the final target content.

[0039] For example, assuming the vehicle is in motion and the user's emotional state is identified as "anxious," the collected environmental data includes: a rainy day and congested traffic, a weekday morning rush hour, and a city commute. Then, a pre-built user music preference library is invoked (containing frequently listened-to soothing and upbeat music, each tagged with "anxiety-appropriate," "rainy day-appropriate," and "commuting-appropriate"). First, the emotional match between each piece of music in the library and the emotion of "anxiety" is calculated, filtering out soothing and calming music with an emotional match ≥80%. Next, the scene match between these pieces and the "rainy day, morning rush hour, city commute" scenario is calculated, filtering out music with a scene match ≥75%. Finally, considering both match scores, the three most highly matched soothing pieces (all types previously saved by the user) are selected as the target content for this push notification, thus simultaneously catering to the user's anxious state of mind and the needs of a rainy morning rush hour commute.

[0040] This embodiment breaks through the limitations of existing in-vehicle music push that only consider user preferences or a single scenario. By using real-time user emotional state and scenario environment data across the entire driving and riding experience as dual matching criteria, target content is filtered from a pre-built user music preference library. This ensures that the selected target content can simultaneously adapt to the user's current mood and actual driving and riding scenario, thereby improving the personalization and accuracy of in-vehicle music push from the source. It effectively solves the technical pain point of the push content being out of touch with the user's real-time needs and driving and riding scenario, laying the core content foundation for subsequent adaptive push, while reducing the probability of pushing invalid music content and improving the user's auditory experience during driving and riding.

[0041] S12. Based on the music playback status in the car, the characteristics of changes in emotional state, and the interaction data between the user and the in-car playback device, determine the push strategy for the target content.

[0042] Among them, the change characteristics are used to indicate the dynamic changes in a user's emotional state over time, the magnitude of fluctuations, and the characteristics of state transitions.

[0043] Specifically, the in-car music playback status reflects the current operating status of the in-car music system, serving as the foundation for developing push strategies. This includes three core states: "playing," "paused," and "not playing," along with auxiliary data such as the current playback progress and volume. The emotional state change characteristics indicate three dynamic attributes of the user's emotional state over time: dynamic attribute changes (e.g., from anxiety to calm, from pleasure to indifference), fluctuation magnitude (e.g., small, medium, large amplitude of emotional fluctuation), and state switching characteristics (e.g., switching frequency, switching direction). This is the core basis for characterizing dynamic changes in user emotions and determining whether the user's music needs have changed. The user's interaction data with the in-car playback device represents the user's real-time operational intent regarding in-car music pushes. This includes user confirmation commands, rejection commands, switching commands, skipping commands, volume adjustment commands, and user-initiated music searches and selections. This is a key basis for adjusting push strategies and aligning with the user's subjective intentions.

[0044] Specifically, a comprehensive analysis is conducted on three types of data: music playback status, emotional state, and user interaction data with the in-vehicle playback device. This analysis uncovers the correlations between the data (e.g., high emotional fluctuation + music not playing + no interaction commands → proactive intervention push notifications are needed; stable emotional state + music playing + no interaction commands → no proactive push notifications required). Based on these data characteristics, the current in-vehicle playback scenario is determined. This scenario characterizes the overall state related to music push strategy formulation within the in-vehicle environment, as well as the corresponding music push intervention adaptation requirements (e.g., whether proactive push notifications are needed, the intensity of push intervention, and the push method). Based on the determined playback scenario and the corresponding push intervention adaptation requirements, a suitable push strategy is selected from a pre-set set of push strategies. Simultaneously, the specific execution rules of the push strategy are clarified to ensure a high degree of adaptation between the push strategy and the current real-time in-vehicle state and user needs. The push strategy must cover the core push methods in different scenarios to ensure flexibility.

[0045] For example, assuming the current music playback status in the car is "not playing," real-time monitoring reveals the following characteristics of the user's emotional state changes: the emotional state changes from "calm" to "anxious" (dynamic attribute change), the fluctuation level is "high fluctuation" (large emotional fluctuation amplitude), and the state switching characteristic is "single switch (calm → anxiety)." Simultaneously, the collected user interaction data with the in-car playback device shows "no real-time interaction commands" (the user did not operate the playback device). Analyzing these three types of data comprehensively reveals: music not playing (basic environment), high user emotional fluctuation shifting towards negative emotions (requiring in-depth push intervention), and no interaction commands (the user did not explicitly refuse). Based on the intent, the current in-car playback scenario is determined to be a "negative emotion, no playback scenario." The corresponding push intervention requirements for this scenario are "proactive push, lightweight reminder, and quick relief of negative emotions." Secondly, based on this scenario, an "automatic playback push strategy" is selected from the preset push strategy set. The execution rule is clearly defined as "first, issue a short push reminder via in-car voice (such as 'Detecting that you are in an anxious state, we are pushing soothing music for you'), and then automatically play the target content without requiring manual confirmation from the user." Finally, this automatic playback push strategy is determined to be the core strategy for this push, ensuring that the push method adapts to the current real-time in-car status and user needs.

[0046] In this embodiment, by comprehensively considering three types of core real-time data in the in-vehicle scenario—the status of in-vehicle music playback, the characteristics of changes in the user's emotional state, and the interaction data between the user and the in-vehicle playback device—a push strategy adapted to the current real-time status of the in-vehicle system is formulated. This allows the push strategy to dynamically adapt to changes in the user's emotions, the running status of the in-vehicle music, and the user's real-time operational intentions, thereby improving the rationality, flexibility, and scenario adaptability of the push strategy. It avoids the decline in user experience caused by blind or ineffective pushes, while reducing the frequency of drivers manually operating the in-vehicle playback device, reducing the risk of driver distraction, and adapting to the core requirements of "safety, convenience, and intelligence" in the in-vehicle scenario, thus achieving adaptive adjustment of the push method.

[0047] S13. Based on the push strategy, push target content to users.

[0048] Specifically, based on the specific type of push strategy, corresponding push operations are executed to ensure a standardized and smooth operation process. Push strategies can be categorized into three execution methods: For automatic playback push strategies, the in-vehicle audio system is directly invoked to automatically play the target content at a preset volume; for push strategies displayed on the in-vehicle interactive interface, the target content is sorted by matching priority and displayed in a list format on the in-vehicle central control screen, dashboard, and other interactive interfaces, with content characteristics (such as emotional and scene adaptation) labeled for user selection; for seamless switching push strategies, the target content is pre-loaded, and during breaks in currently playing music (such as song endings or interludes), the content is switched to without lag or interruption to avoid affecting the user's auditory experience. During push execution, push feedback data (such as whether the user skipped, adjusted the volume, or switched content) is collected in real time. This data serves as auxiliary data for subsequent optimization of push strategies and updating the user's music preference library, further improving the accuracy of adaptive pushes.

[0049] Specifically, if there are multiple target contents (such as a list of target contents), the push will be executed in descending order of the overall matching degree of the target contents, prioritizing the content with the highest suitability; a preset limit on the number of pushes can be set (such as 3-5 items per push) to avoid excessive content from interfering with the user; target content preloading is supported during push execution to ensure seamless switching of push strategies without lag or interruption; when the automatic playback push strategy is executed, an initial volume adapted to the in-vehicle scenario (such as 50% volume) is preset to avoid the volume being too high or too low and affecting the user; during push execution, real-time user feedback data (skip, switch, volume adjustment, confirmation) is collected and synchronized to the user's music preference library and push strategy judgment model for dynamic optimization of subsequent target content selection and push strategy formulation.

[0050] For example, if the push strategy determined in S12 is "auto-play push strategy" and the execution rule is "auto-play after voice reminder", and the target content determined in S11 is 3 soothing light music tracks (sorted from high to low overall matching degree), then this step is specifically executed as follows: First, a short reminder is issued through the in-vehicle voice module: "We have detected that you are in an anxious state, so we are pushing soothing music to relieve your anxiety." After the reminder ends, no manual operation is required from the user. The in-vehicle audio system is automatically activated, and the soothing light music track with the highest overall matching degree is played first from the 3 target tracks at a preset initial volume of 50%. During the playback, user feedback (such as whether to adjust the volume) is collected in real time to complete the auto-play push.

[0051] In one specific embodiment, when the emotional state of the user in the vehicle is identified as a preset anxiety emotion, at least one music adjustment option is generated to adjust the style of the target content; the selection command input by the user based on the music adjustment option is obtained; the corresponding target adjustment music is determined based on the selection command, and the target adjustment music is used as the target content and pushed to the user according to the push strategy.

[0052] Specifically, when the user's emotion is confirmed to be the preset anxiety, at least one music adjustment option is automatically generated to adjust the target content. The scope of the adjustment option is strictly limited to the "style of the target content". That is, the adjustment option does not change the core adaptation attribute of the target content (still meeting the user's anxiety relief needs), and only provides differentiated choices in terms of style and genre, so that the adjusted music can still adapt to the user's current mood and in-car scenario. At the same time, the generated adjustment options need to be clearly marked with the adjustment direction (such as "soothing instrumental music" or "upbeat folk music") to facilitate users' quick understanding and selection. The number of adjustment options can be preset (such as 2-4 types) to avoid too many options interfering with user operation.

[0053] The generated music adjustment options are displayed through the in-vehicle interactive interface (central control screen, instrument panel) or communicated to the user via in-vehicle voice broadcast, allowing the user to input selection commands based on their adjustment needs. Users can input corresponding selection commands through in-vehicle touch operations, voice commands, steering wheel buttons, and other in-vehicle scenario-adapted interaction methods. The system collects these commands in real time, ensuring real-time and convenient command acquisition, adapting to the user's need for quick operation and undistracted driving. The collected user selection commands are parsed to determine the adjustment direction corresponding to the selected music adjustment option. Then, music content that matches this adjustment direction and still meets the user's anxiety relief needs is selected from a pre-built user music preference library as the target adjustment music. Subsequently, this target adjustment music is updated as the target content for this push notification (replacing the originally selected target content or supplementing the original target content), ensuring that the updated target content both meets the user's self-selected adjustment needs and does not deviate from the core purpose of anxiety relief. Finally, using the push strategy determined in S12 of this application (such as automatic playback, interface display, and seamless switching), the updated target adjustment music (new target content) is pushed to the user, so that the push method is in line with the real-time music playback status of the vehicle, the user's emotional change characteristics and interaction habits, and completes the closed loop of "anxiety emotion recognition → autonomous adjustment → adaptive push", realizing the deep integration of anxiety emotion regulation and adaptive push.

[0054] In this embodiment, the push operation of the target content is executed according to the push strategy determined in S12, ensuring that the push execution method strictly corresponds to the push strategy determination result, the real-time in-vehicle scenario, and user needs. This completes a full adaptive push closed loop of content filtering → strategy formulation → strategy execution, eliminating the need for users to manually filter, switch, or play music. This significantly simplifies the operation process of in-vehicle music, reduces the risk of driver distraction caused by manual operation of in-vehicle playback devices, and ensures driving safety. At the same time, through standardized push execution, the target content can efficiently reach users, further enhancing the personalization and adaptability of music push, improving the user's music experience during driving, and realizing the intelligent and automated implementation of in-vehicle music push.

[0055] In the above solution, by combining the emotional state of the user inside the vehicle with scene environment data representing the natural environment of the vehicle, the time dimension, and the spatial travel scenario, the target content is determined from the user's music preference library. This allows the pushed music content to simultaneously match the user's real-time emotions and the actual driving scenario, significantly improving the accuracy of matching music content with the user's current needs. Furthermore, based on the real-time music playback status inside the vehicle, the emotional state change characteristics reflecting the dynamic changes and fluctuations of the user's emotions, and the interaction data between the user and the in-vehicle playback device, a corresponding push strategy is determined. This ensures that the push strategy is tailored to the real-time status of in-vehicle playback, the dynamic changes in the user's emotions, and the user's interaction habits, guaranteeing the rationality and adaptability of the push method. Finally, by pushing the target content to the user according to this push strategy, the entire process of in-vehicle music adaptive intelligent push can be achieved without the need for additional manual operation by the user. This simplifies the operation of in-vehicle music use while effectively improving the intelligence level of in-vehicle music push and the overall driving music experience.

[0056] In some embodiments, such as Figure 2 As shown, the adaptive push method for in-vehicle music also includes: S14. Obtain the status perception data of the users inside the vehicle.

[0057] Among them, state perception data is used to characterize the in-vehicle user's emotion-related features and vehicle driving operation features that can reflect the user's emotions. For example, the user's emotion-related feature data includes at least one of the following: user facial feature data (micro-expressions, facial muscle state), voice feature data (pitch, speech rate, tone), and physiological sign data (steering wheel grip strength, heart rate, respiratory rate); the vehicle driving operation feature data that reflects the user's emotions includes at least one of the following: vehicle speed change data, acceleration / deceleration data, braking frequency data, steering wheel rotation amplitude / frequency data, and lane change frequency data.

[0058] Specifically, considering the unique characteristics of the in-vehicle environment, non-invasive, safe, and high-precision in-vehicle devices are selected. Examples include in-vehicle cameras (for facial features), in-vehicle microphones (for voice features), steering wheel grip force sensors (for grip force data), in-vehicle heart rate monitoring devices (for physiological data), and vehicle OBD (On-Board Diagnostics) interfaces (for driving operation data). Each data acquisition device is controlled to collect data synchronously and in real-time, with the collection frequency adapted to the response speed of emotional changes. This avoids missing details of emotional changes due to an excessively low collection frequency, or wasting system resources due to an excessively high frequency. Simultaneously, the device's operating status is monitored in real-time during the collection process to ensure data continuity and prevent data interruptions from affecting subsequent emotion recognition.

[0059] Specifically, the collected raw data undergoes simple preprocessing, including data denoising, outlier removal, and data standardization (such as unifying the magnitude of values ​​collected from different devices) to ensure the usability of the collected data, reduce interference factors in the raw data, and improve the accuracy of subsequent data classification and analysis.

[0060] In this embodiment, by collecting state perception data that characterizes user emotions and reflects driving operation characteristics, a comprehensive, accurate, and real-time data source is provided for subsequent user emotional state recognition. This solves the technical pain point of inaccurate emotion judgment caused by single and one-sided emotion recognition data in related technologies. At the same time, the representation range of state perception data is clearly defined to ensure that the collected data is strongly correlated with "user emotions", avoiding the waste of system resources caused by invalid data collection and improving the efficiency of the emotion recognition process.

[0061] S15. Classify the state-aware data and determine the different types of state-aware sub-data.

[0062] Among them, the state perception sub-data is used to characterize the user's own emotion-related features, as well as vehicle driving behavior related data related to the user's emotions. For example, the user's own emotion-related features may include at least: facial micro-expression data, voice tone / speech rate data, heart rate / grip strength and other physiological signs data; the vehicle driving behavior related data related to the user's emotions may include at least: vehicle speed changes, braking frequency, steering wheel rotation range, lane change frequency and other driving operation data.

[0063] Specifically, each type of state-aware sub-data is labeled with a corresponding type label, which includes "data type + emotion association method" (e.g., "facial micro-expression sub-data - direct emotion association"). The classified sub-data is stored in the vehicle system's temporary database for a duration not exceeding the current emotion recognition cycle (e.g., 5 minutes) to avoid consuming too much storage resources.

[0064] Specifically, the preprocessed state-aware data from S14 is retrieved, and the validity of the data is verified again (data that is still invalid after preprocessing is removed) to ensure that all data to be classified are strongly correlated with emotions, avoiding invalid data from participating in the classification and improving classification efficiency. According to a preset classification algorithm (such as a data label-based classification algorithm), the data to be classified is divided into corresponding sub-data categories one by one. Each sub-data category is clearly labeled with a type label (such as "facial micro-expression sub-data", "grip strength sub-data", "braking frequency sub-data") to facilitate subsequent identification and extraction of corresponding features in S16. During the classification process, data is ensured to be non-repeated and non-omitted, and each piece of state-aware data corresponds to a unique sub-data category. After classification, different types of state-aware sub-data are output and stored in the system's temporary database.

[0065] For example, retrieve the pre-processed state perception data collected in S14 (including: facial micro-expression data of the driver's tense eyes and drooping mouth, voice data of high pitch and fast speech, physiological data of grip strength of 45N, vehicle speed change data of vehicle speed dropping from 60km / h to 30km / h, and braking frequency data of 3 braking actions within 1 minute); perform classification operations according to the classification standard of "emotion association type": divide the facial micro-expression data, voice data, and grip strength data into "state perception sub-data representing the user's own emotional characteristics", and label them "facial micro-expression sub-data", "voice sub-data", and "grip strength sub-data" respectively; divide the vehicle speed change data and braking frequency data into "state perception sub-data representing vehicle driving behavior association data related to the user's emotions", and label them "vehicle speed change sub-data" and "braking frequency sub-data" respectively; after classification, store these two types of sub-data in the system's temporary database.

[0066] In this embodiment, by classifying the collected state-aware data, the messy comprehensive data is decomposed into different types of state-aware sub-data, making the data structure clearer and more targeted, thus solving the technical problem in related technologies that the data is messy and cannot be used for refined sentiment analysis.

[0067] S16. Perform feature extraction and emotion feature analysis on the state perception sub-data respectively to obtain the emotion analysis results corresponding to each type of state perception sub-data.

[0068] Specifically, for different types of sub-data, feature extraction algorithms adapted to their data characteristics are used to extract core features strongly correlated with user emotions (excluding irrelevant features). For example, for "user's own emotion-related sub-data," features that directly reflect emotions are extracted (such as "eye corner curvature and mouth corner angle" from facial micro-expression sub-data, "pitch frequency and speech rate" from speech sub-data, and "grip strength and grip strength change rate" from grip strength sub-data). For "driving behavior-related sub-data," features that indirectly reflect emotions are extracted (such as "vehicle speed change amplitude and frequency" from vehicle speed change sub-data, and "number of braking times and braking force" from braking frequency sub-data). This ensures that the extracted features can directly or indirectly reflect the user's emotional state. For example, for user's own emotion-related sub-data: facial micro-expression sub-data uses facial feature extraction algorithms (such as the HOG algorithm (Histogram of Oriented Gradients)), and speech sub-data uses Mel-Frequency Cepstral Coefficients). The Coefficients (MFCC) algorithm is used for physiological characteristics sub-data, and the numerical feature extraction algorithm is used for driving behavior correlation sub-data. The statistical feature extraction algorithm is used to extract behavioral features per unit time (such as the number of braking times and the magnitude of sudden speed changes per unit time).

[0069] Specifically, the extracted emotional association features of each sub-data category are compared with a pre-set emotional feature database (such as "features corresponding to anxiety: tight eyes, higher pitch, increased grip strength, and increased braking frequency") to calculate the feature matching degree, thereby obtaining the emotional analysis results corresponding to each sub-data category. The pre-set emotional feature database can contain feature thresholds corresponding to at least three common emotions in in-vehicle scenarios (anxiety, calmness, and pleasure). The emotional analysis results must clearly include "possible emotion type and emotion confidence level" (such as "anxiety, confidence level 86%" for facial micro-expression sub-data analysis and "anxiety, confidence level 82%" for braking frequency sub-data analysis) to ensure that the analysis results are quantifiable and verifiable. The emotional analysis results corresponding to each sub-data category are summarized to form a multi-dimensional emotional analysis dataset (each sub-data category corresponds to an independent emotional analysis result).

[0070] For example, regarding user-related emotion-related sub-data: Facial micro-expression sub-data (tight corners of the eyes, drooping corners of the mouth): the HOG algorithm was used to extract the features of "corner of the eye curvature (-5°) and corner of the mouth angle (-3°)," which were compared with a preset emotion feature database (anxiety corresponds to corner of the eye curvature ≤-3° and corner of the mouth angle ≤-2°). The feature matching degree was 86%, and the analysis result was "anxiety, confidence level 86%"; Speech sub-data (high pitch, fast speech rate): the MFCC algorithm was used to extract the features of "pitch frequency 500Hz and speech rate 180 words / minute," which were compared with anxiety emotion features (pitch frequency ≥450Hz and speech rate ≥160 words / minute). The feature matching degree was 84%, and the analysis result was "anxiety, confidence level 84%"; Grip strength sub-data (grip strength 45N, grip strength change rate 10N / second): grip strength magnitude and change rate features were extracted and compared with anxiety emotion features (grip strength ≥4... For the following sub-data: For the vehicle speed change sub-data (speed suddenly dropping from 60km / h to 30km / h, a sudden change of 30km / h): the sudden change rate feature was extracted and compared with the anxiety feature (sudden change rate ≥25km / h), with a feature matching degree of 82%, resulting in the analysis result "anxiety, confidence level 82%"; For the braking frequency sub-data (3 braking actions within 1 minute): the number of braking actions per unit time was extracted and compared with the anxiety feature (≥2 times / minute), with a feature matching degree of 80%, resulting in the analysis result "anxiety, confidence level 80%"; Multi-dimensional analysis results were output: the analysis results of all sub-data were summarized to form a dataset (all anxiety, confidence level 80%-88%), which was output to S17 for subsequent emotion result fusion.

[0071] In this embodiment, targeted feature extraction and emotional feature analysis are performed on the different types of state perception sub-data after S15 classification, so as to fully explore the emotional correlation features contained in each type of sub-data. This solves the technical problem in related technologies that use a unified analysis method for different types of data, resulting in insufficient emotional feature mining and inaccurate analysis results. At the same time, through categorized analysis, multi-dimensional emotional analysis results can be obtained.

[0072] S17. The emotion analysis results are fused to obtain the emotional state of the users in the car.

[0073] Specifically, from the output of S16, the sentiment analysis results corresponding to all types of state perception sub-data are retrieved in real time, clarifying the "emotion type and sentiment confidence level" of each analysis result, providing a basis for subsequent screening and fusion; according to the preset confidence level threshold, the effective sentiment analysis results are screened out, and low-confidence abnormal results with confidence level <60% are removed (to avoid interfering with the final judgment), and the effective results with confidence level ≥60% are retained; if all analysis results are low-confidence, the process returns to S14 to re-collect data to ensure the reliability of the final emotional state.

[0074] Specifically, a preset fusion strategy is adopted to fuse the selected valid sentiment analysis results. The core logic is that "high confidence results have high weight and low confidence results have low weight". At the same time, the type priority of the sub-data is combined (the weight of sub-data related to the user's own emotions is higher than that of sub-data related to driving behavior) to calculate the comprehensive sentiment confidence. If the sentiment type corresponding to all valid analysis results is consistent, the sentiment type is directly confirmed. If there are multiple sentiment types, the sentiment type with the highest comprehensive confidence is selected as the core sentiment type after fusion.

[0075] Specifically, based on the integrated emotional confidence score, a final judgment is made. When the integrated emotional confidence score is ≥80%, the corresponding emotional state (such as "anxious" or "calm") is directly output. When the integrated confidence score is between 60% and 79%, "suspected + corresponding emotional type" (such as "suspected anxiety") is output, and data is re-collected and verified.

[0076] In this embodiment, by fusing the multi-dimensional emotion analysis results obtained in S16, the emotional information reflected by different types of sub-data is integrated, solving the technical problem in related technologies where single-dimensional emotion analysis results have biases, leading to inaccurate judgment of emotional state. At the same time, through a reasonable fusion strategy, the interference of low-confidence analysis results is weakened, the weight of high-confidence analysis results is strengthened, and the accuracy, robustness and reliability of the final emotion state recognition are improved, ensuring that the output emotion state can truly reflect the user's current mood.

[0077] In the above solution, S14 collects state perception data representing user emotion-related features and emotion-related driving operation features, providing a comprehensive and reliable data source for emotion recognition; S15 classifies and breaks down the data into different types of sub-data to solve the problem of data clutter and inability to perform fine-grained analysis; S16 performs feature extraction and emotion analysis on each sub-data separately to mine multi-dimensional emotion information and avoid misjudgment based on a single data dimension; S17 integrates the results of multi-dimensional emotion analysis to weaken the interference of low-confidence results, improve the accuracy and robustness of emotion state recognition, and comprehensively solve the pain points of existing emotion recognition data being single and inaccurate in judgment, providing an emotional basis for subsequent target content matching and push strategy formulation, laying the core foundation for adaptive push of in-vehicle music.

[0078] In some embodiments, such as Figure 3 As shown, based on emotional state and scene environment data, target content matching the emotional state is determined from a pre-built user music preference library, including: S111. Determine the first match between the emotional state and multiple items in the user's music preference library.

[0079] Specifically, a pre-built user music preference library is retrieved, and multiple pieces of music (not a single piece) are selected as comparison objects. Each selected piece carries a pre-defined emotional matching feature tag (corresponding to the user's emotional state). Emotional feature tags corresponding to the user's emotional state are extracted (e.g., anxiety corresponds to the feature tags "soothing, calm, no strong rhythm"), clarifying the core requirements for emotional matching. Simultaneously, emotional matching feature tags are extracted for each piece of music to be compared in the user music preference library (tags are based on the music's style, rhythm, genre, and the user's historical preferences; for example, a piece of light music might be labeled "soothing, calm, anxiety-matching"). Then, a pre-defined quantitative matching algorithm is used to compare the emotional feature tags of the emotional state with the emotional matching feature tags of each piece of music, calculating the overlap and fit between the two. Finally, the first matching degree for each piece of music is output. The first matching degree uses a numerical range of 0-1 (or 0%-100%); the higher the value, the higher the degree of matching between the music content and the user's current emotional state.

[0080] For example, the user's emotional state (anxiety) output by S17 is retrieved, and its core emotional feature tags are identified as "soothing, calm, and no lyrics." Simultaneously, 30 pieces of music (all frequently listened-to light music and instrumental music) are randomly selected from the user's music preference library as comparison objects. Dual features are extracted: ① Anxiety emotion feature tags: soothing, calm, and no lyrics; ② Emotional fit feature tags for each piece of music (3 examples): Music A (soothing, calm, and no lyrics), Music B (upbeat, with lyrics, and pleasant fit), and Music C (soothing, with lyrics, and calm fit). Perform quantitative matching calculation: A feature overlap matching algorithm is used to calculate the first match degree between each music content and the anxiety emotion: ① Music A: All 3 feature labels completely overlap, first match degree = 3 ÷ 3 × 100% = 100% (1.0); ② Music B: 0 feature labels overlap, first match degree = 0 ÷ 3 × 100% = 0% (0.0); ③ Music C: 2 feature labels overlap (soothing, calm), first match degree = 2 ÷ 3 × 100% ≈ 66.7% (0.67); The remaining 27 music contents are calculated using the same logic to obtain their respective first match degrees (range 0-100%). Output results: A first match degree list is generated, labeling the ID of each music content and its corresponding match degree (e.g., Music A: 1.0, Music C: 0.67, Music B: 0.0…), and output to S112 for subsequent candidate content filtering.

[0081] In this embodiment, by quantifying the first matching degree between the user's emotional state in the vehicle and multiple music contents in the user's music preference library, the abstract "emotional adaptation" is transformed into a quantifiable and comparable numerical indicator, which solves the technical pain points of related technologies, such as the lack of clear judgment criteria for emotion and music adaptation, the subjectivity of the adaptation process, and low accuracy.

[0082] S112. Based on the first matching degree, determine multiple candidate contents.

[0083] Specifically, a first matching degree screening threshold (quantification standard) is preset. The screening threshold can be dynamically adjusted based on user interaction habits. The core screening logic is to "retain music content with a first matching degree ≥ the screening threshold and remove content with a first matching degree < the screening threshold". At the same time, the number of candidate content is limited (preset minimum and maximum values) to avoid too many candidate contents being selected (increasing the amount of subsequent calculations) or too few (failing to meet the selection requirements of scene matching).

[0084] Specifically, according to the preset screening criteria, all music content in the first matching degree list is screened one by one, and music content that meets the conditions is retained; if the number of screened content exceeds the preset maximum value, it is sorted from high to low according to the first matching degree, and the top N items are selected (N is the preset maximum value); if the number of screened content is lower than the preset minimum value, the screening threshold is appropriately lowered (not lower than the minimum threshold, such as 0.5), and the screening is repeated so that the number of candidate content meets the requirements of subsequent scene matching; the multiple candidate contents that have been screened are sorted from high to low according to the first matching degree (for the priority reference of subsequent scene matching), and the ID, first matching degree and emotion adaptation feature label of each candidate content are marked.

[0085] In this embodiment, based on the first matching degree calculated by S111, multiple music contents in the user's music preference library are filtered to select multiple candidate contents with a high degree of matching with the user's real-time emotional state. This solves the technical problems in related technologies, such as the lack of clear standards for emotion matching filtering, messy filtering results, and too much invalid content.

[0086] S113. Determine the second degree of matching between the scene environment data and multiple candidate contents.

[0087] Specifically, extract scene feature tags corresponding to scene environment data (e.g., the natural environment is rainy, the time dimension is night, and the spatial travel scenario is commuting, corresponding to the tags "rainy day, night, commuting") to clarify the core requirements of scene adaptation; extract scene adaptation feature tags for each candidate content (the tags are based on music style, rhythm and in-vehicle scene adaptability, such as a light music track labeled "rainy day, night, commuting adaptable") to ensure that the feature tags of both parties are comparable and that the tags do not conflict with the emotion adaptation tags.

[0088] Specifically, a quantitative matching algorithm adapted to scene features is used to compare the scene feature tags of the scene environment data with the scene adaptation feature tags of each candidate content, calculate the overlap and fit between the two, and finally output the second matching degree corresponding to each candidate content. The second matching degree also uses a numerical range of 0-1 (or 0%-100%). The higher the value, the higher the degree of adaptation of the candidate content to the current driving scenario, ensuring consistency with the quantitative standard of the first matching degree, facilitating subsequent comprehensive reference. The calculated second matching degree corresponding to each candidate content is then compared with the ID of the candidate content. The candidate content identifier, first matching degree, and emotion / scene tag are stored to form a second matching degree list. Among them, the scene feature tags of the scene environment data are extracted based on three types of scene data, with one core tag for each type of scene (e.g., natural environment: rainy / sunny, time dimension: night / day, spatial travel scene: commuting / long distance), for a total of three core scene tags. The scene adaptation feature tags of each candidate content include at least one scene type tag + one style tag for the adapted scene (e.g., "rainy day, soothing adaptation"). The tags are based on the characteristics of the music itself and the usage scenario of the in-vehicle scene and can be dynamically updated.

[0089] In this embodiment, based on the high-emotion-adaptation candidate content selected in S112, the second matching degree between the scene environment data and each candidate content is further quantified and calculated, and "scene adaptation" is also transformed into a quantifiable and comparable numerical indicator. This solves the technical pain point in related technologies that only considers emotion adaptation and ignores driving scene adaptation, resulting in the push content being out of touch with the actual driving scene.

[0090] S114. Based on the second matching degree, determine the target content.

[0091] Specifically, a second matching degree screening threshold (quantitative standard) is preset. The core screening logic is to retain candidate content with a second matching degree ≥ the screening threshold and remove content with a second matching degree < the screening threshold. At the same time, if multiple candidate contents meet the second matching degree threshold requirement, the content with the highest second matching degree and a first matching degree not lower than the preset emotion adaptation threshold is selected as the final target content, based on the first matching degree. If only one candidate content meets the second matching degree threshold, it is directly determined as the target content. If no content meets the threshold, the second matching degree threshold is appropriately lowered (not lower than the minimum threshold, such as 0.5), and the screening is repeated.

[0092] Specifically, according to the preset final screening criteria, all candidate content in the second matching degree list is screened one by one, prioritizing content with high scene adaptability while also considering emotional adaptability; if multiple content that meets the criteria is selected, it is sorted from high to low according to the second matching degree, and the first content is selected as the final target content (if the user prefers multiple content, the first two can be selected to ensure the flexibility of the push strategy); if the selected content still does not meet the requirements, the threshold is lowered according to the rules and the selection is re-screened to ensure that the final effective target content can be determined.

[0093] For example, retrieve the data: retrieve the second matching degree list (8 candidate contents) output by S113, and clarify the first and second matching degree of each content (e.g., ID1 music A: first 1.0, second 1.0; ID2 music D: first 0.95, second 0.67; ID4 music F: first 0.85, second 0.8; ID5 music G: first 0.8, second 0.58, etc.). The preset second matching degree filtering threshold is 0.6 (60%), and the emotional matching baseline is first matching degree ≥ 0.6; the filtering priority is: second matching degree > first matching degree; the default number of target content is 1; the 8 candidate contents are filtered, and the contents with second matching degree < 0.6 are removed (such as ID5 music G, second matching degree 0.58), leaving 7 contents that meet the conditions; sorted by second matching degree from high to low, the first one is ID1 music A (second matching degree 1.0, first matching degree 1.0, both meet the requirements), so ID1 music A is selected as the final target content; the target content (ID1 music A) is output, labeled with its first matching degree 1.0, second matching degree 1.0, emotional label "soothing, stable, no lyrics", and scene label "rainy day, night, commuting".

[0094] In one specific embodiment, the candidate content corresponding to the second matching degree is sorted according to the priority of the target matching degree to generate a target content list, which includes multiple candidate content; based on the push strategy, the candidate content is pushed to the user in the order of the target content list.

[0095] Specifically, the target matching degree is a comprehensive quantitative indicator that integrates the first matching degree (emotional adaptation) and the second matching degree (scene adaptation). It is calculated using a weighted fusion method, with weight allocation prioritizing the needs of in-vehicle scenarios (scene adaptation weight is no less than emotion adaptation weight) to ensure that the target matching degree truly reflects the comprehensive adaptability of the content. The target matching degree also uses a numerical range of 0-1 (or 0%-100%), with higher values ​​indicating stronger comprehensive adaptability. Based on the calculated target matching degree, all candidate content is prioritized and sorted from highest to lowest target matching degree to generate a target content list. The list clearly labels each candidate content with its ID, target matching degree, first matching degree, second matching degree, and dual adaptation tag, making the list information clear and easy for users to understand and select. At the same time, the number of candidate content in the list is limited to avoid the list becoming too long and interfering with user operation.

[0096] Specifically, based on the push strategy determined by S12, the push operation of candidate content is executed according to the priority order of the target content list; different push strategies correspond to different order push logic: autoplay strategy → directly play the first and first content of the list; interface display strategy → display all content in the list order for users to actively select; seamless switching strategy → preload the first two content in the list order for easy switching; during the push process, the list order is synchronized in real time to ensure that the push order is consistent with the overall adaptability priority.

[0097] In this embodiment, based on the second matching degree calculated in S113, the selected high-emotion-adaptation candidate content is further filtered to finally determine the target content that is highly adapted to the user's emotions and driving scenarios. This solves the technical pain point in related technologies that only consider emotions or scenarios, resulting in low overall adaptability of the pushed content. At the same time, the quantitative indicators of scenario adaptation are used as the final selection criteria to ensure that the target content not only matches the user's real-time mood, but also perfectly adapts to the natural environment, time dimension and spatial travel scenario of in-vehicle driving, achieving dual adaptation of emotion and scenario.

[0098] The above solution focuses on dual adaptation of "emotion + scenario". S111 quantifies and calculates the first matching degree between emotional state and content in the preference library. S112 filters candidate content with high emotional matching based on this matching degree. S113 further quantifies and calculates the second matching degree between scenario environment data and candidate content. S114 combines the second matching degree to determine the final target content. This transforms abstract emotion and scenario adaptation into quantifiable indicators, solving the problems of single-dimensional adaptation, inaccurate content matching, and disconnect from the in-vehicle scenario in related technologies. It ensures that the target content not only matches the user's real-time mood but also perfectly adapts to the driving scenario, providing high-quality core content support for the subsequent push strategy formulation and execution, and improving the accuracy of in-vehicle music push.

[0099] In some embodiments, such as Figure 4As shown, based on the music playback status in the car, changes in emotional state, and user interaction data with the in-car playback device, a push strategy for targeted content is determined, including: S121. Based on the characteristics of music playback status, emotional state changes, and interaction data, determine the current playback scenario in the vehicle.

[0100] Among them, the playback scenario is used to characterize the scenarios related to music push strategy formulation in the in-vehicle environment, as well as the music push intervention adaptation requirements corresponding to the changing characteristics.

[0101] Specifically, a comprehensive correlation analysis is conducted on the preprocessed music playback status, change characteristics, and interaction data to uncover the correspondence between "data combinations and push intervention needs." That is, data characteristics are used to determine whether push intervention is needed and what level of intervention is required (e.g., high mood fluctuation + music not playing + no interaction → strong push intervention needed; stable mood + music playing + no interaction → no active intervention needed). At the same time, the core needs corresponding to each data combination are identified and transformed into "push intervention adaptation needs" for the playback scenario, aligning with the representation scope of the playback scenario in the supplementary explanation.

[0102] Specifically, two core representation dimensions of playback scenarios are clearly defined: ① Scenarios related to music recommendation strategy formulation in the in-vehicle environment (such as "not playing + high fluctuation of negative emotions" and "playing + stable emotions"); ② Music recommendation intervention adaptation requirements corresponding to the scenario (such as "requires active automatic playback intervention" and "no active intervention required, only display alternatives"). At the same time, clear classification criteria are set to ensure that the determination of playback scenarios is unique and repeatable, and to avoid scenario confusion (such as using "music playback status + emotional fluctuation level" as the core classification dimension, supplemented by interaction data).

[0103] Specifically, based on the data correlation analysis results and the preset playback scenario classification standards, the corresponding playback scenario in the car is determined, and the core characteristics of the scenario (such as "music not playing, high emotional fluctuation and anxiety, no interactive commands") and the corresponding push intervention adaptation requirements (such as "music to relieve anxiety needs to be actively pushed without user manual confirmation") are clearly marked.

[0104] Specifically, if there are contradictions between the music playback status, change characteristics, and interaction data (such as high emotional fluctuations but the user issues a "reject push" command), the user interaction data will be given priority and the scenario will be determined as "command priority adaptation"; if the data collection is incomplete (such as failure to collect emotional features), the current playback scenario will be inferred based on the remaining two types of data and historical scenario determination records, and the data will be re-collected to ensure that the scenario determination is not interrupted.

[0105] For example, consider the following: music playback status (not playing, no current playback progress), emotional state change characteristics (dynamically changing attributes: calm → anxious, fluctuation level: high fluctuation 75%, state switching characteristic: single switch), and user interaction data (no confirmation, rejection, or switching commands within 3 seconds, considered no interaction). The data is preprocessed to ensure no abnormal data is found, and can be used for scenario determination. Based on weighted allocation (emotion 40%, playback status 35%, interaction data 25%), the correlation between the three types of data is analyzed: high emotional fluctuation (negative anxiety) → strong need for intervention push notifications; music not playing → potential for proactive push notifications. Environmental conditions; no interactive commands → user shows no explicit intention to refuse, proactive push can be executed; comprehensive analysis shows that the current adaptation requirement is "proactive strong intervention push"; based on the preset playback scenario classification standard and combined with the data correlation analysis results, the current in-car playback scenario is determined to be "not playing + high emotional fluctuation (negative) + no interaction: proactive strong intervention push scenario", clarifying that the core characteristics of the current in-car playback scenario are "music not playing, user anxiety and high emotional fluctuation, no interactive commands", and the corresponding push intervention adaptation requirement is "to proactively push music to relieve anxiety, without requiring manual confirmation from the user, and quickly intervene in the user's emotions".

[0106] In this embodiment, by integrating three core real-time data types—music playback status, emotional state change characteristics, and user interaction data—the current playback scenario in the vehicle is defined. This transforms abstract in-vehicle environment data into a concrete scenario representation that is "strongly correlated with the push strategy," thus solving the technical pain points of existing technologies where push strategy formulation lacks clear scenario basis, blindly adapts, and is disconnected from the actual in-vehicle environment.

[0107] S122. Determine the push strategy based on the playback scenario.

[0108] The push strategy includes at least one of the following: automatic playback push strategy, push strategy for displaying target content on the in-vehicle interactive interface, and push strategy for seamlessly switching the music being played.

[0109] Specifically, based on the push intervention adaptation requirements of playback scenarios, a "one-to-one, many-to-one" matching rule is set, meaning that one type of playback scenario corresponds to at least one push strategy, and one push strategy can adapt to multiple playback scenarios with the same intervention requirements. The matching rule prioritizes the correspondence between intervention requirements and strategy execution methods (e.g., strong intervention requirements → automatic playback strategy, lightweight requirements → interface display strategy, switching requirements → seamless switching strategy), while also taking into account the safety and convenience of in-vehicle scenarios. According to the matching rule, combined with the core characteristics of the current playback scenario and the intervention adaptation requirements, the most suitable push strategy is selected from three core push strategies. If the current scenario can adapt to multiple strategies (e.g., a lightweight display scenario can adapt to interface display and simple automatic playback), then the final push strategy is determined by combining the user's historical interaction habits (e.g., if the user prefers manual selection, then the interface display strategy is adapted first). At the same time, the specific execution rules of the final push strategy are clarified (e.g., the initial volume of automatic playback, the number of items in the list displayed on the interface) to ensure that the strategy can be executed directly.

[0110] In this embodiment, the playback scenario determined in S121 is used as the core basis. Combined with the three types of push strategies defined in the supplementary explanation, a push strategy adapted to the current in-vehicle scenario is matched. This solves the technical pain points of related technologies, such as the push strategy being singular and fixed, disconnected from the actual in-vehicle scenario and user needs, and unreasonable intervention methods. At the same time, the scope of push strategy types is clearly defined, ensuring that the strategy formulation has clear guidance and feasibility, and realizing a one-to-one correspondence between "scenario and strategy". This allows the push strategy to not only fit the current in-vehicle environment (such as adapting to automatic playback in a non-playing scenario) but also meet the corresponding push intervention adaptation requirements (such as adapting to active push in a strong intervention scenario). This further improves the flexibility, targeting, and safety of the push strategy, reduces manual operation by the driver, and lowers the risk of driver distraction.

[0111] In the above solution, S121 integrates three core data categories—music playback status, emotional state change characteristics, and user interaction data—to define in-vehicle playback scenarios that include push notification intervention adaptation requirements, addressing the pain point of push notification strategy formulation lacking clear scenario basis and blind adaptation. Based on the in-vehicle playback scenario, S122 matches the appropriate push notification strategy from three categories: automatic playback, interface display, and seamless switching, achieving a one-to-one correspondence between "scenario and strategy." This avoids a single, rigid push notification strategy, ensuring that the strategy aligns with the real-time status of the vehicle and the potential needs of users, reducing invalid and interfering push notifications, lowering the risk of driver distraction during manual operation, providing a clear and reasonable operational basis for subsequent push notification execution, and improving the targeting and safety of the push notification strategy.

[0112] In some embodiments, such as Figure 5 As shown, when multiple users are in the vehicle, the adaptive push method for in-vehicle music also includes: S18. Based on scene environment data and the emotional state of each user, determine the target content corresponding to each user from the pre-built user music preference library.

[0113] Specifically, through in-vehicle cameras, seat pressure sensors, and other in-vehicle devices, the system identifies the number of users currently seated in the vehicle and their seating positions (e.g., front left, front right, rear left, rear right), clearly defining "each user" as all currently seated users identified by the system (including the driver and passengers). Simultaneously, a unique identifier is assigned to each user (e.g., user 1 - driver, user 2 - front right passenger) to ensure the user matching process is unambiguous and traceable. The system retrieves overall in-vehicle scene environment data (natural environment, time dimension, spatial travel scenario), which is shared by all users to ensure that each user's target content aligns with the overall in-vehicle driving scenario. It also retrieves each user's independent emotional state (obtained separately through the emotion recognition processes S14-S17, without interference) and a pre-built user music preference library containing each user's corresponding preference sub-library (stored in partitions by user identifier, including each user's historical playback, favorites, preference tags, etc.).

[0114] Specifically, for each user, a dual matching process of "emotion + scenario" is executed separately (using the matching logic of S111-S114). This involves calculating the first matching degree (emotional fit) between the user's emotional state and multiple contents in their preference sub-library; filtering candidate content based on the first matching degree; calculating the second matching degree (scenario fit) between scenario environment data and the user's candidate content; and determining the unique target content (or target content list) for the user based on the second matching degree. During the matching process, the matching logic and thresholds for each user do not interfere with each other, ensuring that the target content matches the individual needs of each user. The target content for each user is associated with and stored with the user's identifier and location, forming a "user identifier - location - target content" correspondence table, which is synchronously output to S19 as the core basis for zoned channel matching and push notifications, achieving seamless integration with subsequent steps.

[0115] The user music preference library is stored in partitions according to user identifiers, with each user corresponding to an independent preference sub-library. The preference sub-library can be dynamically updated based on the user's real-time interaction data (such as adding favorites and skipping push notifications). When a new user is detected (without a corresponding preference sub-library), a temporary preference sub-library is automatically generated (based on the initialization of in-vehicle general preferences), and the user's preferences will be gradually adapted thereafter.

[0116] For example, by detecting pressure sensors in the car seats, if the pressure on both the left and right front seats is ≥50N, and the camera captures human features, it can be determined that there are two users. Unique identifiers are assigned to User 1 (driver, left front seat) and User 2 (passenger, right front seat). Dual data is retrieved: natural environment (rainy day), time dimension (nighttime 20:30), and spatial travel scenario (urban commuting); User 1's emotional state (anxious, high volatility), corresponding to a preference sub-database (primarily soothing, wordless light music); User 2's emotional state (calm, low volatility), corresponding to a preference sub-database (primarily upbeat folk and pop songs). User 1 (Driver): Using the dual matching logic, target content with an emotion matching degree ≥ 0.7 and a scene matching degree ≥ 0.6 is selected (soothing piano instrumental music, first matching degree 1.0, second matching degree 1.0); User 2 (Passenger): The same dual matching is performed to select target content with an emotion matching degree ≥ 0.7 and a scene matching degree ≥ 0.6 (upbeat folk music, first matching degree 0.85, second matching degree 0.75); The matching processes of the two users do not interfere with each other, ensuring that the content matches their respective preferences and emotions; a corresponding relationship table is formed: User 1 (front row left) → soothing piano instrumental music; User 2 (front row right) → upbeat folk music, output to S19 for subsequent partition channel matching.

[0117] This embodiment breaks through the limitations of existing in-vehicle music push "single-user adaptation". For multiple users in the car, it combines unified scene environment data (overall in-vehicle scene) with each user's independent emotional state to match their respective target content from the user music preference library. This solves the technical pain points of "one-size-fits-all" music push when multiple users are in the same car, which cannot take into account the differences in each user's emotions and preferences. At the same time, it ensures that the target content for each user meets the triple requirements of "scene adaptation + emotion adaptation + preference adaptation". It not only fits the overall in-vehicle driving scene, but also matches each user's real-time mood and long-term music preferences. This greatly improves the personalization, targeting and practicality of music push in multi-user in-vehicle scenarios, improves the multi-scene coverage capability of in-vehicle music adaptive push, and lays the core content foundation for subsequent independent push in different areas.

[0118] S19. Based on the vehicle-mounted zone audio system, each user's seating position is matched with a corresponding independent audio playback channel.

[0119] Specifically, for each user's seating position, a corresponding independent audio channel is matched, following the rule of "one-to-one correspondence between position and channel": front left seating position → channel 1 (driver's seat speaker group); front right seating position → channel 2 (passenger seat speaker group); rear left seating position → channel 3 (rear left speaker group); and so on. Simultaneously, the "seating position - channel identifier" is merged with the correspondence table in S18 to form a complete correspondence of "user identifier - seating position - channel identifier - target content". For each matched independent audio channel, real-time availability testing is performed, including whether the channel is powered on normally, whether the speaker is working properly, and whether the volume is within the adjustable range (e.g., 5%-100%). If a channel abnormality is detected (e.g., speaker malfunction), the user's target content is temporarily matched to an adjacent backup channel (e.g., if the front left channel is abnormal, it is temporarily matched to the rear left channel), and a channel failure prompt (brief voice message) is issued to ensure the push process is not interrupted.

[0120] In this embodiment, an independent audio playback channel is matched to each user's seating position through an in-vehicle zone audio system, which solves the technical pain point of music playback interference between multiple users in the same vehicle and the inability to achieve independent listening. At the same time, a unique correspondence between "seating position and independent audio channel" is established to ensure that each user's target content can be pushed to that user's location without affecting other users, which greatly improves the listening experience and convenience in multi-user in-vehicle scenarios.

[0121] S20. Push the target content for each user to the audio playback channel of the corresponding user's seat.

[0122] Specifically, for each channel identifier, the target content corresponding to that channel is preloaded separately. During the preloading process, the channels do not interfere with each other. If the target content is not fully loaded, the push execution for that channel is delayed, and the user is prompted with "Content loading". For each channel, the push operation is executed independently according to the push strategy for the user corresponding to that channel. If the push strategy is automatic playback: the speaker of that channel is controlled to automatically play the target content at a preset initial volume (40%-60%) without the user's manual confirmation. If the push strategy is interface display: the target content is displayed on the interactive interface that the user can observe (such as the instrument panel for the driver and the split screen of the central control screen for passengers) for the user to actively select to play. If the push strategy is seamless switching: the target content is seamlessly switched during the interval between the currently playing content in that channel. During the push process, the playback status and volume of each channel do not interfere with each other and can be controlled independently.

[0123] Specifically, during the push process, the push status of each channel is monitored in real time (such as whether it plays normally, whether the volume is normal, and whether the user has interactive commands). If a push anomaly is detected (such as content stuttering or channel failure), the target content is reloaded or the system switches to a backup channel. At the same time, push feedback data for each user is collected (such as skipping, pausing, and volume adjustment) to optimize the target content matching and push strategy for that user in the future.

[0124] In one specific embodiment, if a conflict is detected between the push strategies of different users, the target content or push strategy of each user is adjusted based on the push priority rules and the interaction data between the user and the in-vehicle playback device.

[0125] Specifically, before the S20 push is executed, the push strategies of all users are uniformly tested to determine whether there are any conflicts. The conflict definition standard is "the push strategies of different users cause the in-vehicle zone audio system to be unable to execute simultaneously (e.g., two users both request automatic playback, and the channels are adjacent, resulting in excessive audio interference; or multiple users simultaneously request seamless switching, resulting in insufficient system resources)". Among them, conflicts can be divided into two categories for targeted adjustments: ① Resource conflict: The push strategies of multiple users simultaneously occupy the core resources of the in-vehicle system (such as CPU, audio channel bandwidth), resulting in the inability to execute simultaneously (e.g., more than 3 users request automatic playback, resulting in insufficient system resources); ② Interference conflict: The push strategies of different users cause audio to interfere with each other, affecting the independent listening experience or driving safety (e.g., both the driver's channel and the passenger channel request high-volume automatic playback, resulting in excessive interference).

[0126] Specifically, the push priority rules (core basis) are: a clearly defined priority order is preset, prioritizing driving safety and core user needs; user interaction data (auxiliary basis) is: recent interaction data for each user is retrieved (e.g., whether they have rejected push notifications, frequently switched content, or have a clear volume preference) to determine the user's subjective intent, optimize the adjustment plan, and avoid adjustments that violate user needs. Based on conflict type, priority rules, and interaction data, the principle of "prioritizing adjustments for low-priority users" is adopted, with two adjustment methods (either one or a combination): ① Adjust push strategy: adjust the push strategy for low-priority users to a type that does not consume resources or cause interference (e.g., autoplay → interface display, high volume → low volume); ② Adjust target content: adjust the target content for low-priority users to content with lower volume and a smoother rhythm to avoid interfering with high-priority users; during the adjustment process, the push strategy and target content for high-priority users remain unchanged. After the adjustment is completed, re-test the push strategy for all users to confirm that there are no new conflicts; at the same time, send a brief prompt to the users whose strategies were adjusted (such as "Due to push conflict, your interface display mode has been adjusted") to ensure that the users are aware of the changes; if the verification is successful, execute the push operation of S20; if the verification fails, readjust until there are no conflicts.

[0127] In this embodiment, based on the "user-location-channel" correspondence established in S19, the target content of each user is pushed to the independent audio channel of its corresponding location, realizing "independent listening without interference" when multiple users are in the same car. This solves the technical pain point that existing in-vehicle music push cannot take into account the needs of multiple users and that audio interferes with each other. At the same time, the push process follows the push strategy logic of this application, ensuring that the target content push method of each user is tailored to its own needs and in-vehicle scenario. This not only ensures the driving safety of the driver (such as without manual operation) but also improves the listening experience of ordinary passengers, realizing the intelligent, personalized and convenient music push in multi-user in-vehicle scenarios.

[0128] The above solution breaks through the limitation of the existing "single-user adaptation" of in-vehicle music push. It targets multiple users in the car, combines unified scene environment data with each user's independent emotional state, determines target content from the preference library, and pushes independent audio playback channels to each user's seat through the in-vehicle zone audio system. This enables multiple users to "listen independently without interference" when in the same car. It ensures that the target content for each user matches their own emotions, preferences and the overall in-vehicle scene, while also taking into account the driver's driving safety and the listening experience of ordinary passengers. This improves the push loop of multi-user adaptation, expands the scenario coverage of the solution, and enhances the personalization and practicality of music push in multi-user in-vehicle scenarios.

[0129] In addition, such as Figure 6 As shown, Figure 6This is a schematic diagram of the structure of an adaptive in-vehicle music push device 600 provided in an embodiment of this application. The adaptive in-vehicle music push device 600 includes: The first determining module 601 is used to determine target content that matches the emotional state and scene environment data of the user in the vehicle from a pre-built user music preference library. The scene environment data is used to characterize the natural environment, time dimension and spatial travel scenario of the user during the vehicle driving process. The second determining module 602 is used to determine the push strategy for target content based on the music playback status in the vehicle, the change characteristics of the emotional state, and the interaction data between the user and the in-vehicle playback device. The change characteristics are used to indicate the dynamic change attributes, fluctuation level, and state switching characteristics of the user's emotional state over time. The push module 603 is used to push target content to users based on push strategies.

[0130] In the above solution, by combining the emotional state of the user inside the vehicle with scene environment data representing the natural environment of the vehicle, the time dimension, and the spatial travel scenario, the target content is determined from the user's music preference library. This allows the pushed music content to simultaneously match the user's real-time emotions and the actual driving scenario, significantly improving the accuracy of matching music content with the user's current needs. Furthermore, based on the real-time music playback status inside the vehicle, the emotional state change characteristics reflecting the dynamic changes and fluctuations of the user's emotions, and the interaction data between the user and the in-vehicle playback device, a corresponding push strategy is determined. This ensures that the push strategy is tailored to the real-time status of in-vehicle playback, the dynamic changes in the user's emotions, and the user's interaction habits, guaranteeing the rationality and adaptability of the push method. Finally, the target content is pushed to the user according to the push strategy, achieving fully adaptive and intelligent push of in-vehicle music without requiring additional manual intervention from the user. This simplifies the operation of in-vehicle music while effectively improving the intelligence level of in-vehicle music push and the overall driving music experience.

[0131] In one specific embodiment, the in-vehicle music adaptive push device 600 further includes an acquisition module: The acquisition module acquires the state perception data of the users inside the vehicle. The state perception data is used to characterize the emotion-related features of the users inside the vehicle, as well as the vehicle driving operation features that can reflect the user's emotions. The state-aware data is classified to identify different types of state-aware sub-data. The state-aware sub-data is used to characterize the user's own emotion-related features, as well as vehicle driving behavior related data that are related to the user's emotions. Feature extraction and emotion feature analysis were performed on the state-aware sub-data to obtain the emotion analysis results corresponding to each type of state-aware sub-data. The emotional analysis results are fused and processed to obtain the emotional state of the users inside the vehicle.

[0132] In one specific embodiment, the first determining module 601 is further configured to: Determine the initial match between emotional state and multiple items in the user's music preference library; Based on the first matching degree, multiple candidate contents are identified; Determine the second degree of matching between the scene environment data and multiple candidate contents; Based on the second matching degree, the target content is determined.

[0133] In one specific embodiment, the in-vehicle music adaptive push device further includes a first generation module: The first generation module is used to sort the candidate content corresponding to the second matching degree according to the priority of the target matching degree and generate a target content list, which includes multiple candidate contents. Based on the push strategy, candidate content is pushed to users in the order of the target content list.

[0134] In one specific embodiment, the second determining module 602 is further configured to: Based on the characteristics of music playback status, emotional state changes, and interaction data, the current playback scenario in the car is determined. The playback scenario is used to characterize the scenarios related to music push strategy formulation in the in-vehicle environment, as well as the music push intervention adaptation requirements corresponding to the changing characteristics. Based on the playback scenario, a push strategy is determined. The push strategy includes at least one of the following: an automatic playback push strategy, a push strategy for displaying target content on the in-vehicle interactive interface, and a push strategy for seamlessly switching the music being played.

[0135] In one specific embodiment, when multiple users are present in the vehicle, the in-vehicle music adaptive push device 600 further includes a third determining module: The third determination module is used to determine the target content for each user from a pre-built user music preference library based on scene environment data and the emotional state of each user. Based on the in-vehicle zone audio system, an independent audio playback channel is matched for each user's seating position; The target content for each user is pushed to the audio playback channel at the corresponding user's seat location.

[0136] In one specific embodiment, the in-vehicle music adaptive push device 600 further includes an adjustment module: The adjustment module is used to adjust the target content or push strategy of each user based on the push priority rules and the interaction data between the user and the in-vehicle playback device if a conflict is detected between the push strategies of different users.

[0137] In one specific embodiment, the in-vehicle music adaptive push device 600 further includes a second generation module: The second generation module is used to generate at least one music adjustment option for adjusting the target content when the emotional state of the user in the car is identified as a preset anxiety emotion. The music adjustment option is used to adjust the style and / or genre of the target content. Obtain the user's selection command based on music adjustment options; Based on the selection command, the corresponding target music is determined, and the target music is used as the target content and pushed to the user according to the push strategy.

[0138] Regarding the apparatus in the above embodiments, the specific manner in which each unit performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0139] This embodiment also provides a vehicle, including: the vehicle executing any of the optional in-vehicle music adaptive push methods described above, thus achieving the same effect as the above implementation method.

[0140] This embodiment also provides a computer-readable storage medium, including: storing a computer program on the computer-readable storage medium, wherein when the computer program is executed by a processor, it implements any of the optional in-vehicle music adaptive push methods described above.

[0141] The beneficial effects of the above embodiments can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0142] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0143] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0144] In the description of this application, it should be understood that if the terms "upper", "lower", "front", "rear", "left" and "right" are used to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the position or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0145] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0146] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for adaptive music delivery in vehicles, characterized in that, The method includes: Based on the emotional state and scene environment data of the users in the vehicle, target content that matches the emotional state and scene environment data is determined from a pre-built user music preference library. The scene environment data is used to characterize the natural environment, time dimension and spatial travel scenario of the user during the vehicle driving process. Based on the music playback status in the car, the change characteristics of the emotional state, and the interaction data between the user and the in-car playback device, a push strategy for pushing the target content is determined. The change characteristics are used to indicate the dynamic change attributes, fluctuation magnitude, and state switching characteristics of the user's emotional state over time. Based on the aforementioned push strategy, the target content is pushed to the user.

2. The method according to claim 1, characterized in that, The method further includes: Acquire state perception data of the user inside the vehicle. The state perception data is used to characterize the user's emotion-related features and vehicle driving operation features that can reflect the user's emotions. The state perception data is classified to determine different types of state perception sub-data. The state perception sub-data is used to characterize the user's own emotion-related features and vehicle driving behavior association data related to the user's emotions. Feature extraction and emotion feature analysis are performed on the state-aware sub-data respectively to obtain the emotion analysis results corresponding to each type of state-aware sub-data. The emotion analysis results are fused to obtain the emotional state of the user in the vehicle.

3. The method according to claim 1, characterized in that, The step of determining target content matching the emotional state from a pre-built user music preference library based on the emotional state and scene environment data includes: Determine the first degree of match between the emotional state and multiple items in the user's music preference library; Based on the first matching degree, multiple candidate contents are determined; Determine the second matching degree between the scene environment data and multiple candidate contents; The target content is determined based on the second matching degree.

4. The method according to claim 3, characterized in that, The method further includes: The candidate content corresponding to the second matching degree is sorted according to the priority of the target matching degree to generate a target content list, which includes multiple candidate contents. Based on the push strategy, the candidate content is pushed to the user in the order of the target content list.

5. The method according to claim 1, characterized in that, The strategy for pushing the target content is determined based on the music playback status in the vehicle, the changes in the user's emotional state, and the interaction data between the user and the in-vehicle playback device, including: Based on the music playback status, the change characteristics of the emotional state, and the interaction data, the current playback scenario in the vehicle is determined. The playback scenario is used to characterize the scenarios related to the music push strategy formulation in the vehicle environment, as well as the music push intervention adaptation requirements corresponding to the change characteristics. Based on the playback scenario, the push strategy is determined. The push strategy includes at least one of the following: an automatic playback push strategy, a push strategy that displays the target content on the in-vehicle interactive interface, and a push strategy that seamlessly switches the music being played.

6. The method according to claim 1, characterized in that, When there are multiple users in the vehicle, the method further includes: Based on the scene environment data and the emotional state of each user, the target content corresponding to each user is determined from the pre-built user music preference library; Based on the in-vehicle zone audio system, an independent audio playback channel is matched for each user's seating position; The target content for each user is pushed to the audio playback channel at the corresponding user's seat location.

7. The method according to claim 6, characterized in that, The method further includes: If a conflict is detected between the push strategies of different users, the target content or push strategy of each user will be adjusted based on the push priority rules and the interaction data between the user and the in-vehicle playback device.

8. The method according to claim 1, characterized in that, The method further includes: When the emotional state of the user in the vehicle is identified as a preset anxiety, at least one music adjustment option is generated to adjust the style of the target content. Obtain the selection command input by the user based on the music adjustment options; Based on the selection instruction, the corresponding target adjustment music is determined, and the target adjustment music is used as the target content and pushed to the user according to the push strategy.

9. A vehicle, characterized in that, include: The vehicle performs the in-vehicle music adaptive push method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, include: A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the adaptive push method for in-vehicle music as described in any one of claims 1 to 8.