Guest greeting information processing method, vehicle and storage medium

By acquiring the user's image and voice sequences when the vehicle starts, and using a cross-modal attention fusion algorithm to generate personalized welcome information, the problem of lack of personalization in standardized content in existing technologies is solved, more accurate emotion and interest recognition is achieved, and the user experience is improved.

CN120716751APending Publication Date: 2025-09-30GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510868326.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the existing technology, after the vehicle is turned on, it pushes standardized welcome content based on preset rules. It lacks the ability to dynamically generate personalized content and cannot adapt to the user's real-time interest changes. In addition, the misjudgment rate of user emotions through single-modality recognition is high and there is a lack of multi-modal verification.

Method used

By determining whether emotional interest recognition is allowed when the vehicle starts, obtaining the user's image and voice sequences, and using a cross-modal attention fusion algorithm to combine image and voice features, it dynamically generates welcome information, including welcome images, music, and environmental adjustments, to meet the user's emotional and interest needs.

Benefits of technology

It improves the user's personalized driving experience, and the generated welcome information is more in line with the user's real emotions and interest changes, which improves the comfort and personalized experience of the in-car environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120716751A_ABST
    Figure CN120716751A_ABST
Patent Text Reader

Abstract

The invention provides a welcome information processing method, a vehicle and a storage medium, and the method comprises the steps: responding to the starting of the vehicle at a first moment, determining whether emotional interest recognition is allowed, and if yes, obtaining an image sequence and a voice sequence of a user in the vehicle operation process; determining emotion information and interest information of the user according to the image sequence and the voice sequence; and in response to starting of the vehicle at a second moment after the first moment, determining and outputting welcome information according to the emotion information and the interest information of the user. According to the method, the welcome information output after the vehicle is started at the second moment is determined according to the predetermined emotion information and interest information of the user after the first moment, the welcome information related to the emotion and interest of the user can be dynamically generated, the generated welcome information adapts to the interest change of the user, and the personalized driving experience feeling of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of interactive recommendation, and in particular to a welcoming information processing method, a vehicle, and a storage medium. Background Art

[0002] Smart cars are developing rapidly, and people have more and more requirements for the functional configuration of vehicles. Users pay more attention to the comfort of the in-car environment and personalized experience. For companies, they hope to break through the bottleneck of serious homogeneity of cabin products and form intelligent interactive services with brand tone through personalized scenario definition and recommendation algorithms.

[0003] In the prior art, after a vehicle is turned on, standardized welcome content is pushed based on preset rules, and the user's personalized experience is not high. Summary of the Invention

[0004] The purpose of this application is to address the deficiencies in the above-mentioned prior art and provide a welcome information processing method, vehicle and storage medium to improve the user's personalized experience.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present application are as follows: In a first aspect, an embodiment of the present application provides a method for processing welcome information, the method comprising: In response to the vehicle being started at a first moment, determining whether emotion interest recognition is allowed, and if so, obtaining an image sequence and a voice sequence of the user during the operation of the vehicle; Determining the user's emotion information and interest information based on the image sequence and the voice sequence; In response to the vehicle being started at a second moment after the first moment, welcome information is determined and output based on the user's emotional information and interest information, where the welcome information includes at least one of the following: a welcome image, welcome music, and vehicle environment information.

[0006] Optionally, determining whether to allow emotion interest recognition includes: Get the privacy protection level set by the user; If the privacy protection level is lower than a preset level, it is determined that emotion interest recognition is allowed.

[0007] Optionally, the acquiring of an image sequence and a voice sequence of a user during operation of the vehicle includes: Determining the start time of the user's conversation based on monitoring results of the voice collection device on the vehicle; An image sequence and a voice sequence starting from the conversation start time are acquired.

[0008] Optionally, determining the user's emotion information and interest information based on the image sequence and the voice sequence includes: Determining a target image sequence and a target speech sequence according to the image sequence and the speech sequence; Performing feature extraction on the target image sequence and the target speech sequence respectively to obtain an image vector and a speech vector; Performing semantic analysis on the target speech sequence to obtain the interest information; A cross-modal attention fusion algorithm is used to perform weighted vector fusion on the image vector and the speech vector to obtain the emotion information.

[0009] Optionally, determining a target image sequence and a target speech sequence according to the image sequence and the speech sequence includes: Aligning the image sequence and the speech sequence according to time information to obtain an aligned image sequence and an aligned speech sequence; An image sequence of a target period is selected from the aligned image sequence as the target image sequence, and a speech sequence of the target period is selected from the aligned speech sequence as the target speech sequence.

[0010] Optionally, determining and outputting welcome information based on the user's emotional information and interest information includes: Determining whether the emotion information belongs to the emotion type configured by the user that the system is expected to recognize; If so, the welcoming information is determined and output according to the emotion information and interest information of the user.

[0011] Optionally, determining the welcoming information according to the user's emotional information and interest information includes: determining and outputting a welcoming image related to the interest information based on the interest information and the emotion information; Determining and outputting, based on the interest information and / or the emotion information, welcoming music that matches the interest information and / or the emotion information; The parameters of the environment-related devices of the vehicle are adjusted to parameters matching the interest information and / or the emotion information.

[0012] Optionally, it also includes: In response to the user's feedback information regarding the welcome information, a label of the welcome information is updated, wherein the feedback information includes: like, dislike, and favorite.

[0013] Optionally, it also includes: Get multiple historical welcome information; Analyzing the target emotions in the plurality of historical welcome messages to obtain emotion change information, wherein the emotion change information is used to indicate the target emotion of the user at a preset fixed time; The target welcoming information at the preset fixed time is generated according to the emotion change information and the multiple historical welcoming information.

[0014] Optionally, it also includes: According to the positioning data of the interest information, regional cultural elements associated with the interest information are determined.

[0015] Optionally, it also includes: In response to the user's triggering operation for sharing the welcome information, obtaining the welcome terminal device of the target vehicle to be shared by the user; The welcoming information is synchronized to the welcoming terminal device of the target vehicle.

[0016] In the second aspect, an embodiment of the present application also provides a vehicle, comprising: a processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor. When the application is running, the processor and the storage medium communicate through the bus, and the processor executes the program instructions to perform the steps of the welcome information processing method described in the first aspect above.

[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program is read and executes the steps of the welcome information processing method described in the first aspect above.

[0018] The beneficial effects of this application are: The present application provides a welcome information processing method, vehicle, and storage medium. In response to a vehicle being started at a first moment, the method determines whether to allow entry into an emotion and interest recognition process, obtains a user's image sequence and voice sequence during vehicle operation, and determines the user's emotion and interest information based on the image sequence and voice sequence. When the user determines to allow entry into the emotion and interest recognition process after the vehicle is started at the first moment, the method obtains the user's image sequence and voice sequence after the first moment, and determines the user's emotion and interest information based on the user's emotion sequence and voice sequence after the first moment. Specifically, when determining the user's emotion information, the method considers not only the voice sequence in the user's conversation but also the image sequence of the user during the conversation, i.e., the user's facial expression information. This makes the obtained user emotion information more accurate and more consistent with the user's true emotions, avoiding the prior art method of determining the user's emotion through a single modality. Furthermore, when the vehicle is started at a second moment after the first moment, the method determines the welcome information to be output at the second moment based on the predetermined user emotion and interest information after the first moment. This method dynamically generates welcome information related to the user's emotions and interests, allowing the generated welcome information to adapt to changes in the user's interests and enhance the user's personalized driving experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 A schematic diagram of a system architecture for processing welcome information provided in an embodiment of the present application; Figure 2 A flowchart of a method for processing welcome information provided in an embodiment of the present application; Figure 3 A flowchart of a second method for processing welcome information provided in an embodiment of the present application; Figure 4 A flowchart of a third method for processing welcome information provided in an embodiment of the present application; Figure 5 A flowchart of a fourth method for processing welcome information provided in an embodiment of the present application; Figure 6 A schematic diagram of a flow chart of emotion recognition and data processing provided in an embodiment of the present application; Figure 7 A schematic diagram of a personalized experience generation process provided in an embodiment of the present application; Figure 8 A structural block diagram of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0022] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0023] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0024] Existing technologies rely solely on facial expressions or voice for emotion recognition, lacking multimodal validation. This results in a high rate of false positives, as recognition accuracy varies by as much as 30% between laboratory environments and real-world in-car scenarios. Furthermore, existing technologies fail to correlate specific semantic content. For example, recognizing the emotion "happy" may fail to extract key points of interest for a trip to Paris. Existing technologies push standardized welcome content based on pre-set rules, lacking the ability to dynamically generate personalized content. This makes recommendations inappropriate for real-time user interest changes, and lacks long-term memory mechanisms, rendering each interaction an isolated event.

[0025] Therefore, based on the shortcomings of the prior art, the present application provides a method for processing welcome information.

[0026] The user characteristics applicable to the welcome information processing method in this application are: users who pay attention to the comfort of the in-car environment and personalized experience, like intelligent and technologically advanced in-car functions, and pursue a high-quality lifestyle.

[0027] Optionally, the welcome information processing method provided in the embodiment of the present application is applied to a vehicle, which can be a terminal device installed on the vehicle or a server. The electronic device can be, for example, a mobile phone, tablet computer, laptop computer, PDA, desktop computer, or other terminal device with computing and display capabilities. Specifically, it can be applied to an application in a terminal device, such as an application (APP) on a mobile phone, an application system on a computer, etc. The vehicle includes Figure 1 A system for processing welcome information.

[0028] Figure 1 A schematic diagram of a system architecture for processing welcome information provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the system architecture may include: a user interaction layer, an application layer, an ambient light control module, a data processing layer, a data acquisition layer, and a data storage layer.

[0029] The data acquisition layer includes a camera and microphone. The camera can capture user facial images, support low-light imaging, real-time face tracking, and analyze facial features. The microphone can also provide noise reduction, support multi-person voice separation, and perform real-time audio processing.

[0030] The data storage layer includes a memory library that can store historical welcome information. The memory library can include user personalized profiles, historical welcome information, and personalized recommendation model parameters.

[0031] The data processing layer may include an emotion recognition model, a speech recognition model, and a topic extraction and analysis model. The emotion recognition model uses deep learning to analyze facial expressions and generate emotion vectors. The speech recognition model analyzes voice pitch, rhythm, and energy characteristics, and can use the speech content to determine emotion and generate speech vectors. The topic extraction and analysis model uses NLP to perform keyword extraction, topic clustering, user interest mapping, and time series context analysis to obtain interest and emotion information.

[0032] The application layer can include: personalized theme generation, a music playback module, and an ambient lighting control module. The personalized theme generation module includes a theme template library, a dynamic visual content synthesis engine, style transfer, and personalized adjustment. Specifically, it can generate a welcome image based on emotional and interest information. The music playback module includes a mood-music mapping library, a music streaming playback engine, and a copyright content association system. Specifically, it can determine and play welcome music based on emotional and interest information. The ambient lighting control module includes an emotion-color mapping model, a dynamic light effect generation engine, and an in-vehicle zone control system. Specifically, it can determine and output specific ambient lighting based on emotional and interest information.

[0033] The user interaction layer may include: an in-vehicle display screen and a voice interaction interface. The in-vehicle display screen can display welcome information, and the voice interaction interface can perform voice interaction.

[0034] The specific implementation process of the welcoming information processing provided in the embodiment of the present application is explained in detail below.

[0035] Figure 2 This is a flow chart of a method for processing welcome information provided by an embodiment of the present application, the execution subject of the method is the aforementioned vehicle. Figure 2 As shown, the method includes: S101: In response to a vehicle being started at a first moment, determining whether to allow entry into emotion interest recognition.

[0036] Specifically, if it is determined that the emotion and interest recognition is allowed, the following step S102 is executed. If it is determined that the emotion and interest recognition is not allowed, the process ends.

[0037] Specifically, when the user gets in the car and starts the vehicle at the first moment, it can be determined whether to enter the emotion interest recognition based on the user's preset operation.

[0038] S102: Acquire the user's image sequence and voice sequence during vehicle operation.

[0039] The image sequence may include facial images of the user at various moments within a period of time, and the voice sequence may include conversational voices of the user at various moments within a period of time.

[0040] Optionally, when the vehicle is started at the first moment and it is determined that entry into emotion interest recognition is allowed, it means that the user now agrees to recognize the user's image sequence and voice sequence during the operation of the vehicle. At this time, the camera and voice microphone on the vehicle can be activated, and the user's image sequence is collected through the camera on the vehicle, and the user's voice sequence is collected through the voice microphone, and the collected image sequence and voice sequence are sent to the electronic device.

[0041] S103: Determine the user's emotion information and interest information based on the image sequence and the voice sequence.

[0042] Optionally, a predetermined method may be used to determine the user's emotional information and interest information after the user boarded the vehicle at the first moment based on the user's facial expression in each image in the identified user image sequence and the conversation content in the voice sequence. The determined emotional information and interest information may be stored in a memory bank.

[0043] Among them, emotional information can refer to emotions such as happiness, sadness, anger, fear, disgust, surprise, neutrality, etc., and interest information can refer to topics that users like to discuss, etc. Based on the emotional information and interest information, the user's emotions when discussing the interest information can be known.

[0044] S104 . In response to the vehicle being started at a second moment after the first moment, determine and output a welcome message based on the user's emotional information and interest information.

[0045] The welcome information may include at least one of the following: a welcome image, welcome music, and vehicle environment information. The welcome image may refer to the welcome screen and content displayed on the welcome terminal device; the welcome music may refer to the music played on the welcome terminal device; and the vehicle environment information may refer to the color of the vehicle's ambient lighting, the vehicle's operating status, and the scent released by the vehicle's in-vehicle fragrance system.

[0046] Optionally, the second moment refers to the moment when the user re-enters the vehicle and restarts it after the first moment. That is, after the user restarts the vehicle at the second moment, the user's emotional information and interest information determined in step S103 after the vehicle was started at the first moment can be retrieved from the memory bank, and a preset method is used to determine and output a welcome message based on the determined user emotional information and interest information. That is, after the vehicle is restarted at the second moment, the welcome message output to the user is determined based on the user's emotional information and interest information determined after the first moment.

[0047] In this embodiment, in response to the vehicle being started at a first moment, a determination is made as to whether to allow access to emotion and interest recognition. The user's image and voice sequences are obtained during vehicle operation, and the user's emotion and interest information is determined based on the image and voice sequences. If the user determines to allow access to emotion and interest recognition after the vehicle is started at the first moment, the user's image and voice sequences after the first moment are obtained. The user's emotion and interest information after the first moment is determined based on the user's emotion and voice sequences after the first moment. This determination of the user's emotion information not only considers the voice sequence during the user's conversation, but also the image sequence during the conversation, i.e., the user's facial expressions. This makes the resulting user emotion information more accurate and more consistent with the user's true emotions, avoiding the prior art practice of determining user emotions through a single modality. Furthermore, when the vehicle is started at a second moment after the first moment, the welcome information output after the second moment of vehicle startup is determined based on the predetermined user emotion and interest information after the first moment. This allows for the dynamic generation of welcome information related to the user's emotions and interests, allowing the generated welcome information to adapt to changes in the user's interests and enhance the user's personalized driving experience.

[0048] Optionally, the determination of whether to allow entry into emotion interest recognition in S101 may include: Specifically, the privacy protection level set by the user may be obtained, and if the privacy protection level is lower than a preset level, it is determined that the emotion interest recognition is allowed.

[0049] Specifically, multiple privacy protection levels can be set on the welcome terminal device. For example, privacy protection level 1 is the standard mode, privacy protection level 2 is the emotion interest recognition mode off, and privacy protection level 3 is the emotion interest recognition mode on. The level of privacy protection level 1 is higher than that of privacy protection level 2, and the level of privacy protection level 2 is higher than that of privacy protection level 3. The preset level is level 2. When it is lower than privacy protection level 2, it is determined that entry into emotion interest recognition is allowed, that is, when the emotion interest recognition mode is determined to be turned on, it is determined that entry into emotion interest recognition is allowed.

[0050] Among them, the emotion interest recognition mode can be turned on or off manually by the user on the welcome terminal, or it can be controlled by voice. For example, the user can enter the emotion interest recognition by voice "turn on emotion recognition" or "exit private mode", and can also exit the emotion interest recognition by voice "turn off emotion recognition" or "enter private mode".

[0051] Optionally, the above S102, obtaining the user's image sequence and voice sequence during vehicle operation, may include: Specifically, the start time of the user's conversation can be determined based on the monitoring results of the voice collection device on the vehicle, and the image sequence and voice sequence starting at the start time of the conversation can be acquired. That is, when it is determined that emotional interest recognition is allowed, the acquisition of the user's voice sequence and image sequence can be started at the start time of the user's conversation and stopped at the end time of the user's conversation.

[0052] For example, if the user gets in the car and starts the vehicle at time t1, and determines to allow entry into emotion interest recognition, at t1+5min, the user starts chatting with a friend, then the user's starting time is t1+5min, and the user's voice sequence and image sequence are obtained from t1+5min. The user stops talking at t1+10min, and the acquisition of the user's voice sequence and image sequence is stopped at t1+10min.

[0053] Figure 3 The flowchart of the second method for processing welcome information provided in the embodiment of the present application is as follows: Figure 3 As shown, the above S103, determining the user's emotion information and interest information based on the image sequence and the voice sequence, may include: S201: Determine a target image sequence and a target speech sequence according to an image sequence and a speech sequence.

[0054] Specifically, a target image sequence can be determined from an image sequence using a preset method, and the target image sequence is a partial image sequence in the image sequence. A target voice sequence can be determined from a voice sequence using a preset method, and the target voice sequence is a partial voice sequence in the voice sequence.

[0055] For example, if the image sequence is the images between t1+5min and t1+10min mentioned above, and the speech sequence is also the speech between t1+5min and t1+10min mentioned above, using the preset method, it can be determined that the target image sequence is the images between t1+6min and t1+7min, and the target speech sequence is the speech between t1+6min and t1+7min.

[0056] S202 , performing feature extraction on the target image sequence and the target speech sequence respectively to obtain an image vector and a speech vector.

[0057] Specifically, features of facial expressions in each image in the target image sequence may be extracted to obtain an image vector, and features of intonation in the target speech vector may be extracted to obtain a speech vector.

[0058] S203: Perform speech analysis on the target speech sequence to obtain interest information.

[0059] Specifically, the NLP model can be used to perform contextual semantic analysis on the conversation content in the target speech sequence, identify keywords in the target speech sequence, and use the extracted keywords as interest information.

[0060] For example, if the identified keywords are "travel", "Paris", "Eiffel Tower", etc., then "travel", "Paris", and "Eiffel Tower" are respectively used as interest information.

[0061] S204: Use a cross-modal attention fusion algorithm to perform weighted vector fusion on the image vector and the speech vector to obtain emotional information.

[0062] Specifically, vector distance calculation can be performed, and a cross-modal attention fusion algorithm can be used to perform weighted vector fusion on the image vector obtained through expression recognition and the voice vector obtained through intonation analysis to obtain the user's emotional information after getting on the bus for the first time.

[0063] For example, when the user starts the vehicle for the first time and enthusiastically discusses the upcoming Paris trip plan with a friend, by analyzing the image sequence and voice sequence of the user during the conversation, the user's emotional information can be identified as "happy", and the identified interest information can be "travel", "Paris", and "Eiffel Tower".

[0064] For example, when the user starts the vehicle at the first moment, he or she gets into an argument with others, becomes emotional, and raises the tone of voice. By using the cross-modal attention fusion algorithm to perform weighted vector fusion on the image vector and the speech vector, the emotional information obtained is "anger". The content of the quarreling conversation is not recorded in the interest information, and only the emotional information is determined.

[0065] In this embodiment, by combining the facial expressions in the image sequence and the intonation in the voice sequence to perform feature fusion analysis, the user's emotional information during the conversation is obtained, and the user's interest information during the conversation is obtained based on the user's voice content, which can make the obtained emotional information more accurate.

[0066] Optionally, the above S201, determining the target image sequence and the target speech sequence according to the image sequence and the speech sequence, may include: Optionally, the image sequence and the speech sequence can be aligned according to time information to obtain an aligned image sequence and an aligned speech sequence. An image sequence of a target period is selected from the aligned image sequence as the target image sequence, and a speech sequence of a target period is selected from the aligned speech sequence as the target speech sequence. The selected target period is the aforementioned period t1 + 6 min to t1 + 7 min.

[0067] Specifically, the image sequence includes facial images at each moment, and the voice sequence includes voice at each moment. The facial images at each moment can be aligned with the voice at each moment according to the time information, for example, the facial image at t1+1s is aligned with the voice at t1+1s, the facial image at t1+2s is aligned with the voice at t1+2s, and so on, to obtain the aligned image sequence and the aligned voice sequence.

[0068] Optionally, determining and outputting the welcome information according to the user's emotional information and interest information in S104 may include: Optionally, it may be determined whether the emotional information belongs to an emotional type configured by the user that the system is expected to recognize. If so, the welcome information may be determined based on the emotional information and interest information of the user.

[0069] Specifically, it can be determined whether the user's emotional information after starting the vehicle at the first moment belongs to the emotion type configured by the user that the system is expected to recognize. If it is determined that the user's emotional information after starting the vehicle at the first moment belongs to the emotion type configured by the user that the system is expected to recognize, then a welcome message for the user after starting the vehicle at the second moment is determined and output based on the user's emotional information and interest information after starting the vehicle at the first moment.

[0070] If not, that is, the emotional information of the user after starting the vehicle at the first moment does not belong to the emotion type that the user configured to be recognized by the system, then it means that the emotion is an emotion that the user does not need. At this time, there is no need to determine and output the welcome information after the user starts the vehicle at the second moment based on the emotional information and interest information of the user determined after the first moment.

[0071] Among them, the emotion type that the user wants the system to recognize can be one or more of the aforementioned emotions such as happiness, sadness, anger, fear, disgust, surprise, neutrality, etc. The user can set the emotion type that the user wants the system to recognize in advance on the welcome terminal device.

[0072] Figure 4 The flowchart of the third method for processing welcome information provided in the embodiment of the present application is as follows: Figure 4 As shown, the above S104, determining and outputting the welcome information according to the user's emotional information and interest information, may include: S301: Determine and output a welcoming image related to the interest information based on the interest information and the emotion information.

[0073] Optionally, a welcome image related to the interest information may be dynamically generated based on the interest information and the emotion information, wherein the generated welcome image may include a welcome screen and welcome content.

[0074] For example, when the user's conversation information with friends after the first moment determines that the user's interest information is "travel" and "Paris", and the user's emotional information when it comes to "travel" and "Paris" is "happy", then when the user starts the vehicle at the second moment, a startup screen related to Paris travel can be generated.

[0075] S302: Determine and output, based on the interest information and / or the emotion information, welcoming music that matches the interest information and / or the emotion information.

[0076] Specifically, if the interest information contains the interest content of the user after starting the vehicle at the first moment, that is, the conversation content after the user starts the vehicle at the first moment involves the interest content, such as the Paris trip mentioned above, then the welcome music matching the emotional information and the emotional information can be determined and output based on the interest information and the emotional information; if the interest information does not contain the interest content of the user after starting the vehicle at the first moment, that is, the conversation content after the user starts the vehicle at the first moment does not involve the interest content, such as the content of a conversation with others, then the welcome music matching the emotional information can be determined and output based on the emotional information.

[0077] After the welcome music is determined, the determined welcome music can be displayed on the welcome terminal device. When the user needs to play it, the play button can be triggered on the welcome terminal device, so that the determined welcome music can be played in response to the user's play confirmation trigger operation.

[0078] For example, if the emotion information of the user after starting the vehicle at the first moment is "happy" and the interest information is "Paris travel" as mentioned above, then after the user starts the vehicle at the second moment, French cheerful music can be determined and played.

[0079] For example, if the user's emotional information after starting the vehicle at the first moment is "angry" and there is no information of interest, soothing light music may be determined.

[0080] S303: Adjust the parameters of the vehicle's environment-related devices to parameters that match the interest information and / or emotion information.

[0081] The vehicle's environment-related equipment may include, for example, ambient lighting inside the vehicle, the vehicle's operating status, and a fragrance system.

[0082] For example, when the user's emotion information after starting the vehicle at the first moment is "happy" and the interest information is "Paris travel", the vehicle's ambient light can be adjusted to a romantic warm color tone after the user starts the vehicle at the second moment.

[0083] For example, if the user's emotional state after starting the vehicle is "angry," the interior lighting can be adjusted to a calming blue hue. Alternatively, the vehicle can be set to a massage mode, or the fragrance system can be adjusted to release lavender essential oil.

[0084] For example, when the user starts the vehicle at the first moment, the user is alone in the car, listening to depressing music with a melancholy expression. By using the cross-modal attention fusion algorithm to perform weighted vector fusion on the image vector and the voice vector, the emotional information obtained is "sadness", and the vehicle's ambient light can be adjusted to a gentle ambient light.

[0085] In this embodiment, by dynamically determining a welcome image related to the interest information based on the interest information, determining matching welcome music based on the interest information and emotional information, and adjusting the parameters of the vehicle's environmental related equipment to matching parameters, a multi-sensory coordinated response such as images, ambient lights, and music is achieved, thereby enhancing the user experience.

[0086] Optionally, the method may further include: In response to user feedback information regarding the welcome information, the label of the welcome information is updated, wherein the feedback information may include: like, dislike, and favorite.

[0087] Specifically, each interface of the welcome terminal device may include feedback buttons such as "Like," "Dislike," and "Favorite," allowing users to provide feedback on the welcome information determined and output at the second moment. Users may also provide feedback on the welcome information through voice, for example, by saying "I like this welcome screen," "I don't like this music," etc.

[0088] Optionally, upon receiving user feedback regarding the welcome message, the tag of the welcome message can be updated in the aforementioned memory library. The tag may include "like," "dislike," or "favorite." For example, if a user generates soothing light music A for the emotion of "anger," and the user responds "like" when light music A is played, the tag of the light music A can be updated to "like" in the memory library. Subsequently, if the user's emotion is "anger," light music A can be directly selected as the welcome music.

[0089] Figure 5 The flowchart of the fourth method for processing the welcome information provided in the embodiment of the present application is as follows: Figure 5 As shown, the method may include: S401. Obtain multiple historical welcome information.

[0090] Optionally, the welcome information for a second moment determined based on the emotion information and interest information determined at the first moment after the user's emotion information and interest information are determined after the first moment can be stored as a historical welcome information. A historical welcome information then includes the emotion information and interest information determined at the first moment, as well as the welcome information generated at the second moment. A correlation can be established between the emotion information and interest information determined at the first moment and the welcome information generated at the second moment, resulting in a historical welcome information. Multiple historical welcome information can be stored in the memory bank.

[0091] S402: Analyze target emotions in multiple historical welcome messages to obtain emotion change information.

[0092] The emotion change information may indicate the user's target emotion at a preset fixed time.

[0093] Optionally, a plurality of historical welcome messages may include a plurality of different emotions, and analysis may be performed on a target emotion to obtain emotion change information of the target emotion.

[0094] Exemplarily, the target emotion may be "anxiety," and the emotional information in multiple historical welcome messages may be analyzed to determine the changing pattern of the "anxiety" emotion. For example, during the commute every Wednesday, the user is very likely to experience "anxiety."

[0095] S403: Generate target welcoming music at a preset fixed time according to the emotion change information and the plurality of historical welcoming information.

[0096] Specifically, target welcoming information at a preset fixed time may be generated according to the emotion change information and the welcoming information associated with the target emotion in a plurality of historical welcoming information.

[0097] For example, among multiple historical welcome information, the welcome music associated with the emotion of "anxiety" is "soothing music", then "soothing music" can be generated and played during the user's commute every Wednesday.

[0098] In this embodiment, by analyzing the user's continuous emotional fluctuations, welcoming information matching the emotional fluctuations is generated in advance.

[0099] Optionally, the method may further include: Specifically, regional cultural elements associated with the interest information can be determined based on the location data of the interest information. Regional cultural elements matching the interest information can be automatically associated based on GPS positioning, where the cultural elements can be, for example, special music, special restaurants, special attractions, etc.

[0100] For example, when users are discussing Paris travel, you can recommend the local music "La Vie En Rose." You can also recommend special restaurants in Paris to users.

[0101] Optionally, the method may further include: Specifically, in response to a user's triggering action to share the welcome information, the user obtains the welcome terminal device of the target vehicle to which the welcome information is to be shared, and synchronizes the welcome information to the welcome terminal device of the target vehicle. For example, a user can share a Paris travel picture to a family account and synchronize it to the welcome terminals of other vehicles.

[0102] Optionally, a time decay factor may be introduced into the memory library to automatically reduce obsolete interest data, for example, deleting a travel destination discussed half a year ago, thereby reducing the impact of obsolete interest data on the currently generated welcome information.

[0103] Optionally, this application can also stylize the user's interest information with the car company's brand to generate an image that combines both personality and brand characteristics. That is, the car company's brand logo can be added to the generated welcome screen, thus creating a welcome screen that combines personality and brand characteristics.

[0104] Optionally, the present application can also project emotional AR navigation prompts that match the information of interest onto the vehicle's windshield, such as a virtual Eiffel Tower appearing in a Paris travel scene, to enhance the immersive experience.

[0105] Optionally, this application can also use brainwave monitoring instead of cameras: real-time brainwave signals are collected through the brain-computer interface, and the user's emotional state is analyzed in combination with voice recognition to avoid the risk of camera privacy leakage.

[0106] Optionally, the present application can also enhance the accuracy of emotion recognition through wearable device technology, such as using steering wheel grip sensors, heart rate monitoring and other physiological indicators (skin conductivity, breathing rate).

[0107] Figure 6 A flowchart of emotion recognition and data processing provided in an embodiment of the present application is shown as follows: Figure 6 As shown: S501. The user gets on the bus.

[0108] The user getting on the vehicle refers to starting the vehicle at the first moment in the aforementioned step S101.

[0109] S502: Activate the data collection module and identify emotional information and interest information.

[0110] This step is the same as steps S101 to S103 in the aforementioned specific implementation.

[0111] S503: Store to memory.

[0112] After the emotion information and topic information are identified, the emotion information and topic information are stored in a memory bank.

[0113] S504: Update the preference model.

[0114] The preference model is updated based on the identified emotional information and interest information.

[0115] Figure 7 A schematic diagram of a personalized experience generation process provided in an embodiment of the present application is shown as follows: Figure 7 As shown: S601. The user gets on the bus.

[0116] The user getting on the vehicle refers to getting on the vehicle in step S104 in the aforementioned specific implementation manner, that is, starting the vehicle at a second moment after the first moment.

[0117] S602: Read historical data.

[0118] This step refers to reading the emotion information and interest information in S502.

[0119] S603: Generate personalized experience.

[0120] This step refers to determining and outputting the welcome information according to the user's emotional information and interest information in S104 in the aforementioned specific implementation.

[0121] S604: Collect user feedback.

[0122] S605: Update the preference model.

[0123] Specifically, the preference model may be updated according to the user's feedback information, wherein updating the preference model refers to updating the label of the welcome information in the aforementioned specific implementation.

[0124] Figure 8 This is a structural block diagram of a vehicle 700 provided in an embodiment of the present application. Figure 8 As shown, the electronic device may include: a processor 701 and a memory 702.

[0125] Optionally, a bus 703 may also be included, wherein the memory 702 is used to store machine-readable instructions executable by the processor 701. When the vehicle 700 is running, the processor 701 communicates with the memory 702 through the bus 703. When the machine-readable instructions are executed by the processor 701, the method steps in the above method embodiment are performed.

[0126] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method steps in the above-mentioned embodiment of the welcome information processing method are executed.

[0127] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0128] In addition, the functional units in the various embodiments of the present application can be integrated into a single processing unit, each unit can exist physically separately, or two or more units can be integrated into a single unit. If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0129] The above is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.

Claims

1. A method for processing welcoming information, characterized in that: The method comprises: In response to the vehicle being started at a first moment, determining whether emotion interest recognition is allowed, and if so, obtaining an image sequence and a voice sequence of the user during the operation of the vehicle; Determining the user's emotion information and interest information based on the image sequence and the voice sequence; In response to the vehicle being started at a second moment after the first moment, welcome information is determined and output based on the user's emotional information and interest information, where the welcome information includes at least one of the following: a welcome image, welcome music, and vehicle environment information.

2. The method for processing welcome information according to claim 1, wherein: The determining whether to allow emotion interest recognition includes: Get the privacy protection level set by the user; If the privacy protection level is lower than a preset level, it is determined that emotion interest recognition is allowed.

3. The method for processing welcome information according to claim 1, wherein: The acquiring of the image sequence and voice sequence of the user during the operation of the vehicle includes: Determining the start time of the user's conversation based on monitoring results of the voice collection device on the vehicle; An image sequence and a voice sequence starting from the conversation start time are acquired.

4. The method for processing welcome information according to claim 1, wherein: The determining the emotion information and interest information of the user according to the image sequence and the voice sequence includes: Determining a target image sequence and a target speech sequence according to the image sequence and the speech sequence; Performing feature extraction on the target image sequence and the target speech sequence respectively to obtain an image vector and a speech vector; Performing semantic analysis on the target speech sequence to obtain the interest information; A cross-modal attention fusion algorithm is used to perform weighted vector fusion on the image vector and the speech vector to obtain the emotion information.

5. The method for processing welcome information according to claim 4, characterized in that: The determining of a target image sequence and a target speech sequence according to the image sequence and the speech sequence includes: Aligning the image sequence and the speech sequence according to time information to obtain an aligned image sequence and an aligned speech sequence; An image sequence of a target period is selected from the aligned image sequence as the target image sequence, and a speech sequence of the target period is selected from the aligned speech sequence as the target speech sequence.

6. The method for processing welcome information according to claim 1, wherein: The determining and outputting the welcoming information according to the user's emotional information and interest information includes: Determining whether the emotion information belongs to the emotion type configured by the user that the system is expected to recognize; If so, the welcoming information is determined and output according to the emotion information and interest information of the user.

7. The method for processing welcome information according to claim 1 or 6, characterized in that: The determining of the welcoming information according to the user's emotional information and interest information includes: determining and outputting a welcoming image related to the interest information based on the interest information and the emotion information; Determining and outputting, based on the interest information and / or the emotion information, welcoming music that matches the interest information and / or the emotion information; The parameters of the environment-related devices of the vehicle are adjusted to parameters matching the interest information and / or the emotion information.

8. The method for processing welcome information according to claim 1, wherein: Also includes: In response to the user's feedback information regarding the welcome information, a label of the welcome information is updated, wherein the feedback information includes: like, dislike, and favorite.

9. The method for processing welcome information according to claim 1, wherein: Also includes: Get multiple historical welcome information; Analyzing the target emotions in the plurality of historical welcome messages to obtain emotion change information, wherein the emotion change information is used to indicate the target emotion of the user at a preset fixed time; The target welcoming information at the preset fixed time is generated according to the emotion change information and the multiple historical welcoming information.

10. The method for processing welcome information according to claim 7, wherein: Also includes: According to the positioning data of the interest information, regional cultural elements associated with the interest information are determined.

11. The method for processing welcome information according to claim 1, wherein: Also includes: In response to the user's triggering operation for sharing the welcome information, obtaining the welcome terminal device of the target vehicle to be shared by the user; The welcoming information is synchronized to the welcoming terminal device of the target vehicle.

12. A vehicle, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor. When the application is running, the processor and the storage medium communicate via the bus, and the processor executes the program instructions to perform the steps of the welcome information processing method described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Information recommendation and search method, device, equipment and computer-readable medium

    CN107918649A

  • Vehicle welcome method and device, electronic equipment and medium

    CN115534810A

  • Vehicle welcome system and control method

    CN117485241A

  • Vehicle-mounted welcome broadcasting method and system and vehicle

    CN117850584A

  • Intelligent vehicle welcome method and device and storage medium

    CN118445433A