Voice input processing method and device, equipment, medium and product

By identifying the emotional state of the user's voice input and dynamically adjusting the visual elements of the voice input interface, the problem of a single visual feedback mechanism in the prior art is solved, and a richer and more flexible emotional feedback effect is achieved, improving the user experience.

CN120220720AInactive Publication Date: 2025-06-27IFLYTEK CO LTD

Patent Information

Application Number
CN202510694667.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The visual feedback mechanism of the existing voice input system is relatively single, lacks flexibility, and cannot effectively capture and reflect the user's emotional state.

Method used

By identifying the emotional state during user's voice input, and based on the mapping relationship between the preset emotional state and the visual display effect, the matching target visual display effect is determined, and the visual elements of the voice input interface are dynamically adjusted, such as color, waveform and interface deformation, to achieve visual expression of emotions.

Benefits of technology

It improves the diversity and flexibility of visual feedback effects during voice input, enhances the user's interactive experience, and can instantly reflect the user's emotional state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220720A_ABST
    Figure CN120220720A_ABST
Patent Text Reader

Abstract

The invention provides a voice input processing method and apparatus, a device, a medium and a product. The method comprises the steps of identifying an emotional state when a user inputs a voice signal; based on a preset mapping relationship between the emotional state and a visual display effect, determining a target visual display effect matched with the emotional state of the user, the target visual display effect being used for visually displaying the emotional state of the user; and presenting the target visual display effect in a voice input interface corresponding to the voice signal. The diversity and flexibility of visual feedback in the voice input process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech processing, and in particular to a speech input processing method, device, equipment, medium and product. Background Art

[0002] As an important human-computer interaction method, voice input is widely used in various intelligent devices and software systems. In order to enhance the interactivity and fun of voice input, many systems will display dynamic animation effects when users input voice, such as dynamic waveforms or icon animations, to intuitively feedback the strength and change trend of the user's voice signal.

[0003] However, currently these visual feedback mechanisms mostly rely on fixed visual templates or preset simple animation effects, resulting in a relatively simple overall presentation form and low flexibility. Summary of the invention

[0004] Based on the above-mentioned technical status, the present application provides a method, device, equipment, medium and product for processing voice input, which can improve the diversity and flexibility of visual feedback effects during voice input.

[0005] In order to achieve the above technical objectives, this application specifically proposes the following technical solutions: According to the first aspect of an embodiment of the present application, a method for processing voice input is provided, including: identifying the emotional state of a user when inputting a voice signal; determining a target visual display effect that matches the emotional state of the user based on a mapping relationship between a preset emotional state and a visual display effect, wherein the target visual display effect is used to visually display the emotional state of the user; and presenting the target visual display effect in a voice input interface corresponding to the voice signal. In some implementations, the visual display effect includes at least one of a color parameter of an HSL color space, a speed of particles contained in a waveform of the speech signal, a deformation parameter of the speech input interface, and a dynamic texture parameter.

[0006] In some implementations, presenting the target visual display effect in the voice input interface corresponding to the voice signal includes: adjusting the current visual display effect of the voice input interface to gradually change it to the target visual display effect.

[0007] In some implementations, when the target visual display effect includes target color parameters and / or target deformation parameters of the voice input interface; wherein, adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect includes: adjusting the current color parameters of the voice input interface to gradually change to the target color parameters; and / or, adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

[0008] In some implementations, adjusting the current color parameters of the voice input interface to gradually change to the target color parameters includes: determining an adjustment step value for the color parameters; based on the adjustment step value of the color parameters, adjusting the current color parameters of the voice input interface to gradually change to the target color parameters.

[0009] In some implementations, the color parameters include hue parameters and / or saturation parameters; wherein, determining the adjustment step value of the color parameters includes: determining the adjustment step value of the hue parameters based on the product of the difference in hue parameters between the current hue parameters and the target hue parameters of the voice input interface and a preset adjustment ratio; and / or, determining the adjustment step value of the saturation parameters based on a preset saturation adjustment value.

[0010] In some implementations, adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters includes: performing weighted summation on the current deformation parameters and the target deformation parameters based on the respective weights corresponding to the current deformation parameters and the target deformation parameters, and determining the weighted summation result as the adjustment step value of the deformation parameters; based on the adjustment step value of the deformation parameters, adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

[0011] In some implementations, when the target visual display effect includes a target particle velocity, adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect further includes: determining a target number of particles to be pre-loaded based on the performance parameters of the display device; pre-loading the target number of particles through a WebGPU parallel computing framework, and rendering the target visual display effect based on the target particle velocity.

[0012] In some implementations, identifying the emotional state when recognizing the user's input voice signal includes: obtaining input data in each modality, where the input data in each modality includes at least one of the voice signal, the text data corresponding to the voice signal, the video data when the user inputs the voice signal, and the heart rate data; for each modality among the various modalities, extracting features from the input data in that modality to obtain a feature representation in that modality; based on the feature representations in the various modalities, obtaining a comprehensive feature; and according to the comprehensive feature, identifying the emotional state when the user inputs the voice signal.

[0013] In some implementations, the method further includes: obtaining feedback data of the user on the target visual display effect; and updating the mapping relationship based on the feedback data.

[0014] In some implementations, the feedback data includes interaction behavior data of the user on the target visual display effect; wherein, updating the mapping relationship based on the feedback data includes: updating the mapping relationship based on the interaction behavior data of the user on the target visual display effect, so that the mapping relationship conforms to the user's preference.

[0015] According to a second aspect of the embodiments of the present application, there is provided an electronic device, including a memory and a processor; the memory is connected to the processor and is used for storing programs; the processor is used for implementing the processing method of voice input as described in the first aspect by running the programs in the memory.

[0016] According to a third aspect of the embodiments of the present application, there is provided a storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processing method of voice input as described in the first aspect is implemented.

[0017] According to a fourth aspect of the embodiments of the present application, there is provided a computer program product, including computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute: the processing method of voice input as described in the first aspect.

[0018] A method, apparatus, device, medium, and product for processing voice input provided by an embodiment of the present application identify the emotional state of a user during voice input, determine a target visual display effect that matches the user's current emotional state according to a preset mapping relationship between the emotional state and the visual display effect, and then display it on the voice input interface. Since the target visual display effect is determined based on the user's emotional state, it can visually express the user's current emotional state, providing diverse visual effects for the visual feedback mechanism during the voice input process. These visual display effects can also be adjusted in real time as the user's emotions change, so the flexibility of the visual display can be improved and the user's interaction experience can be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0020] Figure 1 It is a flowchart of a method for processing voice input provided by an embodiment of the present application.

[0021] Figure 2 It is a flowchart of identifying the emotional state of a user provided by an embodiment of the present application.

[0022] Figure 3 It is a schematic diagram of the principle of identifying the emotional state of a user based on multimodal data provided by an embodiment of the present application.

[0023] Figure 4 It is a schematic diagram of the process of gradually adjusting the hue parameter provided by an embodiment of the present application.

[0024] Figure 5 It is a schematic diagram of the process of gradually adjusting the interface deformation parameter provided by an embodiment of the present application.

[0025] Figure 6 It is a schematic diagram of the structure of a device for processing voice input provided by an embodiment of the present application.

[0026] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The technical solutions provided in the embodiments of the present application can be exemplarily applied to hardware devices such as processors, electronic devices, servers (including cloud servers), or packaged as software programs to be run. When the hardware device executes the processing procedures of the technical solutions in the embodiments of the present application, or when the above software program is run, the automatic splitting of the target task and the automatic invocation of the application programming interface required for the task can be realized, achieving the purpose of completing the target task. The embodiments of the present application only exemplarily introduce the specific processing procedures of the technical solutions of the present application, and do not limit the specific implementation forms of the technical solutions of the present application. Any technical implementation form that can execute the processing procedures of the technical solutions of the present application can be adopted by the embodiments of the present application.

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.

[0029] Before introducing the solutions of the present application, the related technologies will be introduced first: The current voice input systems mainly focus on voice recognition and text conversion functions, and most of their user interfaces are built based on preset static skins or templates. For example, when some mainstream mobile phone input methods perform voice input, they will display the intensity and changes of the user's voice signal through dynamic waveforms or icon animations. However, these visual feedback mechanisms usually rely on fixed visual templates or simple real-time data-driven implementations. Although they can effectively reflect the basic physical characteristics of sound, such as by displaying the audio waveform or spectrogram in real time, there are still obvious deficiencies in capturing and feedbacking the user's emotional state. Specifically, traditional systems mostly use fixed visual templates or basic animation effects to present the voice input process, without fully considering the personalized and contextual needs of users. For example, they do not make corresponding interface adjustments according to the user's emotional fluctuations.

[0030] In addition, the accuracy of the emotion recognition module in some solutions is relatively low, making it difficult to parse the emotional information contained in the voice in real time and accurately, resulting in a disconnection between the dynamic visualization effect and the actual emotional state of the user.

[0031] At the same time, due to the lack of an efficient data collection, processing, and visualization coordination mechanism, such systems often face high latency problems and cannot meet the user's demand for instant interaction.

[0032] It should be noted that the user's emotion has a significant impact on the efficiency and experience of voice input. However, traditional voice input systems do not fully recognize this factor and lack effective emotion recognition and response capabilities. With the development of affective computing and deep learning technologies, how to utilize these advanced technologies to capture the emotional cues in the user's voice, and accordingly achieve dynamic adjustment of the interface, thereby enhancing the user experience and providing a more rich interactive experience, has become an urgent problem to be solved.

[0033] In view of this, the embodiments of the present application are committed to providing a method, apparatus, device, medium and product for processing voice input. By capturing the emotional state of the user in real time during the voice input process, and according to the mapping relationship between the preset emotional state and the visual display effect, the target visual display effect suitable for the emotional state when the user inputs the voice signal is determined. Subsequently, in the voice input interface corresponding to the voice signal, the target visual display effect is dynamically presented to achieve the visual expression of the user's emotional state. This can not only enhance the diversity and flexibility of the voice input interface, but also enhance the user's interactive experience. Details will be described one by one in the following embodiments.

[0034] Exemplary Method Figure 1 The following is a flowchart of a method for processing voice input provided by an embodiment of the present application. As Figure 1 shown, the method for processing voice input provided in this embodiment includes steps S101 - S103: S101. Identify the emotional state of the user when inputting the voice signal.

[0035] Identifying the emotional state contained in the user's voice input process is one of the key links to achieve emotional interaction. To improve the accuracy of identifying the user's emotional state, in some embodiments, the user's current emotional state can be identified based on the input data in each modality of the user. As Figure 2 shown, it specifically includes the following steps S201 to S204: S201. Obtain the input data in each modality, and the input data in each modality includes at least one of a voice signal, text data corresponding to the voice signal, video data when the user inputs the voice signal, and heart rate data.

[0036] In this embodiment, the user's emotional state can be captured through various modality data. Each modality includes at least one of a voice modality, a text modality, a video modality, and a heart rate modality. The data acquisition process of each modality will be elaborated in detail below: Among them, the input data in the voice modality can be obtained by the voice acquisition module (such as a microphone) to capture the user's voice signal in real time when the user activates the voice input function. For example, in the voice input interfaces of various applications such as chat software and search engines, the user can start the voice input function by clicking the microphone button.

[0037] In some examples, to improve the voice quality, a dual microphone array combined with beamforming technology can be used to suppress ambient noise to ensure the clarity of the voice signal. Among them, the sampling rate can be set to 16 kHz to fully cover the basic frequency range of human voices (80 - 255 Hz), thus ensuring the basic voice quality of the voice input. For example, when the user exclaims "This function is amazing!", the dual microphone array can not only record the basic frequency components in the speech, but also identify and emphasize the parts where the high-frequency energy increases significantly (about 2000 - 4000 Hz).

[0038] Since certain emotional expressions in the sentence are often accompanied by a sudden increase in energy in a specific frequency band. Therefore, through the above sampling rate configuration, the dual microphone array can not only accurately transcribe the user's voice content, but also reflect the emotional state of the speaker, thus helping to capture the emotional information in the voice signal.

[0039] Text data can also reflect the user's emotional state to a certain extent. Therefore, after obtaining the voice signal, the voice signal can also be transcribed into text data to obtain the input data in the text modality, providing multi-dimensional support for identifying the user's emotional state.

[0040] In addition, since facial expressions are also an important part of emotional expression, the user's facial video data can also be collected in real time through a camera to obtain the input data in the video modality.

[0041] In addition to this, heart rate, as one of the physiological indicators, is sensitive to emotional changes and can reflect the user's internal emotional state to a certain extent. Therefore, the user's heart rate data can also be collected through a heart rate monitoring device (such as a smart bracelet or watch) to obtain the input data in the heart rate modality.

[0042] Through the input data of the above various modalities, multiple dimensions of data sources can be provided for identifying the user's emotional state to improve the recognition accuracy of the user's emotional state.

[0043] In some embodiments, to enhance the user's personalized experience, an authorization option for facial video data collection can also be provided. The user can choose whether to enable the camera permission according to personal preferences. When the user selects the authorization option to turn on the camera, the user's facial video data can be collected.

[0044] Since the microphone, camera, and heart rate monitoring device may not be started simultaneously, it may cause an out-of-sync problem among multi-source data. Therefore, after obtaining the multi-source data, a timestamp alignment algorithm can be used to achieve the synchronous processing of multi-source data, so that the multi-source data is synchronized in time.

[0045] Specifically, by adding accurate timestamp marks to each record at the source of data acquisition, the data from different devices can be aligned and synchronized. In this way, it can be ensured that multi-modal data such as voice, video, and heart rate are accurately matched on the time axis, providing a basis for subsequent emotion state recognition.

[0046] S202. For each modality among all modalities, extract features from the input data under this modality to obtain the feature representation under this modality.

[0047] There are corresponding feature extraction methods for different modalities. For the input data under different modalities, the corresponding feature extraction methods can be used for feature extraction. As Figure 3 shown, specifically as follows: In some emotion states, the performance of certain acoustic features will be more prominent. For example, in the expression of anger, the standard deviation of the fundamental frequency will increase significantly (exceeding 40Hz). Therefore, for speech signals, multi-dimensional MFCC (Mel Frequency Cepstral Coefficients), fundamental frequency (F0), and speech rate (syllables / second) can be extracted as acoustic features to obtain the feature representation under the speech modality. In order to balance both the accuracy of feature expression and the efficiency of feature extraction, 13-dimensional MFCC can be selected.

[0048] For text data, the sentiment tendency of the text can be extracted through a pre-trained natural language processing model, such as the RoBERTa model, and a sentiment word density map can be constructed based on the NRC sentiment dictionary to obtain the feature representation under the text modality. For example, if the RoBERTa model recognizes that a sentence contains multiple negative words (such as "terrible", "failed", "disappointed"), these words will be recognized and counted in the statistics of negative sentiment words. For facial video data, multiple (such as 468) facial key points can be detected in real time through a lightweight facial detection model (FaceMesh model) to obtain facial expression features. For example, when it is detected that the corners of the mouth are upturned (AU12 value > 0.7), this indicates the existence of significant positive expression features.

[0049] For heart rate data, key indicators such as average heart rate and heart rate variability (HRV) can be extracted from it as heart rate characteristics reflecting the user's physiological state. These characteristics can reflect the user's emotional fluctuations. For example, when it is detected that the user's average heart rate increases and the heart rate variability decreases, it may indicate that the user is in a tense or anxious emotional state. In this way, emotion-related characteristics in the physiological modality can be obtained, thereby enhancing the comprehensiveness and accuracy of emotion recognition.

[0050] S203. Based on the feature representations in each modality, obtain comprehensive features.

[0051] Continue to refer to Figure 3 , after obtaining the feature representations of each modality, they can be fused through an attention mechanism to generate a comprehensive feature representation. Specifically, by inputting the feature representations of each modality into their respective corresponding linear layers for preliminary processing, the feature representations of different modalities are transformed into a common hidden dimension space for subsequent processing. For example, all feature representations are mapped into a 128-dimensional space.

[0052] Next, the feature representations of each modality after preliminary processing are concatenated together to obtain a new combined feature. This combined feature will further be input into the fusion layer to enhance its expressive ability using an activation function and generate a fused feature representation.

[0053] To further optimize the above fusion process, the weights of the feature representations in each modality can be calculated through an attention mechanism. Specifically, it can be implemented through an attention layer, which can dynamically evaluate the contribution degree of each feature representation according to the current input.

[0054] Finally, based on the calculated weights, the feature representations in each modality are weighted and summed to obtain the final comprehensive features, ensuring that the feature representations that are most critical for understanding the user's emotional state can occupy a more important position in the final result.

[0055] S204. According to the comprehensive features, identify the emotional state of the user when inputting a voice signal.

[0056] Continue to refer to Figure 3 , in some embodiments, an emotion recognition model can be used to identify the emotional state expressed by the user during voice input based on the aforementioned comprehensive features. This model can combine multi-modal information such as voice, text, facial expressions, and physiological signals to perform accurate emotion analysis.

[0057] The emotion recognition model can adopt a hybrid neural network model. For example, it combines a convolutional neural network (CNN) with a recurrent neural network (RNN) or a Transformer module to capture local features and temporal dependencies simultaneously. The model receives the comprehensive features after multi-modal fusion as input and performs multi-level emotion perception and classification processing based on this.

[0058] In terms of emotion classification, the model not only supports the recognition of 7 basic emotions, including: happy, sad, angry, fear, surprised, disgusted, and neutral, but also extends to 20 compound emotion categories, such as disappointed, anxious, frustrated, excited, expectant, ironic, etc., so as to achieve a more fine-grained and realistic emotion expression classification effect.

[0059] What the model finally outputs are emotion state labels with confidence levels, indicating the emotion type and its intensity distribution that the current speech input is most likely to correspond to. For example, when the user inputs the speech: "I simply can't stand this design anymore!", the model will comprehensively analyze information such as intonation changes, keyword semantics, facial micro-expressions, and heart rate fluctuations in the speech and output the following emotion recognition results: Angry (main emotion, confidence level 72%); Disgusted (secondary emotion, confidence level 25%); Disappointed (secondary emotion, confidence level 3%).

[0060] This indicates that the user's current emotion is mainly manifested as anger, accompanied by a certain degree of disgust and slight disappointment. Through this fine-grained emotion recognition mechanism, the actual psychological state of the user can be understood more accurately, providing more personalized feedback and response strategies for subsequent human-computer interaction. For example, dynamically adjusting the color, animation style, or prompt content of the speech input interface according to the recognized emotion, so as to enhance the emotional resonance of the user experience.

[0061] When training the emotion recognition model, a two-stage training strategy can be adopted to improve its accuracy and generalization ability. First, pre-train on the widely used emotion recognition dataset IEMOCAP to help the model learn basic emotion features and patterns. Then, fine-tune on the self-owned dataset in a specific application scenario to make it better adapt to the specific requirements and characteristics in actual applications.

[0062] To further improve the training effect and the robustness of the model, data augmentation processing can also be performed on the dataset. For example, by adding room impulse responses (RIRs) to simulate different acoustic environments. In this way, not only can the data sample size be expanded, but the model can also perform more stably when facing the variable environmental noises in the real world.

[0063] In terms of real-time inference optimization, static quantization technology can be adopted to convert model weights to INT8 precision to reduce computational latency. In this way, not only can high recognition accuracy be maintained, but also the system response time can be ensured to be less than 30 milliseconds (ms), thus enhancing the user experience, especially in application scenarios that require quick feedback. Additionally, it enables the model to run smoothly on resource-constrained devices, expanding its application scope.

[0064] Continuing to refer Figure 1 , after step S101 of the voice input processing method of this embodiment, step S102 may further be included.

[0065] S102. Based on the mapping relationship between the preset emotional state and the visual display effect, determine the target visual display effect that matches the user's emotional state.

[0066] Among them, the target visual display effect is used to visually and intuitively present the user's emotional state, thereby enhancing the emotional expression and interaction experience during the voice input process.

[0067] Each visual display effect corresponds to a set of adjustable visualization parameters. These parameters include at least one of the color parameters in the HSL color space, the speed of particle movement in the voice waveform, the deformation degree of the voice input interface, and the manifestation form of the dynamic texture. By combining and dynamically adjusting these visual elements, rich and diverse emotional feedback effects can be achieved. Among them, HSL is a model that represents colors as hue, saturation, and lightness. Hue (H) is the basic attribute of a color, representing the type of color, such as red, green, blue, etc.; saturation (S) represents the purity of the color, and the higher the saturation, the more vivid the color; lightness (L) controls the brightness and darkness of the color. By adjusting the hue, saturation, and lightness, the color atmosphere of the visual elements can be changed.

[0068] To achieve an intelligent match between the visual display effect and the user's emotional state, a mapping rule table between the emotional state and the visual parameters can be constructed in advance. The mapping rule table includes various emotional states and their corresponding visualization configuration parameters, enabling the automatic invocation of the corresponding visual style for real-time feedback based on the recognized emotional state. The following are example configurations for some emotional states, as shown in Table 1: Table 1 Mapping Rule Table When it is recognized that the user is in a "joyful" mood, the voice input interface can adopt a bright yellow-green tone (HSL value of (60, 90%, 80%)), the particles in the voice waveform jump quickly, the interface is moderately enlarged, and it is combined with a dynamic texture of starlight twinkling to create a relaxed and pleasant atmosphere.

[0069] If the "angry" emotion is detected, the interface will switch to a red color tone (HSL value: (0, 85%, 50%)), the particles in the voice waveform will jitter at high frequency, the interface edges will be sharpened, and a flame ripple effect will be superimposed to enhance the visual communication of intense emotions. For example, in the state of angry emotion, the target visual display effect can be set as follows: the proportion of red is 80%, the serration degree of the interface edge is 5 Px, and the particle speed is 12 Hz.

[0070] For the "sad" emotion, a cool blue color (HSL value: (240, 40%, 30%)) can be adopted. The particles in the voice waveform will jump slowly, the interface will shrink and display, and an animated texture of raindrops falling will be supplemented to convey the feeling of depression and calmness.

[0071] When the user's emotion is "excited", the voice waveform will jump quickly and be brightly colored to express the user's high-spirited emotion. The frequency of the waveform will increase, and bright and saturated colors (such as high brightness and saturation settings in the HSL value) will be adopted to create a visual effect full of vitality and passion.

[0072] When it is detected that the user has an anxiety emotion lasting for 3 seconds, the interface background will gradually change to a dark blue ripple, a breathing light effect will appear at the edge of the voice input box, and a meditation guidance animation will pop up automatically.

[0073] On the contrary, when the user's emotion is "calm", the voice waveform will show the characteristics of being smooth and with a soft color tone. At this time, a milder color configuration (such as a lower saturation and medium brightness HSL value) can be selected, and the dynamic changes of the waveform will be reduced to present a visual feeling of tranquility and harmony, so as to reflect the user's peaceful state of mind. For example, in the neutral emotion state such as "calm", the target visual display effect can be set as follows: the proportion of blue is 50%, the interface edge is smooth, and the particle speed is 6 Hz.

[0074] Through the dynamic visual feedback mechanism based on emotion recognition, it is possible to flexibly adjust the visual display effect of the voice input interface according to the different emotional states of the user, so as to provide a more user-emotion-experience-fitting interaction environment.

[0075] After identifying the user's emotional state, the mapping rule table constructed in advance can be looked up to obtain the interface effect that can visually reflect the user's current emotional state.

[0076] It should be noted that each emotional state can correspond to at least one dynamic texture effect, and the user can select one of them as the dynamic texture effect to be displayed in that emotional state.

[0077] S103. Present the target visual display effect in the voice input interface corresponding to the voice signal.

[0078] After obtaining various visual elements that match the user's current emotional state, the interface can be rendered based on these visual elements to present a target visual display effect in the voice input interface corresponding to the voice signal.

[0079] In this embodiment, after identifying the user's emotional state during the voice input process, based on the mapping relationship between the preset emotional state and the visual display effect, a target visual display effect that matches the user's current emotional state is determined and presented in the user's voice input interface to intuitively reflect the user's current emotional state and provide rich and diverse forms of expression for the visual feedback mechanism of voice input. Since the visual display effect can be dynamically adjusted as the user's emotional state changes, the flexibility of the visual display can be improved, ensuring the diversity and immediacy of the user experience.

[0080] To ensure that the visual changes can be smoothly transitioned and avoid discomfort caused to the user by sudden changes, a dynamic transition scheme can also be adopted in this embodiment.

[0081] Specifically, in step S103, when it is necessary to switch from the current visual display effect to a new visual display effect, a gradient technique can be used to achieve a smooth transition. That is, step S103 specifically includes: adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect.

[0082] For example, if the user's current emotion changes from "calm" to "excited", then by slowly increasing the saturation and brightness of the color, accelerating the speed of the particles in the waveform, and gradually enhancing the dynamic effect of the interface until it fully reaches the target visual display effect corresponding to the "excited" emotion.

[0083] Among them, there can be various specific implementation methods for the gradient adjustment, as follows: In some implementation methods, when the target visual display effect includes target color parameters, adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect includes: adjusting the current color parameters of the voice input interface to gradually change to the target color parameters.

[0084] In some implementation methods, when the target visual display effect includes the target deformation parameters of the voice input interface, adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect includes: adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

[0085] In some implementations, when the target visual display effect includes the target color parameters and the target deformation parameters of the voice input interface, the current visual display effect of the voice input interface is adjusted to gradually change to the target visual display effect, including: adjusting the current color parameters of the voice input interface to gradually change to the target color parameters, and adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

[0086] During the adjustment of the color parameters, by adjusting the current color parameters of the voice input interface to gradually transition to the target color parameters, the smoothness and naturalness of the color change process can be ensured.

[0087] During the adjustment of the interface deformation parameters, by adjusting the current interface layout or element shape according to the predefined target deformation parameters, for example, by precisely controlling the change rate and path of the attributes such as the position, size, and angle of each interface element, to ensure that the entire transition process is smooth and natural, thereby achieving visual dynamic changes. Among them, the deformation can include scaling, rotation, and deformation animation. The deformation animation includes changing a rectangular button into a circular icon.

[0088] In some complex scenarios, the color and deformation can also be adjusted simultaneously. This requires simultaneously handling the smooth transition of colors and synchronously processing the change of the interface shape. For example, while a voice input interface changes from square to circular, the background color also changes from light blue to dark blue. The following will detail the respective gradual adjustment processes of the color parameters and the deformation parameters: In some implementations, when adjusting the current color parameters of the voice input interface to gradually change to the target color parameters, a series of intermediate color values can be determined and applied step by step at a set speed until the final target color is reached to achieve a natural transition effect. Specifically, it includes the following steps a1 and a2: Step a1: Determine the adjustment step value of the color parameters.

[0089] Among them, the color parameters include the hue parameter and / or the saturation parameter; then step a1 determines the adjustment step value of the color parameters, including: determining the adjustment step value of the hue parameter based on the product of the hue parameter difference between the current hue parameter and the target hue parameter of the voice input interface and the preset adjustment ratio; and / or, determining the adjustment step value of the saturation parameter based on the preset saturation adjustment value.

[0090] Among them, the adjustment step value of the hue parameter is the key to realizing the color gradual change process. By reasonably setting the adjustment amplitude of each step, that is, the adjustment step, the speed and smoothness of the change of the hue parameter from the current state to the target state can be controlled.

[0091] Assume the current hue parameter is H current , the target hue parameter is Htarget , the preset adjustment ratio is R H , which ranges from 0 to 1, for example, 0.15. First, calculate the hue parameter difference ∆H = H target - H current . Next, calculate the adjustment step value Step H = ∆H * R H .

[0092] For the saturation parameter, a fixed adjustment value can be set for it, and the value range is [0.1, 0.5], such as 0.1, 0.2, 0.3, 0.4, 0.5, etc.

[0093] Step a2: Based on the adjustment step value of the color parameter, adjust the current color parameter of the voice input interface to gradually change to the target color parameter.

[0094] Through the adjustment step value of the color parameter, gradually adjust the current color parameter of the voice input interface to make it smoothly transition to the target color parameter. For example, if the target is to gradually change from a darker blue (HSL value is (240, 40%, 30%)) to a brighter sky blue (HSL value is (180, 80%, 90%)), then by calculating the adjustment step values of the hue parameter, saturation, and brightness respectively, and gradually adjusting according to the adjustment step value, the visual smoothness during the color conversion process can be ensured.

[0095] As Figure 4 shown, for the hue parameter, it is based on the currently adjusted hue parameter, and each time ∆H * R H is added until the target hue parameter is reached.

[0096] In some implementation manners, adjusting the current deformation parameter of the voice input interface to gradually change to the target deformation parameter includes the following steps b1 and b2: Step b1: Based on the weights corresponding to the current deformation parameter and the target deformation parameter respectively, perform weighted summation on the current deformation parameter and the target deformation parameter, and determine the weighted summation result as the adjustment step value of the deformation parameter.

[0097] During the process of the interface form changing dynamically according to the user's emotion, in order to achieve a smooth transition from the current deformation state to the target deformation state, the step value for controlling the gradual change rhythm can be calculated according to the difference between the current deformation parameter (such as the interface scaling ratio, the deformation degree of the control, etc.) and the target deformation parameter, and in combination with the set weight ratio. Among them, the weights corresponding to the current deformation parameter and the target deformation parameter respectively can be set according to the user's needs, specifically as follows: In some examples, the weights can be set according to the immediate feedback requirements. For example, if the user hopes to quickly adjust the interface form when detecting a change in emotion to obtain immediate feedback. Then the weight of the target deformation parameter should be set relatively high (e.g., 0.8), while the weight of the current deformation parameter is relatively low (e.g., 0.2). In this way, the new interface form can be presented as soon as possible, reducing the transition time and enhancing the user's immediate experience.

[0098] In other examples, the weights can also be set according to the user's requirements for the transition effect. For example, if the user hopes to feel a smooth and natural interface transition when the emotion changes, a more balanced weight distribution can be adopted. For example, a relatively high weight (e.g., 0.7) is set for the current deformation parameter in the initial stage, while the weight of the target deformation parameter is slightly lower (e.g., 0.3). Then, as the transition process progresses, the former is gradually reduced and the latter is increased until it completely transitions to the target deformation parameter, giving the user more time to adapt to the new state.

[0099] In still other examples, the weights can also be set according to the user's requirements for emotional expressiveness. For example, if the user needs to emphasize a certain emotional state (such as "anger" or "joy"), the expressiveness of this emotion can be enhanced by adjusting the weights. For example, when expressing "anger", in order to achieve a strong visual impact effect more quickly, the weight of the target deformation parameter can be set to the maximum value (e.g., 1), ignoring the influence of the current deformation parameter, so as to achieve a rapid and strong change. On the contrary, for some more delicate emotions (such as "calm"), the weight of the target deformation parameter can be appropriately reduced to make the transition slower and more harmonious.

[0100] In still other examples, the weights can also be set according to the user's personalized needs. Specifically, the weights can be customized according to the preferences or historical data of different users. For example, some users may prefer quick and direct emotional feedback. At this time, a higher weight of the target deformation parameter can be configured for such users; while for users who like gentle changes, a more balanced weight setting is adopted. In this way, individual differences can be better met and a personalized user experience can be provided.

[0101] Step b2: Based on the adjustment step value of the deformation parameter, adjust the current deformation parameter of the voice input interface to gradually change to the target deformation parameter.

[0102] Based on the deformation parameter adjustment step value calculated in step b1, gradually iteratively update the current deformation parameter of the voice input interface, so that it smoothly transitions from the current state to the target deformation parameter. Specifically, in each update cycle, the current deformation parameter is finely adjusted according to the set step value, such as increasing or decreasing a certain scaling ratio, bending degree, control displacement amount, etc., until the preset target deformation effect is achieved.

[0103] Such asFigure 5 As shown, taking interface scaling as an example, if the current interface is at its normal size (scaling ratio is 1.0) and the goal is to zoom in to 1.2 times based on the user's emotion recognition result, then in each frame or at a fixed time interval, gradually increase the scaling ratio according to the calculated step value, such as 0.05 times, instead of an instantaneous jump, thus avoiding a visual abruptness.

[0104] In addition, easing functions (such as linear interpolation, ease-in-out function) can be combined to further optimize the transition curve, making the interface change more natural and smooth.

[0105] In this embodiment, by adjusting the current color parameters of the voice input interface to gradually change to the target color parameters, and / or adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters, not only can the accuracy of the interface's response to emotion recognition results be improved, but also the immersion and emotional resonance experience during user interaction can be enhanced. Compared with directly switching to a new visual setting, by gradually adjusting the visual parameters of the current interface (such as color, particle velocity, interface deformation, and dynamic texture, etc.), making it gradually approach the target visual effect. This can not only ensure the coherence and smoothness during the visual conversion process, but also improve the overall comfort and immersion of the user experience. Even in the case of rapid emotional changes, users can still feel a natural and harmonious visual experience, enhancing the emotional resonance of human-computer interaction. Whether it is a subtle change or a significant fluctuation in emotion, it can be delicately and appropriately represented in the interface design.

[0106] To ensure a low-latency visual rendering effect, thus completing the whole process from voice input to visual feedback in a short time, when the target visual display effect includes the target particle velocity, then adjust the current visual display effect of the voice input interface to gradually change to the target visual display effect, and it also includes: determining the target number of pre-loaded particles based on the performance parameters of the display device; pre-loading the target number of particles through the WebGPU parallel computing framework, and rendering the target visual display effect based on the target particle velocity.

[0107] Among them, the performance parameters of the display device include the processing ability of the GPU, memory capacity, etc.

[0108] Among them, determining the target number of pre-loaded particles based on the performance parameters of the display device includes: determining the pre-loaded particle number that matches the display device based on the mapping relationship between the performance parameters of the display device and the pre-loaded particle number.

[0109] In some examples, the mapping relationship between the performance parameters of the display device and the pre-loaded particle number can be a one-to-one relationship. For example, different performance parameters correspond to different pre-loaded particle numbers, and the pre-loaded particle numbers corresponding to each performance parameter are different.

[0110] In some examples, the mapping relationship between the performance parameters of the display device and the number of pre-loaded particles can also be a many-to-one relationship. For example, the performance parameters can be divided into multiple interval ranges, and the corresponding number of pre-loaded particles can be set for each interval range.

[0111] For example, the performance parameters can be divided into the first performance parameter interval range, the second performance parameter interval range, and the third performance parameter interval range; and the first performance parameter interval range corresponds to high-end devices, the second performance parameter interval range corresponds to mid-end devices, and the third performance parameter interval range corresponds to low-end devices; the number of pre-loaded particles corresponding to high-end devices, mid-end devices, and low-end devices decreases in sequence, such as 2000, 800, 500.

[0112] In this embodiment, by automatically adjusting the number of particles according to the actual performance of the display device, it is ensured that the best visual rendering effect can be achieved under various hardware conditions, while maintaining low latency and high response speed.

[0113] In order to make the visual display effect closer to the actual preferences of users, in some embodiments, feedback data on the target visual display effect can be collected, and the above mapping relationship can be updated based on this feedback data.

[0114] Among them, the feedback data includes the interaction behavior data of users on the target visual display effect; based on the feedback data, updating the mapping relationship includes: based on the interaction behavior data of users on the target visual display effect, updating the mapping relationship to make the mapping relationship conform to the preferences of users.

[0115] Among them, the interaction behavior data includes manually adjusting the size of the voice input interface (such as shrinking, enlarging, etc.), frequently switching dynamic texture effects, the usage duration of the target visual display effect, etc.

[0116] After presenting the target visual display effect to the user, if the user is not satisfied with the visual display effect in a certain emotional state, the visual display effect in that emotional state can be manually adjusted. By recording the interaction behavior data of the user, the above mapping relationship can be further optimized to make it more in line with the personal preferences of the user.

[0117] In the embodiment of the present application, the feature extraction part takes 25 ms, the emotion classification and reasoning part takes 35 ms, and the visual rendering part takes 40 ms, and the overall time consumption is less than or equal to 100 ms. In this way, it can ensure instant response for users and provide a smooth and seamless interaction experience. Whether in the application scenario of real-time speech emotion analysis or in an interactive media environment that requires rapid feedback, it is crucial for improving the user experience.

[0118] Exemplary device Corresponding to the above method for processing voice input, an embodiment of the present application further provides a device for processing voice input. Figure 6 It is a schematic structural diagram of a device for processing voice input provided by an embodiment of the present application. As Figure 6 shown, the device for processing voice input provided by an embodiment of the present application includes: an identification unit 601, a determination unit 602, and a presentation unit 603; wherein, the identification unit 601 identifies the emotional state when the user inputs a voice signal; the determination unit 602 is configured to determine a target visual display effect matching the emotional state of the user based on a preset mapping relationship between the emotional state and the visual display effect, and the target visual display effect is used to visually display the emotional state of the user; the presentation unit 603 is configured to present the target visual display effect in the voice input interface corresponding to the voice signal.

[0119] In some embodiments, the visual display effect includes at least one of color parameters in the HSL color space, the speed of particles included in the waveform of the voice signal, the deformation parameters of the voice input interface, and dynamic texture parameters.

[0120] In some embodiments, the presentation unit 603 presenting the target visual display effect in the voice input interface corresponding to the voice signal includes: adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect.

[0121] In some embodiments, when the target visual display effect includes target color parameters and / or target deformation parameters of the voice input interface; wherein, the presentation unit 603 adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect includes: adjusting the current color parameters of the voice input interface to gradually change to the target color parameters; and / or, adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

[0122] In some embodiments, the presentation unit 603 adjusting the current color parameters of the voice input interface to gradually change to the target color parameters includes: determining an adjustment step value of the color parameters; and adjusting the current color parameters of the voice input interface to gradually change to the target color parameters based on the adjustment step value of the color parameters.

[0123] In some embodiments, the color parameters include hue parameters and / or saturation parameters; wherein, the rendering unit 603 determines the adjustment step value of the color parameters, including: determining the adjustment step value of the hue parameters based on the product of the hue parameter difference between the current hue parameter and the target hue parameter of the voice input interface and a preset adjustment ratio; and / or determining the adjustment step value of the saturation parameters based on a preset saturation adjustment value.

[0124] In some embodiments, the rendering unit 603 adjusts the current deformation parameter of the voice input interface to gradually change to the target deformation parameter, including: performing weighted summation on the current deformation parameter and the target deformation parameter based on the respective weights corresponding to the current deformation parameter and the target deformation parameter, and determining the weighted summation result as the adjustment step value of the deformation parameter; adjusting the current deformation parameter of the voice input interface to gradually change to the target deformation parameter based on the adjustment step value of the deformation parameter.

[0125] In some embodiments, when the target visual display effect includes a target particle velocity, the rendering unit 603 adjusts the current visual display effect of the voice input interface to gradually change to the target visual display effect, further including: determining the target number of particles to be pre-loaded based on the performance parameters of the display device; pre-loading the target number of particles through a WebGPU parallel computing framework, and rendering the target visual display effect based on the target particle velocity.

[0126] In some embodiments, the recognition unit 601 recognizes the emotional state when the user inputs a voice signal, including: obtaining input data in each modality, where the input data in each modality includes at least one of the voice signal, the text data corresponding to the voice signal, the video data when the user inputs the voice signal, and the heart rate data; for each modality among the various modalities, extracting features from the input data in that modality to obtain a feature representation in that modality; obtaining a comprehensive feature based on the feature representations in the various modalities; and recognizing the emotional state when the user inputs the voice signal according to the comprehensive feature.

[0127] In some embodiments, the device further includes: an update unit 604, configured to obtain feedback data of the user on the target visual display effect; and update the mapping relationship based on the feedback data.

[0128] In some embodiments, the feedback data includes interaction behavior data of the user with respect to the target visual display effect; wherein, the updating unit 604 updates the mapping relationship based on the feedback data, including: updating the mapping relationship based on the interaction behavior data of the user with respect to the target visual display effect, so that the mapping relationship conforms to the user's preference.

[0129] The processing device for voice input provided in this embodiment belongs to the same inventive concept as the method for processing voice input provided in the foregoing embodiments of the present application, and can execute the method for processing voice input provided in any of the foregoing embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the method for processing voice input. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the method for processing voice input provided in the foregoing embodiments of the present application, and details will not be repeated here.

[0130] The functions implemented by the above recognition unit 601, determination unit 602, presentation unit 603, and updating unit 604 may be implemented by the same or different processors, which is not limited in the embodiments of the present application.

[0131] It should be understood that the units in the above device may be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor may be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory may be a memory inside the device or a memory outside the device. Alternatively, the units in the device may be implemented in the form of a hardware circuit, and the functions of some or all of the units may be implemented through the design of the hardware circuit. The hardware circuit may be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented through the design of the logical relationship of the components in the circuit; again, in another implementation, the hardware circuit may be implemented by a PLD. Taking an FPGA as an example, it may include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file, so as to implement the functions of some or all of the above units. All units of the above device may be implemented entirely in the form of a processor calling software, or entirely in the form of a hardware circuit, or partially in the form of a processor calling software, and the remaining part in the form of a hardware circuit.

[0132] In the embodiments of the present application, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and running capabilities, such as a CPU, microprocessor, GPU, or DSP, etc.; in another implementation, the processor can implement certain functions through the logical relationship of hardware circuits, and the logical relationship of the hardware circuits is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a type of ASIC, such as an NPU, TPU, DPU, etc.

[0133] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0134] In addition, each unit in the above device can be integrated in whole or in part, or can be independently implemented. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC can include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The types of the at least one processor can be different, such as including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0135] Exemplary electronic device Embodiments of the present application propose an electronic device. Refer to Figure 7 As shown, the electronic device includes: A memory 200 and a processor 210; Wherein, the memory 200 is connected to the processor 210 and is used to store programs; The processor 210 is used to implement the voice input processing method disclosed in any of the above embodiments by running the programs stored in the memory 200.

[0136] Specifically, the above electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0137] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them: The bus may include a path for transmitting information between various components of the computer system.

[0138] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or may be an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0139] The processor 210 may include a main processor, and may also include a baseband chip, a modem, etc.

[0140] The memory 200 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, etc.

[0141] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, etc.

[0142] The output device 240 may include a device for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0143] The communication interface 220 may include a device of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0144] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement each step of any one of the voice input processing methods provided in the above embodiments of the present application.

[0145] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored on a memory through the data interface to execute the voice input processing method introduced in any of the above embodiments. For the specific processing process and its beneficial effects, reference may be made to the embodiments of the voice input processing method described above.

[0146] Exemplary Computer Program Product and Storage Medium In addition to the above methods and devices, embodiments of the present application may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the method for processing voice input according to various embodiments of the present application described in any of the above embodiments of this specification.

[0147] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0148] In addition, embodiments of the present application may also be a storage medium on which a computer program is stored. The computer program is executed by a processor to perform the steps in the method for processing voice input according to various embodiments of the present application described in any of the above embodiments of this specification, and specifically may implement the steps of the above methods.

[0149] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps may be in other sequences or performed simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0150] It should be noted that the various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments may be referred to each other. For device embodiments, since they are basically similar to method embodiments, they are described relatively simply, and the relevant parts may refer to the partial description of the method embodiments.

[0151] The steps in the methods of the various embodiments of the present application may be adjusted, combined, and deleted according to actual needs. The technical features recorded in the various embodiments may be replaced or combined.

[0152] The modules and sub-modules in the devices and terminals in the various embodiments of the present application may be combined, divided, and deleted according to actual needs.

[0153] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or modules, and can be in electrical, mechanical, or other forms.

[0154] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] In addition, in each embodiment of the present application, each functional module or sub-module can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.

[0156] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0157] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0158] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0159] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing voice input, characterized in that, including: identifying the emotional state when the user inputs a voice signal; determining a target visual display effect matching the emotional state of the user based on a preset mapping relationship between the emotional state and the visual display effect, where the target visual display effect is used to visually display the emotional state of the user; presenting the target visual display effect in the voice input interface corresponding to the voice signal.

2. The method according to claim 1, wherein The visual display effect includes at least one of color parameters in the HSL color space, the speed of particles included in the waveform of the voice signal, the deformation parameters of the voice input interface, and dynamic texture parameters.

3. The method according to claim 2, wherein Presenting the target visual display effect in the voice input interface corresponding to the voice signal includes: adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect.

4. The method according to claim 3, wherein In the case where the target visual display effect includes target color parameters and / or target deformation parameters of the voice input interface; wherein, adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect includes: adjusting the current color parameters of the voice input interface to gradually change to the target color parameters; and / or, adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

5. The method according to claim 4, wherein Adjusting the current color parameters of the voice input interface to gradually change to the target color parameters includes: determining an adjustment step value of the color parameters; based on the adjustment step value of the color parameters, adjusting the current color parameters of the voice input interface to gradually change to the target color parameters.

6. The method according to claim 5, wherein The color parameters include hue parameters and / or saturation parameters; wherein, determining the adjustment step value of the color parameters includes: determining the adjustment step value of the hue parameters based on the product of the hue parameter difference between the current hue parameter and the target hue parameter of the voice input interface and a preset adjustment ratio; and / or, determining the adjustment step value of the saturation parameters based on a preset saturation adjustment value.

7. The method according to claim 4, characterized in that, Adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters includes: performing weighted summation on the current deformation parameters and the target deformation parameters based on the respective weights corresponding to the current deformation parameters and the target deformation parameters, and determining the weighted summation result as the adjustment step value of the deformation parameters; based on the adjustment step value of the deformation parameters, adjusting the current deformation parameters of the voice input interface to gradually change to the target deformation parameters.

8. The method according to claim 4, characterized in that When the target visual display effect includes a target particle speed, adjusting the current visual display effect of the voice input interface to gradually change to the target visual display effect further includes: determining a target number of pre-loaded particles based on the performance parameters of the display device; pre-loading the target number of particles through a WebGPU parallel computing framework, and rendering the target visual display effect based on the target particle speed.

9. The method according to any one of claims 1-8, characterized in that Identifying the emotional state when the user inputs a voice signal includes: Obtain the input data in each modality, where the input data in each modality includes at least one of the voice signal, the text data corresponding to the voice signal, the video data when the user inputs the voice signal, and the heart rate data; For each modality among the various modalities, perform feature extraction on the input data in that modality to obtain the feature representation in that modality; Based on the feature representations in the various modalities, obtain the comprehensive feature; According to the comprehensive feature, identify the emotional state when the user inputs the voice signal.

10. The method according to any one of claims 1-8, characterized in that, The method further includes: Obtain the feedback data of the user on the target visual display effect; Based on the feedback data, update the mapping relationship.

11. The method according to claim 10, characterized in that, The feedback data includes the interaction behavior data of the user on the target visual display effect; Wherein, the updating of the mapping relationship based on the feedback data includes: Based on the interaction behavior data of the user on the target visual display effect, update the mapping relationship so that the mapping relationship conforms to the user's preference.

12. An electronic device, characterized in that, Comprising a memory and a processor; The memory is connected to the processor and is used for storing programs; The processor is used for implementing the method according to any one of claims 1 to 11 by running the program in the memory.

13. A storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer program product, characterized in that, Comprising computer program instructions, and when the computer program instructions are run by the processor, the processor is caused to implement the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and apparatus for image display control according to viewer factors and responses

    CN102473264A

  • Method and device for adjusting display state of electronic equipment, and storage medium

    CN109688264A

  • Voice-based animation display method and device, computer device and storage medium

    CN110379430A

  • Message display method and device, terminal and computer readable storage medium

    CN111106995A

  • Visual representation method and device of voice emotion and computer storage medium

    CN112037821A

Cited By

  • Multimedia file generation method and device

    CN121547660A