Intelligent sound box response method and system based on multi-dimensional state perception

By using multi-dimensional state perception technology, smart speakers collect voice input signals to extract user identity, emotion, and environmental features, generate a comprehensive state vector, and match the optimal strategy in the response strategy library. This solves the problem of rigid response in existing smart speakers and enables flexible and personalized interaction.

CN120998210APending Publication Date: 2025-11-21SHENZHEN AO LAI DE ELECTRONIC CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511186307.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-23
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing smart speaker response mechanisms lack the ability to perceive multi-dimensional state characteristics such as the user's individual identity, real-time emotional state, and environment, resulting in rigid response methods that cannot be dynamically adjusted according to specific scenarios, thus limiting the user experience.

Method used

By collecting voice input signals for multi-dimensional state perception, including user identification, emotional state quantification, and environmental state feature extraction, a comprehensive state vector is generated. Fuzzy matching is then performed in the response strategy knowledge base to obtain the optimal response strategy and adjust the smart speaker's response content and output parameters.

Benefits of technology

It achieves comprehensive perception of individual user differences and environmental conditions, improves response accuracy and decision-making efficiency, breaks through the limitations of traditional fixed rules, enhances the system's generalization and adaptability, and improves user interaction comfort and personalized experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998210A_ABST
    Figure CN120998210A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent sound box response method and system based on multi-dimensional state perception, and the method comprises the steps: receiving a voice input signal of a user, collecting multi-dimensional state information based on the voice input signal, carrying out the multi-dimensional state perception processing of the multi-dimensional state information, and obtaining structured state data, the multi-dimensional state perception processing comprises user identity recognition, emotional state quantification and environment state feature extraction; fusing the structured state data to generate a comprehensive state vector; based on the comprehensive state vector, executing a fuzzy matching operation in a preset response strategy knowledge base to obtain an optimal response strategy; and according to the optimal response strategy, adjusting response content and output parameters of the intelligent sound box, and completing personalized dynamic voice interaction response. The application has the effect of improving the response flexibility of the intelligent loudspeaker box.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent voice interaction, and in particular to a smart speaker response method and system based on multi-dimensional state perception. Background Technology

[0002] Currently, smart speakers are widely used in homes, offices, and other settings. Their main functions include voice recognition, task execution, information broadcasting, and entertainment control. Users can trigger the speaker to perform operations such as playing music, checking the weather, and controlling home appliances through voice commands, greatly improving the convenience of human-computer interaction.

[0003] Existing smart speaker response mechanisms are mostly based on "command triggering → task execution," relying primarily on voice recognition and command parsing technologies. They lack the ability to perceive the user's individual identity, real-time emotional state, and surrounding environment. Some products have introduced simple emotion recognition functions, roughly distinguishing emotion tags such as "happy" and "angry," but these are often limited to single-dimensional classification, lacking quantitative analysis of emotion intensity and tendency. Simultaneously, environmental conditions are typically simplistically categorized as "quiet" or "noisy," lacking detailed processing such as background sound type recognition and sound source signal-to-noise ratio analysis, resulting in the system's inability to accurately adapt to complex scenarios. Furthermore, even when the system acquires some state information, its response strategy remains fixed, unable to dynamically adjust based on the user's overall state, significantly limiting the user experience.

[0004] The existing technical solutions mentioned above have the following drawbacks: existing smart speakers often process tasks based solely on the voice content itself during the response process, ignoring the multi-dimensional state characteristics behind the voice. The response method is rigid and cannot adjust the output parameters or tone of voice according to specific scenarios, so there is room for improvement. Summary of the Invention

[0005] To improve the responsiveness of smart speakers, this application provides a smart speaker response method and system based on multi-dimensional state perception.

[0006] The above-mentioned objective of this application is achieved through the following technical solution: A smart speaker response method based on multi-dimensional state perception, the method comprising: The system receives the user's voice input signal, collects multi-dimensional state information based on the voice input signal, performs multi-dimensional state perception processing on the multi-dimensional state information, and obtains structured state data. The multi-dimensional state perception processing includes user identity recognition, emotional state quantification, and environmental state feature extraction. The structured state data is fused to generate a comprehensive state vector; Based on the comprehensive state vector, a fuzzy matching operation is performed in the preset response strategy knowledge base to obtain the optimal response strategy; Based on the optimal response strategy, the response content and output parameters of the smart speaker are adjusted to achieve personalized dynamic voice interaction response.

[0007] By adopting the above technical solutions, and collecting multi-dimensional state information based on voice input signals to perform user identification, emotional state quantification, and environmental state feature extraction, a comprehensive perception of individual user differences, subjective emotions, and current environmental states can be achieved, thus providing rich contextual information support for smart speakers. By fusing structured state data to generate a comprehensive state vector, multi-dimensional states can be expressed with a unified data structure, effectively avoiding inconsistencies in responses caused by information fragmentation, thereby improving decision-making efficiency and response accuracy. By performing fuzzy matching operations in the response strategy knowledge base to obtain the optimal response strategy, the limitations of traditional fixed rules can be overcome, enabling dynamic adaptation to novel state combinations, thereby improving the system's generalization and responsiveness. By adjusting the response content and output parameters according to the optimal response strategy, flexible adjustments can be made in multiple dimensions such as tone, content, and volume, thereby enhancing user interaction comfort and personalized experience.

[0008] In one example, this application can be further configured as follows: the step of collecting multi-dimensional state information based on the voice input signal and performing user identification on the multi-dimensional state information includes: Based on the voiceprint recognition algorithm, individual feature parameters for constructing the multidimensional state information are extracted from the voice input signal, and the individual feature parameters are compared with a preset user identity database to obtain the user identity identifier; When the comparison result cannot match a registered user, the current identity is marked as an unfamiliar user, and the corresponding voice data is recorded for subsequent identity modeling.

[0009] By adopting the above technical solution, the system can accurately identify the current speaker's identity by extracting individual features from the speech signal based on the voiceprint recognition algorithm and comparing them with the user identity database, thereby realizing the differentiated adaptation of the response strategy to the user role. By marking unmatched identities as unfamiliar users and recording the speech data, the system can support subsequent identity modeling and expansion, thereby enhancing the system's adaptability and continuous learning ability in new user access scenarios.

[0010] In one example, this application can be further configured as follows: the step of collecting multi-dimensional state information based on the voice input signal and quantifying the emotional state of the multi-dimensional state information includes: Acoustic features are extracted from the speech input signal, and the acoustic feature parameters include pitch, speech rate, energy, and zero-crossing rate. Based on the acoustic feature parameters, the current user's emotion is mapped to a two-dimensional emotion coordinate using a preset emotion mapping model. The X-axis of the two-dimensional emotion coordinate represents the degree of excitement, and the Y-axis represents the emotional tendency.

[0011] By adopting the above technical solution, and by extracting acoustic features such as pitch, speech rate, energy, and zero-crossing rate from speech signals and constructing two-dimensional emotion coordinates, it is possible to continuously quantify and express emotional states. Compared with traditional emotion classification, it has higher resolution and modeling accuracy, thereby improving the system's ability to recognize subtle changes in user emotions and achieving more detailed and humanized response strategy matching.

[0012] In one example, this application can be further configured as follows: the step of collecting multi-dimensional state information based on the voice input signal and extracting environmental state features from the multi-dimensional state information includes: Background audio data during the voice input process is collected using a microphone array; Environmental characteristic parameters are calculated based on the background audio data, including the average background noise decibel value and the human voice signal-to-noise ratio. The background audio data is input into an audio classification model to identify background sound type labels. The background sound type labels are used to indicate the sound source attributes of the current environment. The sound source attributes include quiet, music, television sound, and multi-person conversation. The average background noise decibel value, human voice signal-to-noise ratio, and background sound type label are encoded and combined to form an environmental state vector.

[0013] By adopting the above technical solutions, background noise is collected through a microphone array, and the average noise decibel value and human voice signal-to-noise ratio are calculated. This objectively reflects the noise level and speech clarity of the current environment, thus providing a quantitative basis for adjusting output parameters. By identifying background noise type labels through an audio classification model, the typical sound source attributes of the current scene can be determined, thus providing contextual semantics for response strategy matching. By encoding multiple environmental parameters into environmental state vectors, structured modeling of environmental states can be achieved, thereby improving the system's ability to express and adapt to diverse scenarios.

[0014] In one example, this application can be further configured as follows: Based on the comprehensive state vector, performing a fuzzy matching operation in a preset response strategy knowledge base to obtain the optimal response strategy includes: The similarity between the comprehensive state vector and each preset state template in the response strategy knowledge base is calculated using a nearest neighbor search algorithm based on Euclidean distance to obtain the calculation results; Based on the calculation results, the response strategy corresponding to the state template with the smallest distance from the current comprehensive state vector is selected from the state templates and is taken as the optimal response strategy. When all similarities are below the preset matching threshold, the default response strategy is invoked, and the current comprehensive state vector is recorded in the strategy learning log.

[0015] By adopting the above technical solutions, and using the Euclidean distance nearest neighbor search algorithm to calculate the similarity between the comprehensive state vector and the policy template, the matching degree between the current state and the historical policy can be accurately evaluated, thereby selecting the closest response policy to improve the rationality of the response. By selecting the policy corresponding to the minimum distance as the optimal response scheme, the most suitable one can be selected from multiple candidate policies, thereby avoiding interaction errors caused by unrelated policies being triggered by mistake. By calling the default policy and recording the state data when the similarity is lower than the matching threshold, the system can still have stable response capabilities in unknown states. At the same time, the accumulated data can be used for knowledge base optimization, thereby improving the system's self-learning ability and long-term evolution.

[0016] In one example, this application can be further configured such that adjusting the response content and output parameters of the smart speaker according to the optimal response strategy includes: Based on the optimal response strategy, a content type matching the current user's identity is selected, and the content type includes adult news broadcasts, children's stories, or operation prompts. The speech synthesis parameters are adjusted according to the optimal response strategy. The speech synthesis parameters include intonation type, speech rate level and timbre emotion curve to match the user's current emotional state. Adjusting the volume output parameters according to the optimal response strategy, the adjustment includes speaking at a low speed in noisy environments, simplifying the speech content, or activating a low volume mode.

[0017] By adopting the above technical solutions, and by selecting content types that match the user's identity according to the optimal response strategy, precise distribution of content resources can be achieved, thereby improving the targeting of personalized services. By adjusting the speech synthesis parameters, including tone, speed, and timbre emotion curve, the speech output can be made more in line with the current user's emotional state, thereby enhancing the emotional resonance and acceptability of voice interaction. By using low-speed broadcasting, simplifying content, or activating low-volume mode in noisy environments, communication barriers caused by environmental interference can be effectively suppressed, thereby ensuring the comprehensibility and environmental adaptability of information transmission.

[0018] In one example, this application can be further configured as follows: the smart speaker response method based on multi-dimensional state awareness also includes: After executing the personalized dynamic voice interaction response, a strategy execution record is generated based on the current comprehensive state vector and response result, and the strategy execution record is stored in the strategy learning log. The strategy execution record includes the comprehensive state vector, the selected optimal response strategy identifier, and user interaction feedback information. When the cumulative number of strategy execution records reaches a preset learning threshold, a partial update of the response strategy knowledge base is triggered to improve the matching accuracy and adaptability of subsequent response strategies.

[0019] By adopting the above technical solution, after completing the response, a strategy execution record containing a comprehensive state vector, the optimal strategy identifier, and user feedback can be generated and stored. This allows for the accumulation of state and response relationship data during the interaction process, thus providing data support for knowledge base updates. By triggering a local update when the number of strategy execution records reaches a preset threshold, the knowledge base content can be dynamically adjusted to match changes in user behavior, thereby improving the timeliness and accuracy of strategy matching and enhancing the system's adaptability and interactive performance in long-term operation.

[0020] The second objective of this invention is achieved through the following technical solution: A smart speaker response system based on multi-dimensional state perception, the smart speaker response system based on multi-dimensional state perception includes: The state perception processing module is used to receive the user's voice input signal, collect multi-dimensional state information based on the voice input signal, perform multi-dimensional state perception processing on the multi-dimensional state information to obtain structured state data. The multi-dimensional state perception processing includes user identity recognition, emotional state quantification, and environmental state feature extraction. The state fusion generation module is used to fuse the structured state data to generate a comprehensive state vector; The strategy matching module is used to perform fuzzy matching operations in a preset response strategy knowledge base based on the comprehensive state vector to obtain the optimal response strategy. The response control module is used to adjust the response content and output parameters of the smart speaker according to the optimal response strategy, so as to complete the personalized dynamic voice interaction response.

[0021] By adopting the above technical solutions, and collecting multi-dimensional state information based on voice input signals to perform user identification, emotional state quantification, and environmental state feature extraction, a comprehensive perception of individual user differences, subjective emotions, and current environmental states can be achieved, thus providing rich contextual information support for smart speakers. By fusing structured state data to generate a comprehensive state vector, multi-dimensional states can be expressed with a unified data structure, effectively avoiding inconsistencies in responses caused by information fragmentation, thereby improving decision-making efficiency and response accuracy. By performing fuzzy matching operations in the response strategy knowledge base to obtain the optimal response strategy, the limitations of traditional fixed rules can be overcome, enabling dynamic adaptation to novel state combinations, thereby improving the system's generalization and responsiveness. By adjusting the response content and output parameters according to the optimal response strategy, flexible adjustments can be made in multiple dimensions such as tone, content, and volume, thereby enhancing user interaction comfort and personalized experience.

[0022] In summary, this application includes the following beneficial technical effects: 1. By collecting multi-dimensional state information based on voice input signals and performing user identification, emotional state quantification, and environmental state feature extraction, it is possible to achieve a comprehensive perception of individual user differences, subjective emotions, and current environmental state, thereby providing rich contextual information support for smart speakers; by fusing structured state data to generate a comprehensive state vector, it is possible to express multi-dimensional states with a unified data structure, effectively avoiding inconsistent responses caused by information fragmentation, thereby improving decision-making efficiency and response accuracy; 2. By performing fuzzy matching operations in the response strategy knowledge base to obtain the optimal response strategy, the limitations of traditional fixed rules can be broken, and dynamic adaptation to novel state combinations can be achieved, thereby improving the system's generalization and responsiveness. By adjusting the response content and output parameters according to the optimal response strategy, flexible adjustments can be made in multiple dimensions such as tone, content, and volume, thereby enhancing the user's interactive comfort and personalized experience. Attached Figure Description

[0023] Figure 1 This is a flowchart of a smart speaker response method based on multi-dimensional state perception in one embodiment of this application; Figure 2 This is a flowchart illustrating the implementation of step S10 in a smart speaker response method based on multi-dimensional state perception in one embodiment of this application. Figure 3 This is another implementation flowchart of step S10 in a smart speaker response method based on multi-dimensional state perception in one embodiment of this application; Figure 4 This is another implementation flowchart of step S10 in a smart speaker response method based on multi-dimensional state perception in one embodiment of this application; Figure 5 This is a flowchart illustrating the implementation of step S30 in a smart speaker response method based on multi-dimensional state perception in one embodiment of this application. Figure 6 This is a flowchart illustrating the implementation of step S40 in a smart speaker response method based on multi-dimensional state perception in one embodiment of this application. Figure 7 This is a flowchart of an implementation of a smart speaker response method based on multi-dimensional state perception in one embodiment of this application; Figure 8 This is a principle block diagram of a smart speaker response system based on multi-dimensional state perception in one embodiment of this application. Detailed Implementation

[0024] The present application will be further described in detail below with reference to the accompanying drawings.

[0025] In one embodiment, such as Figure 1 As shown, this application discloses a smart speaker response method based on multi-dimensional state perception, which specifically includes the following steps: S10: Receives the user's voice input signal, collects multi-dimensional state information based on the voice input signal, performs multi-dimensional state perception processing on the multi-dimensional state information, and obtains structured state data. The multi-dimensional state perception processing includes user identity recognition, emotional state quantification, and environmental state feature extraction.

[0026] Specifically, after receiving the user's voice input signal, the execution action collects multiple dimensions of state information in parallel based on the voice input signal, including identity clues, emotional features, and background environment parameters in the audio stream. Among them, the user's identity can be extracted by preliminarily analyzing the voiceprint pattern in the voice signal to extract individual feature codes. The emotional state is initially estimated by calculating the dynamic feature indicators of the voice in the time and frequency domains. The environmental state is constructed based on the overall energy of the audio received by the microphone, the background noise level, and the speech intelligibility index to build a basic parameter set. Then, the collected multi-dimensional state information is input into the multi-dimensional state perception processing flow, and operations such as identity association, coarse classification of emotional attributes, and extraction of environmental semantic features are performed in sequence. Finally, structured state data including identity information, emotion estimation value, and environmental representation results are constructed to support the execution logic of subsequent fusion modeling and response decision-making. For example, in a certain scenario, the speaker is identified as family member A, whose tone is fast, volume is high, and speech rate is fast. The emotional state is estimated to be "excited". At the same time, the environmental background presents a continuous low noise state and is identified as a "quiet" scene. Thus, a set of structured state data is obtained for subsequent processing.

[0027] S20: Fuse the structured state data to generate a comprehensive state vector.

[0028] Specifically, after generating the structured state data, the action performs semantic alignment and feature normalization on each element in the structured state data. By setting a unified data template, user identity attributes, emotion estimation parameters, and environmental representation variables are encoded into fixed positions and normalized and standardized to make them comparable and matching in numerical space. Then, the processed multiple fields are concatenated into a unified multidimensional state vector representation and a timestamp is added to mark the state snapshot of the current interaction scenario. Finally, a comprehensive state vector is formed to express the complete state model of the user's current interaction situation. For example, the user ID field is encoded as 001, the emotion coordinate is (0.8, -0.4), and the environmental state vector is [35dB, 12dB, encoding 03], where encoding 03 represents "TV background sound". The final generated CSV is [001, 0.8, -0.4, 35, 12, 03, 202507311427].

[0029] S30: Based on the comprehensive state vector, perform fuzzy matching operation in the preset response strategy knowledge base to obtain the optimal response strategy.

[0030] Specifically, after obtaining the comprehensive state vector, the action is to input the vector into the matching module of the response strategy knowledge base. The difference between the current state and the historical state template is evaluated through a preset vector similarity calculation mechanism. A distance ranking list between state vectors is constructed using a feature distance metric method. The closest candidate strategy vector is searched in turn and a matching score mapping relationship is established. In actual execution, multiple fuzzy intervals can be set to divide the applicability of different levels of response strategies, and then the response strategy with the best score is selected as the output scheme of the current interaction scenario. For example, when the emotion coordinates of the identified CSV are close to (0.9, -0.7) and the background sound label is "multi-person conversation", the similarity with the strategy template "child users, excited emotions, multi-sound source environment" in the knowledge base is the lowest. The system will select this strategy and give feedback such as "use a gentle tone to guide the content and automatically reduce the volume by 20%". If the similarity between the current state and all templates is lower than the system threshold of 0.3, the general strategy template is returned and the CSV is marked for subsequent iterative learning.

[0031] S40: Based on the optimal response strategy, adjust the smart speaker's response content and output parameters to complete personalized dynamic voice interaction responses.

[0032] Specifically, after selecting the optimal response strategy, the execution action applies the parameters configured in the strategy to the execution flow of the current voice interaction task. First, based on the strategy content, the response resource type suitable for the current user is selected, including audio replies, command feedback, or contextual prompts. Then, the tone, speech rate, and timbre are adjusted according to the speech synthesis parameters defined by the strategy. For example, a calm tone is used to reply to emotionally agitated users, and the speech rate is automatically reduced to enhance semantic clarity in noisy environments. Finally, the volume level is configured according to the output intensity parameters set by the strategy to ensure consistent and comfortable voice output performance in different user states and scenarios. For example, when it is recognized that family member B is in a low mood and the background is a quiet nighttime environment, the output strategy is "to use a soft tone to broadcast a short phrase prompt and limit the volume to 30%". When it is recognized that a stranger is in a positive mood and the background is a conversation among multiple people, the output is "to use a standard speech rate to broadcast a short message and activate the enhanced clarity speech mode", thus completing a round of personalized, state-aware, dynamic voice interaction response process.

[0033] By adopting the above technical solutions, and collecting multi-dimensional state information based on voice input signals to perform user identification, emotional state quantification, and environmental state feature extraction, a comprehensive perception of individual user differences, subjective emotions, and current environmental states can be achieved, thus providing rich contextual information support for smart speakers. By fusing structured state data to generate a comprehensive state vector, multi-dimensional states can be expressed with a unified data structure, effectively avoiding inconsistencies in responses caused by information fragmentation, thereby improving decision-making efficiency and response accuracy. By performing fuzzy matching operations in the response strategy knowledge base to obtain the optimal response strategy, the limitations of traditional fixed rules can be overcome, enabling dynamic adaptation to novel state combinations, thereby improving the system's generalization and responsiveness. By adjusting the response content and output parameters according to the optimal response strategy, flexible adjustments can be made in multiple dimensions such as tone, content, and volume, thereby enhancing user interaction comfort and personalized experience.

[0034] In one embodiment, such as Figure 2 As shown, in step S10, which involves collecting multi-dimensional state information based on the voice input signal and performing user identification on the multi-dimensional state information, the specific steps include: S11: Based on the voiceprint recognition algorithm, extract individual feature parameters from the voice input signal to construct multi-dimensional state information, and compare the individual feature parameters with the preset user identity database to obtain the user identity identifier.

[0035] Specifically, during the user identification process, a set of acoustic feature parameters for individual modeling is extracted from the current voice input signal based on the voiceprint recognition algorithm. These parameters include multiple dimensions such as short-time energy distribution, formant frequency, speaking speed, and pronunciation pattern. After feature encoding, a set of voiceprint vector representations is formed. Then, the voiceprint vector is compared one by one with the registered user voiceprint models in the locally stored user identity database. The similarity score is calculated and the attribution is determined. Finally, the most matching set of identity tags is selected as the user identity identifier for output. For example, in a voice interaction, if the extracted voiceprint vector has a similarity of more than 95% with the template of user A in the database, it is identified as "user A". If the similarity is lower than a set threshold, it is not classified as a known identity for the time being.

[0036] S12: When the comparison result cannot match a registered user, mark the current identity as an unfamiliar user and record the corresponding voice data for subsequent identity modeling.

[0037] Specifically, after voiceprint comparison is completed, if the similarity calculation result fails to establish a valid match with any registered user identity, the action is to mark the current interaction initiator as an unfamiliar user and assign a set of temporary identity identifiers for subsequent interaction recognition. At the same time, the current voice input signal is cached and features are extracted. The voice data is stored in the user modeling pending queue in the form of raw audio and structured features, providing a corpus basis for subsequent identity registration or automatic learning modeling. For example, in a certain interaction, it is identified as "Visitor 001" and its voiceprint features are recorded for the system to determine whether it is a frequent interaction object, thereby triggering the automatic profile building or personalized suggestion process.

[0038] In one embodiment, such as Figure 3 As shown, in step S10, which involves collecting multi-dimensional state information based on the voice input signal and quantifying the emotional state of the multi-dimensional state information, the specific steps include: S13: Extract acoustic features from the speech input signal. Acoustic feature parameters include pitch, speech rate, energy, and zero-crossing rate.

[0039] Specifically, after receiving the voice input signal, the system performs short-time framing and windowing processing on the voice signal, and extracts acoustic feature parameters within each frame. These parameters include pitch as a fundamental frequency variation curve to reflect intonation fluctuations, speech rate to assess speaking rhythm by the number of phonemes per unit time, energy representing the amplitude intensity of each frame of speech to assess the level of emotional activity, and zero-crossing rate reflecting the frequency of signal changes to measure tension. These acoustic parameters are smoothed and statistically analyzed in the time and frequency domains, ultimately forming a parameter set containing feature dimensions such as mean, variance, and maximum value. For example, if a voice input is detected to have a high average pitch, high energy, fast speech rate, and frequent fluctuations in the zero-crossing rate, it is initially inferred to be a manifestation of emotional excitement, providing input basis for subsequent emotion modeling.

[0040] S14: Based on acoustic feature parameters, the current user's emotion is mapped to a two-dimensional emotion coordinate using a preset emotion mapping model. The X-axis of the two-dimensional emotion coordinate represents the degree of excitement, and the Y-axis represents the emotional tendency.

[0041] Specifically, after obtaining the acoustic feature parameter set, the action is executed to input the set of parameters into a preset emotion mapping model for mapping operations. The emotion mapping model learns and constructs a nonlinear mapping relationship between acoustic parameters and emotion labels through training samples, and outputs an emotion estimate represented by two-dimensional coordinates using regression or neural network methods. The X-axis represents the level of excitement from -1 to 1, with a larger value indicating greater excitement, and the Y-axis represents the emotional tendency from -1 to 1, with a higher value indicating a more positive attitude. After the model outputs the coordinates, it can be smoothly adjusted in combination with the current dialogue history or context state. For example, if a set of acoustic features is detected and the output coordinates are (0.85, -0.4), it indicates that the current user's emotional state is excited but leaning towards negative. This coordinate will be used as an important component of the subsequent comprehensive state vector in response strategy matching.

[0042] Furthermore, in the construction and training phase of the emotion mapping model, a training sample set containing multiple types of real user speech data and subjective emotion labels is first collected. The speech data covers various speaking styles, language rhythms, and emotional states. The subjective labels are annotated by multiple annotators using two-dimensional emotion coordinates based on emotion dimension standards, where the X-axis represents the level of arousal and the Y-axis represents the emotional tendency. Subsequently, the speech samples are processed by frame segmentation and acoustic feature parameters are extracted, including pitch, speech rate, energy, zero-crossing rate, harmonic distortion ratio, etc., which constitute the input feature vector set of the model. A dual-channel regression neural network structure is adopted as the modeling framework. The forward channel is used to learn the prediction of the level of arousal (X-axis), and the backward channel is used to learn the prediction of the emotional tendency (Y-axis). The mean square error of the two-dimensional coordinate output is optimized by a joint loss function. Dropout and Batch Normalization strategies are introduced during training to improve the generalization performance and stability of the model. Finally, a lightweight emotion mapping model for deployment is formed, which can quickly map continuous two-dimensional emotion coordinate values ​​after receiving new speech input features, ensuring that the response strategy has sufficient discriminative ability and response accuracy in different emotional states.

[0043] In one embodiment, such as Figure 4 As shown, in step S10, which involves acquiring multi-dimensional state information based on the voice input signal and extracting environmental state features from the multi-dimensional state information, the specific steps include: S15: Acquire background audio data during the voice input process using a microphone array.

[0044] Specifically, while performing voice input parsing, the system continuously collects raw audio signals during the input process through a microphone array. This audio data includes not only the voice content spoken by the user but also background sound information present in the environment. The array-type multi-channel microphone input can capture the directionality of the sound source, the degree of spatial diffusion, and the sound intensity distribution. The audio data is buffered in the acoustic analysis module for subsequent processing. For example, when the user issues the voice command "play the news," the system simultaneously captures the bass output of the TV in the room, the sound reflected from the walls, and the dominant voice track input directly in front of the user, forming a complete background audio data sample.

[0045] S16: Calculate environmental characteristic parameters based on background audio data. Environmental characteristic parameters include average background noise in decibels and human voice signal-to-noise ratio.

[0046] Specifically, after completing the background audio data acquisition, the system performs frame-level segmentation and feature statistics on the data. Based on the frame energy change, it calculates the average background noise decibel value of the entire segment to reflect the overall sound pressure level of the environment. At the same time, by separating the user's main speech track and non-speech track, it calculates the human voice signal-to-noise ratio index using the power ratio. This index is used to measure the relative intensity between the user's speech and the background noise, serving as a key reference for subsequent speech output clarity adjustment and response strategy selection. For example, in a certain input, the analysis shows that the average background noise is 42dB and the human voice signal-to-noise ratio is 9dB, indicating that the user's voice is still identifiable in a moderately noisy environment.

[0047] S17: Input the background audio data into the audio classification model to identify the background sound type label. The background sound type label is used to indicate the sound source attributes of the current environment. Sound source attributes include quiet, music, TV sound, and multi-person conversation.

[0048] Specifically, after extracting the basic environmental features, the system performs an action to input the collected background audio data into a preset audio classification model. This model is based on a convolutional neural network architecture and has been trained on samples from multiple home environment scenarios. It can identify the acoustic patterns of the main sound sources in the background noise and output corresponding label codes. The labels are used to indicate the main sound source attributes of the environment. Typical categories include "quiet", "continuous music", "TV playback", "multi-person conversation", etc. During the processing, a confidence threshold is set for each label to avoid misjudgment. For example, in a certain detection, the audio classification model outputs the label "multi-person conversation" with a confidence score of 0.91. Based on this, the system determines that the current environment is a complex speech interference scenario.

[0049] Furthermore, in the construction and training of the audio classification model, background audio data samples covering various typical home and office scenarios were first collected. These samples included various real-world usage scenarios such as quiet environments, background music playback, TV program clips, natural conversations among multiple people, and kitchen noise. Annotators labeled the audio sources according to their types, forming a training corpus. The audio samples were then preprocessed, including operations such as uniform sampling rate, frame-by-frame windowing, removal of silent segments, and amplitude normalization. Acoustic features such as Mel-frequency cepstral coefficients (MFCC), spectral centroid, zero-crossing rate, and short-time energy were extracted and used as the model's input feature vector. The corresponding sound source labels are then encoded into multi-class output labels. The training model adopts a multi-layer one-dimensional convolutional neural network (1D-CNN) architecture. Its input layer receives multi-dimensional time series features, and after several convolution-pooling-normalization combination layers, it is connected to a fully connected layer to output the sound source probability distribution. The loss function uses multi-class cross-entropy and introduces a class balance factor to solve the problem of imbalanced sample size. An early stopping mechanism is used during training to prevent overfitting. After training is completed, the deployment model is exported, which can quickly identify the main background sound type labels in the current environment in the real-time speech input processing process. As a key component of environmental state perception, it improves the scene adaptability of policy decision-making.

[0050] S18: Encode and combine the average background noise decibel value, human voice signal-to-noise ratio, and background sound type label to form an environmental state vector.

[0051] Specifically, after acquiring the background noise decibel value, human voice signal-to-noise ratio, and background sound type label, the action is executed to uniformly encode the three types of features. Fixed-point quantization is used to normalize continuous value parameters (such as noise value and signal-to-noise ratio) to a unified vector dimension range. At the same time, the sound source type label is mapped to discrete integer encoding. Then, the three types of data are concatenated in a predefined order to form an environment state vector. This vector serves as a structured representation describing the current interactive environment and will directly participate in the subsequent fusion of comprehensive state vectors and strategy matching process. For example, the combination [42,9,03] represents a scenario where the current environment has a moderate noise level, moderate human voice clarity, and the background sound is "multi-person conversation" encoded as 03, thus improving the adaptability of subsequent responses and environmental robustness.

[0052] In one embodiment, such as Figure 5 As shown, in step S30, based on the comprehensive state vector, a fuzzy matching operation is performed in the preset response strategy knowledge base to obtain the optimal response strategy, specifically including: S31: The nearest neighbor search algorithm based on Euclidean distance is used to calculate the similarity between the comprehensive state vector and each preset state template in the response strategy knowledge base, and the calculation results are obtained.

[0053] Specifically, after obtaining the current comprehensive state vector, the action is executed to input the vector into the response strategy matching engine. All preset state templates in the response strategy knowledge base are traversed, and the distance difference between the current vector and each historical template is evaluated using a similarity calculation method based on Euclidean distance. The vector distance in the multidimensional space is calculated by performing the square difference, summation, and square root on the feature values ​​of each dimension, thereby forming a complete similarity score list. Each state template corresponds to a distance value, which is used to measure its matching degree with the current state. For example, if the current comprehensive state vector is [001,0.8,-0.4,42,9,03], the Euclidean distance with a certain template [001,0.7,-0.3,40,10,03] is 1.87. The system records this result as a candidate match and participates in subsequent sorting and filtering.

[0054] S32: Based on the calculation results, select the response strategy corresponding to the state template with the smallest distance from the current integrated state vector from the state templates, and use it as the optimal response strategy.

[0055] Specifically, after calculating the similarity of all state templates, the action will sort the distance results and select the state template with the smallest distance as the best matching target for the current interaction state. At the same time, the response strategy configuration associated with this template will be extracted as the optimal response strategy for output. This strategy usually includes a combination of multiple parameters such as response content type, speech synthesis parameters and volume adjustment rules. For example, if the minimum distance selected in an interaction is 1.22, its corresponding template is "User ID=001, Emotion=Excited and negative, Environment=Multi-person conversation", and its binding strategy is "Use a soothing tone to reply and enable noise reduction control". This strategy will be used as the adaptation scheme for this speech response for personalized output generation.

[0056] S33: When all similarities are below the preset matching threshold, the default response strategy is invoked, and the current comprehensive state vector is recorded in the strategy learning log.

[0057] Specifically, after traversing all state templates and completing distance sorting, if the system detects that the Euclidean distance of all matching items is higher than the preset matching threshold (e.g., 2.5), meaning that the overall similarity between the current state and the historical records is low, then the system executes an action to call a general default response strategy as a safe response output. This strategy is a pre-set robust strategy template that ensures that the system still has reasonable interactive feedback capabilities in unidentified states. At the same time, the current comprehensive state vector, along with the timestamp and response result, is recorded in the strategy learning log for subsequent iterative training and knowledge expansion of the strategy library. For example, if an interaction state is determined to be a rare combination, the system responds using the default strategy "standard broadcast + moderate speech rate" and writes the state vector into the log for analysis, classification, and strategy reconstruction during the next batch model training.

[0058] In one embodiment, such as Figure 6 As shown, in step S40, the response content and output parameters of the smart speaker are adjusted according to the optimal response strategy, specifically including: S41: Select a content type that matches the current user's identity based on the optimal response strategy. The content type may include adult news broadcasts, children's stories, or operation prompts.

[0059] Specifically, after determining the optimal response strategy, the execution action is based on the user identity tags and content preference configuration specified in the strategy. It selects content types from the content resource pool that match the current user identity and outputs the response. If the user identity is identified as an adult family member, practical information content such as news broadcasts and schedule reminders are selected first. If the user is identified as a child user, interesting content such as children's stories, animated explanations, or simple Q&A are matched. If the user is identified as a visitor or an unidentified user, neutral operation prompts are selected by default. For example, when the user ID is identified as 001 and the bound identity is "parent", after issuing the command "Give me something relaxing", the system will select the daily life information broadcast content as the response content and execute it.

[0060] S42: Adjust the speech synthesis parameters according to the optimal response strategy. The speech synthesis parameters include tone type, speech rate level and timbre emotion curve to match the user's current emotional state.

[0061] Specifically, after content type filtering is completed, the execution action dynamically adjusts the speech engine according to the speech synthesis parameters specified in the optimal response strategy. This includes setting the tone type (e.g., formal tone, gentle tone, pleasant tone), speech rate level (e.g., slow, normal, fast), and timbre emotion curve (e.g., steady, gradually increasing, and emphasis at the end) to make the final output speech more in line with the current user's emotional state and content context. For example, when the emotion coordinate is (-0.6, 0.2), indicating slightly excited but positive tendencies, the strategy recommends using a gentle tone and medium speech rate, and raising the tone at the end to express encouragement, thus forming a speech response style of "Hello, this is a lighthearted life tip that I hope will be helpful to you."

[0062] S43: Adjust volume output parameters according to the optimal response strategy, including adjusting the speech rate in noisy environments, simplifying the speech content, or activating the low volume mode.

[0063] Specifically, after configuring the speech content and synthesis parameters, the system dynamically sets the output volume level and auxiliary presentation of the speech broadcast based on the volume control rules configured in the environmental state parameters and response strategy. If the system detects that the current ambient noise intensity is high and the human voice signal-to-noise ratio is low, it can automatically reduce the speech speed to avoid semantic loss, or semantically compress the broadcast content to output only key information. If necessary, it can enable a low volume mode and prompt the user to use headphones or move closer to the device for interaction. For example, when the current background noise is 52dB, the system recognizes a "multi-person conversation" scenario and the user is relatively calm, it will use a low speech speed, shorten the broadcast time, and prompt "The environment is a bit noisy. I'll tell you briefly: Today's temperature is 32 degrees Celsius. Remember to wear sunscreen when you go out."

[0064] In one embodiment, such as Figure 7 As shown, this smart speaker response method based on multi-dimensional state perception also includes: S50: After executing the personalized dynamic voice interaction response, generate a strategy execution record based on the current comprehensive state vector and response result, and store the strategy execution record in the strategy learning log. The strategy execution record includes the comprehensive state vector, the selected optimal response strategy identifier, and user interaction feedback information.

[0065] Specifically, after completing the personalized voice response output, the execution action generates a strategy execution record based on the comprehensive state vector used in the current interaction and the executed response strategy. This record contains three core data items: first, a snapshot of the comprehensive state vector during this interaction, used to fully describe the user's identity, emotional state, and environmental state; second, the identifier code of the selected optimal response strategy, used to track the strategy selection process; and third, user interaction feedback information, which can be collected through keyword recognition, tone of voice, or explicit user replies. For example, if the user actively says "This is good" or "Say it again" after the voice broadcast, it will be parsed as positive or neutral feedback and embedded in the record structure. The final generated strategy execution record is as follows: [CSV:001,0.8,-0.4,42,9,03;Strategy ID:S012;Feedback:Positive], and is appended to the strategy learning log file as the basic data source for subsequent knowledge base optimization.

[0066] S60: When the cumulative number of policy execution records reaches the preset learning threshold, a partial update of the response policy knowledge base is triggered to improve the matching accuracy and adaptability of subsequent response policies.

[0067] Specifically, when the number of accumulated strategy execution records in the strategy learning log reaches the system's preset learning trigger threshold, the execution action will initiate a partial update process for the response strategy knowledge base. First, cluster analysis and similarity summarization are performed on the state vectors and feedback results stored in the log to identify state types in the current knowledge base that have insufficient response or high misjudgment frequency. Then, the strategy configuration for these states is adjusted and optimized based on the quality of user feedback. For example, if it is found that when a certain environment is a combination of "TV sound + neutral mood", the user feedback obtained by the original response strategy is mostly "the volume is too high", then the volume parameter of the corresponding strategy will be reduced by 20% and a reminder will be set in the update. The updated strategy will replace the original entry or be added to the strategy base as a new entry to improve the response effect and personalized adaptability of subsequent matching.

[0068] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0069] In one embodiment, a smart speaker response system based on multi-dimensional state perception is provided, which corresponds one-to-one with the smart speaker response method based on multi-dimensional state perception in the above embodiments. For example... Figure 8 As shown, this smart speaker response system based on multi-dimensional state perception includes a state perception processing module, a state fusion generation module, a strategy matching module, and a response control module. Detailed descriptions of each functional module are as follows: The state perception processing module is used to receive the user's voice input signal, collect multi-dimensional state information based on the voice input signal, perform multi-dimensional state perception processing on the multi-dimensional state information, and obtain structured state data. Multi-dimensional state perception processing includes user identity recognition, emotional state quantification, and environmental state feature extraction. The state fusion generation module is used to fuse structured state data to generate a comprehensive state vector; The strategy matching module is used to perform fuzzy matching operations in a preset response strategy knowledge base based on the comprehensive state vector to obtain the optimal response strategy. The response control module is used to adjust the smart speaker's response content and output parameters according to the optimal response strategy, so as to complete personalized dynamic voice interaction response.

[0070] Optionally, the state-aware processing module includes: The voiceprint recognition submodule is used to extract individual feature parameters from the voice input signal based on the voiceprint recognition algorithm to construct multi-dimensional state information, and compare the individual feature parameters with the preset user identity database to obtain the user identity identifier. The identity modeling submodule is used to mark the current identity as an unfamiliar user when the comparison result cannot match a registered user, and to record the corresponding voice data for subsequent identity modeling. The acoustic feature extraction submodule is used to extract acoustic features from the speech input signal. The acoustic feature parameters include pitch, speech rate, energy, and zero-crossing rate. The emotion mapping submodule is used to map the current user's emotion into two-dimensional emotion coordinates based on acoustic feature parameters and a preset emotion mapping model. The X-axis of the two-dimensional emotion coordinates represents the degree of excitement, and the Y-axis represents the emotional tendency. The background audio acquisition submodule is used to acquire background audio data during the voice input process via a microphone array. The environmental parameter calculation submodule is used to calculate environmental characteristic parameters based on background audio data. The environmental characteristic parameters include the average background noise decibel value and the human voice signal-to-noise ratio. The sound source recognition submodule is used to input background audio data into the audio classification model and identify the background sound type label. The background sound type label is used to indicate the sound source attributes of the current environment. Sound source attributes include quiet, music, TV sound, and multi-person conversation. The vector encoding submodule is used to encode and combine the average background noise decibel value, human voice signal-to-noise ratio, and background sound type label to form an environmental state vector.

[0071] Optionally, the strategy matching module includes: The similarity calculation submodule is used to calculate the similarity between the comprehensive state vector and each preset state template in the response strategy knowledge base using the nearest neighbor search algorithm based on Euclidean distance, and obtain the calculation result. The optimal strategy selection submodule is used to select the response strategy corresponding to the state template with the smallest distance from the current integrated state vector from the state templates based on the calculation results, and use it as the optimal response strategy. The default strategy fallback submodule is used to invoke the default response strategy and record the current comprehensive state vector to the strategy learning log when all similarities are below the preset matching threshold.

[0072] Optionally, the response control module includes: The content type matching submodule is used to select the content type that matches the current user's identity based on the optimal response strategy. The content type includes adult news broadcasts, children's stories, or operation prompts. The speech synthesis adjustment submodule is used to adjust the speech synthesis parameters according to the optimal response strategy. The speech synthesis parameters include tone type, speech rate level and timbre emotion curve to match the user's current emotional state. The volume control submodule is used to adjust the volume output parameters according to the optimal response strategy, including adjusting the speech rate in noisy environments, simplifying the speech content, or activating a low volume mode.

[0073] Optionally, this smart speaker response system based on multi-dimensional state perception also includes: The strategy execution record module is used to generate a strategy execution record based on the current comprehensive state vector and response result after executing the personalized dynamic voice interaction response, and store the strategy execution record in the strategy learning log. The strategy execution record includes the comprehensive state vector, the selected optimal response strategy identifier, and user interaction feedback information. The strategy knowledge base update module is used to trigger a partial update of the response strategy knowledge base when the cumulative number of strategy execution records reaches a preset learning threshold, so as to improve the matching accuracy and adaptability of subsequent response strategies.

[0074] For specific limitations regarding a smart speaker response system based on multi-dimensional state perception, please refer to the limitations of a smart speaker response method based on multi-dimensional state perception mentioned above, which will not be repeated here. Each module in the aforementioned smart speaker response system based on multi-dimensional state perception can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0075] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0076] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A smart speaker response method based on multi-dimensional state perception, characterized in that, The aforementioned smart speaker response method based on multi-dimensional state perception includes: The system receives the user's voice input signal, collects multi-dimensional state information based on the voice input signal, performs multi-dimensional state perception processing on the multi-dimensional state information, and obtains structured state data. The multi-dimensional state perception processing includes user identity recognition, emotional state quantification, and environmental state feature extraction. The structured state data is fused to generate a comprehensive state vector; Based on the comprehensive state vector, a fuzzy matching operation is performed in the preset response strategy knowledge base to obtain the optimal response strategy; Based on the optimal response strategy, the response content and output parameters of the smart speaker are adjusted to achieve personalized dynamic voice interaction response.

2. The smart speaker response method based on multi-dimensional state perception according to claim 1, characterized in that, The step of collecting multi-dimensional state information based on the voice input signal and performing user identification on the multi-dimensional state information includes: Based on the voiceprint recognition algorithm, individual feature parameters for constructing the multidimensional state information are extracted from the voice input signal, and the individual feature parameters are compared with a preset user identity database to obtain the user identity identifier; When the comparison result cannot match a registered user, the current identity is marked as an unfamiliar user, and the corresponding voice data is recorded for subsequent identity modeling.

3. The smart speaker response method based on multi-dimensional state perception according to claim 1, characterized in that, The step of collecting multi-dimensional state information based on the voice input signal and quantifying the emotional state of the multi-dimensional state information includes: Acoustic features are extracted from the speech input signal, and the acoustic feature parameters include pitch, speech rate, energy, and zero-crossing rate. Based on the acoustic feature parameters, the current user's emotion is mapped to a two-dimensional emotion coordinate using a preset emotion mapping model. The X-axis of the two-dimensional emotion coordinate represents the degree of excitement, and the Y-axis represents the emotional tendency.

4. The smart speaker response method based on multi-dimensional state perception according to claim 1, characterized in that, The step of collecting multi-dimensional state information based on the voice input signal and extracting environmental state features from the multi-dimensional state information includes: Background audio data during the voice input process is collected using a microphone array; Environmental characteristic parameters are calculated based on the background audio data, including the average background noise decibel value and the human voice signal-to-noise ratio. The background audio data is input into an audio classification model to identify background sound type labels. The background sound type labels are used to indicate the sound source attributes of the current environment. The sound source attributes include quiet, music, television sound, and multi-person conversation. The average background noise decibel value, human voice signal-to-noise ratio, and background sound type label are encoded and combined to form an environmental state vector.

5. The smart speaker response method based on multi-dimensional state perception according to claim 1, characterized in that, The step of performing a fuzzy matching operation in a preset response strategy knowledge base based on the comprehensive state vector to obtain the optimal response strategy includes: The similarity between the comprehensive state vector and each preset state template in the response strategy knowledge base is calculated using a nearest neighbor search algorithm based on Euclidean distance to obtain the calculation results; Based on the calculation results, the response strategy corresponding to the state template with the smallest distance from the current comprehensive state vector is selected from the state templates and is taken as the optimal response strategy. When all similarities are below the preset matching threshold, the default response strategy is invoked, and the current comprehensive state vector is recorded in the strategy learning log.

6. The smart speaker response method based on multi-dimensional state perception according to claim 1, characterized in that, The step of adjusting the smart speaker's response content and output parameters according to the optimal response strategy includes: Based on the optimal response strategy, a content type matching the current user's identity is selected, and the content type includes adult news broadcasts, children's stories, or operation prompts. The speech synthesis parameters are adjusted according to the optimal response strategy. The speech synthesis parameters include intonation type, speech rate level and timbre emotion curve to match the user's current emotional state. Adjusting the volume output parameters according to the optimal response strategy, the adjustment includes speaking at a low speed in noisy environments, simplifying the speech content, or activating a low volume mode.

7. The smart speaker response method based on multi-dimensional state perception according to claim 1, characterized in that, The smart speaker response method based on multi-dimensional state perception also includes: After executing the personalized dynamic voice interaction response, a strategy execution record is generated based on the current comprehensive state vector and response result, and the strategy execution record is stored in the strategy learning log. The strategy execution record includes the comprehensive state vector, the selected optimal response strategy identifier, and user interaction feedback information. When the cumulative number of strategy execution records reaches a preset learning threshold, a partial update of the response strategy knowledge base is triggered to improve the matching accuracy and adaptability of subsequent response strategies.

8. A smart speaker response system based on multi-dimensional state perception, characterized in that, The aforementioned smart speaker response system based on multi-dimensional state perception includes: The state perception processing module is used to receive the user's voice input signal, collect multi-dimensional state information based on the voice input signal, perform multi-dimensional state perception processing on the multi-dimensional state information, and obtain structured state data. The multi-dimensional state perception processing includes user identity recognition, emotional state quantification, and environmental state feature extraction. The state fusion generation module is used to fuse the structured state data to generate a comprehensive state vector; The strategy matching module is used to perform fuzzy matching operations in a preset response strategy knowledge base based on the comprehensive state vector to obtain the optimal response strategy. The response control module is used to adjust the response content and output parameters of the smart speaker according to the optimal response strategy, so as to complete the personalized dynamic voice interaction response.

Citation Information

Patent Citations

  • Response strategy generation and execution method and device, equipment and medium

    CN119541461A

  • LLM-based client intention identification and response system, method and device, and medium

    CN119808789A

  • Intelligent interactive enterprise management simulation system and method thereof

    CN120257840A

  • Rapid intention insight and intelligent response method and system for intelligent equipment

    CN120448977A

  • AI intelligent auxiliary question answering system for student practice

    CN120450914A