Interactive equipment scheduling method and related product
By using real-time emotion recognition and voice activity detection, combined with intelligent scheduling of voice input devices, the problem of poor user experience in existing voice interaction technologies has been solved, achieving more accurate and timely voice responses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-14
AI Technical Summary
Existing voice interaction technologies suffer from poor user experience, such as long speech without feedback leading to uncertainty, overload of single-round information broadcasting, misjudging network latency as lag, and lack of unified scheduling for multi-capability collaboration, resulting in low execution rate and insufficient anxiety recognition.
By acquiring voice data collected by the target voice input device in real time, emotion recognition and voice activity detection are performed to generate appropriate voice response results. Combined with preset priority and proximity state parameters, intelligent scheduling of voice input devices is carried out.
It enables intelligent scheduling of multiple voice input devices, improves the accuracy and timeliness of voice interaction, reduces the negative experience caused by interaction interruptions, and provides responses that meet user needs.
Smart Images

Figure CN121862104A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of voice interaction technology, specifically to an interactive device scheduling method and related products. Background Technology
[0002] Currently, voice interaction technology has been deeply integrated into real-world scenarios such as cycling, commuting, and meetings, and related technologies for multi-device, multimodal interaction have a certain foundation for exploration. However, existing voice interaction technologies still face some problems, such as: from a user experience perspective, during voice interaction, long speeches without feedback can easily cause user uncertainty or accidental interruptions; single-round information recitation overload exceeds auditory working memory, leading to low execution rates; there is a lack of recognition and adaptive response to user emotions such as anxiety and anger; and the lack of placeholder feedback during network latency periods can easily be misjudged as lag or offline. From a technical collaboration perspective, the aforementioned capabilities such as echoing, density adjustment, emotion recognition, and filler words, when used in isolation, are mutually restrictive, lacking a solution that can unify these capabilities under the same timing and arbitration rules, making it impossible to achieve multi-capability collaborative work to form a complete closed loop of dialogue experience.
[0003] Therefore, it is necessary to propose an interactive device scheduling method to solve at least one of the above-mentioned technical problems. Summary of the Invention
[0004] The embodiments of this disclosure propose an interactive device scheduling method and related products.
[0005] Firstly, this disclosure provides an interactive device scheduling method, including: The system acquires voice data collected by the target voice input device in real time, and performs emotion recognition on the collected voice data to obtain the user's emotion category. The collected voice data is subjected to voice activity detection to obtain the user's current interaction state; The voice response result of the target voice input device is generated based on the user's emotion category and the interaction state.
[0006] In some alternative implementations, the target voice input device is predetermined through the following target voice input device determination steps: Obtain the preset priority and proximity status parameters of each voice input device; For each of the voice input devices, a weighted sum of the preset priority and proximity state parameters of the voice input device is obtained to obtain the voice input device adaptation score of the voice input device. The proximity state parameters are used to characterize the physical distance between the voice input device and the user. The voice input device with the highest voice input device adaptation score among all the aforementioned voice input devices is determined as the target voice input device.
[0007] In some optional implementations, the real-time acquisition of voice data collected by the target voice input device, and the performance of emotion recognition on the collected voice data to obtain the user's emotion category, includes: Extract the voice features and interaction behavior features of the voice data from the target voice input device; Based on the speech features and the interaction behavior features, the speech data is processed for emotion recognition to obtain the emotion representation of the speech data. The user emotion category is generated from the emotion representation of the voice data.
[0008] In some optional implementations, the step of performing voice activity detection on the collected voice data to obtain the user's current interaction state includes: Extract speech pause data from the speech data of the target speech input device; If the duration of the pause in the voice pause data exceeds the pause duration threshold, the current interaction state is determined to be an interrupted input state. If the duration of the pause in the voice pause data does not exceed the pause duration threshold, the current interaction state is determined to be a continuous input state.
[0009] In some optional implementations, generating the voice response result of the target voice input device based on the user's emotion category and the interaction state includes: In response to the current interaction state being the continuous input state, a response emotion category is determined based on the user emotion category; A non-overlapping voice response result is generated for the voice data based on the response emotion category.
[0010] In some optional implementations, generating the voice response result of the target voice input device based on the user's emotion category and the interaction state further includes: In response to the current interaction state being the interrupted input state, the information density of the voice data is calculated; Obtain the current communication connection speed with the target voice input device; In response to the information density exceeding a preset information density threshold and the current communication connection speed exceeding a preset minimum communication connection speed threshold, a semantic transition speech response result is generated for the speech data based on the user emotion category and the information density.
[0011] Secondly, this disclosure provides an interactive device scheduling apparatus, comprising: The voice data acquisition unit is used to acquire voice data collected by the target voice input device in real time, and to perform emotion recognition on the collected voice data to obtain the user's emotion category. The detection unit is used to detect voice activity in the collected voice data to obtain the user's current interaction state; A response unit is used to generate a voice response result of the target voice input device based on the user's emotion category and the interaction state.
[0012] In some alternative implementations, the target voice input device is predetermined through the following target voice input device determination steps: Obtain the preset priority and proximity status parameters of each voice input device; For each of the voice input devices, a weighted sum of the preset priority and proximity state parameters of the voice input device is obtained to obtain the voice input device adaptation score of the voice input device. The proximity state parameters are used to characterize the physical distance between the voice input device and the user. The voice input device with the highest voice input device adaptation score among all the aforementioned voice input devices is determined as the target voice input device.
[0013] In some optional implementations, the voice data acquisition unit is further configured to: Extract the voice features and interaction behavior features of the voice data from the target voice input device; Based on the speech features and the interaction behavior features, the speech data is processed for emotion recognition to obtain the emotion representation of the speech data. The user emotion category is generated from the emotion representation of the voice data.
[0014] In some alternative implementations, the detection unit is further configured to: Extract speech pause data from the speech data of the target speech input device; If the duration of the pause in the voice pause data exceeds the pause duration threshold, the current interaction state is determined to be an interrupted input state. If the duration of the pause in the voice pause data does not exceed the pause duration threshold, the current interaction state is determined to be a continuous input state.
[0015] In some optional implementations, the response unit is further configured to: In response to the current interaction state being the continuous input state, a response emotion category is determined based on the user emotion category; A non-overlapping voice response result is generated for the voice data based on the response emotion category.
[0016] In some optional implementations, the response unit further includes: In response to the current interaction state being the interrupted input state, the information density of the voice data is calculated; Obtain the current communication connection speed with the target voice input device; In response to the information density exceeding a preset information density threshold and the current communication connection speed exceeding a preset minimum communication connection speed threshold, a semantic transition speech response result is generated for the speech data based on the user emotion category and the information density.
[0017] Thirdly, this disclosure provides an electronic device, including: One or more processors; Storage device, on which one or more programs are stored, When the above-described one or more programs are executed by the above-described one or more processors, the above-described one or more processors implement the method as described in any embodiment of the first aspect of this disclosure.
[0018] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method described in any embodiment of the first aspect of this disclosure.
[0019] Fifthly, this disclosure provides a computer program product including a computer program / instructions that, when executed by a processor, implement the method described in any embodiment of the first aspect of this disclosure.
[0020] The interactive device scheduling method, apparatus, electronic device, and storage medium provided in the embodiments of this disclosure first acquire voice data collected by a target voice input device in real time, and then perform emotion recognition on the collected voice data to obtain the user's emotion category; next, perform voice activity detection on the collected voice data to obtain the user's current interaction state; finally, generate a voice response result for the target voice input device based on the user's emotion category and interaction state. In this way, by recognizing the user's emotion category, detecting the interaction state, and generating an appropriate voice response result in real time, intelligent scheduling and optimal selection of multiple voice input devices are achieved. This enables voice interaction to accurately adapt to the user's emotional state, providing the user with a response that meets their needs, effectively improving the accuracy and timeliness of voice interaction, and reducing the negative experience caused by interaction interruptions. Attached Figure Description
[0021] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the invention. In the drawings: Figure 1This is an exemplary system architecture diagram to which one embodiment of this disclosure can be applied; Figure 2A This is a flowchart of one embodiment of the interactive device scheduling method disclosed herein; Figure 2B This is an exploded flowchart of one embodiment of step 201 of the present disclosure; Figure 2C This is an exploded flowchart of one embodiment of step 202 of the present disclosure; Figure 2D This is an exploded flowchart of one embodiment of step 203 of this disclosure; Figure 3 This is a flowchart of one embodiment of the target voice input device determination step 300 of this disclosure; Figure 4 This is a schematic diagram of the structure of an embodiment of the interactive device scheduling apparatus according to the present disclosure; Figure 5 This is a schematic diagram of the structure of a computer system suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0022] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] Figure 1 An exemplary system architecture 100 is shown, to which embodiments of the interactive device scheduling methods, apparatuses, electronic devices and storage media of this disclosure can be applied.
[0025] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0026] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as interactive device scheduling applications, voice interaction applications, video conferencing applications, short video social applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0027] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with microphones and speakers, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), portable computers, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0028] Server 105 can be a server providing various services, such as acquiring voice data collected by the target voice input device in real time on terminal devices 101, 102, and 103, and performing emotion recognition on the collected voice data to obtain the user's emotion category. The backend server can perform voice activity detection on the collected voice data to obtain the user's current interaction state.
[0029] In some cases, the interactive device scheduling method provided in this disclosure can be jointly executed by terminal devices 101, 102, 103 and server 105. For example, the step of "real-time acquisition of voice data collected by the target voice input device and emotion recognition of the collected voice data to obtain the user's emotion category" can be executed by terminal devices 101, 102, 103, and the step of "voice activity detection of the collected voice data to obtain the user's current interaction state" can be executed by server 105. This disclosure does not limit this. Correspondingly, the interactive device scheduling device can also be respectively set in terminal devices 101, 102, 103 and server 105.
[0030] In some cases, the interactive device scheduling method provided in this disclosure can be executed by server 105. Accordingly, the interactive device scheduling device can also be set in server 105. In this case, the system architecture 100 may not include terminal devices 101, 102, and 103.
[0031] In some cases, the interactive device scheduling method provided in this disclosure can be executed by terminal devices 101, 102, and 103. Correspondingly, the interactive device scheduling device can also be set in terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include server 105.
[0032] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (for example, used to provide distributed services), or as a single software program or software module. No specific limitations are made here.
[0033] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0034] Continue to refer to Figure 2A , Figure 2A A flow 200 is shown as an embodiment of the interactive device scheduling method according to the present disclosure. Figure 2A The interactive device scheduling method shown can be applied to Figure 1 The terminal device or server shown. The process 200 includes the following steps: Step 201: Acquire voice data collected by the target voice input device in real time, and perform emotion recognition on the collected voice data to obtain the user's emotion category.
[0035] In this embodiment, the execution body of the interactive device scheduling method (e.g.) Figure 1The server 105 can first acquire the voice data collected by the target voice input device in real time. Here, the target voice input device refers to the voice input device with the highest voice input device adaptation score selected from multiple voice input devices with voice acquisition and data transmission capabilities, after weighted calculation based on preset priority, proximity parameters such as physical distance from the user, etc. The purpose of selecting the target voice input device is to ensure the accuracy, efficiency, and adaptability of voice interaction in scenarios where multiple voice input devices coexist. It prioritizes voice input devices that are closer to the user, respond more promptly, and have better acquisition quality, avoiding voice data conflicts or redundancy caused by simultaneous acquisition by multiple voice input devices. At the same time, it makes voice data acquisition more targeted, reduces resource consumption caused by invalid voice data transmission, and provides high-quality raw data support for subsequent emotion recognition, interaction state detection, and adaptive response generation, ultimately achieving a more natural, smooth, and user-friendly voice interaction experience. The core function of the target voice input device is to start voice acquisition through specific modes such as voice wake-up, button triggering, and touch sensing, to capture the user's voice content, acoustic features, and interaction features in real time, and transmit the voice data to the execution subject to support emotion recognition, interaction state detection, and the generation of adapted voice responses. It is the core input carrier for achieving accurate and efficient voice interaction.
[0036] Voice input can be activated in various modes, including voice wake-up (e.g., the user speaks a preset wake-up word to trigger voice capture), physical operation (e.g., pressing the voice capture button on the device, touching a designated interaction area), and gesture control (e.g., double-tap gestures on smartwatches, touch gestures on headphones), and wear detection linkage (the device automatically initiates voice capture preparation after detecting the user's wearing status). This ensures users can easily trigger voice input in different scenarios. The acquired voice data is complete user voice information captured by the voice input device through microphones and other acquisition components. It includes not only the semantic content expressed by the user, such as inquiries and commands, but also voice-related acoustic and interactive features, such as speech rate, volume, pitch changes, pause duration, and intonation. This voice data comprehensively reflects the user's expressive intent and emotional state. The core purpose of acquiring voice data is to provide raw data support for subsequent emotion recognition and interaction state detection. Based on these analysis results, voice responses that meet user needs can be generated. Its key role is to build an effective interactive bridge between the user and the voice input device, enabling the voice input device to accurately understand the user's intentions and perceive the user's emotions, thus avoiding aimless responses.
[0037] In practice, after the target voice input device activates the voice acquisition function through the corresponding modality, it captures the user's voice in real time and converts it into transmittable digital voice data. Then, through a preset communication connection method (such as Bluetooth, Wi-Fi, mobile network, etc.), the acquired voice data is uploaded in real time to the execution subject of the interactive device scheduling method. The execution subject continuously listens for and receives the voice data, while ensuring the real-time and integrity of the data transmission, laying the foundation for subsequent processing steps such as emotion recognition and interaction state detection.
[0038] Then, the aforementioned executing entity can perform emotion recognition on the collected voice data to determine the user's emotion category. Here, emotion recognition refers to the technical process of extracting voice features and interaction behavior features from the collected voice data, analyzing and processing these features through a pre-set emotion recognition model, and then determining the user's current emotion category. Common emotion categories include anxiety, anger, depression, calmness, and joy. The core purpose of emotion recognition is to accurately perceive the user's emotional state, providing a basis for generating appropriate voice responses.
[0039] By clearly defining the user's emotion category, the accuracy of the voice input source and the precise perception of the user's emotional state are achieved. This lays a high-quality data foundation for generating adaptive voice responses based on emotions and interaction states, effectively avoiding interaction deviations caused by multi-device interference and emotional insensitivity, and improving the targeting of the interaction and the comfort of the user experience.
[0040] In some alternative implementations, step 201 may include, for example: Figure 2B The following steps 2011 to 2013 are shown: Step 2011: Extract the speech features and interaction behavior features of the speech data from the target speech input device.
[0041] Here, speech features refer to acoustic properties extracted from speech data; interactive behavior features refer to features that reflect the user's behavioral patterns during speech interaction. These speech features and interactive behavior features together serve as the core input data for emotion recognition, used to analyze and determine the user's current emotion category (such as anxiety, pleasure, etc.) through a pre-set model or algorithm, providing a key basis for generating an emotion-appropriate speech response.
[0042] In specific implementations, speech features may include, for example, the following features: Speech rate: the number of syllables pronounced per unit of time; Volume: The loudness or softness of a sound; Pitch: The variation in the highness or lowness of a sound; intonation: the rise and fall of a voice; Spectral characteristics: The distribution of sound frequencies.
[0043] As examples, some corresponding extraction methods for the above speech features are given below: Speech rate: The effective speech segments are located by speech activity detection technology, the number of syllables or words per unit time (such as words per minute) is counted, and the speech rate value is determined by combining the speech pause boundaries; Volume: Short-time energy analysis is used to calculate the energy value of each frame of the speech signal. The relative volume is obtained through normalization, or the peak and average values of the signal amplitude are directly extracted to characterize the volume. Pitch: Based on fundamental frequency estimation, the fundamental frequency value of periodic vibrations in the speech signal is extracted using the autocorrelation method, cepstral method or YIN algorithm, and the pitch is reflected by the curve of fundamental frequency change over time; intonation variation: Analyze the joint variation characteristics of fundamental frequency and energy, calculate the standard deviation of fundamental frequency, the fluctuation amplitude of energy and the co-variation rate of the two within the speech segment, and quantify the degree of intonation. Spectral characteristics: The time-domain speech signal is converted to the frequency domain by fast Fourier transform, and features such as Mel frequency cepstral coefficients, spectral centroid, and spectral bandwidth are extracted, or the energy distribution characteristics of different frequency bands are obtained by using Mel filter banks.
[0044] Interactive behavior characteristics may include, for example, the following features: Rhythm: The regularity of intervals between speech segments; Sentence completeness: the fluency of spoken expression; Repetition frequency: The number of times a specific word or phrase is repeated.
[0045] As examples, the following are some corresponding methods for extracting the above-mentioned interactive behavior features: Expression rhythm: The start and end timestamps of each speech segment are obtained through speech frame processing, the time interval distribution of adjacent speech segments and the alternation period of speech and pauses are calculated, and the speed and rhythm of expression are quantitatively analyzed. Sentence completeness: Natural language processing technology is used to analyze the text structure after speech-to-text conversion, detect the grammatical completeness, semantic coherence and missing key information (subject, predicate, object), and make a comprehensive judgment in combination with the fluency of the speech signal (no frequent interruptions or repetitions); Repetition frequency: By converting speech data into text through speech-to-text conversion, string matching algorithms or word frequency statistics tools are used to identify and count the number of times specific words, phrases or sentences are repeated in the text, and the timestamps of the speech segments are combined to locate the repeated interaction scenarios.
[0046] By accurately extracting speech features and interaction behavior features from the speech data collected by the target speech input device, multi-dimensional and in-depth analysis of the speech data is achieved. This provides comprehensive and high-quality feature input support for subsequent emotion recognition, thereby ensuring the accuracy and reliability of emotion recognition results. It lays a solid data foundation for generating personalized speech responses that match the user's emotions, and further improves the accuracy of voice interaction and user experience.
[0047] Step 2012: Perform emotion recognition processing on the speech data based on speech features and interaction behavior features to obtain the emotion representation of the speech data.
[0048] Here, emotion representation refers to the core information carrier that accurately depicts a user's emotional state, obtained through model analysis and quantification of extracted speech features and interactive behavior features. It includes not only the category attributes of the emotion but also its intensity dimension, and may also include the dynamic trends of emotional change. Its core purpose is to provide a clear emotional basis for the subsequent generation of adapted speech responses.
[0049] In practice, emotional representations may include, for example, the following: Emotional categories: such as anxiety, anger, depression, peace, and joy, which are the core emotional types of users; Emotional intensity: The degree of intensity of emotions such as mild anxiety and moderate anger, usually expressed as a quantitative score or level; Emotional dynamics: such as the trajectory of emotions gradually shifting from calm to anxiety, and information on changes in the frequency of emotional fluctuations over time.
[0050] As examples, some corresponding methods for extracting the above emotional representations are given below: Emotion Category: The extracted voice features and interaction behavior features are input into a pre-trained emotion classification model (such as Support Vector Machine (SVM), Recurrent Neural Network (RNN), Transformer model). The model learns the mapping relationship between features and emotion categories, and outputs the user's core emotion type (such as anxiety, anger, peace, etc.). Cross-validation is used to optimize model parameters and improve classification accuracy. Emotional intensity: Using regression analysis, the extracted speech features and interactive behavior features are input into the emotional intensity regression model (such as logistic regression, gradient boosting tree GBRT). The model outputs a quantitative score of 0-100 or a "mild / moderate / severe" level. Alternatively, a rule base can be established based on feature thresholds (such as high speech speed + high volume corresponding to high anger intensity) to determine the intensity. Emotional dynamic features: Based on a sliding time window (e.g., 500ms-1s window), feature sequences are continuously extracted. The changing trend of features over time is captured by a time series model (e.g., Long Short-Term Memory Network LSTM, Gated Recurrent Unit GRU), generating an emotion change trajectory curve. The slope, fluctuation amplitude and peak interval of the curve are calculated to quantify the dynamic change pattern of emotions.
[0051] By performing emotion recognition processing on voice data based on voice features and interaction behavior features to obtain emotion representations, a multi-dimensional and accurate characterization of users' emotional states is achieved. This provides quantitative and detailed emotional basis for subsequent interaction decisions, thereby breaking through the limitations of single semantic understanding, deeply adapting to users' emotional needs, laying the core foundation for generating empathetic and personalized voice responses, and significantly improving the emotional fit of voice interaction and user experience satisfaction.
[0052] Step 2013: Generate user emotion categories from voice data based on emotion representations.
[0053] Here, user emotion category refers to the classification result extracted from emotion representation, which is used to clearly define the core attributes of user emotion. It is a general definition of the user's current dominant emotion, which is used to provide clear emotion guidance for subsequent interactive device scheduling and voice response generation, and ensure that the interaction strategy can be adjusted in a targeted manner.
[0054] In practice, user emotion categories may include, for example, the following categories: Anxiety: Manifested as rapid speech and frequent pauses; Anger: manifested as a sudden increase in volume and a sharp, shrill tone; Low spirits: manifested as slow speech and a low tone; Peaceful: Characterized by a steady and even speech rhythm; Pleasant: Accompanied by an upward tone and a light, cheerful manner.
[0055] As examples, the following are some corresponding methods for extracting the above user emotion categories: Anxiety: By detecting speech features such as increased speech rate, fluctuating volume, and high-frequency pitch fluctuations, combined with interactive behavior features such as irregular and frequent pauses and fragmented sentences (repetition or interruption), when these feature combinations meet the preset anxiety threshold, it is determined to be an anxiety category. Anger: When the volume of speech features is significantly increased, the tone is sharp and explosive, and the tone fluctuates wildly, and when the interactive behavior features are short and forceful sentences with high repetition frequency (such as repeated questioning words), it is classified as anger based on the feature intensity exceeding the anger judgment threshold. Low-pitched: When capturing speech features such as slow speech rate, low volume, low tone and lack of fluctuation, and interactive behavior features such as prolonged pauses, sluggish expression rhythm and low sentence completeness (often accompanied by sighing pauses), and the output results of the low-pitched feature model are matched, it is determined to be the low-pitched category. Peaceful: When the speech features are moderate speech rate, stable volume, and gentle tone changes, and the interactive behavior features include regular pause duration, even expression rhythm, and coherent and complete sentences, and all features fall within the peaceful feature range, it is judged as the peaceful category. Pleasant: When extracting speech features such as rising pitch, light and energetic tone, moderate volume and rhythm, and interactive behavior features such as fluent expression, short and regular pauses, and matching the pleasant feature template or model to output a high confidence result, it is classified as pleasant.
[0056] By extracting clear user emotion categories based on emotional representations, the system achieves precise focus and clear definition of users' core emotional states, avoiding the ambiguity and generalization of emotional information. This provides a direct and clear decision-making basis for the personalized adaptation of subsequent voice responses, further enhancing the empathy and fit of voice interaction, and improving users' interactive experience and acceptance.
[0057] Step 202: Perform voice activity detection on the collected voice data to obtain the user's current interaction state.
[0058] In some alternative implementations, step 202 may include, for example: Figure 2C The following steps are shown from 2021 to 2023: Step 2021: Extract speech pause data from the speech data of the target speech input device.
[0059] Here, speech pause data refers to information related to silent intervals during user speech expression extracted from speech data collected from the target speech input device. This includes the start and end timestamps of the pauses, their duration, frequency, and interval patterns. Its core purpose is to serve as a key basis for determining the user's current interaction state (continuous input / interrupted input). The extraction aims to accurately distinguish between valid speech and silent periods by analyzing pause characteristics, providing data support for subsequent interaction state determination.
[0060] In practice, speech activity detection (VAD) technology can be used for extraction. First, the speech data is processed by framing. Silent frames are identified by calculating acoustic parameters such as energy and zero-crossing rate of each frame. Then, continuous silent frames are integrated into pause segments. Pause-related data are calculated by combining timestamps. At the same time, invalid silences caused by short-term noise are filtered out to ensure data accuracy.
[0061] By extracting speech pause data from the speech data of the target speech input device, the system accurately captures and quantifies the silent information in the rhythm of the user's speech expression, providing key feature support for subsequent interaction state detection. This avoids the one-sidedness of judging the interaction state solely based on the speech content, ensuring that the judgment of the interaction state is more objective and accurate, laying the foundation for generating responses that fit the rhythm of the user's input, and improving the coherence and adaptability of voice interaction.
[0062] Step 2022: In response to the fact that the duration of the pause in the voice pause data exceeds the pause duration threshold, the current interaction state is determined to be an interrupted input state.
[0063] Here, the duration of a pause refers to the continuous length of a single pause segment in the speech pause data from start to end. It is mainly used to distinguish between temporary pauses and active termination of user speech input. Interrupted input state refers to the state in which the user actively stops speech input due to reasons such as completing the expression or temporary interruption.
[0064] In practice, the duration of a single pause segment is first determined based on voice pause data. Then, the duration of the pause is compared with a preset pause duration threshold (set according to common voice interaction scenarios and user expression habits). When the duration of the pause exceeds the threshold, the current interaction state can be determined as an interrupted input state, thus accurately distinguishing between temporary pauses caused by the user taking a short breath or thinking and the termination of input.
[0065] By comparing the duration of the pause with a pause duration threshold to determine the interruption status of input, the system achieves accurate identification of user input termination behavior, avoiding misjudging temporary pauses as input interruptions or missing real interruption signals. This provides a reliable basis for timely response and avoids invalid waiting or interruption of user input, ensuring the timeliness and continuity of voice interaction and improving the smoothness of the user interaction experience.
[0066] Step 2023: In response to the fact that the duration of the pause in the voice pause data does not exceed the pause duration threshold, the current interaction state is determined to be a continuous input state.
[0067] Here, continuous input state refers to an interactive state in which the user has not stopped speaking and the current pause is only a brief breath, thinking, or a natural interval between sentences.
[0068] In practice, the duration of the current pause segment is first obtained based on the voice pause data. The pause duration is then compared with a pause duration threshold. If the pause duration does not exceed the pause duration threshold, and it is determined that the user is still in the process of voice input, the voice acquisition state will be maintained, and the voice data transmitted by the target voice input device will be continuously received without initiating the response generation process, until the pause duration is detected to exceed the pause duration threshold or the complete voice input is received.
[0069] By comparing the duration of pauses with a pause duration threshold to determine the continuous input status, the system achieves precise adaptation to the rhythm of user voice input. This avoids interrupting the expression by accidentally initiating a response during a short pause, thus ensuring the integrity and coherence of user voice input. It provides comprehensive data support for generating accurate and adapted responses, further improving the fluency of voice interaction and user experience.
[0070] Step 203: Generate the voice response result of the target voice input device based on the user's emotion category and interaction state.
[0071] In some alternative implementations, step 203 may include, for example: Figure 2D The following steps 2031 to 2032 are shown: Step 2031: In response to the current interaction state being a continuous input state, determine the response emotion category based on the user's emotion category.
[0072] Here, the response emotion category refers to the emotion type determined based on the user's current emotion category and the interaction needs under continuous input, used to match the user's emotions and adapt to subsequent responses. Its core is to enhance the fit of the interaction through emotional resonance, without interfering with the user's continuous input.
[0073] In practice, response emotion categories may include, for example, the following categories: When the user is anxious, the emotional response category is mild reassurance; When the user is happy, the emotional response category is positive and consistent. When a user is feeling down, the emotional response category is warm and encouraging. When the user's emotions are stable or the intensity of their emotions is low, the response emotion category is peaceful and friendly. When a user experiences significant emotional fluctuations, the response emotion category is "soothing and gentle".
[0074] Determining the response emotion category based on the user's emotion category requires first establishing an adaptation relationship between the user's emotion and the response emotion through an emotion mapping rule base. Then, the adaptation accuracy is optimized by combining machine learning models (such as an emotion matching model based on Transformer). At the same time, the interaction requirements of continuous input state are taken into account to ensure that the response emotion not only matches the user's emotional state, but also does not interfere with the user's subsequent voice input, thus balancing empathy and interaction continuity.
[0075] By determining the response emotion category based on the user's emotion category during continuous input, voice interaction has been upgraded from simply responding to content to accurately empathizing with emotions. This not only matches the user's current emotional state but also avoids interference with the user's continuous expression due to emotional misalignment. At the same time, it conveys a sense of being cared for, alleviating the uncertainty of the interaction, thereby effectively bridging the psychological distance between human and computer interaction and enhancing the user's sense of identification and comfort with the interaction.
[0076] Step 2032: Generate non-overlapping voice response results for the voice data based on the response emotion category.
[0077] Here, non-intercepting voice response results refer to auxiliary voice feedback generated based on the response emotion category when the user is in a continuous input state. This feedback does not preempt the voice interaction round and does not interrupt the user's subsequent voice input. It is used to convey empathy and interaction confirmation without interfering with the user's continuous expression, allowing the user to perceive that the emotion has been captured and is being responded to in real time, while maintaining the continuity and fluency of the interaction.
[0078] In practice, generating non-turn-taking speech responses can be achieved using natural language generation (NLG) tools (such as dialogue generation frameworks based on pre-trained models like the GPT series and BART) combined with an emotion-adaptation template library (containing preset speech structures for different response emotion categories). First, the model generates a context-appropriate text response based on the response emotion category and the semantic information of the user's speech data. Then, text-to-speech (TTS) tools (such as Google Text-to-Speech and Baidu Speech Synthesis) convert the text into speech. Simultaneously, speech parameter adjustment tools (such as controlling the speech rate to slow down and the tone to be softer) ensure that the speech style matches the response emotion. Finally, a short, lightweight non-turn-taking speech response is output without interrupting the user's input.
[0079] As an example, the following are some non-take-ahead speech response results for the above-mentioned emotion categories: Gentle reassurance: "I'm listening," "Don't worry"; Positive alignment: "That's great," "That's awesome!" Warm encouragement: "Keep going!" "Keep going!" Peaceful and friendly: "Hmm," "Okay"; Shu Huanhuanchong: "Slow down," "Take it easy."
[0080] By generating short, quick, and non-repeating voice response results based on the response emotion category, an interactive effect that is uninterrupted, responsive, and highly empathetic during continuous input is achieved. This allows users to perceive attention and emotional fit in real time without interfering with the rhythm of subsequent expression, thereby significantly reducing the uncertainty and loneliness of users during continuous input and making human-computer interaction more harmonious and warm.
[0081] In some alternative implementations, step 203 may also include, for example: Figure 2D The following steps 2033 to 2035 are shown: Step 2033: In response to the current interaction state being an interrupted input state, calculate the information density of the voice data.
[0082] Here, the information density of voice data refers to the total amount of effective core information contained in voice data per unit time (or unit voice length). It is used to measure the information richness and key information ratio of user voice input, and to provide a basis for generating accurate and adapted voice responses in the future.
[0083] In practice, when calculating information density, natural language processing (NLP) tools can be used. First, speech data is converted into text using speech-to-text (ASR) tools. Then, keyword extraction tools (such as TF-IDF and TextRank algorithms) and semantic analysis models (such as BERT and RoBERTa) are used to extract core information and key phrases, count the number of effective information units, and calculate information density by combining speech duration (or text length). At the same time, redundant and repetitive content is filtered to ensure that the calculation results accurately reflect the concentration of effective information.
[0084] By calculating the information density of voice data during interrupted input, a quantitative assessment of the quality of user voice input is achieved. This accurately distinguishes between input scenarios with high and low information density, providing data support for the level of detail and core focus of subsequent response content. This avoids the problem of mismatch between response content and information density, making voice responses more targeted and efficient, and improving the accuracy of interaction and user satisfaction.
[0085] Step 2034: Obtain the current communication connection speed with the target voice input device.
[0086] Here, the current communication connection speed refers to the network transmission rate (including upload rate, download rate, and latency) between the target voice input device and the executing entity for real-time transmission of voice data, used to evaluate the stability of the communication link and the efficiency of data transmission.
[0087] In practice, when obtaining the current communication connection speed, network status detection tools (such as the NetworkInfo API on the device side, the ping command at the system level, and TCP / UDP bandwidth testing tools) can be used to collect indicators such as transmission latency, packet loss rate, and data transmission volume per unit time in real time, which can be converted into connection speed data at the Mbps (megabits per second) level. At the same time, the device signal strength (such as the number of Wi-Fi signal bars and the cellular network signal value) can be used to assist in the verification.
[0088] By leveraging the current communication connection speed, the system achieves real-time perception of the interactive link transmission capability, accurately identifying different network scenarios such as high-speed stability and low-speed lag. This provides a network adaptation basis for the generation of subsequent semantic transition voice response results (such as adjusting the voice compression ratio and controlling the amount of response data), avoiding response delays, lag, or transmission failures due to insufficient connection speed. This ensures the real-time performance and smoothness of voice responses, further improving the reliability of interaction and user experience under different network environments.
[0089] Step 2035: In response to the information density exceeding a preset information density threshold and the current communication connection speed exceeding a preset minimum communication connection speed threshold, generate a semantic transition speech response result for the speech data based on the user's emotion category and information density.
[0090] Here, the semantic transition voice response result refers to the voice response generated in scenarios where the user interrupts input, the voice information density is high, and the communication connection is stable. It has both emotional adaptation and semantic connection functions, and is used to carry over the core information expressed by the user and naturally transition to subsequent interactions, avoiding dialogue gaps.
[0091] In practice, the emotional tone of the response is first determined based on the user's emotional category (e.g., the tone is understanding and comforting when the user is angry, and recognition and agreement when the user is happy). Then, core keywords and key semantics are selected by combining high information density features. The emotional expression and semantic points are integrated through natural language generation (NLG) tools. The response is generated by adopting a structure of emotional continuation + core information extraction + transition guidance, while controlling the language to be concise and coherent, in line with the natural rhythm of oral communication.
[0092] By combining user emotion categories and information density to generate semantic transition voice response results in high information density and stable communication scenarios, the system achieves the dual effects of emotional empathy and semantic connection. This allows users to feel that their emotions are understood and core information is captured, while the natural transition promotes the orderly progress of the interaction. This avoids the emptiness of simple emotional feedback or the stiffness of mechanical semantic responses, thereby significantly improving the coherence and logic of human-computer dialogue, strengthening users' trust and satisfaction with the interaction, and making the response more targeted and practical.
[0093] In some alternative implementations, the target voice input device is via, for example... Figure 3 The target voice input device determination step 300 shown is predetermined and includes the following steps 301 to 303: Step 301: Obtain the preset priority and proximity status parameters of each voice input device.
[0094] Here, voice input devices encompass a wide range of voice input devices with voice acquisition and interaction capabilities. Specifically, these may include open / semi-open headphones equipped with array microphones, smartwatches with integrated voice acquisition modules, and smartphones with microphone components. They can also be extended to smart speakers, in-vehicle voice devices, and other devices that support voice input. All of these devices have the hardware foundation and data transmission capabilities to acquire user voice information in real time.
[0095] Preset priority refers to the device response priority sorting pre-set according to the usage scenario, permission level or user-defined rules of the voice input device (such as primary device higher than auxiliary device, frequently used device higher than temporary device). When setting, it can be completed through the system configuration interface or backend parameters in combination with factors such as device type, user frequency of use, and function permissions.
[0096] The proximity state parameter is used to characterize the physical distance between the voice input device and the user, as well as quantitative indicators of the signal (such as distance value, signal strength value, and sensing trigger level), and is used to determine whether the device is within the effective interaction range.
[0097] In practice, when obtaining the proximity status parameters of each voice input device, data is collected through the distance sensor (such as infrared or ultrasonic sensor), Bluetooth / Wi-Fi signal strength detection module, or positioning module built into the voice input device. The preset priority is obtained by reading the pre-configured device priority list or user-defined stored parameters.
[0098] By acquiring the preset priority and proximity state parameters of each voice input device, the effectiveness and response order of devices in multi-device interaction scenarios are accurately defined, providing a key basis for subsequent device selection and interaction scheduling. This avoids conflicts caused by simultaneous responses from multiple voice input devices, ensures that voice input devices within the effective interaction range and with high priority are activated first, and improves the orderliness, accuracy, and smoothness of multi-voice input device interaction.
[0099] Step 302: For each voice input device, the preset priority and proximity state parameters of the voice input device are weighted and summed to obtain the voice input device adaptation score of the voice input device.
[0100] Here, the voice input device adaptation score is a comprehensive quantitative value obtained by weighting and summing preset priority and neighboring state parameters. Its core purpose is to measure the degree of matching between the voice input device and the current user's interaction needs, providing a direct basis for selecting the optimal response device in multi-device scenarios.
[0101] In practice, first, based on actual interaction needs, assign reasonable weights to preset priorities (such as high weights for main devices and frequently used devices) and proximity state parameters (such as higher weights for closer devices). Then, multiply the preset priority quantization value and proximity state parameter quantization value of a single voice input device by their corresponding weights. Finally, add the two weighted results to obtain the voice input device adaptation score of the device.
[0102] By weighted summing of preset priority and proximity state parameters to obtain the voice input device compatibility score, an objective quantitative assessment of the matching degree of multiple devices is achieved, avoiding device selection bias caused by single-dimensional judgment. This enables the rapid selection of the voice input device that best matches the user's interaction needs, ensuring priority response to the voice input device with the highest compatibility, and reducing multi-device conflicts and invalid responses.
[0103] Step 303: Determine the voice input device with the highest voice input device adaptation score among all voice input devices as the target voice input device.
[0104] By identifying the voice input device with the highest compatibility score as the target voice input device, the optimal interaction device is accurately selected in multi-device scenarios. This directly locks the voice input device that best matches the user's needs, thereby reducing the situation where invalid devices occupy interaction resources.
[0105] The interactive device scheduling method provided in the above embodiments of this disclosure acquires voice data collected by the target voice input device in real time, performs emotion recognition on the collected voice data to obtain the user's emotion category, then performs voice activity detection on the collected voice data to obtain the user's current interaction state, and finally generates the voice response result of the target voice input device based on the user's emotion category and interaction state. In this way, by recognizing the user's emotion category, detecting the interaction state, and generating an appropriate voice response result in real time, intelligent scheduling and optimal selection of multiple voice input devices are achieved, enabling voice interaction to accurately adapt to the user's emotional state, providing the user with a response that meets their needs, effectively improving the accuracy and timeliness of voice interaction, and reducing the unpleasant experience caused by interaction interruption.
[0106] Further reference Figure 4 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an interactive device scheduling apparatus, which is similar to... Figure 2ACorresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0107] like Figure 4 As shown, the interactive device scheduling device 400 of this embodiment includes: a voice data acquisition unit 401, a detection unit 402, and a response unit 403. The voice data acquisition unit 401 is configured to acquire voice data collected by the target voice input device in real time, and perform emotion recognition on the collected voice data to obtain the user's emotion category; the detection unit 402 is configured to perform voice activity detection on the collected voice data to obtain the user's current interaction state; and the response unit 403 is configured to generate a voice response result for the target voice input device based on the user's emotion category and interaction state.
[0108] In this embodiment, the specific processing of the voice data acquisition unit 401, detection unit 402, and response unit 403 of the interactive device scheduling device 400 and the resulting technical effects can be referred to respectively. Figure 2A The relevant descriptions of steps 201, 202, and 203 in the corresponding embodiments will not be repeated here.
[0109] In some alternative implementations, the target voice input device is predetermined through the following target voice input device determination steps: Obtain the preset priority and proximity status parameters of each voice input device; For each voice input device, the preset priority and proximity state parameters of the voice input device are weighted and summed to obtain the voice input device adaptation score of the voice input device. The proximity state parameters are used to characterize the physical distance between the voice input device and the user. The voice input device with the highest voice input device compatibility score among all voice input devices is identified as the target voice input device.
[0110] In some alternative implementations, the voice data acquisition unit 401 may be further configured as follows: Extract speech features and interactive behavior features from the speech data of the target speech input device; Emotion recognition processing is performed on speech data based on speech features and interactive behavior features to obtain the emotion representation of the speech data. User emotion categories are generated from voice data based on emotion representations.
[0111] In some alternative implementations, the detection unit 402 may be further configured as follows: Extract speech pause data from the target speech input device; If the duration of the pause in the voice pause data exceeds the pause duration threshold, the current interaction state is determined to be an interrupted input state. If the duration of the pause in the voice pause data does not exceed the pause duration threshold, the current interaction state is determined to be a continuous input state.
[0112] In some alternative implementations, the response unit 403 may be further configured as follows: In response to the current interaction state being a continuous input state, the response emotion category is determined based on the user's emotion category; Generate non-repeating voice response results for voice data based on the response emotion category.
[0113] In some alternative implementations, the response unit 403 further includes: In response to the current interaction state being an interrupted input state, calculate the information density of the voice data; Obtain the current communication connection speed with the target voice input device; In response to an information density exceeding a preset information density threshold and a current communication connection speed exceeding a preset minimum communication connection speed threshold, a semantic transition speech response result is generated for the speech data based on the user's emotion category and information density.
[0114] It should be noted that the implementation details and technical effects of each module and unit in the interactive device scheduling device provided in the embodiments of this disclosure can be referred to the descriptions of other embodiments in this disclosure, and will not be repeated here.
[0115] The following is for reference. Figure 5 It shows a schematic diagram of the structure of a computer system 500 suitable for implementing the terminal device of this disclosure. Figure 5 The computer system 500 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0116] like Figure 5 As shown, the computer system 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the computer system 500. The processing device 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0117] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows computer system 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 A computer system 500 with various electronic devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0119] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0120] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0121] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following functions: Figure 2A The embodiments shown and their alternative implementations illustrate an interactive device scheduling method.
[0122] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0124] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not necessarily limiting in certain circumstances; for example, a voice data acquisition unit can also be described as "a unit that acquires voice data collected by a target voice input device in real time and performs emotion recognition on the collected voice data to obtain the user's emotion category."
[0125] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. A method for scheduling interactive devices, characterized in that, The method includes: The system acquires voice data collected by the target voice input device in real time, and performs emotion recognition on the collected voice data to obtain the user's emotion category. The collected voice data is subjected to voice activity detection to obtain the user's current interaction state; The voice response result of the target voice input device is generated based on the user's emotion category and the interaction state.
2. The method according to claim 1, characterized in that, The target voice input device is predetermined through the following target voice input device determination steps: Obtain the preset priority and proximity status parameters of each voice input device; For each of the voice input devices, a weighted sum of the preset priority and proximity state parameters of the voice input device is obtained to obtain the voice input device adaptation score of the voice input device. The proximity state parameters are used to characterize the physical distance between the voice input device and the user. The voice input device with the highest voice input device adaptation score among all the aforementioned voice input devices is determined as the target voice input device.
3. The method according to claim 1, characterized in that, The real-time acquisition of voice data collected by the target voice input device, and the performance of emotion recognition on the collected voice data to obtain the user's emotion category, includes: Extract the voice features and interaction behavior features of the voice data from the target voice input device; Based on the speech features and the interaction behavior features, the speech data is processed for emotion recognition to obtain the emotion representation of the speech data. The user emotion category is generated from the emotion representation of the voice data.
4. The method according to claim 1, characterized in that, The step of detecting voice activity in the collected voice data to obtain the user's current interaction state includes: Extract speech pause data from the speech data of the target speech input device; If the duration of the pause in the voice pause data exceeds the pause duration threshold, the current interaction state is determined to be an interrupted input state. If the duration of the pause in the voice pause data does not exceed the pause duration threshold, the current interaction state is determined to be a continuous input state.
5. The method according to claim 4, characterized in that, The process of generating the voice response result of the target voice input device based on the user's emotion category and the interaction state includes: In response to the current interaction state being the continuous input state, a response emotion category is determined based on the user emotion category; A non-overlapping voice response result is generated for the voice data based on the response emotion category.
6. The method according to claim 5, characterized in that, The step of generating the voice response result of the target voice input device based on the user's emotion category and the interaction state further includes: In response to the current interaction state being the interrupted input state, the information density of the voice data is calculated; Obtain the current communication connection speed with the target voice input device; In response to the information density exceeding a preset information density threshold and the current communication connection speed exceeding a preset minimum communication connection speed threshold, a semantic transition speech response result is generated for the speech data based on the user emotion category and the information density.
7. An interactive device scheduling apparatus, comprising: The voice data acquisition unit is used to acquire voice data collected by the target voice input device in real time, and to perform emotion recognition on the collected voice data to obtain the user's emotion category. The detection unit is used to detect voice activity in the collected voice data to obtain the user's current interaction state; A response unit is used to generate a voice response result of the target voice input device based on the user's emotion category and the interaction state.
8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.
10. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method as described in any one of claims 1-6.