Artificial intelligence-based robot intelligent question answering method and system, and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本申请提供基于人工智能的机器人智能问答方法、系统及存储介质,解决了现有技术在进行智能问答时,难以自适应地通过回答技巧缓解口吃患者紧张情绪的技术问题
1.本发明在整体架构上采用了“多模态感知—语义重构—自适应检测—个性化共情”的四阶段串联闭环设计,这一架构的设计逻辑并非简单地将语音截断、语义误解和交互冷漠三个问题各自独立解决,而是通过精心设计的数据流转关系,使每个阶段的输出自然成为下一阶段的关键输入,从而形成了一条从信号采集到情感回应的完整价值链。具体而言,步骤S1通过融合声学信息与嘴唇动作、呼吸频率等辅助识别信息对口吃类型进行精准识别,其产出的口吃类型标签和非流利语音片段直接为步骤S2的语义清洗提供了清洗策略的选择依据——不同口吃类型对应不同的冗余去除逻辑,音节重复型保留首个重复单元而删除后续重复,无声阻塞型则忽略静音时段而保留前后文本,清洗策略与识别结果深度绑定而非泛化处理;步骤S2输出的目标语义文本则为步骤S3的回复语音合成提供了内容基础;而步骤S3中计算出的卡顿影响因子同时驱动两个关键机制,一方面自适应调整语音活动检测阈值以解决截断问题,另一方面加载个性化交互数据以解决交互冷漠问题;最终,步骤S4基于自适应阈值完成语音检测并输出共情回复,而用户在感受到被理解和尊重后紧张情绪得以舒缓,口吃程度随之减轻,下一轮交互中的语音信号质量自然提升,从而形成“感知更准→理解更精→回应更暖→用户更放松→信号更清晰”的良性正向循环。这种闭环架构使得本发明不是停留在单点修复的层面,而是从系统工程的视角实现了对口吃用户交互体验的全链路优化,随着交互的持续迭代,各环节的参数和策略不断贴合用户的个体节律,真正实现了“越用越懂你,越用越流畅”的进化效果。
Smart Images

Figure CN122551791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to AI-based intelligent question-answering methods, systems, and storage media for robots. Background Technology
[0002] Currently, with the rapid development of artificial intelligence technology, intelligent chatbot question-and-answer systems have been widely applied in customer service, education, and healthcare. Existing intelligent question-and-answer systems typically rely on speech activity detection technology to determine the start and end points of a user's speech, and utilize automatic speech recognition and natural language processing technologies for semantic understanding and response generation. However, for users who stutter or have speech disorders, their speech streams are often accompanied by non-fluent phenomena such as syllable repetition, word repetition, prolonged sounds, or silent blockages. At the speech acquisition level, existing intelligent question-and-answer systems often use a fixed VAD threshold. When a user pauses or blocks for a long time, the system is prone to misjudging the end of the user's speech and prematurely truncating the speech, resulting in the inaccurate acquisition of the user's complete intent. At the semantic understanding level, existing semantic understanding models struggle to effectively identify and remove redundant and repetitive information from the speech of stuttering users, easily leading to semantic misunderstandings. More importantly, the speech disorder of stuttering users is not simply an abnormality in the speech signal, but a complex phenomenon closely related to their psychological state. Stuttering users generally experience high levels of anxiety and sensitivity to speech during interaction. Their stuttering severity fluctuates significantly with the pressure of the interaction environment. When faced with indifferent and impatient machine responses, users' tension is further aroused, and stuttering symptoms worsen, forming a vicious cycle of "increased interaction pressure → heightened tension → worsened stuttering → system misjudgment → more mechanical response".
[0003] However, existing intelligent question-answering systems employ a uniform, undifferentiated response pattern in their interaction strategies. They fail to consider the differences in users' stuttering styles, nor do they dynamically adjust the response speed, sentence complexity, or emotional support based on the user's current emotional state. Furthermore, they lack a personalized strategy learning mechanism based on the historical interaction experiences of similar users. This impersonal, "one-size-fits-all" approach not only fails to provide necessary psychological buffering and emotional support for stutterers but also continuously exacerbates their tension and frustration, resulting in a poor user experience. Consequently, many stutterers abandon voice interaction services.
[0004] Therefore, this invention discloses an AI-based intelligent question-answering method, system, and storage medium for robots, in order to solve the above-mentioned technical problems. Summary of the Invention
[0005] This application provides an AI-based robotic intelligent question-answering method, system, and storage medium, which solves the technical problem that existing technologies struggle to adaptively alleviate the anxiety of stuttering patients through answering techniques during intelligent question-answering.
[0006] To achieve the above objectives, this application adopts the following technical solution: Firstly, it provides AI-based intelligent question-answering methods for robots, including: S1: Obtain the user's original speech and auxiliary recognition information, and extract acoustic information and stuttering type from the original speech; among which, acoustic information includes non-fluent speech segments, abnormal word ratio and Chinese character spacing; stuttering type labels include at least one or more combinations of syllable repetition, word repetition, sound prolongation, and silent blockage; auxiliary recognition information includes the number of lip movements and breathing frequency; S2: Perform semantic cleaning and reconstruction on the original speech, remove redundant information marked as non-fluent speech segments, and generate target semantic text that conforms to grammatical norms by combining the context. S3: When the user's original speech contains at least one type of stuttering, acquire the user's existing acoustic information and auxiliary recognition information, identify the user's stuttering impact factor based on the acoustic information and auxiliary recognition information, determine the speech activity detection threshold based on the stuttering impact factor, load personalized interaction data associated with the user based on the stuttering impact factor and stuttering type label, and synthesize reply speech based on personalized interaction data and target semantic text; wherein, personalized interaction data includes speech rhythm and the proportion of concerned sentences; S4: Detects the user's original voice based on the voice activity detection threshold, and outputs the reply voice through the robot to complete the question-and-answer interaction.
[0007] In conjunction with the first aspect mentioned above, one possible implementation involves acquiring the user's original speech and auxiliary recognition information, and extracting acoustic information and stuttering type from the original speech, including: The system collects the user's original voice through a microphone array, and simultaneously acquires video stream data containing the user's facial features through an image acquisition device. The system extracts the user's lip movement trajectory features from the video stream to obtain the number of lip movements at each time point. The system also acquires the user's breathing frequency through an airflow sensor. The image acquisition device includes a high-definition RGB camera and a depth camera. The original speech is detected using a non-fluent speech detection model to identify non-fluent speech segments in the speech stream; the proportion of the duration of the same character in the non-fluent speech segment to the total duration of the corresponding sentence is calculated to obtain the abnormal character proportion; the silence duration between adjacent characters is calculated to obtain the Chinese character interval; the non-fluent speech segments, the abnormal character proportion, and the Chinese character interval are summarized into acoustic information; the acoustic information is time-aligned with auxiliary recognition information; wherein, the non-fluent speech detection model is implemented using a convolutional neural network and / or a time-delay neural network structure; A multimodal stuttering classification model is used to determine the obstruction pattern of the non-fluent speech segment and generate corresponding stuttering type labels. Specifically, when the same phoneme appears continuously within a preset short period of time, it is marked as syllable repetition; when a complete word appears continuously, it is marked as word repetition; when the duration of a vowel or consonant exceeds a preset duration threshold, it is marked as sound prolongation; when an acoustic signal interruption is detected and the lip movement trajectory features show that the user has made a vocalization action and the breathing frequency shows that the airflow is obstructed, it is marked as silent obstruction. The multimodal stuttering classification model is implemented through a bidirectional long short-term memory network, a temporal segmentation network, and / or a recurrent neural network structure.
[0008] In conjunction with the first aspect mentioned above, one possible implementation involves semantic cleaning and reconstruction of the original speech, removing redundant information marked as non-fluent speech segments, and generating grammatically correct target semantic text based on the context, including: The original speech is subjected to speech recognition to generate a preliminary text sequence containing words and timestamp information; wherein, the timestamp information is used to mark the start and end times of each word or syllable in the original speech; Align the timestamps of the non-fluent speech segments with the word timestamps in the preliminary text sequence to determine the positions of redundant words or abnormal characters corresponding to the non-fluent speech segments in the text; Based on the stuttering type tag, redundant words or abnormal characters are selectively removed to obtain the cleaned intermediate text. The specific cleaning logic includes: if the stuttering type tag is a syllable repetition or word repetition, the first repetitive unit is retained and subsequent repetitive words are deleted; if the stuttering type tag is a sound prolongation, the prolonged character is truncated to a single character corresponding to the standard pronunciation duration; if the stuttering type tag is a silent blockage, the silent period corresponding to the segment is ignored, and the text content before and after the segment is directly retained. The intermediate text is input into a pre-trained semantic reconstruction model. Combined with the context of historical dialogues, the cleaned text is subjected to grammatical error correction and semantic completion to obtain target semantic text that conforms to grammatical norms. The semantic reconstruction model is implemented using a Transformer encoder-decoder network structure to predict and fill in semantic related words that may be missing due to the removal of redundant information, and adjusts the word order to conform to grammatical norms, ultimately generating the target semantic text.
[0009] In conjunction with the first aspect mentioned above, one possible implementation involves identifying the user's stuttering impact factors based on acoustic information and auxiliary recognition information, including: The time ratio between the duration of existing non-fluent speech segments and the duration of the corresponding complete sentences is obtained, and the average value of the time ratio is marked as the basic stuttering factor. ; Get the average percentage of abnormal words in a single statement from a user's statements within the last n minutes. Spacing between Chinese characters Based on Chinese character spacing The interval regularity factor for each statement is determined by formula (1). Where n is manually set and can take a value of 15; Let be the numbers of each statement within the last n minutes, and The range of values is , The maximum value of statement numbers within the last n minutes; This is the number of the Chinese character in the current sentence, and The range of values is , This represents the maximum value of the Chinese character index in the current statement; Get the end time of each statement made by the user within the last n minutes. Based on end time The attenuation weight of each statement is determined by calculating formula (2) using auxiliary recognition information. ; The average percentage of abnormal words in each statement. Interval regularity factor and decay weight The emotional fluctuation blocking factor is determined by formula (3). ; through the basic stuttering factor and mood fluctuation blocking factors Determine the impact factor of stuttering Among them, the lag impact factor The formula for calculation is: , To adjust the parameters, they can be obtained by fitting multiple rounds of test data during the user's first use, or preset to [0.6, 0.8] based on experience, to balance the contribution of inherent defects and on-the-spot emotions to stuttering; stuttering influence factor It combines the effects of long-term physiological stuttering characteristics with short-term psychological fluctuations; The calculation formula (1) is: ; In the formula, The average value of the spacing between existing Chinese characters used by the user; The calculation formula (2) is: ; In the formula, For the current time, The standard duration is set based on experience. This is the attenuation coefficient used to adjust for the effect of time distance, and The value range is [0,1]; This represents the number of lip movements corresponding to the current statement. It is the average number of lip movements for all existing statements by the current user; This is the amplitude adjustment coefficient used to adjust the importance of the number of lip movements, and The value range is [0,1]; The breathing frequency corresponding to the current statement. It is the average breathing frequency of all statements currently in use by the user; This is the amplitude adjustment coefficient used to adjust the importance of respiratory rate, and The value range of is [0,1], and , and The ratio between any two is limited to the range [0.67, 1.5], used for limitation. , and The imbalance caused by any coefficient being too large; The calculation formula (3) is: ; In the formula, and Weighting coefficients, set according to different stuttering types, are used to balance the average proportion of abnormal words. and interval regularity factor The importance of setting stuttering type tags, especially for users whose stuttering is mainly manifested in syllable repetition and word repetition, is highlighted. To highlight the proportion of abnormal words; for users whose main symptoms are silent blocking and prolonged sound, adjust the settings. To highlight abnormal interval patterns; the system presets several sets of parameter combinations and automatically matches them according to the current user's stuttering type tag.
[0010] In conjunction with the first aspect mentioned above, one possible implementation involves determining the speech activity detection threshold based on the stuttering impact factor, including: Extracting the impact factor of stuttering Based on the lag impact factor Baseline silence tolerance range and average stuttering impact factor of other users in history. Determine the threshold for voice activity detection Among them, the voice activity detection threshold The formula for calculation is: ; This represents the median of the baseline noise tolerance range. This represents the maximum value of the baseline noise tolerance range. The baseline silence tolerance range is the minimum value of the inter-personal dialogue between two participants. Specifically, it is the duration of silence between the end of one person's speech and the beginning of the other's speech. Participants included both non-stutterers and stutterers. Natural speech data was collected, and the duration of silence in all dialogues was calculated and distributed from smallest to largest. The 30th percentile value was taken as the minimum value of the baseline silence tolerance range. The 70th percentile value is taken as the maximum value. The median as .
[0011] In conjunction with the first aspect above, in one possible implementation, personalized interaction data associated with the user is loaded based on the stuttering impact factor and stuttering type tag, including: Extract the lag impact factors closest to the current user's current time. And stuttering type tags, extract users with the same stuttering type tags as the current user from the historical data repository, and mark these users as similar users; obtain the stuttering influence factor of similar users in history, and place the stuttering influence factor of the current user at the same level. The historical impact factor within the fluctuation range is marked as the reference impact factor. Among them, the lag impact factor The fluctuation range is based on the current user's lag impact factor. Based on the preset fixed percentage Forming a relative fluctuation range Fixed percentage It is determined based on the number of users for the current type of tag, for example: , This represents the total number of similar users corresponding to the current stuttering type tag. For the preset scaling constant, such as ; The reference impact factor number; Extracting reference impact factor The corresponding next reference impact factor It will be greater than the next reference impact factor. Reference impact factor Integrate into target factors, and obtain each target factor and its corresponding next reference impact factor. Personalized interaction data of the robot over time; among which, personalized interaction data includes voice rhythm and the proportion of caring statements.
[0012] In conjunction with the first aspect mentioned above, one possible implementation involves synthesizing a response speech based on personalized interaction data and target semantic text, including: Several sets of speech rhythms, proportions of relevant statements, target semantic text, and response speech were extracted from a speech reference database. The speech reference database includes several speech rhythms, proportions of relevant statements, target semantic text, and response speech set by experts based on the speech rhythms, proportions of relevant statements, and target semantic text. The speech rhythm, the proportion of relevant sentences, the target semantic text, and the response speech are integrated into several sets of training data and test data. The training data is used to train the artificial intelligence model, and the test data is used to test the trained artificial intelligence model. The artificial intelligence model is adjusted according to the test results. Finally, a speech synthesis model is obtained with the input of speech rhythm, the proportion of relevant sentences, and the target semantic text, and the output of the response speech. The artificial intelligence model is a non-linear regression model based on neural networks, and is implemented using a BP neural network structure and / or an RBF neural network structure. The speech rhythm and proportion of caring statements in the personalized interaction data corresponding to the current user, as well as the target semantic text, are input into the speech synthesis model to obtain the response speech.
[0013] In conjunction with the first aspect above, one possible implementation involves detecting the user's original speech based on a speech activity detection threshold, including: Extract the voice activity detection threshold corresponding to the current user. When the user's pause time exceeds the voice activity detection threshold, identify the received original voice and generate a response voice.
[0014] Secondly, this application provides an AI-based intelligent question-answering system for robots, including: a communication unit and a processing unit; The communication unit is used to acquire the user's original speech and auxiliary recognition information, and to extract acoustic information and stuttering type from the original speech; wherein, the acoustic information includes non-fluent speech segments, abnormal word ratio, and Chinese character spacing; the stuttering type label includes at least one or more combinations of syllable repetition, word repetition, sound prolongation, and silent blockage; the auxiliary recognition information includes the number of lip movements and breathing frequency; The processing unit is used to perform semantic cleaning and reconstruction on the original speech, remove redundant information marked as non-fluent speech segments, and generate target semantic text that conforms to grammatical norms by combining the context. When the user's original speech contains at least one type of stuttering, it acquires the user's existing acoustic information and auxiliary recognition information, identifies the user's stuttering influence factor based on the acoustic information and auxiliary recognition information, and determines the speech activity detection threshold based on the stuttering influence factor. Based on the stuttering influence factor and stuttering type label, it loads personalized interaction data associated with the user, synthesizes reply speech based on personalized interaction data and target semantic text, detects the user's original speech based on the speech activity detection threshold, and outputs the reply speech through the robot to complete the question-and-answer interaction.
[0015] Thirdly, this application provides a storage medium storing instructions that, when executed on an AI-based robotic intelligent question-answering system, cause the AI-based robotic intelligent question-answering system to perform the methods described in the first aspect and any possible implementation thereof.
[0016] This application provides an AI-based intelligent question-answering method, system, and storage medium for robots, with the following benefits: 1. The present invention adopts a four-stage serial closed-loop design of "multimodal perception - semantic reconstruction - adaptive detection - personalized empathy" in its overall architecture. The design logic of this architecture is not to solve the three problems of speech truncation, semantic misunderstanding and interactive indifference independently, but to make the output of each stage naturally become the key input of the next stage through carefully designed data flow relationship, thus forming a complete value chain from signal acquisition to emotional response. Specifically, step S1 accurately identifies stuttering types by fusing acoustic information with auxiliary recognition information such as lip movements and breathing frequency. The resulting stuttering type labels and non-fluent speech segments directly provide the basis for selecting cleaning strategies for semantic cleaning in step S2. Different stuttering types correspond to different redundancy removal logics. For syllable repetition types, the first repetition unit is retained while subsequent repetitions are deleted. For silent blocking types, the silent periods are ignored while the preceding and following texts are retained. The cleaning strategy is deeply bound to the recognition results rather than being a generalized process. The target semantic text output by step S2 provides the content foundation for the response speech synthesis in step S3. The stuttering impact factor calculated in step S3 drives two key mechanisms: on the one hand, it adaptively adjusts the speech activity detection threshold to solve the truncation problem; on the other hand, it loads personalized interaction data to solve the problem of cold interaction. Finally, step S4 completes speech detection based on the adaptive threshold and outputs an empathetic response. After the user feels understood and respected, their tension is relieved, the degree of stuttering is reduced, and the quality of the speech signal in the next round of interaction is naturally improved, thus forming a virtuous cycle of "more accurate perception → more precise understanding → warmer response → more relaxed user → clearer signal". This closed-loop architecture allows the invention to go beyond single-point repair, achieving end-to-end optimization of the user experience for stutterers from a systems engineering perspective. As the interaction continues to iterate, the parameters and strategies of each link constantly adapt to the user's individual rhythm, truly achieving an evolutionary effect of "the more you use it, the better it understands you, and the smoother it becomes."
[0017] 2. This invention embodies the innovative ideas of "decoupling and integrating physiological baseline and psychological fluctuations, adaptive perception of stuttering type, and multimodal temporal synergy" in the design concept of the stuttering influencing factor. Traditional solutions treat stuttering as a single abnormality in speech signals and treat all users with fixed parameters in a one-size-fits-all manner. This invention, however, fundamentally re-examines the nature of stuttering, recognizing that it is not a one-dimensional phenomenon but a complex state determined by the user's long-term inherent physiological speech impairment characteristics and short-term psychological fluctuations driven by the dialogue context. Based on this understanding, this invention creatively decomposes the stuttering influencing factor into two sub-dimensions: one characterizes the severity of the user's long-term stable stuttering baseline, acting as a "baseline calibrator" to ensure that the system always remembers the user's inherent impairment level and is not misled by single fluctuations; the other characterizes the user's emotional fluctuations and tension level in the current dialogue, acting as a "context regulator" to capture changes in pre-context states such as anxiety and tension in real time. The two sub-dimensions are then fused into a unified factor through adjustable weights, thereby simultaneously achieving the stability of long-term memory and the sensitivity of short-term response within the same variable. This "decoupling before fusion" design concept enables the system to accurately identify the user's inherent stuttering patterns while also sensitively detecting temporary fluctuations exacerbated by interactive pressure during dialogue. This fundamentally avoids the dilemma of traditional solutions that either "only look at the past and ignore the present" or "only look at the present and forget the past." At the feature selection level, this invention breaks through the inherent limitations of relying solely on acoustic information of speech by incorporating auxiliary recognition information such as the number of lip movements and breathing frequency into the factor construction. The design concept is that stuttering is not only manifested as speech interruptions or repetitions but is also often accompanied by abnormal facial muscle movements and disordered breathing rhythms. These non-acoustic signals are precisely important external signs of the user being in a high-stress state. Incorporating them into the factor calculation allows the system to automatically enhance its ability to perceive high-anxiety states from a multimodal perspective. In terms of time dimension, this invention introduces a "recency priority" attenuation mechanism, inspired by the psychological recency effect. A user's recent stuttering performance better reflects their current emotional state and language organization ability. Observations closer to the current moment have a greater weight in influencing factors, ensuring that factors are always synchronized with the user's real-time state rather than lagging behind historical averages. In terms of type adaptation, this invention designs differentiated feature emphasis strategies for different stuttering types. For users primarily exhibiting syllable and word repetition, the proportion of abnormal words better reveals the nature of their stuttering; for users primarily exhibiting silent blocking and sound prolongation, the regularity of word spacing anomalies is more prominent. The factor calculation logic automatically switches the emphasis dimension according to the stuttering phenotype, rather than applying the same weight combination to all users. This concept significantly improves the accuracy of factor adaptation for different subtypes of stuttering users. In terms of balance constraints, this invention sets proportional constraints on the influence of each dimension, forcibly preventing any dimension—time decay, lip movements, or breathing frequency—from becoming excessively dominant, ensuring the balance of multi-source information during the fusion process.In terms of application, the stuttering impact factor serves as the core of this invention, simultaneously driving the adaptive adjustment of the voice activity detection threshold and the precise loading of personalized interactive data. This factor connects the two dimensions of "detection" and "interaction," making the voice detection threshold no longer a fixed parameter but a personalized value that dynamically floats with the user's real-time status. The response voice is no longer a standardized output but an empathetic response where the speech rate, sentence complexity, and emotional support are all adaptively adjusted according to the degree of stuttering, achieving a fundamental transformation from "one-size-fits-all" to "one-person-one-custom."
[0018] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0019] Figure 1 A schematic diagram illustrating the steps of an AI-based intelligent question-answering method for robots provided in this application embodiment; Figure 2 A schematic diagram illustrating the steps for identifying user lag impact factors provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of an AI-based intelligent question-answering system provided in an embodiment of this application. Detailed Implementation
[0020] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0021] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0022] like Figure 1 As shown in the embodiments of this application, the AI-based intelligent question-answering method for robots includes: S1: Obtain the user's original speech and auxiliary recognition information, and extract acoustic information and stuttering type from the original speech; among which, acoustic information includes non-fluent speech segments, abnormal word ratio and Chinese character spacing; stuttering type labels include at least one or more combinations of syllable repetition, word repetition, sound prolongation, and silent blockage; auxiliary recognition information includes the number of lip movements and breathing frequency; S2: Perform semantic cleaning and reconstruction on the original speech, remove redundant information marked as non-fluent speech segments, and generate target semantic text that conforms to grammatical norms by combining the context. S3: When the user's original speech contains at least one type of stuttering, acquire the user's existing acoustic information and auxiliary recognition information, identify the user's stuttering impact factor based on the acoustic information and auxiliary recognition information, determine the speech activity detection threshold based on the stuttering impact factor, load personalized interaction data associated with the user based on the stuttering impact factor and stuttering type label, and synthesize reply speech based on personalized interaction data and target semantic text; wherein, personalized interaction data includes speech rhythm and the proportion of concerned sentences; S4: Detects the user's original voice based on the voice activity detection threshold, and outputs the reply voice through the robot to complete the question-and-answer interaction.
[0023] It is worth noting that this invention innovatively solves the prominent problems of speech truncation, semantic misunderstanding, and cold interaction in existing intelligent question-answering systems when serving users with stuttering and speech disorders by integrating multimodal perception, semantic reconstruction, and adaptive interaction mechanisms. First, this invention calculates a dynamic and personalized stuttering impact factor. This factor essentially characterizes the severity and pattern of the user's current speech stuttering, thereby driving adaptive adjustment of the speech activity detection threshold. For example, when the system detects that the user's lips are still moving and their breathing exhibits the rapid rhythm characteristic of pre-vocalization, but there is a silent blockage at the audio end, it will automatically extend the waiting time according to the stuttering impact factor, tolerating a longer period of silence without triggering truncation. This mechanism completely reverses the situation of "machines not waiting for stutterers," ensuring that the user's complete pronunciation and true intention are preserved. The complete collection of redundant information such as syllable repetition and inserted phrases provides ample raw data for subsequent cleaning and reconstruction, improving recognition integrity several times over.
[0024] Furthermore, this invention creates a deeply personalized interactive empathy experience, significantly reducing users' psychological barriers to interaction. Stuttering is not a singular phenomenon; different individuals or situations exhibit significant differences in stuttering types. Some people exhibit syllable repetition, while others primarily experience silent blocking. Their psychological sensitivity and interaction preferences also vary. In step S3 of this invention, personalized interaction data associated with the user is loaded based on stuttering influencing factors and specific stuttering type labels. This data can encompass customized response speech rate, pause strategies, encouraging feedback phrases, and even soothing background audio. When the system synthesizes a response speech, it does not coldly output standard sentences but rather plays them at a pace adapted to the user's tolerance. For example, for anxious, repetitive stuttering users, the system automatically slows down the speech rate, uses shorter sentences, and appropriately inserts soothing natural language fragments such as "Don't rush, please speak slowly" into the speech stream. This empathetic interactive loop provides significant emotional support beyond its functional aspects, making users feel respected and understood, easing their tension, and in turn reducing their stuttering, thus creating a positive cycle that significantly enhances user stickiness and willingness to continue using the service.
[0025] Finally, this invention forms a dynamic closed loop of "detection—cleaning—adaptation—personalized empathy." As interaction continues, the system iteratively updates the user's stuttering impact factors and interaction preference data, making the VAD threshold, semantic cleaning strategy, and personalized response style increasingly aligned with the user's individual rhythm and habits, achieving an evolutionary effect of "the more you use it, the better it understands you; the more you use it, the smoother it becomes." In summary, this invention not only provides a high-precision voice question-answering technology solution but also uses artificial intelligence to build an accessible and warm communication bridge for the stuttering community, possessing both significant industrial application value and profound social and humanistic care.
[0026] In one possible implementation of this application embodiment, the above-mentioned S1 can be implemented by the following S101, which will be described in detail below: S101: Acquire the user's original speech and auxiliary recognition information, extract acoustic information and stuttering type from the original speech, including: The system captures the user's raw voice using a microphone array, and simultaneously acquires video stream data containing the user's facial features using an image acquisition device. The system extracts the user's lip movement trajectory features from the video stream to determine the number of lip movements at various time points. A breath sensor is used to acquire the user's breathing rate. The image acquisition device includes a high-definition RGB camera and a depth camera. It should be understood that, in addition to RGB and depth cameras, infrared cameras or thermal imaging devices can also be used in other embodiments to capture facial features, as long as the lip movement trajectory features can be extracted.
[0027] The original speech is detected using a non-fluent speech detection model to identify non-fluent speech segments in the speech stream; the proportion of the duration of the same character in the non-fluent speech segment to the total duration of the corresponding sentence is calculated to obtain the abnormal character proportion; the silence duration between adjacent characters is calculated to obtain the Chinese character interval; the non-fluent speech segments, the abnormal character proportion, and the Chinese character interval are summarized into acoustic information; the acoustic information is time-aligned with auxiliary recognition information; wherein, the non-fluent speech detection model is implemented using a convolutional neural network and / or a time-delay neural network structure; A multimodal stuttering classification model is used to determine the obstruction pattern of the non-fluent speech segment and generate corresponding stuttering type labels. Specifically, when the same phoneme appears continuously within a preset short period of time, it is marked as syllable repetition; when a complete word appears continuously, it is marked as word repetition; when the duration of a vowel or consonant exceeds a preset duration threshold, it is marked as sound prolongation; when an acoustic signal interruption is detected and the lip movement trajectory features show that the user has made a vocalization action and the breathing frequency shows that the airflow is obstructed, it is marked as silent obstruction. The multimodal stuttering classification model is implemented through a bidirectional long short-term memory network, a temporal segmentation network, and / or a recurrent neural network structure.
[0028] It should be noted that the percentage of time occupied by the same word in the non-fluent speech segment is calculated as the proportion of the total time of the corresponding sentence. The time occupied by the same word here includes: the time when the same word is repeated in a sentence, for example, in "I...I...I want", the time between the first "I" and "want" is the duration; and the time between the same word and the nearest non-same word in a sentence, for example, in "I...(pause)...want", the time between the first "I" and "want" is the duration, including the pause duration.
[0029] In one possible implementation of this application embodiment, the above-mentioned S2 can be implemented by the following S201, which will be described in detail below: S201: Perform semantic cleaning and reconstruction on the original speech, remove redundant information marked as non-fluent speech segments, and generate target semantic text that conforms to grammatical rules by combining the context, including: The original speech is subjected to speech recognition to generate a preliminary text sequence containing words and timestamp information; wherein, the timestamp information is used to mark the start and end times of each word or syllable in the original speech; Align the timestamps of the non-fluent speech segments with the word timestamps in the preliminary text sequence to determine the positions of redundant words or abnormal characters corresponding to the non-fluent speech segments in the text; Based on the stuttering type tag, redundant words or abnormal characters are selectively removed to obtain the cleaned intermediate text. The specific cleaning logic includes: if the stuttering type tag is a syllable repetition or word repetition, the first repetitive unit is retained and subsequent repetitive words are deleted; if the stuttering type tag is a sound prolongation, the prolonged character is truncated to a single character corresponding to the standard pronunciation duration; if the stuttering type tag is a silent blockage, the silent period corresponding to the segment is ignored, and the text content before and after the segment is directly retained. The intermediate text is input into a pre-trained semantic reconstruction model. Combined with the context of historical dialogues, the cleaned text is subjected to grammatical error correction and semantic completion to obtain target semantic text that conforms to grammatical norms. The semantic reconstruction model is implemented using a Transformer encoder-decoder network structure to predict and fill in semantic related words that may be missing due to the removal of redundant information, and adjusts the word order to conform to grammatical norms, ultimately generating the target semantic text.
[0030] In one possible implementation of this application embodiment, the above-mentioned S3 can be implemented by the following S301, S302, S303 and S304, which are described in detail below: like Figure 2 As shown, S301: Identifies the user's stuttering impact factors based on acoustic information and auxiliary recognition information, including: The time ratio between the duration of existing non-fluent speech segments and the duration of the corresponding complete sentences is obtained, and the average value of the time ratio is marked as the basic stuttering factor. Basic stuttering factor It reflects the user's long-term, stable stuttering severity, and is equivalent to a "baseline calibrator"; Get the average percentage of abnormal words in a single statement from a user's statements within the last n minutes. Spacing between Chinese characters Based on Chinese character spacing The interval regularity factor for each statement is determined by formula (1). Where n is manually set and can take a value of 15; Let be the numbers of each statement within the last n minutes, and The range of values is , The maximum value of statement numbers within the last n minutes; This is the number of the Chinese character in the current sentence, and The range of values is , This represents the maximum value of the Chinese character index in the current statement; Get the end time of each statement made by the user within the last n minutes. Based on end time The attenuation weight of each statement is determined by calculating formula (2) using auxiliary recognition information. Decay weight The design incorporates time decay factors and multimodal feature deviations, utilizing the "recency effect" to make statements closer to the current moment contribute more to the stuttering impact factor. At the same time, it enhances the response to high stress states by measuring the deviation of lip movements and breathing frequency from the baseline. The average percentage of abnormal words in each statement. Interval regularity factor and decay weight The emotional fluctuation blocking factor is determined by formula (3). Emotional fluctuation blocking factor By capturing users' short-term states such as anxiety and tension in real time through multiple dynamic observation indicators, it is equivalent to a "situation regulator"; By basic stuttering factor and mood fluctuation blocking factors Determine the impact factor of stuttering Among them, the lag impact factor The formula for calculation is: , To adjust the parameters, they can be obtained by fitting multiple rounds of test data during the user's first use, or preset to [0.6, 0.8] based on experience, to balance the contribution of inherent defects and on-the-spot emotions to stuttering; stuttering influence factor It combines the effects of long-term physiological stuttering characteristics with short-term psychological fluctuations; The calculation formula (1) is: ; In the formula, The average value of the spacing between existing Chinese characters used by the user; The calculation formula (2) is: ; In the formula, For the current time, The standard duration is set based on experience. This is the attenuation coefficient used to adjust for the effect of time distance, and The value range is [0,1]; This represents the number of lip movements corresponding to the current statement. It is the average number of lip movements for all existing statements by the current user; This is the amplitude adjustment coefficient used to adjust the importance of the number of lip movements, and The value range is [0,1]; The breathing frequency corresponding to the current statement. It is the average breathing frequency of all statements currently in use by the user; This is the amplitude adjustment coefficient used to adjust the importance of respiratory rate, and The value range of is [0,1], and , and The ratio between any two is limited to the range [0.67, 1.5], used for limitation. , and The imbalance caused by any coefficient being too large; The calculation formula (3) is: ; In the formula, and Weighting coefficients, set according to different stuttering types, are used to balance the average proportion of abnormal words. and interval regularity factor The importance of setting stuttering type tags, especially for users whose stuttering is mainly manifested in syllable repetition and word repetition, is highlighted. To highlight the proportion of abnormal words; for users whose main symptoms are silent blocking and prolonged sound, adjust the settings. To highlight abnormal interval patterns; the system presets several sets of parameter combinations and automatically matches them according to the current user's stuttering type tag.
[0031] It is worth noting that this step constructs a "physiological" A stuttering influencing factor calculation framework of "parallel psychological dual-factor and type-aware adaptive weighting" is proposed. This design does not simply treat stuttering as a single-dimensional abnormality of speech signal, but rather decouples and then fuses the user's inherent physiological stuttering characteristics with psychological fluctuations driven by present-day emotions. The beneficial effect of this hierarchical design is that the basic stuttering factor reflects the user's long-term, stable stuttering severity, equivalent to a "baseline calibrator"; while the emotional fluctuation blocking factor captures the user's short-term states such as anxiety and tension in real time through multiple dynamic observation indicators, equivalent to a "situational regulator." Both are integrated through adjustable parameters. By employing weighted fusion, the entire model can maintain a long-term memory of individual characteristics while also responding sensitively to state changes during the dialogue process, thus achieving a balance between stability and sensitivity in the design.
[0032] In feature selection and fusion design, this invention breaks through the limitations of relying solely on acoustic information, creatively incorporating auxiliary recognition information such as the number of lip movements and breathing frequency into the modeling. Stuttering is not only manifested as interruptions or repetitions in speech, but is also often accompanied by abnormal facial muscle movements and changes in breathing rhythm. This auxiliary information is integrated into the decay weights in the form of an exponential weight function. The design utilizes a nonlinear mapping relationship where "the greater the deviation from the norm, the stronger its contribution to current emotional fluctuations," enabling the model to automatically enhance its response to high-stress states from multimodal signals. Simultaneously, the design normalizes the number of lip movements and respiratory rate relative to their respective historical averages, eliminating dimensional differences between different users and giving the module cross-user versatility.
[0033] At the temporal modeling design level, this invention introduces a weight calculation method based on time decay. This is achieved by setting the current time... With statement end time The time difference, and using the exponential decay factor This design makes the contribution of sentences closer to the current moment to the stuttering impact factor greater. This design is faithful to the psychological "recency effect," that is, a user's recent stuttering performance is more likely to reflect their current emotional fluctuations and language organization ability. It is combined with the deviation terms of lip movements and breathing frequency in the same weighting formula (2), and restrictions are imposed. The ratios among the three are in the range of [0.67, 1.5]. The design forcibly avoids any one type of feature from excessively dominating the weights, ensuring the balanced fusion of multi-source information in the time dimension.
[0034] Another design innovation is the parameter adaptation mechanism for different stuttering types. This includes addressing the emotional fluctuation blocking factor. In the calculation, the system automatically matches based on the user's stuttering type tags. and The weighting combination is as follows. For users whose speech patterns primarily involve syllable repetition and word repetition, the percentage of abnormal characters better reflects the nature of their lag; therefore, the design... For users who primarily experience silent blocking and prolonged sound, the abnormal regularity of Chinese character spacing is even more pronounced. This design makes the entire logic for identifying stuttering influencing factors no longer a "one-size-fits-all" approach, but rather deeply tied to the user's stuttering phenotype, significantly improving the algorithm's adaptability and fitting accuracy for different subtypes of stuttering users.
[0035] In summary, this step achieves multiple innovations at the design level, including multimodal feature decoupling and fusion, adaptive temporal decay, type-aware weight self-matching, and coefficient balance constraints. These designs together ensure that the stuttering impact factor can objectively quantify inherent stuttering defects and dynamically capture on-site emotional fluctuations, laying a solid design foundation for the personalized setting of subsequent speech activity detection thresholds and the adjustment of the naturalness of response speech.
[0036] It should be noted that if there are no abnormal word proportions or Chinese character intervals in the user's recent n minutes, the standard emotional fluctuation blocking factor is used. The standard emotional fluctuation blocking factor is determined based on the historical data corresponding to the user's current stuttering type label. Specifically, other users corresponding to the user's current stuttering type label are marked as reference users, the ratio of the reference user's emotional fluctuation blocking factor to the corresponding basic stuttering factor in history is marked as the reference ratio, the average of several reference ratios is marked as the feature reference ratio, and the standard emotional fluctuation blocking factor is obtained by multiplying the feature reference ratio by the current user's basic stuttering factor.
[0037] S302: Determine the speech activity detection threshold based on the stuttering impact factor, including: Extracting the impact factor of stuttering Based on the lag impact factor Baseline silence tolerance range and average stuttering impact factor of other users in history. Determine the threshold for voice activity detection Among them, the voice activity detection threshold The formula for calculation is: ; This represents the median of the baseline noise tolerance range. This represents the maximum value of the baseline noise tolerance range. The baseline silence tolerance range is the minimum value of the inter-personal dialogue between two participants. Specifically, it is the duration of silence between the end of one person's speech and the beginning of the other's speech. Participants included both non-stutterers and stutterers. Natural speech data was collected, and the duration of silence in all dialogues was calculated and distributed from smallest to largest. The 30th percentile value was taken as the minimum value of the baseline silence tolerance range. The 70th percentile value is taken as the maximum value. The median as It should be understood that the speech activity detection threshold... This is a key parameter for the system to determine whether a user has finished speaking. Traditional intelligent question-answering systems typically use a fixed threshold, such as setting a silence period of more than 500 milliseconds to indicate the end of speech. However, for users who stutter, their speech is often accompanied by longer pauses or blockages. A fixed threshold can easily lead to misjudgment by the system, prematurely truncating the user's speech and thus losing crucial information. This embodiment introduces a dynamic adjustment mechanism, allowing the threshold to adaptively adjust according to the user's real-time status.
[0038] It's worth noting that this step breaks away from traditional methods that rely on fixed thresholds or single user histories, introducing the ratio of the "lag impact factor" to the "group average lag factor" as the core of dynamic adjustment. When a user's current lag level is higher than the group average, the system automatically adjusts the baseline mute tolerance range to the median. Adjusting the window upwards lengthens the silence detection window, preventing stuttering-related pauses from being misinterpreted as the end of speech; conversely, adjusting it downwards shortens the window to improve interaction efficiency. Simultaneously, through... Limiting the threshold clamp to within the baseline tolerance range ensures sensitivity while avoiding interference from extreme values. This design allows the robot to adaptively adjust the sentence segmentation sensitivity based on the user's real-time stuttering severity, significantly reducing accidental interruptions or unnecessary waiting during conversations.
[0039] It should be noted that if the average stuttering impact factor If it is 0, then let .
[0040] S303: Load personalized interaction data associated with the user based on the stuttering impact factor and stuttering type label, including: Extract the lag impact factors closest to the current user's current time. The system uses stuttering type tags to extract users with the same stuttering type tags as the current user from a historical data repository and marks these users as similar users. It should be understood that stuttering type tags are an important basis for classifying user interaction patterns. For example, "syllable repetition type" users are often more sensitive to speech speed, while "silent blocking type" users may need to wait patiently for prompts. By filtering by type, a group with similar interactive psychological characteristics can be identified, avoiding interference from invalid data between different types of users. Obtain the lag impact factor of similar users in history, and then select the lag impact factor of the current user. The historical impact factor within the fluctuation range is marked as the reference impact factor. Among them, the lag impact factor The fluctuation range is based on the current user's lag impact factor. Based on the preset fixed percentage Forming a relative fluctuation range Fixed percentage It is determined based on the number of users for the current type of tag, for example: , This represents the total number of similar users corresponding to the current stuttering type tag. For the preset scaling constant, such as ; The reference impact factor number; Extracting reference impact factor The corresponding next reference impact factor It will be greater than the next reference impact factor. Reference impact factor Integrate into target factors, and obtain each target factor and its corresponding next reference impact factor. The system collects personalized interaction data from the robot over time, including voice rhythm and the proportion of caring statements. The core of this filtering logic is to find "positive feedback cases": users with high stuttering factors (poor performance) in the past who, after using a specific robot interaction strategy, showed a decrease in stuttering factors in the next moment. By learning from these historical interaction data that successfully alleviated user anxiety, the system can develop the most effective reassurance strategy for the current user.
[0041] It is worth noting that this step employs a dual screening strategy of "same type + similar stuttering factors". First, it extracts similar user groups with the same stuttering type tag as the current user. Then, based on the current user's fluctuation range, it filters out reference factors whose historical stuttering factors fall within this range. Then, the next time-step factors corresponding to these reference factors are extracted. (Representing the trend of state change) greater than The core of this design lies in its association with personalized interaction data. It not only relies on the current state but also references the historical patterns of similar users experiencing similar lag, learning effective robot interaction strategies from these patterns. It adaptively adjusts with the total number of similar users, expanding the search scope when the number of users is small and finely targeting when the number of users is large, ensuring robustness in cold start and large-scale scenarios.
[0042] It should be noted that the impact factor is used as a reference. The corresponding next reference impact factor It is the current reference impact factor After the corresponding voice reply, the user should refer to the impact factor when responding; if the reference impact factor is... and the corresponding next reference impact factor If the time interval between them is greater than the interval threshold, the current reference impact factor will be removed in subsequent analyses. The interval threshold is set manually, and is usually set to 10 minutes.
[0043] It should be noted that the speech rhythm refers to the average speech rate of the response, and the proportion of caring statements refers to the percentage of characters in encouraging or comforting statements embedded in the target semantic text relative to the total number of characters in the response.
[0044] S304: Synthesize response speech based on personalized interaction data and target semantic text, including: Several sets of speech rhythms, proportions of concern statements, target semantic text, and response speech were extracted from a speech reference database. The speech reference database includes several speech rhythms, proportions of concern statements, target semantic text, and response speech set by experts based on the speech rhythms, proportions of concern statements, and target semantic text. It should be understood that "expert settings" here refers to response schemes carefully designed by psychology experts and senior customer service personnel for specific pause states and semantic content. For example, a slow-speed response designed for anxious users, or a response containing encouraging words designed for blocked users. These high-quality samples constitute the "gold standard" for model training. The speech rhythm, the proportion of relevant sentences, the target semantic text, and the response speech are integrated into several sets of training data and test data. The training data is used to train the artificial intelligence model, and the test data is used to test the trained artificial intelligence model. The artificial intelligence model is adjusted according to the test results. Finally, a speech synthesis model is obtained with the input of speech rhythm, the proportion of relevant sentences, and the target semantic text, and the output of the response speech. The artificial intelligence model is a non-linear regression model based on neural networks, and is implemented using a BP neural network structure and / or an RBF neural network structure. The speech rhythm and proportion of caring statements in the personalized interaction data corresponding to the current user, as well as the target semantic text, are input into the speech synthesis model to obtain the response speech.
[0045] In one possible implementation of this application embodiment, the above-mentioned S4 can be implemented by the following S401, which will be described in detail below: S401: Detects the user's raw speech based on a speech activity detection threshold, including: The system extracts the speech activity detection threshold for the current user. When the user's pause time exceeds the threshold, it identifies the received raw speech and generates a response speech. This step ensures that the response is timely, neither interrupting the user's difficult expression nor waiting unnecessarily after the user has finished speaking. This embodiment achieves a qualitative leap from "mechanical response" to "empathic interaction" by constructing a mechanism of "learning from historical success cases + neural network speech synthesis," significantly improving the interactive experience and psychological comfort of stuttering users.
[0046] The foregoing mainly describes the solutions of the embodiments of this application from the perspective of device implementation. It is understood that each device, such as an AI-based intelligent question-answering system, includes at least one of the hardware structures and software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0047] This application embodiment can divide the AI-based robot intelligent question-answering system into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0048] When using integrated units, Figure 3 A possible structural schematic diagram of the AI-based robot intelligent question-answering system (referred to as communication system 30) involved in the above embodiments is shown. The communication system 30 includes a processing unit 301 and a communication unit 302, and may also include a storage medium (referred to as storage unit 303). Figure 3 The structural diagram shown can be used to illustrate the structure of the AI-based intelligent question-and-answer system for robots involved in the above embodiments.
[0049] when Figure 3 The structural diagram shown is used to illustrate the structure of the AI-based robot intelligent question-and-answer system involved in the above embodiments. The processing unit 301 is used to control and manage the actions of the AI-based robot intelligent question-and-answer system, the communication unit 302 is used for the AI-based robot intelligent question-and-answer system to communicate with other devices, and the storage unit 303 is used to store the program code and data of the AI-based robot intelligent question-and-answer system.
[0050] For example, communication unit 302 is used to acquire the user's original speech and auxiliary recognition information, and extract acoustic information and stuttering type from the original speech; wherein, the acoustic information includes non-fluent speech segments, abnormal word ratio and Chinese character spacing; the stuttering type label includes at least one or more combinations of syllable repetition, word repetition, sound prolongation, and silent blockage; the auxiliary recognition information includes the number of lip movements and breathing frequency; Processing unit 301: performs semantic cleaning and reconstruction on the original speech, removes redundant information marked as non-fluent speech segments, and generates target semantic text that conforms to grammatical norms by combining the context; when the user's original speech contains at least one type of stuttering, it acquires the user's existing acoustic information and auxiliary recognition information, identifies the user's stuttering influence factor based on the acoustic information and auxiliary recognition information, and determines the speech activity detection threshold based on the stuttering influence factor; loads personalized interaction data associated with the user based on the stuttering influence factor and stuttering type label, synthesizes reply speech based on personalized interaction data and target semantic text; detects the user's original speech based on the speech activity detection threshold, and outputs the reply speech through the robot to complete the question-and-answer interaction.
[0051] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0052] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative descriptions of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
[0053] Some of the data in the above calculation formula are obtained by removing dimensions and taking their numerical values. The calculation formula is a calculation formula that is closest to the real situation, obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the calculation formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.
Claims
1. An AI-based intelligent question-answering method for robots, characterized in that, include: S1: Obtain the user's original speech and auxiliary recognition information, and extract acoustic information and stuttering type from the original speech; among which, acoustic information includes non-fluent speech segments, abnormal word ratio and Chinese character spacing; stuttering type labels include at least one or more combinations of syllable repetition, word repetition, sound prolongation, and silent blockage; auxiliary recognition information includes the number of lip movements and breathing frequency; S2: Perform semantic cleaning and reconstruction on the original speech, remove redundant information marked as non-fluent speech segments, and generate target semantic text that conforms to grammatical norms by combining the context. S3: When the user's original speech contains at least one type of stuttering, acquire the user's existing acoustic information and auxiliary recognition information, identify the user's stuttering impact factor based on the acoustic information and auxiliary recognition information, determine the speech activity detection threshold based on the stuttering impact factor, load personalized interaction data associated with the user based on the stuttering impact factor and stuttering type label, and synthesize reply speech based on personalized interaction data and target semantic text; wherein, personalized interaction data includes speech rhythm and the proportion of concerned sentences; S4: Detects the user's original voice based on the voice activity detection threshold, and outputs the reply voice through the robot to complete the question-and-answer interaction.
2. The AI-based intelligent question-answering method for robots according to claim 1, characterized in that, The process of acquiring the user's original speech and auxiliary recognition information, and extracting acoustic information and stuttering type from the original speech, includes: The system collects the user's original voice through a microphone array, and simultaneously acquires video stream data containing the user's facial features through an image acquisition device. The system extracts the user's lip movement trajectory features from the video stream to obtain the number of lip movements at each time point. The system also acquires the user's breathing frequency through an airflow sensor. The image acquisition device includes a high-definition RGB camera and a depth camera. The original speech is detected using a non-fluent speech detection model to identify non-fluent speech segments in the speech stream; the proportion of the duration of the same character in the non-fluent speech segment to the total duration of the corresponding sentence is calculated to obtain the abnormal character proportion; the silence duration between adjacent characters is calculated to obtain the Chinese character interval; the non-fluent speech segments, the abnormal character proportion, and the Chinese character interval are summarized into acoustic information; the acoustic information is time-aligned with auxiliary recognition information; wherein, the non-fluent speech detection model is implemented using a convolutional neural network and / or a time-delay neural network structure; A multimodal stuttering classification model is used to determine the obstruction pattern of the non-fluent speech segment and generate corresponding stuttering type labels. Specifically, when the same phoneme appears continuously within a preset short period of time, it is marked as syllable repetition; when a complete word appears continuously, it is marked as word repetition; when the duration of a vowel or consonant exceeds a preset duration threshold, it is marked as sound prolongation; when an acoustic signal interruption is detected and the lip movement trajectory features show that the user has made a vocalization action and the breathing frequency shows that the airflow is obstructed, it is marked as silent obstruction. The multimodal stuttering classification model is implemented through a bidirectional long short-term memory network, a temporal segmentation network, and / or a recurrent neural network structure.
3. The AI-based intelligent question-answering method for robots according to claim 1, characterized in that, The process of semantic cleaning and reconstruction of the original speech, removing redundant information marked as non-fluent speech segments, and generating target semantic text that conforms to grammatical rules by combining the context, includes: The original speech is subjected to speech recognition to generate a preliminary text sequence containing words and timestamp information; wherein, the timestamp information is used to mark the start and end times of each word or syllable in the original speech; Align the timestamps of the non-fluent speech segments with the word timestamps in the preliminary text sequence to determine the positions of redundant words or abnormal characters corresponding to the non-fluent speech segments in the text; Based on the stuttering type tag, redundant words or abnormal characters are selectively removed to obtain the cleaned intermediate text. The specific cleaning logic includes: if the stuttering type tag is a syllable repetition or word repetition, the first repetitive unit is retained and subsequent repetitive words are deleted; if the stuttering type tag is a sound prolongation, the prolonged character is truncated to a single character corresponding to the standard pronunciation duration; if the stuttering type tag is a silent blockage, the silent period corresponding to the segment is ignored, and the text content before and after the segment is directly retained. The intermediate text is input into a pre-trained semantic reconstruction model, and combined with the context of historical dialogue, the cleaned text is subjected to grammatical error correction and semantic completion to obtain target semantic text that conforms to grammatical norms; wherein, the semantic reconstruction model is implemented using the encoder-decoder network structure of Transformer.
4. The AI-based intelligent question-answering method for robots according to claim 1, characterized in that, The factors affecting user lag based on acoustic information and auxiliary recognition information include: The time ratio between the duration of existing non-fluent speech segments and the duration of the corresponding complete sentences is obtained, and the average value of the time ratio is marked as the basic stuttering factor. ; Get the average percentage of abnormal words in a single statement from a user's statements within the last n minutes. Spacing between Chinese characters Based on Chinese character spacing The interval regularity factor for each statement is determined by formula (1). ;in, Let be the numbers of each statement within the last n minutes, and The range of values is , The maximum value of statement numbers within the last n minutes; This is the number of the Chinese character in the current sentence, and The range of values is , This represents the maximum value of the Chinese character index in the current statement; Get the end time of each statement made by the user within the last n minutes. Based on end time The attenuation weight of each statement is determined by calculating formula (2) using auxiliary recognition information. ; The average percentage of abnormal words in each statement. Interval regularity factor and decay weight The emotional fluctuation blocking factor is determined by formula (3). ; through the basic stuttering factor and mood fluctuation blocking factors Determine the impact factor of stuttering Among them, the lag impact factor The formula for calculation is: , To adjust the parameters, they can be obtained by fitting multiple rounds of test data when the user uses it for the first time, or preset to [0.6, 0.8] based on experience, in order to balance the contribution of inherent defects and on-the-spot emotions to stuttering; The calculation formula (1) is: ; In the formula, The average value of the spacing between existing Chinese characters used by the user; The calculation formula (2) is: ; In the formula, For the current time, For standard duration, This is the attenuation coefficient used to adjust for the effect of time distance, and The value range is [0,1]; This represents the number of lip movements corresponding to the current statement. It is the average number of lip movements for all existing statements by the current user; This is the amplitude adjustment coefficient used to adjust the importance of the number of lip movements, and The value range is [0,1]; The breathing frequency corresponding to the current statement. It is the average breathing frequency of all statements currently in use by the user; This is the amplitude adjustment coefficient used to adjust the importance of respiratory rate, and The value range of is [0,1], and , and The ratio between any two is limited to the range [0.67, 1.5]. The calculation formula (3) is: ; In the formula, and Weighting coefficients, set according to different stuttering types, are used to balance the average proportion of abnormal words. and interval regularity factor The importance of.
5. The AI-based intelligent question-answering method for robots according to claim 4, characterized in that, The method for determining the speech activity detection threshold based on the stuttering impact factor includes: Extracting the impact factor of stuttering Based on the lag impact factor Baseline silence tolerance range and average stuttering impact factor of other users in history. Determine the threshold for voice activity detection Among them, the voice activity detection threshold The formula for calculation is: ; This represents the median of the baseline noise tolerance range. This represents the maximum value of the baseline noise tolerance range. This is the minimum baseline silence tolerance range; the baseline silence tolerance range is determined based on the dialogue between two participants in a group of experimenters.
6. The AI-based intelligent question-answering method for robots according to claim 1, characterized in that, The personalized interaction data associated with the user, loaded based on the stuttering impact factor and stuttering type tag, includes: Extract the lag impact factors closest to the current user's current time. And stuttering type tags, extract users with the same stuttering type tags as the current user from the historical data repository, and mark these users as similar users; obtain the stuttering influence factor of similar users in history, and place the stuttering influence factor of the current user at the same level. The historical impact factor within the fluctuation range is marked as the reference impact factor. Among them, the lag impact factor The fluctuation range is based on the current user's lag impact factor. Based on the preset fixed percentage Forming a relative fluctuation range Fixed percentage It is determined based on the number of users for the current type of tag; The reference impact factor number; Extracting reference impact factor The corresponding next reference impact factor It will be greater than the next reference impact factor. Reference impact factor Integrate into target factors, and obtain each target factor and its corresponding next reference impact factor. Personalized interaction data of the robot over time; among which, personalized interaction data includes voice rhythm and the proportion of caring statements.
7. The AI-based intelligent question-answering method for robots according to claim 1, characterized in that, The speech response synthesized based on personalized interaction data and target semantic text includes: Several sets of speech rhythms, proportions of relevant statements, target semantic text, and response speech were extracted from a speech reference database. The speech reference database includes several speech rhythms, proportions of relevant statements, target semantic text, and response speech set by experts based on the speech rhythms, proportions of relevant statements, and target semantic text. The speech rhythm, the proportion of relevant sentences, the target semantic text, and the response speech are integrated into several sets of training data and test data. The training data is used to train the artificial intelligence model, and the test data is used to test the trained artificial intelligence model. The artificial intelligence model is adjusted according to the test results. Finally, a speech synthesis model is obtained with the input of speech rhythm, the proportion of relevant sentences, and the target semantic text, and the output of the response speech. The artificial intelligence model is a non-linear regression model based on neural networks, and is implemented using a BP neural network structure and / or an RBF neural network structure. The speech rhythm and proportion of caring statements in the personalized interaction data corresponding to the current user, as well as the target semantic text, are input into the speech synthesis model to obtain the response speech.
8. The AI-based intelligent question-answering method for robots according to claim 1, characterized in that, The method of detecting the user's original speech based on a speech activity detection threshold includes: Extract the voice activity detection threshold corresponding to the current user. When the user's pause time exceeds the voice activity detection threshold, identify the received original voice and generate a response voice.
9. An AI-based intelligent question-answering system for robots, used to run the AI-based intelligent question-answering method according to any one of claims 1 to 8, characterized in that, include: Communication unit and processing unit; The communication unit is used to acquire the user's original speech and auxiliary recognition information, and to extract acoustic information and stuttering type from the original speech; wherein, the acoustic information includes non-fluent speech segments, abnormal word ratio, and Chinese character spacing; the stuttering type label includes at least one or more combinations of syllable repetition, word repetition, sound prolongation, and silent blockage; the auxiliary recognition information includes the number of lip movements and breathing frequency; The processing unit is used to perform semantic cleaning and reconstruction on the original speech, remove redundant information marked as non-fluent speech segments, and generate target semantic text that conforms to grammatical norms by combining the context. When the user's original speech contains at least one type of stuttering, it acquires the user's existing acoustic information and auxiliary recognition information, identifies the user's stuttering influence factor based on the acoustic information and auxiliary recognition information, and determines the speech activity detection threshold based on the stuttering influence factor. Based on the stuttering influence factor and stuttering type label, it loads personalized interaction data associated with the user, synthesizes reply speech based on personalized interaction data and target semantic text, detects the user's original speech based on the speech activity detection threshold, and outputs the reply speech through the robot to complete the question-and-answer interaction.
10. A storage medium storing instructions that, when executed on an AI-based robot intelligent question-answering system, cause the AI-based robot intelligent question-answering system to perform the AI-based robot intelligent question-answering method according to any one of claims 1 to 8.