Intelligent follow-up visit voice robot multi-round dialogue state tracking system
By using a multi-turn dialogue state tracking system for intelligent follow-up voice robots, and employing hierarchical parsing and disambiguation processing techniques, the problem of loose integration between dialogue semantic processing and historical state in existing technologies is solved. This enables real-time tracking and efficient management of dialogue state, improving the coherence and accuracy of interaction.
Patent Information
- Application Number
- CN202511538664.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies for tracking the state of multi-turn dialogues in follow-up voice robots suffer from poor interaction coherence due to a lack of close integration between dialogue semantic processing and historical states. Furthermore, the coordination between state tracking and data management is insufficient, making it difficult to meet the need for real-time adjustment of interaction strategies.
Employing a hierarchical semantic state decoder, a context-disambiguated state optimization model, an intent evolution state tracking algorithm, a follow-up dialogue streaming analysis platform, and a multi-turn dialogue state storage module, this system achieves accurate tracking and management of dialogue states through layered parsing, disambiguation processing, and intent tracking, combined with efficient access to real-time and historical data.
It enables real-time capture of dialogue status and efficient management of historical data, ensuring interaction continuity and information accuracy, reducing interaction interruptions and information collection deviations, and improving the efficiency of follow-up work.
Smart Images

Figure CN121687104A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice robot dialogue technology, and in particular to a multi-turn dialogue state tracking system for an intelligent follow-up voice robot. Background Technology
[0002] With the continuous expansion and intelligent upgrading of medical follow-up scenarios, follow-up voice robots have become an important tool for improving follow-up efficiency and ensuring follow-up continuity. Accurate tracking of the robot's state during multi-turn dialogues is a core prerequisite for effective robot interaction. Current follow-up dialogues are characterized by multiple dialogue rounds, broad information dimensions, and dynamic changes in patient statements. Traditional dialogue processing methods struggle to capture state fluctuations in real time, failing to provide the robot with coherent and accurate interaction data, leading to problems such as interaction interruptions and information collection errors during follow-up. To address these issues, a system capable of adapting to multi-turn dialogue scenarios, processing dialogue data in real time, and stably tracking dialogue states is urgently needed to meet the demands of medical follow-up for smooth interaction and accurate information, supporting the efficient conduct of follow-up work.
[0003] Existing technologies for tracking the state of multi-turn dialogues in follow-up voice robots have two significant shortcomings: First, the integration of dialogue semantic processing and historical states is not close enough. Most technologies only process the content of a single turn of dialogue and do not fully connect it with key information in historical dialogues. This makes it difficult to accurately determine the current dialogue state when there are connections or semantic extensions in the dialogue content, affecting the continuity of subsequent interactions. Second, the coordination between dialogue state tracking and data management is insufficient. Existing technologies lack effective processing and retrieval mechanisms for data generated during the dialogue process. They cannot efficiently link real-time tracked state data with historical stored data, which makes it easy for state tracking to deviate or lag during the advancement of multi-turn dialogues, making it difficult to meet the robot's need to adjust its interaction strategies in real time. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of existing technologies, this invention provides a multi-turn dialogue state tracking system for intelligent follow-up voice robots.
[0005] The technical solution adopted in this invention is a multi-turn dialogue state tracking system for an intelligent follow-up voice robot, comprising: a hierarchical semantic state decoder module receiving dialogue voice signals collected by the follow-up voice robot, performing hierarchical parsing of the semantic units contained in the signals, and outputting semantic state feature vectors to a context disambiguation state optimization model module; the context disambiguation state optimization model module receiving the semantic state feature vectors output by the hierarchical semantic state decoder module, performing disambiguation processing in conjunction with historical dialogue state data, generating disambiguated state parameters, and transmitting them to an intent evolution state tracking algorithm module; the intent evolution state tracking algorithm module constructing an intent evolution trajectory based on the disambiguated state parameters output by the context disambiguation state optimization model module, calculating the intent evolution rate and direction parameters, and sending the results to a follow-up dialogue streaming analysis platform; and the follow-up dialogue streaming analysis platform... The analysis platform receives the intent evolution trajectory, rate, and direction parameters output by the intent evolution state tracking algorithm module, performs dynamic analysis on the real-time dialogue stream, extracts dialogue state association features, and synchronizes the analysis results and feature data to the multi-turn dialogue state storage module. The multi-turn dialogue state storage module classifies and stores the analysis results and feature data transmitted by the follow-up dialogue streaming analysis platform, establishes a dialogue state index library, and provides data support for each module to call. The robot interaction control module retrieves dialogue state data from the multi-turn dialogue state storage module, combines it with the real-time analysis results of the follow-up dialogue streaming analysis platform, generates robot interaction control signals, and performs bidirectional data interaction with the hierarchical semantic state decoder module, the context disambiguation state optimization model module, the intent evolution state tracking algorithm module, the follow-up dialogue streaming analysis platform, and the multi-turn dialogue state storage module.
[0006] Furthermore, the hierarchical semantic state decoder module calculates the semantic state feature vector using the following formula: In the formula This is the semantic state feature vector of layer I. Input the number of semantic units for the current layer. The weight coefficient of the i-th semantic unit. Let be the weight matrix of the i-th semantic unit in layer I. For the original data of the i-th semantic unit, For the bias term of the i-th semantic unit in layer I, The historical state influence coefficient. The dimension of the semantic state feature vector of the previous layer. The weight parameter for the k-th historical state in the first layer. is the kth dimension value of the semantic state feature vector of layer I-1.
[0007] Furthermore, the context disambiguation state optimization model module performs disambiguation processing using the following formula: In the formula Let be the state parameters after disambiguation at time t. The activation function adjustment coefficient, The dimension of the semantic state feature vector is... The weight of the j-th semantic state feature is... Let j be the value of the j-th dimension of the semantic state feature vector. The impact coefficient of historical dialogue. As a dimension of historical dialogue status, Let q be the weight of the q-th historical dialogue state. This represents the q-th dimension value of the historical dialogue state.
[0008] Furthermore, the intent evolution state tracking algorithm module calculates the intent evolution trajectory parameters using the following formula: In the formula Let t be the parameters of the intended evolution trajectory. for The trajectory parameters are intended to evolve at any given moment. This represents the influence coefficient of the real-time disambiguation state. For the current round of dialogue, The weights of the disambiguation states in the s-th round are... The state parameters after disambiguation in the s-th round are... To evolve the fluctuation coefficient, The intended evolution direction angle at time t. For the dimension of disambiguation state change, The weight for the change of the u-th disambiguation state. Let be the change in the u-th dimension of the disambiguation state.
[0009] Furthermore, the follow-up dialogue streaming analysis platform extracts dialogue state association features using the following formula: In the formula Let be the dialogue state associated feature vector at time t. For the dimension of the intended evolution trajectory parameters, For the weights of the trajectory parameters of the k-th intention, Let k be the value of the intended evolution trajectory parameter. These are the convolution feature extraction coefficients. The amount of disambiguation state data, The convolution weights for the z-th disambiguation state are... This is the convolution operation function. Let z be the parameters of the z-th convolutional kernel. These are the pooling feature extraction coefficients. The number of semantic state feature vectors. The pooling weights for the a-th semantic state are... This is a pooling operation function. Let be the semantic state feature vector of the a-th term.
[0010] Furthermore, the robot interaction control module generates interaction control signals using the following formula: In the formula Let be the robot interaction control signal vector at time t. The influence coefficient of the dialogue state-related features. Associating features with dialogue state The weights of the features associated with the c-th dialogue state are: The c-th dimension value of the dialogue state-associated feature vector. The influence coefficient of the intended evolution trajectory parameters. The number of trajectory parameters to be evolved. For the weights of the trajectory parameters of the e-th intention, Let the e-th intentional evolution trajectory parameter value be... The influence coefficient of the state parameters after disambiguation. The dimension of the state parameters after disambiguation. The weights of the h-th disambiguated state parameters are... This represents the h-th dimension value of the state parameters after disambiguation.
[0011] Furthermore, the intent evolution state tracking algorithm module includes an intent feature extraction unit, an evolution rate calculation unit, an evolution direction determination unit, and a trajectory correction unit. The intent feature extraction unit receives the disambiguated state parameters output by the context disambiguation state optimization model module, filters the intent-related features contained in the parameters, and extracts the disambiguated state parameters of consecutive dialogue rounds through a sliding window mechanism to extract the feature change patterns. The evolution rate calculation unit calculates the change amplitude of intent features between adjacent dialogue rounds based on the feature change patterns output by the intent feature extraction unit, and obtains the intent evolution rate value per unit time by combining it with the time interval parameter, thus establishing a rate change curve. The evolution direction determination unit receives the rate change curve output by the evolution rate calculation unit, calculates the slope of the curve, and determines the direction vector of intent evolution by combining it with the dimensional change trend of the disambiguated state parameters, and marks the direction change nodes. The trajectory correction unit adjusts the initially constructed intent evolution trajectory based on the direction vector and change nodes output by the evolution direction determination unit, eliminates abnormal fluctuation points, makes the trajectory more consistent with the actual intent evolution process, and outputs the corrected trajectory parameters to the follow-up dialogue streaming analysis platform.
[0012] Furthermore, the follow-up dialogue streaming analysis platform includes a dialogue stream receiving unit, a real-time analysis unit, a feature extraction unit, and a data synchronization unit. The dialogue stream receiving unit receives the real-time dialogue data stream transmitted by the follow-up voice robot, performs frame parsing on the data stream, extracts the dialogue content and timestamp information from each frame of data, and establishes a dialogue stream data queue. The real-time analysis unit retrieves the dialogue stream data queue from the dialogue stream receiving unit, combines the intent evolution trajectory parameters output by the intent evolution state tracking algorithm module, performs semantic matching analysis on each frame of dialogue data, and identifies dialogue state change points. Based on the dialogue state change points identified by the real-time analysis unit, the feature extraction unit extracts dialogue feature data before and after the change points, obtains feature vectors related to the dialogue state through a feature association algorithm, and performs dimensionality reduction processing on the feature vectors. The data synchronization unit receives the reduced feature vectors output by the feature extraction unit, associates them with the analysis results of the real-time analysis unit, adds timestamp identifiers, and synchronously transmits them to the multi-turn dialogue state storage module to ensure the timeliness and accuracy of data transmission.
[0013] Furthermore, the multi-turn dialogue state storage module includes a data classification unit, an index building unit, a storage management unit, and a data retrieval unit. The data classification unit receives the analysis results and feature data transmitted from the follow-up dialogue streaming analysis platform, and classifies the data according to the dialogue topic, time period, and patient identification information, establishing different data category directories. The index building unit extracts keywords from the data under each category based on the category directories established by the data classification unit, generates data index items, and constructs a dialogue state index library. The index items include data storage address, category identifier, and time information. The storage management unit allocates storage for the data divided by the data classification unit, allocates different storage areas according to the importance of the data and access frequency, performs periodic verification of the stored data, and repairs data corruption issues. The data retrieval unit receives data retrieval requests from each module, parses the data requirement information in the request, searches for the corresponding index item in the dialogue state index library, retrieves the required data according to the storage address in the index item, and feeds it back to the requesting module, recording the data retrieval log.
[0014] A multi-turn dialogue state tracking system for an intelligent follow-up voice robot includes the following steps: S1, receiving dialogue voice signals collected by the follow-up voice robot and transmitting them to a hierarchical semantic state decoder module, where the decoder performs hierarchical parsing of semantic units in the voice signal and outputs semantic state feature vectors; S2, transmitting the semantic state feature vectors to a context disambiguation state optimization model module, where the model combines historical dialogue state data to disambiguate the feature vectors and generate disambiguated state parameters; S3, inputting the disambiguated state parameters into an intent evolution state tracking algorithm module, where the algorithm constructs a state tracking algorithm based on the parameters. S4. The intention evolution trajectory, evolution rate, and direction parameters are calculated; S5. The intention evolution trajectory, rate, and direction parameters are sent to the follow-up dialogue streaming analysis platform, which dynamically analyzes the real-time dialogue stream and extracts dialogue state correlation features; S6. The analysis results and feature data are transmitted to the multi-turn dialogue state storage module, which classifies and stores the data and establishes a dialogue state index library; S7. Dialogue state data is retrieved from the multi-turn dialogue state storage module, combined with the real-time analysis results of the follow-up dialogue streaming analysis platform, to generate robot interaction control signals, which are transmitted to the follow-up voice robot for multi-turn dialogue state tracking.
[0015] Beneficial Effects: This invention proposes a multi-turn dialogue state tracking system for an intelligent follow-up voice robot. Through the coherent collaboration of semantic processing, disambiguation optimization, intent tracking, and streaming analysis, it can capture dynamically changing semantics and intents in real time during dialogue. Combined with the classification management and efficient retrieval of historical data by the multi-turn dialogue state storage module, it solves the problem of weak integration between dialogue semantic processing and historical states in existing technologies, avoiding state judgment bias caused by single-turn processing and ensuring interaction continuity. On the other hand, the bidirectional data interaction mechanism between modules enables real-time tracked state data to quickly link with stored historical data. The dynamic analysis of the follow-up dialogue streaming analysis platform and the instantaneous response of the robot interaction control module form a closed loop, making up for the lack of coordination between state tracking and data management in existing technologies, eliminating deviations and lags in state tracking, ensuring that the robot can adjust its interaction strategy in real time, meeting the needs of medical follow-up for smooth interaction and accurate information. At the same time, through systematic state tracking and data management, it improves the efficiency of follow-up work and reduces interaction interruptions and information collection biases. Attached Figure Description
[0016] Figure 1 This is a diagram showing the system module composition of the present invention; Figure 2 This is a flowchart of the system operation steps of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] like Figure 1 As shown, an intelligent follow-up voice robot multi-turn dialogue state tracking system includes: a hierarchical semantic state decoder module, a context disambiguation state optimization model module, an intent evolution state tracking algorithm module, a follow-up dialogue streaming analysis platform, a multi-turn dialogue state storage module, and a robot interaction control module.
[0019] The hierarchical semantic state decoder module receives the dialogue speech signal collected by the follow-up speech robot, performs hierarchical parsing on the semantic units contained in the signal, and outputs the semantic state feature vector to the context disambiguation state optimization model module. Specifically, the hierarchical semantic state decoder module receives dialogue speech signals with a sampling rate of 16kHz and quantization precision collected by the follow-up speech robot. It employs a three-layer parsing architecture for hierarchical processing: the first layer segments the speech signal into frames with a frame length of 20ms and a frame shift of 10ms, extracting Mel-frequency cepstral coefficients (MFCC) features for each frame, resulting in 13 MFCC parameters; the second layer dynamically expands the MFCC features in the time dimension, calculating the first and second differences to form a 39-dimensional feature vector, and smoothing the features using a sliding window of size 5; the third layer uses an attention mechanism to assign weights to the smoothed feature vectors, with weight values ranging from 0.1 to 0.9. Feature vectors with a weight sum greater than 0.6 are mapped to semantic units, transforming the speech signal into a semantic state feature vector containing information such as patient symptom descriptions and follow-up needs, with an output dimension of 256. This module ensures the accuracy of semantic feature extraction through hierarchical parsing, providing high-quality input for subsequent disambiguation processing. During implementation, the processing latency at each level must not exceed 100ms to meet the requirements of real-time dialogue.
[0020] The context disambiguation state optimization model module receives the semantic state feature vector output by the hierarchical semantic state decoder module, combines it with historical dialogue state data for disambiguation processing, generates disambiguated state parameters, and transmits them to the intent evolution state tracking algorithm module. Specifically, the context disambiguation state optimization model module receives a 256-dimensional semantic state feature vector output from the hierarchical semantic state decoder, and simultaneously retrieves historical dialogue state data from the last 5 rounds in the multi-turn dialogue state storage module. This historical data includes semantic features, intent tags, and timestamp information for each round of dialogue. The model first calculates the similarity between the semantic state feature vector and the historical data, setting a similarity threshold of 0.7. Feature vectors with similarity below 0.7 are marked as potentially ambiguous data. Then, a context association algorithm is used, combined with the patient's expression habits in the historical dialogue (such as colloquial expressions for specific symptoms), to correct the potentially ambiguous data. During the correction process, the frequency of occurrence of the same semantic units in the historical data is considered, with semantic units appearing more than 3 times being prioritized for correction. Finally, a 128-dimensional disambiguated state parameter is generated, with the parameter error controlled within ±5%. This module eliminates semantic ambiguity by combining historical data with context association correction, providing accurate data for intent evolution tracking. During implementation, it is necessary to ensure that the historical data retrieval response time does not exceed 50ms to guarantee model processing efficiency.
[0021] The intent evolution state tracking algorithm module constructs the intent evolution trajectory based on the disambiguated state parameters output by the context disambiguation state optimization model module, calculates the intent evolution rate and direction parameters, and sends the results to the follow-up dialogue streaming analysis platform. Specifically, the intent evolution state tracking algorithm module receives 128-dimensional disambiguated state parameters output by the context disambiguation state optimization model. It first initializes the intent evolution trajectory, setting the initial point to the basic intent parameters at the start of the dialogue (e.g., the initial parameter value corresponding to "inquire about medication"). Then, based on a 1-second time step interval, it calculates the change in disambiguated state parameters within each time step, using the Euclidean distance formula. Time steps with changes greater than 0.2 are marked as intent change nodes. Next, based on the parameter change trend between change nodes, it calculates the intent evolution rate, measured in "parameter dimensions / second," with a rate range of 0.05-0.3 parameter dimensions / second. Simultaneously, it calculates the intent evolution direction using vector angles, with an angle range of 0-360°. Finally, it constructs a 3D intent evolution trajectory containing time, rate, and direction information, with a trajectory data sampling frequency of 1Hz. This module accurately captures the intent evolution process through dynamic calculation of parameter changes and trajectory construction, providing key trajectory data for streaming analysis. During implementation, it is crucial to ensure that the trajectory update frequency matches the dialogue stream transmission frequency to avoid data asynchrony.
[0022] The follow-up dialogue streaming analysis platform receives the intent evolution trajectory, rate and direction parameters output by the intent evolution state tracking algorithm module, performs dynamic analysis on the real-time dialogue stream, extracts dialogue state association features, and synchronizes the analysis results and feature data to the multi-turn dialogue state storage module. Specifically, the follow-up dialogue streaming analysis platform receives 3D intent evolution trajectory data output by the intent evolution state tracking algorithm, and simultaneously receives real-time dialogue streams transmitted by the follow-up voice robot. The dialogue stream transmission rate is 2Mbps, and the data format is JSON. The platform first parses the dialogue stream, extracting the text content, speaker identifier (patient / robot), and duration information for each sentence, with a parsing accuracy of over 99%. Then, it combines the intent evolution trajectory data to perform semantic matching on the text content, using a cosine similarity algorithm with a matching threshold of 0.65 to identify dialogue content segments that match the intent evolution trajectory. Next, it extracts dialogue state association features from the segments, including semantic similarity, intent matching degree, and dialogue duration percentage, extracting a total of 20 features, with feature values normalized to the 0-1 range. Finally, it fuses the features to generate a 64-dimensional dialogue state feature vector. This module associates trajectory data with dialogue content through dialogue stream parsing and feature extraction, providing feature data for the storage and control modules. During implementation, the dialogue stream parsing latency must be guaranteed to be no more than 200ms to meet real-time analysis requirements.
[0023] The multi-turn dialogue state storage module classifies and stores the analysis results and feature data transmitted by the follow-up dialogue streaming analysis platform, establishes a dialogue state index library, and provides data support for each module to call. Specifically, the multi-turn dialogue state storage module receives the 64-dimensional dialogue state feature vector and analysis results output by the follow-up dialogue streaming analysis platform. It adopts a distributed storage architecture with a storage capacity of 10TB and a read / write speed of no less than 500MB / s. The module first categorizes the received data based on patient ID (using an 18-bit numeric code), follow-up date (in YYYY-MM-DD format), and dialogue topic (such as "follow-up reminder" or "symptom feedback"), resulting in 10 topic categories. Next, it creates an index for each data category, including the data storage path, feature vector summary, and access timestamp, with an index update frequency of once per minute. Then, it manages the stored data using a hot / cold data separation strategy, storing the most recent 30 days of hot data on SSDs and data older than 30 days on HDDs. Data is verified every 24 hours using the MD5 hash algorithm to ensure data integrity. Finally, it responds to data retrieval requests from various modules, with a response time of no more than 30ms. Through categorized storage and efficient management, this module provides stable data support for all modules in the system. During implementation, a data backup mechanism must be configured, performing a full backup daily at 2 AM to prevent data loss.
[0024] The robot interaction control module retrieves dialogue state data from the multi-turn dialogue state storage module, combines it with the real-time analysis results from the follow-up dialogue streaming analysis platform, generates robot interaction control signals, and performs bidirectional data interaction with the hierarchical semantic state decoder module, the context disambiguation state optimization model module, the intent evolution state tracking algorithm module, the follow-up dialogue streaming analysis platform, and the multi-turn dialogue state storage module.
[0025] Specifically, the robot interaction control module retrieves a 64-dimensional dialogue state feature vector from the multi-turn dialogue state storage module, and simultaneously receives real-time analysis results from the follow-up dialogue streaming analysis platform. It uses a microcontroller (model STM32F407) as the control core, with an operating frequency of 168MHz. The module first fuses the retrieved feature vectors with the real-time analysis results using a weighted summation method. The feature vector weight is set to 0.6, and the real-time analysis result weight is set to 0.4. After fusion, control parameters are generated. Then, robot interaction control signals are generated based on the control parameters. The signal types include voice output control signals (sampling rate 44.1kHz, bit rate 128kbps) and dialogue logic control signals (including selection of the next dialogue node and determination of the question content). The control signals are then transmitted to the follow-up voice robot using the RS485 communication protocol at a transmission rate of 115200bps, with a transmission delay controlled within 50ms. Simultaneously, the module receives the execution status signals fed back by the robot and adjusts the control signals in real time. The adjustment range is calculated based on the execution status error. No adjustment is made when the error is less than 10%, and the error is corrected proportionally when it is greater than 10%. This module achieves precise interaction between the robot and the patient by generating and adjusting control signals. During implementation, the stability of the control signals must be ensured to avoid interaction interruption due to signal fluctuations.
[0026] Preferably, the hierarchical semantic state decoder module calculates the semantic state feature vector using the following formula: In the formula This is the semantic state feature vector of layer I. Input the number of semantic units for the current layer. The weight coefficient of the i-th semantic unit. Let be the weight matrix of the i-th semantic unit in layer I. For the original data of the i-th semantic unit, For the bias term of the i-th semantic unit in layer I, The historical state influence coefficient. The dimension of the semantic state feature vector of the previous layer. The weight parameter for the k-th historical state in the first layer. is the kth dimension value of the semantic state feature vector of layer I-1.
[0027] Specifically, in the semantic state feature vector calculation process of the hierarchical semantic state decoder module, the values of each parameter are determined first: the number of input semantic units in the current layer is set to 32, and the weight coefficient of each semantic unit is allocated according to semantic importance, ranging from 0.15 to 0.85, among which the weight of core semantic units (such as units related to patient symptom description) is not less than 0.6; the dimension of the weight matrix of the i-th semantic unit in the l-th layer is set to 64×32, and the bias term ranges from -0.5 to 0.5, and is optimized after 5000 rounds of sample training after random initialization; the historical state influence coefficient is set to 0.3, the dimension of the semantic state feature vector of the previous layer is 256, and the weight parameter of the k-th historical state in the l-th layer is processed by L2 regularization, with a value range from 0 to 0.4. During implementation, the original data of each semantic unit is first processed with its corresponding weight matrix and bias term. Then, the ReLU activation function is used to filter out output values less than 0. Subsequently, the processing results of all semantic units are summed by weight coefficients. Simultaneously, the sum of the products of the semantic state feature vector of the previous layer and the corresponding historical state weight parameters is calculated, multiplied by the historical state influence coefficient, and added to the aforementioned weighted sum to obtain the semantic state feature vector of the current layer. This process ensures that the semantic features are integrated with historical information, improving feature accuracy. During implementation, the calculation time for each round must be controlled to not exceed 80ms, and the feature vector output dimension is uniformly 256 dimensions to meet the data input requirements of subsequent modules.
[0028] Preferably, the context disambiguation state optimization model module performs disambiguation processing using the following formula: In the formula Let be the state parameters after disambiguation at time t. The activation function adjustment coefficient, The dimension of the semantic state feature vector is... The weight of the j-th semantic state feature is... Let j be the value of the j-th dimension of the semantic state feature vector. The impact coefficient of historical dialogue. As a dimension of historical dialogue status, Let q be the weight of the q-th historical dialogue state. This represents the q-th dimension value of the historical dialogue state.
[0029] Specifically, in the context disambiguation state optimization model module, the disambiguation processing calculation is implemented with an activation function adjustment coefficient of 0.7. This coefficient is determined through testing with 1000 sets of ambiguous samples to ensure the discriminative power of ambiguous data. The semantic state feature vector has a dimension of 256, and the weight of each semantic state feature is normalized by the Softmax function, with a value range between 0.002 and 0.015. Among them, the weight of features with high relevance to historical dialogues is no less than 0.008. The historical dialogue influence coefficient is set to 0.45, and the historical dialogue state has a dimension of 128. The weight of the q-th historical dialogue state decays according to the dialogue round. The weight of the most recent dialogue round is 0.3, and the weight decreases by 0.05 for each round backward, with a minimum of no less than 0.1. The process first calculates the sum of the products of the semantic state feature vector and its corresponding weight, then calculates the sum of the products of the historical dialogue states and their corresponding weights. These two results are then multiplied by the activation function adjustment coefficient and the historical dialogue influence coefficient, respectively, and summed to obtain the comprehensive input value. This comprehensive input value is then substituted into the Sigmoid activation function to output the disambiguated state parameters. The parameter values range from 0 to 1, where values above 0.6 are considered definite states, values below 0.4 are considered states requiring further verification, and values between 0.4 and 0.6 require secondary disambiguation based on more historical data. This process achieves precise disambiguation through parameter adjustment, ensuring a disambiguation accuracy of at least 92% and a processing time of less than 60ms per round.
[0030] Preferably, the intent evolution state tracking algorithm module calculates the intent evolution trajectory parameters using the following formula: In the formula Let t be the parameters of the intended evolution trajectory. for The trajectory parameters are intended to evolve at any given moment. This represents the influence coefficient of the real-time disambiguation state. For the current round of dialogue, The weights of the disambiguation states in the s-th round are... The state parameters after disambiguation in the s-th round are... To evolve the fluctuation coefficient, The intended evolution direction angle at time t. For the dimension of disambiguation state change, The weight for the change of the u-th disambiguation state. Let be the change in the u-th dimension of the disambiguation state.
[0031] Specifically, the intent evolution trajectory parameter calculation of the intent evolution state tracking algorithm module determines the values of each parameter during implementation: the real-time disambiguation state influence coefficient is set to 0.5, which is determined by analyzing 200 sets of follow-up dialogue intent evolution data to balance the influence of real-time and historical data; the maximum number of current dialogue rounds is 20 rounds, and the weight of the disambiguation state in each round increases with the round, with the weight of the first round being 0.03, and the weight increasing by 0.02 for each additional round, not exceeding 0.35; the intent evolution fluctuation coefficient is set to 0.2 to control the degree of influence of intent fluctuation on the trajectory; the intent evolution direction angle at time t is calculated based on the trend of the disambiguation state changes in the previous 3 rounds, with a value range between 0° and 360°, and a change exceeding 45° is marked as an intent reversal; the disambiguation state change dimension is 64-dimensional, and the weight of each disambiguation state change is determined by variance analysis, with the weight of dimensions with a change variance greater than 0.1 not less than 0.04. During implementation, the intent evolution trajectory parameters from the previous moment are first obtained. Then, the sum of the products of the disambiguation state parameters and their corresponding weights from all previous rounds is calculated and multiplied by the real-time disambiguation state influence coefficient. Simultaneously, the sum of the products of the changes in each dimension of the disambiguation state and their corresponding weights is calculated, multiplied by the sine of the intent evolution direction angle, and then multiplied by the intent evolution fluctuation coefficient. The two results are added to the trajectory parameters from the previous moment to obtain the intent evolution trajectory parameters for the current moment. This process enables dynamic updating of the intent trajectory. During implementation, the trajectory parameter update frequency is synchronized with the dialogue rounds, with each round's update time not exceeding 70ms, ensuring that the trajectory closely matches actual intent changes.
[0032] Preferably, the follow-up dialogue streaming analysis platform extracts dialogue state association features using the following formula: In the formula Let be the dialogue state associated feature vector at time t. For the dimension of the intended evolution trajectory parameters, For the weights of the trajectory parameters of the k-th intention, Let k be the value of the intended evolution trajectory parameter. These are the convolution feature extraction coefficients. The amount of disambiguation state data, The convolution weights for the z-th disambiguation state are... This is the convolution operation function. Let z be the parameters of the z-th convolutional kernel. These are the pooling feature extraction coefficients. The number of semantic state feature vectors. The pooling weights for the a-th semantic state are... This is a pooling operation function. Let be the semantic state feature vector of the a-th term.
[0033] Specifically, in the follow-up dialogue streaming analysis platform, the dialogue state association feature extraction calculation is implemented with the following parameters set: the intent evolution trajectory parameter dimension is 64-dimensional, and the weight of each intent evolution trajectory parameter is allocated according to the influence of the trajectory on the dialogue state. The weight of core trajectory parameters (such as intent transition related parameters) is not less than 0.08, and the weight of other parameters ranges from 0.01 to 0.06; the convolution feature extraction coefficient is set to 0.35, the amount of disambiguation state data is calculated as a data segment every 500ms, the convolution weight dimension of each disambiguation state is 3×3, the convolution kernel parameters are fine-tuned after pre-training on the ImageNet dataset, and a total of 16 convolution kernels are set; the pooling feature extraction coefficient is set to 0.25, the number of semantic state feature vectors is consistent with the dialogue round, and the pooling weight of each semantic state is determined according to the semantic integrity of the feature vector. The weight of feature vectors with integrity scores higher than 0.8 is not less than 0.05. The process first calculates the sum of the products of the intent evolution trajectory parameters and their corresponding weights. Then, it performs convolution operations on each disambiguation state data segment, extracts local features, multiplies them by the corresponding convolution weights, sums all results, and multiplies by the convolution feature extraction coefficient. Simultaneously, it performs max pooling on each semantic state feature vector, retains key features, multiplies them by the corresponding pooling weights, sums them, and multiplies by the pooling feature extraction coefficient. Finally, it adds the three results to obtain the dialogue state-related feature vector. This process integrates multi-dimensional features, with the feature vector dimension uniformly set to 128 dimensions during implementation. The extraction time is less than 90ms, providing comprehensive feature support for subsequent data storage and control signal generation.
[0034] Preferably, the robot interaction control module generates the interaction control signal using the following formula: In the formula Let be the robot interaction control signal vector at time t. The influence coefficient of the dialogue state-related features. Associating features with dialogue state The weights of the features associated with the c-th dialogue state are: The c-th dimension value of the dialogue state-associated feature vector. The influence coefficient of the intended evolution trajectory parameters. The number of trajectory parameters to be evolved. For the weights of the trajectory parameters of the e-th intention, Let the e-th intentional evolution trajectory parameter value be... The influence coefficient of the state parameters after disambiguation. The dimension of the state parameters after disambiguation. The weights of the h-th disambiguated state parameters are... This represents the h-th dimension value of the state parameters after disambiguation.
[0035] Specifically, the interaction control signal generation and calculation of the robot interaction control module determines the values of each parameter during implementation: the influence coefficient of dialogue state association features is set to 0.5, the dimension of dialogue state association features is 128, and the weight of each dialogue state association feature is optimized by the gradient descent algorithm, with a value range between 0.005 and 0.02. Among them, the weight of features directly related to the interaction strategy (such as the urgency of patient needs) is not less than 0.015; the influence coefficient of intent evolution trajectory parameters is set to 0.3, the number of intent evolution trajectory parameters corresponds to the dialogue rounds, and the weight of each intent evolution trajectory parameter decays with the parameter update time. The weight of the latest updated parameter is 0.04, and the weight decreases by 0.005 every round; the influence coefficient of disambiguated state parameters is set to 0.2, the dimension of disambiguated state parameters is 128, and the weight of each disambiguated state parameter is determined according to the disambiguation confidence of the parameter. The weight of parameters with a confidence higher than 0.9 is not less than 0.008. During implementation, the sum of the products of dialogue state association features and their corresponding weights (multiplied by the corresponding influence coefficient), the sum of the products of intent evolution trajectory parameters and their corresponding weights (multiplied by the corresponding influence coefficient), and the sum of the products of disambiguated state parameters and their corresponding weights (multiplied by the corresponding influence coefficient) are calculated separately. These three results are then summed to obtain the robot interaction control signal vector. This process ensures that the control signal integrates data from multiple modules. During implementation, the control signal vector has a 64-dimensional dimension, a generation time of no more than 50ms, and a bit error rate controlled within 10% during signal transmission. -6 The following ensures the accurate execution of the robot's interactive actions.
[0036] Preferably, the intent evolution state tracking algorithm module includes an intent feature extraction unit, an evolution rate calculation unit, an evolution direction determination unit, and a trajectory correction unit. The intent feature extraction unit receives the disambiguated state parameters output by the context disambiguation state optimization model module, filters the intent-related features contained in the parameters, and extracts the disambiguated state parameters of consecutive dialogue rounds through a sliding window mechanism to extract the feature change patterns. The evolution rate calculation unit calculates the change amplitude of intent features between adjacent dialogue rounds based on the feature change patterns output by the intent feature extraction unit, and obtains the intent evolution rate value per unit time by combining it with the time interval parameter, thus establishing a rate change curve. The evolution direction determination unit receives the rate change curve output by the evolution rate calculation unit, calculates the slope of the curve, and determines the direction vector of intent evolution by combining it with the dimensional change trend of the disambiguated state parameters, and marks the direction change nodes. The trajectory correction unit adjusts the initially constructed intent evolution trajectory based on the direction vector and change nodes output by the evolution direction determination unit, eliminates abnormal fluctuation points, makes the trajectory more consistent with the actual intent evolution process, and outputs the corrected trajectory parameters to the follow-up dialogue streaming analysis platform.
[0037] Specifically, the intent evolution state tracking algorithm module comprises four units, each with clearly defined parameters and processes during implementation: The intent feature extraction unit receives 128-dimensional disambiguated state parameters output from the context disambiguation state optimization model. It uses a sliding window of size 5 to extract parameter data from six consecutive rounds of dialogue, with a window step size of 1 round. The mean and standard deviation of parameters within each window are calculated, and parameters with a mean deviation exceeding 0.15 are marked as key intent features, forming a 64-dimensional feature vector. The evolution rate calculation unit, based on the extracted feature vectors, calculates the Euclidean distance between feature vectors from two adjacent rounds of dialogue. The distance is divided by the time interval between the two rounds of dialogue (in seconds, ranging from 1 to 5 seconds) to obtain the evolution rate value. The rate threshold is set to 0.3; values exceeding this threshold are marked as rapid evolution. The system simultaneously plots the rate change curve with a time granularity of 1 second. The evolution direction determination unit calculates the slope of the rate change curve. A slope with an absolute value greater than 0.2 is considered a directional change. Combining the change trends of 10 core dimensions in the disambiguated state parameters, a three-dimensional direction vector is generated (x, y, and z axes correspond to the dimensions of symptom description, treatment needs, and follow-up arrangements, respectively). When the angle between the direction vectors changes by more than 30°, it is marked as an evolution direction node. The trajectory correction unit uses the 3σ criterion to identify abnormal fluctuation points in the rate change curve (points that deviate from the mean by more than 3 times the standard deviation). The abnormal points are replaced with the mean of 5 adjacent normal points to smooth the initial trajectory. The corrected trajectory data is updated every 200ms and output to the follow-up dialogue streaming analysis platform to ensure that the intended trajectory is consistent with the actual evolution.
[0038] Preferably, the follow-up dialogue streaming analysis platform includes a dialogue stream receiving unit, a real-time analysis unit, a feature extraction unit, and a data synchronization unit. The dialogue stream receiving unit receives the real-time dialogue data stream transmitted by the follow-up voice robot, performs frame parsing on the data stream, extracts the dialogue content and timestamp information from each frame, and establishes a dialogue stream data queue. The real-time analysis unit retrieves the dialogue stream data queue from the dialogue stream receiving unit, combines the intent evolution trajectory parameters output by the intent evolution state tracking algorithm module, performs semantic matching analysis on each frame of dialogue data, and identifies dialogue state change points. The feature extraction unit extracts dialogue feature data before and after the change points based on the dialogue state change points identified by the real-time analysis unit, obtains feature vectors related to the dialogue state through a feature association algorithm, and performs dimensionality reduction processing on the feature vectors. The data synchronization unit receives the reduced feature vectors output by the feature extraction unit, associates them with the analysis results of the real-time analysis unit, adds timestamp identifiers, and synchronously transmits them to the multi-turn dialogue state storage module to ensure the timeliness and accuracy of data transmission.
[0039] Specifically, the four units of the follow-up dialogue streaming analysis platform have the following parameters and operating procedures: The dialogue stream receiving unit receives the real-time dialogue data stream transmitted by the follow-up voice robot. The data stream transmission rate is 2Mbps, encapsulated using the RTP protocol. The unit parses the data in 100ms frames, extracting the text content (encoded in UTF-8), speaker identifier (1-byte binary code, 00 for patient, 01 for robot), and duration information (accurate to 10ms) from each frame. The parsing error rate is controlled below 0.1%. The parsed data is stored in a 100MB cache queue. The real-time analysis unit retrieves data from the cache queue and, combined with the 64-dimensional intent evolution trajectory parameters output by the intent evolution state tracking algorithm, uses a cosine similarity algorithm to calculate the matching degree between the text content and the trajectory parameters. The matching degree threshold is set to 0.6. 5. Text segments below the threshold are marked as candidate points for state changes. These candidate points are then verified using a dynamic time warping algorithm. Once confirmed, they are marked as dialogue state change points. The feature extraction unit extracts features from the text data 3 seconds before and after the change point, including 20-dimensional features such as word frequency (counting the occurrence of 500 commonly used medical terms), semantic similarity (average similarity with the previous 3 rounds of dialogue), and sentence structure (proportion of declarative / interrogative sentences). Principal component analysis is used to reduce the feature dimensions to 16 dimensions, and the feature variance contribution rate after reduction is no less than 90%. The data synchronization unit adds millisecond-level timestamps to the reduced feature vectors and analysis results, and transmits them to the multi-turn dialogue state storage module via TCP protocol. The transmission delay is controlled within 50ms, and a transmission log (including transmission time, data volume, and checksum) is recorded to ensure the accuracy and traceability of data synchronization.
[0040] Preferably, the multi-turn dialogue state storage module includes a data classification unit, an index building unit, a storage management unit, and a data retrieval unit. The data classification unit receives the analysis results and feature data transmitted from the follow-up dialogue streaming analysis platform, and classifies the data according to the dialogue topic, time period, and patient identification information to establish different data category directories. The index building unit extracts keywords from the data under each category based on the category directories established by the data classification unit, generates data index items, and constructs a dialogue state index library. The index items include data storage address, category identifier, and time information. The storage management unit allocates storage for the data divided by the data classification unit, allocates different storage areas according to the importance of the data and access frequency, performs periodic verification of the stored data, and repairs data corruption issues. The data retrieval unit receives data retrieval requests from each module, parses the data requirement information in the request, searches for the corresponding index item in the dialogue state index library, retrieves the required data according to the storage address in the index item, and feeds it back to the requesting module, recording the data retrieval log.
[0041] Specifically, the four units of the multi-turn dialogue state storage module have clearly defined parameters and operations during implementation: The data classification unit receives data (including 16-dimensional feature vectors and analysis results) transmitted from the follow-up dialogue streaming analysis platform. It performs three-level classification based on dialogue topic (divided into 8 categories such as disease consultation, medication guidance, and follow-up appointment), time period (divided by day, storing nearly 90 days of data), and patient identifier (18-digit numerical code). The classification accuracy is required to reach over 99.5%. After classification, the data is stored in a directory structure of "topic-date-patient ID," with each directory containing subfolders with a maximum capacity of 1GB. The index building unit extracts keywords from the data in each directory (5-8 keywords per data segment, including disease name, treatment plan, etc.) and generates an index containing the stored data. The system stores index entries containing the storage path (absolute path length not exceeding 256 characters), category identifier (3-digit numeric code), and timestamp (accurate to the second). These index entries are sorted in ascending order by patient ID, and a B+ tree index is constructed. The query response time for this index is controlled within 10ms. The storage management unit employs a hybrid storage architecture. Hot data (accessed more than 5 times / day) for approximately 30 days is stored on SSDs (read / write speed not less than 500MB / s), while cold data (30-90 days) is stored on HDDs (read / write speed not less than 100MB / s). CRC32 checks are performed on the stored data daily from 2-4 AM. Data failing the check is recovered from backup files (using a RAID5 backup strategy, with a backup capacity 1.5 times the original data). The data loss rate is controlled within 10%. -6 The data retrieval unit receives retrieval requests from various modules (the request includes patient ID, time range, and data type). After parsing the request, it queries the corresponding index item in the index library and retrieves the data according to the storage path in the index item. Data compression (compression rate not less than 30%) is used to reduce the amount of data transmitted during retrieval. At the same time, a retrieval log (including the requesting module, retrieval time, and data volume) is recorded. The log is retained for 30 days to ensure the efficiency and monitorability of data retrieval.
[0042] The hierarchical semantic state decoder is the core module in this invention responsible for converting the speech signals collected by the follow-up speech robot into structured semantic features. Essentially, it achieves accurate extraction of speech semantics through a hierarchical parsing architecture. Its implementation requires receiving a dialogue speech signal with a 16kHz sampling rate and 16-bit quantization precision, employing a three-layer parsing process: the first layer segments the signal with a 20ms frame length and a 10ms frame shift, extracting 13-dimensional Mel-frequency cepstral coefficients (MFCC) features; the second layer expands to 39-dimensional dynamic features and smooths them with a 5-frame window; the third layer maps semantic units through an attention mechanism with weights of 0.1-0.9, outputting a 256-dimensional semantic state feature vector, with each layer's processing latency not exceeding 100ms. This module transforms unstructured speech signals into semantic features that can be subsequently processed, filtering out invalid speech information and retaining key content such as patient symptom descriptions and follow-up needs. It provides high-quality data input for subsequent context disambiguation and intent tracking, avoiding processing deviations in subsequent modules due to inaccurate semantic extraction, ensuring the reliability of basic data for multi-turn dialogue state tracking, and meeting the dual requirements of real-time and accuracy in semantic parsing in medical follow-up scenarios.
[0043] The context-disambiguation state optimization model is a key module for eliminating semantic feature ambiguity and improving the accuracy of state data. Essentially, it's a processing unit that combines historical dialogue data to correct and optimize semantic features. Its implementation requires receiving a 256-dimensional semantic feature vector output from a hierarchical semantic state decoder, while simultaneously retrieving historical data from the last five rounds of dialogue state storage. First, it calculates the similarity between the feature vector and the historical data (threshold 0.7), marking potentially ambiguous data below the threshold. Then, based on semantic units appearing more than three times in the historical data, it corrects ambiguous data by considering the patient's expression habits, ultimately generating 128-dimensional disambiguated state parameters (error ±5%), with a historical data retrieval response time of no more than 50ms. This module solves the ambiguity problem caused by the patient's colloquial expressions and vague descriptions in semantic features, ensuring that the output state parameters accurately reflect the patient's true intentions. It avoids interference from ambiguous data in subsequent intention evolution tracking, reduces robot interaction deviations caused by semantic misunderstandings, and provides a guarantee for accurately capturing patient needs in medical follow-ups. It serves as a crucial bridge connecting semantic parsing and intention tracking, improving the entire system's adaptability to complex dialogue scenarios.
[0044] The intent evolution state tracking algorithm is the core algorithm module for dynamically capturing the trajectory of changes in patient dialogue intent. Essentially, it's a computational unit that constructs the intent evolution path based on disambiguated state parameters. Its implementation requires receiving 128-dimensional disambiguated state parameters. First, it initializes the starting point of the intent trajectory. Then, it calculates the parameter change (Euclidean distance) with a 1-second time step, marking intent change nodes with changes exceeding 0.2. Next, it calculates the evolution rate of 0.05-0.3 parameter dimensions / second and the evolution direction of 0-360°, constructing a three-dimensional intent trajectory. The trajectory sampling frequency is 1Hz, and the update frequency is synchronized with the dialogue flow. This algorithm tracks the entire process of patient intent from its initial state to dynamic changes in real time, such as the intent shift from "consulting about medication" to "scheduling a follow-up visit," accurately marking intent turning points. It provides dynamic intent data for the follow-up dialogue streaming analysis platform, avoiding the shortcomings of traditional static intent recognition in adapting to changes in intent during multi-turn dialogues. This allows the robot to promptly perceive changes in patient needs, providing decision support for subsequent interaction strategy adjustments and ensuring the continuity and relevance of dialogue interactions during medical follow-ups.
[0045] The follow-up dialogue streaming analysis platform is a core processing platform that correlates intent trajectories with real-time dialogue streams and extracts dialogue state features. Essentially, it's a comprehensive unit that integrates multi-source data to generate structured state features. Its implementation requires receiving three-dimensional intent evolution trajectory data and a JSON-formatted dialogue stream with a transmission rate of 2Mbps. First, it parses the dialogue stream in 100ms frames (with an accuracy exceeding 99%), extracting text content, speaker identifiers, and duration. Then, using a threshold of 0.65, it matches the text with the intent trajectory using a cosine similarity algorithm to identify state change points. Subsequently, it extracts 20-dimensional features, reduces them to 16 dimensions, adds millisecond-level timestamps, and synchronizes them to the multi-turn dialogue state storage module, with a transmission latency of no more than 50ms. This platform deeply correlates dynamic intent trajectories with real-time dialogue content, extracting dialogue state features including semantic similarity and intent matching, thus achieving the transition from dynamic tracking to structured storage of data. Its significance lies in opening up the data flow channel between intent tracking, data storage, and interactive control, providing a classification and storage basis for the multi-turn dialogue state storage module, and providing real-time feature support for the robot interaction control module. It is a key hub to ensure efficient data flow and accurate state tracking of the entire system, and directly affects the smoothness of medical follow-up interaction and the accuracy of information collection.
[0046] like Figure 2As shown, a multi-turn dialogue state tracking system for an intelligent follow-up voice robot includes the following steps: S1, receiving dialogue voice signals collected by the follow-up voice robot and transmitting them to a hierarchical semantic state decoder module, where the decoder performs hierarchical parsing of semantic units in the voice signal and outputs semantic state feature vectors; S2, transmitting the semantic state feature vectors to a context disambiguation state optimization model module, where the model combines historical dialogue state data to disambiguate the feature vectors and generate disambiguated state parameters; S3, inputting the disambiguated state parameters into an intent evolution state tracking algorithm module, where the algorithm is based on parameter configuration... S4. Construct the intention evolution trajectory and calculate the evolution rate and direction parameters; S5. Send the intention evolution trajectory, rate, and direction parameters to the follow-up dialogue streaming analysis platform, which dynamically analyzes the real-time dialogue stream and extracts dialogue state correlation features; S6. Transmit the analysis results and feature data to the multi-turn dialogue state storage module, which classifies and stores the data and establishes a dialogue state index library; S7. Retrieve dialogue state data from the multi-turn dialogue state storage module, combine it with the real-time analysis results from the follow-up dialogue streaming analysis platform, generate robot interaction control signals, and transmit them to the follow-up voice robot for multi-turn dialogue state tracking.
[0047] Meanwhile, the modules interact bidirectionally to ensure the continuity and accuracy of dialogue state tracking. After each round of dialogue, the multi-round dialogue state storage module updates the stored data, the intent evolution state tracking algorithm module adjusts the evolution trajectory parameters based on the new data, the context disambiguation state optimization model module updates the historical dialogue state data, the hierarchical semantic state decoder module prepares to receive the next round of dialogue voice signals, the follow-up dialogue streaming analysis platform clears the current dialogue stream data queue, and the robot interaction control module resets the control signal parameters to prepare for the next round of dialogue state tracking.
[0048] A multi-turn dialogue state tracking system for an intelligent follow-up voice robot, through the coherent collaboration of multiple modules, can capture the dynamically changing semantics and intentions in multi-turn dialogues in real time. Simultaneously, relying on a multi-turn dialogue state storage module for the classified management and efficient retrieval of historical data, it can fully correlate key information from different rounds of dialogue, avoiding the limitations of judgments caused by relying solely on single-turn content. This advantage precisely overcomes the shortcomings of existing technologies where "dialogue semantic processing and historical state are not closely integrated," effectively eliminating state judgment biases caused by a lack of historical data support. It ensures that the robot can accurately grasp the current dialogue state even when dialogue content connects or semantics extend, guaranteeing the continuity of multi-turn interactions and providing a stable basis for subsequent interaction strategy adjustments.
[0049] The system establishes a two-way data interaction mechanism between its modules, enabling real-time tracking of dialogue status data to quickly link with stored historical data. Simultaneously, the dynamic analysis function of the follow-up dialogue streaming analysis platform and the real-time response of the robot's interaction control module form a closed loop, allowing for the synchronous processing of dialogue stream data and status parameters. This advantage specifically addresses the deficiency of insufficient coordination between status tracking and data management in existing technologies, eliminating deviations and lags in the status tracking process. It allows the robot to adjust its interaction strategy promptly based on real-time updated status data, meeting the demands of medical follow-up for smooth interaction and accurate information. Furthermore, through systematic status tracking and data management, it reduces interaction interruptions and information collection errors, further improving the overall efficiency of follow-up work.
[0050] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-turn dialogue state tracking system for an intelligent follow-up voice robot, characterized in that, include: The hierarchical semantic state decoder module receives the dialogue speech signal collected by the follow-up voice robot, performs hierarchical parsing of the semantic units contained in the signal, and outputs the semantic state feature vector to the context disambiguation state optimization model module. The context disambiguation state optimization model module receives the semantic state feature vector output by the hierarchical semantic state decoder module, performs disambiguation processing in combination with historical dialogue state data, generates disambiguated state parameters, and transmits them to the intent evolution state tracking algorithm module. The intent evolution state tracking algorithm module constructs the intent evolution trajectory based on the disambiguated state parameters output by the context disambiguation state optimization model module, calculates the intent evolution rate and direction parameters, and sends the results to the follow-up dialogue streaming analysis platform. The follow-up dialogue streaming analysis platform receives the intent evolution trajectory, rate, and direction parameters output by the intent evolution state tracking algorithm module, performs dynamic analysis on the real-time dialogue stream, extracts dialogue state association features, and synchronizes the analysis results and feature data to the multi-turn dialogue state storage module. The multi-turn dialogue state storage module classifies and stores the analysis results and feature data transmitted from the follow-up dialogue streaming analysis platform, and establishes a dialogue state index library to provide data support for each module to call. The robot interaction control module retrieves dialogue state data from the multi-turn dialogue state storage module, combines it with the real-time analysis results of the follow-up dialogue streaming analysis platform, generates robot interaction control signals, and conducts bidirectional data interaction with the hierarchical semantic state decoder module, the context disambiguation state optimization model module, the intent evolution state tracking algorithm module, the follow-up dialogue streaming analysis platform, and the multi-turn dialogue state storage module.
2. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The hierarchical semantic state decoder module calculates the semantic state feature vector using the following formula: In the formula This is the semantic state feature vector of layer I. Input the number of semantic units for the current layer. The weight coefficient of the i-th semantic unit. Let be the weight matrix of the i-th semantic unit in layer I. For the original data of the i-th semantic unit, For the bias term of the i-th semantic unit in layer I, The historical state influence coefficient. The dimension of the semantic state feature vector of the previous layer. The weight parameter for the k-th historical state in the first layer. is the kth dimension value of the semantic state feature vector of layer I-1.
3. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The context disambiguation state optimization model module performs disambiguation processing using the following formula: In the formula Let be the state parameters after disambiguation at time t. The activation function adjustment coefficient, The dimension of the semantic state feature vector is... The weight of the j-th semantic state feature is... Let j be the value of the j-th dimension of the semantic state feature vector. The impact coefficient of historical dialogue. As a dimension of historical dialogue status, Let q be the weight of the q-th historical dialogue state. This represents the q-th dimension value of the historical dialogue state.
4. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The intent evolution state tracking algorithm module calculates the intent evolution trajectory parameters using the following formula: In the formula Let t be the parameters of the intended evolution trajectory. for The trajectory parameters are intended to evolve at any given moment. This represents the influence coefficient of the real-time disambiguation state. For the current round of dialogue, The weights of the disambiguation states in the s-th round are... The state parameters after disambiguation in the s-th round are... To evolve the fluctuation coefficient, The intended evolution direction angle at time t. For the dimension of disambiguation state change, The weight for the change of the u-th disambiguation state. Let be the change in the u-th dimension of the disambiguation state.
5. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The follow-up dialogue streaming analysis platform extracts dialogue state association features using the following formula: In the formula Let be the dialogue state associated feature vector at time t. For the dimension of the intended evolution trajectory parameters, For the weights of the trajectory parameters of the k-th intention, Let k be the value of the intended evolution trajectory parameter. These are the convolution feature extraction coefficients. The amount of disambiguation state data, The convolution weights for the z-th disambiguation state are... This is the convolution operation function. Let z be the parameters of the z-th convolutional kernel. These are the pooling feature extraction coefficients. The number of semantic state feature vectors. The pooling weights for the a-th semantic state are... This is a pooling operation function. Let be the semantic state feature vector of the a-th term.
6. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The robot interaction control module generates interaction control signals using the following formula: In the formula Let be the robot interaction control signal vector at time t. The influence coefficient of the dialogue state-related features. Associating features with dialogue state The weights of the features associated with the c-th dialogue state are: The c-th dimension value of the dialogue state-associated feature vector. The influence coefficient of the intended evolution trajectory parameters. The number of trajectory parameters to be evolved. For the weights of the trajectory parameters of the e-th intention, For the e-th intentional evolution trajectory parameter value, The influence coefficient of the state parameters after disambiguation. The dimension of the state parameters after disambiguation. The weights of the h-th disambiguated state parameters are... This represents the h-th dimension value of the state parameters after disambiguation.
7. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The intent evolution state tracking algorithm module includes an intent feature extraction unit, an evolution rate calculation unit, an evolution direction determination unit, and a trajectory correction unit. The intent feature extraction unit receives the disambiguated state parameters output by the context disambiguation state optimization model module, filters the intent-related features contained in the parameters, and extracts the disambiguated state parameters of consecutive dialogue rounds through a sliding window mechanism to extract the feature change patterns. The evolution rate calculation unit calculates the change amplitude of intent features between adjacent dialogue rounds based on the feature change pattern output by the intent feature extraction unit, and obtains the intent evolution rate value per unit time by combining the time interval parameter, and establishes the rate change curve. The evolution direction determination unit receives the rate change curve output by the evolution rate calculation unit, calculates the slope of the curve, and determines the direction vector of the intention evolution by combining the dimensional change trend of the state parameters after disambiguation, and marks the direction change nodes. The trajectory correction unit adjusts the initially constructed intention evolution trajectory based on the direction vector and change nodes output by the evolution direction determination unit, eliminates abnormal fluctuation points, makes the trajectory more in line with the actual intention evolution process, and outputs the corrected trajectory parameters to the follow-up dialogue streaming analysis platform.
8. The intelligent follow-up voice robot multi-turn dialogue state tracking system according to claim 1, characterized in that, The follow-up dialogue streaming analysis platform includes a dialogue stream receiving unit, a real-time analysis unit, a feature extraction unit, and a data synchronization unit. The dialogue stream receiving unit receives the real-time dialogue data stream transmitted by the follow-up voice robot, performs frame parsing on the data stream, extracts the dialogue content and timestamp information in each frame of data, and establishes a dialogue stream data queue. The real-time analysis unit retrieves the dialogue stream data queue from the dialogue stream receiving unit and, in conjunction with the intent evolution trajectory parameters output by the intent evolution state tracking algorithm module, performs semantic matching analysis on each frame of dialogue data to identify dialogue state change points. Based on the dialogue state change points identified by the real-time analysis unit, the feature extraction unit extracts dialogue feature data before and after the change points, obtains feature vectors related to the dialogue state through a feature association algorithm, and performs dimensionality reduction processing on the feature vectors. The data synchronization unit receives the reduced feature vectors output by the feature extraction unit, associates them with the analysis results of the real-time analysis unit, adds a timestamp, and synchronously transmits them to the multi-turn dialogue state storage module to ensure the timeliness and accuracy of data transmission.
9. A multi-turn dialogue state tracking system for an intelligent follow-up voice robot according to claim 1, characterized in that, The multi-turn dialogue state storage module includes a data classification unit, an index building unit, a storage management unit, and a data retrieval unit. The data classification unit receives analysis results and feature data transmitted from the follow-up dialogue streaming analysis platform, and classifies the data according to the dialogue topic, time period, and patient identification information, establishing different data category directories. The index building unit, based on the category directories established by the data classification unit, extracts keywords from the data under each category, generates data index items, and constructs a dialogue state index library. Each index item includes the data storage address, category identifier, and time information. The storage management unit allocates storage for the data divided by the data classification unit, allocating different storage areas according to data importance and access frequency, and periodically verifies the stored data to repair data corruption issues. The data retrieval unit receives data retrieval requests from each module, parses the data requirement information in the request, searches for the corresponding index item in the dialogue state index library, retrieves the required data according to the storage address in the index item, and feeds it back to the requesting module, recording the data retrieval log.
10. A multi-turn dialogue state tracking system for an intelligent follow-up voice robot according to claims 1-9, characterized in that, The system operation steps include: S1. Receiving the dialogue voice signal collected by the follow-up voice robot and transmitting it to the hierarchical semantic state decoder module. The decoder performs hierarchical parsing of the semantic units in the voice signal and outputs semantic state feature vectors; S2. Transmitting the semantic state feature vectors to the context disambiguation state optimization model module. This model combines historical dialogue state data to disambiguate the feature vectors and generate disambiguated state parameters; S3. Inputting the disambiguated state parameters into the intent evolution state tracking algorithm module. This algorithm constructs the intent evolution trajectory based on the parameters and calculates the evolution rate and direction parameters; S4. Sending the intent evolution trajectory, rate, and direction parameters to the follow-up dialogue streaming analysis platform. This platform dynamically analyzes the real-time dialogue stream and extracts dialogue state association features; S5. Transmitting the analysis results and feature data to the multi-turn dialogue state storage module. This module classifies and stores the data and establishes a dialogue state index library; S6. Retrieving dialogue state data from the multi-turn dialogue state storage module, combining it with the real-time analysis results of the follow-up dialogue streaming analysis platform, generating robot interaction control signals, and transmitting them to the follow-up voice robot for multi-turn dialogue state tracking.