AI multi-technology fusion intelligent dialogue model construction method and system for elderly care
By combining the AI intent recognition model with dynamic semantic features and cross-modal vector sequences, the problem of low efficiency of multimodal information fusion in elderly care scenarios is solved, accurate recognition of the elderly's intentions and compliance of dialogue responses are achieved, and the interactive experience is improved.
Patent Information
- Application Number
- CN202510804147.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing intelligent dialogue technologies for elderly care scenarios mostly rely on a single modality and fail to effectively combine multimodal information such as semantics, voice, and vision, making it difficult to fully capture the user's true intentions. This is especially true when the elderly express themselves vaguely or use non-verbal cues, and misjudgments are prone to occur.
By receiving conversation data streams, extracting dynamic semantic features and combining them with sound signals and mouth movement video signals, a cross-modal vector sequence is generated, an AI intent recognition model is constructed, and the synergy of multi-dimensional features is utilized to achieve the organic combination of multimodal information and cross-modal feature association.
It improves the ability to capture user intentions, reduces the misjudgment rate, adapts to the vague expressions and expression barriers of the elderly, and provides more intelligent and considerate conversation services.
Smart Images

Figure CN120317381B_ABST
Abstract
Description
Technical Field
[0001] The present invention proposes a method and system for constructing an AI multi-technology fusion intelligent dialogue model for elderly care, relating to the technical field of model construction based on AI technology. Background Art
[0002] Smart elderly care services have emerged as a crucial approach to addressing these challenges, creating an increasingly urgent need for intelligent conversational technologies in these settings. However, existing intelligent conversational technologies for elderly care rely solely on a single modality, such as text or voice, and fail to fully integrate multimodal information, including semantics, speech, and vision. This makes it difficult to fully capture the true intent of users in these settings. For example, elderly individuals may experience decreased expressiveness and need to communicate their needs through supplementary information such as tone and lip shape, which cannot be accurately recognized using a single modality. Even some technologies that incorporate multimodality often rely on simple splicing or independent processing, lacking effective cross-modal feature correlation mechanisms. This results in inefficient feature information fusion and limited intent recognition accuracy. Traditional models struggle to effectively capture the dynamic contextual changes in semantics and the temporal correlation of multimodal features during dynamic conversations, leading to significant errors in intent recognition. This is particularly true when elderly individuals express themselves vaguely or include non-verbal cues, making misjudgment more likely. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention proposes a method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care, including:
[0004] S1. Receive a conversation data stream input by a user, and extract a dynamic semantic feature sequence from the conversation data stream;
[0005] S2. Simultaneously collecting a user voice signal and a mouth movement video signal, extracting a voice feature vector from the voice signal, extracting a mouth movement feature vector from the mouth movement video signal, and concatenating the voice feature vector and the mouth movement feature vector according to timestamps to generate a cross-modal vector sequence;
[0006] S3. Build an AI intent recognition model, input the dynamic semantic feature sequence and the cross-modal vector sequence into the AI intent recognition model, and output a user intent category label;
[0007] S4. Generate a response by using the mapping relationship between the intent category label and the dialogue process and the large language model corresponding to the state transition probability matrix output.
[0008] In a preferred embodiment, in step S3, the dynamic semantic feature sequence and the cross-modal vector sequence are fused to obtain a fusion matrix; an AI intent recognition model is constructed, and the fusion matrix is input into the AI intent recognition model to obtain the intent category to which the fusion matrix belongs.
[0009] In a preferred embodiment, the dynamic semantic feature sequence composed of T semantic feature vectors is YD={c1,c2,…c t ,…,c T},c t Represents the semantic feature vector at each time point t, and T cross-modal vectors constitute a cross-modal vector sequence dal={d1,d2,…d t ,…,d T}, d t Represents the cross-modal vector at each time point t; let W q is the query transformation matrix of the semantic feature vector, W k is the key transformation matrix of the cross-modal vector, and the fusion matrix JU is:
[0010] .
[0011] In a preferred embodiment, the AI intent recognition model includes a time series feature formation unit, a pooling unit, and a classification unit; the fusion matrix JU is input into the time series feature formation unit, and the global time series information H is output:
[0012] H=TransformerEncoder(JU);
[0013] Among them, TransformerEncoder is the encoding function, JU={JU1,JU2,…JU t ,…,JU T}, JU t is the tth fusion feature among T fusion features;
[0014] Input the global temporal information H into the pooling unit to obtain the global feature h g :
[0015] h g =MeanPooling(H);
[0016] Among them, MeanPooling is the average pooling operation;
[0017] Calculate the global feature h g The corresponding probability P(y) is the intent category label y:
[0018] ;
[0019] Among them, softmax is the activation function and the weight matrix is Q y , the bias vector is b y .
[0020] In a preferred embodiment, in step S4, a dialog node network T is defined for each intent category label y. y={s1, s2, ...s i …、s V}, where s i is the dialogue node network T mapped based on the intent category label y for the i-th dialogue state node y The v dialogue state nodes in the construct v×v state transition probability matrix Z y , select the dialogue state node pair with the highest state transition probability matrix element value, and combine multiple dialogue state node pairs to form a dialogue process as the backbone process of the currently executed large language model.
[0021] In a preferred embodiment, the sound wave signal is framed and Hamming windowed to obtain a windowed signal x win (n); split the windowed signal x win (n) Perform Fourier transform to obtain the sound wave spectrum signal X(k), and calculate the power spectrum P(k) of the sound wave spectrum signal X(k); n represents the sampling point index, and k is the frequency index;
[0022] Filter the power spectrum through L frequency band filters to extract the energy of each frequency band :
[0023] ;
[0024] in, For the The energy of the frequency band, For the The filtering function of the frequency band filter, N is the total number of sampling points in the frame;
[0025] The energy of each frequency band Perform discrete cosine transform on the logarithm of , extract the cepstral coefficients, and retain the first C cepstral coefficients M(c) as the sound wave feature vector:
[0026]
[0027] in, ; c is the cepstral coefficient index.
[0028] In a preferred embodiment, in step S1, the original session data stream is cleaned by preprocessing, static semantics are extracted using TF-IDF and a pre-trained model, and semantic feature sequences containing dynamic contextual relationships are captured using LSTM.
[0029] In a preferred embodiment, in step S2, the time alignment of the sound signal and the mouth movement video signal is ensured by the trigger signal and the timestamp, and the sampling rate of the sound signal is f a and the video frame rate f of the mouth movement video signal vThe alignment error is less than the frame interval 1 / f of the sound signal a .
[0030] In a preferred embodiment, in step S2, the mouth movement video signal I is cut. mouth , a single-frame image is obtained, and the spatial features of the single-frame image are gradually extracted through the 2D convolution layer; the spatial dimension is reduced through downsampling in the pooling layer; and the features after the spatial dimension reduction are compressed into a low-dimensional mouth action feature vector z through the fully connected layer.
[0031] The present invention also proposes an AI multi-technology fusion intelligent dialogue model construction system for elderly care, which is used to implement the above-mentioned AI multi-technology fusion intelligent dialogue model construction method for elderly care, including: an input interaction module, a multimodal data processing module, an intention recognition module, and a dialogue process generation module;
[0032] The input interaction module is used to receive the conversation data stream input by the user through the remote conversation terminal and extract the dynamic semantic feature sequence in the conversation data stream;
[0033] The multimodal data processing module is used to collect the user's voice signal and mouth movement video signal, extract the voice feature vector from the voice signal, extract the mouth movement feature vector from the mouth movement video signal, and splice the voice feature vector and mouth movement feature vector according to the timestamp to generate a cross-modal vector sequence;
[0034] The intent recognition module is used to build an AI intent recognition model, input dynamic semantic feature sequences and cross-modal vector sequences into the AI intent recognition model, and output user intent category labels;
[0035] The dialogue process generation module uses the mapping relationship between intent category labels and dialogue processes, combined with the large language model corresponding to the state transition probability matrix output, to generate responses.
[0036] Compared with the prior art, the present invention has the following beneficial technical effects:
[0037] 1. By simultaneously extracting dynamic semantic features, vocal features, and mouth movement features from conversation data streams and splicing them together by timestamp to generate a cross-modal vector sequence, this system organically combines text, voice, and visual multimodal information, comprehensively covering both verbal and non-verbal cues in user interactions and improving the ability to capture user intent.
[0038] 2. Input multimodal feature sequences into the AI intent recognition model and utilize the synergistic effect of multi-dimensional features to effectively resolve the ambiguity problem of semantic understanding under a single modality. This is especially true for the ambiguous expressions, tone changes, and lip-assisted expressions of the elderly in elderly care scenarios. This model can more accurately identify the true intention and reduce the misjudgment rate.
[0039] 3. By extracting dynamic semantic feature sequences, the model can track changes in the conversation context in real time. Combined with the temporal information of cross-modal vector sequences, it can better understand the dynamic logic of the conversation, make the conversation response more in line with scenario requirements, and enhance the interactive experience of the elderly.
[0040] 4. By introducing the processing of sound signals and mouth movement video signals, it is specifically adapted to the expression barriers that may exist in the elderly. By using non-verbal modalities to assist semantic understanding, the practicality and robustness of the model in elderly care scenarios are enhanced, providing the elderly with smarter and more intimate conversation services. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0042] Figure 1 This is a schematic diagram of the process of constructing the AI multi-technology fusion intelligent dialogue model for elderly care of the present invention;
[0043] Figure 2 Schematic diagram of the alignment process of the sound wave signal and the mouth movement video signal according to the present invention;
[0044] Figure 3 Schematic diagram of extracting acoustic feature vectors from acoustic signals according to the present invention;
[0045] Figure 4 A three-layer architecture diagram of the intelligent dialogue model construction system of the present invention;
[0046] Figure 5 A line graph for comparative analysis of the present invention and the prior art. DETAILED DESCRIPTION
[0047] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0048] Example 1
[0049] like Figure 1 The figure shows a flowchart of the AI multi-technology fusion intelligent dialogue model construction process for elderly care in this embodiment.
[0050] Specifically, the method for building an AI multi-technology integrated intelligent dialogue model for elderly care includes the following steps:
[0051] S1. Receive a conversation data stream input by a user through a remote conversation terminal, and extract a dynamic semantic feature sequence from the conversation data stream.
[0052] In a remote session terminal environment, the conversation data stream input by users contains rich semantic information. This step cleans the raw conversation data stream through preprocessing, extracts static semantics by combining TF-IDF with a pre-trained model, and then uses LSTM to capture temporal dependencies, ultimately generating a semantic feature sequence that contains dynamic contextual relationships.
[0053] S2. Simultaneously collect the user's vocal signal and mouth movement video signal, extract the vocal feature vector from the vocal signal, extract the mouth movement feature vector from the mouth movement video signal, and concatenate the vocal feature vector and mouth movement feature vector according to timestamps to generate a cross-modal vector sequence.
[0054] S21. Collect the user's voice signal and mouth movement video signal, and align the voice signal and mouth movement video signal.
[0055] like Figure 2 As shown, a microphone array is used to collect sound signals, and a high-speed camera is used to collect mouth movement video signals. The trigger signal and time stamp are used to ensure that the time of the sound signal and the mouth movement video signal are aligned, and the sampling rate of the sound signal is also ensured. a and the video frame rate f of the mouth movement video signal v The alignment error must be less than the frame interval 1 / f of the sound signal a .
[0056] S22. Extracting a sound wave feature vector from the sound wave signal.
[0057] like Figure 3 As shown, first, the sound wave signal is framed and a Hamming window is applied to reduce spectrum leakage. Preferably, the frame length is set to 25ms and the frame shift is set to 10ms.
[0058] After adding the Hamming window to the input framed sound wave signal, the windowed signal x is obtained. win (n):
[0059] ;
[0060] x(n) is the discrete time representation of the nth sampling point of the input sound wave signal, x win (n) is the result of adding a Hamming window to the input signal x(n); N is the total number of sampling points in the frame, 0≤n≤N−1; the coefficient 1 and 2 is the weighting parameter of the Hamming window, which is used to smooth the signal edges.
[0061] Preferably, when sampling at 16 kHz, N=400, then each frame contains N=400 samples, and x(n) is one of the 400 samples.
[0062] Secondly, for the windowed signal x win (n) Perform Fourier transform to obtain the sound wave spectrum signal X(k), and calculate the power spectrum P(k) of the sound wave spectrum signal X(k). The power spectrum represents the energy distribution of the signal in the frequency domain.
[0063]
[0064] Where k is the frequency index, 0≤k≤N / 2.
[0065] The power spectrum is then filtered through L band filters to extract the energy of each band :
[0066]
[0067] in, For the The energy of the frequency band, For the The filter function of the frequency band filter.
[0068] The energy of each frequency band Perform discrete cosine transform on the logarithm of , extract the cepstral coefficients, and retain the first C cepstral coefficients M(c) as the sound wave feature vector:
[0069]
[0070] c is the cepstral coefficient index, c=1, 2, ..., C; L is the number of filters, L is preferably 20-40; C is the number of retained cepstral coefficients, C is preferably 12-13.
[0071] It should be noted that the sound wave feature vector is the short-time spectrum feature of the sound wave signal, which characterizes the shape of the vocal tract and the characteristics of the sound source. The cepstral coefficients that are not used as sound wave feature vectors are high-order coefficients that characterize details or noise and are discarded.
[0072] S23. Extracting a mouth movement feature vector from the mouth movement video signal.
[0073] Cropping mouth movement video signal I mouth , get a single frame image as the input of the feature extraction unit.
[0074] The feature extraction unit first performs feature compression on the single-frame image, and uses 2D convolutional layers, pooling layers, and fully connected layers to extract feature vectors from the single-frame image that reflects static mouth features.
[0075] Specifically, the 2D convolutional layer gradually extracts the spatial features of a single frame image. The spatial features include: extracting the key points of the lip contour, calculating the distance between the left and right mouth corners, the distance between the midpoints of the upper and lower lips, the pixel area of the lip region, etc., realizing an incremental increase in the number of channels, for example, 16→32→64.
[0076] Then through the pooling layer, the spatial dimension is reduced by downsampling.
[0077] Finally, through the fully connected layer: the features with reduced spatial dimensions are compressed into a low-dimensional mouth movement feature vector z, thereby capturing the global features of the mouth geometry.
[0078] In a specific embodiment, a 64×64×3 single-frame image is input, and after three convolution-pooling operations, 8×8×64 features are obtained, and a 64-dimensional mouth action feature vector z is output through a fully connected layer.
[0079] S24. Concatenate the sound wave feature vector and the mouth movement feature vector according to the timestamps to generate a cross-modal vector sequence.
[0080] Let the sound wave feature vector corresponding to each time point t be v t , the corresponding mouth movement feature vector is k t . Concatenate the sound wave feature vector and the mouth movement feature vector according to the timestamp to generate a cross-modal vector concatenation and generate a cross-modal vector sequence dal:
[0081] dal={d1,d2,…d t ,…,d T},d t =[v t ;k t ];
[0082] The cross-modal vector sequence dal is the sound wave feature vector v containing T time points output in chronological order t and mouth movement feature vector k t time series.
[0083] S3. Build an AI intent recognition model, input the dynamic semantic feature sequence and cross-modal vector sequence into the AI intent recognition model, and output the user intent category label.
[0084] S31. Fuse the dynamic semantic feature sequence and the cross-modal vector sequence to obtain a fusion matrix.
[0085] Assume that the dynamic semantic feature sequence is: YD={c1,c2,…c t ,…,cT}, represents the semantic feature vector c at each time point t t There are T semantic feature vectors in the dynamic semantic feature sequence; the cross-modal vector sequence is dal={d1,d2,…d t ,…,d T}.
[0086] The dynamic semantic feature sequence YD and the cross-modal vector sequence dal are fused to obtain the fusion matrix JU.
[0087] Specifically, let W q is the query transformation matrix of the semantic feature vector, W k is the key transformation matrix across modal vectors,
[0088] ;
[0089] The fusion matrix JU is:
[0090] .
[0091] is the fusion feature at the t-th time point.
[0092] S33. Build an AI intent recognition model, input the fusion matrix into the AI intent recognition model, and obtain the intent category to which the fusion matrix belongs.
[0093] The AI intent recognition model includes a time series feature formation unit, a pooling unit, and a classification unit.
[0094] First, the fusion matrix JU is input into the time series feature formation unit. The input parameters include T fusion features. The time series feature formation unit outputs the global time series information H based on the TransformerEncoder:
[0095] H=TransformerEncoder(JU),JU={JU1,JU2,…JU t ,…,JU T};
[0096] TransformerEncoder is the encoding function of the time series feature formation unit, which encodes the input parameters. H is the global time series information output after TransformerEncoder processing, which reflects the global dependency of the input fusion matrix JU.
[0097] Preferably, TransformerEncoder is encoded by a multi-head self-attention mechanism and a feedforward neural network. In the multi-head self-attention mechanism, the input fusion matrix JU={JU1,JU2,…JU t ,…,JU T Linearly project the information into multiple query, key, and value matrices, calculate the attention weight in each head, concatenate the results of each head, and linearly transform them to capture global information, thereby outputting global temporal information H that can represent the global situation.
[0098] Input the global timing information H into the pooling unit: h g =MeanPooling(H);
[0099] The pooling unit performs average pooling operation through MeanPooling. The function of MeanPooling is to average pool the input global time series information H and compress the global time series information H into a representative global feature h g , providing fixed-dimensional input to the classification unit.
[0100] The classification unit calculates the probability of the global feature h given the input g In the case of , calculate the global feature h g The probability of the corresponding intent category label y.
[0101] The specific calculation process is to first transform the global feature h g Through the weight matrix Q y Perform a linear transformation and add the bias vector b y , and then the result is converted into a probability distribution through the softmax activation function. It should be noted that the sum of the probabilities of all intent category labels is 1.
[0102] Specifically, the probability calculation formula is:
[0103]
[0104] Among them, y is the intention category label, and the softmax activation function is used to obtain the input global feature h g The probability P(y) of belonging to the intent class label y.
[0105] This step can significantly improve the accuracy of user intent recognition in noisy environments or fuzzy semantic scenes by deeply fusing semantics, sound information, and mouth shape information. It is suitable for intent understanding tasks in intelligent interactive systems and provides an efficient solution for intent recognition in complex scenarios.
[0106] S4. Generate a response by using the mapping relationship between the intent category label and the dialogue process and the large language model corresponding to the state transition probability matrix output.
[0107] S41: Define the intent category label-response mapping table and construct the state transition probability matrix.
[0108] (1) Define the mapping relationship between intent category labels and dialogue processes.
[0109] Define a dialog node network T for each intent category label y y ={s1, s2, ...s i …、s V}, where s i is the conversation content corresponding to the conversation state node i, such as "greeting", "asking for needs", "confirming information", "providing results", "ending conversation", etc., and v is the total number of conversation state nodes.
[0110] This mapping relationship standardizes the static correspondence between intent category labels and standardized dialog nodes, providing a structured framework for the subsequent construction of state transition probability matrix and process execution.
[0111] (2) Construct the state transition probability matrix based on the mapping relationship.
[0112] Dialogue node network T based on intent category label y mapping y The conversation content corresponding to the v dialogue state nodes in the construct v×v state transition probability matrix Z y , where the matrix elements Z ij Indicates the intention category label y, from the conversation content s corresponding to the conversation state node i to s j The transition probability.
[0113] It should be noted that the state transition probability matrix Z y It can be obtained in the following ways:
[0114] Statistical analysis: Based on historical conversation data, calculate the conversation content s corresponding to the conversation state node i to s j The frequency ratio of .
[0115] Reinforcement learning: Through interactive learning between the agent and the environment, the optimal state transition strategy is optimized.
[0116] This step injects dynamic rules into the static correspondence between intent category labels and standardized dialogue node paths. yIt reflects the patterns of branching, jumping, or ending of the dialogue process within the standard path caused by user behavior or system decision-making in actual dialogue interactions. For example, the probability of the user jumping directly to the end state after confirming the information in advance is the key basis for the dynamic execution of subsequent processes.
[0117] S42: Use the state transition probability matrix to output the corresponding large language model.
[0118] According to the matrix element Z of the state transition probability matrix ij The dialog state node pair corresponding to the element position with the highest matrix element value is selected, that is, the dialog state node pair with the highest state transition probability, to drive the actual flow of the dialog state nodes. Multiple dialog state node pairs are combined to form the corresponding dialog flow as the main flow of the currently executed large language model.
[0119] In a specific embodiment, suppose the system is in the current dialogue state node i dialogue content s i , according to the matrix element Z ij The state transition probability represented by determines the conversation content s of the next dialogue state node j to be transferred to j , execute the conversation content s of conversation state node j j ; Repeat this process until the flow end state is triggered.
[0120] This step clearly demonstrates the complete logical chain from intent to conversation process mapping, to dynamic rule definition within the conversation, and finally to content matching and execution based on state transition probability.
[0121] Example 2
[0122] This embodiment aims to meet the diverse communication needs of the elderly in elderly care scenarios and forms an intelligent dialogue model construction system that integrates input layer, processing layer and output layer. Figure 4 As shown in the figure, the three-layer architecture diagram of the intelligent dialogue model construction system.
[0123] Input layer: Receives user input in three modalities: text, voice, and vision.
[0124] Processing layer: Processes text, voice, and visual data separately, then integrates the information through a multimodal fusion module, and performs intent recognition and dialogue management.
[0125] Output layer: Generates text responses, speech synthesis, and action commands.
[0126] Based on the ideas in the above architecture diagram, in the preferred embodiment, the intelligent dialogue model construction system is specifically designed as the following structural units, including: input interaction module, multimodal data processing module, intent recognition module, and dialogue process generation module. Four core functional units are used. Each module realizes information flow and collaborative work through a standardized data interface, forming a complete processing chain from user input to dialogue response.
[0127] (1) Input interaction module
[0128] The input interaction module includes: a conversation data stream acquisition unit and a multimodal signal acquisition unit.
[0129] The session data stream collection unit is responsible for receiving session data streams input by users via remote session terminals. It supports access from a variety of terminal devices, such as smart speakers, elderly care monitoring terminals, and mobile apps. A real-time data transmission protocol is used to ensure low-latency data transmission. An integrated data preprocessing submodule cleans the raw session data stream to remove noise, duplicate information, and invalid characters. The TF-IDF algorithm is used to extract static semantic features of keywords in the text. Deep semantic encoding is performed using a pretrained language model (such as BERT). An LSTM network is then used to capture dynamic semantic relationships within the context, generating a dynamic semantic feature sequence that includes temporal dependencies.
[0130] The multimodal signal acquisition unit is equipped with a highly sensitive microphone array and a high-definition camera to capture user voice signals and mouth movement video signals in real time. The microphone array supports 360-degree sound acquisition, effectively reducing ambient noise interference. The camera uses infrared fill light technology to ensure clear capture of mouth movements in low-light environments. A hardware trigger signal and a high-precision timestamp synchronization mechanism achieve time alignment between the voice signal and the mouth movement video signal.
[0131] (2) Multimodal Data Processing Module
[0132] The multimodal data processing module includes: sound wave feature extraction submodule, mouth movement feature extraction submodule, and cross-modal vector fusion submodule.
[0133] The sound wave feature extraction submodule performs frame segmentation and windowing on the collected sound wave signal. It then filters the power spectrum using a Fourier transform and a filter to extract the energy of each frequency band. The logarithm of the energy of each frequency band is then subjected to a discrete cosine transform to extract multiple sound wave feature vectors, forming a sequence of sound wave feature vectors. The mouth movement feature extraction submodule uses computer vision technology to process the mouth movement video signal. It first locates the facial region using a face detection algorithm, then accurately tracks the lip contour and extracts dynamic feature parameters such as lip opening, angle of mouth corner lift, and lip movement speed. These parameters are normalized to generate a sequence of mouth movement feature vectors.
[0134] The cross-modal vector fusion submodule concatenates the vocal feature vector and the mouth movement feature vector according to timestamps to generate a cross-modal vector sequence. This cross-modal vector sequence serves as an important input for subsequent intent recognition, addressing the limitations of single text semantics and enhancing the ability to capture user emotions and expressive intent.
[0135] (3) Intent Recognition Module
[0136] The intent recognition module includes: feature fusion unit and intent classification unit.
[0137] The feature fusion unit constructs an interaction mechanism between semantic features and cross-modal features, introduces an attention mechanism to achieve deep fusion of the two, and outputs global temporal information containing contextual semantics and multimodal information.
[0138] The intent classification unit performs average pooling on global temporal information to generate a global feature vector. Using a fully connected layer and a softmax function, it calculates the probability distribution of the intent category labels corresponding to these global features. The intent classification unit supports dynamic updates, expanding intent categories based on new user needs in elderly care scenarios and improving classification adaptability.
[0139] (4) Dialogue process generation module
[0140] The dialogue process generation module includes a dialogue node network construction unit, a state transition probability calculation unit, and a large language model adaptation unit.
[0141] The dialogue node network construction unit predefines a dialogue node network for each intent category. The dialogue node network includes multiple dialogue state nodes. The dialogue node network covers common needs in elderly care scenarios, such as health consultation, life service appointments, emergency assistance, emotional companionship, etc.
[0142] The state transition probability calculation unit calculates the transition frequency between dialogue state nodes for each intent category based on historical conversation data. This unit constructs a state transition probability matrix, where each element represents the probability of transitioning from one state node to another. This matrix smoothes the transition probability and avoids zero-probability issues. When generating the conversation flow, starting from the initial state node corresponding to the current intent, the node pairs with the highest state transition probability are selected in sequence to form a conversation flow chain, which serves as the backbone of the large language model. This flow supports branch selection, dynamically adjusting the conversation path based on subsequent user input to achieve a personalized interactive experience.
[0143] The large language model adaptation unit adapts the generated conversation flow to the large language model, using control information within the process nodes to guide the large language model in generating responses that meet scenario requirements. Simultaneously, responses are semantically verified using a knowledge base in the elderly care field to ensure accuracy and security.
[0144] Example 3
[0145] In AI conversational systems for senior care, conversational accuracy primarily measures the model's ability to understand user intent and generate accurate and contextually relevant responses. Key metrics include intent recognition accuracy (the percentage of correctly identified user intent), response relevance (the degree to which the response matches the intent), contextual coherence (how well it handles multiple rounds of conversation), robustness (tolerance to noise, ambiguous input, or speech problems unique to the elderly), and scenario-specific adaptability (optimization for senior care needs).
[0146] like Figure 5 As shown, the intelligent dialogue model of the present invention performs better on the elderly care scenario test data.
[0147] In actual deployment, data quality and training algorithms will affect accuracy. This intelligent dialogue model construction system significantly improves the ability to understand the complex expressions of the elderly by integrating multimodal information such as text semantics, sound characteristics, and mouth movements. It is especially suitable for special elderly groups such as those with hearing impairments and unclear language expression. The dialogue process generation mechanism based on the state transition probability matrix enables the system to dynamically adjust the interaction strategy according to user needs, providing a more natural and smooth conversation experience. In actual applications, this intelligent dialogue model construction system can be deployed in scenarios such as smart elderly care communities and home-based elderly care monitoring equipment to achieve multi-functional integration of health management, life services, emotional companionship, etc., providing all-round intelligent interaction support for the elderly and facilitating the digitalization and intelligent upgrading of elderly care services.
[0148] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. The database involved in the embodiments provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on a regional block chain, etc., but is not limited to this. The processor involved in the embodiments provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited to this.
[0149] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0150] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care, characterized by: include: S1. Receive a conversation data stream input by a user, and extract a dynamic semantic feature sequence from the conversation data stream; S2. Simultaneously collecting a user voice signal and a mouth movement video signal, extracting a voice feature vector from the voice signal, extracting a mouth movement feature vector from the mouth movement video signal, and concatenating the voice feature vector and the mouth movement feature vector according to timestamps to generate a cross-modal vector sequence; S3. Build an AI intent recognition model, input the dynamic semantic feature sequence and the cross-modal vector sequence into the AI intent recognition model, and output a user intent category label; S4. Generate a response by using the mapping relationship between intent category labels and conversation flow, combined with the large language model corresponding to the state transition probability matrix output; Define a dialog node network T for each intent category label y y ={s1, s2, ...s i …、s V }, where s i is the conversation content corresponding to the conversation state node i, and the conversation node network T mapped based on the intent category label y y The v dialogue state nodes in the construct v×v state transition probability matrix Z y , select the dialogue state node pair corresponding to the matrix element position with the highest state transition probability matrix element value, and combine multiple dialogue state node pairs to form a dialogue process as the backbone process of the currently executed large language model.
2. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 1 is characterized in that: In step S3, the dynamic semantic feature sequence and the cross-modal vector sequence are fused to obtain a fusion matrix; an AI intent recognition model is constructed, and the fusion matrix is input into the AI intent recognition model to obtain the intent category to which the fusion matrix belongs.
3. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 2 is characterized in that: Suppose the dynamic semantic feature sequence composed of T semantic feature vectors is YD={c1,c2,…c t ,…,c T },c t Represents the semantic feature vector at each time point t, and T cross-modal vectors constitute a cross-modal vector sequence dal={d1,d2,…d t ,…,d T }, d t Represents the cross-modal vector at each time point t; let W q is the query transformation matrix of the semantic feature vector, W k is the key transformation matrix of the cross-modal vector, and the fusion matrix JU is: 。 4. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 3 is characterized in that: The AI intent recognition model includes a time series feature formation unit, a pooling unit, and a classification unit. The fusion matrix JU is input into the time series feature formation unit, and the global time series information H is output: H=TransformerEncoder(JU); Among them, JU={JU1,JU2,…JU t ,…,JU T }, JU t is the tth fused feature among T fused features, TransformerEncoder is the encoding function; Input the global temporal information H into the pooling unit to obtain the global feature h g : h g =MeanPooling(H); Among them, MeanPooling is the average pooling operation; Calculate the global feature h g The corresponding probability P(y) is the intent category label y: ; Among them, softmax is the activation function and the weight matrix is Q y , the bias vector is b y .
5. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 1 is characterized in that: The sound wave signal is framed and processed with a Hamming window to obtain the windowed signal x win (n); split the windowed signal x win (n) Perform Fourier transform to obtain the sound wave spectrum signal X(k), and calculate the power spectrum P(k) of the sound wave spectrum signal X(k); n represents the sampling point index, and k is the frequency index; Filter the power spectrum through L frequency band filters to extract the energy of each frequency band : ; in, For the The energy of the frequency band, For the The filtering function of the frequency band filter, N is the total number of sampling points in the frame; The energy of each frequency band Perform discrete cosine transform on the logarithm of , extract the cepstral coefficients, and retain the first C cepstral coefficients M(c) as the sound wave feature vector: ; in, ; c is the cepstral coefficient index.
6. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 1 is characterized in that: In step S1, the original session data stream is cleaned by preprocessing, static semantics are extracted using TF-IDF and a pre-trained model, and semantic feature sequences containing dynamic contextual relationships are captured using LSTM.
7. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 1 is characterized in that: In step S2, the time alignment of the sound signal and the mouth movement video signal is ensured by the trigger signal and the timestamp, and the sampling rate of the sound signal is f a and the video frame rate f of the mouth movement video signal v The alignment error is less than the frame interval 1 / f of the sound signal a .
8. The method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care according to claim 1 is characterized in that: In step S2, the mouth movement video signal I is cut. mouth , get a single frame image, and gradually extract the spatial features of the single frame image through the 2D convolution layer; The spatial dimension is reduced by downsampling through the pooling layer; the features after spatial dimension reduction are compressed into a low-dimensional mouth movement feature vector z through the fully connected layer.
9. An AI multi-technology fusion intelligent dialogue model construction system for elderly care, characterized by: A method for constructing an AI multi-technology fusion intelligent dialogue model for elderly care as described in any one of claims 1 to 8, comprising: an input interaction module, a multimodal data processing module, an intent recognition module, and a dialogue process generation module; The input interaction module is used to receive the conversation data stream input by the user through the remote conversation terminal and extract the dynamic semantic feature sequence in the conversation data stream; The multimodal data processing module is used to collect the user's voice signal and mouth movement video signal, extract the voice feature vector from the voice signal, extract the mouth movement feature vector from the mouth movement video signal, and splice the voice feature vector and mouth movement feature vector according to the timestamp to generate a cross-modal vector sequence; The intent recognition module is used to build an AI intent recognition model, input dynamic semantic feature sequences and cross-modal vector sequences into the AI intent recognition model, and output user intent category labels; The dialogue process generation module uses the mapping relationship between intent category labels and dialogue processes, combined with the large language model corresponding to the state transition probability matrix output, to generate responses.
Citation Information
Patent Citations
AI intelligent customer service response method and system based on remote digital service
CN119719319A
Multi-modal large language model dialogue generation method based on natural language understanding
CN119989268A
Intelligent customer service dialogue generation method and system based on user intention recognition
CN119990149A
Cited By
Dynamic intention recognition and multi-agent seamless switching session system and method based on large language model
CN121301230A