Intelligent screening system for chronic obstructive pulmonary disease based on voice interaction

CN122822332APending Publication Date: 2026-09-25PEKING UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610983611.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有的慢阻肺筛查与管理技术在面向社区及老年人群时,存在明显的局限性:

Benefits of technology

[0018]本发明提供的基于语音交互的慢阻肺智能筛查系统,通过用户端对用户语音进行识别对用户数据进行收集,无需复杂手动操作的无障碍交互方式,并基于云端风险评估与数据管理平台对收集的数据进行处理,完成慢阻肺的筛查,从而降低筛查门槛,达到提高筛查效率与可及性的目的。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822332A_ABST
    Figure CN122822332A_ABST
Patent Text Reader

Abstract

The application provides a kind of slow obstructive lung intelligent screening system based on voice interaction, comprising: including user end, cloud risk assessment and data management platform and collaborative end;The user end is used to obtain user voice answer, slow obstructive lung screening questionnaire and symptom scale collection data and test data;The cloud risk assessment and data management platform are used to calculate risk score based on the data received from the user end and generate screening grading results;The collaborative end is used to view the screening records, grading results and high-risk population list of the cloud risk assessment and data management platform, and support result correction and follow-up management.Through the user end, the user data is collected by recognizing the user voice, without complex manual operation barrier-free interaction mode, and based on the cloud risk assessment and data management platform, the collected data is processed to complete the screening of slow obstructive lung, thereby reducing the screening threshold and achieving the purpose of improving screening efficiency and accessibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart medical technology, and in particular relates to a smart screening system for COPD based on voice interaction. Background Technology

[0002] Chronic obstructive pulmonary disease (COPD) is a common, preventable and treatable chronic airway disease. Early screening, diagnosis and intervention are crucial for slowing the progression of COPD, improving patients' quality of life and reducing the disease burden.

[0003] Existing COPD screening and management technologies have significant limitations when applied to community and elderly populations: First, existing human-computer interaction methods are not user-friendly for the elderly. For example, a community-based early screening and intervention system for chronic obstructive pulmonary disease (COPD) relies primarily on residents manually filling out screening questionnaires and scales such as CAT and mMRC on mobile devices. This poses a significant barrier to use for elderly people with vision loss, dexterity issues, or unfamiliarity with the complex operation of smartphones, resulting in low screening participation rates, incomplete data collection, and difficulty in truly covering high-risk groups.

[0004] Second, existing technologies lack intelligent and barrier-free screening methods for people with low operational capabilities. Although some advanced systems have introduced voice interaction technology, their core application scenario is for long-term disease management of diagnosed patients, focusing on patient education and data tracking through natural language dialogue, rather than for initial screening. Their voice functions have not been optimized for the standardized collection of screening questionnaires, especially lacking support for the recognition of multiple dialects, which are extremely common among the elderly. This limits the feasibility of widespread application of such technologies in grassroots communities.

[0005] Third, the existing community screening process lacks sufficient automation and intelligence. Current technologies, such as a community-based early screening and intervention system for chronic obstructive pulmonary disease (COPD), establish a management pathway from the community to specialist physicians. However, the initial assessment of screening results, risk stratification, and subsequent action recommendations (such as which screening tools to use and whether referral is necessary) still heavily rely on manual judgment by community doctors or specialists. The system fails to achieve automatic and rapid risk stratification and personalized recommendations based on multi-source data (such as questionnaire responses and simple physiological tests). This not only increases the workload of medical staff but also makes it difficult to achieve efficient large-scale initial screening in grassroots settings with relatively scarce medical resources.

[0006] In summary, existing technologies have limitations in collecting information from specific groups such as those with poor eyesight or who are illiterate, which makes it difficult to screen for COPD in a timely and accurate manner. Summary of the Invention

[0007] To address the problems existing in the prior art, the present invention provides a voice-interactive intelligent screening system for COPD.

[0008] This disclosure provides a voice-interactive intelligent screening system for COPD, comprising: This includes the user end, cloud-based risk assessment and data management platform, and collaboration end; The user terminal is used to obtain user voice responses, COPD screening questionnaires and symptom scale data, and test data. The cloud-based risk assessment and data management platform is used to calculate risk scores and generate screening and grading results based on data received from user terminals. The collaborative terminal is used to view the screening records, classification results, and high-risk population list of the cloud-based risk assessment and data management platform, and supports result correction and follow-up management. The user terminal includes a voice interaction module, a questionnaire and scale collection module, and an auxiliary test collection module; The voice interaction module is used to play screening questions by voice and receive user voice responses, and convert the voice responses into response text; The questionnaire and scale collection module is used to collect data according to the preset COPD screening questionnaire and symptom scale. The auxiliary test acquisition module is used to acquire simplified breathing test and walking test data; The cloud-based risk assessment and data management platform includes a speech recognition and semantic parsing module and a joint risk classification engine; The speech recognition and semantic parsing module is used to recognize, clean, and structure the response text from the user terminal. The joint risk grading engine is used to perform fusion analysis on the response text processed by the speech recognition and semantic parsing modules and auxiliary test data, automatically calculate risk scores and generate screening and grading results.

[0009] Optionally, the user terminal also includes an offline caching module, which is used to store screening records and voice data in a network-free environment and automatically encrypt and synchronize them to the cloud after connecting to the network.

[0010] Optionally, the voice interaction module is also used to send the converted response text to the user and generate a confirmation response text based on the user's instructions; The voice interaction module includes dialect recognition, and the dialect recognition module includes: The segment length is dynamically calculated by calling the Fourier spectrum continuity parameter and the zero crossover rate, and long speech is segmented into dialect segments; High-pass filtering and wavelet shift denoising are applied to the segmented dialect data to remove environmental background noise. Feature vectors of dialect speech are extracted through encoder and attention mechanism. Based on the acquired geographical location, the target dialect type is matched in the preset dialect lexicon based on the feature vector of the dialect speech. Based on the matched target dialect type, the corresponding dialect conversion model is used to perform cascade fusion of voiceprint vector features and phonological features to obtain structured text information. The structured text information is then converted into specific scores of the COPD scale according to preset regulations.

[0011] Optionally, the voice interaction module supports languages ​​including Mandarin and dialects, and achieves age-friendly interaction through voice broadcasting, large fonts, and large button interfaces.

[0012] Optionally, the cloud-based risk assessment and data management platform may also include a questionnaire and scale management module and a data storage and statistical analysis module; The questionnaire and scale management module is used to manage questionnaires and scales. The data storage and statistical analysis module connects the user to the joint risk classification engine, stores the data from the joint risk classification engine, and performs statistical analysis on the data from the joint risk classification engine.

[0013] Optionally, the joint risk grading engine includes a pre-trained WH-ESPC model, which is used to process the acquired CT image data to obtain screening results. The WH-ESPC model includes a feature extraction module based on the CaFormer architecture, a multi-sequence CT feature extraction module based on historical memory and load balancing hybrid expert network, and a classification decision module based on bidirectional LSTM. The feature extraction module based on the CaFormer architecture is used to extract abstract features from a single CT image to obtain a multi-dimensional feature vector. The hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing is used to adaptively extract key lesion features from multi-sequence CT images. The bidirectional LSTM-based classification decision module is used to aggregate sequence features based on an attention mechanism and output screening results.

[0014] Optionally, the hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing includes a hybrid expert model basic architecture, a historical information backtracking module, and a regularized load balancing adjustment. The hybrid expert model infrastructure includes shared experts and routing experts; The data processing flow of the hybrid expert model infrastructure includes: Basic features are extracted from the input features by sharing experts; Based on a two-stage gated routing mechanism, the probability of routing to each routing expert is calculated according to the extracted basic features; Select k routing experts based on probability results; The feature representation is obtained by fusing the outputs of the shared expert with the outputs of the selected k routing experts; The formula for calculating the probability of routing to each routing expert is as follows: , in, Let be the probability of being assigned to the k-th routing expert. For the first A load balancing solution from a routing expert. Based on the feature vector, These are the weight parameters.

[0015] Optionally, in the two-stage gated routing mechanism, the load balancing item of the routing expert is updated based on the historical information backtracking module, which integrates the load data of historical iterations through a sliding window mechanism to smooth training fluctuations.

[0016] Optionally, the update formula for the load balancing item is: , in, and These are the load balancing bias terms for the k-th routing expert in the t-th and t+1-th iterations, respectively. Let the total workload allocated to the k-th routing expert in the t-th iteration be _____. Let the average load of the routing expert be the value in the t-th iteration. is the regularization coefficient.

[0017] Optionally, the classification decision module based on bidirectional LSTM outputs screening results, including: The input sequence features are processed by a two-layer bidirectional LSTM network to obtain the forward and backward hidden states at each time step; Attention-weighted aggregation of the hidden states at all time steps is performed to generate a global feature representation; The global features are mapped to the classification space through a fully connected layer and a softmax function, and a probability distribution is generated.

[0018] The COPD intelligent screening system provided by this invention collects user data by recognizing user voice on the user terminal. It is an accessible interactive method that does not require complicated manual operation. The collected data is processed based on a cloud-based risk assessment and data management platform to complete the COPD screening, thereby lowering the screening threshold and improving screening efficiency and accessibility. Attached Figure Description

[0019] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0020] Figure 1 A schematic diagram of the COPD intelligent screening system based on voice interaction provided in the embodiments of this disclosure; Figure 2 A schematic diagram illustrating the principle framework of the WH-ESPC model provided in this embodiment of the disclosure; Figure 3 A schematic diagram of the principle of the hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing provided in the embodiments of this disclosure. Detailed Implementation

[0021] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0022] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0023] It should be noted that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice method. Furthermore, this device and / or practice method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.

[0024] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The illustrations only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0025] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0026] like Figure 1 As shown, this embodiment discloses a voice-interactive intelligent screening system for COPD, characterized in that it includes: This includes the user end, cloud-based risk assessment and data management platform, and collaboration end; The user terminal is used to obtain user voice responses, COPD screening questionnaires and symptom scale data, and test data. The cloud-based risk assessment and data management platform is used to calculate risk scores and generate screening and grading results based on data received from user terminals. The collaborative terminal is used to view the screening records, classification results, and high-risk population list of the cloud-based risk assessment and data management platform, and supports result correction and follow-up management. The user terminal includes a voice interaction module, a questionnaire and scale collection module, and an auxiliary test collection module; The voice interaction module is used to play screening questions by voice and receive user voice responses, and convert the voice responses into response text; The questionnaire and scale collection module is used to collect data according to the preset COPD screening questionnaire and symptom scale. The auxiliary test acquisition module is used to acquire simplified breathing test and walking test data; the auxiliary test acquisition module uses the mobile phone microphone and sensors to acquire simplified breathing test and walking test data; The cloud-based risk assessment and data management platform includes a speech recognition and semantic parsing module and a joint risk classification engine; The speech recognition and semantic parsing module is used to recognize, clean, and structure the response text from the user terminal. The joint risk grading engine is used to perform fusion analysis on the response text processed by the speech recognition and semantic parsing modules and auxiliary test data, automatically calculate risk scores and generate screening and grading results.

[0027] The user terminal also includes an offline caching module, which stores screening records and voice data in offline environments and automatically encrypts and synchronizes them to the cloud upon reconnection. The offline caching module completes voice prompt playback, answer recording, and preliminary scoring in offline conditions, and automatically uploads the cached data to the cloud upon reconnection, where the cloud re-evaluates and updates the screening results.

[0028] The voice interaction module is also used to send the converted response text to the user and generate a confirmation response text based on the user's instructions. That is, the converted text is sent to the user; if the conversion is incorrect, the user can re-convert, or modify the incorrect parts and confirm via interaction.

[0029] After the system starts up, the user terminal obtains the user's voice response and simultaneously uses the mobile phone's sensors to obtain geographical location information and auxiliary test data (such as acceleration for walking tests).

[0030] The system monitors network status in real time. If the network is unavailable, it enters "offline caching" mode to complete voice segmentation and preliminary scoring locally; once the network is restored, it is automatically encrypted and synchronized to the cloud to ensure data continuity and integrity.

[0031] In response to the characteristics of shortness of breath and irregular pauses in COPD patients, especially the elderly, the Fourier spectrum continuity parameter and zero crossover rate are used to dynamically calculate the segment length, and long speech is segmented into typical dialect segments suitable for recognition.

[0032] High-pass filtering and wavelet shift denoising are applied to the segmented data to remove environmental background noise. Then, feature vectors of typical dialect speech are extracted through an encoder and attention mechanism.

[0033] Based on the acquired geographic location, the system accurately matches the target dialect type within a pre-defined dialect lexicon. A dialect conversion model is used to cascade and fuse voiceprint vector features and phonological features, outputting highly accurate structured text information. The identified dialect descriptions (such as "not able to breathe" or "chest tightness") are then converted into specific scores on the standard COPD scale (CAT or mMRC).

[0034] The joint risk grading engine integrates and analyzes voice scale scores with auxiliary test data (breathing audio characteristics, gait stability) to automatically calculate a comprehensive risk score. Based on the risk level, the system generates personalized screening recommendations. Results are synchronized to both doctors and family members.

[0035] Collaborators (family members or doctors) correct any erroneous text. The corrected data, along with the original speech, is then fed back to the dialect lexicon and the recognition model as a reward signal for reinforcement learning, driving the system to continuously improve its recognition accuracy.

[0036] The voice interaction module supports Mandarin and dialects, and achieves age-friendly interaction through voice broadcasting, large fonts and large buttons, allowing users to complete the screening without manual input or text reading.

[0037] Additionally, the voice interaction module can support languages ​​such as English and French, depending on the requirements. The cloud-based risk assessment and data management platform also includes a questionnaire and scale management module and a data storage and statistical analysis module. The questionnaire and scale management module is used to manage questionnaires and scales. The data storage and statistical analysis module allows users to connect to the joint risk classification engine, store the data from the joint risk classification engine, and perform statistical analysis on the data.

[0038] The collaborative platform includes a family member app and a doctor app, used to view screening records, triage results, and lists of high-risk groups, and supports result correction and follow-up management. The family member app is used to correct the identified text and associate the corrected content with the original record to improve the accuracy of screening responses. The doctor app automatically compiles lists of high-risk individuals within the community and generates screening statistical reports based on age, region, and risk level for subsequent follow-up and management.

[0039] This implementation allows users to install the client on mobile phones and other electronic devices via a mini-program.

[0040] This embodiment proposes an improved load balancing method in the WH-ESPC model and introduces a historical memory information module, which allows the MoE module to take into account past expert selection, thereby improving the model's performance and stability and providing a new approach for feature extraction from multi-sequence medical image data.

[0041] like Figure 3 As shown, the WH-ESPC model uses the softsign regularization mechanism, which increases the training stability of the model, enabling the model to obtain better training results more stably.

[0042] Mixture of Experts (MoE) models are a highly efficient sparse neural network architecture. The core idea is to dynamically select which expert sub-models should process the input data through a gating network, thereby achieving conditional computation and significantly increasing the upper limit of model parameters while maintaining low computational cost. Early MoE models employed computationally intensive methods, such as the Jacobs et al. MoE ensemble framework, which constructed multiple parallel expert sub-networks and used a gating network to learn to assign weights to the outputs of all experts for weighted summation. Since all experts need to process every input, the computational cost increases linearly with the number of experts. This characteristic makes it a computationally intensive model, thus limiting its scalability. To address this bottleneck, modern MoE architectures introduce a revolutionary dynamic sparse routing mechanism. Shazeer et al. redefined the function of the gating network, moving beyond weight assignment to dynamically selecting one or a few of the most relevant experts for activation computation for each input token. This strategy demonstrates significant advantages when combined with the Transformer architecture: most of the model's structure is shared by all tokens, with only the computationally intensive feedforward network layers replaced by sparsely routed MoE layers. The SwitchTransformer is a prime example of this "shared + sparse routing" strategy. By activating only a single expert's minimal routing, it scales model parameters to the trillions level with almost no change in computational cost, demonstrating the enormous potential of modern MoE architectures in building ultra-large-scale models. Building upon the SwitchTransformer, researchers have proposed many important improvements. For instance, the DeepSeek team introduced the concept of shared experts, enhancing model stability. The DeepSeek team also proposed a variance-based load balancing method, achieving an unsupervised load balancing adjustment mechanism. However, the application of MoE methods in medical imaging is still limited. This is because medical images have a low sample size and complex case features, making it difficult to provide a sufficiently large dataset for the model to learn from during the training phase.

[0043] like Figure 2As shown, the WH-ESPC model comprises three key modules: The first is a Caformer-based feature extraction module. This module extracts abstract features from each CT image sequentially using a pre-trained Caformer image feature extractor, providing the foundation for the next step of building a multi-sequence feature extraction module. The second is a hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing. It designs a hybrid expert network based on a two-stage gated routing mechanism and proposes a load-balanced weighted model based on historical information backtracking and softsign regularization. This module can adaptively extract key information from multi-sequence CT image data. The third is an LSTM-based classification decision module, which designs a bidirectional LSTM network to achieve classification output. Details are as follows: Caformer is an image feature extraction model based on a separable convolution and multi-head self-attention (MSA) mechanism. It first extracts features using separable convolution, then processes the input sequence through MSA and a feedforward neural network (FFN). Here, i represents the sequence position index. "LayerNorm" represents layer normalization, which standardizes the input features along the feature dimension to stabilize the training process. The formula for each Caformer encoder layer can be expressed as: (1), in, This represents separable convolution, where the features extracted by separable convolution are input into a multi-head self-attention mechanism to calculate the relationship between each element and other elements in the sequence. The MSA formula can be further expanded as follows: (2), in, It is a matrix of queries, keys, and values. , The self-attention mechanism can be formally expressed as: (3), The weight matrix is ​​a learnable matrix. The output of the MSA is residually connected to the input, and then normalized again. The output after the residual connection and normalization is then further transformed nonlinearly through a feedforward neural network, which can be expressed as: (4), in These are learnable weights and biases. The output of the encoder is added to the result of the residual concatenation, and the encoder output is mapped to the class space using a classification head to obtain the final feature extraction result. Feature extraction is performed independently for each image, where each image is initially processed by a Caformer. Finally, a 512-dimensional feature vector is computed for each image.

[0044] Generally speaking, raw CT image data can be defined as... ,in It is the number of samples. This represents the number of CT images contained in each sample. After feature extraction by the Caformer module, it becomes... .

[0045] The hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing is as follows: Hybrid expert model basic framework: This application constructs a hybrid architecture comprising one shared expert and n routing experts, expressed as formula (5): (5), The shared expert is responsible for processing all basic input features, while the routing expert is responsible for capturing specific types of pathological features (such as centrilobular emphysema, bronchial wall thickening, etc.). The shared expert uses a sparse activation method, and only one shared expert is activated in each epoch. Sixteen shared experts are defined in this paper, and four of them are activated each time. Specifically, this application designs a two-stage gated routing mechanism, specifically expressed as formulas (6)-(8). Phase 1: Shared Expert Screening - All Input Tokens First, common features are extracted using shared experts: (6), in The token feature output by the Transformer encoder.

[0046] Phase Two: Routing Expert Assignment. Based on the output of the shared experts, calculate the probability of routing to each expert. (7), in For the first The load balancing entries for each routing expert (initially 0). The ultimate goal is to achieve expert output fusion, that is, to share the expert's output with the selected expert. The outputs of several routing experts are merged: (8), In the process of routing expert allocation, the most important thing is to achieve load balancing, which is a key step to ensure that the model can have good learning performance and stable convergence. Existing load balancing methods can be expressed by formulas (9)-(11).

[0047] Define the load of the routing expert For the first In the next iteration, it is allocated to the... The sum of the token weights of each expert: (9), in This is the batch size. Next, the average load is calculated and the load balancing entries are updated to ensure load balancing across all routing experts: (10) (11), Historical information review module: Existing load balancing update methods only consider model parameters in the current epoch and cannot utilize training information from previous epochs. This may lead to model instability in the early stages of training, thus affecting the final convergence result. Therefore, to further improve model performance, a historical information backtracking module based on a sliding window mechanism is proposed. A sliding window mechanism is introduced to integrate load data from the past three iterations to smooth out fluctuations. and Updated to formulas (12)-(13): (12) (13) In terms of code implementation, a cell is defined. This module will save the load information from the past three epochs and keep it updated on a rolling basis. If there are fewer than three epochs, the module will not be activated. The formula remains unchanged at this point.

[0048] in This is an adjustment coefficient used to control the bias update rate.

[0049] Softsign-based regularized load balancing adjustment: In actual experiments, due to differences in the distribution of data from different centers, The computational results are not very stable, which may lead to significant fluctuations in the model during multi-center training. Furthermore, this may also reduce the model's generalization ability. Therefore, a method based on... The regularization term of the function controls the output of load balancing. In actual experiments, it can help the model converge more stably and improve model performance.

[0050] The formula for the function can be expressed as formula (14), and its output range is between -1 and 1. It is a variant of the hyperbolic tangent function, relative to the tanh function. The function's curve is smoother, therefore it is less prone to the gradient vanishing problem. (14) in, These represent the model's raw calculated values ​​for load balancing. These values ​​vary considerably, and without constraints, it's possible that some routing experts will be perpetually selected. The function compresses it to ( For the interval 1,1), perform a bounded, zero-centered, differentiable regularization.

[0051] In addition, a regularization parameter was added to better control the output range of the load. .final, The formula is expressed as shown in (15): (15) This concludes the definition of the hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing. For ease of expression, this module is defined as follows: Even after the data has undergone weighted calculations in this module, the output size remains the same. That is, expressed as formula (16): (16) LSTM-based classification decision module: Finally, a bidirectional LSTM feature extraction module is defined for classification computation, and a two-layer bidirectional LSTM network is deployed to process sequence features. The hidden state dimension of each LSTM layer is [dimension missing]. The calculation process is as follows: (17)-(19) (17) (18) (19) in , for the first Input features at each time step, and , respectively, are the hidden states of the forward and backward LSTMs. The features are fused. Attention aggregation is performed on the hidden states of all time steps in the LSTM output to capture features of key lesion regions: (20) (twenty one), in For the first Attention weights at each time step The aggregated global features are then mapped to a three-class classification space using a fully connected layer, and... Function generation probability distribution: (twenty two), in This is the classification probability vector, corresponding to the predicted probabilities of three categories: normal person, PRISM, and COPD.

[0052] The WH-ESPC model effectively addresses the challenge of feature extraction in multi-sequence CT by employing a two-level gated routing approach (including shared experts and routing experts). Shared experts extract basic features, while routing experts model features specific to particular lesion types, achieving a good balance between computational efficiency and feature representation capabilities. Furthermore, two innovative modules are introduced to improve existing MoE (Expert Hybrid) models. The first module is a history memory module, which fuses expert load data from the previous three iterations through a sliding window, smoothing fluctuations during training and making the model's convergence more stable on multi-center data. Ablation experiments show that removing this module leads to an approximately 8% decrease in the F1 score for the PRISm category. The second module is a regularized load balancing module based on the Softsign function. This module maps expert load rate bias to the interval [-1, 1], addressing the instability caused by the variance of different center data distributions, thereby improving the model's generalization ability on the external validation set and resulting in an approximately 4% performance improvement. Visualization analysis results show that in CT images (such as the right lower lobe tower region), the model's attention weight for bronchial wall thickening areas is significantly different from the baseline model, further validating its ability to identify and monitor lesions.

[0053] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A voice-interactive intelligent screening system for COPD, characterized in that, include: This includes the user end, cloud-based risk assessment and data management platform, and collaboration end; The user terminal is used to obtain user voice responses, COPD screening questionnaires and symptom scale data, and test data. The cloud-based risk assessment and data management platform is used to calculate risk scores and generate screening and grading results based on data received from user terminals. The collaborative terminal is used to view the screening records, classification results, and high-risk population list of the cloud-based risk assessment and data management platform, and supports result correction and follow-up management. The user terminal includes a voice interaction module, a questionnaire and scale collection module, and an auxiliary test collection module; The voice interaction module is used to play screening questions by voice and receive user voice responses, and convert the voice responses into response text; The questionnaire and scale collection module is used to collect data according to the preset COPD screening questionnaire and symptom scale. The auxiliary test acquisition module is used to acquire simplified breathing test and walking test data; The cloud-based risk assessment and data management platform includes a speech recognition and semantic parsing module and a joint risk classification engine; The speech recognition and semantic parsing module is used to recognize, clean, and structure the response text from the user terminal. The joint risk grading engine is used to perform fusion analysis on the response text processed by the speech recognition and semantic parsing modules and auxiliary test data, automatically calculate risk scores and generate screening and grading results.

2. The intelligent COPD screening system based on voice interaction according to claim 1, characterized in that, The user terminal also includes an offline caching module, which is used to store screening records and voice data in a network-free environment and automatically encrypt and synchronize them to the cloud after connecting to the network.

3. The intelligent COPD screening system based on voice interaction according to claim 1, characterized in that, The voice interaction module is also used to send the converted response text to the user and generate a confirmation response text based on the user's instructions; The voice interaction module includes dialect recognition, and the dialect recognition module includes: The segment length is dynamically calculated by calling the Fourier spectrum continuity parameter and the zero crossover rate, and long speech is segmented into dialect segments; High-pass filtering and wavelet shift denoising are applied to the segmented dialect data to remove environmental background noise. Feature vectors of dialect speech are extracted through encoder and attention mechanism. Based on the acquired geographical location, the target dialect type is matched in the preset dialect lexicon based on the feature vector of the dialect speech. Based on the matched target dialect type, the corresponding dialect conversion model is used to perform cascade fusion of voiceprint vector features and phonological features to obtain structured text information. The structured text information is then converted into specific scores of the COPD scale according to preset regulations.

4. The intelligent COPD screening system based on voice interaction according to claim 1, characterized in that, The voice interaction module supports languages ​​including Mandarin and dialects, and achieves age-friendly interaction through voice broadcasting, large fonts, and large button interfaces.

5. The intelligent COPD screening system based on voice interaction according to claim 1, characterized in that, The cloud-based risk assessment and data management platform also includes a questionnaire and scale management module and a data storage and statistical analysis module; The questionnaire and scale management module is used to manage questionnaires and scales. The data storage and statistical analysis module connects the user to the joint risk classification engine, stores the data from the joint risk classification engine, and performs statistical analysis on the data from the joint risk classification engine.

6. The intelligent COPD screening system based on voice interaction according to claim 1, characterized in that, The joint risk grading engine includes a pre-trained WH-ESPC model, which is used to process the acquired CT image data to obtain screening results. The WH-ESPC model includes a feature extraction module based on the CaFormer architecture, a multi-sequence CT feature extraction module based on historical memory and load balancing hybrid expert network, and a classification decision module based on bidirectional LSTM. The feature extraction module based on the CaFormer architecture is used to extract abstract features from a single CT image to obtain a multi-dimensional feature vector. The hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing is used to adaptively extract key lesion features from multi-sequence CT images. The bidirectional LSTM-based classification decision module is used to aggregate sequence features based on an attention mechanism and output screening results.

7. The intelligent COPD screening system based on voice interaction according to claim 6, characterized in that, The hybrid expert network multi-sequence CT feature extraction module based on historical memory and load balancing includes a hybrid expert model basic framework, a historical information backtracking module, and a regularized load balancing adjustment. The hybrid expert model infrastructure includes shared experts and routing experts; The data processing flow of the hybrid expert model infrastructure includes: Basic features are extracted from the input features by sharing experts; Based on a two-stage gated routing mechanism, the probability of routing to each routing expert is calculated according to the extracted basic features; Select k routing experts based on probability results; The feature representation is obtained by fusing the outputs of the shared expert with the outputs of the selected k routing experts; The formula for calculating the probability of routing to each routing expert is as follows: , in, Let be the probability of being assigned to the k-th routing expert. For the first A load balancing solution from a routing expert. Based on the feature vector, These are the weight parameters.

8. The intelligent COPD screening system based on voice interaction according to claim 7, characterized in that, In the dual-stage gated routing mechanism, the load balancing item of the routing expert is updated based on the historical information backtracking module. The historical information backtracking module integrates the load data of historical iterations through a sliding window mechanism to smooth training fluctuations.

9. The intelligent COPD screening system based on voice interaction according to claim 8, characterized in that, The update formula for the load balancing item is: , in, and These are the load balancing bias terms for the k-th routing expert in the t-th and t+1-th iterations, respectively. Let the total workload allocated to the k-th routing expert in the t-th iteration be _____. Let the average load of the routing expert be the value in the t-th iteration. is the regularization coefficient.

10. The intelligent COPD screening system based on voice interaction according to claim 6, characterized in that, The classification decision module based on bidirectional LSTM outputs screening results, including: The input sequence features are processed by a two-layer bidirectional LSTM network to obtain the forward and backward hidden states at each time step; Attention-weighted aggregation of the hidden states at all time steps is performed to generate a global feature representation; The global features are mapped to the classification space through a fully connected layer and a softmax function, and a probability distribution is generated.