An artificial intelligence-based guidance and pre-diagnosis system

By employing asynchronous and parallel information acquisition and structured processing methods, combined with acoustic feature analysis and adaptive interaction, the problems of disrupted narrative continuity and omission of key information in existing systems have been solved, achieving efficient information recording and structured processing.

CN121031601BActive Publication Date: 2026-01-27GUANGDONG HAUCI NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511586957.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-27
Estimated Expiration
2045-11-03

AI Technical Summary

Technical Problem

Existing AI-based triage and pre-diagnosis systems can disrupt the integrity of users' original narrative context during information collection and structured processing. Furthermore, the rigid interaction process restricts users' autonomous statements and may result in the omission of key information.

Method used

The system employs a narrative flow-based continuous acquisition module, a background entity and relation pre-annotation module, and a non-invasive summary confirmation module to achieve asynchronous parallel processing of information acquisition and structured processing in time. It assigns semantic weights to clinical named entities through acoustic feature analysis and generates a summary for confirmation when the user pauses. It also uses semantic entropy calculation to enable adaptive switching of interaction modes.

Benefits of technology

While ensuring the continuity of the narrative flow, the system fully records user information and extracts key clinical points, improving the efficiency and accuracy of information structuring. The adaptive interaction mode switching enhances the effectiveness of information collection in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121031601B_ABST
    Figure CN121031601B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical care information science and discloses a guide and pre-consultation system based on artificial intelligence, which comprises a narrative flow uninterrupted collection module, a background entity and relation pre-annotation module and a non-invasive abstract confirmation module. The system receives user voice narratives uninterruptedly, synchronously extracts acoustic features, ascertains the semantic weight of the entity based on the background asynchronous identification of clinical named entities, generates an abstract based on the entity with the weight at the narrative pause point for user confirmation, and establishes a new working mode of asynchronous parallel information collection and structured processing. The destructive interrogation is replaced by daemon listening, so that the integrity and fidelity of the original medical information are ensured, and the probability of early clinical risk discovery is improved through quantitative analysis of the narrative rhythm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an artificial intelligence-based triage and pre-diagnosis system, belonging to the field of healthcare informatics technology. Background Technology

[0002] Currently, in AI-based triage and pre-diagnosis systems, a common technical practice is to adopt a dialogue process based on decision trees or serialized state machines. This process guides users through a series of pre-set questions and performs real-time parsing and structured filling of the user's answers. This approach has engineering value in improving the efficiency of subsequent medical information processing. However, when this approach is applied to scenarios with complex conditions or where the user's statements deviate from the pre-set consultation path, its mechanism of simultaneous information collection and real-time structured processing can lead to the disruption of the integrity of the user's original narrative context. The user's coherent statements are repeatedly interrupted and forcibly incorporated into pre-set data fields during this process. Information that may have clinical diagnostic value but is not touched by the pre-set questions is at risk of being overlooked at the source of information collection.

[0003] Besides the aforementioned limitations in interactive process design, existing technologies in other areas also reflect a tendency to deviate from the core objective of improving the quality of clinical information collection. For example, Chinese invention patent CN118238163B discloses an artificial intelligence-based pre-consultation system and method, which aims to solve the physical stability problem of intelligent consultation equipment that may become unbalanced during hydraulic lifting due to uneven weight distribution between the host and storage cabinet. This solution deploys a hardware detection device that includes lifting, vibration, horizontal, and tilt states to achieve step-by-step monitoring and judgment of physical faults in the equipment. Although the patent is named an artificial intelligence pre-consultation system, its technical contribution is entirely focused on the electromechanical structural safety and fault diagnosis of the equipment. It does not provide any technical inspiration for the fundamental task of the pre-consultation system: how to optimize the dialogue and interaction between doctors and patients and ensure the integrity and fidelity of clinical information collection. This precisely highlights the current technical bottleneck in this field: even when artificial intelligence technology is applied, its focus may be limited to peripheral hardware security rather than deeply solving the core problem of clinical information interaction.

[0004] Specifically, the root cause of this problem lies not in the performance limitations of natural language processing algorithms, but in the design of the information collection process itself. The synchronous coupling of information collection and information structuring in time means that the real-time processing requirements of information structuring inevitably sacrifice the continuity of the original narrative flow, creating a technical constraint that is difficult to avoid under current technological approaches. This constraint manifests in two ways: 1. The information collection process structurally interferes with the user's original narrative context, affecting the fidelity of the original information; 2. The rigid interrogation-style interaction process restricts the user's autonomous description of atypical symptoms, potentially leading to the omission of crucial information. Therefore, how to design an information collection and processing method that can fully preserve the context of the user's free narrative during information collection while simultaneously achieving efficient structured processing of clinical information, avoiding the mutual constraint between the two processing objectives due to limitations in process design, becomes the technical problem this invention aims to solve. Summary of the Invention

[0005] This invention provides an AI-based triage and pre-diagnosis system. Its main purpose is to solve the problem that the existing technology, which involves simultaneous information collection and information structuring, leads to the destruction of the integrity of the user's original narrative context and creates a technical constraint between information integrity and structuring efficiency.

[0006] To achieve the above objectives, this invention provides an artificial intelligence-based triage and pre-diagnosis system, which includes a narrative flow continuous acquisition module, a background entity and relation pre-annotation module, and a non-intrusive summary confirmation module.

[0007] The narrative stream uninterrupted acquisition module is configured to receive free-formatted voice narrative input from users and continuously perform two synchronous operations: transcribing the content of the voice narrative input into a timestamped original narrative track, and extracting acoustic features from the voice narrative input to generate an acoustic prosodic track that is time-aligned with the original narrative track;

[0008] The background entity and relation pre-annotation module is configured to establish a user-personalized acoustic baseline based on the acoustic features of non-diagnostic statements at the beginning of the dialogue. While the narrative flow uninterrupted acquisition module is working, it processes the original narrative track in parallel and asynchronously to identify clinical named entities and generate a temporary structured information sketch. Furthermore, it is configured to compare the acoustic features in the acoustic prosodic track that are time-aligned with the clinical named entity with the acoustic baseline when a clinical named entity is identified to calculate a deviation value, and assign the deviation value as a semantic weight to the clinical named entity.

[0009] The non-invasive summary confirmation module is configured to generate a natural language summary based on clinical named entities and semantic weights in the structured information sketch when a natural pause point defined by the user's voice pause duration exceeding a preset duration threshold is detected in the original narrative track. Only after receiving the user's confirmation instruction will the confirmed clinical named entities in the structured information sketch be written into the final structured pre-diagnosis report.

[0010] Preferably, the non-invasive summary verification module is configured to prioritize the integration of clinical named entities with high semantic weights when generating natural language summaries, and present them to users in an interactive language that explores clinical focus, so as to guide users to make supplementary statements around the clinical named entities.

[0011] Preferably, the system further includes a semantic entropy calculation unit and an interaction mode switching unit; the semantic entropy calculation unit is configured to calculate a semantic entropy value in real time based on the confidence of clinical named entities in a structured information sketch within a sliding time window and the number of contradictory relationships between clinical named entities; the interaction mode switching unit is configured to automatically switch the interaction mode of the narrative flow uninterrupted acquisition module from a listening mode that mainly receives user-defined voice narrative input to a guided mode that provides users with focused options to clarify core ambiguities when the semantic entropy value exceeds an entropy threshold.

[0012] Preferably, the semantic entropy calculation unit is configured to calculate the semantic entropy value in real time. ,in ;in, This represents the average confidence level of all clinical named entities in the structured information sketch within the sliding time window. This represents the number of conflicting relationships between clinical named entities in a structured information sketch within a sliding time window. and The preset weighting coefficients; the interaction mode switching unit, configured to be used when... When the entropy threshold is exceeded, the interaction mode is switched.

[0013] Preferably, the narrative stream uninterrupted acquisition module is also configured to distinguish the extraction of acoustic features into processing a first frequency band containing human voice and processing a second frequency band not containing human voice; the background entity and relation pre-labeling module is also configured to compensate the acoustic features of the first frequency band based on the acoustic features of the second frequency band before assigning semantic weights to clinical named entities.

[0014] Preferably, the background entity and relation pre-labeling module is configured to perform compensation in the following way: subtract a compensation amount proportional to the energy value of the second frequency band from the energy value of the first frequency band to obtain a compensated energy value, and assign semantic weights to clinical named entities based on the compensated energy value.

[0015] Preferably, the narrative flow uninterrupted acquisition module is configured to prioritize open-ended and unguided interaction strategies to encourage users to make continuous, free statements.

[0016] Preferably, the non-invasive summary confirmation module is configured to highlight or prioritize clinical named entities that are pre-determined to be related to high-risk diseases based on the medical knowledge base and have been marked with high semantic weight when generating natural language summaries.

[0017] Preferably, the system also includes a risk escalation module; the risk escalation module is configured to activate when the information confirmed by the user contains a predefined combination of high-risk symptoms consisting of multiple clinical named entities, so as to send an alarm message to the terminal device of the human agent.

[0018] Preferably, the background entity and relation pre-labeling module is also connected to an external electronic medical record database; the background entity and relation pre-labeling module is also configured to perform weighted adjustment of the confidence of clinical named entities identified in the structured information sketch based on the user's historical medical record information obtained from the electronic medical record database.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] 1. By setting up a continuous narrative flow acquisition module for uninterrupted reception of user-formatted narrative input, a background entity and relation pre-annotation module for parallel processing of narrative input while the acquisition module is working, and a non-invasive summary confirmation module for generating summaries based on the processing results of the background pre-annotation module at natural pauses in the narrative input and presenting them to the user for confirmation, a working mode that separates the information acquisition process from the information structuring process in time is established. In this mode, the two mutually constraining goals in the traditional consultation process—the complete recording of the user's narrative context and the structured extraction of clinical information from it—are decomposed into two asynchronous and non-interfering processes. The structured processing results in the background are used to support the phased summary in the front end, and the information confirmed by the front end is then formally written into the final report. This process ensures that the information structuring process does not come at the cost of disrupting the continuity of the original narrative flow, avoiding the risk of losing key clinical details due to premature intervention in structured question answering.

[0021] 2. Based on the existing working method, the system synchronously extracts acoustic features from user narrative input by configuring a continuous narrative flow acquisition module. A background entity and relation pre-labeling module is also configured to assign semantic weights to identified clinical named entities based on these acoustic features, thus linking the user's acoustic prosodic information with clinical semantic information on the same timeline. When the background module identifies a clinical entity, it simultaneously examines the deviation between the user's acoustic features describing that entity and their personalized acoustic baseline within the same time period, converting this deviation into the entity's semantic weight. This processing method allows the system to record not only the content of the user's statement but also the state of the statement itself, enabling the final generated structured information sketch to transcend plain text records and instead carry key, clinically significant hierarchical information.

[0022] 3. By adding a unit to the backend entity and relationship pre-annotation module for real-time calculation of the semantic entropy of temporary structured information sketches, and using the calculated entropy value as the basis for switching system interaction modes, an adaptive interaction control loop based on the quality of the information flow itself was established. When the structured information sketches generated in the backend exhibit a high entropy state due to low confidence of clinical entities or contradictory relationships between entities, the system will determine that the current user's narrative is confused and automatically switch the frontend interaction from a listening mode that primarily receives free narratives to a guiding mode that provides focused options to clarify core ambiguities, until the semantic entropy of the sketches falls below a preset threshold before switching back to the listening mode. This mechanism enables the system to autonomously switch between the roles of protective listening and structured guidance when facing users with different narrative qualities. This invention maintains the effectiveness of information collection in complex interactive scenarios. Furthermore, it differentiates the extraction of acoustic features into processing a first channel containing human voice frequencies and a second channel primarily consisting of non-human voice frequencies. Before assigning semantic weights to clinical named entities, the acoustic features of the second channel are used to compensate for the acoustic features of the first channel. This compensation mechanism leverages the characteristic that environmental noise typically affects both channels simultaneously. By subtracting a compensation amount proportional to the second channel signal from the first channel signal, it suppresses the interference of sudden environmental noise on acoustic feature analysis. This design enables the judgment of user narrative prosody to effectively distinguish between changes in the user's own state and interference from the external environment, providing a more reliable signal basis for subsequent semantic weight assignment based on acoustic features. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the overall functional architecture and core data flow of the system of the present invention;

[0024] Figure 2 This is a schematic diagram illustrating the semantic entropy value changes and mode switching threshold triggering under different dialogue qualities according to the present invention;

[0025] Figure 3 This is the timing diagram for the adaptive interaction mode switching closed-loop control based on semantic entropy of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] The present invention discloses an AI-based triage and pre-diagnosis system. Its system architecture includes a continuous narrative stream acquisition module, a background entity and relation pre-annotation module, and a non-intrusive summary confirmation module. The system employs an asynchronous parallel approach to information acquisition and structured processing. The continuous narrative stream acquisition module performs complete and uninterrupted reception and recording of the user's free-formatted voice narrative. Simultaneously, the background entity and relation pre-annotation module analyzes and performs structured preprocessing on the recorded narrative data stream in parallel and asynchronously. Finally, the non-intrusive summary confirmation module analyzes and preprocesses the user's free-formatted voice narrative. The natural pause point of the narrative is used to generate a summary based on the processing results of the backend module for user confirmation, thereby achieving structured processing of medical information while ensuring the integrity of the original narrative context. In current healthcare informatics applications, the interrogation-style interaction adopted by automated consultation systems to improve the efficiency of information structuring often leads to the interruption of the user's coherent description of their condition, resulting in the destruction or loss of narrative context information valuable for diagnosis. To solve this problem, the narrative flow uninterrupted acquisition module of this invention is configured as a front-end application deployed on the user's terminal device in specific engineering implementations. This program calls... The device's microphone interface continuously receives users' free-format voice input. Its internal processor is configured to perform two simultaneous operations: first, using a speech-to-text engine to transcribe the received voice input into a timestamped text stream in real time, and recording this stream completely to form an original narrative track; second, while transcribing the voice, a parallel acoustic feature extraction submodule processes the same original audio stream to extract objective acoustic feature parameters such as speech rate, pause frequency, and volume energy envelope, and combines these parameters with synchronized timestamps to form a track consistent with the original narrative. The acoustic rhythmic tracks are parallel to the orbital tracks. It should be noted that, to improve the reliability of acoustic analysis in practical application environments, the extraction of acoustic features is divided into processing a first frequency band containing human voices (e.g., 50Hz to 8kHz) and processing a second high-frequency band not containing human voices (e.g., 10kHz to 16kHz), with the latter data used as a reference for ambient noise. In terms of interaction strategy, this module is limited to prioritizing open and unguided speech to encourage users to make continuous free statements. Through the above configuration, this module ensures the integrity and fidelity of the original medical information flow at the source of information acquisition.

[0028] Given that the acquired raw narrative stream is unstructured and its simple text content cannot reflect the emphasis of clinical key points, the system employs a background entity and relation pre-annotation module for asynchronous processing. This module, running as a parallel background process, processes the text data stream in the raw narrative track while the narrative stream acquisition module continues operating. It invokes a medical-oriented natural language processing service to identify clinical named entities from the text and preliminarily infer the relationships between entities, thereby dynamically generating a temporary structured information sketch. A key technical aspect is that this module is configured to calculate and establish a user-personalized acoustic baseline at the beginning of the dialogue, based on the acoustic features corresponding to non-diagnostic statements in the acoustic prosodic track. Subsequently, when the module... arrive When a clinical named entity is identified, it queries the acoustic feature parameters of the acoustic prosodic track within the same time period and first performs noise compensation. This involves subtracting a compensation amount proportional to the energy value of the second frequency band from the energy value of the first frequency band to obtain a compensated energy value. Subsequently, the system compares this compensated acoustic feature with a pre-established acoustic baseline, calculates a deviation value, and assigns this deviation value as a semantic weight to the clinical named entity. Specifically, the entity and relation pre-labeling module uses the compensated energy value when performing noise compensation. Through formula The calculation shows that, and The energy values ​​and compensation coefficients for the first and second frequency bands within the same time window are respectively. Before system deployment, linear regression analysis is performed on background audio samples collected in the target application scenario to calculate the statistical correlation coefficient between the energy values ​​of the first and second frequency bands. Subsequently, the standardized deviation values ​​of the acoustic features are then analyzed. Convert to semantic weights When using a linear mapping relationship Calculations are performed, in which, The value is the difference between the compensated acoustic feature value and the mean of the user's personalized acoustic baseline, divided by the standard deviation of that baseline. This is a preset gain coefficient, whose function is to... The distribution range is linearly scaled to a preset weight range that facilitates priority judgment by the subsequent summary generation module. For example, if the standard deviation of the compensated energy envelope of the time period corresponding to the clinical named entity chest tightness exceeds 1.8 times the corresponding value of the user's acoustic baseline, the system assigns a high semantic weight value to the chest tightness entity. In this way, by associating acoustic prosodic information with clinical semantic information on the time axis, the generated structured information sketch not only records the content of the statement, but also carries the hierarchical information of clinical focus in a quantitative way.

[0029] After the system obtains a weighted structured information sketch through background processing, a non-intrusive summary confirmation module is set up to complete information confirmation without interfering with the user and to address the risk of interaction failure due to user narrative confusion. This module continuously monitors the original narrative trajectory. When it detects a natural pause point defined by the user's voice pause duration exceeding a preset duration threshold (e.g., 3 seconds), the module is activated. After activation, it generates a natural language summary based on the clinical named entities and semantic weights in the structured information sketch. During summary generation, it is configured to prioritize the integration of clinical named entities marked with high semantic weights and present them to the user with interactive language that explores clinical key points to guide the user to make supplementary statements. Only after receiving the user's confirmation instruction are the confirmed clinical named entities written into the final structured pre-diagnosis report. To further improve the system's effectiveness in complex interaction scenarios, the system also includes a semantic entropy calculation unit and an interaction mode switching unit. The semantic entropy calculation unit is configured to calculate a semantic entropy value in real time within a sliding time window based on the average confidence of clinical named entities in the structured information sketch and the number of contradictory relationships between entities. Its calculation follows the formula In this formula, the number of contradictory relationships between clinical named entities within the sliding time window is... The statistics are compiled using a standardized decision-making procedure executed by a backend entity and relation pre-labeling module. This procedure employs three parallel decision-making methods: First, based on a built-in medical antonym library, when two entities or their modifiers identifying antonyms for the same symptom location within the same sliding window (e.g., headache and no headache), the number of contradictory relationships is incremented by one. Second, based on a medical knowledge graph, when two identified clinical named entities are predefined as mutually exclusive in the graph (e.g., pregnant and normal menstrual cycle), the number of contradictory relationships is incremented by one. Third, for clinical named entities containing quantitative values, when their values ​​conflict with the qualitative description of another entity (e.g., body temperature 39.1°C), the number of contradictory relationships is incremented by one. If there is no fever, the number of contradictory relationships increases by one; among them, This represents the average confidence level of all clinically named entities within the sliding time window. This represents the number of conflicting relationships between clinical named entities within a sliding time window. and The preset weighting coefficients are used; the interaction mode switching unit is configured to switch when the calculated weighting coefficients are used. When the preset entropy threshold is exceeded, the interactive mode of the narrative flow uninterrupted acquisition module is automatically switched from the listening mode, which mainly receives the user's free-format voice narrative input, to the guidance mode, which provides the user with focused options to clarify core ambiguities, until the semantic entropy value falls back below the threshold and then switches back to the listening mode. Through this process and adaptive control mechanism, the information structuring process is completed without disrupting the continuity of the narrative flow, and the system has the ability to maintain the effectiveness of information acquisition in different interactive environments.

[0030] To enable immediate response to identified potentially high-risk clinical information, the system also includes a risk escalation module. This module is functionally configured to continuously monitor clinical named entities (NMOs) confirmed by a non-invasive summary verification module and written into the final structured pre-consultation report, and compare them in real-time with an internally pre-defined risk knowledge base consisting of a series of red flag symptom combinations. When an NMO is detected containing one or more events that perfectly match predefined high-risk symptom combinations in the risk knowledge base—for example, when chest pain and radiating pain in the left shoulder are simultaneously confirmed in the same report—the risk escalation module is activated. Upon activation, the module performs two parallel operations: first, it generates an alert message and immediately pushes it to the pre-consultation report via a secure network interface. First, the configured operator's terminal device contains the user's de-identified ID, the name of the triggered high-risk symptom combination, and a unique data index pointing to the user's complete original narrative trajectory and structured information sketch. The clinical named entities that triggered the alarm are highlighted or prioritized on the operator's terminal device interface. Second, a command is sent to the non-invasive summary confirmation module. This command causes the non-invasive summary confirmation module to highlight or prioritize clinical named entities identified as being related to high-risk diseases in the next or currently being generated natural language summary. For example, this could be achieved by displaying the word "chest pain" in a different colored font in the text summary presented to the user, or by emphasizing it in the audio summary through added stress or pre-voice prompts.

[0031] Example 1: This example demonstrates the operation of a disclosed AI-based triage and pre-diagnosis system in a specific application scenario. In an online medical service platform application, a user initiates a pre-diagnosis process due to persistent fatigue. Initially, their description of symptoms is divergent, including details about work stress and sleep habits alongside the main discomfort. After system startup, the narrative flow continuous acquisition module guides the user's free narration with open-ended statements and, according to the procedures in the specific implementation, simultaneously generates a timestamped original narrative track and an acoustic prosodic track. During the user's continuous narration, the system does not interrupt with questions. Simultaneously, the background entity and relation pre-labeling module processes the continuously input original narrative track asynchronously in parallel, identifying multiple clinical named entities such as fatigue, shortness of breath while climbing stairs, and high work stress, and constructs a temporary structured information sketch based on these entities. During the process, after describing a series of details about work fatigue, the user mentioned that recently they need to elevate their head with two pillows to feel comfortable when sleeping at night. For a system that only analyzes text keywords, this information would be given a low priority due to the lack of clear disease-specificity. The background entity and relation pre-labeling module of this invention, when identifying the clinical named entity "elevating the pillow," detected from the time-aligned acoustic prosodic track that the user's speech rate was significantly slower when describing this detail, and the standard deviation of the energy envelope deviated from the personalized acoustic baseline established at the beginning of the conversation by more than a predetermined multiple. Based on this, the system assigned a high semantic weight to the clinical named entity "elevating the pillow." This process is the result of the synergistic effect of continuous narrative flow acquisition and background acoustic feature analysis. That is, the former provides a recording basis for the emergence of low-signal-intensity clinical details, while the latter provides an objective quantitative basis for identifying potential clinical priorities from these details.

[0032] When a user's narrative reaches a natural pause exceeding a preset time threshold, the non-invasive summary confirmation module is activated. This module generates a natural language summary based on a pre-generated structured information sketch with semantic weights, and focuses the content according to these weights. The message presented to the user is: "Okay, let me confirm. You mentioned feeling fatigued lately, especially when climbing stairs. Also, you mentioned that using a higher pillow at night is more comfortable. Could you elaborate on this?" The user's attention is guided to this clinical detail, which the system has assigned a high weight to, and they provide supplementary information. The confirmed information is then written into the structured pre-diagnosis report. This asynchronous parallel approach to information collection and structured processing resolves the conflict between information integrity and structuring efficiency. Through quantitative analysis of acoustic prosody, the system can identify potential clinical risk signals from a continuous narrative flow.

[0033] To further verify the decisive technical effect of the present invention in capturing key weak clinical signal information by assigning semantic weights to clinical named entities based on acoustic features, the following comparative example 1 is set up.

[0034] Comparative Example 1: The system used in this comparative example has the same hardware environment, software architecture, and all modules except the background entity and relation pre-annotation module as the system in the aforementioned embodiments. The only difference is that the background entity and relation pre-annotation module in this comparative example is configured to perform only text-based clinical named entity recognition, without performing semantic weight labeling based on acoustic features. All identified clinical named entities are assigned a default, equal priority. Using the system in this comparative example, free-form speech narratives input by the same user, identical to those in the aforementioned embodiments, are processed. During the processing, the system's narrative stream uninterrupted acquisition module normally receives user input. The system continuously inputs the original narrative track and generates an acoustic prosodic track. The background entity and relation pre-annotation module processes the continuously input original narrative track in parallel asynchronously. It accurately identifies clinical named entities such as fatigue, shortness of breath when climbing stairs, high work stress, and raising the pillow, and constructs a temporary structured information sketch. However, due to the lack of acoustic feature analysis function, the module cannot perceive the acoustic changes such as the user's speech rate slowing down and the standard deviation of the energy envelope increasing when stating the detail of raising the pillow. Therefore, the clinical named entity of raising the pillow is assigned the same default weight value (e.g., priority set to medium) in the structured information sketch, along with other entities such as high work stress.

[0035] When a user's narrative reaches a natural pause exceeding a preset duration threshold, the non-intrusive summary confirmation module is activated. Based on a structured information sketch generated in the background without differentiated semantic weights, this module uses its built-in default summary generation strategy, which integrates content according to the order of entity appearance, to generate the following natural language summary, which is then presented to the user: Okay, let me confirm. You mentioned that you've been feeling fatigued lately, and you get short of breath when climbing stairs, and you also feel a lot of work pressure. Is there anything else you'd like to add regarding these situations? The results showed that, because the key weak signal clinical detail of pillow elevation was not prioritized and explored in the summary, the user's attention was not drawn to this potentially high-risk information, and the subsequent supplementary statements also failed to focus on this core issue. Although the final structured pre-diagnosis report recorded the entity of pillow elevation, it was not marked as a high priority, resulting in the system's failure to effectively detect potential early risk signals related to heart failure. The experimental results of this comparative example confirm that, under the premise of using the same asynchronous information flow architecture as this invention, the lack of a semantic weighting mechanism based on acoustic features will cause the system to be unable to effectively distinguish and focus on key weak clinical signals from free narratives containing a large amount of background information, thereby reducing the probability of early clinical risk detection.

[0036] Example 2: To objectively verify the technical effectiveness of the AI-based triage and pre-diagnosis system of the present invention in terms of information collection integrity and interaction robustness, the following comparative experiment was set up. The experimental platform was based on a server environment, which deployed the system using the technical solution of the present invention (the sample group of the present invention), and two systems for comparison. The experimental data came from a dataset containing 500 anonymized pre-recorded audio samples. This dataset simulated real doctor-patient dialogues, including 100 cases with key weak signal information pre-annotated by clinical experts, and 50 cases with highly chaotic narrative content. Two control groups were set up in the experiment. Control group 1 adopted the same asynchronous information flow architecture as the sample group of the present invention, but its background entity and relation pre-annotation module was configured not to perform semantic weight labeling based on acoustic features. Control group 2 adopted a synchronous interrogation and parsing dialogue system based on decision trees.

[0037] In the first phase of the experiment, 100 audio recordings containing key weak signal information from the dataset were input into the experimental group and two control groups for processing. The key clinical information capture rate was used as the evaluation index, defined as the percentage of cases with pre-annotated key clinical information recorded and confirmed in the final structured pre-diagnosis report generated by the system out of the total number of test cases. During the operation of the experimental group, for key weak signal clinical named entities such as raising the pillow and waking up gasping for air at night, the system assigned high semantic weights to them based on features such as slowed speech rate and energy envelope standard deviation exceeding 1.5 times the acoustic baseline in the corresponding acoustic prosodic track, and prioritized integration and exploratory questioning in the non-invasive summary confirmation stage. Control group 1 failed to distinguish and process these weak signal information because it did not use the weighting mechanism. Control group 2 did not provide users with the opportunity to state information other than such preset questions because it adopted a fixed closed questioning process. The statistical data of the experimental results are shown in Table 1.

[0038] Table 1: Comparison of capture rates of key weak signal clinical information by different systems.

[0039]

[0040] In the second phase of the experiment, audio inputs from 50 high-confusion cases were provided to three separate systems for processing. The effective interaction completion rate was used as the evaluation metric, defined as the percentage of cases in which the system could generate a structured report that was confirmed by the user and whose content was logically consistent. In the processing flow of the sample group of this invention, the semantic entropy calculation unit detected that the average confidence level of clinical named entities in the structured information draft was consistently below 0.6, and the number of contradictory relationships exceeded three, resulting in a lower calculated semantic entropy value. The system exceeded a preset entropy threshold of 0.85, automatically triggering an interaction mode switch from listening mode to a guided mode providing focused options. Control groups 1 and 2, lacking this adaptive switching mechanism, repeatedly performed ineffective information collection or looped within the same level of the decision tree when processing chaotic narratives, making it difficult to generate effective final reports. Their effective interaction completion rates were 12% and 6%, respectively, while the effective interaction completion rate of the sample group of this invention was 92%. Experimental data shows that the technical solution using asynchronous information flow combined with acoustic features for semantic weighting has a higher capture rate for key clinical information not actively stated in users' free narratives than asynchronous systems without acoustic feature analysis and traditional synchronous interrogation systems. Simultaneously, the semantic entropy-based adaptive interaction mode switching mechanism enables the system to maintain the effectiveness of information collection when processing low-quality, highly chaotic narrative inputs.

[0041] Example 3: This example combines Figures 1 to 3This describes an AI-based triage and pre-diagnosis system, such as... Figure 1 As shown in the diagram, after a user inputs their medical condition and feelings through free-form voice narration, the narrative flow continuous acquisition module synchronously generates a data stream called the original narrative track (text + timestamp) and a data stream called the acoustic prosody track (acoustic features + timestamp). These two tracks are sent in parallel to the background entity and relation pre-labeling module. While asynchronously recognizing clinical entities, this module assigns semantic weights to the entities based on acoustic features, thereby generating a temporary structured information sketch containing semantic weights. This sketch is used by the non-invasive summary confirmation module at the narrative pause point to generate a summary and guide the user to confirm. The structured medical information confirmed by the user is finally written into the final structured pre-diagnosis report. At the same time, a risk escalation module monitors the confirmed high-risk symptom combinations and can send alerts to human agents when necessary. A semantic entropy calculation unit calculates the confidence and contradiction of the information sketch in real time to output semantic entropy. This entropy value is used as a judgment basis by the interaction mode switching unit. When the semantic entropy value exceeds the threshold, it switches to the guidance mode.

[0042] like Figure 2 As shown, the horizontal axis represents time, and the vertical axis represents semantic entropy value. The curve for normal dialogue shows that its semantic entropy value remains at a low level throughout the dialogue, while the curve for chaotic dialogue shows that its semantic entropy value increases significantly at a certain stage of the dialogue and continuously exceeds the switching threshold of 0.85 marked in the legend. This comparison clearly illustrates that the system can objectively distinguish between clear and chaotic states of dialogue based on changes in semantic entropy value, thus providing a reliable decision-making basis for adaptive switching of interaction modes. Figure 3 As shown, the process begins with the user providing a disorganized description of their condition. The narrative flow acquisition module transmits the narrative trajectory to the background entity annotation module. After identifying clinical entities with low confidence and multiple contradictory relationships, the background module transmits the draft data to the semantic entropy calculation unit, which then calculates the average confidence level. Quantity of contradictory relationships with statistics Then, the semantic entropy value was calculated using a formula. It also reports that the entropy value exceeds the threshold, and the interaction mode switching unit performs a judgment. After the condition is judged, the system switches to guided mode and drives the narrative flow acquisition module to provide focused options to the user. After the user selects a specific option, the system transmits the clarified information. The background module recalculates the entropy value. If the entropy value decreases, the interaction mode switching unit immediately switches the interaction mode back to listening mode. If the entropy value is still high, the system continues to execute guided mode, thus completing an adaptive interaction adjustment cycle.

[0043] Example 4: This example discloses a series of standardized engineering calibration procedures for determining the key operating parameters required before deployment of the system of the present invention, so as to maintain the stability and effectiveness of information collection and processing when facing different users and interaction environments; in order to calibrate the natural pause duration threshold used by the non-intrusive summary confirmation module, a calibration dataset containing 1000 anonymized pre-diagnosis dialogue audios is used; the first step of the calibration procedure is to process the audio streams of all dialogues in the dataset, automatically detect and record the duration of all speech pauses after the user completes a semantically complete sentence and before starting the next sentence, and then... The duration data is collected; secondly, all recorded duration data are plotted into a frequency distribution histogram, which exhibits a bimodal distribution, where the first peak corresponds to short pauses within a sentence, and the second peak corresponds to long pauses between sentences; thirdly, a duration value is determined at the trough of this bimodal distribution, which is the statistical dividing point between the two types of pauses. In this calibration process, this trough value was determined to be 3.2 seconds. Therefore, the duration threshold is set to 3 seconds as the basis for the system to determine that the user has completed a stage of the narrative and can intervene for summary confirmation; this serves as the acoustic baseline and bias used by the background entity and relation pre-annotation module when calculating semantic weights. The deviation threshold was determined using an acoustic baseline calibration corpus comprised of 50 test subjects of different genders and ages reading neutral, non-medical standard text. The first step of the calibration procedure was to collect and process all audio data from this corpus, and calculate the mean and standard deviation of a series of acoustic features, such as speech rate and the standard deviation of volume energy envelope, across the entire corpus. These statistical values ​​were then set as the initial group acoustic baseline for the system. When the system interacts with a specific user, it first uses this initial group acoustic baseline. During the initial non-diagnostic statement phase of the dialogue, the user's acoustic feature data is collected to further refine the initial group acoustic baseline. The system performs three steps: First, it generates a personalized acoustic baseline for the user. Second, to determine the deviation, the system divides the absolute value of the difference between the acoustic feature value of a clinical named entity within a given time period and the mean of its personalized acoustic baseline by the standard deviation of that baseline, thus obtaining a standardized deviation value. Third, by statistically analyzing the deviation values ​​of all data in the calibration corpus, values ​​outside the 95% confidence interval are defined as high-weight deviations. Calculations show that this value corresponds to a deviation of 1.5 times the standard deviation. Accordingly, the system is configured to assign a high semantic weight to a clinical named entity when its standardized deviation value is greater than 1.5.

[0044] To calibrate the semantic entropy threshold used by the interaction mode switching unit and the weighting coefficients in the semantic entropy calculation formula, an entropy calibration dataset containing 100 narratively coherent dialogues and 100 narratively disordered dialogues, pre-annotated by clinical experts, was used. The first step in the calibration procedure was to determine the weighting coefficients. and The value of is determined by taking values ​​for in a two-dimensional parameter space. and Traverse from 0.1 to 1.0 in increments of 0.1, for each ( , All combinations use formulas. Calculate the average semantic entropy value of the narrative coherence group and the average semantic entropy value of the narrative incoherence group, and select the group that maximizes the difference between the average entropy values ​​of the two groups. , The combination of these factors serves as the optimal coefficients, and through this optimization process, The value of was determined to be 0.7. The value of was determined to be 0.5; in the second step, based on the determined weight coefficient, the semantic entropy values ​​of all 200 dialogues in the entropy calibration dataset were calculated, and the subject operating characteristic curve was plotted. The semantic entropy value corresponding to the coordinate point closest to the upper left corner (0,1) on the curve was selected as the switching threshold. Through this process, the semantic entropy threshold was determined to be 0.85; through the above calibration procedure, the basis for the system to make judgments on mode switching during operation was determined.

[0045] Example 5: This example discloses the pre-configuration and verification procedures performed by the system of the present invention before deployment in a specific medical service scenario to ensure the stable operation of its risk identification and information processing functions. In an application scenario where the system is to be deployed at a hospital emergency triage station, in order to enable the system to accurately perform semantic weight calibration based on acoustic features in a high-noise environment, a set of environmental acoustic calibration procedures were performed before the system was officially put into use. The first step of this procedure is to continuously collect 24 hours of ambient background audio at the deployment location using the system's own microphone array to form an acoustic environment sample library for that specific scenario. The second step is for the system to process the sample library, analyze the statistical correlation between the energy envelopes of the non-human voice frequency band (10kHz to 16kHz) and the human voice frequency band (50Hz to 8kHz), and calculate a compensation coefficient based on this. The coefficient is used in the calculation process of compensating the acoustic features of the first frequency band in the specific implementation; the third step is that the system uses a test audio set containing standard pronunciation and preset acoustic changes (such as stress, slowing down speech) to play back and collect in a compensated environment, and verifies whether its background entity and relation pre-labeling module can identify the preset acoustic changes and label the corresponding high semantic weights. When the recognition accuracy is greater than 95%, the acoustic calibration in this environment is completed.

[0046] To construct and maintain the Red Flag Symptom Combination Knowledge Base upon which the system's built-in Risk Upgrade Module relies, the system implements a standardized knowledge base management procedure. The first step of this procedure involves medical informatics experts extracting and structuring a series of symptom combinations pointing to acute and critical illnesses, based on publicly available clinical guidelines. For example, chest pain accompanied by radiating pain in the left shoulder is defined as a high-risk event consisting of two clinically named entities: chest pain and radiating pain in the left shoulder. This combination and its corresponding risk level are stored in the Risk Upgrade Module's database. The second step defines a periodic update mechanism: the system is configured to automatically retrieve and compare updated versions of mainstream clinical guideline databases quarterly. When any changes to guideline content related to Red Flag symptoms are detected, the system automatically generates an update task awaiting review and prompts the system administrator for manual review and confirmation. Furthermore, all events where user-confirmed information contains Red Flag symptom combinations and triggers the Risk Upgrade Module to send alerts to human operator terminals are recorded in a security audit log. This log is used to continuously verify and update the effectiveness of the risk identification rules.

[0047] Example 6: This example discloses a series of standardized model construction and online adaptation procedures performed before deployment of the system of the present invention to ensure that its core algorithm module has reproducible recognition performance and can adapt to specific users and data environments during runtime; to enable the background entity and relation pre-annotation module to identify clinical named entities from the original narrative trajectory, the system uses a corpus containing millions of anonymized electronic medical record texts and public medical literature abstracts to train a sequence labeling model based on bidirectional long short-term memory network and conditional random field (BiLSTM-CRF) before deployment; the first step of this training procedure is to preprocess all text in the corpus, including The process involves three steps: First, word segmentation and word vector embedding are performed, and the text is labeled with entities according to the BIO annotation system. Second, the processed dataset is divided into training, validation, and test sets. The network parameters of the sequence labeling model are iteratively optimized using the training set, with the optimization objective being to minimize the prediction loss function on the validation set. Third, when the model's F1 score on the validation set fails to improve for several consecutive training cycles, training is terminated, and the entity recognition performance of the final model is evaluated using the test set. Only when the model's F1 score on the test set is not lower than the preset 0.92 is the model solidified and deployed in the background entity and relation pre-annotation module as its core algorithm for recognizing clinical named entities.

[0048] To enable the system to establish a personalized acoustic baseline for each user during interaction, the system executes an online baseline adaptive adjustment procedure during runtime. This procedure is configured to start at the initial stage of user interaction, i.e., after the system collects the first 30 seconds of non-diagnostic statement speech. At this stage, the system uses the initial group acoustic baseline determined in the calibration procedure of Example 3 as a temporary baseline. After the procedure starts, the system calculates the mean values ​​of various acoustic feature parameters (such as speech rate and energy envelope standard deviation) for the user within these 30 seconds using a backward moving average method. A weighted update algorithm is then used to fuse the initial group acoustic baseline with the user's actual acoustic feature mean values ​​from the first 30 seconds to generate a personalized acoustic baseline adapted to the user. This update algorithm ensures that the personalized acoustic baseline serves as a more representative reference in subsequent semantic weight calibration. To enable the background entity and relation pre-labeling module to utilize... The system adjusts the confidence level of clinical named entities using the user's historical medical record information, while ensuring system robustness in the event of missing information. A built-in data fusion and rollback logic is employed. When the system identifies a clinical named entity from the original narrative, it uses authorized identification information obtained from the user to initiate an asynchronous query request to an external electronic medical record database. If the query is successful and the user's historical medical record information contains a record matching the currently identified clinical named entity, the system multiplies the entity's confidence level by a preset weighting coefficient of 1.2. If the query fails to return valid information due to network interruption, database unresponsiveness, or lack of matching records in the dataset, the system skips the weighting adjustment step and directly uses the original confidence level output by the model for subsequent processing. This procedure allows the system to improve identification accuracy when data is available and seamlessly roll back to the standard processing flow when data is unavailable.

[0049] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A triage and pre-diagnosis system based on artificial intelligence, characterized in that, The system includes a narrative stream continuous acquisition module, a background entity and relation pre-annotation module, and a non-intrusive summary verification module. The narrative stream uninterrupted acquisition module is configured to receive free-formatted voice narrative input from users and continuously perform two synchronous operations: transcribing the content of the voice narrative input into a timestamped original narrative track, and extracting acoustic features from the voice narrative input to generate an acoustic prosodic track that is time-aligned with the original narrative track; The background entity and relation pre-annotation module is configured to establish a user-personalized acoustic baseline based on the acoustic features of non-diagnostic statements at the beginning of the dialogue. While the narrative flow uninterrupted acquisition module is working, it processes the original narrative track in parallel and asynchronously to identify clinical named entities and generate a temporary structured information sketch. It is also configured to compare the acoustic features in the acoustic prosodic track that are time-aligned with the clinical named entity with the acoustic baseline when a clinical named entity is identified to calculate a deviation value, and assign the deviation value as a semantic weight to the clinical named entity. The non-invasive summary confirmation module is configured to generate a natural language summary based on clinical named entities and semantic weights in the structured information sketch when a natural pause point defined by the user's voice pause duration exceeding a preset duration threshold is detected in the original narrative track. Only after receiving the user's confirmation instruction will the confirmed clinical named entities in the structured information sketch be written into the final structured pre-diagnosis report.

2. The AI-based triage and pre-diagnosis system according to claim 1, characterized in that, The non-invasive summary verification module is configured to integrate clinical named entities with high semantic weights when generating natural language summaries, and present them to users in an interactive dialogue that explores clinical focus, so as to guide users to make supplementary statements around the clinical named entities.

3. The AI-based triage and pre-diagnosis system according to claim 1, characterized in that, The system also includes a semantic entropy calculation unit and an interaction mode switching unit. The semantic entropy calculation unit is configured to calculate a semantic entropy value in real time based on the confidence of clinical named entities in a structured information sketch within a sliding time window and the number of contradictory relationships between clinical named entities. The interaction mode switching unit is configured to automatically switch the interaction mode of the narrative flow uninterrupted acquisition module from a listening mode that mainly receives user-initiated free-format voice narrative input to a guided mode that provides users with focused options to clarify core ambiguities when the semantic entropy value exceeds an entropy threshold.

4. The AI-based triage and pre-diagnosis system according to claim 3, characterized in that, The semantic entropy calculation unit is configured to calculate semantic entropy values ​​in real time. ,in ;in, This represents the average confidence level of all clinical named entities in the structured information sketch within the sliding time window. This represents the number of conflicting relationships between clinical named entities in a structured information sketch within a sliding time window. and The preset weighting coefficients; the interaction mode switching unit, configured to be used when... When the entropy threshold is exceeded, the interaction mode is switched.

5. The AI-based triage and pre-diagnosis system according to claim 1, characterized in that, The narrative stream uninterrupted acquisition module is also configured to distinguish the extraction of acoustic features into processing a first frequency band containing human voices and processing a second frequency band that does not contain human voices; the background entity and relation pre-labeling module is also configured to compensate the acoustic features of the first frequency band based on the acoustic features of the second frequency band before assigning semantic weights to clinical named entities.

6. The AI-based triage and pre-diagnosis system according to claim 5, characterized in that, The background entity and relation pre-labeling module is configured to perform compensation in the following way: subtract a compensation amount proportional to the energy value of the second frequency band from the energy value of the first frequency band to obtain a compensated energy value, and assign semantic weights to clinical named entities based on the compensated energy value.

7. The AI-based triage and pre-diagnosis system according to claim 1, characterized in that, The narrative stream uninterrupted acquisition module is configured to employ an open and unguided interaction strategy.

8. The AI-based triage and pre-diagnosis system according to claim 1, characterized in that, The non-invasive summary confirmation module is configured to highlight or prioritize clinical named entities that are pre-determined to be related to high-risk diseases based on the medical knowledge base and have been marked with high semantic weight when generating natural language summaries.

9. A triage and pre-diagnosis system based on artificial intelligence according to claim 8, characterized in that, The system also includes a risk escalation module; The risk escalation module is configured to send an alert to the operator's terminal device when the information confirmed by the user contains a predefined combination of high-risk symptoms consisting of multiple clinical named entities.

10. A triage and pre-diagnosis system based on artificial intelligence according to claim 1, characterized in that, The background entity and relation pre-labeling module is also connected to an external electronic medical record database; the background entity and relation pre-labeling module is also configured to perform weighted adjustments on the confidence of clinical named entities identified in the structured information sketch based on the user's historical medical record information obtained from the electronic medical record database.

Citation Information

Patent Citations

  • A pre-diagnosis system and method based on artificial intelligence

    CN118238163B

  • Multilingual speech recognition method and device and electronic equipment

    CN112185348A

  • Method and system for generating electronic medical record based on outpatient inquiry dialogue of large model

    CN117637097A