Intelligent medical image acquisition and report generation system based on voice interaction

The intelligent medical image acquisition and report generation system based on voice interaction solves the problem of image acquisition and report separation, realizes the synchronous generation and real-time transfer of images and reports, improves operational efficiency and report quality, and promotes the real-time nature of remote teaching and consultation.

CN121862293APending Publication Date: 2026-04-14HEFEI DVL ELECTRON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-27
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The existing medical image acquisition and diagnosis process suffers from fragmented operational procedures, poor information synchronization, and low efficiency in remote teaching and consultation. In particular, image acquisition control relies on foot switches and physical buttons, and the images and reports are separated, lacking a real-time communication mechanism.

Method used

An intelligent medical image acquisition and report generation system based on voice interaction is adopted. Through a voice acquisition module, a voice recognition module, an image acquisition and preprocessing module, a structured report automatic generation module, a hospital information system interface module, and a remote teaching and consultation communication module, an intelligent closed loop is realized for the entire process of image, voice, semantics, and report. Highly robust voice acquisition and recognition technology, delay summation beamforming algorithm, Wiener filtering noise reduction, CTC unaligned recognition, and BiLSTM + Softmax model are used to synchronously bind and generate images and reports.

Benefits of technology

It enables simultaneous image acquisition and report generation, reduces operational errors, improves the completeness and efficiency of reports, and facilitates real-time transfer of images and reports, thus promoting the real-time nature and professionalism of remote teaching and consultation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862293A_ABST
    Figure CN121862293A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent medical image acquisition and report generation system based on voice interaction, and relates to the technical field of intelligent medical treatment. Comprising a voice acquisition module, a voice recognition module, an image acquisition and preprocessing module, a structured report automatic generation module, a hospital information system interface module, a remote teaching and consultation communication module and an intelligent case knowledge base. And the voice acquisition module converts a natural language instruction and diagnosis description of a doctor in an examination process into equipment control behaviors and structured medical semantic data in real time through a high-robustness voice acquisition and recognition technology, so that image acquisition control and diagnosis information generation are synchronously completed on the same time axis. According to the invention, the human-computer interaction mode of traditional image acquisition and diagnosis is fundamentally reconstructed, and the whole-process intelligent closed loop of images, voice, semantics, reports and hospital information systems is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart healthcare technology, and in particular to a smart medical image acquisition and report generation system based on voice interaction. Background Technology

[0002] In current clinical practice, medical imaging equipment has become a core tool for disease diagnosis, surgical navigation, and efficacy evaluation. However, current image acquisition and control still mainly rely on foot switches, physical buttons, or touch controls. When performing ultrasound examinations or endoscopic procedures, doctors need to precisely control the probe or endoscope with both hands, while also using foot switches or manual methods to perform operations such as freezing, acquisition, and annotation. This fragmented workflow easily leads to problems such as acquisition delays, image omissions, and inaccurate annotations.

[0003] Currently, the diagnostic reporting process still relies primarily on doctors recalling information and manually entering it. There is a clear temporal separation between image acquisition and diagnostic writing, which not only results in a large workload but also makes it prone to omissions of medical information due to memory bias. Furthermore, the lack of semantic-level synchronization between image data and text reports hinders accurate traceability and intelligent analysis.

[0004] Furthermore, current medical information systems are clearly fragmented. While basic interfaces exist between HIS, PACS, and imaging workstations, they largely remain at the level of "post-event archiving," lacking an automatic communication mechanism based on the real-time acquisition process. Remote teaching and remote consultations typically rely on independent video conferencing systems, failing to achieve deep integration with image acquisition, audio explanation, and diagnostic processes. This significantly limits the real-time nature and professionalism of teaching and consultations, thus leaving room for improvement. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a voice-interactive intelligent medical image acquisition and report generation system. Its advantage lies in fundamentally reconstructing the traditional human-computer interaction method for image acquisition and diagnosis, achieving a fully intelligent closed loop encompassing image, voice, semantics, report, and hospital information systems.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: The intelligent medical image acquisition and report generation system based on voice interaction includes a voice acquisition module, a voice recognition module, an image acquisition and preprocessing module, a structured report automatic generation module, a hospital information system interface module, a remote teaching and consultation communication module, and an intelligent case knowledge base. The voice acquisition module uses robust voice acquisition and recognition technology to convert doctors' natural language commands and diagnostic descriptions during the examination process into equipment control behaviors and structured medical semantic data in real time, so that image acquisition control and diagnostic information generation are completed synchronously on the same time axis. The speech recognition module automatically identifies and structurally encodes the location, image features, lesion attributes, and diagnostic conclusions contained in the speech, achieving precise semantic binding between images and reports; the system achieves bidirectional data connectivity with the hospital's HIS and PACS, enabling a complete closed loop of real-time flow of patient information, examination requests, image data, and diagnostic reports; The remote teaching and consultation communication module extends the voice-driven mechanism to remote teaching and consultation scenarios, putting voice control, image acquisition, real-time explanation and expert annotation in the same voice interaction system, truly realizing a fully intelligent working mode of "what is said is what is controlled, what is controlled is what is recorded, what is recorded is what is taught, and what is taught is what is learned".

[0007] The invention is further configured such that, during the inspection process, the voice acquisition module performs voice interaction through a microphone array installed around the workstation or equipment. To adapt to inspection scenarios with high noise levels, a delay-summing beamforming algorithm is used. ;in, For the m-th microphone signal, To maximize noise suppression by optimizing the delay in the direction of the doctor, directional optimization is used.

[0008] The present invention is further configured such that the voice interaction uses Wiener filtering for noise reduction, and the noise reduction formula is: Used to improve the quality of speech signals before recognition.

[0009] The present invention is further configured such that the speech signal acquired by the speech recognition module is input through an acoustic model and a language model, and CTC is used to achieve unaligned recognition, as shown in the formula: The formula for identifying the optimal output sequence is: .

[0010] The present invention is further configured such that the text obtained from the speech recognition is processed by a medical-specific neural network for entity recognition and intent extraction, using a BiLSTM + Softmax sequence labeling model: , It identifies medical entities based on anatomical locations, imaging features, and lesion attributes, providing structured data for automated report generation.

[0011] The present invention is further configured such that, while performing voice control, the image acquisition and preprocessing module performs real-time encoding and key frame extraction on the acquired video stream, and completes synchronous binding with the voice semantic data through a unified timestamp, so that any voice description can be accurately matched with the corresponding image frame, fundamentally solving the industry problem of "images and reports not being synchronized".

[0012] The present invention is further configured such that the system is deeply connected with HIS and PACS systems through a standardized interface, and the patient's basic information and examination request are automatically synchronized to the workstation before the examination begins. After the image acquisition is completed, the image data and structured report are automatically sent back for archiving, avoiding manual re-entry.

[0013] The present invention is further configured such that, in the remote teaching and consultation scenario of the system, real-time image streams, voice explanations, and AI analysis results are synchronously pushed to remote terminals through an encrypted link, and experts can remotely annotate, explain, and guide through voice, thereby achieving true "voice-driven remote collaboration".

[0014] The present invention is further configured such that, during the long-term operation of the system, all voice, image and report data are entered into the intelligent case knowledge base after de-identification processing, which is used to construct searchable, analyzable and learnable medical corpus and image data assets, providing a data foundation for subsequent AI model training and assisted diagnosis.

[0015] The beneficial effects of this invention are as follows: 1. By establishing voice interaction as the core control channel in the medical image acquisition and diagnosis process, the traditional operation mode relying on foot pedals, buttons, and mice has been completely changed, enabling doctors to maintain a continuous, natural, and uninterrupted operating experience during the examination. The doctor's verbal behavior directly drives the device's behavior, achieving true "hands-free operation" and significantly reducing human error and operator fatigue.

[0016] 2. Through speech semantic analysis and an automatic structured report generation mechanism, the simultaneous completion of image acquisition and report generation is achieved, transforming the diagnostic process from "post-event recall and recording" to "real-time process generation." This not only significantly improves report writing efficiency but also enhances the completeness, standardization, and traceability of report content. The strong semantic correlation between images, speech, and reports provides a reliable data foundation for subsequent refined quality control, scientific research analysis, and teaching review.

[0017] 3. Through deep automatic connectivity with HIS and PACS systems, a real-time closed loop is formed between image acquisition, diagnostic generation, and intra-hospital information flow, avoiding the efficiency losses and data risks caused by repeated manual data handling in traditional systems. The integrated voice-driven remote teaching and consultation modules ensure that teaching and consultation are no longer supplementary steps after image acquisition, but rather synchronous activities during the acquisition process, effectively promoting the dissemination of high-quality medical resources to grassroots levels. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the voice-driven system of the intelligent medical image acquisition and report generation system based on voice interaction proposed in this invention. Figure 2 This is a flowchart of the speech processing and medical semantic parsing of the intelligent medical image acquisition and report generation system based on voice interaction proposed in this invention. Figure 3 This is a timing diagram of the voice-image-time synchronization binding of the intelligent medical image acquisition and report generation system based on voice interaction proposed in this invention; Figure 4 This is a flowchart illustrating the security verification and fault tolerance process of the voice-driven device control in the intelligent medical image acquisition and report generation system based on voice interaction proposed in this invention. Detailed Implementation

[0019] The technical solution of this patent will be further described in detail below with reference to specific embodiments.

[0020] The embodiments of this patent are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this patent, and should not be construed as limiting this patent.

[0021] Reference Figure 1-4 The intelligent medical image acquisition and report generation system based on voice interaction includes a voice acquisition module, a voice recognition module, an image acquisition and preprocessing module, a structured report automatic generation module, a hospital information system interface module, a remote teaching and consultation communication module, and an intelligent case knowledge base. The voice acquisition module uses robust voice acquisition and recognition technology to convert doctors' natural language commands and diagnostic descriptions during the examination process into equipment control behaviors and structured medical semantic data in real time, so that image acquisition control and diagnostic information generation are completed synchronously on the same time axis. The speech recognition module automatically identifies and structurally encodes the location, image features, lesion attributes, and diagnostic conclusions contained in the speech, achieving precise semantic binding between images and reports; the system achieves bidirectional data connectivity with the hospital's HIS and PACS, enabling a complete closed loop of real-time flow of patient information, examination requests, image data, and diagnostic reports; The remote teaching and consultation communication module extends the voice-driven mechanism to remote teaching and consultation scenarios, putting voice control, image acquisition, real-time explanation and expert annotation in the same voice interaction system, truly realizing a fully intelligent working mode of "what is said is what is controlled, what is controlled is what is recorded, what is recorded is what is taught, and what is taught is what is learned".

[0022] In clinical applications, doctors wear or use multi-microphone arrays deployed on workstations for voice interaction. The system effectively eliminates equipment noise, environmental noise, and echo interference through beamforming, adaptive noise reduction, and echo suppression technologies, ensuring stable capture of voice signals even in the complex sound fields of operating rooms and examination rooms. During examinations, doctors issue commands in natural language, such as "freeze image," "start acquisition," "zoom in on right lobe," and "mark lesion." The voice signals are sent to the speech recognition module in real time and converted into text information.

[0023] In this embodiment, the voice acquisition module performs voice interaction during the inspection process using a microphone array installed around the workstation or equipment. To adapt to noisy inspection scenarios, a delay-summing beamforming algorithm is used. ;in, For the m-th microphone signal, To minimize the delay in the direction of the doctor, directional optimization is used to maximize noise suppression. Wiener filtering is used for noise reduction in voice interaction; the noise reduction formula is: Used to improve the quality of speech signals before recognition.

[0024] The speech signal acquired by the speech recognition module is input through the acoustic model and the language model, and CTC is used to achieve unaligned recognition. The formula is as follows: The formula for identifying the optimal output sequence is: .

[0025] The use of a medical lexicon to enhance the language model significantly improved the recognition accuracy of medical terms such as "gallbladder," "echo," and "sound shadow."

[0026] Furthermore, the voice text is fed into a medical semantic understanding engine, which distinguishes and parses operational instructions and diagnostic descriptions based on a medical-specific semantic model. When device control semantics are identified, the system automatically generates standardized control instructions, directly driving the ultrasound device or endoscope to perform the corresponding actions through the device control interface. When diagnostic semantics are identified, the system automatically extracts lesion location, imaging features, qualitative descriptions, and diagnostic tendency information, and fills in the information in a structured manner according to the standard template for medical reports.

[0027] The text obtained from speech recognition is processed by a medical-specific neural network for entity recognition and intent extraction, using a BiLSTM + Softmax sequence labeling model: , It identifies medical entities based on anatomical locations, imaging features, and lesion attributes, providing structured data for automated report generation.

[0028] While executing voice control, the image acquisition and preprocessing module performs real-time encoding and keyframe extraction on the acquired video stream, and synchronizes it with the voice semantic data through a unified timestamp. This ensures that any voice description can be accurately mapped to the corresponding image frame, fundamentally solving the industry problem of "images and reports not being synchronized".

[0029] Notably, the system achieves deep integration with HIS and PACS systems through standardized interfaces. Patient basic information and examination requests are automatically synchronized to the workstation before the examination begins. After image acquisition, image data and structured reports are automatically returned for archiving, avoiding repetitive manual data entry. For remote teaching and consultation, the system uses encrypted links to simultaneously push real-time image streams, audio explanations, and AI analysis results to remote terminals. Experts can remotely annotate, explain, and guide via voice, achieving true "voice-driven remote collaboration." During long-term operation, all audio, image, and report data, after anonymization, are entered into an intelligent case knowledge base to build searchable, analyzable, and learnable medical corpora and image data assets, providing a data foundation for subsequent AI model training and assisted diagnosis.

[0030] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A voice-interactive intelligent medical image acquisition and report generation system, characterized in that, It includes a voice acquisition module, a voice recognition module, an image acquisition and preprocessing module, a structured report automatic generation module, a hospital information system interface module, a remote teaching and consultation communication module, and an intelligent case knowledge base; The voice acquisition module uses robust voice acquisition and recognition technology to convert doctors' natural language commands and diagnostic descriptions during the examination process into equipment control behaviors and structured medical semantic data in real time, so that image acquisition control and diagnostic information generation are completed synchronously on the same time axis. The speech recognition module automatically identifies and structurally encodes the location, image features, lesion attributes, and diagnostic conclusions contained in the speech, achieving precise semantic binding between images and reports; the system achieves bidirectional data connectivity with the hospital's HIS and PACS, enabling a complete closed loop of real-time flow of patient information, examination requests, image data, and diagnostic reports; The remote teaching and consultation communication module extends the voice-driven mechanism to remote teaching and consultation scenarios, putting voice control, image acquisition, real-time explanation and expert annotation in the same voice interaction system, truly realizing a fully intelligent working mode of "what is said is what is controlled, what is controlled is what is recorded, what is recorded is what is taught, and what is taught is what is learned".

2. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 1, characterized in that, During the inspection process, the voice acquisition module uses a microphone array installed around the workstation or equipment for voice interaction. To adapt to noisy inspection scenarios, a delay-summing beamforming algorithm is used. ;in, For the m-th microphone signal, To maximize noise suppression by optimizing the delay in the direction of the doctor, directional optimization is used.

3. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 2, characterized in that, The voice interaction uses Wiener filtering for noise reduction, and the noise reduction formula is as follows: ; Used to improve the quality of speech signals before recognition.

4. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 1, characterized in that, The speech signal acquired by the speech recognition module is input through the acoustic model and the language model, and CTC is used to achieve unaligned recognition, as shown in the formula: The formula for identifying the optimal output sequence is: 。 5. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 4, characterized in that, The text obtained from speech recognition is processed by a medical-specific neural network for entity recognition and intent extraction, using a BiLSTM + Softmax sequence labeling model: , It identifies medical entities based on anatomical locations, imaging features, and lesion attributes, providing structured data for automated report generation.

6. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 1, characterized in that, While executing voice control, the image acquisition and preprocessing module performs real-time encoding and keyframe extraction on the acquired video stream, and synchronizes it with the voice semantic data through a unified timestamp, so that any voice description can be accurately matched with the corresponding image frame, fundamentally solving the industry problem of "images and reports not being synchronized".

7. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 1, characterized in that, The system is deeply connected to HIS and PACS systems through standardized interfaces. Patient basic information and examination requests are automatically synchronized to the workstation before the examination begins. After image acquisition is completed, image data and structured reports are automatically sent back for archiving, avoiding manual re-entry.

8. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 7, characterized in that, In the remote teaching and consultation scenarios described in the system, real-time image streams, voice explanations, and AI analysis results are synchronously pushed to remote terminals through encrypted links. Experts can remotely annotate, explain, and guide via voice, achieving true "voice-driven remote collaboration".

9. The intelligent medical image acquisition and report generation system based on voice interaction according to claim 8, characterized in that, During the long-term operation of the system, all voice, image, and report data are anonymized and then entered into an intelligent case knowledge base. This base is used to construct searchable, analyzable, and learnable medical corpus and image data assets, providing a data foundation for subsequent AI model training and assisted diagnosis.