Intelligent CT (Computed Tomography) vehicle using method based on multi-modal data fusion

Multimodal data fusion technology is used to automatically identify medical staff and patients on CT ambulances and generate electronic medical records, which solves the problems of immediacy and accuracy in medical record generation on CT ambulances, simplifies the diagnosis process, and improves diagnostic efficiency and accuracy.

CN120809047AInactive Publication Date: 2025-10-17RENMIN HOSPITAL OF WUHAN UNIVERSITY (HUBEI GENERAL HOSPITAL)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511311193.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult to generate medical records instantly and accurately on a CT ambulance, especially when recording a large amount of diagnostic information on board. Due to limited space and the number of personnel, the traditional manual input method is inefficient and cannot be uploaded to the hospital server in a timely manner.

Method used

Using multimodal data fusion technology, video and audio acquisition devices combined with voiceprint recognition and coordinate system association, it automatically identifies medical staff and patients, generates electronic medical records, imports the patient's medical history, CT scan results and diagnosis content into the medical records through a multi-data fusion solution, and generates medical records through AI noise reduction and voice recognition.

Benefits of technology

It has achieved the instant and accurate generation of medical records on the CT ambulance, simplified the diagnosis process, improved the reliability and accuracy of diagnosis, ensured that medical records were uploaded to the hospital in a timely manner, and gained more time for rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809047A_ABST
    Figure CN120809047A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent CT vehicle using method based on multi-modal data fusion, a vehicle main body is included, a stretcher, brain CT equipment, an operation panel and a display terminal are arranged in a compartment of the vehicle main body, and a plurality of video acquisition devices and audio acquisition devices are further arranged in the compartment of the vehicle main body. The problem that the medical record file is generated on the CT ambulance immediately and accurately is solved, the medical record file can be updated in real time according to a CT scanning result and uploaded to a hospital immediately, and more rescue time is won for patients.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of ambulances, in particular to a method for using an intelligent CT ambulance based on multi-modal data fusion. BACKGROUND

[0002] Stroke is usually divided into ischemic and hemorrhagic types, and the treatment methods of the two types are different. It is difficult to determine which first-aid method should be used without using CT scanning. Since the golden rescue time for stroke is usually 3-4.5 hours, if CT scanning is performed after the patient arrives at the hospital and the scanning result is waited for, the rescue time may be missed. Therefore, in the prior art, a new ambulance integrating a brain CT device into a rescue device is provided, which performs CT scanning at the first time nearby after receiving the patient, thereby advancing the rescue.

[0003] The traditional ambulance can only perform brief first-aid treatment, and therefore needs to judge and record less content. However, for the ambulance with a CT device, since diagnosis is advanced, a large amount of diagnosis information needs to be recorded on the vehicle. Since the ambulance has a limited size, only a small number of medical staff can be carried, and after the CT result is obtained, the medical staff still needs to perform targeted treatment in the vehicle, and it is difficult to free up personnel to record information in the traditional manual input medical record way. SUMMARY

[0004] The application provides a method for using an intelligent CT ambulance based on multi-modal data fusion, which solves the problem of instant and accurate generation of medical record files on the CT ambulance.

[0005] To solve the above technical problems, the technical solution adopted by the application is as follows: an intelligent CT ambulance, comprising a vehicle main body, a stretcher, a brain CT device, an operation panel and a display terminal arranged in the vehicle compartment of the vehicle main body, and a plurality of video acquisition devices and audio acquisition devices arranged in the vehicle compartment of the vehicle main body.

[0006] In the preferred scheme, the method comprises the following steps: establishing a key word library of oral medical records; collecting voiceprint information of medical staff and establishing a voiceprint library; establishing a sound field coordinate system and a visual field coordinate system, and associating the two coordinate systems; performing CT scanning on the patient, and after the scanning is completed, the medical staff and the patient's family members get on the vehicle, and the vehicle departs; the audio acquisition device collects sound information, and shields sound sources outside the vehicle compartment and vehicle vibration noise; calculating the position coordinates of the sound sources in the vehicle compartment; the video acquisition device collects image information, collects the facial image of the patient, and identifies other person subjects in the vehicle compartment, and calculates the position coordinates of the person subjects in the vehicle compartment; The coordinates of each character subject position are compared with the coordinates of the sound source position, and the matching character subject and sound source subject in the image are bound; The sound source is compared with the voiceprint library, and if it is confirmed as a medical staff, the background system starts to establish the dictation medical record, and the terminal displays a prompt for the medical staff; The mouth position of the character subject who speaks is locked, and the key sound area is determined with the mouth position as the center. The moving state of the character subject is monitored, the path of the key sound area is predicted, and the center position of the area is fed back to the background system. The audio collection device strengthens the directional sound pickup of the key sound area at the next moment; The key sound area is denoised, the human voice is extracted, and the keyword monitoring is performed; The electronic medical record template is called, and the filling position in the electronic medical record template is pre-called according to the keyword semantics. After the effective content is collected, it is filled in the filling position; According to the collected patient face image, the patient's past medical history data is imported into the electronic medical record; The CT scan result data is imported into the electronic medical record; The electronic medical record is generated and displayed on the display terminal; The medical staff analyzes the CT scan result and inputs the diagnosis conclusion through the operation panel, and synchronously updates the electronic medical record; The electronic medical record is uploaded to the hospital server.

[0007] In the preferred scheme, one audio collection device is installed at each of the eight vertices in the car compartment to form a three-dimensional space positioning; During the voice recording stage of the dictation medical record, the indicator light is on, and the eight audio collection devices start to collect sound signals; According to the time sequence of the sound source received by the audio collection device, the audio inside the car compartment is filtered out and stored, and the audio outside the car compartment is removed.

[0008] In the preferred scheme, any two audio collection devices and are selected; The signals in the audio collection device are preprocessed; , ; According to the collected time domain signals and , the frequency domain signals and are obtained by Fourier transform; The generalized cross-correlation function is calculated: ; Where, is the audio collection deviceThe generalized cross-correlation function of is the Fourier transform, is the Fourier transform, and are the time-domain signals collected by the first and the second audio acquisition devices, represents the time delay, denotes the conjugate complex number, and the sound source arrival time difference is determined by finding the peak position of ; The sound source arrival time difference calculated according to the above formula is used to locate the three-dimensional coordinates of the sound source based on the spherical wave propagation model, and the formula of the spherical wave propagation model is: ; wherein, is the time difference of the sound source arriving at the first and the second audio acquisition devices, is the three-dimensional coordinates of the sound source, and are the three-dimensional coordinates of the first and the second audio acquisition devices, is the sound speed (343 m / s is taken); A plurality of sets of time difference data are obtained by the eight audio acquisition devices to form an overdetermined equation set, and the overdetermined equation set is solved by the least square method to calculate the three-dimensional coordinates of the sound source ; The audio inside the vehicle compartment is screened out, the audio signal located in the center area of the sound source is retained, and is recorded and stored.

[0009] In the preferred scheme, an AI noise reduction model integrating the spectrum characteristics of the engine low-frequency noise of the ambulance and the high-frequency noise of the CT device is established. After the audio outside the vehicle compartment is shielded, the AI noise reduction model is used to suppress specific noise of the real-time voice data, wherein the specific noise includes external environmental noise and in-vehicle noise.

[0010] In the preferred scheme, the method for extracting the effective content of the pure voice segment includes: The pure voice segment is input into the trained BioClinicalBERT model; The key sentences are extracted according to the pure voice segment; The medical terminology library is called to map the key sentences into standard medical terminology texts.

[0011] The beneficial effects of the present application are: the audio acquisition device cooperates with the intelligent algorithm to accurately identify the content, assists the medical staff to record the patient information, and can update the medical record archives in real time according to the CT scan results, upload the hospital in time, and strive for more rescue time for the patient; the audio acquisition device is arranged in three dimensions, the carriage is divided into an effective collection space, the sound outside the carriage is shielded, invalid sound sources are avoided, and the difficulty of voice extraction is simplified; the video acquisition device collects image data, and the moving direction of the sound source is predicted in advance and fed back to the background system, the system guides the audio acquisition device to move the key collection area synchronously, and the inaccuracy of voice recognition caused by audio acquisition lag is avoided; the multi-element data fusion scheme is adopted, the patient's past medical history, CT scan results, disease state image and diagnosis content of the medical staff are imported and electronic medical records are generated, the diagnosis process is simplified, and the reliability and accuracy of diagnosis are improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] The present application will be further described below in combination with the drawings and examples.

[0013] Figure 1 is a side view of the interior of the ambulance.

[0014] Figure 2 is a top view of the interior of the ambulance.

[0015] Figure 3 is a flowchart of generating and uploading spoken medical records.

[0016] In the figure: vehicle body 1; video acquisition device 2; audio acquisition device 3; stretcher 4; brain CT device 5; display terminal 6; operation panel 7. DETAILED DESCRIPTION

[0017] Example 1: As Figures 1-2 , an intelligent CT ambulance includes a vehicle body 1, a stretcher 4, a brain CT device 5, an operation panel 7 and a display terminal 6 are arranged in the carriage of the vehicle body 1, and a plurality of video acquisition devices 2 and audio acquisition devices 3 are arranged in the carriage of the vehicle body 1.

[0018] The ambulance is divided into a driving cabin and a carriage, and basic rescue instruments such as a defibrillator, a breathing machine, a cardiopulmonary resuscitation instrument, an electrocardiograph, an oxygen cylinder, a biochemical analyzer and a blood cell analyzer are also integrated in the carriage, and a ventilation and lighting system is provided. In addition to the control system, the vehicle body is also integrated with a battery, a communication system, a positioning system, a navigation system and the like.

[0019] Example 2: As Figure 3 , a method for using an intelligent CT vehicle based on multi-modal data fusion: Establish a spoken medical record keyword library; Collect the voiceprint information of medical staff, and establish a voiceprint library; Establish a sound field coordinate system and a field of view coordinate system, and coordinate the two coordinate systems; CT scanning is performed on the patient, after the scanning is completed, the medical staff and the patient's family members get on the vehicle, and the vehicle starts; The audio acquisition device 3 acquires sound information, and shields the sound source outside the vehicle compartment and the vibration noise of the vehicle; The position coordinates of the sound source in the vehicle compartment are calculated; The main principle of voice recognition sound source position is to use the time difference of sound source propagation to different acquisition devices. Because the speed of sound propagation is limited, the medical staff needs to check the patient, check the diagnosis result, dispense medicine, etc. in the ambulance, so there is frequent movement when speaking. If only multiple audio acquisition devices are used to predict the movement of the sound source, the target will be predicted to the next position within the sound propagation time, that is, the area calculated at the next time is lagging behind the actual acquisition area.

[0020] The video acquisition device 2 acquires image information, acquires patient face images, and identifies other character subjects in the vehicle compartment, and calculates the position coordinates of each character subject in the vehicle compartment; Compare the position coordinates of each character subject with the position coordinates of the sound source, and bind the matching character subject in the image with the sound source subject; Compare the sound source with the voiceprint library, if it is confirmed as medical staff, the background system starts to establish the spoken medical record, and the display terminal 6 prompts the medical staff; Lock the mouth position of the character subject who speaks, and define the key sound area centered on the mouth position, monitor the movement state of the character subject, predict the path of the key sound area, and feed back the center position of the area to the background system. The audio acquisition device 3 strengthens the directional sound pickup of the key sound area at the next time; Noise reduction is performed on the key sound area, human voice is extracted, and keyword monitoring is performed; Electronic medical record template calling, according to the keyword semantics, the filling position in the electronic medical record template is called in advance, and after the effective content is collected, it is filled into the filling position; According to the collected patient face image, the patient's past medical history data is imported into the electronic medical record; The CT scan result data is imported into the electronic medical record; Generate an electronic medical record and display it on the display terminal 6; The medical staff analyzes the CT scan result and inputs the diagnosis conclusion through the operation panel 7, and synchronously updates the electronic medical record; Upload the electronic medical record to the hospital server.

[0021] Install indicator lights in the vehicle compartment of the vehicle body 1, and install an audio acquisition device at each of the eight vertices in the vehicle compartment to form a three-dimensional space positioning; wherein the eight vertices include four corner points of the bottom of the vehicle cabin and four corner points of the top of the vehicle cabin, covering front-rear, left-right, and up-down directions; As shown in the figure, taking the bottom left corner vertex of the vehicle cabin as the origin Figure 1 , the length, width, and height of the vehicle cabin are , , , , then the coordinates of the eight vertices are: , , , , , , , .

[0022] Specifically, through the spatial distribution of the eight vertices, a three-dimensional sound field coverage is formed, providing geometric parameters for subsequent TDOA calculation.

[0023] The medical staff pre-records the voiceprint information in the vehicle cabin to form a voiceprint library; During the oral medical record sound recording stage, the indicator light is turned on, and the eight audio collection devices start collecting sound signals; wherein, according to the strength of the same audio signal in the eight audio collection devices, the position of the sound source is preliminarily judged, since the closer the audio collection device is to the sound source, the stronger the signal it receives during sound propagation, by comparing the signal strength received by each device, the direction and range of the sound source can be roughly determined. Provide an initial direction reference for the following calculation of time difference; The audio signals collected by the eight vertex audio collection devices in the vehicle cabin are determined by the generalized cross-correlation-phase transform (GCC-PHAT) algorithm to determine the time difference of the sound source arrival; According to the time sequence of the sound source received by each audio collection device, the audio inside the vehicle cabin is filtered out and recorded and stored; The specific calculation process is as follows: Collect any two audio collection devices and ; Preprocess the signals in the audio collection devices , , , that is, eliminate the direct current component and suppress low-frequency noise; Perform Fourier transform on the time domain signals and collected by them to obtain frequency domain signals and .

[0024] Then calculate the generalized cross-correlation function : ; By finding The peak position of the sound source can be used to determine the time difference between the two audio collection devices. The algorithm eliminates the influence of signal amplitude by normalizing the cross-power spectrum, so that the time difference can be detected more accurately even in a low signal-to-noise ratio environment, providing key data for the subsequent accurate three-dimensional coordinate positioning of the sound source.

[0025] in, For audio acquisition device The generalized cross-correlation function of is the Fourier transform, and They are and The time domain signal collected by the audio collection device 3, represents the time delay, Representing the complex conjugate, by finding The peak position of the sound source is used to determine the arrival time difference ; The time difference of arrival of the sound source calculated according to the above formula , the three-dimensional coordinates of the sound source are located based on the spherical wave propagation model. The formula of the spherical wave propagation model is: ; in, The sound source reaches and The time difference of the audio collection device, are the three-dimensional coordinates of the sound source, and They are and The three-dimensional coordinates of the audio acquisition device, is the speed of sound (taken as 343m / s).

[0026] Specifically, since there are eight audio acquisition devices, multiple groups of time difference data can be obtained, thereby forming an overdetermined set of equations.

[0027] By solving this overdetermined set of equations using the least squares method, we can obtain the three-dimensional coordinates of the sound source. ; Assume that the three-dimensional coordinates of the eight audio collection devices (microphones) in the ambulance are ( ), select the first microphone as the reference point ( ); right Seven independent equations can be established to form a nonlinear over-determined equation set about :

[0028] wherein, ; Specifically, each equation describes the relationship between the distance difference and time difference of the sound source to two microphones, and 8 microphones provide 7 equations (reference point minus 1), 3 unknowns, and the number of equations exceeds the number of unknowns, forming an over-determined system, which needs to be solved by the least square method to obtain the optimal solution.

[0029] In this embodiment, the least square method solving step is solved by nonlinear iterative optimization, specifically as follows: First, define the error function as the residual sum of squares of the left and right sides of the equation set, and the goal is to minimize this function:

[0030] By minimizing the residual sum of squares, it is ensured that the calculated time difference is closest to the measured value, which is suitable for robust estimation in a noisy environment.

[0031] Then, Taylor expansion and linearization are performed on each at the initial estimate value :

[0032] wherein the gradient vector is calculated as: ; Similarly, ; Specifically, the nonlinear equation is approximated as a linear equation, which is convenient for iterative solution. The initial value can be set as the geometric center of the microphone array or the coordinates of the previous frame.

[0033] Next, the Jacobian matrix is constructed, and the i-th row corresponds to the gradient of : ; wherein the residual vector is: ; Subsequently, the Gauss-Newton method is used to solve the linearized equation set: ; wherein is the coordinate correction amount. Solving this linear equation set obtains , and the sound source coordinates are updated as:

[0034] Specifically, if the matrix singular, damping factor can be introduced (such as Levenberg-Marquardt algorithm): ; wherein, is a damping coefficient to avoid divergence of iteration.

[0035] Finally, when the norm of the correction is less than a threshold, i.e. or the number of iterations reaches an upper limit, the iteration is terminated, and the optimal solution is obtained.

[0036] Using multiple sets of time difference data, the spatial position of the sound source in the vehicle cabin is accurately calculated.

[0037] According to the three-dimensional coordinates of the sound source, the sound source center area of the vehicle cabin is determined. First, multiple frames (such as frames) of time difference data are collected in real time, and the sound source coordinate sequence is obtained by least squares method.

[0038] Calculate the mean and standard deviation of the coordinates: ; ; ; Define the sound source center area as the range of mean ± 1.5 times the standard deviation: ; Specifically, by statistically suppressing random noise, 1.5 times the standard deviation covers about 93.3% of the data points, and in actual testing, it converges to (bed and medical operation area).

[0039] The range of this area is 1.25m×0.9m×1.1m, which is used to obtain effective speech data. The audio signals in this area are the main source of effective speech data, and the sound source center area provides a clear target range for subsequent audio signal enhancement and screening; Spatial filtering of audio signals in the sound source center area, based on sound source center area directional collection of real-time speech data, using minimum variance distortionless response (MVDR) beamforming algorithm to enhance the audio signal of medical staff, and through AI noise reduction model to dynamically suppress the external environmental noise and vehicle noise of real-time speech data; The process of solving the weight of the audio acquisition device by the MVDR beamforming algorithm is as follows: According to the three-dimensional coordinates of the sound source, the target direction vector is calculated, and then the optimization problem is solved as follows: ; wherein, is the weight of the mth audio acquisition device, is the propagation time delay of the sound source to the mth audio acquisition device, is the signal frequency. The weight of each audio acquisition device is obtained by solving this optimization problem , and the enhanced signal is synthesized

[0040] , achieving a main lobe gain increase of 12 dB and a side lobe suppression of -20 dB or lower, thereby directionally enhancing the audio signal of the medical staff. Further, if the coordinate solution exceeds the range of 1.25m x 0.9m x 1.1m, it is marked as invalid noise. Specifically, the AI noise reduction model integrates the frequency spectrum feature library of ambulance engine low-frequency noise (20-200Hz) and CT device high-frequency noise (>4kHz).

[0041] The AI noise reduction model adopts a dual-branch structure of Conv-TasNet (Time Domain Audio Separation Network) combined with RNN (Recurrent Neural Network). Among them, the Conv-TasNet branch is used for time domain signal separation, and the time domain features of speech and noise are extracted through a convolutional encoder-decoder structure. The RNN branch combines the frequency spectrum feature library to dynamically match and suppress the frequency domain noise features, eliminating steady-state and non-steady-state noise. The audio inside the vehicle compartment is screened out, and the audio signal located in the center of the sound source is retained and recorded. Through the previous sound source positioning and audio enhancement processing, this step further screens out the required audio signal, providing high-quality data for the next step of voiceprint recognition and speech processing.

[0042] The voiceprint library is called to determine the pre-stored voiceprint features of the medical staff, and the retained audio signal is compared with the voiceprint through the GMM-UBM algorithm. The GMM-UBM algorithm calculates the similarity between the detected audio signal and the pre-stored voiceprint features, and when the similarity meets a certain threshold, it is determined to be matched. The matched audio signal is marked as "valid human voice", and the unmatched audio signal is marked as "non-human voice" and removed, and the output is a pure speech segment with a signal-to-noise ratio of ≥18dB.

[0043] Specifically, through voiceprint recognition technology, non-medical staff speech interference is further excluded, ensuring the accuracy of subsequent processed speech data.

[0044] The indicator light of the display terminal is triggered synchronously, and when the voice acquisition is started, the green light is always on, the green light flashes during the acquisition process, and the green light is turned off when the acquisition is terminated.

[0045] The indicator light of the display terminal is triggered synchronously, and when the voice acquisition is started, the green light is always on, the green light flashes during the acquisition process, and the green light is turned off when the acquisition is terminated. ​

[0046] The reserved audio signal is subjected to DTW dynamic time warping with the pre-stored voiceprint features of the medical staff, voice recognition is started when the similarity is greater than or equal to 0.9, and the formula of the DTW dynamic time warping is as follows: ; wherein, is a feature sequence of a to-be-detected audio signal, is a pre-stored voiceprint feature sequence, is a distance between two feature points, is a weight on an alignment path.

[0047] Voice recognition is started when the calculated similarity is greater than or equal to 0.9, the DTW algorithm solves the stretching problem of the audio signal on the time axis, accurately determines whether to start voice recognition, and improves the accuracy and reliability of voice recognition.

[0048] When the green light is extinguished and the "end recording" voice instruction is detected, an encrypted voice data packet is generated and a time stamp is attached.

[0049] The pure voice segment is input into the pre-trained BioClinicalBERT model, the pure voice segment in the continuous dialogue is recognized, and the key sentences are extracted in combination with the role labels, including symptom description, suggestion and medication record; A medical terminology library is established, and semantic ambiguity is resolved through a context attention mechanism, the BioClinicalBERT pre-training model is called to map the key sentences to the medical terminology library, and a standard medical terminology text is generated.

[0050] Specifically, the medical terminology library contains standard clinical terms, and the development of BioClinicalBERT starts from further optimization of BioBERT (a biomedical model trained based on PubMed literature). The applicant found that although BioBERT performs excellently on biomedical texts, there is still room for improvement in term understanding and context capture in clinical records (such as electronic health records, EHR). Therefore, the team based on the weights of BioBERT conducted secondary training on the MIMICIII database (containing about 880M words of ICU patient clinical notes), forming a hybrid architecture of Bio+ClinicalBERT.

[0051] Based on the BioBERT-Basev1.0 (PubMed200K+PMC270K) initialization model parameters, the training configuration is batch size 32, maximum sequence length 128, learning rate 5e-5, training step number 150,000 steps, and the semantic and syntactic rules of clinical texts are learned through the mask language model (MLM) task.

[0052] BioClinicalBERT has outstanding performance in named entity recognition (NER) tasks, such as: accurately identifying drug names like “ibuprofen” and “amoxicillin”.

[0053] distinguishing between disease entities like “acute myocardial infarction” and “chronic heart failure”.

[0054] recognizing professional terms like “serum creatinine level”.

[0055] In natural language inference (NLI) tasks, the model can capture the logical relationships in clinical text. For example: identifying the causal chain in “the patient was admitted due to diabetic ketoacidosis and improved after insulin treatment”.

[0056] understanding the negative semantics in “no fever, cough, but there is dyspnea”.

[0057]

[0058] The above examples are only preferred technical solutions of the present application, and should not be regarded as a limitation of the present application. The protection scope of the present application should be based on the technical solutions recited in the claims, including equivalent replacement solutions of the technical features recited in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present application.

Claims

1. A method for using an intelligent CT vehicle based on multimodal data fusion, characterized by: A stretcher (4), a brain CT device (5), an operation panel (7), a display terminal (6), and a plurality of video acquisition devices (2) and audio acquisition devices (3) are arranged in the compartment of the vehicle body (1); Establish a keyword database for oral medical records; Collect voiceprint information of medical staff and establish a voiceprint database; Establish the sound field coordinate system and the field of view coordinate system, and associate the two coordinate systems; The patient underwent a CT scan. After the scan, the medical staff and the patient's family members boarded the vehicle and the vehicle departed. The audio collection device (3) collects sound information and shields the sound sources outside the vehicle compartment and the vehicle vibration noise; Calculate the position coordinates of the sound source in the car; The video acquisition device (2) acquires image information, acquires a facial image of the patient, identifies other human subjects in the carriage, and calculates the position coordinates of each human subject in the carriage; Compare the position coordinates of each person subject with the position coordinates of the sound source, and bind the matching person subject in the image with the sound source subject; The sound source is compared with the voiceprint database. If it is confirmed to be a medical staff, the background system starts to create an oral medical record, and the display terminal (6) prompts the medical staff; The mouth position of the person making the sound is locked, and a key sounding area is delineated with the mouth position as the center. The movement state of the person is monitored, and the path of the key sounding area is predicted. The center position of the area is fed back to the background system, and the audio collection device (3) performs directional sound pickup and reinforcement on the key sounding area at the next moment. Perform noise reduction on key vocal areas, extract human voices, and perform keyword monitoring; The electronic medical record template is called, and the filling position in the electronic medical record template is pre-called according to the keyword semantics, and the valid content is collected and filled in the filling position; Importing the patient's medical history data into the electronic medical record based on the collected patient facial image; Importing CT scan data into electronic medical records; Generate electronic medical records and display them on the display terminal (6); Medical staff analyze the CT scan results and input the diagnosis conclusion through the operation panel (7), and simultaneously update the electronic medical record; Electronic medical records are uploaded to the hospital server; An audio collection device (3) is installed at each of the eight vertices in the carriage to form a three-dimensional spatial positioning; During the audio recording phase of the oral medical history, the indicator lights up and the eight audio collection devices (3) begin to collect audio signals; According to the time sequence of the sound sources received by the audio acquisition device (3), the audio inside the vehicle is screened out, recorded and stored, and the audio outside the vehicle is discarded; Collect any two audio collection devices and ; Audio acquisition device in 、 Preprocess the signal; According to the collected time domain signal and Perform Fourier transform to obtain frequency domain signal and ; Calculate the generalized cross-correlation function : ; in, For audio acquisition device (3) The generalized cross-correlation function of is the Fourier transform, and They are and The time domain signal collected by the audio collection device (3) is represents the time delay, Representing the conjugate complex number, by finding The peak position of the sound source is used to determine the arrival time difference ; According to the time difference of arrival of the sound source , the three-dimensional coordinates of the sound source are located based on the spherical wave propagation model. The formula of the spherical wave propagation model is: ; in, The sound source reaches and The time difference of the audio collection device, are the three-dimensional coordinates of the sound source, and They are and The three-dimensional coordinates of the audio acquisition device, The speed of sound is 343m / s; Eight audio acquisition devices obtain multiple sets of time difference data to form an overdetermined equation system, which is solved by the least squares method to calculate the three-dimensional coordinates of the sound source. ; Filter out the audio inside the car, retain the audio signal in the center of the sound source, and record and store it.

2. The method for using the intelligent CT vehicle based on multimodal data fusion according to claim 1 is characterized by: Build an AI noise reduction model that integrates the spectrum feature library of low-frequency noise from ambulance engines and high-frequency noise from CT equipment; After the audio outside the car is shielded, the AI ​​noise reduction model is used to suppress the specific noise of the real-time voice data, where the specific noise includes external environmental noise and noise inside the car.

3. The method for using the intelligent CT vehicle based on multimodal data fusion according to claim 1 is characterized by: Including methods for extracting effective content of pure speech segments: Input the clean speech segment into the trained BioClinicalBERT model; Extract key sentences based on clean speech segments; The medical terminology database is called to map key sentences into standard medical terminology texts.

Citation Information

Patent Citations

  • Multi-robot cooperative 3D sound source identification and positioning method

    CN112379330A

  • 5G ambulance supporting voice transcription and patient remote physical examination

    CN116407405A

  • Method for generating medical record report based on doctor-patient dialogue

    CN119964717A

  • Traditional Chinese medicine intelligent inquiry method and system based on AI big language model

    CN120067279A

  • Ambulance with CT equipment

    CN210991237U