Processing Method, Device, Electronic Device and Storage Medium for Lung Pathological Images

By combining the processing methods of lung pathological images and vocalprint sequences, the problem of low efficiency and accuracy of pathological images in the prior art is solved, and more efficient and accurate detection of the region of interest is achieved.

CN114693911BActive Publication Date: 2025-06-17SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011584096.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-28
Publication Date
2025-06-17
Estimated Expiration
2040-12-28

AI Technical Summary

Technical Problem

The existing manual observation pathological images have relatively low video reading efficiency and accuracy, which are time-consuming and labor-intensive and subjective, and may cause misjudgment.

Method used

By using the lung pathological images of the target personnel and the voiceprint sequence when the target personnel vocalizes according to the specified content, the lung pathological images are detected, and the voiceprints are used as auxiliary detection, which improves the accuracy of the detection.

Benefits of technology

The accuracy of detecting areas of interest in pathological images is improved, the need for manual observation is reduced, and the video reading efficiency of pathological images is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693911B_ABST
    Figure CN114693911B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method for processing lung pathological images, including: obtaining a lung pathological image and a voiceprint sequence of a target person, where the voiceprint sequence is collected when the target person makes a sound according to specified content, the lung pathological image includes a region of interest, and the voiceprint sequence includes time-domain information and frequency-domain information; extracting features from the lung pathological image through a preset image feature extraction network to obtain a lung pathological feature map; extracting features from the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map; decoding the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network to obtain a decoding result as the processing result of the lung pathological image, and the processing result of the lung pathological image includes a region of interest. The efficiency of reading lung pathological images is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method, device, electronic device and storage medium for processing pulmonary pathological images. Background Art

[0002] Pathological examination is a method for examining the pathological morphology in organs, tissues or cells of an organism. Currently, it is usually used to diagnose whether a person has cancer. The method of making a pathological diagnosis by simply observing morphology is a purely qualitative and morphological method, which can only make a rough quantitative estimation. For example, the malignancy of a malignant tumor is judged according to the number of nuclear divisions of tumor cells, especially pathological nuclear divisions. The examination method of pathological morphology is to first observe the pathological changes of the specimen, then cut a certain size of diseased tissue, make a pathological section by pathological histological method, and further examine the lesion under a microscope. Specifically, the tissue to be examined is sectioned and then stained to obtain different stained images, such as immunohistochemical stained images. Doctors in the pathology department complete the diagnosis of a case of pathology by observing the whole and local regions of interest in the stained images under a microscope. However, observing stained images manually, such as looking for diseased regions under a microscope, is time-consuming and laborious, with low reading efficiency, and there is a large degree of subjectivity, and misjudgment may occur. Therefore, the existing reading efficiency and accuracy of manually observing pathological images are relatively low. Summary of the Invention

[0003] An embodiment of the present invention provides a method for processing pulmonary pathological images, which can process pulmonary pathological images by using the pulmonary pathological images of a target person and the voiceprint sequence when the target person vocalizes according to specified content, so as to detect the regions of interest in the pulmonary pathological images. Since voiceprint is added as an auxiliary detection, the accuracy of detecting the regions of interest in pathological images is improved, and there is no need to observe pathological images manually, thereby improving the reading efficiency of pathological images.

[0004] In a first aspect, an embodiment of the present invention provides a method for processing pulmonary pathological images, the method comprising:

[0005] Obtaining a pulmonary pathological image and a voiceprint sequence of a target person, the voiceprint sequence being collected when the target person vocalizes according to specified content, the pulmonary pathological image including regions of interest, and the voiceprint sequence including time-domain information and frequency-domain information;

[0006] Performing feature extraction on the pulmonary pathological image through a preset image feature extraction network to obtain a pulmonary pathological feature map;

[0007] Performing feature extraction on the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map;

[0008] Decode the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network, and obtain the decoding result as the processing result of the lung pathological image. The processing result of the lung pathological image includes the region of interest.

[0009] Optionally, the preset image feature extraction network includes a preset global feature extraction network and a preset local feature extraction network. The lung pathological feature map includes a lung pathological global feature map and a local pathological feature map. The step of obtaining the lung pathological feature map by extracting features from the lung pathological image through the preset image feature extraction network includes:

[0010] Input the lung pathological image into the preset global feature extraction network to obtain a lung pathological global feature map;

[0011] Randomly slice the lung pathological image to obtain a plurality of local lung pathological images, and sequentially input the plurality of local lung pathological images into the preset local feature extraction network to obtain a local pathological feature map.

[0012] Optionally, the step of inputting the lung pathological image into the global feature extraction network to obtain a lung pathological global feature map includes:

[0013] Sequentially extract a first global feature map, a second global feature map, and a third global feature map at different depths of the preset global feature extraction network;

[0014] The scale resolution of the first global feature map is greater than that of the second global feature map, and the scale resolution of the second global feature map is greater than that of the third global feature image.

[0015] Optionally, the step of obtaining the time-frequency voiceprint feature map by extracting features from the voiceprint sequence through a preset voiceprint feature extraction network includes:

[0016] Sequentially extract a first voiceprint feature map, a second voiceprint feature map, and a third voiceprint feature map at different depths of the preset voiceprint feature extraction network;

[0017] The scale resolution of the first voiceprint feature map is the same as that of the first global feature map, the scale resolution of the second voiceprint feature map is the same as that of the second global feature map, and the scale resolution of the third voiceprint feature map is the same as that of the third global feature map.

[0018] Optionally, the scale resolution of the local pathological feature map is the same as that of the first global feature map. The process of decoding the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network to obtain a decoding result as the processing result of the lung pathological image includes:

[0019] Fusing the third voiceprint feature map and the third global feature map through a preset first fusion method to obtain a first fused feature map;

[0020] Performing a first upsampling on the first fused feature map to upsample the first fused feature map to the scale resolution of the second global feature map, obtaining a first upsampled feature map;

[0021] Fusing the second voiceprint feature map, the second global feature map, and the first upsampled feature map through a preset second fusion method to obtain a second fused feature map;

[0022] Performing a second upsampling on the second fused feature map to upsample the second fused feature map to the scale resolution of the first global feature map, obtaining a second upsampled feature map;

[0023] Fusing the first voiceprint feature map, the first global feature map, and the second upsampled feature map through a preset third fusion method to obtain a third fused feature map;

[0024] Performing a third upsampling on the third fused feature map to upsample the third fused feature map to the scale resolution of the lung pathological image, obtaining a decoding result as the processing result of the lung pathological image, and the processing result of the lung pathological image includes a region of interest.

[0025] Optionally, the obtaining of the voiceprint sequence of the target person includes:

[0026] Obtaining a first voiceprint sequence collected when the target person makes a sound according to the specified content;

[0027] Denosing the first voiceprint sequence to obtain a second voiceprint sequence;

[0028] Converting the second voiceprint sequence from time-domain information to frequency-domain information, and obtaining the voiceprint sequence of the target person according to the time-domain information and the frequency-domain information of the second voiceprint sequence.

[0029] Optionally, the converting of the second voiceprint sequence from time-domain information to frequency-domain information includes:

[0030] Performing frame segmentation on the second voiceprint sequence to obtain a second voiceprint sequence after frame segmentation;

[0031] Perform windowing processing on the second voiceprint sequence after frame segmentation processing to obtain a windowed second voiceprint sequence;

[0032] Perform fast Fourier transform point number processing on the windowed second voiceprint sequence to convert the second voiceprint sequence from time-domain information to frequency-domain information.

[0033] Optionally, the image feature extraction network, the voiceprint feature extraction network, and the decoding network are trained through the same data set. The steps of the training include:

[0034] Construct a training data set, which includes sample lung pathological images, sample voiceprint sequences, and corresponding region-of-interest annotation data. The sample lung pathological images are the lung pathological images of sample personnel, the region-of-interest annotation data are the annotated regions of the lung pathological images of sample personnel, and the sample voiceprint sequences are the voiceprint sequences collected when the sample personnel vocalize according to specified content;

[0035] Through the training data set, jointly train the image feature extraction network, the voiceprint feature extraction network, and the decoding network.

[0036] Optionally, the jointly training the image feature extraction network, the voiceprint feature extraction network, and the decoding network through the training data set includes:

[0037] Adjust the network parameters of the image feature extraction network, the voiceprint feature extraction network, and the decoding network by performing backpropagation by calculating the error loss of the image feature extraction network.

[0038] In a second aspect, an embodiment of the present invention further provides a processing device for lung pathological images. The device includes:

[0039] An acquisition module, configured to acquire a lung pathological image and a voiceprint sequence of a target person. The voiceprint sequence is collected when the target person vocalizes according to specified content. The lung pathological image includes a region of interest, and the voiceprint sequence includes time-domain information and frequency-domain information;

[0040] A first extraction module, configured to extract features from the lung pathological image through a preset image feature extraction network to obtain a lung pathological feature map;

[0041] A second extraction module, configured to extract features from the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map;

[0042] A decoding module, configured to decode the pulmonary pathological feature map and the time-frequency voiceprint feature map through a preset decoding network, and obtain a decoding result as a processing result of the pulmonary pathological image, where the processing result of the pulmonary pathological image includes a region of interest.

[0043] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the method for processing a pulmonary pathological image provided by the embodiment of the present invention are implemented.

[0044] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for processing a pulmonary pathological image provided by the embodiment of the invention are implemented.

[0045] In the embodiment of the present invention, a pulmonary pathological image and a voiceprint sequence of a target person are obtained. The voiceprint sequence is collected when the target person vocalizes according to a specified content. The pulmonary pathological image includes a region of interest, and the voiceprint sequence includes time-domain information and frequency-domain information; a pulmonary pathological feature map is obtained by performing feature extraction on the pulmonary pathological image through a preset image feature extraction network; a time-frequency voiceprint feature map is obtained by performing feature extraction on the voiceprint sequence through a preset voiceprint feature extraction network; the pulmonary pathological feature map and the time-frequency voiceprint feature map are decoded through a preset decoding network, and a decoding result is obtained as a processing result of the pulmonary pathological image, where the processing result of the pulmonary pathological image includes a region of interest. It is possible to process the pulmonary pathological image by using the pulmonary pathological image of the target person and the voiceprint sequence when the target person vocalizes according to the specified content, so as to detect the region of interest in the pulmonary pathological image. Since the voiceprint is added as an auxiliary detection, the accuracy of detecting the region of interest in the pathological image is improved, and there is no need to manually observe the pathological image, thereby improving the reading efficiency of the pathological image. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a flowchart of a method for processing a pulmonary pathological image provided by an embodiment of the present invention;

[0048] Figure 2It is a flowchart of a method for obtaining a voiceprint sequence provided by an embodiment of the present invention;

[0049] Figure 2a It is a schematic diagram of the relationship between time-domain information and frequency-domain information provided by an embodiment of the present invention;

[0050] Figure 3 It is a flowchart of converting time-domain information into frequency-domain information provided by an embodiment of the present invention;

[0051] Figure 4 It is a flowchart of extracting a pulmonary pathological feature map provided by an embodiment of the present invention;

[0052] Figure 5 It is an overall schematic diagram of a processing system for pulmonary pathological images provided by an embodiment of the present invention;

[0053] Figure 6 It is a flowchart of training a processing model for pulmonary pathological images provided by an embodiment of the present invention;

[0054] Figure 7 It is a structural schematic diagram of a processing device for pulmonary pathological images provided by an embodiment of the present invention;

[0055] Figure 8 It is a structural schematic diagram of a first extraction module provided by an embodiment of the present invention;

[0056] Figure 9 It is a structural schematic diagram of a decoding module provided by an embodiment of the present invention;

[0057] Figure 10 It is a structural schematic diagram of an acquisition module provided by an embodiment of the present invention;

[0058] Figure 11 It is a structural schematic diagram of a conversion sub-module provided by an embodiment of the present invention;

[0059] Figure 12 It is a structural schematic diagram of another processing device for pulmonary pathological images provided by an embodiment of the present invention;

[0060] Figure 13 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for processing lung pathological images provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:

[0063] 101. Obtain the lung pathological image and voiceprint sequence of the target person.

[0064] In the embodiment of the present invention, the above voiceprint sequence is collected when the target person makes a sound according to the specified content. The above lung pathological image includes the region of interest, and the voiceprint sequence includes time-domain information and frequency-domain information. The above lung pathological image can be an image of a pathological section obtained by an electron microscope or an optical microscope, or an image taken by an endoscope. In the pathological image, after the section is processed by a reagent, there will be a differential color state between the diseased cells and the normal cells, such as the color difference and shape difference of the cell nuclei.

[0065] The above region of interest refers to the region where the lung diseased cells are located, such as the region where the lung tumor cells are located.

[0066] The above voiceprint sequence can be obtained by a voice sensor (such as a microphone). The above voiceprint sequence can be framed according to a preset framing rule. For example, the preset framing rule includes a frame rate n and a duration t, that is, there are n voiceprint frames per second in the voiceprint sequence, the duration of the voiceprint sequence is t seconds, and the total number of voiceprint frames N in the voiceprint sequence = n×t. The above specified content can be the text content uttered during exhalation, such as coughing sounds, or text content with vowels "e", "o", "ong", etc., so that the voiceprint sequence includes the implicit information during exhalation through the lungs.

[0067] It should be noted that the above lung pathological image can be one frame or a continuous multi-frame. In the embodiment of the present invention, the above lung pathological image is preferably one frame, and the above voiceprint sequence is preferably 1 second or 2 seconds.

[0068] Of course, when the lung pathological image is a continuous multi-frame image, the above voiceprint sequence can have the same preset duration as the engine voiceprint sequence. For example, if the lung pathological image sequence is 3 seconds, then the voiceprint sequence is also 3 seconds. Further, it can be understood that the number of frames of the above lung pathological image sequence is the same as the number of frames of the voiceprint sequence, and the interval time between two adjacent frames in the lung pathological image sequence is the same as the interval time between two adjacent frames in the voiceprint sequence.

[0069] The above time-domain information is the voiceprint information originally collected by the sound sensor, which can be understood as the voiceprint information described along the time axis. The above frequency-domain information can be obtained by converting the time-domain information, and the above frequency-domain information can be understood as the voiceprint information described along the frequency axis. The conversion of dynamic voiceprint information from the time domain to the frequency domain can be achieved through Fourier series and Fourier transform.

[0070] Optionally, please refer to Figure 2 , Figure 2 which is a flowchart of a method for obtaining a voiceprint sequence provided by an embodiment of the present invention. As Figure 2 shown, the conversion method includes the following steps:

[0071] 201. Obtain a first voiceprint sequence collected when a target person makes a sound according to a specified content.

[0072] In an embodiment of the present invention, when obtaining a pulmonary pathological image of a target person, the target person can be allowed to make a sound according to a specified content, and the voiceprint information of the target person can be collected in real time through a sound sensor as the first voiceprint sequence. At this time, the pulmonary pathological image and the first voiceprint sequence are matched in time. It can also be that the target person is separately allowed to make a sound according to a specified content, and the voiceprint information of the target person can be collected in real time through a sound sensor as the first voiceprint sequence. At this time, the pulmonary pathological image and the first voiceprint sequence are not matched in time, but still belong to the same target person. In an embodiment of the present invention, the time interval between obtaining the pulmonary pathological image and the first voiceprint sequence does not exceed a preset time threshold. For example, the time interval between obtaining the pulmonary pathological image and the first voiceprint sequence does not exceed 10 minutes.

[0073] 202. Denoise the first voiceprint sequence to obtain a second voiceprint sequence.

[0074] In an embodiment of the present invention, the first voiceprint sequence can be denoised through autocorrelation denoising processing to eliminate the environmental noise of the first voiceprint sequence and obtain a second voiceprint sequence.

[0075] Currently, the above step 202 is optional. In some environments with good sound collection conditions, noise elimination may not be required. For example, when collecting a voiceprint sequence in an environment provided by a hospital specifically for collecting voiceprint sequences, the above step 202 can be skipped and directly proceed to step 203.

[0076] 203. Convert the second voiceprint sequence from time-domain information to frequency-domain information, and obtain the voiceprint sequence of the target person according to the time-domain information and frequency-domain information of the second voiceprint sequence.

[0077] In an embodiment of the present invention, the voiceprint sequence of the above target person may include both time-domain information and frequency-domain information, and the conversion of the second voiceprint sequence from time-domain information to frequency-domain information may be performed through short-time Fourier transform. After obtaining the frequency-domain information of the second voiceprint sequence, the frequency-domain information of the second voiceprint sequence is fused with the time-domain information of the second voiceprint sequence. The obtained voiceprint sequence after fusion has both a time dimension and a frequency dimension. The above voiceprint sequence may be as Figure 2a shown.

[0078] Specifically, please refer to Figure 3 , Figure 3 which is a flowchart for converting time-domain information to frequency-domain information provided by an embodiment of the present invention. As Figure 3 shown, it includes the following steps:

[0079] 301. Perform frame division on the second voiceprint sequence to obtain the second voiceprint sequence after frame division.

[0080] In an embodiment of the present invention, the above frame division may be to divide the second voiceprint sequence according to a preset frame division plan. When the pulmonary pathological image is a continuous multi-frame image, it may be divided according to the number of frames in the pulmonary pathological image, that is, the number of frames included in the pulmonary pathological image, the second voiceprint sequence will be divided into that many frames. For example, if the pulmonary pathological image sequence includes 100 frames, the second voiceprint sequence will be divided into 100 frames. It should be noted that the duration of the pulmonary pathological image sequence is the same as that of the second voiceprint sequence. Of course, in the case where the pulmonary pathological image is one frame, the voiceprint sequence may be divided according to user needs. For example, it is divided into n frames per second, where n is greater than or equal to 1.

[0081] Further, the length of the above second voiceprint sequence may be represented by the following formula:

[0082] N = f s × t

[0083] where N is the length of the second voiceprint sequence, f s is the sampling frequency of the second voiceprint sequence (the sampling frequency of the sound sensor), and t is the sampling duration of the second voiceprint sequence.

[0084] After frame division, the above second voiceprint sequence of the engine is:

[0085] {y1, y2, …, y m}

[0086] where the above y m is the mth voiceprint frame in the second voiceprint sequence after frame division, and the above m is the total number of voiceprint frames in the second voiceprint sequence after frame division.

[0087] In a possible embodiment, to make the transition between acoustic feature frames in the second acoustic feature sequence after frame segmentation smooth and maintain the continuity of the acoustic feature data, an overlapping segmentation method can be used to perform frame segmentation on the second acoustic feature sequence. The overlapping segmentation method can be understood as having an overlapping part between the current acoustic feature frame and the previous acoustic feature frame, and an overlapping part between the current acoustic feature frame and the next acoustic feature frame. At this time, assuming that the length of each acoustic feature frame in the second acoustic feature sequence after frame segmentation is nfft, and the overlapping length between two adjacent acoustic feature frames in the second acoustic feature sequence after frame segmentation is overlap, the length L of the second acoustic feature sequence after the above frame segmentation can be expressed by the following formula:

[0088] L = (nfft - overlap) × m

[0089] 302. Perform windowing processing on the second acoustic feature sequence after frame segmentation to obtain the windowed second acoustic feature sequence.

[0090] In the embodiment of the present invention, the initial window type, window length, and sliding step size can be set first to determine the window parameters to be added. Further, perform windowing processing on the second acoustic feature sequence after frame segmentation to obtain the windowed second acoustic feature sequence. Specifically, the window function y can be calculated according to the adaptive time scale rule window and the window length, and multiply the segmented acoustic feature data by the window function to obtain the windowed second acoustic feature sequence:

[0091] y i = y i × y window

[0092] where y i is the data of the i-th acoustic feature frame after segmentation, and y i is the corresponding i-th windowed acoustic feature frame.

[0093] 303. Perform fast Fourier transform point number processing on the windowed second acoustic feature sequence to convert the second acoustic feature sequence from time domain information to frequency domain information.

[0094] In an embodiment of the present invention, after obtaining the windowed second voiceprint sequence, the windowed second voiceprint sequence can be processed by the number of points of the Fast Fourier Transform (FFT). Through the processing of the number of points of the Fast Fourier Transform, the second voiceprint sequence is converted from time-domain information to frequency-domain information, and the frequency and amplitude information corresponding to each moment of the second voiceprint sequence are obtained, that is, the time-frequency characteristics of the second voiceprint sequence in the embodiment of the present invention are obtained. Specifically, first, the voiceprint data is processed with the initial window type, window length, sliding step, and the number of points of the Fast Fourier Transform. Among them, the above window type can select a window with better sidelobe suppression, the above sliding step can adopt a 100% sliding step, and the above number of points of the Fast Fourier Transform can select a smaller value. In this way, the voiceprint data can be processed quickly, and the processing speed of the voiceprint data can be improved. Then, according to the high and low frequencies of the voiceprint data, the window type, window length, sliding step, and the number of points of the Fast Fourier Transform are adjusted to meet the requirements of time-frequency analysis. Specifically, a window with a narrower main lobe width can be used instead, and at the same time, a 40% sliding step is adopted, and the number of points of the Fast Fourier Transform is increased for adjustment. After the above adjustment of the adaptive time scale, the window type, window length, sliding step, and the number of points of the Fast Fourier Transform that meet the requirements of time-frequency analysis are obtained. Then, the voiceprint data is processed by short-time Fourier transform according to the window type, window length, sliding step, and the number of points of the Fast Fourier Transform that meet the requirements of time-frequency analysis.

[0095] In an embodiment of the present invention, since the voiceprint sequence of the specified content implicitly includes the lung condition of the target person, the diseased lung cells (region of interest) will affect the normal state of the lungs, and this state can be amplified by the voice of the specified content and hidden in the voiceprint sequence after being captured by the sound sensor. Therefore, the voiceprint sequence of the target person can be used as an auxiliary detection of the region of interest.

[0096] 102. Feature extraction is performed on the lung pathological image through a preset image feature extraction network to obtain a lung pathological feature map.

[0097] In an embodiment of the present invention, the above preset image feature extraction network can be understood as a pre-trained image feature extraction network, and the above preset image feature extraction network can be constructed based on a convolutional neural network. If the lung pathological image is a continuous multi-frame image, then the above image feature extraction network may include a three-dimensional convolution module for extracting spatio-temporal features in the lung pathological image.

[0098] Optionally, in an embodiment of the present invention, the above preset image feature extraction network may include a preset global feature extraction network and a preset local feature extraction network, and the above lung pathological feature map includes a lung pathological global feature map and a local pathological feature map. Please refer to Figure 4 ,Figure 4 This is a flowchart for extracting a lung pathological feature map provided by an embodiment of the present invention. As Figure 4 shown, it includes the following steps:

[0099] 401. Input the lung pathological image into a preset global feature extraction network to obtain a lung pathological global feature map.

[0100] In the embodiment of the present invention, the above-mentioned preset global feature extraction network can be understood as a pre-trained feature extraction network. The above-mentioned global feature extraction network can be constructed based on a deep convolutional network or a residual neural network, and feature maps with different scale resolutions can be extracted at different depths. For example, low-level feature maps with a larger scale resolution are extracted in the shallow network, medium-level feature maps with a medium scale resolution are extracted in the middle network, and high-level feature maps with a smaller scale resolution are extracted in the deep network.

[0101] Specifically, the first global feature map, the second global feature map, and the third global feature map can be sequentially extracted according to the different depths of the preset global feature extraction network. Among them, the scale resolution of the above-mentioned first global feature map is greater than that of the second global feature map, and the scale resolution of the second global feature map is greater than that of the third global feature map.

[0102] For example, the scale resolution of the lung pathological image is 1024×1024. The scale resolution of the first global feature map can be 512×512, the scale resolution of the second global feature map can be 256×256, and the scale resolution of the third global feature map can be 64×64.

[0103] It should be noted that the first global feature map is a shallow feature map and can retain relatively rich local details. The second global feature map is a middle feature map. Compared with the shallow feature map, it is more advanced in terms of semantics than the shallow feature map, but the local details retained are less than those of the first global feature. The third global feature is a deep feature map, which is an abstract expression of high-level semantics but has very few local details. In a possible embodiment, the above-mentioned global feature extraction network can be constructed based on a residual neural network, and the first global feature map, the second global feature, and the third global feature can all contain the residuals of their upper layer, thereby retaining more local details.

[0104] 402. Randomly slice the lung pathological image to obtain a number of lung pathological local images, and sequentially input the number of lung pathological local images into a preset local feature extraction network to obtain local pathological feature maps.

[0105] In the embodiment of the present invention, the above-mentioned random slicing can be that the lung pathological image is randomly sliced into the original Figure 1Local images with a scale resolution of 1 / 2. For example, the scale resolution of a lung pathological image is 1024×1024, and several local images of 512×512 are randomly sliced. Generally speaking, the region of interest is often located in the local area of the lung pathological image. When extracting features from the lung pathological image, some information will be lost during the downsampling process. Therefore, in the embodiments of the present invention, by randomly slicing the lung pathological image, more abundant local details can be retained on the basis of the information of the original lung pathological image, thereby improving the detection accuracy of the region of interest.

[0106] Through a preset local feature extraction network, a local feature map with the same scale resolution as the first global feature map can be extracted. Taking the above example for illustration, a local feature map with a scale resolution of 512×512 can be obtained, and there is no information loss due to downsampling.

[0107] 103. Extract features from the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map.

[0108] In the embodiments of the present invention, the above-mentioned preset voiceprint feature extraction network can be understood as a pre-trained voiceprint feature extraction network, and the above-mentioned preset voiceprint feature extraction network can be constructed based on a convolutional neural network or a residual neural network. The above-mentioned voiceprint feature extraction network may include a three-dimensional convolution module for extracting the time-frequency voiceprint features of the voiceprint sequence. In the voiceprint feature extraction network, feature maps with different scale resolutions can be extracted according to different depths. For example, low-level feature maps with a relatively large scale resolution are extracted in the shallow network, medium-level feature maps with a medium scale resolution are extracted in the middle network, and high-level feature maps with a relatively small scale resolution are extracted in the deep network.

[0109] Specifically, the first voiceprint feature map, the second voiceprint feature map, and the third voiceprint feature map can be extracted in sequence according to the different depths of the preset voiceprint feature extraction network. Among them, the scale resolution of the above-mentioned first voiceprint feature map is greater than that of the above-mentioned second voiceprint feature map, and the scale resolution of the above-mentioned second voiceprint feature map is greater than that of the above-mentioned third voiceprint feature map.

[0110] For example, by linearly encoding the voiceprint sequence, the voiceprint sequence can be encoded into an input map with a scale resolution of 1024×1024. The scale resolution of the first voiceprint feature map can be 512×512, the scale resolution of the second voiceprint feature map can be 256×256, and the scale resolution of the third voiceprint feature map can be 64×64.

[0111] It should be noted that the first voiceprint feature map is a shallow feature map, which can retain relatively rich local details. The second voiceprint feature map is a middle-layer feature map. Compared with the shallow feature map, it is more advanced in semantics than the shallow feature map, but the local details retained are less than those of the first voiceprint feature. The third voiceprint feature is a deep feature map, which is an abstract expression of high-level semantics but has very few local details. In a possible embodiment, the above voiceprint feature extraction network can be constructed based on a residual neural network. The above first voiceprint feature map, second voiceprint feature, and third voiceprint feature can all contain the residuals of their upper layer, so that more local details can be retained.

[0112] 104. Decode the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network, and use the decoding result as the processing result of the lung pathological image.

[0113] In the embodiment of the present invention, the processing result of the above lung pathological image includes the region of interest. The above decoding includes an upsampling process and a feature fusion process. The above upsampling can be upsampling based on deconvolution or upsampling based on interpolation.

[0114] Optionally, please refer to Figure 5 , Figure 5 is the overall schematic diagram of a lung pathological image processing system provided by an embodiment of the present invention. As Figure 5 shown, the above third voiceprint feature map and the above third global feature map can be fused through a preset first fusion method to obtain a first fusion feature map. The above first fusion method can be fusion through a convolution method. For example, through a 1×1 convolution method, the third voiceprint feature map and the above third global feature map can be fused. The scale resolution of the first fusion feature map is 64×64.

[0115] Perform first upsampling on the above first fusion feature map to upsample the above first fusion feature map to the scale resolution of the above second global feature map to obtain a first upsampled feature map. The above first upsampling can be 4-fold upsampling, and the first fusion feature map with a scale resolution of 64×64 is upsampled to a second upsampled feature map with a scale resolution of 512×512.

[0116] Fuse the above second voiceprint feature map, the above second global feature map, and the above first upsampled feature map through a preset second fusion method to obtain a second fusion feature map. The above second fusion method can be fusion through a convolution method. For example, through a 3×3 convolution method, the above second voiceprint feature map, the above second global feature map, and the above first upsampled feature map can be fused. The scale resolution of the second fusion feature map is 256×256.

[0117] Perform a second upsampling on the above-mentioned second fused feature map to upsample the second fused feature map to the scale resolution of the first global feature map, obtaining a second upsampled feature map. The above-mentioned second upsampling can be a 2-fold upsampling, upsampling the second fused feature map with a scale resolution of 256×256 to a second upsampled feature map with a scale resolution of 512×512.

[0118] Fuse the above-mentioned first voiceprint feature map, the first global feature map, and the second upsampled feature map through a preset third fusion method to obtain a third fused feature map. The above-mentioned third fusion method can be fusion through a convolution method. For example, through a 3×3 convolution method, fuse the above-mentioned first voiceprint feature map, the third global feature map, and the second upsampled feature map. The scale resolution of the third fused feature map is 512×512.

[0119] Perform a third upsampling on the above-mentioned third fused feature map to upsample the third fused feature map to the scale resolution of the lung pathological image, obtaining a decoding result as the processing result of the lung pathological image. The processing result of the lung pathological image includes the region of interest. The above-mentioned third upsampling can be a 2-fold upsampling, upsampling the third fused feature map with a scale resolution of 512×512 to a decoding result with a scale resolution of 1024×1024.

[0120] In an embodiment of the present invention, a lung pathological image and a voiceprint sequence of a target person are obtained. The voiceprint sequence is collected when the target person vocalizes according to a specified content. The lung pathological image includes a region of interest, and the voiceprint sequence includes time-domain information and frequency-domain information; feature extraction is performed on the lung pathological image through a preset image feature extraction network to obtain a lung pathological feature map; feature extraction is performed on the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map; decoding is performed on the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network to obtain a decoding result as the processing result of the lung pathological image. The processing result of the lung pathological image includes the region of interest. It is possible to process the lung pathological image by using the lung pathological image of the target person and the voiceprint sequence when the target person vocalizes according to the specified content, thereby detecting the region of interest in the lung pathological image. Since the voiceprint is added as an auxiliary detection, the accuracy of detecting the region of interest in the pathological image is improved, and there is no need for manual observation of the pathological image, thereby improving the reading efficiency of the pathological image.

[0121] It should be noted that the method for processing a lung pathological image provided in the embodiment of the present invention can be applied to devices such as mobile phones, monitors, computers, and servers that can process lung pathological images.

[0122] Optionally, see Figure 6 , Figure 6 which is a training flow chart of a processing model for pulmonary pathological images provided by an embodiment of the present invention. The above-mentioned processing model for pulmonary pathological images includes an image feature extraction network, a voiceprint feature extraction network, and a decoding network. The output of the image feature extraction network is connected to the input of the decoding network, and the output of the voiceprint feature extraction network is connected to the input of the decoding network. As Figure 4 shown, it includes the following steps:

[0123] 601. Construct a training data set.

[0124] In the embodiment of the present invention, the above-mentioned training data set includes sample pulmonary pathological images, sample voiceprint sequences, and corresponding region-of-interest annotation data. The above-mentioned sample pulmonary pathological images are pulmonary pathological images of sample personnel, the above-mentioned region-of-interest annotation data are annotated regions of the pulmonary pathological images of sample personnel, and the above-mentioned sample voiceprint sequences are voiceprint sequences collected when sample personnel vocalize according to specified content. The above-mentioned sample voiceprint sequences include time-domain information and frequency-domain information.

[0125] The above-mentioned sample voiceprint sequences include time-domain information and frequency-domain information. Specifically, the above-mentioned sample voiceprint sequences can be obtained through the Figure 2 method of the embodiment, and the frequency-domain information in the above-mentioned sample engine voiceprint sequences can be obtained through the Figure 3 conversion method of the embodiment.

[0126] 602. Through the training data set, jointly train the image feature extraction network, the voiceprint feature extraction network, and the decoding network.

[0127] In the embodiment of the present invention, specifically, the above-mentioned sample pulmonary pathological images can be used by the image feature extraction network to be trained to extract features to obtain sample spatio-temporal features.

[0128] The above-mentioned image feature extraction network to be trained includes a global feature extraction network and a local feature extraction network. The sample pulmonary pathological images are input into the image feature extraction network to obtain a sample pulmonary pathological global feature map. The sample pulmonary pathological images are randomly sliced to obtain a number of sample pulmonary pathological local images, and the number of sample pulmonary pathological local images are sequentially input into the local feature extraction network to obtain a sample local pathological feature map.

[0129] Optionally, the network parameters of the image feature extraction network, the voiceprint feature extraction network, and the decoding network can be adjusted by calculating the error loss of the image feature extraction network and performing backpropagation.

[0130] Further optionally, the network parameters of the image feature extraction network, the voiceprint feature extraction network, and the decoding network can be adjusted by calculating the error loss of the local feature extraction network and performing backpropagation. Since the local feature extraction network uses random slicing to obtain a large number of sample local lung pathological images, even if the number of sample lung pathological images is small, a large number of sample local lung pathological images can be obtained. Therefore, the error loss provided by the local feature extraction network can more quickly fit the network parameters during the gradient descent process, thereby improving the training speed. Moreover, it is equivalent to increasing the number of samples, which can improve the accuracy of the trained image feature extraction network, voiceprint feature extraction network, and decoding network.

[0131] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a processing device for lung pathological images provided by an embodiment of the present invention. As Figure 7 shown, the device includes:

[0132] An acquisition module 701, configured to acquire a lung pathological image and a voiceprint sequence of a target person. The voiceprint sequence is acquired when the target person makes a sound according to a specified content. The lung pathological image includes a region of interest, and the voiceprint sequence includes time domain information and frequency domain information;

[0133] A first extraction module 702, configured to extract features from the lung pathological image through a preset image feature extraction network to obtain a lung pathological feature map;

[0134] A second extraction module 703, configured to extract features from the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map;

[0135] A decoding module 704, configured to decode the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network to obtain a decoding result as the processing result of the lung pathological image. The processing result of the lung pathological image includes a region of interest.

[0136] Optionally, as Figure 8 shown, the preset image feature extraction network includes a preset global feature extraction network and a preset local feature extraction network. The lung pathological feature map includes a lung pathological global feature map and a local pathological feature map. The first extraction module 702 includes:

[0137] A global extraction sub-module 7021, configured to input the lung pathological image into the preset global feature extraction network to obtain a lung pathological global feature map;

[0138] The local extraction sub-module 7022 is used to randomly slice the lung pathological image to obtain a plurality of local lung pathological images, and sequentially input the plurality of local lung pathological images into the preset local feature extraction network to obtain local pathological feature maps.

[0139] Optionally, the global extraction sub-module 7021 is further configured to sequentially extract a first global feature map, a second global feature map, and a third global feature map at different depths of the preset global feature extraction network;

[0140] The scale resolution of the first global feature map is greater than that of the second global feature map, and the scale resolution of the second global feature map is greater than that of the third global feature image.

[0141] Optionally, the second extraction module 703 is further configured to sequentially extract a first voiceprint feature map, a second voiceprint feature map, and a third voiceprint feature map at different depths of the preset voiceprint feature extraction network;

[0142] The scale resolution of the first voiceprint feature map is the same as that of the first global feature map, the scale resolution of the second voiceprint feature map is the same as that of the second global feature map, and the scale resolution of the third voiceprint feature map is the same as that of the third global feature map.

[0143] Optionally, as Figure 9 shown, the scale resolution of the local pathological feature map is the same as that of the first global feature map, and the decoding module 704 includes:

[0144] The first integration sub-module 7041 is used to fuse the third voiceprint feature map and the third global feature map through a preset first fusion method to obtain a first fusion feature map;

[0145] The first upsampling module 7042 is used to perform first upsampling on the first fusion feature map to upsample the first fusion feature map to the scale resolution of the second global feature map to obtain a first upsampled feature map;

[0146] The second fusion sub-module 7043 is used to fuse the second voiceprint feature map, the second global feature map, and the first upsampled feature map through a preset second fusion method to obtain a second fusion feature map;

[0147] The second upsampling module 7044 is used to perform second upsampling on the second fusion feature map to upsample the second fusion feature map to the scale resolution of the first global feature map to obtain a second upsampled feature map;

[0148] The third fusion sub-module 7045 is configured to fuse the first voiceprint feature map, the first global feature map, and the second upsampled feature map through a preset third fusion method to obtain a third fusion feature map;

[0149] The third upsampling sub-module 7046 is configured to perform third upsampling on the third fusion feature map to upsample the third fusion feature map to the scale resolution of the lung pathological image, and obtain a decoding result as the processing result of the lung pathological image, where the processing result of the lung pathological image includes a region of interest.

[0150] Optionally, as Figure 10 shown, the obtaining module 701 includes:

[0151] An obtaining sub-module 7011, configured to obtain a first voiceprint sequence collected when a target person makes a sound according to specified content;

[0152] A noise reduction sub-module 7012, configured to perform noise reduction on the first voiceprint sequence to obtain a second voiceprint sequence;

[0153] A conversion sub-module 7013, configured to convert the second voiceprint sequence from time-domain information to frequency-domain information, and obtain the voiceprint sequence of the target person according to the time-domain information and the frequency-domain information of the second voiceprint sequence.

[0154] Optionally, as Figure 11 shown, the conversion sub-module 7013 includes:

[0155] A framing unit 70131, configured to perform framing processing on the second voiceprint sequence to obtain a framed second voiceprint sequence;

[0156] A windowing unit 70132, configured to perform windowing processing on the framed second voiceprint sequence to obtain a windowed second voiceprint sequence;

[0157] A conversion unit 70133, configured to perform fast Fourier transform point number processing on the windowed second voiceprint sequence to convert the second voiceprint sequence from time-domain information to frequency-domain information.

[0158] Optionally, as Figure 12 shown, the image feature extraction network, the voiceprint feature extraction network, and the decoding network are trained through the same data set, and the device further includes:

[0159] A construction module 705 for constructing a training data set, where the training data set includes sample lung pathological images, sample voiceprint sequences, and corresponding region-of-interest annotation data. The sample lung pathological images are lung pathological images of sample personnel, the region-of-interest annotation data are annotated regions of the lung pathological images of the sample personnel, and the sample voiceprint sequences are voiceprint sequences collected when the sample personnel vocalize according to specified content;

[0160] A training module 706 for jointly training the image feature extraction network, the voiceprint feature extraction network, and the decoding network through the training data set.

[0161] Optionally, the training module 706 is further configured to adjust the network parameters of the image feature extraction network, the voiceprint feature extraction network, and the decoding network by performing backpropagation through calculating the error loss of the image feature extraction network.

[0162] It should be noted that the processing device for lung pathological images provided in the embodiments of the present invention can be applied to devices such as mobile phones, monitors, computers, and servers that can process lung pathological images.

[0163] The processing device for lung pathological images provided in the embodiments of the present invention can implement each process implemented by the processing method of lung pathological images in the above method embodiments and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.

[0164] See Figure 13 , Figure 13 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. As Figure 13 shown, it includes: a memory 1302, a processor 1301, and a computer program stored on the memory 1302 and executable on the processor 1301, where:

[0165] The processor 1301 is configured to call the computer program stored in the memory 1302 and execute the following steps:

[0166] Obtain a lung pathological image and a voiceprint sequence of a target person. The voiceprint sequence is collected when the target person vocalizes according to specified content. The lung pathological image includes a region of interest, and the voiceprint sequence includes time domain information and frequency domain information;

[0167] Perform feature extraction on the lung pathological image through a preset image feature extraction network to obtain a lung pathological feature map;

[0168] Perform feature extraction on the voiceprint sequence through a preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map;

[0169] Decode the lung pathological feature map and the time-frequency voiceprint feature map through a preset decoding network, and obtain a decoding result as the processing result of the lung pathological image. The processing result of the lung pathological image includes a region of interest.

[0170] Optionally, the preset image feature extraction network includes a preset global feature extraction network and a preset local feature extraction network. The lung pathological feature map includes a lung pathological global feature map and a local pathological feature map. The step of the processor 1301 performing feature extraction on the lung pathological image through the preset image feature extraction network to obtain a lung pathological feature map includes:

[0171] Input the lung pathological image into the preset global feature extraction network to obtain a lung pathological global feature map;

[0172] Randomly slice the lung pathological image to obtain a plurality of lung pathological local images, and sequentially input the plurality of lung pathological local images into the preset local feature extraction network to obtain a local pathological feature map.

[0173] Optionally, the step of the processor 1301 performing inputting the lung pathological image into the global feature extraction network to obtain a lung pathological global feature map includes:

[0174] Sequentially extract a first global feature map, a second global feature map, and a third global feature map at different depths of the preset global feature extraction network;

[0175] The scale resolution of the first global feature map is greater than that of the second global feature map, and the scale resolution of the second global feature map is greater than that of the third global feature image.

[0176] Optionally, the step of the processor 1301 performing feature extraction on the voiceprint sequence through the preset voiceprint feature extraction network to obtain a time-frequency voiceprint feature map includes:

[0177] Sequentially extract a first voiceprint feature map, a second voiceprint feature map, and a third voiceprint feature map at different depths of the preset voiceprint feature extraction network;

[0178] The scale resolution of the first voiceprint feature map is the same as that of the first global feature map, the scale resolution of the second voiceprint feature map is the same as that of the second global feature map, and the scale resolution of the third voiceprint feature map is the same as that of the third global feature map.

[0179] Optionally, the scale resolution of the local pathological feature map is the same as that of the first global feature map. The processor 1301 executes decoding the pulmonary pathological feature map and the time-frequency voiceprint feature map through a preset decoding network, and obtaining a decoding result as the processing result of the pulmonary pathological image, including:

[0180] Fusing the third voiceprint feature map and the third global feature map through a preset first fusion method to obtain a first fused feature map;

[0181] Performing a first upsampling on the first fused feature map to upsample the first fused feature map to the scale resolution of the second global feature map, obtaining a first upsampled feature map;

[0182] Fusing the second voiceprint feature map, the second global feature map, and the first upsampled feature map through a preset second fusion method to obtain a second fused feature map;

[0183] Performing a second upsampling on the second fused feature map to upsample the second fused feature map to the scale resolution of the first global feature map, obtaining a second upsampled feature map;

[0184] Fusing the first voiceprint feature map, the first global feature map, and the second upsampled feature map through a preset third fusion method to obtain a third fused feature map;

[0185] Performing a third upsampling on the third fused feature map to upsample the third fused feature map to the scale resolution of the pulmonary pathological image, obtaining a decoding result as the processing result of the pulmonary pathological image, and the processing result of the pulmonary pathological image includes a region of interest.

[0186] Optionally, the processor 1301 executes obtaining the voiceprint sequence of the target person, including:

[0187] Obtaining a first voiceprint sequence collected when the target person makes a sound according to a specified content;

[0188] Denosing the first voiceprint sequence to obtain a second voiceprint sequence;

[0189] Converting the second voiceprint sequence from time-domain information to frequency-domain information, and obtaining the voiceprint sequence of the target person according to the time-domain information and the frequency-domain information of the second voiceprint sequence.

[0190] Optionally, the processor 1301 executes converting the second voiceprint sequence from time-domain information to frequency-domain information, including:

[0191] Perform frame processing on the second voiceprint sequence to obtain the second voiceprint sequence after frame processing;

[0192] Perform windowing processing on the second voiceprint sequence after frame processing to obtain the second voiceprint sequence after windowing;

[0193] Perform fast Fourier transform point number processing on the second voiceprint sequence after windowing to convert the second voiceprint sequence from time-domain information to frequency-domain information.

[0194] Optionally, the image feature extraction network, the voiceprint feature extraction network, and the decoding network are trained through the same data set, and the processor 1301 further executes including:

[0195] Construct a training data set, the training data set includes sample lung pathological images, sample voiceprint sequences, and corresponding region-of-interest annotation data, the sample lung pathological images are lung pathological images of sample personnel, the region-of-interest annotation data is the annotated region of the lung pathological images of sample personnel, and the sample voiceprint sequences are voiceprint sequences collected when sample personnel vocalize according to specified content;

[0196] Through the training data set, jointly train the image feature extraction network, the voiceprint feature extraction network, and the decoding network.

[0197] Optionally, the processor 1301 executes the joint training of the image feature extraction network, the voiceprint feature extraction network, and the decoding network through the training data set, including:

[0198] Adjust the network parameters of the image feature extraction network, the voiceprint feature extraction network, and the decoding network by performing backpropagation by calculating the error loss of the image feature extraction network.

[0199] It should be noted that the above electronic device can be a mobile phone, a monitor, a computer, a server, etc. that can be applied to the processing of lung pathological images.

[0200] The electronic device provided by the embodiment of the present invention can implement each process implemented by the processing method of the lung pathological image in the above method embodiment, and can achieve the same beneficial effects. To avoid repetition, it will not be elaborated here.

[0201] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the processing method of the lung pathological image provided by the embodiment of the present invention, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0202] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0203] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.

Claims

1. A method for processing pulmonary pathological images, characterized in that, Including the following steps: Obtain the pulmonary pathological image and voiceprint sequence of the target person. The voiceprint sequence is collected when the target person vocalizes according to the specified content. The pulmonary pathological image includes the region of interest. The voiceprint sequence includes time-domain information and frequency-domain information. The time interval between obtaining the pulmonary pathological image and the voiceprint sequence does not exceed a preset time threshold. The number of frames of the pulmonary pathological image sequence is the same as that of the voiceprint sequence, and the time interval between two adjacent frames in the pulmonary pathological image sequence is the same as that between two adjacent frames in the voiceprint sequence; Extract the first global feature map, the second global feature map, and the third global feature map successively at different depths of a preset global feature extraction network. The scale resolution of the first global feature map is greater than that of the second global feature map, and the scale resolution of the second global feature map is greater than that of the third global feature image. Randomly slice the pulmonary pathological image to obtain a number of local pulmonary pathological images, and input the number of local pulmonary pathological images into a preset local feature extraction network successively to obtain a local pathological feature map. The scale resolution of the local pathological feature map is the same as that of the first global feature map; Extract the first voiceprint feature map, the second voiceprint feature map, and the third voiceprint feature map successively at different depths of a preset voiceprint feature extraction network. The scale resolution of the first voiceprint feature map is the same as that of the first global feature map, the scale resolution of the second voiceprint feature map is the same as that of the second global feature map, and the scale resolution of the third voiceprint feature map is the same as that of the third global feature map; Fuse the third voiceprint feature map and the third global feature map through a preset first fusion method to obtain a first fusion feature map; Perform a first upsampling on the first fusion feature map to upsample the first fusion feature map to the scale resolution of the second global feature map to obtain a first upsampled feature map. Fuse the second voiceprint feature map, the second global feature map, and the first upsampled feature map through a preset second fusion method to obtain a second fusion feature map; Perform a second upsampling on the second fusion feature map to upsample the second fusion feature map to the scale resolution of the first global feature map to obtain a second upsampled feature map; Fuse the first voiceprint feature map, the first global feature map, and the second upsampled feature map through a preset third fusion method to obtain a third fusion feature map; Perform a third upsampling on the third fusion feature map to upsample the third fusion feature map to the scale resolution of the pulmonary pathological image to obtain a decoding result as the processing result of the pulmonary pathological image. The processing result of the pulmonary pathological image includes the region of interest.

2. The method according to claim 1, characterized in that, The obtaining of the voiceprint sequence of the target person includes: Obtain the first voiceprint sequence collected when the target person vocalizes according to the specified content; Denoise the first voiceprint sequence to obtain a second voiceprint sequence; Convert the second voiceprint sequence from time-domain information to frequency-domain information, and obtain the voiceprint sequence of the target person based on the time-domain information and the frequency-domain information of the second voiceprint sequence.

3. The method according to claim 2, characterized in that, The conversion of the second voiceprint sequence from time-domain information to frequency-domain information includes: Perform frame segmentation on the second voiceprint sequence to obtain the second voiceprint sequence after frame segmentation; Perform windowing on the second voiceprint sequence after frame segmentation to obtain the second voiceprint sequence after windowing; Perform fast Fourier transform point number processing on the second voiceprint sequence after windowing to convert the second voiceprint sequence from time-domain information to frequency-domain information.

4. The method according to claim 1, characterized in that, The image feature extraction network, the voiceprint feature extraction network, and the decoding network are trained through the same data set. The steps of the training include: Construct a training data set, which includes sample lung pathological images, sample voiceprint sequences, and corresponding region of interest annotation data. The sample lung pathological images are the lung pathological images of sample persons, the region of interest annotation data are the annotated regions of the lung pathological images of sample persons, and the sample voiceprint sequences are the voiceprint sequences collected when sample persons vocalize according to specified content; Through the training data set, jointly train the image feature extraction network, the voiceprint feature extraction network, and the decoding network.

5. The method according to claim 4, characterized in that, The joint training of the image feature extraction network, the voiceprint feature extraction network, and the decoding network through the training data set includes: Adjust the network parameters of the image feature extraction network, the voiceprint feature extraction network, and the decoding network by performing backpropagation by calculating the error loss of the image feature extraction network.

6. A device for processing pulmonary pathological images, characterized in that, The device includes: An acquisition module, configured to acquire the lung pathological image and the voiceprint sequence of the target person. The voiceprint sequence is collected when the target person vocalizes according to the specified content. The lung pathological image includes a region of interest. The voiceprint sequence includes time-domain information and frequency-domain information. The time interval between the acquisition of the lung pathological image and the voiceprint sequence does not exceed a preset time threshold. The number of frames of the lung pathological image sequence is the same as the number of frames of the voiceprint sequence, and the time interval between two adjacent frames in the lung pathological image sequence is the same as the time interval between two adjacent frames in the voiceprint sequence; A first extraction module, configured to sequentially extract a first global feature map, a second global feature map, and a third global feature map at different depths of a preset global feature extraction network; the scale resolution of the first global feature map is greater than the scale resolution of the second global feature map, and the scale resolution of the second global feature map is greater than the scale resolution of the third global feature image; randomly slice the lung pathological image to obtain a plurality of local lung pathological images, and sequentially input the plurality of local lung pathological images into a preset local feature extraction network to obtain local pathological feature maps, and the scale resolution of the local pathological feature maps is the same as the scale resolution of the first global feature map; A second extraction module, configured to sequentially extract a first voiceprint feature map, a second voiceprint feature map, and a third voiceprint feature map at different depths of a preset voiceprint feature extraction network; the scale resolution of the first voiceprint feature map is the same as that of the first global feature map, the scale resolution of the second voiceprint feature map is the same as that of the second global feature map, and the scale resolution of the third voiceprint feature map is the same as that of the third global feature map; A decoding module, configured to fuse the third voiceprint feature map and the third global feature map through a preset first fusion method to obtain a first fused feature map; perform a first upsampling on the first fused feature map to upsample the first fused feature map to the scale resolution of the second global feature map to obtain a first upsampled feature map; fuse the second voiceprint feature map, the second global feature map, and the first upsampled feature map through a preset second fusion method to obtain a second fused feature map; perform a second upsampling on the second fused feature map to upsample the second fused feature map to the scale resolution of the first global feature map to obtain a second upsampled feature map; fuse the first voiceprint feature map, the first global feature map, and the second upsampled feature map through a preset third fusion method to obtain a third fused feature map; perform a third upsampling on the third fused feature map to upsample the third fused feature map to the scale resolution of the lung pathological image to obtain a decoding result as the processing result of the lung pathological image, and the processing result of the lung pathological image includes a region of interest.

7. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the processing method of the lung pathological image according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the processing method of the lung pathological image according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Electronic device, identification method based on face image and voiceprint information, and storage medium

    CN108446674A

  • Image recognition method and related device

    CN111126258A