Scanning image processing method and device, equipment, storage medium and product
By introducing a combination of visual feature extractor and Transformer decoder in the scan image processing, the spatial and temporal correlation of visual features in the scan image is solved, and the problem of insufficient image processing accuracy in traditional methods is achieved, and higher processing accuracy is achieved.
Patent Information
- Application Number
- CN202510585581.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
In the prior art, the processing methods of scanning images obtained by scanning probes are mostly based on traditional image processing techniques, and ignore the deeper correlation between various features of the scanning probe system during the scanning process, resulting in insufficient image processing accuracy.
A scanning image processing method is proposed, and the scanning image is extracted through a feature extraction module. The module includes a visual feature extractor and a Transformer decoder. The latter uses a self-attention mechanism to capture the spatial and temporal correlation between visual features and generate a second feature sequence.
By fully exploring the spatial and temporal correlation in visual features, the accuracy of scanning image processing tasks is improved, and the demand for image processing accuracy is met.
Smart Images

Figure CN120107618A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a scanned image processing method, device, equipment, storage medium and product. Background Art
[0002] Among the related technologies, scanning probe microscopy (SPM) technology is widely used in materials science, biology, chemistry and other fields, especially in the study of surface morphology and material properties at the nanoscale.
[0003] However, the current processing methods for scanning images obtained by scanning probes are mostly based on traditional image processing technology, which usually relies on a single feature in a single frame image or image sequence, while ignoring the deeper correlation between various features in the scanning probe system during the scanning process. As a result, the accuracy of related tasks in scanning images in related technologies cannot meet the requirements.
[0004] Therefore, how to improve the image processing accuracy of scanned images is a problem that needs to be solved urgently. Summary of the invention
[0005] The main purpose of this application is to provide a scanned image processing method, device, equipment, storage medium and product, aiming to solve the technical problem of how to improve the image processing accuracy of scanned images.
[0006] To achieve the above purpose, the present application proposes a scanned image processing method, which includes: Acquire a scanned image; the scanned image includes a plurality of scanned segments obtained by sequential scanning; The feature extraction module of the scanned image processing model is used to extract features from the scanned image; the feature extraction module includes a visual feature extractor and a Transformer decoder, the visual feature extractor is used to extract visual features from the scanned image to obtain a first feature sequence, the first feature sequence includes visual features corresponding to each of the plurality of scan segments; the Transformer decoder is used to capture the spatiotemporal correlation between the visual features in the first feature sequence to obtain a second feature sequence; The task processing module of the scanning image processing model performs preset task processing on the second feature sequence to obtain a scanning image processing result.
[0007] In some embodiments, before acquiring the scanned image, the scanned image processing method further includes: Inputting the sample scanned image into the initial feature extraction module to obtain a second feature sequence sample output by the initial feature extraction module; the sample scanned image includes a plurality of sample scan segments obtained by sequential scanning, and the initial feature extraction module includes the visual feature extractor and the Transformer decoder; the second feature sequence sample includes sample image features corresponding to the plurality of sample scan segments one by one, and the last sample image feature in the second feature sequence sample is a visual feature predicted for the next scan segment of the corresponding sample scan segment; Obtaining a predicted image according to the features of the last sample image; Calculating the loss between the predicted image and the real image data of the next scanning segment; Based on the loss, the parameters of the initial feature extraction module are adjusted to train the feature extraction module.
[0008] In some embodiments, calculating the loss between the predicted image and the real image data of the next scanning segment includes: The loss is calculated by expression 1; The expression 1 is: ; Wherein, L is the loss, is the predicted image, is the real image data.
[0009] In some embodiments, before acquiring the scanned image, the scanned image processing method further includes: The visual feature extractor is pre-trained using an image sample set related to the task application domain.
[0010] In some embodiments, the pre-training the visual feature extractor comprises: The visual feature extractor is pre-trained by at least one of an autoencoder model, contrastive learning, and a masked autoencoder.
[0011] In some embodiments, the image sample set is an unlabeled image sample set.
[0012] In some embodiments, the task processing module is a probe quality detection module; the task processing module through the scanning image processing model performs preset task processing on the second feature sequence to obtain a scanning image processing result, including: Performing pooling processing on the second feature sequence to obtain a one-dimensional feature vector sequence; The one-dimensional feature vector sequence is input into a classification model in the probe quality detection module, and based on the output of the classification model, a probe quality assessment result corresponding to the scanned image is determined.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a scanning image processing device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the scanning image processing method described above.
[0014] In addition, to achieve the above objectives, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the scanning image processing method described above are implemented.
[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the scanning image processing method described above are implemented.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: Since the scanned image has the characteristic that its pixel areas are scanned sequentially, the feature extraction module uses the self-attention mechanism of the decoder of the Transformer architecture to extract the correlation between visual features based on the visual features extracted by the visual feature extractor, so that the second feature sequence extracted by the feature extraction module can fully explore the spatiotemporal correlation in the visual features. When applied to downstream scanned image processing tasks, the accuracy of task processing can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A schematic diagram of a process flow of a scanned image processing method provided by an embodiment of the present application is shown; Figure 2 A schematic diagram of a scanning image acquisition trajectory provided by an exemplary embodiment of the present application is shown; Figure 3 A schematic diagram showing channel data contained in a scanned image provided by an exemplary embodiment of the present application is shown; Figure 4 A schematic diagram showing an abnormal image before and after processing provided by an exemplary embodiment of the present application is shown; Figure 5 A schematic diagram showing a missing region before and after processing provided by an exemplary embodiment of the present application is shown; Figure 6 A schematic diagram before and after normalization processing provided by an exemplary embodiment of the present application is shown; Figure 7 A schematic diagram showing a partial training process of a scanning image processing model provided by an exemplary embodiment of the present application is shown; Figure 8 A schematic diagram of the structure of a scanning image processing device provided in an embodiment of the present application is shown.
[0020] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0022] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0023] The main solution of the embodiment of the present application is: acquiring a scanned image; the scanned image includes multiple scan segments obtained by scanning in sequence; performing feature extraction on the scanned image through a feature extraction module of a scanned image processing model; the feature extraction module includes a visual feature extractor and a Transformer decoder, the visual feature extractor is used to extract visual features of the scanned image to obtain a first feature sequence, the first feature sequence includes visual features that correspond one to one to the multiple scan segments; the Transformer decoder is used to capture the spatiotemporal correlation between each visual feature in the first feature sequence to obtain a second feature sequence; performing preset task processing on the second feature sequence through a task processing module of the scanned image processing model to obtain a scanned image processing result.
[0024] Among the related technologies, scanning probe microscopy (SPM) technology is widely used in materials science, biology, chemistry and other fields, especially in the study of surface morphology and material properties at the nanoscale.
[0025] However, the current processing methods for scanning images obtained by scanning probes are mostly based on traditional image processing technology, which usually relies on a single feature in a single frame image or image sequence, while ignoring the deeper correlation between various features in the scanning probe system during the scanning process. As a result, the accuracy of related tasks in scanning images in related technologies cannot meet the requirements.
[0026] Therefore, how to improve the image processing accuracy of scanned images is a problem that needs to be solved urgently.
[0027] Based on this, the present application provides a solution, which enables in-depth extraction of the spatiotemporal correlation between each scan segment based on the original characteristics of the scanned image during model training, thereby achieving better efficiency and accuracy when processing downstream tasks with spatiotemporal dependencies.
[0028] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a scanning image processing device capable of realizing the above functions, etc. The following takes the scanning image processing device as an example to illustrate this embodiment and the following embodiments.
[0029] Reference Figure 1 , Figure 1 The flowchart of a scanning image processing method provided by an embodiment of the present application is shown. The scanning image processing method can be applied to a scanning image processing device, and includes the following steps S110 to S130: Step S110, obtaining a scanned image.
[0030] The image capture method generally includes photographic imaging and scanning imaging. When photographic imaging is used to capture an image, usually only one capture is required to obtain a complete image, and each pixel in the complete image is captured at the same time. When acquiring an image, the scanning imaging method requires scanning different areas of the captured object in a certain order, thereby gradually forming a scanned image. That is, the scanned image described in this embodiment refers to: an image obtained by scanning imaging.
[0031] For ease of understanding, the scanning probe system is used as an example to explain the scanning image processing method provided in this embodiment. The scanning probe system refers to a microscope system that scans a sample to be tested by a scanning probe to obtain a scanning image. Figure 2 As shown, when the scanning probe is actually scanning, it can be Figure 2 The scanning is performed in the order shown by the arrows in the figure, and finally a complete scanned image is obtained. That is, the scanned image may include multiple scanned segments obtained by scanning in sequence.
[0032] In this embodiment, the scanned image can be expressed as H×W×C, where H refers to the number of pixel rows of the scanned image, W refers to the number of pixel columns of the scanned image, and C refers to the number of channels corresponding to the scanned image. It should be noted that when the scanning probe scans the sample to be tested, the channels are different depending on the type of the scanning probe system, that is, the recorded information is different. As an example, the scanned image collected by the scanning tunneling microscope can contain data of three channels: height, current, and phase (for example Figure 3 In the process of collecting the scanned image, the scanning may be performed pixel by pixel, row by row, or column by column, which is not limited in this embodiment.
[0033] In this embodiment, each row of pixels may be defined as a scanning segment; in some other feasible implementations, each scanning segment may also be a part of a row of pixels or multiple rows of pixels. For ease of understanding, each row is used as a scanning segment for explanation below.
[0034] As mentioned above, since the scanning probe scans line by line, for the scanning probe system, each scanning segment has a strong temporal and spatial correlation with the scanning segments before and after it. This correlation is specifically based on three dimensions: (1) The sample to be tested itself reflects the correlation, which can be roughly understood as a common picture, and the parts of the object in the picture can be correlated; (2) There is a correlation between similar scanning results, because the scanning process is performed line by line, and the result of each line is closely related to the current state of the scanning probe system. (3) Usually, the scanning probe system records the information of multiple channels at the same time. The data between these channels is correlated, representing the state of the current system or the sample to be tested.
[0035] It is understandable that, in other scanning imaging technology fields except the scanning probe system, the scanned images acquired therefrom all have the characteristic of temporal and spatial correlation, which will not be elaborated in this embodiment.
[0036] Step S120, extracting features from the scanned image using a feature extraction module of the scanned image processing model.
[0037] Among them, the feature extraction module includes a visual feature extractor and a Transformer decoder.
[0038] The visual feature extractor is a module used to perform the first feature extraction on the scanned image, which is usually a multi-layer neural network, a convolutional neural network or a visual transformer. The input of the visual feature extractor is the scanned image. By performing visual feature extraction on the scanned image, the scanned image can be mapped from the dimension of the image data to the feature dimension, thereby obtaining the first feature sequence corresponding to the scanned image. In this embodiment, the first feature sequence can be a one-dimensional feature sequence.
[0039] The Transformer decoder is a module for further extracting the deep correlation between features. Specifically, the Transformer decoder can be an autoregressive model based on the Transformer architecture, which can be established based on a commonly used encoder-decoder model (encoder-decoder) or a pure decoder (decoder) model. In this embodiment, the pure decoder model is used as an example to explain the subsequent embodiments.
[0040] After receiving the first feature sequence input by the visual feature extractor, the Transformer decoder can combine the starting vector (START vector, used to mark the starting position of the sequence) to capture and fuse the correlation between each element through the self-attention mechanism built into the Transformer decoder architecture, thereby outputting a second feature sequence containing spatiotemporal correlations. In the second feature sequence, each element (i.e., feature) contains a spatiotemporal correlation with other elements.
[0041] It is worth mentioning that the feature extraction module can be obtained by pre-training. The feature extraction module can be trained uniformly as a whole, or the visual feature extractor can be pre-trained separately and then used as a part of the feature extraction module to train the entire feature extraction module as a whole.
[0042] For ease of understanding, the following explanation is given by taking the example of first pre-training the visual feature extractor separately and then training the entire feature extraction module as a whole.
[0043] Specifically, the visual feature extractor may be pre-trained separately using an image sample set related to the task application field, so as to ensure the accuracy of the visual feature extractor when performing feature extraction in the subsequent overall training process for the entire feature extraction module.
[0044] For the scanning probe scenario, a commonly used convolutional neural network (CNN) can be used to pre-train the CNN skeleton through the image data set of the scanning probe, thereby increasing the feature recognition ability of the corresponding CNN skeleton for the image data in the scanning probe field. In the process of training the visual feature extractor, some fine-tuning operations are also involved. The relevant parameters of the visual feature extractor can be fine-tuned based on one or more of the autoencoder model (Autoencoder), contrast learning, and masked autoencoder to ensure higher feature recognition capabilities. In other feasible implementations, a discrete variational autoencoder (discrete VAE) can also be used to train the visual feature extraction model on the aforementioned scanning probe image data set, which is not limited in this embodiment.
[0045] It should be noted that, in this embodiment, the image data set of the scanning probe is obtained by preprocessing the original sample set, that is, the scanned images in the original sample set have all undergone the process of removing abnormal images, removing incomplete parts, and standardizing and normalizing.
[0046] Specifically, Figure 4 As shown, for the original sample set, the maximum and minimum values of each scanning segment (for example, each row) in each original sample image in the original sample set can be calculated, and based on the difference between the maximum and minimum values and the mean square error, a comparison is performed through a preset difference threshold and a mean square error threshold. If the difference is less than the difference threshold, and the mean square error is less than the mean square error threshold, it can be considered that there is no abnormality in this scanning segment in the original sample image; if the difference is not less than the difference threshold, or the mean square error is not less than the mean square error threshold, it can be considered that there is an abnormality in this scanning segment in the original sample image. If the number of scanning segments in the original sample image that exceed the abnormality reaches a preset reference ratio, it is considered that the original sample image is an abnormal image, and the original sample image can be removed from the original sample set.
[0047] like Figure 5 As shown, for the original sample set, the incomplete parts of each original sample image in the original sample set can also be removed. For example, in actual application scenarios, there are some types of scanning probe systems that may only scan a part of the sample to be tested. In this case, the scanning probe system of this model may complete the remaining unscanned area (even if the scan is not actually performed). For the incomplete scanned image obtained in this case, the mean square error can be checked one by one according to the scan segment, and compared with the pre-set integrity threshold. If the mean square error of the scan segment is less than the integrity threshold, it can be considered that this scan segment is the completed data, and the scan segment can be removed from the original sample image.
[0048] like Figure 6 As shown, after completing Figure 4-5 After the processing shown, the processed original sample image can be standardized and normalized. The maximum and minimum values of the original sample image can be scaled to [0,1]. This can be obtained by subtracting the minimum value from all pixel values of the original sample image and then dividing it by the difference between the maximum and the minimum value. The pixel values of the entire original sample image can also be changed to conform to a distribution with a mean of 0 and a standard deviation of 1. Specifically, the mean and variance of all original sample images can be counted, and then the pixels of each original sample image can be subtracted from the mean and divided by the variance.
[0049] Thus, a scanning probe image data set (i.e., the image sample set described in this embodiment) that has been processed and meets the input conditions of the scanning image processing model is obtained. Based on this, pre-training the visual feature extractor can make the visual feature extractor more suitable for the characteristics of the scanning probe field, thereby extracting features more accurately.
[0050] In this embodiment, after the pre-training of the visual feature extractor is completed, the feature extraction module including the visual feature extractor and the Transformer decoder can be trained as a whole. It should be noted that in this embodiment, the Transformer decoder can also be connected to the reconstruction module.
[0051] Specifically, when processing an image, the decoder based on the Transformer architecture can predict the features of the n+1th token through the first n tokens in the input sequence, while capturing changes in the system state. When specifically applied to the scanning probe image processing field of this embodiment, the Transformer architecture can be pre-trained with a large amount of unlabeled data and can efficiently model the spatiotemporal correlation between features.
[0052] The autoregressive model based on the Transformer architecture has a self-attention mechanism. The self-attention mechanism refers to a mechanism that enables the model to automatically pay attention to the relationship between elements at different positions in the sequence when processing sequence data. In the actual processing of the self-attention mechanism, it includes two parts: calculating the attention score and calculating the attention weight. For each element in the sequence (the first feature sequence), it is used as a query (Query), and the similarity is calculated with the keys (Key) of all elements in the sequence (including itself) to obtain the attention score. The commonly used similarity calculation method is the dot product, and then the attention score is normalized by the softmax function to obtain the attention weight of each element. The attention weight is weighted and summed with the value (Value) of the corresponding element to obtain a new representation of the current element. This new representation integrates the information of other elements in the sequence and reflects the degree of association between the current element and other elements (for example, when applied in the field of scanning probes, such as the spatiotemporal correlation described in this embodiment).
[0053] It is understandable that in this embodiment, due to the existence of the self-attention mechanism, the sample data used in this embodiment can be a large amount of unlabeled data when pre-training the decoder, because the self-attention mechanism can automatically focus on the relationship between different elements without the need for manual labeling for training. Therefore, the decoder using the Transformer architecture can further reduce the process of manual labeling, reduce labor costs, and improve training efficiency.
[0054] While the self-attention mechanism captures the spatiotemporal correlation, the characteristics of the staggered output of the Transformer architecture enable the Transformer decoder to predict the next Token. In this embodiment, the visual feature extracted by the visual feature extractor for each scan segment is used as a Token, and the Transformer decoder can predict the next visual feature based on all the visual features that have appeared in the first feature sequence.
[0055] After obtaining the predicted next visual feature, the predicted visual feature can be input into a reconstruction module (reconstructor), and the reconstruction module can reconvert the visual feature into the scanning image space of the scanning probe to generate a predicted image. The network architecture of the reconstruction module can generally be a multi-layer neural network or a deconvolution network, which is not limited in this embodiment.
[0056] It is understandable that during the training process of the model, the accuracy of the model generation result can be known by comparing the difference between the predicted image and the real image, and the parameters in the model can be adjusted based on this. Specifically, in this embodiment, the loss between the predicted image and the real image data of the next scan segment can be calculated; based on the loss, the parameters in the initial feature extraction module (i.e., the untrained model architecture) are adjusted to train the feature extraction module.
[0057] In some embodiments, the mean square error loss function can be used Calculate the loss, where L is the loss, To predict the image, is the real image data.
[0058] In summary, the overall training process of the feature extraction module can be described as follows: Figure 7 As shown. Figure 7 In the example, the tokens output by the visual feature extractor can be grouped into a feature sequence (with a total of i+1 elements), represented from the 0th element to the ith element as: {S, h1, …, hi}, where S refers to the "START vector" used to mark the starting position 0 of the sequence for recognition by the Transformer decoder.
[0059] ④Task processing module ( Figure 7 (not shown) For the parts not described in detail here, please refer to the specific description in step S130.
[0060] Step S130, performing preset task processing on the second feature sequence through the task processing module of the scan image processing model to obtain a scan image processing result.
[0061] In this embodiment, the Transformer decoder can be connected to the task processing module. The task processing module is used to process downstream image processing tasks. The specific processing tasks involved are different, and the contents contained in the task processing module are also different. For example, the downstream image processing tasks may include the classification tasks of scanned images, the tasks of searching for specific content, the tasks of image quality detection, etc. The following are some typical examples to explain the application of the scanned image processing model provided in this embodiment in downstream tasks.
[0062] In some embodiments, when the scanned image processing model is applied to the image classification task, the scanned image obtained by the actual scan (assuming that it includes n scan segments) can be input into the scanned image processing model, and after the visual feature extractor and the Transformer decoder, a second feature vector sequence {Z1,…,Zn} is obtained. Then this feature sequence is converted into a one-dimensional vector sequence zavg=AvgPool({Z1,…,Zn}) through a pooling operation. In this example, the task processing module may include an image classification model fhead1 (usually a multi-layer fully connected network), and the one-dimensional vector sequence is input into the image classification model fhead1, so that the scanned image can be classified and the category of the scanned image can be output. Among them, fhead1 can also be obtained by pre-training with a pre-collected sample data set, and its classification task can be set by the staff, for example, it can be to judge the quality of the scanned image, etc., which is not limited in this embodiment. It can be understood that in the pre-training process, the backbone network (i.e., the feature extraction module) can be frozen, and only fhead1 can be trained; or all model parameters can be trained together during the training of the feature extraction module, which is not limited in this embodiment.
[0063] In other embodiments, when the scanning image processing model is applied to the probe quality detection task, the image portion that has been scanned can be input into the scanning image processing model, and the visual feature extractor can extract the following features: Figure 7 The feature sequence {S, h1, ..., hi} shown in the figure is then input into the Transformer decoder in sequence. The Transformer decoder processes h1 to hi to obtain the feature vector sequence Zi = fenc (h1, ..., hi). Similar to the previous example, Zi can be input into a classification model fhead2 that is also pre-trained to classify Zi. For example, the classification task is to judge whether the quality of the current latest row is good or bad.
[0064] It can be understood that the image classification task and the probe quality detection task belong to different dimensions. The image classification task belongs to the dimension of the entire image, and its input is a complete scanned image; while the probe quality detection task belongs to the scan segment dimension, and its input is an image being scanned, which may be only a part of the complete scanned image.
[0065] For other tasks such as object detection and segmentation, as an example, a decoder can be used as the backbone network, and a suitable decoding network can be added after the backbone network. For example, for the object detection model, the confidence of the detection box, coordinate regression and classification detection head are added, which will not be described in detail in this embodiment.
[0066] This embodiment provides a scanned image processing method. Since the scanned image has the characteristic that its pixel areas are scanned sequentially, the feature extraction module uses the self-attention mechanism of the decoder of the Transformer architecture to extract the correlation between the visual features based on the visual features extracted by the visual feature extractor, so that the second feature sequence extracted by the feature extraction module can fully explore the spatiotemporal correlation in the visual features. When applied to downstream scanned image processing tasks, the accuracy of task processing can be improved.
[0067] The present application provides a scanning image processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the scanning image processing method in the above-mentioned embodiment one.
[0068] Reference below Figure 8 , which shows a schematic diagram of the structure of a scanning image processing device suitable for implementing the embodiment of the present application. The scanning image processing device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The scanning image processing device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0069] like Figure 8As shown, the scanning image processing device 200 may include a processing device 210 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 220 or a program loaded from a storage device 230 to a random access memory (RAM: Random Access Memory) 240. In RAM240, various programs and data required for the operation of the scanning image processing device are also stored. The processing device 210, ROM220 and RAM240 are connected to each other through a bus 250. An input / output (I / O) interface 260 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 260: an input device 270 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 280 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 230 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 290. The communication device 290 can allow the scanning image processing device to communicate with other devices wirelessly or by wire to exchange data. Although the scanning image processing device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0070] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through a communication device, or installed from the storage device 230, or installed from the ROM 220. When the computer program is executed by the processing device 210, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0071] The scanning image processing device provided by the present application adopts the scanning image processing method in the above embodiment, which can solve the technical problem of how to improve the processing efficiency of the scanned image and improve the accuracy of image processing. Compared with the related art, the beneficial effects of the scanning image processing device provided by the present application are the same as the beneficial effects of the scanning image processing method provided by the above embodiment, and the other technical features in the scanning image processing device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0072] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0073] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0074] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the scan image processing method in the above-mentioned embodiment.
[0075] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0076] The computer-readable storage medium may be included in the scanning image processing device; or may exist independently without being assembled into the scanning image processing device.
[0077] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the scanning image processing device, the scanning image processing device can write computer program codes for performing the operations of the present application in one or more programming languages or a combination thereof. The programming languages include object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0078] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0079] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0080] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned scanned image processing method, and can solve the technical problem of how to improve the processing efficiency of scanned images and improve the accuracy of image processing. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the scanned image processing method provided by the above-mentioned embodiment, and will not be repeated here.
[0081] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned scanning image processing method when executed by a processor.
[0082] The computer program product provided by the present application can solve the technical problem of how to improve the processing efficiency of scanned images and improve the accuracy of image processing. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the scanned image processing method provided by the above embodiment, which will not be repeated here.
[0083] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A scanned image processing method, characterized in that: The scanned image processing method comprises: Acquire a scanned image; the scanned image includes a plurality of scanned segments obtained by sequential scanning; The feature extraction module of the scanned image processing model is used to extract features from the scanned image; the feature extraction module includes a visual feature extractor and a Transformer decoder, the visual feature extractor is used to extract visual features from the scanned image to obtain a first feature sequence, the first feature sequence includes visual features corresponding to each of the plurality of scan segments; the Transformer decoder is used to capture the spatiotemporal correlation between the visual features in the first feature sequence to obtain a second feature sequence; The task processing module of the scanning image processing model performs preset task processing on the second feature sequence to obtain a scanning image processing result.
2. The scanned image processing method according to claim 1, characterized in that: Before acquiring the scanned image, the scanned image processing method further includes: Inputting the sample scanned image into the initial feature extraction module to obtain a second feature sequence sample output by the initial feature extraction module; the sample scanned image includes a plurality of sample scan segments obtained by sequential scanning, and the initial feature extraction module includes the visual feature extractor and the Transformer decoder; the second feature sequence sample includes sample image features corresponding to the plurality of sample scan segments one by one, and the last sample image feature in the second feature sequence sample is a visual feature predicted for the next scan segment of the corresponding sample scan segment; Obtaining a predicted image according to the features of the last sample image; Calculating the loss between the predicted image and the real image data of the next scanning segment; Based on the loss, the parameters of the initial feature extraction module are adjusted to train the feature extraction module.
3. The scanned image processing method according to claim 2, characterized in that: The calculating the loss between the predicted image and the real image data of the next scanning segment comprises: The loss is calculated by expression 1; The expression 1 is: ; Wherein, L is the loss, is the predicted image, is the real image data.
4. The scanned image processing method according to claim 1, wherein: Before acquiring the scanned image, the scanned image processing method further includes: The visual feature extractor is pre-trained using an image sample set related to the task application domain.
5. The scanned image processing method according to claim 4, characterized in that: The pre-training of the visual feature extractor comprises: The visual feature extractor is pre-trained by at least one of an autoencoder model, contrastive learning, and a masked autoencoder.
6. The scanned image processing method according to claim 4, characterized in that: The image sample set is an unlabeled image sample set.
7. The scanned image processing method according to claim 1, characterized in that: The task processing module is a probe quality detection module; the task processing module of the scanning image processing model performs preset task processing on the second feature sequence to obtain a scanning image processing result, including: Performing pooling processing on the second feature sequence to obtain a one-dimensional feature vector sequence; The one-dimensional feature vector sequence is input into a classification model in the probe quality detection module, and based on the output of the classification model, a probe quality assessment result corresponding to the scanned image is determined.
8. A scanning image processing device, characterized in that: The scanned image processing device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the scanned image processing method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the scan image processing method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the scanned image processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Pulmonary nodule benign and malignant auxiliary diagnosis system based on computed tomography
CN113902702A
Retina optical coherence tomography image detection method, device and terminal
CN115205410A
Image processing methods, storage medium and computer terminal
WO2024146649A1