Three-dimensional image data bounding box prediction method and data transmission method and device
By projecting the three-dimensional image data into a two-dimensional image and using a prediction model to predict the bounding box, the problem of slow transmission speed of three-dimensional image data in medical image processing is solved, and the data volume is reduced and the transmission speed is improved, which is suitable for multi-center collaboration scenarios.
Patent Information
- Application Number
- CN202510533934.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In medical image processing, the large size and large amount of data are caused by slow data transmission speed, reducing the real-time nature of data processing. The problem is more prominent when multi-center transmission and multiple image processing algorithms are used.
By projecting the three-dimensional image data into a two-dimensional image, its feature vector is determined, and the task label is determined based on the feature vector, and the prediction model corresponding to the task label is used to predict the bounding box of the target area, thereby extracting the data contained in the bounding box for transmission.
It significantly reduces the amount of data to be transmitted, improves the data transmission speed, improves the efficiency and accuracy of bounding box prediction, and is suitable for data processing needs in multi-center collaboration scenarios.
Smart Images

Figure CN120047937A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of computer-aided medical treatment and image processing, and particularly relate to a method for predicting a bounding box of three-dimensional image data, a data transmission method, and a device. Background Art
[0002] The technology of medical image processing has developed rapidly. For example, three-dimensional image data such as computed tomography (CT) and magnetic resonance imaging (MRI) have been widely used in various disease diagnoses and surgical scenarios. However, whether doctors observe the data to complete a diagnosis or use image processing algorithms to extract the targets in the data, the transmission of three-dimensional image data needs to be completed first. Due to the large volume and large amount of data of three-dimensional image data in the medical field, the data transmission speed is slow, reducing the real-time performance of data processing. Especially in the case of multi-center transmission and dependence on multiple image processing algorithms, this problem is particularly prominent. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a method, a device, and a computer program product for predicting a bounding box of three-dimensional image data, so as to improve the efficiency and accuracy of predicting the bounding box of the target area. Furthermore, by extracting the data included in the bounding box for transmission, the amount of data to be transmitted can be significantly reduced, and the data transmission speed can be increased.
[0004] The first aspect of the embodiments of the present application provides a method for predicting a bounding box of three-dimensional image data, including: When receiving three-dimensional image data, projecting the three-dimensional image data into a two-dimensional image; Determining a feature vector of the two-dimensional image, where the feature vector includes a first vector along a first direction axis and a second vector along a second direction axis; Determining a task label of the three-dimensional image data according to the feature vector of the two-dimensional image; Using a prediction model corresponding to the task label to predict a bounding box of a target area of the three-dimensional image data, where the input data of the prediction model when performing the bounding box prediction is the task label and size information of the three-dimensional image data, and an independent prediction model is pre-trained for any type of task label.
[0005] Optionally, the determining the feature vector of the two-dimensional image includes: Respectively accumulating the pixel values of each data point in the two-dimensional image along the first direction axis and the second direction axis of the two-dimensional image to obtain the pixel accumulation values of each data point along the first direction axis and the second direction axis; Linear interpolation sampling is respectively performed using the pixel accumulation values of each data point to obtain a first vector along the first direction axis and a second vector along the second direction axis.
[0006] Optionally, before determining the feature vector of the two-dimensional image, it further includes: Determine the data fingerprint information of the three-dimensional image data, where the data fingerprint information includes a first position value and a second position value; Using the first position value as the lower threshold and the second position value as the upper threshold, perform data clamping on the two-dimensional image, so that the pixel values of each data point of the clamped two-dimensional image are within the range formed by the lower threshold and the upper threshold.
[0007] Optionally, the performing data clamping on the two-dimensional image with the first position value as the lower threshold and the second position value as the upper threshold includes: Determine the pixel values of each data point in the two-dimensional image; Adjust the pixel values of the data points less than the lower threshold to be equal to the lower threshold, and adjust the pixel values of the data points greater than the upper threshold to be equal to the upper threshold.
[0008] Optionally, the projecting the three-dimensional image data into a two-dimensional image includes: Determine the maximum value of each data point in the three-dimensional image data in the direction of the third direction axis; Project the three-dimensional image data along the coronal plane according to the maximum value of each data point in the direction of the third direction axis to obtain a two-dimensional image.
[0009] Optionally, the determining the task label of the three-dimensional image data according to the feature vector of the two-dimensional image includes: Process the feature vector of the two-dimensional image using a trained classification model, and output the corresponding task label of the three-dimensional image data. The classification model is obtained by training a support vector machine model using a feature vector data set marked with task labels. The feature vector data set includes multiple groups of a first vector along the first direction axis and a second vector along the second direction axis of a two-dimensional sample image, and any one of the two-dimensional sample images is obtained by projecting three-dimensional sample image data.
[0010] Optionally, the predicting the bounding box of the target region of the three-dimensional image data using the prediction model corresponding to the task label includes: Process the size information of the 3D image data using the prediction model corresponding to the task label, and output the coordinate values of the bounding box of the target area, where the coordinate values include the starting coordinate values and ending coordinate values of the bounding box on the first direction axis, the second direction axis, and the third direction axis.
[0011] Optionally, before predicting the bounding box of the target area of the 3D image data using the prediction model corresponding to the task label, it further includes: Construct a linear model; Train the linear model using 3D sample image data to obtain prediction models corresponding to any type of task label respectively. Any 3D sample image data has corresponding size information and task labels, and the 3D sample image data corresponding to any type of task label includes one or more bounding boxes.
[0012] A second aspect of the embodiments of the present application provides a method for transmitting 3D image data, including: When receiving 3D image data, determine the bounding boxes of at least one target area in the 3D image data; Extract the data included in the bounding boxes from the 3D image data; Transmit the data included in the extracted bounding boxes; Wherein, the bounding boxes of at least one target area are predicted by the method described in any item of the first aspect above.
[0013] A third aspect of the embodiments of the present application provides a bounding box prediction device for 3D image data, including: A projection module, configured to project the 3D image data into a 2D image when receiving the 3D image data; A determination module, configured to determine the feature vector of the 2D image, where the feature vector includes a first vector along the first direction axis and a second vector along the second direction axis; A classification module, configured to determine the task label of the 3D image data according to the feature vector of the 2D image; A prediction module, configured to predict the bounding box of the target area of the 3D image data using the prediction model corresponding to the task label. The input data for the prediction model when performing bounding box prediction is the task label and size information of the 3D image data, and an independent prediction model is pre-trained for any type of task label.
[0014] A fourth aspect of the embodiments of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the computer device implements the method according to any one of the first aspect and / or the second aspect as described above.
[0015] A fifth aspect of the embodiments of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a computer, the method according to any one of the first aspect and / or the second aspect as described above is implemented.
[0016] A sixth aspect of the embodiments of the present application provides a computer program product, including a computer program. When the computer program runs, the method according to any one of the first aspect and / or the second aspect as described above is executed.
[0017] Compared with the prior art, the embodiments of the present application have the following beneficial effects: In the embodiments of the present application, when receiving three-dimensional image data, the computer device can project the three-dimensional image data into a two-dimensional image. By determining the feature vector of the two-dimensional image, the computer device can determine the task label of the three-dimensional image data according to the feature vector, and use the prediction model corresponding to the task label to predict the bounding box of the target region of the three-dimensional image data, and output the coordinate value of the bounding box. Using the above method provided by the embodiments of the present application for predicting the bounding box can improve the efficiency and accuracy of bounding box prediction.
[0018] The embodiments of the present application can calculate the bounding box of the target region based on a multi-level simple linear model, without using computationally intensive deep learning methods, but instead using a lightweight linear prediction model, which significantly reduces the computational complexity. By introducing a preprocessing step of projection operation, the amount of data processing can be greatly reduced. In addition, the embodiments of the present application simultaneously utilize a two-stage strategy of support vector machine (SVM) classification combined with linear regression prediction, which can ensure the accuracy of locating the target region of interest and further improve the accuracy of bounding box prediction. The lightweight classification model and prediction model design provided by the embodiments of the present application not only make the model easy to be deployed to various medical devices, but also are particularly suitable for the data processing requirements in the multi-center collaboration scenario, providing an efficient and feasible solution for medical image processing. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of a method for predicting the bounding box of three-dimensional image data provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the overall process of a method for predicting the bounding box of three-dimensional image data provided by an embodiment of the present application; Figure 3 It is a schematic diagram of a process for obtaining data fingerprint information provided by an embodiment of the present application; Figure 4 It is a schematic diagram of a process for training a classification model provided by an embodiment of the present application; Figure 5 It is a schematic diagram of a process for training a prediction model provided by an embodiment of the present application; Figure 6 It is a schematic diagram of a method for transmitting three-dimensional image data provided by an embodiment of the present application; Figure 7 It is a schematic diagram of a device for predicting the bounding box of three-dimensional image data provided by an embodiment of the present application; Figure 8 It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0021] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0022] The following uses specific embodiments to illustrate the technical solutions of the present application.
[0023] Referring to Figure 1 , a schematic diagram of a method for predicting the bounding box of three-dimensional image data provided by an embodiment of the present application is shown, which may specifically include the following steps: S101. When receiving three-dimensional image data, project the three-dimensional image data into a two-dimensional image.
[0024] This method can be applied to a computer device, that is, the execution subject of the embodiments of this application is a computer device. By executing the various steps of this method, the computer device can accurately and efficiently predict the bounding box of the target area from the three-dimensional image data. The above computer device can be a device with data processing capabilities, such as a desktop computer, a cloud server, a computer-aided medical device, etc. The embodiments of this application do not limit the type of computer device.
[0025] The three-dimensional image data can be medical image data. For example, CT data, MRI data, etc. The three-dimensional image data can be obtained by scanning the whole body or a specific part of the patient's body with a corresponding scanning device in the pre-operative stage or the examination stage.
[0026] In one example, the computer device that is the execution subject of the embodiments of this application can be a scanning device. That is, corresponding data processing functions can be configured in the scanning device, so that after the scanning device completes the scanning of the patient's body part and obtains the three-dimensional image data, it uses the configured data processing functions to process it, thereby predicting the bounding box of the target area in the three-dimensional image data.
[0027] In another example, the computer device that is the execution subject of the embodiments of this application can also be an independent device communicatively connected to the scanning device. That is, after the scanning device scans and obtains the three-dimensional image data of the patient, it can transmit the three-dimensional image data to an independent computer device for processing. The computer device can predict the bounding box of the target area from the received three-dimensional image data by adopting the various steps of the method provided by the embodiments of this application.
[0028] In the embodiments of this application, when the computer device receives the three-dimensional image data, it can first project the three-dimensional image data into a two-dimensional image. By projecting the three-dimensional image data into a two-dimensional image, the amount of data to be processed in subsequent steps can be reduced, and the data processing efficiency can be accelerated.
[0029] In a possible implementation of the embodiment of the present application, the computer device may project the three-dimensional image data along the coronal plane in the maximum projection mode. Specifically, the computer device may determine the maximum value of each data point in the three-dimensional image data in the direction of the third axis, and then project the three-dimensional image data along the coronal plane according to the maximum value of each data point in the direction of the third axis to obtain a two-dimensional image. The above-mentioned third axis is orthogonal to the first axis and the second axis in pairs, and together they form the spatial coordinate system of the three-dimensional image data. The direction indicated by the third axis may represent the depth direction of the three-dimensional image data. Exemplarily, a single three-dimensional image data may be represented as V, and its size information is represented as H×W×D or H×W×L, where H represents the height of the three-dimensional image data V, W represents the width, and D or L represents the depth. The two-dimensional image obtained by projection along the coronal plane in the maximum projection mode may be represented as I, and its size information is H×W. That is, compared with the three-dimensional image data V, the depth direction data of the projected two-dimensional image I is reduced.
[0030] S102. Determine the feature vector of the two-dimensional image, where the feature vector includes a first vector along the first axis and a second vector along the second axis.
[0031] In the embodiment of the present application, the feature vector of the two-dimensional image may include vectors in two dimensional directions, that is, a first vector along the first axis and a second vector along the second axis. The above-mentioned first axis may be the x-axis in the image coordinate system of the two-dimensional image, and the second axis may be the y-axis in the image coordinate system. Therefore, after the computer device completes the projection to obtain the two-dimensional image, it may determine the first vector and the second vector of the two-dimensional image in the two directions of the x-axis and the y-axis of the image coordinate system respectively.
[0032] In a possible implementation of the embodiment of the present application, the computer device may accumulate the pixel values of each data point in the two-dimensional image along the first axis and the second axis of the two-dimensional image respectively to obtain the pixel accumulation values of each data point along the first axis and the second axis. Then, linear interpolation sampling is respectively performed using the pixel accumulation values of each data point to obtain a first vector along the first axis and a second vector along the second axis.
[0033] Exemplarily, taking the determination of the first vector in the first direction axis, i.e., the x-axis direction, as an example, the computer device can accumulate the pixel values of each data point in the x-axis direction of the two-dimensional image to obtain the pixel accumulation values of each data point in the x-axis direction. This process can be expressed as Ix[i] = sum(I[i,:]). In the above formula, x represents the x-axis direction, i represents each data point, and Ix[i] = sum(I[i,:]) represents the pixel accumulation value corresponding to each data point i. A pixel accumulation value can be calculated for each data point i. Correspondingly, for determining the second vector in the second direction axis, i.e., the y-axis direction, Iy[j] = sum(I[j,:]) can be used to represent the pixel accumulation values of each data point j in the y-axis direction. A pixel accumulation value can also be calculated for each data point j.
[0034] Then, the computer device can obtain two vectors in the x-axis direction and the y-axis direction, namely the first vector and the second vector, through linear interpolation sampling.
[0035] In a possible implementation manner of the embodiment of the present application, after calculating the pixel accumulation values of each data point respectively, in order to reduce the data size, the above pixel accumulation values can be divided by a factor γ, and interpolation sampling is performed with the value obtained after dividing by the factor γ. The above factor γ can be any non-zero integer multiple of 10.
[0036] Therefore, the above formula can be further expressed as: Ix[i] = sum(I[i,:]) / γ, Iy[j] = sum(I[j,:]) / γ. The computer device can perform linear interpolation sampling on Ix and Iy respectively to obtain the first vector I’x and the second vector I’y.
[0037] In the embodiment of the present application, when the computer device performs linear interpolation sampling, Ix and Iy can be sampled into vectors of a certain dimension. For example, 128-dimensional, that is, the obtained first vector I’x and the second vector I’y can be 128-dimensional feature vectors.
[0038] In a possible implementation manner of the embodiment of the present application, in order to improve the accuracy of data processing and reduce the influence of noise and artifacts in the data on subsequent bounding box prediction, after projecting to obtain a two-dimensional image, the computer device can, through data clamping, clamp each data point in the two-dimensional image within a certain range, thereby suppressing the maximum and minimum values in the two-dimensional image and removing the noise and artifacts in the data.
[0039] In the embodiments of the present application, the data fingerprint of the three-dimensional image data can be used to clamp the range of each data in the two-dimensional image. Specifically, the computer device can determine the data fingerprint information of the three-dimensional image data, where the data fingerprint information can include a first position value and a second position value. Then, the computer device uses the first position value as the lower threshold and the second position value as the upper threshold to perform data clamping on the two-dimensional image, so that the pixel values of each data point in the clamped two-dimensional image are within the range composed of the above lower threshold and upper threshold.
[0040] The above data fingerprint can be data that can represent the specific characteristics of the three-dimensional image data calculated according to a certain standard. Exemplarily, the data fingerprint of the three-dimensional image data can include the mean, standard deviation (std), maximum value (max_value), minimum value (min_value), etc. of each data point in the foreground data part of the current three-dimensional image data. Let fp represent the data fingerprint of the three-dimensional image data, then fp = {max_value, min_value, std, mean} can be obtained. Among them, the maximum value (max_value) can refer to the pixel value of the data point located at the 99.5% position after sorting the pixel values of each data point in the foreground data, and the minimum value (min_value) can be the pixel value of the data point located at the 0.5% position after the above sorting.
[0041] In a possible implementation manner of the embodiments of the present application, when using the data fingerprint fp to perform data clamping on the two-dimensional image, the first position value can be the minimum value (min_value) in the above data fingerprint fp, and the second position value can be the maximum value (max_value).
[0042] When using the first position value, that is, the minimum value (min_value), as the lower threshold and the second position value, that is, the maximum value (max_value), as the upper threshold, to perform data on the two-dimensional image, the pixel value of each data point in the two-dimensional image can be determined, so that the pixel value of the data point smaller than the lower threshold is adjusted to be equal to the lower threshold, and the pixel value of the data point larger than the upper threshold is adjusted to be equal to the upper threshold, so as to limit each data point in the two-dimensional image within the range of [min_value, max_value].
[0043] Determining the feature vector of the two-dimensional image can be performed after completing the above data clamping step.
[0044] S103. Determine the task label of the three-dimensional image data according to the feature vector of the two-dimensional image.
[0045] In the embodiments of the present application, the task label can be used to represent the target area of the three-dimensional image data, and the prediction of the bounding box is to predict the bounding box of the target area. For example, if the task label represents the stomach, then after processing, the bounding box of the stomach area needs to be predicted from the three-dimensional image data. The computer device can determine the task label of the three-dimensional image data according to the feature vector of the two-dimensional image.
[0046] In a possible implementation manner of the embodiments of the present application, a trained classification model can be used to process the feature vector of the two-dimensional image and output the corresponding task label of the three-dimensional image data.
[0047] Specifically, a classification model can be pre-trained and configured in the computer device. After the computer completes the processing of the foregoing steps and obtains the feature vector of the two-dimensional image, the classification model can be called to process the feature vector and output the corresponding task label. That is, the feature vector of the two-dimensional image can be used as the input data of the classification model, and after being processed by the classification model, the task label is output.
[0048] In the embodiments of the present application, the foregoing classification model can be obtained by training an SVM model using a feature vector data set marked with task labels. The feature vector data set, as the training data for training the classification model, can include multiple groups of first vectors along the first direction axis of the two-dimensional sample image and second vectors along the second direction axis. Any of the foregoing two-dimensional sample images can be obtained by projecting three-dimensional sample image data.
[0049] Specifically, multiple three-dimensional image data can be collected as sample data to form three-dimensional sample image data. The computer device can project the three-dimensional sample image data according to the steps described above to obtain multiple two-dimensional sample images. For each two-dimensional sample image, the computer device determines the first vector of the image along the first direction axis x-axis and the second vector along the second direction axis y-axis to form multiple vector groups. Each vector group includes a set of feature vectors of the two-dimensional sample image, that is, the first vector I’x and the second vector I’y determined in the foregoing steps. For each three-dimensional sample image data, it can be pre-marked with the corresponding task label. For example, the task label can be the stomach area, both hands area, left hand area, or right hand area, etc. The computer device can input multiple vector groups into the SVM model for model training, output the corresponding task label, and compare it with the true task label of the three-dimensional sample image data corresponding to each marked vector group. Through multiple iterations, a classification model that can be used subsequently is obtained.
[0050] In this way, after the training of the classification model is completed, the classification model can be configured in the computer device to process the feature vector of the two-dimensional image determined each time and output the task label for bounding box prediction.
[0051] S104. Use the prediction model corresponding to the task label to predict the bounding box of the target region of the three-dimensional image data. The input data for the prediction model during bounding box prediction is the task label and size information of the three-dimensional image data. An independent prediction model is pre-trained for each type of task label.
[0052] In an embodiment of the present application, in combination with the task label determined in the foregoing steps, the computer device can call the pre-trained prediction model for prediction and output the bounding box of the target region. Among them, an independent prediction model can be pre-trained for each type of task label. In this way, by training an independent prediction model for each task label, after determining the task label of the three-dimensional image data, the computer device can directly call the prediction model corresponding to the task label for processing, and efficiently and accurately complete the bounding box prediction of the target region.
[0053] In a possible implementation manner of the embodiment of the present application, the prediction model can be a linear model. By constructing a linear model, the computer device can use three-dimensional sample image data to train the linear model and obtain prediction models corresponding to each type of task label respectively.
[0054] In the embodiment of the present application, the constructed linear model can be expressed as y = wx + k, where w is the model parameter and k is the bias. After constructing the linear model in the above form, the computer device can use three-dimensional sample image data for model training. Among them, any three-dimensional sample image data has corresponding size information and task label, and the three-dimensional sample image data corresponding to each type of task label can include one bounding box or multiple bounding boxes. Exemplarily, for the three-dimensional image data of the two-handed hand region collected, the left and right hand regions belong to different bounding boxes respectively.
[0055] In a possible implementation of the embodiment of the present application, when training the linear model in the form of y = wx + k, x = [H, W, D] represents the size information of three-dimensional sample image data, and y = [x1, x2, y1, y2, z1, z2]*n represents the coordinates of n bounding boxes in the three-dimensional sample image data. By using multiple three-dimensional sample image data, algorithms such as the least squares method can be used to optimize and obtain the model parameters w and the bias value k of the prediction model corresponding to the corresponding task label. After calculating the values of the model parameters w and the bias value k, the prediction model y = wx + k can be determined. By configuring the prediction model in a computer device, after receiving the size information and task label of the three-dimensional image data, the prediction model corresponding to the task label can be used to process the task label and size information, and output the coordinate values of the bounding box of the target area. The coordinate values can include the starting coordinate values and ending coordinate values of the bounding box on the first direction axis, i.e., the x-axis, the second direction axis, i.e., the y-axis, and the third direction axis, i.e., the z-axis. Exemplarily, [x1, x2] can be used to represent the starting coordinate values and ending coordinate values of a bounding box on the x-axis, [y1, y2] can be used to represent the starting coordinate values and ending coordinate values of the bounding box on the y-axis, and [z1, z2] can be used to represent the starting coordinate values and ending coordinate values of the bounding box on the z-axis. When there are multiple target areas included in the three-dimensional image data, the corresponding output bounding boxes are also multiple, and the coordinate values of each bounding box include the starting coordinate values and ending coordinate values of the bounding box in the form of [x1, x2, y1, y2, z1, z2] on the x-axis, y-axis, and z-axis. Among them, the number of bounding boxes output by the prediction model is equal to the number of target areas in the three-dimensional image data.
[0056] In a possible implementation of the embodiment of the present application, after the computer device completes the prediction of the bounding box and outputs the respective coordinate values of the corresponding bounding box in the three-dimensional image data, the data in the bounding box can be extracted from the three-dimensional image data according to the coordinate values. In this way, the computer device can transmit the extracted data to other devices for use. For example, it can be transmitted to the device used by a doctor for the doctor to diagnose and treat diseases based on the three-dimensional image data. By using the steps provided by this method to predict the bounding box and extract data for transmission, it is possible to avoid transmitting all the three-dimensional image data and only transmit the data of a specific area therein, which can improve the speed of data transmission.
[0057] In the embodiments of the present application, when receiving three-dimensional image data, a computer device may project the three-dimensional image data into a two-dimensional image. By determining the feature vector of the two-dimensional image, the computer device may determine the task label of the three-dimensional image data according to the feature vector, and use the prediction model corresponding to the task label to predict the bounding box of the target region of the three-dimensional image data, and output the coordinate values of the bounding box. Using the above method provided by the embodiments of the present application for bounding box prediction can improve the efficiency and accuracy of bounding box prediction.
[0058] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0059] For ease of understanding, the following introduces the bounding box prediction method provided by the embodiments of the present application in combination with a complete example.
[0060] As Figure 2 shown, it is a schematic diagram of the overall process of the bounding box prediction method for three-dimensional image data provided by the embodiments of the present application. Figure 2 The three-dimensional data V in Figure 2 is the three-dimensional image data V introduced in the foregoing embodiments. For example, it may be CT data or MRI data, etc. Any three-dimensional data V has corresponding size information, that is, w × h × l shown in w × h (H×W×D or H×W×L represented in the foregoing embodiments). For the three-dimensional data V, it can be converted into a two-dimensional image I through data projection, and its size can be expressed as
[0061] Figure 2 The clamp shown in
[0062] is a numerical limit function. For the two-dimensional image I, the computer device may limit each value in the image I to an interval with a lower limit of fp0 and an upper limit of fp1 through data clamping in combination with the data fingerprint fp. The lower limit fp0 may be the first position value (i.e., the smaller value, min_value) in the foregoing embodiments, and the upper limit fp1 may be the second position value (i.e., the larger value, max_value) in the foregoing embodiments.
[0062] As Figure 3 shown, it is a schematic diagram of a data fingerprint information acquisition process provided by the embodiments of the present application. According to Figure 3As shown, a three-dimensional dataset may include multiple three-dimensional image data, such as three-dimensional image data V1, …, three-dimensional image data Vn, etc. For each three-dimensional image data, corresponding foreground data G can be obtained through image segmentation, and the mean value, standard deviation, minimum value (min_value) that can be used as the lower limit value fp0, maximum value (max_value) that can be used as the upper limit value fp1, etc. can be calculated. Among them, Figure 3 Q in 0.005 (G) may represent the pixel value of the data point at the 0.5% position after sorting all data points in the foreground data G by pixel value size, Q 0.995 (G) represents the pixel value of the data point at the 99.5% position after sorting all data points in the foreground data G by pixel value size.
[0063] According to Figure 3 After calculating the data fingerprint of the three-dimensional image data according to the shown process, this data fingerprint can be used to limit the pixel value of each data point in the two-dimensional image I within the range of [fp0 = min_value, fp1 = max_value].
[0064] Return to refer to Figure 2 , after completing the data clamping of the two-dimensional image I, the computer device can respectively determine the feature vectors along the x-axis and y-axis of the image, and sample to obtain the first vector I’x and the second vector I’y.
[0065] As Figure 4 shown, it is a schematic diagram of a classification model training process provided by an embodiment of the present application. Figure 4 During the training process of the classification model shown, the process of the computer device processing the three-dimensional sample image data and sampling to obtain the first vector I’x along the x-axis and the second vector I’y along the y-axis of the corresponding two-dimensional image is similar to Figure 2 the shown process. Multiple groups of vector groups determined from multiple three-dimensional sample image data can be used as training data for training the classification model, so as to train the classification model, and this classification model can be Figure 4 the SVM model shown in
[0066] On the other hand, in addition to the classification model, a prediction model also needs to be trained. As Figure 5 shown, it is a schematic diagram of a prediction model training process provided by an embodiment of the present application. The prediction model can be trained based on a linear model. For each task label, an independent linear prediction model can be trained for subsequent bounding box prediction. The trained classification model and prediction model can be configured in the computer device.
[0067] Return to refer to Figure 2 , for the three-dimensional data V to be processed, the computer device according toFigure 2 After the corresponding first vector I’x and second vector I’y are determined according to the process shown, the computer device can call the SVM model trained according to the Figure 4 process shown to process the feature vectors and output the task label task_label of the three-dimensional data, that is, Figure 2 the category c shown in. When calling the SVM model to process the feature vectors, the input data of the SVM are the first vector I’x along the x-axis of the image and the second vector I’y along the y-axis of the image in the two-dimensional image obtained by projecting the current three-dimensional data V to be processed. The data output by the SVM is the task label task_label of the corresponding three-dimensional data V.
[0068] Then, the computer device can call the prediction model trained according to the Figure 5 process shown to perform bounding box prediction. Among them, the input data of the prediction model are the size information of the three-dimensional data V w × h × l and the corresponding task label task_label, and the output data of the prediction model are the coordinate values of the bounding box of at least one target region in the three-dimensional data V.
[0069] In this way, at least one target region can be efficiently and accurately determined from the three-dimensional data, and the position where the bounding box of the target region is located can be determined through the coordinate values output by the prediction model.
[0070] Combined with the bounding box prediction methods introduced in the foregoing embodiments, the embodiments of the present application also provide a method for transmitting three-dimensional image data. Referring to Figure 6 , which is a schematic diagram of a method for transmitting three-dimensional image data provided by the embodiments of the present application, may specifically include the following steps: S601. When receiving three-dimensional image data, determine the bounding box of at least one target region in the three-dimensional image data.
[0071] S602. Extract the data included in the bounding box from the three-dimensional image data.
[0072] S603. Transmit the data included in the extracted bounding box.
[0073] Among them, the bounding box of at least one target region can be predicted by the method introduced in the foregoing embodiments of the bounding box prediction method.
[0074] The present method can be applied to a computer device. For example, the computer device can be the scanning device in the foregoing embodiments. That is, after the scanning device obtains three-dimensional image data by scanning, it can extract the data included in the bounding box of the target region in the three-dimensional image data according to the steps described in this method, and transmit the data. Alternatively, the computer device can also be other devices communicating with the scanning device, such as computer-aided medical devices, etc. After receiving the three-dimensional image data scanned by the scanning device, the computer device can extract the data included in the bounding box of the target region in the three-dimensional image data according to the steps described in this method, and transmit the data. For example, the extracted data is distributed to multiple other devices that need to use the data. The embodiments of the present application do not limit this.
[0075] By using the transmission method provided in the embodiments of the present application, the data included in the bounding box of the target region can be extracted from the three-dimensional image data in combination with the steps described in the foregoing various bounding box prediction methods. Thus, in a scenario where three-dimensional image data needs to be transmitted, only the data of the target region is transmitted, reducing the amount of data to be transmitted and improving the data transmission efficiency.
[0076] Referring to Figure 7 , a schematic diagram of a bounding box prediction device for three-dimensional image data provided in the embodiments of the present application is shown. Specifically, it may include a projection module 701, a determination module 702, a classification module 703, and a prediction module 704, where: The projection module 701 is configured to project the three-dimensional image data into a two-dimensional image when receiving the three-dimensional image data; The determination module 702 is configured to determine a feature vector of the two-dimensional image, where the feature vector includes a first vector along a first direction axis and a second vector along a second direction axis; The classification module 703 is configured to determine a task label of the three-dimensional image data according to the feature vector of the two-dimensional image; The prediction module 704 is configured to predict a bounding box of a target region of the three-dimensional image data by using a prediction model corresponding to the task label. The input data of the prediction model when performing bounding box prediction is the task label and size information of the three-dimensional image data, and an independent prediction model is pre-trained for any type of task label.
[0077] In the embodiments of the present application, the determination module 702 may specifically be configured to: Accumulate the pixel values of each data point in the two-dimensional image along the first direction axis and the second direction axis of the two-dimensional image respectively to obtain the pixel accumulation values of each data point along the first direction axis and the second direction axis; Linear interpolation sampling is respectively performed using the pixel accumulation values of each data point to obtain a first vector along the first direction axis and a second vector along the second direction axis.
[0078] In a possible implementation manner of the embodiment of the present application, the determining module 702 may further be configured to: Determine the data fingerprint information of the three-dimensional image data, where the data fingerprint information includes a first position value and a second position value; Taking the first position value as the lower threshold and the second position value as the upper threshold, perform data clamping on the two-dimensional image, so that the pixel values of each data point of the clamped two-dimensional image are within the range formed by the lower threshold and the upper threshold.
[0079] In another possible implementation manner of the embodiment of the present application, the determining module 702 may further be configured to: Determine the pixel values of each data point in the two-dimensional image; Adjust the pixel values of the data points less than the lower threshold to be equal to the lower threshold, and adjust the pixel values of the data points greater than the upper threshold to be equal to the upper threshold.
[0080] In the embodiment of the present application, the projection module 701 may specifically be configured to: Determine the maximum value of each data point in the three-dimensional image data in the direction of the third direction axis; Project the three-dimensional image data along the coronal plane according to the maximum value of each data point in the direction of the third direction axis to obtain a two-dimensional image.
[0081] In the embodiment of the present application, the classification module 703 may specifically be configured to: Process the feature vector of the two-dimensional image using a trained classification model, and output the task label of the corresponding three-dimensional image data. The classification model is obtained by training a support vector machine model using a feature vector data set marked with task labels. The feature vector data set includes multiple groups of a first vector along the first direction axis and a second vector along the second direction axis of a two-dimensional sample image. Any one of the two-dimensional sample images is obtained by projecting three-dimensional sample image data.
[0082] In the embodiment of the present application, the prediction module 704 may specifically be configured to: Process the size information of the three-dimensional image data using the prediction model corresponding to the task label, and output the coordinate values of the bounding box of the target area. The coordinate values include the starting coordinate values and ending coordinate values of the bounding box on the first direction axis, the second direction axis, and the third direction axis.
[0083] In a possible implementation manner of the embodiment of the present application, the device may further include a training module, and specifically, the training module may be used for: Construct a linear model; Use three-dimensional sample image data to train the linear model to obtain a prediction model corresponding to any type of task label respectively. Any piece of the three-dimensional sample image data has corresponding size information and a task label, and the three-dimensional sample image data corresponding to any type of task label includes one or more bounding boxes.
[0084] A bounding box prediction device for three-dimensional image data provided by an embodiment of the present application. By applying this device, each step in the foregoing method embodiments can be implemented.
[0085] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the description in the method embodiment section.
[0086] Referring to Figure 8 , a schematic diagram of a computer device provided by an embodiment of the present application is shown. As Figure 8 shown, the computer device 800 in the embodiment of the present application includes: a processor 810, a memory 820, and a computer program 821 stored in the memory 820 and executable on the processor 810. When the processor 810 executes the computer program 821, the steps in each of the foregoing embodiments of the bounding box prediction method and / or data transmission method for three-dimensional image data are implemented, such as Figure 1 the steps S101 to S104 shown in Figure 6 the steps S601 to S603 shown in Figure 7 . Or, when the processor 810 executes the computer program 821, the functions of each module / unit in the foregoing device embodiments are implemented, such as
[0087] Exemplarily, the computer program 821 may be divided into one or more modules / units. The one or more modules / units are stored in the memory 820 and executed by the processor 810 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments may be used to describe the execution process of the computer program 821 in the computer device 800. For example, the computer program 821 may be divided into a projection module, a determination module, a classification module, and a prediction module, and the specific functions of each module are as follows: The projection module is used to project the three-dimensional image data into a two-dimensional image when receiving the three-dimensional image data; A determination module, configured to determine a feature vector of the two-dimensional image, where the feature vector includes a first vector along a first direction axis and a second vector along a second direction axis; A classification module, configured to determine a task label of the three-dimensional image data according to the feature vector of the two-dimensional image; A prediction module, configured to predict a bounding box of a target region of the three-dimensional image data by using a prediction model corresponding to the task label, where input data for the prediction model during bounding box prediction is the task label and size information of the three-dimensional image data, and an independent prediction model is pre-trained for each type of task label.
[0088] The computer device 800 may be a device capable of implementing the corresponding steps in the foregoing various method embodiments. The computer device 800 may be a desktop computer, a cloud server, or the like. For example, the computer device 800 may be a computer-aided medical device. The computer device 800 may include, but is not limited to, a processor 810 and a memory 820. Those skilled in the art can understand that Figure 8 This is only an example of the computer device 800 and does not constitute a limitation on the computer device 800. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device 800 may further include input / output devices, network access devices, a bus, etc.
[0089] The processor 810 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0090] The memory 820 may be an internal storage unit of the computer device 800, such as a hard disk or memory of the computer device 800. The memory 820 may also be an external storage device of the computer device 800, such as a plug-in hard disk equipped on the computer device 800, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, and so on. Further, the memory 820 may also include both the internal storage unit and the external storage device of the computer device 800. The memory 820 is used to store the computer program 821 and other programs and data required by the computer device 800. The memory 820 may also be used to temporarily store data that has been output or is to be output.
[0091] An embodiment of the present application also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the methods described in the foregoing various embodiments are implemented.
[0092] An embodiment of the present application also discloses a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a computer, the methods described in the foregoing various embodiments are implemented.
[0093] An embodiment of the present application also discloses a computer program product, including a computer program. When the computer program runs on a computer, the computer is caused to execute the methods described in the foregoing various embodiments.
[0094] The foregoing embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for predicting a bounding box of three-dimensional image data, characterized in that: include: When receiving the three-dimensional image data, projecting the three-dimensional image data into a two-dimensional image; Determine a feature vector of the two-dimensional image, wherein the feature vector includes a first vector along a first direction axis and a second vector along a second direction axis; Determining a task label of the three-dimensional image data according to a feature vector of the two-dimensional image; The prediction model corresponding to the task label is used to predict the bounding box of the target area of the three-dimensional image data. The input data of the prediction model when predicting the bounding box is the task label and size information of the three-dimensional image data. Any type of task label has an independent prediction model pre-trained.
2. The method according to claim 1, characterized in that: Determining the feature vector of the two-dimensional image includes: Accumulating pixel values of each data point in the two-dimensional image along a first direction axis and a second direction axis of the two-dimensional image respectively to obtain pixel accumulation values of each data point along the first direction axis and along the second direction axis; Linear interpolation sampling is performed using the pixel accumulation values of each data point to obtain a first vector along the first direction axis and a second vector along the second direction axis.
3. The method according to claim 1, characterized in that Before determining the feature vector of the two-dimensional image, the method further includes: Determine data fingerprint information of the three-dimensional image data, wherein the data fingerprint information includes a first position value and a second position value; Determining a pixel value of each data point in the two-dimensional image; With the first position value as the lower threshold and the second position value as the upper threshold, data clamping is performed on the two-dimensional image, and the pixel values of the data points smaller than the lower threshold are adjusted to be equal to the lower threshold, and the pixel values of the data points larger than the upper threshold are adjusted to be equal to the upper threshold, so that the pixel values of each data point of the clamped two-dimensional image are within the interval formed by the lower threshold and the upper threshold.
4. The method according to claim 3, characterized in that The projecting the three-dimensional image data into a two-dimensional image comprises: Determine the maximum value of each data point in the three-dimensional image data in the direction of the third direction axis; According to the maximum value of each data point in the third axial direction, the three-dimensional image data is projected along the coronal plane to obtain a two-dimensional image.
5. The method according to any one of claims 1 to 4, characterized in that: The step of determining the task label of the three-dimensional image data according to the feature vector of the two-dimensional image includes: The feature vector of the two-dimensional image is processed using a trained classification model, and the corresponding task label of the three-dimensional image data is output. The classification model is obtained by training a support vector machine model using a feature vector data set marked with the task label. The feature vector data set includes multiple groups of first vectors along the first direction axis of the two-dimensional sample image and second vectors along the second direction axis. Any of the two-dimensional sample images is obtained by projecting the three-dimensional sample image data.
6. The method according to claim 5, characterized in that The predicting of the bounding box of the target area of the three-dimensional image data by using the prediction model corresponding to the task label includes: The prediction model corresponding to the task label is used to process the size information of the three-dimensional image data, and the coordinate values of the bounding box of the target area are output, wherein the coordinate values include the starting coordinate values and the ending coordinate values of the bounding box on the first direction axis, the second direction axis and the third direction axis.
7. The method according to claim 6, characterized in that Before using the prediction model corresponding to the task label to predict the bounding box of the target area of the three-dimensional image data, the method further includes: Construct linear models; The linear model is trained using three-dimensional sample image data to obtain prediction models corresponding to any type of task label. Any of the three-dimensional sample image data has corresponding size information and task labels. The three-dimensional sample image data corresponding to any type of task label includes one or more bounding boxes.
8. A method for transmitting three-dimensional image data, characterized in that: include: When receiving the three-dimensional image data, determining a bounding box of at least one target area in the three-dimensional image data; Extracting data contained in the bounding box from the three-dimensional image data; Transmitting the data contained in the extracted bounding box; Wherein, at least one bounding box of the target area is predicted by the method according to any one of claims 1-7.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the computer device is caused to implement the method according to any one of claims 1 to 7 or 8.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed, the method according to any one of claims 1 to 7 or 8 is executed.
Citation Information
Patent Citations
Three-dimensional CT image focus identification method and system based on big data and deep learning
CN118710973A
CT sequence scanning range efficient classification method based on maximum density projection
CN119107492A
Image processing device, image processing model generation device, learning data generation device, and program
WO2021177374A1
Cited By
Two-dimensional projection image generation method, computer equipment and computer program product
CN120833403A
Method for generating a two-dimensional projection image, computer device and computer program product
CN120833403B