An image data processing method based on cloud computing
By using a cloud-based image data processing method, combining a target library and a data detection center with deep learning and residual networks, the problem of labor and time consumption in image annotation is solved, and efficient image data processing and recognition are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZTE (WENZHOU) RAILWAY COMM TECH LTD
- Filing Date
- 2022-12-05
- Publication Date
- 2026-05-19
AI Technical Summary
Current technologies for image annotation are labor-intensive and time-consuming, and image data processing is inefficient.
The cloud-based image data processing method establishes a target library and a data detection center. The target library is used to store and label image information. Deep learning and residual networks are combined for image data processing, and fully convolutional neural networks are used for semantic segmentation and target detection.
It improves the efficiency and accuracy of image data processing, reduces the consumption of human resources, and enables rapid analysis and efficient image target recognition.
Smart Images

Figure CN115757842B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image data processing technology, specifically a cloud computing-based image data processing method. Background Technology
[0002] Image processing technology is a technique that uses computers to process image information. It mainly includes image digitization, image enhancement and restoration, image data encoding, image segmentation and image recognition. Semantic segmentation requires a large amount of labeled image data for model training before it can be used. However, image labeling is a very time-consuming and labor-intensive task. To save time and manpower, we propose an image data processing method based on cloud computing. Summary of the Invention
[0003] The purpose of this invention is to solve the problems existing in the prior art by proposing a cloud computing-based image data processing method.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] A cloud-based image data processing method includes the following steps:
[0006] S1: Image acquisition, based on image data uploaded by staff or users or stored in the database, or can be connected to the Internet to connect to and receive data on the Internet, or select the image processing data source to be used according to its own needs;
[0007] S2: Target clustering processing, which preprocesses the image data collected from the Internet, collects the targets assigned to the images on the Internet, and if there is image data uploaded by staff or users, it is determined whether the staff assigns targets to it, establishes a target library, collects the targets uploaded by staff or users and the targets of image data obtained from the Internet, establishes a target pyramid, performs clustering processing on the image data, and obtains the cluster centers of the corresponding targets.
[0008] S3: Initial clustering of image data. Based on the target image library obtained in S2, cluster the experimental image data found on the Internet, establish a test center, receive the target types obtained in S2, analyze images similar to the target, conduct data verification and comparison, randomly collect different images on the Internet, and label the images according to the target features provided in S2, and import the labeled image data into the sample library.
[0009] S4: Feature enhancement processing, where staff evaluate the data in the sample library detected by S3 and feed the detection results back to the S3 sample library for expansion and improvement, while also annotating the target library with information such as (location, volume);
[0010] S5: Data processing. Establish a data detection center to perform image recognition on image data uploaded by staff or users or stored in the database through the target library. Arrange the images according to the proportion of each target on the image. Use semantic segmentation to identify the range of the target. Label and analyze each object in the image. Use object detection as a precondition for semantic segmentation to identify and label the bounding boxes of each object. This is also the process of pairing the target with each object in the image. Use fully convolutional neural network technology to test the test image. Through pixel-to-pixel mapping and end-to-end mapping, detect and analyze the range of the image and the distance between image objects to obtain image data.
[0011] S6: Data Detection and Optimization: Compare the data obtained from S5 with the data obtained by users or staff to obtain the image data accuracy, and feed it back to the S5 data detection center for optimization.
[0012] Preferably, after the image is acquired in S1, the acquired data uploaded by the staff is subjected to image enhancement. Image enhancement is to operate on an image so that the result is more suitable for processing in a specific application than the original image. According to existing methods, appropriate methods are selected to enhance different numbers of images. There is no universal theory for image enhancement methods. There are many different image enhancement methods, and special cases are treated specially.
[0013] Preferably, during S2 image preprocessing, when constructing the target pyramid using clustering, the direct features of each target are used as the starting point of the pyramid distribution map. The entire pyramid distribution map for different types of images is constructed from bottom to top. The correlation weights between all targets are calculated, and multiple direct features with the same weight are merged into a lower-level pyramid cluster item of the corresponding target pyramid. The correlation weights between all lower-level pyramid cluster items are calculated, and multiple lower-level pyramid cluster items with large correlation weights are merged into middle-level pyramid cluster items. This operation is repeated until the correlation weight between the two top-level pyramid cluster items is greater than zero. Then, a pyramid classification layer tree diagram is constructed. If the correlation weight between the two top-level cluster items is equal to zero, two or more pyramid classification layer tree diagrams are constructed respectively. This is combined with fully convolutional neural network technology to complete data processing.
[0014] Preferably, during image preprocessing in S2, data annotations are added to the target, such as the size of the relative space and the actual volume, to provide data support for abstracting and layering during semantic segmentation in the data detection center in S5.
[0015] Preferably, in step S5, the data detection center establishes a pixel evaluation index library, which includes evaluation indices such as pixel accuracy, average pixel accuracy, average crossover ratio, and weighted frequency crossover ratio.
[0016] Preferably, the data detection center in S5 is optimized using deep learning methods. Based on deep learning methods, multi-class targets can be predicted directly, and multiple targets can be predicted. In conjunction with the target library, a residual network (ResNet) is used as a feature extractor to extract different patterns from the input image by extending the network. Then, these feature maps are input into the pyramid pooling module to distinguish patterns at different scales. They are aggregated at four different scales, each scale corresponding to a pyramid layer, and processed by 1×1 convolutional layers to reduce their dimensionality. The output of the pyramid layer is upsampled and connected with the initial feature map to capture local and global contextual information. Finally, convolutional layers are used to generate pixel-wise predictions.
[0017] Preferably, the accuracy of S5 semantic segmentation refers to the accuracy measurement of pixel-by-pixel labeling. Assuming there are k classes (from l0 to lk, one of which belongs to the background), Pij represents the number of pixels that belong to class i but are predicted to be class j, and Pii represents the number of truly correct class classifications. Pij and Piji are called false positive samples and false negative samples, respectively.
[0018] Compared with existing technologies, the present invention provides a cloud computing-based image data processing method, which has the following beneficial effects:
[0019] 1. This invention establishes a target library so that the targets in the image information are stored in the target library before the image data is detected. The data in the target library is organized and the information of the targets in the target library, such as volume and distance, is labeled. The target data is enriched in a targeted manner, so as to quickly analyze the data on the image targets during the later image processing.
[0020] 2. By setting up a detection center and using deep learning methods to optimize it, deep learning methods can directly predict multiple categories of targets. Multiple targets can be predicted. In conjunction with the target library, a residual network (ResNet) is used as a feature extractor. It is combined with and verified with the data in the target library. The target library pyramid classification layer tree diagram and the residual network (ResNet) are used as feature extractors. Different patterns are extracted from the input image by expanding the network. Then these feature maps are input into the pyramid pooling module to improve the detection quality and speed. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0022] In the attached diagram:
[0023] Figure 1 This is a flowchart of an image data processing method based on cloud computing proposed in this invention. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0025] Example 1:
[0026] Please see Figure 1 A cloud computing-based image data processing method includes the following steps:
[0027] S1: Image acquisition. Based on image data uploaded by staff or users or stored in the database, or connected to the Internet, data can be received and connected to the Internet. Alternatively, the required image processing data source can be selected according to the needs of the user. Image enhancement is the operation on an image to make the result more suitable for processing in a specific application than the original image. The appropriate method is selected to enhance different images based on existing methods. There is no universal theory for image enhancement methods. There are many different image enhancement methods, and special cases are treated specially.
[0028] S2: Target Clustering Processing. This involves preprocessing image data collected from the internet, collecting targets assigned to images from the internet, and, if images are uploaded by staff or users, assigning targets based on staff input. A target database is established, collecting targets from both staff / user uploads and images obtained from the internet. A target pyramid is built, and the image data is clustered to obtain cluster centers for each target. When building the target pyramid using clustering methods, the direct features of each target are used as the starting point of the pyramid distribution map. The entire pyramid distribution map for different types of images is built from the bottom up. The association weights between all targets are calculated, and multiple direct features with the same weight are grouped together. Features are merged into a lower-level pyramid cluster item corresponding to the target pyramid. The association weights between all lower-level pyramid cluster items are calculated, and multiple lower-level pyramid cluster items with large association weights are merged into middle-level pyramid cluster items. This operation is repeated until the association weight between the two top-level pyramid cluster items is greater than zero. Then, a pyramid classification layer tree diagram is built. If the association weight between the two top-level cluster items is equal to zero, two or more pyramid classification layer tree diagrams are built respectively. This is combined with fully convolutional neural network technology to complete data processing, such as data annotation of the target during preprocessing, such as the size of the relative space, actual volume, etc., and provides data support for abstracting and layering during the semantic segmentation process of the data detection center in S5.
[0029] S3: Initial clustering of image data. Based on the target image library obtained in S2, cluster the experimental image data found on the Internet, establish a test center, receive the target types obtained in S2, analyze images similar to the target, conduct data verification and comparison, randomly collect different images on the Internet, and label the images according to the target features provided in S2, and import the labeled image data into the sample library.
[0030] S4: Feature enhancement processing, where staff evaluate the data in the sample library detected by S3 and feed the detection results back to the S3 sample library for expansion and improvement, while also annotating the target library with information such as (location, volume);
[0031] S5: Data Processing. A data detection center is established to perform image recognition on image data uploaded by staff or users, or stored in the database, using a target library. The images are then ordered according to the proportion of each target in the image. Semantic segmentation is used to identify the range of each target. Each object in the image is labeled and analyzed. Object detection is used as a pre-segmentation of semantic segmentation to identify and label the bounding boxes of each object, which is also the process of pairing targets with objects in the image. Fully convolutional neural network technology is used to test the images. Through pixel-to-pixel and end-to-end mapping, the range of the image and the distance between image objects are detected and analyzed to obtain image data. The data detection center establishes a pixel evaluation index library. Evaluation metrics such as library pixel accuracy, average pixel accuracy, average intersection-over-union ratio (IoU), and weighted frequency IoU are optimized by the data detection center using deep learning methods. Based on deep learning methods, multi-class targets can be predicted directly, and multiple targets can be predicted. In conjunction with the target library, a residual network (ResNet) is used as a feature extractor. Different patterns are extracted from the input image by extending the network. Then, these feature maps are input into the pyramid pooling module to distinguish patterns at different scales. They are aggregated at four different scales, each scale corresponding to a pyramid layer, and processed by 1×1 convolutional layers to reduce their dimensionality. The output of the pyramid layer is upsampled and concatenated with the initial feature map to capture local and global contextual information. Finally, convolutional layers are used to generate pixel-wise predictions.
[0032] S6: Data Detection and Optimization: Compare the data obtained from S5 with that of users or staff to obtain the image data accuracy, and then feed it back to the S5 data detection center for optimization.
[0033] The accuracy of semantic segmentation refers to the precision measurement of pixel-by-pixel labeling. Assuming there are k classes (from l0 to lk, one of which belongs to the background), Pij represents the number of pixels that belong to class i but are predicted to be class j, and Pii represents the number of truly correct class classifications. Pij and Piji are called false positive and false negative samples, respectively.
[0034] This invention establishes a target library, so that the targets in the image information are stored in the target library before the image data is detected. The data in the target library is organized and the information of the targets in the target library, such as volume and distance, is labeled. The target data is enriched in a targeted manner, so as to quickly analyze the data of the targets in the image during the later image processing.
[0035] By setting up a detection center and using deep learning methods to optimize it, the deep learning method can directly predict multiple categories of targets. It can predict multiple targets and use ResNet as a feature extractor in conjunction with the target library. It is combined with and verified with the data in the target library. The target library pyramid classification layer tree diagram and ResNet as a feature extractor are used to extract different patterns from the input image by expanding the network. Then, these feature maps are input into the pyramid pooling module to improve the detection quality and speed.
[0036] This invention continuously optimizes the target library and data detection center, thereby improving the efficiency and speed of image data detection during continuous use.
[0037] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0038] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A cloud computing-based image data processing method, characterized in that: Includes the following steps; S1: Image acquisition, based on image data uploaded by staff or users or stored in the database, or can be connected to the Internet to connect to and receive data on the Internet, or select the image processing data source to be used according to its own needs; S2: Target clustering processing, which preprocesses the image data collected from the Internet, collects the targets assigned to the images on the Internet, and if there is image data uploaded by staff or users, it is determined whether the staff assigns targets to it, establishes a target library, collects the targets uploaded by staff or users and the targets of image data obtained from the Internet, establishes a target pyramid, performs clustering processing on the image data, and obtains the cluster centers of the corresponding targets. When constructing the target pyramid using clustering, the direct features of each target are used as the starting point of the pyramid distribution map. The entire pyramid distribution map for different types of images is constructed from bottom to top. The association weights between all targets are calculated, and multiple direct features with the same weight are merged into a lower-level pyramid cluster item of the corresponding target pyramid. The association weights between all lower-level pyramid cluster items are calculated, and multiple lower-level pyramid cluster items with large association weights are merged into middle-level pyramid cluster items. This operation is repeated until the association weight between the two top-level pyramid cluster items is greater than zero. Then, a pyramid classification layer tree diagram is constructed. If the association weight between the two top-level cluster items is equal to zero, two or more pyramid classification layer tree diagrams are constructed respectively. This is combined with fully convolutional neural network technology to complete data processing. S3: Initial clustering of image data. Based on the target image library obtained in S2, cluster the experimental image data found on the Internet, establish a test center, receive the target types obtained in S2, analyze images similar to the target, conduct data verification and comparison, randomly collect different images on the Internet, and label the images according to the target features provided in S2, and import the labeled image data into the sample library. S4: Feature enhancement processing, where staff evaluate the data of the sample library detected by S3 and feed the detection results back to the S3 sample library for expansion and improvement, while labeling the target library with location and volume information; S5: Data processing. Establish a data detection center to perform image recognition on image data uploaded by staff or users or stored in the database through the target library. Arrange the images according to the proportion of each target on the image. Use semantic segmentation to identify the range of the target. Label and analyze each object in the image. Use object detection as a precondition for semantic segmentation to identify and label the bounding boxes of each object. This is also the process of pairing the target with each object in the image. Use fully convolutional neural network technology to test the test image. Through pixel-to-pixel mapping and end-to-end mapping, detect and analyze the range of the image and the distance between image objects to obtain image data. S6: Data Detection and Optimization: Compare the data obtained from S5 with the data obtained by users or staff to obtain the image data accuracy, and feed it back to the S5 data detection center for optimization.
2. The image data processing method based on cloud computing according to claim 1, characterized in that: After the image is acquired in S1, image enhancement is performed on the data uploaded by the staff.
3. The image data processing method based on cloud computing according to claim 1, characterized in that: In S2, during image preprocessing, the target data is annotated with the relative spatial size and actual volume data, providing data support for the abstraction and layering process in the semantic segmentation of the data detection center in S5.
4. The image data processing method based on cloud computing according to claim 1, characterized in that: The data detection center described in S5 establishes a pixel evaluation index library, which includes evaluation indices such as pixel accuracy, average pixel accuracy, average cross-union ratio, and weighted frequency cross-union ratio.
5. The image data processing method based on cloud computing according to claim 1, characterized in that: In S5, the data detection center uses deep learning methods to optimize it. In conjunction with the target library, it uses ResNet as a residual network as a feature extractor. By expanding the network, it extracts different patterns from the input image. Then, these feature maps are input into the pyramid pooling module to distinguish patterns at different scales. They are aggregated at four different scales, each scale corresponding to a pyramid layer, and processed by 1×1 convolutional layers to reduce their dimensionality. The output of the pyramid layer is upsampled and concatenated with the initial feature map to capture local and global contextual information. Finally, convolutional layers are used to generate pixel-wise predictions.
6. The image data processing method based on cloud computing according to claim 1, characterized in that: The accuracy of S5 semantic segmentation refers to the accuracy measurement of pixel-by-pixel labeling. Assuming there are k classes, from l0 to lk, one of which belongs to the background, Pij represents the number of pixels that belong to class i but are predicted to be class j, and Pii represents the number of truly correct class classifications. Pij and Pji are called false positive and false negative samples, respectively.