Machine-vision-oriented image processing system and training method therefor

By designing an image processing system with a multi-layer network structure, the low coding efficiency and compatibility problems of the existing machine vision system are solved, the compatibility of efficient generation of human eye and machine vision images is achieved, and the generalization ability and practicality of the system are improved.

WO2025208906A1PCT designated stage Publication Date: 2025-10-09CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136904
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2024-12-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

The existing machine vision system uses JPEG encoding, which is inefficient and cannot be directly compatible with existing systems and meet the visual needs of the human eye.

Method used

An image processing system for machine vision is designed, including a feature extraction network, an encoder, a decoder, a first image reconstruction network, a second image reconstruction network, a task network, a feature adaptation network and a feature task network. Feature extraction, encoding, decoding and reconstruction processing are performed through a multi-layer network structure to generate images for human eyes and machine vision.

Benefits of technology

It achieves the simultaneous generation of high-quality outputs for human vision and machine vision in one system, improves processing efficiency, is compatible with existing machine vision systems, meets the various needs of human vision and machine vision, and enhances the system's generalization ability and practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136904_09102025_PF_FP_ABST
    Figure CN2024136904_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A machine-vision-oriented image processing system and a training method therefor, which relate to the technical field of machine vision. The system comprises a feature extraction network, an encoder, a decoder, a first image reconstruction network for reconstructing a human-eye-vision-oriented image, a second image reconstruction network for reconstructing a machine-vision-oriented image, a feature adaptation network, a feature task network, and a task network. The system can perform feature extraction processing, encoding and decoding processing, human-eye-vision-oriented image reconstruction processing, and machine vision task processing on an image to be processed, so as to obtain a human-eye-vision-oriented reconstructed image and a machine vision task result. The machine-vision-oriented image processing system can execute machine vision tasks, is compatible with existing machine vision systems, can also obtain reconstructed human-eye-vision-oriented images that are easily understood and observed, and can meet the requirements of human eye vision.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing system and training method for machine vision

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application number 202410390046.2, filed on April 1, 2024, entitled “Image processing system and training method for machine vision”. The entire contents of this Chinese patent application are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of machine vision technology, and in particular to an image processing system for machine vision, a training method for an image processing system for machine vision, a training device for an image processing system for machine vision, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0004] Existing machine vision systems use the JPEG (Joint Photographic Experts Group) encoding method. Since this encoding method is inefficient, it is necessary to adopt advanced image coding technologies to improve coding efficiency, such as feature coding technologies for machine vision with high coding efficiency.

[0005] However, feature encoding technology for machine vision has system compatibility issues during deployment. That is, the existing machine vision system needs to be modified to use feature input, which is not directly compatible with the existing machine vision system and cannot meet the visual needs of the human eye.

[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0007] The present disclosure provides an image processing system for machine vision, a training method for an image processing system for machine vision, a training device for an image processing system for machine vision, an electronic device, a computer-readable storage medium, and a computer program product to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.

[0008] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0009] According to one aspect of the present disclosure, a machine vision-oriented image processing system is provided, comprising: a feature extraction network, an encoder, a decoder, a first reconstructed image network, a second reconstructed image network, a task network, a feature adaptation network, and a feature task network, wherein the first reconstructed image network is used to reconstruct images for human vision, and the second reconstructed image network is used to reconstruct images for machine vision; wherein the feature extraction network is used to extract features of an image to be processed and obtain the original features of the image to be processed; the encoder and the decoder are used to perform encoding and decoding processing on the original features of the image to be processed and obtain reconstructed features of the image to be processed; the first reconstructed image network is used to perform image reconstruction processing on the reconstructed features of the image to be processed and obtain a reconstructed image for human vision corresponding to the image to be processed; the second reconstructed image network and the task network are used to process the reconstructed features of the image to be processed and obtain a first machine vision task result corresponding to the image to be processed; the feature adaptation network and the feature task network are used to process the reconstructed features of the image to be processed and obtain a second machine vision task result corresponding to the image to be processed.

[0010] In some embodiments, the second reconstructed image network is further used to perform image reconstruction processing on the reconstructed features of the image to be processed to obtain a machine vision-oriented reconstructed image corresponding to the image to be processed, and input the machine vision-oriented reconstructed image corresponding to the image to be processed into the task network; the task network is also used to process the machine vision-oriented reconstructed image corresponding to the image to be processed to obtain the first machine vision task result.

[0011] In some embodiments, the feature adaptation network is further used to perform feature transformation processing on the reconstructed features of the image to be processed to obtain the transformed features of the image to be processed, and input the transformed features of the image to be processed into the feature task network; the feature spy network is further used to process the transformed features of the image to be processed to obtain the second machine vision task result.

[0012] According to another aspect of the present disclosure, a training method for a machine vision-oriented image processing system is provided, comprising: obtaining an initial feature extraction network, an initial encoder, an initial decoder, an initial first reconstructed image network, an initial second reconstructed image network, an initial feature adaptation network, a task network, and a feature-task network; training the initial feature extraction network and the initial feature adaptation network based on the feature-task network to obtain a preliminary trained feature extraction network and a preliminary trained feature adaptation network; training the preliminary trained feature extraction network, the preliminary trained feature adaptation network, the initial encoder, and the initial decoder based on the feature-task network to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder; training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the first reconstructed image network; training the initial second reconstructed image network based on the feature extraction network, the encoder, and the task network to obtain the second reconstructed image network; obtaining the machine vision-oriented image processing system based on the feature extraction network, the encoder, the decoder, the first reconstructed image network, the second reconstructed image network, the task network, the feature adaptation network, and the feature-task network.

[0013] In some embodiments, the initial feature extraction network and the initial feature adaptation network are trained based on the feature task network to obtain a preliminary trained feature extraction network and a preliminary trained feature adaptation network, including: obtaining a first sample image; inputting the first sample image into the feature task network to obtain a first eigenvalue of the first sample image passing through each feature layer in the feature task network; inputting the first sample image into the initial feature extraction network, the initial feature adaptation network and the feature task network in sequence to obtain a second eigenvalue of the first sample image passing through each feature layer in the feature task network; calculating a first loss based on the first eigenvalue of the first sample image passing through each feature layer in the feature task network and the second eigenvalue of the first sample image passing through each feature layer in the feature task network; adjusting the parameters of the initial feature extraction network and the initial feature adaptation network based on the first loss to obtain the preliminary trained feature extraction network and the preliminary trained feature adaptation network.

[0014] In some embodiments, the feature extraction network, the feature adaptation network, the initial encoder, and the initial decoder are trained based on the feature task network to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder, including: obtaining a second sample image; inputting the second sample image into the feature task network to obtain the first feature value of each feature layer of the feature task network; inputting the second sample image into the initially trained feature extraction network for feature extraction to obtain the original features of the second sample image; inputting the original features of the second sample image into the initial encoder for encoding processing to obtain the encoded features of the second sample image; inputting the encoded features of the second sample image into the initial decoder for decoding processing, Obtain reconstructed features of the second sample image; input the reconstructed features of the second sample image into the preliminarily trained feature adaptation network and the feature task network in sequence to obtain second feature values ​​of the second sample image passing through each feature layer in the feature task network; calculate a second loss based on the original features of the second sample image, the encoded features of the second sample image, the reconstructed features of the second sample image, the first feature values ​​of the second sample image passing through each feature layer in the feature task network, and the second feature values ​​of the second sample image passing through each feature layer in the feature task network; adjust the parameters of the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder according to the second loss to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder.

[0015] In some embodiments, the second loss is calculated based on the original features of the second sample image, the encoded features of the second sample image, the reconstructed features of the second sample image, the first feature values ​​of the second sample image passing through each feature layer in the feature task network, and the second feature values ​​of the second sample image passing through each feature layer in the feature task network, including: calculating the encoding loss based on the original features of the second sample image and the encoded features of the second sample image; calculating the feature reconstruction loss based on the original features of the second sample image and the reconstructed features of the second sample image; calculating the first multi-level feature loss based on the first feature values ​​of the second sample image passing through each feature layer in the feature task network and the second feature values ​​of the second sample image passing through each feature layer in the feature task network; and calculating the second loss based on the encoding loss, the feature reconstruction loss, and the first multi-level feature loss.

[0016] In some embodiments, the initial first reconstructed image network is trained based on the feature extraction network and the encoder to obtain the first reconstructed image network, including: obtaining a third sample image; inputting the third sample image into the feature extraction network, the encoder, the decoder and the initial first reconstructed image network in sequence to obtain a reconstructed image for human vision corresponding to the third sample image; calculating a third loss based on the third sample image and the reconstructed image for human vision corresponding to the third sample image; adjusting the parameters of the initial first reconstructed image network based on the third loss to obtain the first reconstructed image network.

[0017] In some embodiments, the method further includes: in the process of adjusting the parameters of the initial first reconstructed image network, adjusting the parameters of the decoder according to the third loss to obtain a new decoder, so as to use the new decoder to obtain the machine vision oriented image processing system.

[0018] In some embodiments, the training of the initial second reconstructed image network based on the feature extraction network, the encoder, and the task network to obtain the second reconstructed image network includes: obtaining a fourth sample image; inputting the fourth sample image into the task network for processing to obtain a first eigenvalue of the fourth sample image passing through each feature layer in the task network; inputting the fourth sample image into the feature extraction network, the encoder, the decoder, and the initial second reconstructed image network in sequence to obtain a machine vision-oriented reconstructed image corresponding to the fourth sample image; inputting the machine vision-oriented reconstructed image corresponding to the fourth sample image into the task network for processing to obtain a second eigenvalue of the fourth sample image passing through each feature layer in the task network; calculating a fourth loss based on the four sample images, the machine vision-oriented reconstructed image corresponding to the fourth sample image, the first eigenvalues ​​of the four sample images passing through each feature layer in the task network, and the second eigenvalues ​​of the fourth sample image passing through each feature layer in the task network; and adjusting the parameters of the initial second reconstructed image network based on the fourth loss to obtain the second reconstructed image network.

[0019] In some embodiments, the method further includes: in the process of adjusting the parameters of the initial second reconstructed image network, adjusting the parameters of the decoder according to the fourth loss to obtain a new decoder, so as to use the new decoder to obtain the machine vision oriented image processing system.

[0020] In some embodiments, the calculating of the fourth loss based on the four sample images, the machine vision-oriented reconstructed image corresponding to the fourth sample image, the first eigenvalues ​​of the four sample images passing through each feature layer in the task network, and the second eigenvalues ​​of the fourth sample image passing through each feature layer in the task network includes: calculating the image reconstruction loss based on the four sample images and the machine vision-oriented reconstructed image corresponding to the fourth sample image; calculating the second multi-level feature loss based on the first eigenvalues ​​of the four sample images passing through each feature layer in the task network and the second eigenvalues ​​of the four sample images passing through each feature layer in the task network; and calculating the fourth loss based on the image reconstruction loss and the second multi-level feature loss.

[0021] According to another aspect of the present disclosure, a training device for an image processing system for machine vision is provided, comprising: an initial network acquisition module for obtaining an initial feature extraction network, an initial encoder, an initial decoder, an initial first reconstructed image network, an initial second reconstructed image network, an initial feature adaptation network, a task network, and a feature task network; a first training module for training the initial feature extraction network and the initial feature adaptation network based on the feature task network to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network; a second training module for training the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder based on the feature task network. The system comprises a first training module, a second training module, a first encoder, and a second decoder, and a second training module for training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder; a third training module, for training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the first reconstructed image network; a fourth training module, for training the initial second reconstructed image network based on the feature extraction network, the encoder, and the task network to obtain the second reconstructed image network; a system generation module, for obtaining the image processing system for machine vision based on the feature extraction network, the encoder, the decoder, the first reconstructed image network, the second reconstructed image network, the task network, the feature adaptation network, and the feature task network.

[0022] According to another aspect of the present disclosure, an electronic device is also provided, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above-mentioned training method of the image processing system for machine vision by executing the executable instructions.

[0023] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the image processing system for machine vision is implemented.

[0024] According to another aspect of the present disclosure, a computer program product is further provided, including a computer program, which implements the above-mentioned training method for a machine vision-oriented image processing system when executed by a processor.

[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0027] FIG1 shows an architecture diagram of an image processing system for machine vision according to an embodiment of the present disclosure;

[0028] FIG2 shows a flow chart of a training method for a machine vision-oriented image processing system according to an embodiment of the present disclosure;

[0029] FIG3 shows a process diagram of preliminary training of a feature extraction network and a feature adaptation network in an embodiment of the present disclosure;

[0030] FIG4 shows a process diagram of training a feature extraction network, a feature adaptation network, an encoder, and a decoder in an embodiment of the present disclosure;

[0031] FIG5 shows a process diagram of training the first image reconstruction network in an embodiment of the present disclosure;

[0032] FIG6 shows a process diagram of training the second image reconstruction network in an embodiment of the present disclosure;

[0033] FIG7 shows a flowchart of an image processing method applied to an image processing system for machine video according to an embodiment of the present disclosure;

[0034] FIG8 shows a block diagram of a training device for an image processing system for machine vision according to an embodiment of the present disclosure;

[0035] FIG9 shows a structural block diagram of an electronic device according to an embodiment of the present disclosure;

[0036] FIG10 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0038] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0039] Figure 1 shows the architecture of an image processing system for machine vision according to an embodiment of the present disclosure. As shown in Figure 1 , the image processing system 100 for machine vision includes a feature extraction network 110, an encoder 120, a decoder 130, a first image reconstruction network 140, a second image reconstruction network 150, a task network 160, a feature adaptation network 170, and a feature-task network 180. The first image reconstruction network 140 is used to reconstruct images for human vision, while the second image reconstruction network 150 is used to reconstruct images for machine vision.

[0040] Specifically, the feature extraction network 110 can be used to extract features from the image to be processed and obtain the original features of the image to be processed. The feature extraction network 110 is an important part of the image processing system for machine vision, and its performance directly affects the quality of subsequent image reconstruction and machine vision tasks.

[0041] For example, the network before the stem layer in the ResNet X101 backbone network can be used as the feature extraction network 110. ResNet X101 is a deep residual network structure. The part before the stem layer generally includes several convolutional layers and a downsampling operation. It can efficiently extract low-level to mid-level general features from the original image. These features are universal for various machine vision tasks and can be applied to multiple fields such as target detection, image classification, and semantic segmentation.

[0042] For another example, the first two downsampling networks in an end-to-end image coding network (such as the design proposed by Cheng2020) can be used as the feature extraction network 110. End-to-end image coding networks are usually designed to efficiently compress images while retaining key information so that high-quality images can be reconstructed during decoding. In this process, the network gradually extracts multi-scale features of the image by downsampling layer by layer. The first two downsampling networks can often capture the low-frequency and high-frequency information of the image, which is crucial for subsequent image processing tasks. Therefore, the first two downsampling networks in the end-to-end image coding network are used as the feature extraction network 110. Through the first two downsampling operations, a compact and information-rich feature representation can be generated, which is very beneficial for subsequent image processing tasks (such as image reconstruction, target detection, classification, etc.).

[0043] The encoder 120 and the decoder 130 may be configured to perform encoding and decoding processing on original features of the image to be processed to obtain reconstructed features of the image to be processed.

[0044] As shown in Figure 1, the feature extraction network 110 extracts features from the image to be processed. After obtaining the original features of the image to be processed, the original features of the image to be processed can be passed to the encoder 120. The encoder 120 can then encode the original features of the image to be processed to obtain encoded features of the image to be processed, and then pass the encoded features of the image to be processed to the decoder 130. The decoder 130 can then decode the encoded features of the image to be processed to obtain reconstructed features of the image to be processed.

[0045] The encoder 120 may include a feature transformation layer, a feature quantization layer, a feature residual learning layer, a residual probability estimation layer, a residual probability estimation layer and an entropy coding layer. Among them, the feature transformation layer is responsible for mapping the input features to another feature space, which can be implemented through convolution, fully connected layers, etc. to extract higher-level feature representations. The feature quantization layer can quantize the transformed features and convert them into a discrete, finite set representation to reduce storage and transmission overhead. The feature residual learning layer can learn the residual between the input features and the quantized features, which helps to capture more subtle changes and may improve coding efficiency. For certain coding strategies, the residual probability estimation layer may need to estimate the probability distribution of the residual to support subsequent entropy coding steps. The entropy coding layer can perform lossless data compression based on the quantized features and the residual probability distribution, and encode the features into a compact binary representation.

[0046] The decoder 130 may include an entropy decoding layer, a feature inverse quantization layer, a feature residual learning layer, and a feature inverse transformation layer. Among them, the entropy decoding layer corresponds to the entropy coding layer of the encoder, and is responsible for decoding the compact binary code back to the original quantized feature representation. The feature inverse quantization layer can convert the decoded discrete feature representation back to a continuous feature space, which is the inverse process of the quantization step. The feature residual learning layer means that if the encoder uses feature residual learning, the decoder also needs a corresponding layer (i.e., a feature residual layer) to reconstruct the residual and combine it with the inverse quantized feature. The feature inverse transformation layer can map the decoded and reconstructed features back to the original feature space or the desired output format, involving the opposite operation of the feature transformation layer in the encoder.

[0047] For example, the portion of an end-to-end image coding network (such as the design proposed by Cheng2020) that follows the first two downsampling networks and precedes the final two upsampling networks can serve as the encoder and decoder. The encoder is responsible for downsampling the original features and gradually extracting high-level, abstract feature representations. This portion of the network structure typically includes multiple convolutions and pooling operations. Each downsampling step reduces the spatial dimension while increasing the channel dimension, thereby capturing more global and semantically rich information. In an end-to-end image coding network (such as the design proposed by Cheng2020), the portion after the first two downsampling steps is defined as the encoder stage, which compresses the original image data into a low-dimensional feature vector. The decoder, corresponding to the encoder, has the core task of restoring the compact feature maps generated by the encoder into spatial information at the original resolution. This process typically involves upsampling operations. In an end-to-end image coding network (such as the design proposed by Cheng2020), the portion before the final two upsampling networks serves as the decoder, whose goal is to reconstruct a high-quality image from the encoded features that is as close to the original input as possible.

[0048] The first reconstructed image network 140 can be used to perform image reconstruction processing on the reconstructed features of the image to be processed, and obtain a reconstructed image corresponding to the image to be processed that is oriented to human vision.

[0049] The first image reconstruction network 140 is used to reconstruct images for human vision. As shown in Figure 1, after obtaining the reconstructed features of the image to be processed, the decoder 130 can pass the reconstructed features of the image to be processed to the first image reconstruction network 140. The first image reconstruction network 140 can then perform image reconstruction processing on the reconstructed features of the image to be processed to obtain a reconstructed image corresponding to the image to be processed for human vision.

[0050] Specifically, the first reconstructed image network 140 may include a channel transformation module, a feature learning module, a feature upsampling module, and a post-processing module. The channel transformation module can adjust or convert the number of channels or arrangement of features to meet the requirements of subsequent processing steps. The feature learning module can further extract and learn key features in the image to optimize the reconstruction effect. The feature upsampling module can enlarge the spatial size of the feature map to the target resolution to match the size of the original image. The post-processing module can further optimize and adjust the reconstructed image to improve the visual effect and quality of the image.

[0051] In the disclosed embodiment, the first image reconstruction network 140 can be obtained based on an end-to-end image coding network (such as the design proposed by Cheng2020) or a Transformer network training, and can be specifically obtained by training according to the RD (Rate-Distortion) loss function. The RD loss function is intended to balance the bit rate or compression ratio (Rate) after encoding with the difference or distortion (Distortion) between the reconstructed image and the original image, ensuring that the visual quality of the original image is maintained as much as possible under limited bandwidth or storage resources.

[0052] The second reconstructed image network 150 and the task network 160 can be used to process the reconstructed features of the image to be processed to obtain a first machine vision task result corresponding to the image to be processed.

[0053] In some embodiments, the second reconstructed image network 150 is further used to: perform image reconstruction processing on the reconstructed features of the image to be processed, obtain a machine vision-oriented reconstructed image corresponding to the image to be processed, and input the machine vision-oriented reconstructed image corresponding to the image to be processed into the task network; the task network 160 is further used to: process the machine vision-oriented reconstructed image corresponding to the image to be processed to obtain the first machine vision task result.

[0054] The second reconstructed image network 150 is used to reconstruct images for machine vision. As shown in Figure 1, after obtaining the reconstructed features of the image to be processed, the decoder 130 can pass the reconstructed features of the image to be processed to the second reconstructed image network 150. The second reconstructed image network 150 can then perform image reconstruction processing on the reconstructed features of the image to be processed to obtain a reconstructed image for machine vision corresponding to the image to be processed. The second reconstructed image network 150 can then pass the reconstructed image for machine vision corresponding to the image to be processed to the task network 160. The task network 160 then performs a machine learning task based on the reconstructed image for machine vision corresponding to the image to be processed to obtain the first machine vision task result.

[0055] Specifically, the second image reconstruction network 150 may be similar to the first image reconstruction network 140 , including a channel transformation module, a feature learning module, a feature upsampling module, and a post-processing module.

[0056] In the disclosed embodiment, the design of the second image reconstruction network 150 can refer to the first image reconstruction network 140, that is, it can be based on an end-to-end image coding network (such as the design proposed by Cheng2020) or Transformer network training, and adjusted to the needs of machine vision tasks. For example, in the original RD loss function, distortion terms for machine vision tasks, such as target detection, semantic segmentation, depth estimation and other task-related indicators, can be added to ensure that the reconstructed image can better serve the execution of the machine vision algorithm.

[0057] Task network 160 can be understood as a common machine vision task network, such as the object detection network YOLO, the segmentation network UNet, and the tracking network JDE. These networks receive reconstructed features after encoding and decoding as input and output results related to the task. For example, in object detection, the task network outputs the location coordinates and category labels of objects in the image; in image segmentation, the task network outputs the classification label of each pixel in the image.

[0058] The feature adaptation network 170 and the feature task network 180 may be used to process the reconstructed features of the image to be processed to obtain a second machine vision task result corresponding to the image to be processed.

[0059] In some embodiments, the feature adaptation network 170 is further used to perform feature transformation processing on the reconstructed features of the image to be processed, obtain the transformed features of the image to be processed, and input the transformed features of the image to be processed into the feature task network; the feature spy network 180 is further used to process the transformed features of the image to be processed to obtain the second machine vision task result.

[0060] As shown in Figure 1, after obtaining the reconstructed features of the image to be processed, the decoder 130 can pass the reconstructed features of the image to be processed to the feature adaptation network 170, and then the feature adaptation network 170 can perform feature transformation processing on the reconstructed features of the image to be processed, such as channel transformation processing, to obtain the transformed features of the image to be processed. Then the feature adaptation network 170 can pass the transformed features of the image to be processed to the feature task network 180, and then the feature task network 180 performs a specific machine learning task based on the transformed features of the image to be processed to obtain a second machine vision task result.

[0061] The feature adaptation network 170 plays an important role in machine vision-oriented image processing systems, especially when processing tasks in different fields, different scenarios, or different data distributions. Its main function is to convert or adapt the input features so that these features can better meet the requirements of the subsequent feature task network 180. Specifically, the feature adaptation network 170 may include a channel transformation module for converting the input three-channel features into multi-channel features, thereby capturing and representing more image information to meet the requirements of subsequent machine vision tasks.

[0062] Feature-task network 180 is a machine vision task network that takes features as input. It can be a subnetwork structure that further processes reconstructed features to adapt or optimize them for another type of machine vision task. Unlike task network 160, feature-task network 180 focuses more on feature-level analysis and processing rather than directly generating task results. It can be used for tasks such as feature transformation, feature fusion, and feature comparison to support other machine vision tasks or systems.

[0063] The feature task network 180 adjusts the feature expression according to specific application requirements to make it more suitable for the needs of the second type of machine vision tasks, such as target detection of different scales, fine-grained classification, action recognition, or other tasks that require in-depth mining of image features.

[0064] The image processing system for machine vision provided in the embodiments of the present disclosure includes a feature extraction network, an encoder, a decoder, a first reconstructed image network for reconstructing human-eye vision, a second reconstructed image network for reconstructing machine vision, a feature adaptation network, a feature task network and a task network; the feature extraction network, the encoder and the decoder are used to extract features and perform encoding and decoding processing on the image to be processed to obtain reconstructed features of the image to be processed, and these reconstructed features are not only used for image reconstruction for human-eye vision, but also for processing machine vision tasks, so that the system can simultaneously generate multiple outputs for human eye and machine vision in one processing, thereby realizing full utilization of features and improving processing efficiency; the reconstructed features are processed by the first reconstructed image network Image reconstruction processing obtains high-quality, easy-to-understand reconstructed images for human vision to meet the needs of human vision; and the reconstructed features are processed through the second reconstructed image network and the task network to obtain the first machine vision task results corresponding to the image to be processed, which can generate accurate task results, meet the needs of machine vision applications, and be compatible with existing machine vision systems; and the reconstructed features are processed through the feature adaptation network and the feature task network to obtain the second machine vision task results corresponding to the image to be processed, which can perform deeper analysis and optimization on the reconstructed features in order to perform machine vision tasks corresponding to the feature task network, so that the entire system can efficiently and accurately process a variety of machine vision tasks, thereby improving the generalization ability and practicality of the system.

[0065] Furthermore, each component in the machine vision-oriented image processing system in the embodiment of the present disclosure can be independently optimized and upgraded, thereby enhancing the performance and functionality of the entire system. The system can also be easily integrated into other machine vision systems to achieve a wider range of applications.

[0066] Figure 2 shows a flowchart of a method for training a machine vision-oriented image processing system according to an embodiment of the present disclosure. This training method can be performed by any electronic device with computing processing capabilities, such as a server or terminal device. In some embodiments, this training method can be implemented by a training device for a machine vision-oriented image processing system. As shown in Figure 2, a machine vision-oriented image processing system can be trained by following the following steps.

[0067] Step S210: obtaining an initial feature extraction network, an initial encoder, an initial decoder, an initial first reconstructed image network, an initial second reconstructed image network, an initial feature adaptation network, a task network, and a feature-task network.

[0068] The task network and feature task network are predefined networks used to perform feature-based machine vision tasks. During training, they do not need to be trained. Therefore, in this step, the initial network structure of the task network, feature task network, and all other components of the machine vision-oriented image processing system are obtained, specifically including the initial feature extraction network, initial encoder, initial decoder, initial first image reconstruction network, initial second image reconstruction network, and initial feature adaptation network.

[0069] Step S220 , based on the feature task network, the initial feature extraction network and the initial feature adaptation network are trained to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network.

[0070] In this step, the initial feature extraction network and initial feature adaptation network are trained based on the feature task network, resulting in a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network. It should be noted that while the feature task network itself is not trained, performance feedback from the feature task network can guide training during the training process. The goal of this step is to optimize these two networks so that they can extract and adapt features that are useful to the feature task network.

[0071] Step S230: Based on the feature task network, the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder are trained to obtain a feature extraction network, a feature adaptation network, an encoder, and a decoder.

[0072] After obtaining the preliminarily trained feature extraction network and preliminarily trained feature adaptation network, the initial encoder and initial decoder are added and further training is performed based on the feature task network to obtain the trained feature extraction network, feature adaptation network, encoder, and decoder. The goal of this step is to ensure that the entire feature extraction, encoding, decoding, and feature adaptation process can work together to produce features that are useful to the feature task network.

[0073] Step S240: Training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain a first reconstructed image network.

[0074] Based on the trained feature extraction network and encoder, the initial first reconstruction network is trained. The goal of the first reconstruction network is to generate reconstructed images that are oriented towards human vision, which is very useful for debugging and visualization.

[0075] Step S250: Based on the feature extraction network, the encoder, and the task network, the initial second reconstructed image network is trained to obtain a second reconstructed image network.

[0076] The initial second image reconstruction network is trained by combining the trained feature extraction network, encoder, and task network. The goal of this second image reconstruction network is to generate machine vision-oriented reconstructed images that will be directly used in subsequent machine vision tasks. It should be noted that while the task network itself is not trained, it provides task-related guidance during training.

[0077] Step S260: obtaining an image processing system for machine vision according to the feature extraction network, the encoder, the decoder, the first image reconstruction network, the second image reconstruction network, the task network, the feature adaptation network, and the feature task network.

[0078] In this step, all trained feature extraction networks, encoders, decoders, first image reconstruction networks, second image reconstruction networks, and feature adaptation networks are integrated with pre-defined task networks and feature task networks to form a complete image processing system for machine vision. This system can perform feature extraction, encoding and decoding, image reconstruction for human vision, image reconstruction for machine vision, machine vision task processing through the task network, feature adaptation, and specific machine vision task processing through the feature task network.

[0079] It can be seen from steps S210 to S260 that the training method for the machine vision-oriented image processing system provided by the embodiment of the present disclosure mainly includes the following four training stages.

[0080] The network structure of the first training stage includes feature extraction network, feature adaptation network and feature task network.

[0081] The network structure of the second training phase includes a feature extraction network, an encoder, a decoder, a feature adaptation network, and a feature task network. The feature extraction network and the feature adaptation network can load the weights generated in the first training phase.

[0082] The network structure of the third training phase includes a feature extraction network, an encoder, a decoder, and a first image reconstruction network. The feature extraction network and encoder can load the weights generated in the second training phase and freeze the weights during this phase, meaning that the feature extraction network and encoder are not trained in the third phase. Furthermore, the decoder can load the weights generated in the second training phase and freeze the weights during this phase, meaning that the decoder can be trained or not trained in the third phase.

[0083] The network structure of the fourth training phase includes a feature extraction network, an encoder, a decoder, a second image reconstruction network, and a task network. The feature extraction network and encoder can load the weights generated in the second training phase and freeze the weights during this phase, meaning that the feature extraction network and encoder are not trained in the fourth phase. Furthermore, the decoder can load the weights generated in the second training phase and freeze the weights during this phase, meaning that the decoder can be trained or not trained in the fourth phase.

[0084] Next, the above four training stages are described in detail with reference to FIG3 to FIG6 .

[0085] Figure 3 illustrates a process diagram for initially training a feature extraction network and a feature adaptation network in accordance with an embodiment of the present disclosure. As shown in Figure 3 , the process of training an initial feature extraction network and an initial feature adaptation network based on a feature task network to obtain a initially trained feature extraction network and a initially trained feature adaptation network may include the following steps.

[0086] Step S310: Obtain a first sample image.

[0087] A plurality of images are selected from a training data set as first sample images.

[0088] Step S320: input the first sample image into the feature task network to obtain the first feature value of the first sample image after passing through each feature layer in the feature task network.

[0089] The first sample image is input into a predefined feature task network, which includes multiple feature layers. Each feature layer can extract different features of the image, and obtain the first feature value of the first sample image passing through each feature layer in the feature task network.

[0090] Step S330: input the first sample image into the initial feature extraction network, the initial feature adaptation network, and the feature task network in sequence to obtain the second feature value of the first sample image after passing through each feature layer in the feature task network.

[0091] The first sample image is passed through the initial feature extraction network to extract features, obtaining extracted features. These extracted features are then passed to the initial feature adaptation network for further feature transformation, obtaining transformed features. The transformed features are then input into the feature task network to obtain the second feature values ​​of the first sample image at each feature layer in the feature task network.

[0092] Step S340 , calculating a first loss according to the first eigenvalues ​​of the first sample image passing through each feature layer in the feature task network and the second eigenvalues ​​of the first sample image passing through each feature layer in the feature task network.

[0093] Among them, the calculation formula of the first loss is: L 1i =MSE(F 1(0,i), F 1(1,i) ) (2)

[0094] In the formula, Loss1 is the first loss; j is the number of feature layers in the feature task network; ω i represents the weight corresponding to the i-th feature layer of the feature task network; F 1(0,i) is the first eigenvalue of the first sample image after passing through the i-th feature layer of the feature task network; F 1(1,i) is the second eigenvalue of the first sample image after passing through the i-th feature layer of the feature task network, that is, the second eigenvalue of the first sample image after passing through the initial feature extraction network, the initial feature adaptation network, and then through the i-th feature layer of the feature task network; MSE() is the calculation of the mean square error.

[0095] Step S350: Adjust the parameters of the initial feature extraction network and the initial feature adaptation network according to the first loss to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network.

[0096] Based on the calculated first loss, the backpropagation algorithm is used to adjust the parameters of the initial feature extraction network and the initial feature adaptation network to reduce the loss value. After multiple iterations, the two networks will gradually converge to obtain the initially trained feature extraction network and the initially trained feature adaptation network.

[0097] Figure 4 illustrates a process diagram for training a feature extraction network, a feature adaptation network, an encoder, and a decoder in accordance with an embodiment of the present disclosure. As shown in Figure 4 , the process of training a pre-trained feature extraction network, a pre-trained feature adaptation network, an initial encoder, and an initial decoder based on a feature task network to obtain a feature extraction network, a feature adaptation network, an encoder, and a decoder may include the following steps.

[0098] Step S410: Obtain a second sample image.

[0099] A plurality of images are selected from the training data set as second sample images, wherein the second sample images may be the same as or different from the first sample images, which is not limited in the embodiment of the present disclosure.

[0100] Step S420: input the second sample image into the feature task network to obtain the first feature value of the second sample image after passing through each feature layer in the feature task network.

[0101] Step S430: Input the second sample image into the preliminarily trained feature extraction network to perform feature extraction to obtain original features of the second sample image.

[0102] Step S440: Input the original features of the second sample image into the initial encoder for encoding processing to obtain the encoded features of the second sample image.

[0103] Step S450: Input the encoded features of the second sample image into the initial decoder for decoding processing to obtain reconstructed features of the second sample image.

[0104] Step S460: input the reconstructed features of the second sample image into the preliminarily trained feature adaptation network and feature task network in sequence to obtain the second feature values ​​of the second sample image after passing through each feature layer in the feature task network.

[0105] The reconstructed features of the second sample image are input into the preliminarily trained feature adaptation network for further feature transformation to obtain the transformed features. The transformed features are then input into the feature task network to obtain the second feature values ​​of the second sample image after passing through each feature layer in the feature task network.

[0106] Step S470, calculate the second loss based on the original features of the second sample image, the encoded features of the second sample image, the reconstructed features of the second sample image, the first feature values ​​of the second sample image passing through each feature layer in the feature task network, and the second feature values ​​of the second sample image passing through each feature layer in the feature task network.

[0107] In some embodiments, the second loss can be calculated as follows: the encoding loss is calculated based on the original features of the second sample image and the encoding features of the second sample image; the feature reconstruction loss is calculated based on the original features of the second sample image and the reconstructed features of the second sample image; the first multi-level feature loss is calculated based on the first feature value of the second sample image passing through each feature layer in the feature task network and the second feature value of the second sample image passing through each feature layer in the feature task network; the second loss is calculated based on the encoding loss, the feature reconstruction loss and the first multi-level feature loss.

[0108] Among them, the calculation formula of the second loss is: D=MSE(F0,F rec ) (4) L 2i =MSE(F 2(0,i), F 2(1,i) ) (5)

[0109] In the formula, Loss2 is the second loss; L bppis the encoding loss, which can be calculated by the original features of the second sample image and the encoding features of the second sample image; D is the feature reconstruction loss, which can be calculated by the original features F0 of the second sample and the reconstructed features F0 of the second sample. rec Calculated; α and β are the set weight values, which can be set according to requirements; j is the number of feature layers in the feature task network; ω i represents the weight corresponding to the i-th feature layer of the feature task network; F 2(0,i) is the first eigenvalue of the second sample image after the i-th feature layer of the feature task network; F 2(1,i) is the second eigenvalue of the second sample image after the i-th feature layer of the feature task network; MSE() is the calculation of the mean square error.

[0110] Step S480: Adjust the parameters of the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder according to the second loss to obtain a feature extraction network, a feature adaptation network, an encoder, and a decoder.

[0111] Based on the calculated second loss, the backpropagation algorithm adjusts the parameters of the initially trained feature extraction network, the initially trained feature adaptation network, the initial encoder, and the initial decoder to reduce the loss. After multiple iterations, the various components gradually converge, resulting in the trained feature extraction network, feature adaptation network, encoder, and decoder.

[0112] Figure 5 shows a process diagram of training the first reconstructed image network in an embodiment of the present disclosure. As shown in Figure 5, the process of training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the first reconstructed image network may include the following steps.

[0113] Step S510: Obtain a third sample image.

[0114] A plurality of images are selected from the training data set as third sample images, wherein the third sample image may be the same as or different from the first sample image or the second sample image, which is not limited in the embodiment of the present disclosure.

[0115] Step S520: input the third sample image into the feature extraction network, the encoder, the decoder, and the initial first reconstructed image network in sequence to obtain a reconstructed image corresponding to the third sample image for human vision.

[0116] The third sample image is input into the feature extraction network for feature extraction, obtaining the original features of the third sample image. The original features are then passed to the encoder, which encodes the original features to obtain encoded features of the third sample image. The encoded features are then passed to the decoder, which decodes the encoded features to obtain reconstructed features of the third sample image. Finally, the reconstructed features are passed to the initial first reconstruction network, which performs image reconstruction on the reconstructed features to obtain a human-viewing reconstructed image corresponding to the third sample image.

[0117] Step S530 : Calculate a third loss based on the third sample image and the human-eye-oriented reconstructed image corresponding to the third sample image.

[0118] Among them, the calculation formula of the third loss is:

[0119] In the formula, Loss3 is the third loss, I 10 is the third sample image, I 1rec is the reconstructed image for human vision corresponding to the third sample image, L is the loss function, which can be the L1 loss function (also known as the absolute error loss), and the loss function is the loss between the third sample image I0 and the reconstructed image for human vision corresponding to the third sample image I rec The sum of the absolute values ​​of the pixel value differences between them.

[0120] Step S540: Adjust the parameters of the initial first reconstructed image network according to the third loss to obtain the first reconstructed image network.

[0121] Based on the calculated third loss, the back-propagation algorithm is used to adjust the parameters of the initial first image reconstruction network to reduce the loss value. After multiple iterations, the network will gradually converge to obtain the trained first image reconstruction network.

[0122] In some embodiments, the method further includes: in the process of adjusting the parameters of the initial first image reconstruction network, adjusting the parameters of the decoder according to the third loss to obtain a new decoder, so as to use the new decoder to obtain an image processing system for machine vision.

[0123] In the disclosed embodiments, during the training of the first image reconstruction network, the decoder weights can be frozen, i.e., the decoder parameters are not adjusted. Alternatively, the decoder weights can be frozen, i.e., the decoder parameters are adjusted. Specifically, based on the calculated third loss, the parameters of the initial first image reconstruction network and the decoder are adjusted via a backpropagation algorithm to reduce the loss. Finally, a trained first image reconstruction network and a new decoder are obtained. The new decoder can then be used to generate an image processing system for machine vision.

[0124] Figure 6 shows a process diagram of training the second image reconstruction network in an embodiment of the present disclosure. As shown in Figure 6, the process of training the initial second image reconstruction network based on the feature extraction network, the encoder and the task network to obtain the second image reconstruction network may include the following steps.

[0125] Step S610: Obtain a fourth sample image.

[0126] A plurality of images are selected from the training data set as fourth sample images, wherein the fourth sample image may be the same as or different from the first sample image, the second sample image, or the third sample image, and this is not limited in the embodiment of the present disclosure.

[0127] Step S620: Input the fourth sample image into the task network for processing to obtain the first eigenvalues ​​of the fourth sample image after passing through each feature layer in the task network.

[0128] The fourth sample image is input into a predefined task network, which includes multiple feature layers, each of which can extract different features of the image, and obtains the first feature value of the fourth sample image passing through each feature layer in the task network.

[0129] Step S630: Input the fourth sample image into the feature extraction network, the encoder, the decoder, and the initial second reconstructed image network in sequence to obtain a machine vision-oriented reconstructed image corresponding to the fourth sample image.

[0130] The fourth sample image is input into the feature extraction network for feature extraction, obtaining the original features of the fourth sample image. The original features are then passed to the encoder, which encodes the original features to obtain encoded features of the fourth sample image. The encoded features are then passed to the decoder, which decodes the encoded features to obtain reconstructed features of the fourth sample image. Finally, the reconstructed features are passed to the initial second image reconstruction network, which performs image reconstruction on the reconstructed features to obtain a machine vision-oriented reconstructed image corresponding to the fourth sample image.

[0131] Step S640: Input the machine vision-oriented reconstructed image corresponding to the fourth sample image into the task network for processing to obtain the second eigenvalues ​​of the fourth sample image passing through each feature layer in the task network.

[0132] Step S650, calculate the fourth loss based on the four sample images, the machine vision-oriented reconstructed image corresponding to the fourth sample image, the first eigenvalues ​​of the four sample images passing through each feature layer in the task network, and the second eigenvalues ​​of the fourth sample image passing through each feature layer in the task network.

[0133] In some embodiments, calculating the fourth loss based on the four sample images, the machine vision-oriented reconstructed image corresponding to the fourth sample image, the first eigenvalues ​​of the four sample images through each feature layer in the task network, and the second eigenvalues ​​of the fourth sample image through each feature layer in the task network can include: calculating the image reconstruction loss based on the four sample images and the machine vision-oriented reconstructed image corresponding to the fourth sample image; calculating the second multi-level feature loss based on the first eigenvalues ​​of the four sample images through each feature layer in the task network and the second eigenvalues ​​of the four sample images through each feature layer in the task network; calculating the fourth loss based on the image reconstruction loss and the second multi-level feature loss.

[0134] Among them, the calculation formula of the fourth loss is: L 3i =MSE(F 3(0,i), F 3(1,i) ) (8)

[0135] In the formula, Loss4 is the fourth loss; I 20 is the fourth sample image, I 2rec is the reconstructed image for machine vision corresponding to the fourth sample image, L is the loss function, which can be the L1 loss function (also known as the absolute error loss), and the fourth sample image I is calculated. 20 The reconstructed image I for machine vision corresponding to the fourth sample image 2rec The sum of the absolute values ​​of the pixel value differences between them, θ is the set weight value, which can be set according to requirements; j is the number of feature layers in the task network; ω i represents the weight corresponding to the i-th feature layer of the task network; F 3(0,i) is the first eigenvalue of the fourth sample image after passing through the i-th feature layer of the feature task network; F 3(1,i) is the second eigenvalue of the fourth sample image after the i-th feature layer of the feature task network; MSE() is the calculation of the mean square error.

[0136] Step S660: Adjust the parameters of the initial second reconstructed image network according to the fourth loss to obtain a second reconstructed image network.

[0137] Based on the calculated fourth loss, the back-propagation algorithm is used to adjust the parameters of the initial second image reconstruction network to reduce the loss value. After multiple iterations, the network will gradually converge to obtain the trained second image reconstruction network.

[0138] In some embodiments, the method further includes: in the process of adjusting the parameters of the initial second image reconstruction network, adjusting the parameters of the decoder according to the fourth loss to obtain a new decoder, so as to use the new decoder to obtain an image processing system for machine vision.

[0139] In the disclosed embodiments, during the training of the second image reconstruction network, the decoder weights can be frozen, i.e., the decoder parameters are not adjusted. Alternatively, the decoder weights can be frozen, i.e., the decoder parameters are adjusted. Specifically, based on the calculated fourth loss, the parameters of the initial second image reconstruction network and the decoder are adjusted via a backpropagation algorithm to reduce the loss. Finally, a trained second image reconstruction network and a new decoder are obtained. The new decoder can then be used to generate an image processing system for machine vision.

[0140] The training method for a machine vision-oriented image processing system provided by the embodiment of the present disclosure first utilizes a feature task network to optimize a feature extraction network and a feature adaptation network during the training process to obtain a preliminarily trained feature extraction network and feature adaptation network. The encoder, decoder, and preliminarily trained feature extraction network and feature adaptation network are then trained. Finally, the first image reconstruction network and the second image reconstruction network are trained based on the preliminarily trained feature extraction network and feature adaptation network. This step-by-step iterative approach helps to ensure the effectiveness of information transfer and gradient feedback between the various parts, so that each network structure can reach the optimal state within the overall framework. In addition, the task network and the feature task network directly use pre-defined models, avoiding the huge computing resources and time cost required for training from scratch, while also utilizing existing learning results to improve the generalization ability and stability of the machine vision-oriented image processing system.

[0141] Figure 7 shows a flowchart of an image processing method applied to an image processing system for machine videos according to an embodiment of the present disclosure. As shown in Figure 7 , the image processing method may include the following steps.

[0142] Step S710: obtaining an image to be processed, inputting the image to be processed into a feature extraction network for feature extraction, and obtaining original features of the image to be processed;

[0143] Step S720: inputting the original features of the image to be processed into an encoder for encoding processing to obtain encoded features of the image to be processed;

[0144] Step S730: inputting the encoded features of the image to be processed into a decoder for decoding to obtain reconstructed features of the image to be processed;

[0145] Step S740: inputting the reconstructed features of the image to be processed into a first image reconstruction network for image reconstruction processing to obtain a reconstructed image corresponding to the image to be processed and adapted to human vision, so as to display the reconstructed image to a user;

[0146] Step S750: inputting the reconstructed features of the image to be processed into a second image reconstruction network for image reconstruction processing to obtain a machine vision-oriented reconstructed image corresponding to the image to be processed;

[0147] Step S760: inputting the machine vision-oriented reconstructed image corresponding to the image to be processed into the task network to perform the machine vision task related to the task network, thereby obtaining a first machine vision task result;

[0148] Step S770: inputting the reconstructed features of the image to be processed into a feature adaptation network for feature transformation processing to obtain the transformed features of the image to be processed;

[0149] Step S780: Input the transformed features of the image to be processed into the feature task network to perform the machine vision task related to the feature task network, and obtain a second machine vision task result.

[0150] The image processing method of the disclosed embodiment applied to the image processing system for machine vision generates reconstructed images that adapt to different needs through reconstructed image networks for human vision and machine vision. The reconstructed images for human vision focus more on visual effects and user experience, while the reconstructed images for machine vision focus more on the accuracy and completeness of image information, which helps to efficiently execute machine vision tasks; it can execute machine vision tasks related to the task network and machine vision tasks related to the feature task network, flexibly adapt to different application scenarios and needs, and improve the practicality and scalability of the system; the reconstructed features are transformed through the feature adaptation network, and the feature representation is further adjusted and optimized to adapt to the needs of different machine vision tasks. This feature adaptation capability enhances the flexibility and adaptability of the system, enabling the system to better cope with complex and changeable machine vision tasks.

[0151] Figure 8 shows a block diagram of a training device for a machine vision image processing system according to an embodiment of the present disclosure. As shown in Figure 8 , the training device 800 includes an initial network acquisition module 810 , a first training module 820 , a second training module 830 , a third training module 840 , a fourth training module 850 , and a system generation module 860 .

[0152] Specifically, the initial network acquisition module 810 can be used to obtain an initial feature extraction network, an initial encoder, an initial decoder, an initial first reconstructed image network, an initial second reconstructed image network, an initial feature adaptation network, a task network, and a feature task network; the first training module 820 can be used to train the initial feature extraction network and the initial feature adaptation network based on the feature task network to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network; the second training module 830 can be used to train the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder based on the feature task network. , obtaining a feature extraction network, a feature adaptation network, an encoder and a decoder; the third training module 840 can be used to: based on the feature extraction network and the encoder, train the initial first reconstructed image network to obtain the first reconstructed image network; the fourth training module 850 can be used to: based on the feature extraction network, the encoder and the task network, train the initial second reconstructed image network to obtain the second reconstructed image network; the system generation module 860 can be used to: according to the feature extraction network, the encoder, the decoder, the first reconstructed image network, the second reconstructed image network, the task network, the feature adaptation network and the feature task network, obtain an image processing system for machine vision.

[0153] In some embodiments, the first training module 820 can also be used to: obtain a first sample image; input the first sample image into the feature task network to obtain the first eigenvalue of the first sample image passing through each feature layer in the feature task network; input the first sample image into the initial feature extraction network, the initial feature adaptation network and the feature task network in sequence to obtain the second eigenvalue of the first sample image passing through each feature layer in the feature task network; calculate the first loss based on the first eigenvalue of the first sample image passing through each feature layer in the feature task network and the second eigenvalue of the first sample image passing through each feature layer in the feature task network; adjust the parameters of the initial feature extraction network and the initial feature adaptation network according to the first loss to obtain a preliminary trained feature extraction network and a preliminary trained feature adaptation network.

[0154] In some embodiments, the second training module 830 may further be used to: obtain a second sample image; input the second sample image into the feature task network to obtain a first feature value of the second sample image passing through each feature layer in the feature task network; input the second sample image into the preliminarily trained feature extraction network for feature extraction to obtain the original features of the second sample image; input the original features of the second sample image into the initial encoder for encoding processing to obtain the encoded features of the second sample image; input the encoded features of the second sample image into the initial decoder for decoding processing to obtain the reconstructed features of the second sample image; input the reconstructed features of the second sample image into the preliminarily trained feature adaptation network and feature task network in sequence to obtain the second feature values ​​of the second sample image passing through each feature layer in the feature task network; calculate a second loss based on the original features of the second sample image, the encoded features of the second sample image, the reconstructed features of the second sample image, the first feature values ​​of the second sample image passing through each feature layer in the feature task network, and the second feature values ​​of the second sample image passing through each feature layer in the feature task network; and adjust the parameters of the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder based on the second loss to obtain a feature extraction network, a feature adaptation network, an encoder, and a decoder.

[0155] In some embodiments, the second training module 830 can also be used to: calculate the coding loss based on the original features of the second sample image and the coding features of the second sample image; calculate the feature reconstruction loss based on the original features of the second sample image and the reconstructed features of the second sample image; calculate the first multi-level feature loss based on the first feature values ​​of the second sample image passing through each feature layer in the feature task network and the second feature values ​​of the second sample image passing through each feature layer in the feature task network; calculate the second loss based on the coding loss, the feature reconstruction loss and the first multi-level feature loss.

[0156] In some embodiments, the third training module 840 can also be used to: obtain a third sample image; input the third sample image into the feature extraction network, encoder, decoder and initial first reconstruction image network in sequence to obtain a reconstructed image corresponding to the third sample image for human eye vision; calculate a third loss based on the third sample image and the reconstructed image corresponding to the third sample image for human eye vision; adjust the parameters of the initial first reconstruction image network based on the third loss to obtain a first reconstructed image network.

[0157] In some embodiments, the third training module 840 can also be used to: in the process of adjusting the parameters of the initial first reconstruction image network, adjust the parameters of the decoder according to the third loss to obtain a new decoder, so as to use the new decoder to obtain an image processing system for machine vision.

[0158] In some embodiments, the fourth training module 850 can also be used to: obtain a fourth sample image; input the fourth sample image into the task network for processing to obtain the first eigenvalue of the fourth sample image passing through each feature layer in the task network; input the fourth sample image into the feature extraction network, encoder, decoder and initial second reconstructed image network in sequence to obtain a machine vision-oriented reconstructed image corresponding to the fourth sample image; input the machine vision-oriented reconstructed image corresponding to the fourth sample image into the task network for processing to obtain the second eigenvalue of the fourth sample image passing through each feature layer in the task network; calculate the fourth loss based on the four sample images, the machine vision-oriented reconstructed image corresponding to the fourth sample image, the first eigenvalue of the four sample images passing through each feature layer in the task network and the second eigenvalue of the fourth sample image passing through each feature layer in the task network; adjust the parameters of the initial second reconstructed image network according to the fourth loss to obtain a second reconstructed image network.

[0159] In some embodiments, the fourth training module 850 can also be used to: in the process of adjusting the parameters of the initial second reconstruction image network, adjust the parameters of the decoder according to the fourth loss to obtain a new decoder, so as to use the new decoder to obtain an image processing system for machine vision.

[0160] In some embodiments, the fourth training module 850 can also be used to: calculate the image reconstruction loss based on the machine vision-oriented reconstructed images corresponding to the four sample images and the fourth sample image; calculate the second multi-level feature loss based on the first eigenvalues ​​of each feature layer in the task network and the second eigenvalues ​​of each feature layer in the task network of the four sample images; calculate the fourth loss based on the image reconstruction loss and the second multi-level feature loss.

[0161] Since the principle of solving the problem in this device embodiment is similar to the above-mentioned training method embodiment for the image processing system for machine vision, the implementation of this device embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0162] The electronic device 900 according to this embodiment of the present disclosure is described below with reference to Figure 9. The electronic device 900 shown in Figure 9 is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0163] As shown in Figure 9, electronic device 900 is implemented as a general-purpose computing device. Components of electronic device 900 may include, but are not limited to, the aforementioned at least one processing unit 910, the aforementioned at least one storage unit 920, and a bus 930 connecting various system components (including storage unit 920 and processing unit 910).

[0164] The storage unit stores program codes, which can be executed by the processing unit 910, so that the processing unit 910 performs the steps described in the above “Exemplary Method” section of this specification according to various exemplary embodiments of the present disclosure.

[0165] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 9201 and / or a cache memory unit 9202 , and may further include a read-only memory unit (ROM) 9203 .

[0166] The storage unit 920 may also include a program / utility 9204 having a set (at least one) of program modules 9205, such program modules 9205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0167] Bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0168] The electronic device 900 can also communicate with one or more external devices 940 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 900, and / or any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 950. Furthermore, the electronic device 900 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 900, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0169] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0170] In particular, according to an embodiment of the present disclosure, the process described in the above reference flowchart can be implemented as a computer program product, which includes: a computer program, which, when executed by a processor, implements any one of the above-mentioned video encoding methods for machine vision tasks; or any one of the above-mentioned video decoding methods for machine vision tasks.

[0171] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is further provided, which may be a readable signal medium or a readable storage medium. FIG10 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure. As shown in FIG10 , a program product capable of implementing the above-mentioned method of the present disclosure is stored on the computer-readable storage medium 1000. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.

[0172] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0173] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0174] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0175] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0176] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0177] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0178] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure. Industrial Applicability

[0179] The present disclosure is applicable to the field of machine vision technology. For different machine vision tasks, the reconstructed features can be processed through the second reconstructed image network and the task network, and the reconstructed features can be processed through the feature adaptation network and the feature task network, which can generate accurate task results, meet the needs of machine vision applications, and be compatible with existing machine vision systems; the first reconstructed image network is used to perform image reconstruction processing based on the visual characteristics of the human eye, and high-quality, easy-to-understand reconstructed images can be generated to meet the needs of human vision; the feature extraction network, encoder and decoder are used to extract features and perform encoding and decoding processing on the processed image to obtain reconstructed features. These reconstructed features are not only used for image reconstruction for human eye vision, but also for processing machine vision tasks, so that the system can simultaneously generate multiple outputs for human eye and machine vision in one processing, realize full utilization of features, and improve processing efficiency; and each component in the system can be independently optimized and upgraded, thereby enhancing the performance and function of the entire system. The system is also easy to be integrated into other machine vision systems to achieve wider application.

[0180] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. An image processing system for machine vision, wherein: include: A feature extraction network, an encoder, a decoder, a first reconstructed image network, a second reconstructed image network, a task network, a feature adaptation network, and a feature task network. The first reconstructed image network is used to reconstruct images for human vision, and the second reconstructed image network is used to reconstruct images for machine vision. The feature extraction network is used to extract features from the image to be processed to obtain the original features of the image to be processed; The encoder and the decoder are used to perform encoding and decoding processing on the original features of the image to be processed to obtain reconstructed features of the image to be processed; The first image reconstruction network is used to perform image reconstruction processing on the reconstructed features of the image to be processed to obtain a reconstructed image corresponding to the image to be processed that is oriented to human vision; The second reconstructed image network and the task network are used to process the reconstructed features of the image to be processed to obtain a first machine vision task result corresponding to the image to be processed; The feature adaptation network and the feature task network are used to process the reconstructed features of the image to be processed to obtain a second machine vision task result corresponding to the image to be processed.

2. The system according to claim 1, wherein: The second reconstructed image network is further configured to perform image reconstruction processing on the reconstructed features of the image to be processed, obtain a machine vision-oriented reconstructed image corresponding to the image to be processed, and input the machine vision-oriented reconstructed image corresponding to the image to be processed into the task network; The task network is further used to process the machine vision-oriented reconstructed image corresponding to the image to be processed to obtain the first machine vision task result.

3. The system according to claim 1, wherein: The feature adaptation network is further configured to perform feature transformation processing on the reconstructed features of the image to be processed, obtain the transformed features of the image to be processed, and input the transformed features of the image to be processed into the feature task network; The feature spy network is further used to process the transformed features of the image to be processed to obtain the second machine vision task result.

4. A training method for an image processing system for machine vision, wherein: include: Obtaining an initial feature extraction network, an initial encoder, an initial decoder, an initial first reconstructed image network, an initial second reconstructed image network, an initial feature adaptation network, a task network, and a feature task network; Based on the feature task network, the initial feature extraction network and the initial feature adaptation network are trained to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network; Based on the feature task network, the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder are trained to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder; Training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the first reconstructed image network; Training the initial second reconstructed image network based on the feature extraction network, the encoder, and the task network to obtain the second reconstructed image network; The machine vision-oriented image processing system is obtained according to the feature extraction network, the encoder, the decoder, the first reconstructed image network, the second reconstructed image network, the task network, the feature adaptation network and the feature task network.

5. The method according to claim 4, wherein The step of training the initial feature extraction network and the initial feature adaptation network based on the feature task network to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network includes: obtaining a first sample image; Inputting the first sample image into the feature task network to obtain first feature values ​​of the first sample image after passing through each feature layer in the feature task network; Inputting the first sample image into the initial feature extraction network, the initial feature adaptation network, and the feature task network in sequence, and obtaining a second feature value of the first sample image after passing through each feature layer in the feature task network; Calculating a first loss according to first eigenvalues ​​of the first sample image passing through each feature layer in the feature task network and second eigenvalues ​​of the first sample image passing through each feature layer in the feature task network; Parameters of the initial feature extraction network and the initial feature adaptation network are adjusted according to the first loss to obtain the preliminarily trained feature extraction network and the preliminarily trained feature adaptation network.

6. The method according to claim 4, wherein: The method of training the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder based on the feature task network to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder includes: obtaining a second sample image; Inputting the second sample image into the feature task network to obtain first feature values ​​of the second sample image after passing through each feature layer in the feature task network; Inputting the second sample image into the preliminarily trained feature extraction network to perform feature extraction to obtain original features of the second sample image; Inputting the original features of the second sample image into the initial encoder for encoding processing to obtain the encoded features of the second sample image; Inputting the encoded features of the second sample image into the initial decoder for decoding to obtain reconstructed features of the second sample image; inputting the reconstructed features of the second sample image into the preliminarily trained feature adaptation network and the feature task network in sequence, and obtaining second feature values ​​of the second sample image after passing through each feature layer in the feature task network; Calculating a second loss according to the original features of the second sample image, the encoded features of the second sample image, the reconstructed features of the second sample image, the first feature values ​​of the second sample image passing through each feature layer in the feature task network, and the second feature values ​​of the second sample image passing through each feature layer in the feature task network; Adjust the parameters of the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder according to the second loss to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder.

7. The method according to claim 6, wherein: The calculating the second loss according to the original features of the second sample image, the encoded features of the second sample image, the reconstructed features of the second sample image, the first feature values ​​of the second sample image passing through each feature layer in the feature task network, and the second feature values ​​of the second sample image passing through each feature layer in the feature task network includes: Calculating a coding loss according to the original features of the second sample image and the encoded features of the second sample image; Calculating a feature reconstruction loss according to the original features of the second sample image and the reconstructed features of the second sample image; Calculating a first multi-level feature loss based on the first eigenvalue of each feature layer in the feature task network and the second eigenvalue of each feature layer in the feature task network after the second sample image passes through the feature task network; The second loss is calculated according to the encoding loss, the feature reconstruction loss and the first multi-level feature loss.

8. The method according to claim 4, wherein The step of training the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the first reconstructed image network includes: obtaining a third sample image; Inputting the third sample image into the feature extraction network, the encoder, the decoder, and the initial first reconstructed image network in sequence to obtain a reconstructed image corresponding to the third sample image for human vision; calculating a third loss according to the third sample image and a reconstructed image corresponding to the third sample image and adapted to human vision; The parameters of the initial first reconstructed image network are adjusted according to the third loss to obtain the first reconstructed image network.

9. The method according to claim 8, wherein The method further comprises: In the process of adjusting the parameters of the initial first image reconstruction network, the parameters of the decoder are adjusted according to the third loss to obtain a new decoder, so as to obtain the machine vision-oriented image processing system using the new decoder.

10. The method according to claim 4, wherein: The step of training the initial second reconstructed image network based on the feature extraction network, the encoder, and the task network to obtain the second reconstructed image network includes: obtaining a fourth sample image; Inputting the fourth sample image into the task network for processing, and obtaining first feature values ​​of the fourth sample image after passing through each feature layer in the task network; Inputting the fourth sample image into the feature extraction network, the encoder, the decoder, and the initial second reconstructed image network in sequence to obtain a machine vision-oriented reconstructed image corresponding to the fourth sample image; Inputting the machine vision-oriented reconstructed image corresponding to the fourth sample image into the task network for processing, and obtaining second eigenvalues ​​of the fourth sample image after passing through each feature layer in the task network; Calculating a fourth loss based on the four sample images, a machine vision-oriented reconstructed image corresponding to the fourth sample image, first eigenvalues ​​of the four sample images passing through each feature layer in the task network, and second eigenvalues ​​of the fourth sample image passing through each feature layer in the task network; The parameters of the initial second reconstructed image network are adjusted according to the fourth loss to obtain the second reconstructed image network.

11. The method according to claim 10, wherein: The method further comprises: In the process of adjusting the parameters of the initial second image reconstruction network, the parameters of the decoder are adjusted according to the fourth loss to obtain a new decoder, so as to obtain the machine vision-oriented image processing system using the new decoder.

12. The method according to claim 10, wherein: The calculating the fourth loss according to the four sample images, the machine vision-oriented reconstructed image corresponding to the fourth sample image, the first eigenvalues ​​of the four sample images passing through each feature layer in the task network, and the second eigenvalues ​​of the fourth sample image passing through each feature layer in the task network includes: calculating an image reconstruction loss according to the four sample images and a machine vision-oriented reconstructed image corresponding to the fourth sample image; Calculating a second multi-level feature loss according to first eigenvalues ​​of the four sample images passing through each feature layer in the task network and second eigenvalues ​​of the four sample images passing through each feature layer in the task network; The fourth loss is calculated according to the image reconstruction loss and the second multi-level feature loss.

13. A training device for an image processing system for machine vision, wherein: include: An initial network acquisition module is used to obtain an initial feature extraction network, an initial encoder, an initial decoder, an initial first reconstructed image network, an initial second reconstructed image network, an initial feature adaptation network, a task network, and a feature task network; A first training module is used to train the initial feature extraction network and the initial feature adaptation network based on the feature task network to obtain a preliminarily trained feature extraction network and a preliminarily trained feature adaptation network; A second training module is configured to train the preliminarily trained feature extraction network, the preliminarily trained feature adaptation network, the initial encoder, and the initial decoder based on the feature task network to obtain the feature extraction network, the feature adaptation network, the encoder, and the decoder; A third training module is configured to train the initial first reconstructed image network based on the feature extraction network and the encoder to obtain the first reconstructed image network; a fourth training module, configured to train the initial second reconstructed image network based on the feature extraction network, the encoder, and the task network to obtain the second reconstructed image network; A system generation module is used to obtain the machine vision-oriented image processing system based on the feature extraction network, the encoder, the decoder, the first reconstructed image network, the second reconstructed image network, the task network, the feature adaptation network and the feature task network.

14. An electronic device, wherein: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the training method for a machine vision-oriented image processing system according to any one of claims 4 to 12 by executing the executable instructions.

15. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the training method for a machine vision-oriented image processing system according to any one of claims 4 to 12 is implemented.

16. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the training method for a machine vision-oriented image processing system according to any one of claims 4 to 12.

Citation Information

Patent Citations

  • Extremely low bit rate man-machine cooperation image coding training method and coding and decoding method

    CN113949880A

  • Video coding and decoding method and device for machine vision task, equipment and medium

    CN116366852A

  • Video coding method, device, system and equipment oriented to man-machine mixing and medium

    CN116546214A

  • Machine vision-oriented image processing system and training method thereof

    CN118298275A

  • Method and apparatus for coding feature map based on deep learning in multitasking system for machine vision

    US20240054686A1