Image processing method and device based on task identifier, equipment and medium

CN120655525BActive Publication Date: 2026-09-11PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510722349.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2026-09-11
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

[0007]本发明的主要目的在于提供一种基于任务标识符的图像处理方法、装置、设备及存储介质,旨在解决现有技术中不同图像处理任务所依赖的模型结构和数据处理方式差异显著,缺乏统一、可配置的任务选择机制,难以实现多任务图像处理的高效集成与自动化执行的技术问题

Benefits of technology

[0022]有益效果:本发明涉及图像处理技术领域,可应用于金融科技、医疗健康等业务场景中,公开了一种基于任务标识符的图像处理方法、装置、设备及介质,包括:获取原始图像数据并转换为预设图像格式以生成标准化图像数据;将标准化图像数据与任务标识符绑定形成复合图像数据;依据任务标识符从任务库中选择感知模块,处理复合图像数据生成第一中间特征数据;将第一中间特征数据输入图像处理核心网络生成第二中间特征数据;根据任务标识符选择输出模块处理第二中间特征数据生成目标图像数据。本发明通过任务标识符驱动的模块选择机制,实现了在统一图像处理流程中自动适配多种图像处理任务,避免了现有技术中因模型结构差异导致的无法集成和复用问题,提升了任务配置灵活性和图像处理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655525B_ABST
    Figure CN120655525B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, can be applied to business scenes such as financial technology and medical health, and discloses an image processing method and device based on a task identifier, equipment and a medium, which comprises the following steps: acquiring original image data and converting the original image data into a preset image format to generate standardized image data; binding the standardized image data with a task identifier to form composite image data; selecting a perception module from a task library according to the task identifier, processing the composite image data to generate first intermediate feature data; inputting the first intermediate feature data into an image processing core network to generate second intermediate feature data; and selecting an output module according to the task identifier to process the second intermediate feature data to generate target image data. Through the module selection mechanism driven by the task identifier, the application realizes automatic adaptation of various image processing tasks in a unified image processing process, and improves the task configuration flexibility and image processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image processing method, apparatus, device, and storage medium based on task identifiers. Background Technology

[0002] In practical image processing applications, particularly in fintech, healthcare, and marketing material management, there is a common need for improved image quality and enhanced adaptability.

[0003] Taking the fintech business as an example, documents such as insurance claims materials, customer identity verification images, and remote credit review certificates often require enhancement processing of user-uploaded on-site photos or document images, such as sharpening, denoising, and structural restoration, to meet compliance requirements and ensure the accuracy of intelligent recognition systems. However, these image sources typically suffer from inconsistent resolution, poor lighting conditions, raindrop occlusion, or motion blur, severely impacting the efficiency of subsequent image recognition and model inference. In practical applications, while mature image super-resolution, rain removal, and denoising models exist to address these issues, the structure and processing methods of each model are highly customized, leading to complex system integration, fragmented processing flows, and requiring extensive repetitive architectural design and engineering modifications when expanding to new tasks.

[0004] In the healthcare field, the accuracy and interpretability of image processing are particularly critical. For example, in scenarios such as skin lesion detection, ultrasound image enhancement, CT reconstruction, or endoscopic image analysis, clinical decision support often relies on the accurate restoration and enhancement of image details. However, most existing systems adopt a distributed model stacking architecture, with different tasks calling different models. The lack of a unified data format and modular mechanism leads to high task migration costs and weak model scheduling capabilities, which not only increases the reliance on algorithm experts but also affects deployment efficiency and real-time performance.

[0005] Similar bottlenecks exist in other application scenarios related to image acquisition and material processing, such as the acquisition of outdoor shooting materials for marketing campaigns. Material images often suffer from uneven lighting, slight occlusion, or insufficient resolution due to significant variations in the on-site environment, making them unsuitable for direct use in promotional materials. Although numerous image enhancement models exist for image restoration or quality improvement, most are specifically designed for repairing a particular type of image defect. These models suffer from inconsistent parameter structures, non-standardized interfaces, and a lack of cross-model fusion and scheduling capabilities, hindering efficient reuse in real-world material production chains. Furthermore, because model training relies on data from specific task domains, they lack task generalization and rapid adaptation capabilities, failing to meet the flexibility and scalability requirements of multi-scenario image processing.

[0006] In summary, existing image processing technologies have significant shortcomings in multi-task integration, module standardization, and flexible scheduling. The lack of an image processing system with good model compatibility, task configurability, and process automation capabilities has become a key technological bottleneck limiting image processing efficiency and business scalability. Summary of the Invention

[0007] The main objective of this invention is to provide an image processing method, apparatus, device, and storage medium based on task identifiers, aiming to solve the technical problem in the prior art that different image processing tasks rely on significantly different model structures and data processing methods, lack a unified and configurable task selection mechanism, and are difficult to achieve efficient integration and automated execution of multi-task image processing.

[0008] To achieve the above objectives, the present invention provides an image processing method based on task identifiers, comprising:

[0009] Acquire the raw image data to be processed, and convert the raw image data into a preset image format to generate standardized image data;

[0010] The standardized image data is bound to a task identifier that specifies the target task type for image processing to generate composite image data;

[0011] The corresponding perception module is selected from the task library according to the task identifier, and the composite image data is input into the perception module. The perception module processes the composite image data to generate the first intermediate feature data.

[0012] The first intermediate feature data is input into the image processing core network, and the image processing core network processes the first intermediate feature data to generate the second intermediate feature data.

[0013] The corresponding output module is selected from the task library according to the task identifier, and the second intermediate feature data is input into the output module. The output module processes the second intermediate feature data to generate target image data.

[0014] Furthermore, to achieve the above objectives, the present invention provides an image processing apparatus based on task identifiers, comprising:

[0015] The image preprocessing module is used to acquire the raw image data to be processed and convert the raw image data into a preset image format to generate standardized image data;

[0016] The task binding module is used to bind the standardized image data with a task identifier that specifies the target task type for image processing, thereby generating composite image data.

[0017] The perception module is used to select the corresponding perception module from the task library according to the task identifier, input the composite image data into the perception module, and process the composite image data through the perception module to generate first intermediate feature data;

[0018] The core feature network module is used to input the first intermediate feature data into the image processing core network, and process the first intermediate feature data through the image processing core network to generate the second intermediate feature data.

[0019] The output module is used to select a corresponding output module from the task library according to the task identifier, input the second intermediate feature data into the output module, and process the second intermediate feature data through the output module to generate target image data.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a determining machine device, the determining machine device including a memory, a processor, and a task identifier-based image processing program stored in the memory and executable on the processor, wherein when the task identifier-based image processing program is executed by the processor, it implements the steps of the task identifier-based image processing method as described above.

[0021] Furthermore, to achieve the above objectives, the present invention also provides a deterministic machine-readable storage medium storing an image processing program based on a task identifier, wherein the image processing program based on the task identifier, when executed by a processor, implements the steps of the image processing method based on the task identifier as described above.

[0022] Beneficial Effects: This invention relates to the field of image processing technology and can be applied to business scenarios such as fintech and healthcare. It discloses an image processing method, apparatus, device, and medium based on task identifiers, comprising: acquiring raw image data and converting it into a preset image format to generate standardized image data; binding the standardized image data with task identifiers to form composite image data; selecting a perception module from a task library according to the task identifier to process the composite image data and generate first intermediate feature data; inputting the first intermediate feature data into an image processing core network to generate second intermediate feature data; and selecting an output module according to the task identifier to process the second intermediate feature data and generate target image data. This invention, through a task identifier-driven module selection mechanism, achieves automatic adaptation to multiple image processing tasks within a unified image processing flow, avoiding the integration and reuse problems caused by differences in model structures in existing technologies, and improving task configuration flexibility and image processing efficiency. Attached Figure Description

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0024] Figure 1 This is a schematic diagram of an application environment for an image processing method based on task identifiers according to an embodiment of the present invention;

[0025] Figure 2 This is a flowchart illustrating an embodiment of the image processing method based on task identifiers according to the present invention;

[0026] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image processing device based on task identifiers of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of the device in one embodiment of the present invention;

[0028] Figure 5 This is another structural schematic diagram of the device in one embodiment of the present invention. Detailed Implementation

[0029] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0030] The image processing method based on task identifiers provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain raw image data from the user terminal and convert it into a preset image format to generate standardized image data; bind the standardized image data with a task identifier to form composite image data; select a perception module from the task library based on the task identifier to process the composite image data and generate first intermediate feature data; input the first intermediate feature data into the image processing core network to generate second intermediate feature data; and select an output module based on the task identifier to process the second intermediate feature data to generate target image data. This invention, through a task identifier-driven module selection mechanism, achieves automatic adaptation to multiple image processing tasks within a unified image processing flow, avoiding the integration and reuse problems caused by differences in model structure in existing technologies, and improving task configuration flexibility and image processing efficiency. The user terminal can be, but is not limited to, various personal digital cameras, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0031] Please see Figure 2 , Figure 2This is a flowchart illustrating an embodiment of the image processing method based on task identifiers provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0032] like Figure 2 As shown, the image processing method based on task identifiers proposed in this invention includes the following steps:

[0033] S10: Obtain the raw image data to be processed, and convert the raw image data into a preset image format to generate standardized image data;

[0034] In this embodiment, to achieve multi-task processing of image data, it is first necessary to standardize the format and representation of the image data. Image data is typically acquired from external acquisition devices, such as cameras, image sensors, drones, or portable shooting terminals. Acquiring image data introduces various uncontrollable conditions, including inconsistent resolution, inconsistent color channel configurations, diverse image file formats, and different pixel arrangements. Therefore, it is necessary to unify the raw image data into a format that meets the requirements of downstream processing. During processing, the raw image data is first read. Raw image data refers to the image content without any format conversion or feature normalization processing. It may exist in formats such as JPEG, BMP, PNG, and RAW, or be in a temporary storage structure within the image acquisition buffer. The read operation can be completed through API interfaces, memory mapping, edge node upload interfaces, etc.

[0035] The primary format type of raw image data is a combination of image encoding format, resolution representation, and color model, such as PNG format under the RGB color model, or H.264 frame data under the YUV color model. Verifying its conformity to the preset image format involves identifying its file header, number of channels, bit depth, and color encoding standard. Image format requirements may include the channel order of the pixel matrix, resolution range, compression encoding specifications, etc. If a discrepancy with the preset format is found during verification, a format conversion operation is required. Format conversion can be performed using image codecs, color space conversion functions provided in the OpenCV library, or custom matrix mapping methods to convert the image from its original format to a standard format.

[0036] After converting the image data, it needs to be adjusted to a uniform resolution to adapt to a unified image perception and feature processing model. The resolution adjustment module can be implemented using methods such as bilinear interpolation, region averaging, and content-aware scaling. The input image is scaled according to the ratio between its original size and the target size, ensuring the output image has a consistent spatial dimension. After adjusting the image data, color space standardization is also required. Standardization not only unifies different color spaces (such as HSV and BGR) to the RGB space but also normalizes image channels, such as linearly mapping pixel values ​​from the [0,255] range to [0,1] or [-1,1] to meet the input constraints of the neural network model regarding numerical scale.

[0037] After format conversion, resolution adjustment, and color standardization, the image data is encapsulated into standardized image data with structured identifiers. This encapsulation process is achieved by adding format identifier fields, resolution identifier fields, and color identifier fields to the image data structure. These fields can be designed as JSON format, binary header information, or key-value pair structures. Specific field content may include image width and height, number of channels, color model (e.g., sRGB), and image dimension encoding order (e.g., CHW or HWC). The encapsulated image data can be used for model processing or stored as standard task input, facilitating data exchange and workflow adaptation between systems.

[0038] In some applications, image acquisition devices output compressed video frames, which can be unpacked into a processable pixel matrix using a frame decoder. For example, medical imaging terminal devices may output 16-bit grayscale DICOM format images, which need to be converted to RGB format and mapped to the visible area using a decoding library. In other implementations, the definition of a standardized image format not only includes uniformity in resolution and color channels, but also requires that the image data must have a specific structural order. For instance, when using a specific deep learning inference acceleration engine (such as TensorRT), the image data needs to be arranged in CHW format and aligned to 16-byte boundaries.

[0039] In different scenarios, the generation method of standardized image data can be parameterized. For example, in traffic monitoring applications, the uniform resolution of the input image can be set to 1920×1080, while in low-power processing scenarios on mobile devices, it can be configured to 640×480 to reduce computational burden. Color standardization processing in financial scenarios can adopt a normalization method that enhances contrast to improve the recognition accuracy of images such as receipts and ID cards, while in medical scenarios, contrast stretching and gamma correction can be selected to enhance tissue boundary features.

[0040] Example: In the healthcare field, images captured by hospital terminal devices often exhibit inconsistencies in color models and resolution. For instance, ultrasound images are represented using grayscale, while endoscopic images are high-resolution color data. Standardizing these images allows them to be uniformly input into a multi-task recognition system, enabling consistent annotation of body parts, detection of abnormal areas, and other subsequent operations.

[0041] In the fintech sector, marketing campaign materials may include mobile photos and scanned images of varying resolutions and compression ratios, some also containing redundant backgrounds, noisy watermarks, and other unwanted content. Standardizing format and resolution conversions, along with color standardization, can improve image clarity and contrast, facilitating subsequent image enhancement, content recognition, and text / image structure analysis. This process, within an automated material processing platform, can significantly improve the efficiency of image data flow and the processing accuracy of downstream recognition models.

[0042] This embodiment performs format conversion, resolution adjustment, and color space standardization on the original image data, so that image inputs from different sources and in different formats are uniformly converted into a standardized image data format that can be recognized and adapted by a unified processing system. This significantly improves the compatibility of image preprocessing and the efficiency of task access, and reduces the input adaptation cost of multi-model access.

[0043] S20, bind the standardized image data with a task identifier for specifying the target task type of image processing to generate composite image data;

[0044] In this embodiment, to achieve multi-task selection and adaptation in the image processing workflow, task selection information needs to be appended to the input data to ensure that the subsequent system can load the matching processing module based on this information. First, a pre-built task type library is used as the source of task identifiers. The task type library contains several task definitions corresponding to the image processing objectives, with each task type corresponding to a unique code. This task code is an abstract integer or string identifier and does not directly participate in model calculations; therefore, it needs to be converted into a vector expression that can be input into the model along with the image data. The conversion operation transforms the discrete task identifier into a binary encoded vector with specific dimensions and numerical structure. This encoded vector can be a fixed-length sparse vector, a one-hot vector, an embedded vector, etc., and its dimensional design should match the number of tasks and scalability of the system.

[0045] To embed the encoded vector into the input data, a metadata tag structure needs to be constructed. This structure can contain fields such as the task encoding itself, encoding dimension description, and encoding version number, which describe how the task information exists in the image data. The association between the metadata tag and the standardized image data can be achieved by tagging file header information, external index tables, or constructing structured data objects in memory, ensuring that the task information remains synchronized during data flow.

[0046] Converting task information from logical labels to a model-recognizable structure requires directly embedding the encoded vector into the image data. To achieve this, the encoded vector is added as an independent data channel and concatenated with the pixel channels of the normalized image. A one-dimensional structure is added to the image's channel dimension as the task channel. This channel maintains the same spatial dimension as the original image, and its values ​​are either copied encoded vectors or broadcast-enlarged two-dimensional matrices. After adding the task channel, the image data becomes a composite image structure containing both standard pixel channels and the task channel. Subsequent model inputs can then be based on this structure to implement task selection. This channel concatenation operation requires the image processing model to have the ability to input multi-channel tensors and to incorporate feature awareness of the task channel into the model's front-end structure.

[0047] In practical applications, task encoding can be designed as 16-bit or 32-bit binary encoding vectors, using embedding matrices to map and encode different tasks. Task types such as image deraining, image magnification, and image denoising can be mapped to different encoding bits, forming multi-task sparse encoding. During processing, this encoding vector is copied into a matrix with the same height and width as the image and concatenated to form a new channel for the image data. The concatenation operation can be implemented using tensor operation functions such as `concat` and `stack` in frameworks like NumPy, PyTorch, and TensorFlow. If the task library contains a large number of tasks, low-rank embedding can be used to compress the task encoding into a fixed-dimensional embedding vector, and a position broadcasting mechanism can be used before the image data input to expand the embedding vector to the same spatial size as the image. In system deployment, task identifiers can also be selected by the user or dynamically generated by an automatic decision-making module, and then the server completes the binding and concatenation of the image and the encoding.

[0048] Example Explanation: In healthcare scenarios, doctors can select different image enhancement tasks for images of lesion areas, such as edge sharpening and image deblurring. By adding task identifiers as channels to the image data, processing methods can be dynamically switched within the same model without needing to change the model architecture.

[0049] In fintech business scenarios, contract images uploaded by users may have issues such as watermarks or blurriness. By selecting an identifier and binding it to a task such as "remove watermark" or "enhance contrast," the backend model can automatically perform the target processing operation based on the identifier, thereby achieving automated and consistent image quality improvement.

[0050] This embodiment achieves structured embedding of task selection information by concatenating standardized image data with task identifiers in the channel dimension. This enables subsequent image processing models to achieve multi-task adaptation and dynamic branch loading while maintaining input consistency, thereby improving the uniformity and scalability of the overall model structure and avoiding the resource overhead of repeatedly training multiple models for different tasks.

[0051] S30, select the corresponding perception module from the task library according to the task identifier, input the composite image data into the perception module, and process the composite image data through the perception module to generate the first intermediate feature data;

[0052] In this embodiment, the key to applying task identifiers to the model processing flow lies in implementing a dynamic selection mechanism for the modular structure. First, task channels are extracted from the composite image data. Each channel contains a task encoding matrix of the same dimension as the image channels. This matrix replicates the vector representation of the task identifier through uniform padding or broadcasting mechanisms for subsequent decoding operations. The model system needs to be able to recognize this channel information and perform logical parsing; therefore, a task decoding module needs to be constructed. This module extracts task encoding vectors from individual channels and restores them to recognizable task identifier names or indices by querying a preset task type mapping table. This mapping table can be implemented based on hash mapping, encoding tables, or embedding retrieval mechanisms.

[0053] After obtaining the task identifier, the module selection process begins. The task library contains multiple pre-defined perception modules, each stored as an independent structure configuration file and corresponding weight parameter set. The structure configuration file records the network layer structure of the perception module, including the convolutional layer order, kernel size, activation function, and normalization layer settings. The system locates the target module in the module registry based on the task identifier and obtains the structure definition file path and parameter file path of the perception module through the index table.

[0054] Module loading consists of two phases: structure initialization and parameter injection. The structure initialization process parses the structure definition file into a network structure instance in memory, constructing the computational graph topology of the perceptron module. The parameter injection operation loads the weight tensors of the convolutional layers into the kernel weight storage area of ​​each layer, and loads the scaling and offset factors of the batch normalized layers into the corresponding parameter storage area, completing the module entity construction. The loading process must verify the consistency between the structure and parameters, such as matching the number of channels, consistent layer order, and dimension alignment.

[0055] Composite image data structurally comprises standard image channels and task channels. When providing input to the perception module, image channel separation is required. The channel separation operation extracts all channels except the task channel to form a new tensor as standardized image input, while retaining task information for subsequent modules. This input data stream is fed into the initialized perception module, where a sequence of convolutional layers is executed for feature extraction. The multi-level convolution process achieves progressive abstraction from low-level pixel-level features to high-level semantic features. After each convolutional layer output, batch normalization is performed, using loaded parameters to scale and shift the feature map, stabilizing distribution fluctuations during training and inference.

[0056] After all convolutional and normalization layers, the final output feature tensor is the normalized feature data. Since the output of the perceptual module may have high dimensionality, channel compression is required to adapt to the input specifications of the subsequent core network. Channel compression uses 1×1 convolutional layers, channel average pooling, or projection mapping to compress the number of channels to the target dimension while preserving the expressive power of the main features. The low-dimensional output features are the first intermediate feature data, which serves as the input to the subsequent image processing core network.

[0057] In implementation, the task recognition module can be implemented using a custom layer based on PyTorch. It extracts the task encoding vector from the last channel of the image tensor and maps it to a task label using a lookup table. The configuration file for the perception module records the hierarchical structure in JSON format and constructs a subclass of `nn.Module` through a loading function. Parameter loading uses a weight mapping dictionary to map weight files (e.g., .pt or .h5 format) in the storage path to specified layers in the model structure. After the perception module is built, a normalized image tensor (with task channels removed) can be used as input to perform the model's forward propagation operation. Non-linear activation functions such as ReLU or GELU are used between convolutional layers, and batch normalization parameters (γ and β) are loaded into the BN layer instance according to each layer dimension. The output feature tensor undergoes channel compression using a 1×1 convolution kernel and is returned as input to downstream modules.

[0058] Example Explanation: In a healthcare scenario, for two tasks—diabetic retinopathy detection and skin spot classification—although both inputs are structured image data, their feature extraction requirements differ significantly. Diabetic retinopathy detection relies more on the ability to identify microbleeds under low contrast, while skin spot classification requires enhanced local perception of texture and edge contours. After embedding task identifiers into the uniformly formatted image data, the system automatically selects the corresponding multi-scale low-pass sensing module from the task library for the former, strengthening the hierarchical extraction of lesion micro-features; while for the latter, it triggers a texture sensing module based on a combination of high-pass filters, prioritizing the capture of edge texture structures in the skin. Different sensing modules complete parameter loading through structure registration and independently process the composite image data, outputting first intermediate feature data that meets the accuracy requirements of subsequent diagnosis, without requiring manual intervention in the model switching process.

[0059] In fintech scenarios, user-uploaded image data may fall into different task types, such as bank card photography, document photography, or contract scanning. Although the format may be the same, the focus on detail perception differs significantly. The system automatically matches the corresponding perception module in the task library using a bound task identifier. For example, the bank card task triggers a perception module with local character enhancement paths, prioritizing the extraction of boundary information in the card number area; the document photography task activates modules with high-weight perception paths for background removal and stamp areas, ensuring that blurred watermarks and paper creases do not interfere with structure extraction. The selection and structure loading of perception modules for each task type are accomplished through an identifier-driven mechanism. Composite image data, as a unified input, requires no format conversion, and the first intermediate feature data output possesses clear structural adaptability, ensuring the scalability and accuracy of the financial document processing process.

[0060] This embodiment constructs a task identifier-driven perception module selection mechanism, enabling dynamic matching of different perception front-end structures to image data according to task requirements before it enters the processing flow, thereby improving the system's adaptability to multi-task image processing. This design supports independent deployment and updates of modules, reduces model coupling, and improves training efficiency and inference flexibility. Simultaneously, through channel separation and compression operations, it ensures the feasibility and efficiency of the processing flow in resource-constrained environments.

[0061] S40, the first intermediate feature data is input into the image processing core network, and the first intermediate feature data is processed by the image processing core network to generate the second intermediate feature data;

[0062] In this embodiment, the first intermediate feature data is the feature representation result obtained after processing composite image data in the preceding perception module. It has structural features specific to different image tasks, typically retaining the spatial arrangement and task-sensitive dimension information in the image, but it does not yet possess the deep abstraction capability of a unified representation structure. The image processing core network receives the first intermediate feature data as input and undertakes the function of performing high-order processing on cross-task shared features. When constructing the image processing core network, it includes at least two basic processing units: a feature transformation unit and a semantic mapping unit. The former performs cross-channel fusion on the first intermediate feature data based on a multi-layer convolutional structure or a graph convolutional structure, gradually reducing the feature map size by combining a spatial compression mechanism to improve computational efficiency and aggregate the global context; the latter performs dimensional transformation on the compressed feature vector through a fully connected mapping mechanism, so that its output meets the semantic consistency requirements of subsequent modules.

[0063] The processing first requires initializing the parameters within the core image processing network, including the weights of all convolutional and fully connected layers, and performing a parameter freezing operation to prevent interference with the network's fundamental structure during subsequent multi-task expansion. The freezing operation is achieved by disabling gradient backpropagation for all learnable parameters within the network. Its purpose is to treat the network as a general abstract encoder for feature transformation, unaffected by the specific requirements of downstream tasks. The first intermediate feature data is typically preserved as a two-dimensional or four-dimensional tensor before input. The processing flow converts this into fused features and then performs multi-scale dimensionality reduction step by step. After feature compression, it enters a normalization process to further unify the data distribution characteristics. Finally, a fully connected structure completes the dimensionality mapping output, forming the second intermediate feature data. Compared to the first intermediate feature data, the second intermediate feature data has stronger expressive power in terms of structural compactness, semantic abstraction level, and task adaptability.

[0064] In practical deployments, the core image processing network can be constructed as a frozen ResNet backbone network, with the first few layers performing channel fusion and spatial compression operations, and a lightweight fully connected structure added at the end to complete the output encoding. Alternatively, a Transformer structure can be used as the core network's processing architecture, encoding the first intermediate feature data into a token sequence before input, and using a self-attention mechanism to complete cross-location dependency modeling. In the feature fusion operation, 1×1 convolutions can be used to enhance inter-channel interactions, or attention mechanisms can be used to generate channel weights, enhancing the expressive power of key regions while maintaining feature integrity. Dimensionality reduction can be achieved through stride convolutions or pooling operations; multi-level dimensionality reduction structures can be designed using inverted pyramid feature paths. In the standardization stage, batch normalization or layer normalization techniques can be used, with the specific choice determined based on the sensitivity requirements of the downstream task. In the dimension mapping stage, single-layer linear mapping can be used, or non-linear transformations can be achieved by stacking activation functions.

[0065] Example Explanation: In a healthcare scenario, when processing electrocardiogram (ECG) images and chest X-ray images, the first intermediate feature data output after feature extraction using different perception modules at the input stage exhibits different spatial distribution characteristics and texture representations. In this stage, the system uniformly inputs the first intermediate feature data corresponding to both images into the image processing core network for standardization. Its spatial scale is compressed using a frozen convolutional structure, and a fully connected structure is used to output second intermediate feature data of the same dimension. This satisfies the requirements of the subsequent unified anomaly detection module regarding input dimension and data distribution, improving the model's consistency and performance stability in multi-disease image processing.

[0066] In fintech scenarios, bank card images and handwritten signature images exhibit significant differences in resolution and structural noise due to variations in acquisition methods, resulting in inconsistent first intermediate feature data structures. By inputting these features into the core image processing network, the system automatically performs dimensionality reduction and fusion of features from unstructured interference regions, outputting second intermediate feature data with a unified structure. This allows subsequent image risk scoring or image compliance recognition modules to directly interface with the system, avoiding data preprocessing redundancy caused by model switching between different tasks and effectively improving the task scheduling efficiency and scalability of the image recognition system.

[0067] This embodiment effectively isolates structural conflicts between the perception module and the output module caused by task differences by inputting the first intermediate feature data into the image processing core network. This network acts as a frozen structure to maintain stability, preventing perception module-specific features from interfering with other task paths during the output stage. The first intermediate feature data achieves a unified distribution through channel fusion and dimensionality reduction standardization, and after dimensionality reorganization in a fully connected structure, it outputs the second intermediate feature data. This process ensures the uniformity of image features in both spatial and semantic information dimensions, helping to improve processing consistency and system scalability across different tasks, thereby enhancing the reusability and functional integration capabilities of the image processing system.

[0068] S50, select the corresponding output module from the task library according to the task identifier, input the second intermediate feature data into the output module, and process the second intermediate feature data through the output module to generate target image data.

[0069] In this embodiment, the role of the task identifier is not limited to the logical entry point for model selection; it also needs to achieve precise mapping with different output modules in the task library. The task identifier consists of a unique code that can be parsed into an index value for a specific module. Internally, it serves as a parameter retrieval pointer to locate the metadata structure of the output module configuration. The module configuration index table maintains the mapping relationship in key-value pairs, where the key is the index code corresponding to the task identifier, and the value is the storage path of the module file and the parameter reference tag. The module loading action requires reading the structure definition file and parameter set from the storage system after index parsing is complete. This file can be represented using a serialized structure, such as Protocol Buffers or ONNX format.

[0070] After the module is loaded, the second intermediate feature data needs to undergo structural reconstruction to adapt to the input dimension requirements of the target output module. This structural reconstruction manifests as channel dimension expansion, which transforms the lower-dimensional tensor into a higher-channel representation to match the input feature size of the deconvolution path within the module. The expansion operation can be implemented based on 1×1 convolution or tensor copying techniques, combined with an attention factor to redistribute channel activation values, thereby improving the effective representation ratio between channels.

[0071] The output module's deconvolution process employs a weight-parameter-driven inverse convolution operation to spatially expand the second intermediate feature data. This process is typically accomplished through a series of sequentially connected deconvolution layers. The weights for each layer are derived from a preloaded parameter set, and the kernel size and stride are set by the module structure definition file. After each layer's output, activation functions or normalization operations can be inserted to adjust the feature distribution. After the deconvolution operation, the upsampling ratio parameter is used as a reference coefficient for spatial interpolation. Interpolation algorithms can include linear interpolation, bicubic interpolation, or subpixel shifting. The interpolation operation further enlarges the image tensor size in the spatial dimension while preserving structural continuity.

[0072] After interpolation, the image tensor enters the numerical compression stage, where pixel values ​​are proportionally adjusted through normalization. This operation targets a specific format specification (such as 32-bit floating-point or 8-bit integer) and remaps the data according to a preset standard range. A truncation mechanism can be introduced during this process to prevent calculation errors from causing out-of-bounds representations, while ensuring that the image output conforms to the data interface standards required by downstream tasks (such as visual display, image file generation, and subsequent detection processing). The final generated target image data is the tensor result after structural decoding and format processing by the output module. The entire processing chain ensures continuity and controllability from feature representation to image reconstruction, and task-driven module calls ensure the specificity of the processing path in terms of structure and parameters.

[0073] In practical implementation, the output module invocation process can be achieved through a query-based indexing mechanism. The system maps task identifiers to key-value pairs and retrieves model configuration files containing deconvolution kernel groups and upsampling parameters from the database. The deconvolution kernel parameter groups can adopt multi-scale structures, such as parallel paths of 3×3, 5×5, and 7×7, to adapt to the image edge detail reconstruction needs of varying complexity. Channel expansion can be implemented using 1×1 convolution, while adding a channel attention structure to increase the effective channel weight ratio. Upsampling operations can be performed using bilinear interpolation or subpixel rearrangement modules, with the specific selection dynamically adjusted based on image restoration speed and quality requirements. Regarding image format normalization, the system supports automatic identification of the output format type. If the target image is to be used in a high-precision post-processing module, it is normalized to float32 format; if it is used for visualization or exported as an image file, it is encapsulated in uint8 format. A caching mechanism is supported after module loading to avoid performance impact from multiple loadings. The interpolation process can also dynamically adjust the sampling interval and interpolation function according to the task accuracy level, adapting to various output tasks such as super-resolution, deblurring, and image completion.

[0074] Example Description: In healthcare scenarios, lung CT images, after processing, need to output structural reconstruction images at different resolutions depending on the requirements of different institutions. The system identifies the task identifier and loads the corresponding output module. The deconvolution structure is designed to enhance features at the lung edges, and the upsampling ratio is set to 2.0 to match the input requirements of medical image interpretation systems. The final generated image meets the diagnostic accuracy requirements of doctors in terms of structural detail restoration.

[0075] In fintech scenarios, customer ID card images are processed for two different output tasks: face blurring and anti-fraud watermark synthesis. The system loads two output modules using two different task identifiers. The former employs standard resolution reconstruction and face region separation as the output structure, while the latter uses interpolation and magnification to pad the image edges to fit the watermark area. The final output structure fully meets the dual requirements of identity information security and compliance auditing. This structure allows the image system to automatically adapt to the image processing path of different output tasks without manual intervention.

[0076] This embodiment uses task identifiers to drive output module selection and a parameterized structure to complete image decoding and spatial reconstruction, enabling rapid switching and automatic processing of multiple task output structures. Channel dimension expansion enhances the versatility of detail reconstruction, while deconvolution kernels and upsampling operations ensure accurate spatial information restoration, and normalization guarantees image format consistency. This processing decouples task configuration from the output flow, giving the system modularity, pluggability, and multi-task compatibility, supporting diverse image generation needs while maintaining processing consistency. Ultimately, this achieves a balance between efficiency and accuracy in the entire image processing workflow, improving the scalability and configuration flexibility of the image production system.

[0077] This invention relates to the field of image processing technology and can be applied to business scenarios such as fintech and healthcare. It discloses an image processing method, apparatus, device, and medium based on task identifiers, comprising: acquiring raw image data and converting it into a preset image format to generate standardized image data; binding the standardized image data with task identifiers to form composite image data; selecting a perception module from a task library according to the task identifier to process the composite image data and generate first intermediate feature data; inputting the first intermediate feature data into an image processing core network to generate second intermediate feature data; and selecting an output module according to the task identifier to process the second intermediate feature data and generate target image data. This invention, through a task identifier-driven module selection mechanism, achieves automatic adaptation to multiple image processing tasks within a unified image processing workflow, avoiding the integration and reuse problems caused by differences in model structures in existing technologies, and improving task configuration flexibility and image processing efficiency.

[0078] In one embodiment, step S10 includes:

[0079] S101, Read raw image data from the image acquisition device, the raw image data containing a first format type;

[0080] S102, verify whether the first format type of the original image data meets the preset image format requirements;

[0081] S103, when the first format type does not meet the preset image format requirements, perform an image format conversion operation to generate converted image data of the second format type;

[0082] S104, input the converted image data or the original image data conforming to the preset image format into the resolution adjustment module to generate adjusted image data with the preset resolution;

[0083] S105, Perform color space normalization processing on the adjusted image data to generate standardized color image data;

[0084] S106, the standardized color image data is encapsulated into standardized image data containing format identifier, resolution identifier and color identifier.

[0085] In this embodiment, the image data acquisition stage is applicable to various image acquisition devices, whose output may be in different formats such as JPEG, PNG, TIFF, and BMP, with different compression strategies, channel order, bit depth, and metadata structures. The raw image data reading operation includes image decoding, memory mapping, channel order unification, and bit depth conversion to ensure that the subsequent verification module can completely parse its structural information. Image format type verification is based on file header tags, image metadata parsers, or media description fields in the acquisition protocol. This verification process must match the preset image format requirements. Preset format requirements typically include indicators such as the number of color channels (e.g., RGB, YUV), encoding structure (compressed or uncompressed), and pixel precision (e.g., 8-bit, 10-bit).

[0086] When the original format differs from the target format, an image format conversion operation is triggered. This operation calls an image codec library or image processing framework to perform pixel format transformation, color channel rearrangement, or compression structure unpacking. For example, a PNG image will be decoded into a three-channel RGB image, and then further converted into a unified float32 array format. The image data generated by this conversion becomes the input basis for subsequent image reconstruction and structural processing.

[0087] After image format conversion, its spatial structure needs to be standardized. The resolution adjustment module is used to generate image sizes that match the system's processing standards. Its core is to perform a scale mapping operation on the image's width and height, using algorithms such as bilinear interpolation, Lanczos resampling, or nearest neighbor interpolation. The adjusted image data maintains spatial dimensions consistent with the input interface in the model structure, ensuring a reasonable distribution of the model's receptive field for the input data.

[0088] After resolution alignment, the color space normalization process begins. This process unifies the color representation of images from different sources, such as converting the BGR channel order to RGB, or converting the YCbCr space to standard linear RGB. Normalization is accompanied by pixel normalization, mapping pixel values ​​from 0-255 to the [0,1] or [-1,1] interval. This step directly impacts the stability of the input layer of the convolutional neural network. Normalization parameters can be static (such as the ImageNet preset mean and variance) or dynamically set based on the characteristics of the acquisition device.

[0089] Finally, the processed image data is encapsulated into standardized image data with structured identifiers. This encapsulation structure embeds fields such as format identifiers (e.g., uniformly labeled RGB-float32), resolution identifiers (e.g., 1024×768), and color identifiers (e.g., linear RGB), enabling subsequent modules to quickly identify the image's feature attributes during decoding or parallel processing, thus achieving a self-describing structure of the image processing workflow.

[0090] This embodiment achieves format compatibility and unified expression of image data from different sources through structured processing paths such as image format verification and conversion, resolution alignment, and color standardization. This ensures that the image data meets the structural input requirements of subsequent processing modules of the system, avoids processing failures or model errors caused by heterogeneous input, and improves the robustness and accuracy of the image processing workflow.

[0091] In one embodiment, step S20 above includes:

[0092] S201, Generate a unique task code corresponding to the target task type based on a predefined task type library;

[0093] S202, convert the unique task code into a binary code vector of a preset dimension;

[0094] S203, Create a metadata tag containing the binary encoded vector;

[0095] S204, the metadata tags are associated with and stored with the standardized image data;

[0096] S205, the binary encoded vector is superimposed as an independent data channel onto the channel dimension of the standardized image data to generate the bound composite image data.

[0097] In this embodiment, the predefined task type library is a task enumeration mechanism used to establish a mapping relationship between different image processing tasks and their identifiers. Task types include operations such as image denoising, super-resolution reconstruction, image enhancement, style transfer, and object inpainting. Each task type has a unique index location and descriptive information in the library. When selecting a task, it is necessary to retrieve an entry matching the current target task from the task type library and extract its index or logical code in the library as the basic task code. This task code exists in integer or character form and is not directly used for image channel operations.

[0098] To embed this task encoding into the image data for model computation, the encoding result needs to be converted into a binary encoded vector. The dimension of this encoded vector is set according to the system's general design requirements, such as 8-dimensional, 16-dimensional, 32-dimensional, or higher. Dimensional conversion is achieved through one-hot encoding, fixed-length binary representation, Gray encoding, or hash function mapping. This binary encoded vector is then used as labeling information in subsequent image feature representation.

[0099] To enhance the readability and structured management of task identification information, the generated encoded vector is packaged into a metadata tag. The metadata tag is a structured object, typically containing task encoding fields, dimension description fields, and task tag text fields, used to annotate the semantics and structure of the binary vector. The structure of the metadata tag can be based on JSON, XML, or a custom protocol format, facilitating subsequent task library matching and model inference logic interpretation.

[0100] The association and storage of metadata tags and standardized image data can be achieved through data coupling or separation. One approach is to encapsulate the metadata and image data into a composite data structure, appending a labeling information area to the header of the image data. Another approach is to maintain the reference relationship between the two in the system through mapping indexes or key-value mechanisms, such as indexing image paths and label paths using task IDs. This approach provides flexible data structure support for data management and batch processing.

[0101] After the structural hierarchy binding is completed, the binary encoded vector is embedded into the actual data dimensions of the standardized image data, specifically by adding one or more channel dimensions. Taking a three-channel RGB image as an example, if the task encoding vector is 8-dimensional, the image tensor will expand from the original shape [H,W,3] to [H,W,3+8]. The numerical padding rules for the newly added channels strictly correspond to the binary encoded values ​​(0 or 1), and the vector is copied at all spatial locations in the image, forming a tensor with constant spatial dimensions but expanded channel dimensions, ultimately forming the composite image data with completed task binding.

[0102] This embodiment achieves a structural fusion of image data and task semantics by binding task encoding with image data. This allows the same model framework to select the appropriate processing path or activate different model branches based on the input encoding channels when processing different tasks, thus improving the system's generalization ability. This approach not only simplifies the system overhead of switching between multiple models but also enhances the flexibility of task scheduling and the consistency of data processing flows.

[0103] In one embodiment, step S30 above includes:

[0104] S301, extract the binary encoding vector from the independent data channels of the composite image data;

[0105] S302, Perform a decoding operation on the binary encoded vector to generate a target task type identifier;

[0106] S303, query the module registry in the task library according to the target task type identifier, and match the corresponding perception module configuration parameters. The perception module configuration parameters include convolution kernel weight recombination and batch normalization layer parameters.

[0107] S304, parse the network structure definition file in the configuration parameters of the perception module, and generate the initialization structure of the perception module;

[0108] S305, the convolution kernel weight reorganization is injected into the kernel weight storage area of ​​each convolutional layer of the initialization perception module structure, and the batch normalization layer parameters are written into the batch normalization layer parameter storage area of ​​the initialization perception module structure to generate a perception module with loaded parameters;

[0109] S306, Separate the standardized image data and binary encoded vector from the composite image data to generate the separated standardized image data;

[0110] S307, the separated standardized image data is input into the perception module, multi-level convolution operation is performed through the convolution kernel weight reorganization, and the output features are batch standardized through the batch standardization layer parameters to generate standardized feature data;

[0111] S308, Perform channel dimension compression on the standardized feature data to generate first intermediate feature data.

[0112] In this embodiment, the independent data channels embedded in the composite image data are used to carry task identification information. These channels exist as additional dimensions in the image tensor structure, for example, adding several channels to an RGB image to store binary encoded vectors. The process of extracting the encoded vector from these channels is based on tensor dimension indexing operations. It is necessary to determine the channel index range where the task encoding is located and read the vector value at all spatial locations. Considering that the task identification is globally consistent across the entire image, dimensionality reduction or averaging operations can be performed on the spatial dimensions to extract a unique vector representation.

[0113] The encoded vector is a binary vector of a preset dimension, and its decoding operation is performed according to a predefined task type mapping table in the task library. The mapping method can be static dictionary lookup, hash mapping, or optimal matching based on vector similarity. The decoding result is a target task type identifier with semantic description, which serves as a unique identifier for subsequent module selection and drives the perception module loading process.

[0114] The module registry in the task repository is a structured index collection, where each record binds a task type identifier to the configuration parameters of the perception module. The matching operation retrieves the corresponding perception module configuration parameters by searching the task type identifier field, including convolutional kernel weight reassembly and batch normalized layer parameters. Convolutional kernel weight reassembly is the set of kernel parameter tensors for each layer in a multi-layer convolutional structure, and batch normalized layer parameters include scaling factors and offset coefficients, all recorded in a structured manner.

[0115] A network structure definition file is a data file representing the network topology of a sensing module. It typically uses a structured configuration language such as JSON or YAML to describe information such as the network type of each layer, inter-layer connections, tensor dimensions, and activation functions. Parsing this definition file allows the construction of an empty module skeleton structure—a computation graph structure where each layer exists but its parameters are not yet assigned—for parameter injection.

[0116] In this initialization structure, convolutional kernel weights are injected sequentially into the kernel weight storage area of ​​the corresponding convolutional layer according to the layer number or name. Batch normalized layer parameters are injected into the γ and β parameter sites of the normalized layer. During parameter injection, dimensionality consistency and data format matching should be ensured; dimensionality transformation or data transposition should be performed if necessary. After the above injection is completed, a usable perception module is formed, and its structure and parameter loading state are initialized.

[0117] The composite image data consists of two parts: normalized image data and task identifier encoding vector. The data separation process splits the tensor according to the channel dimension, retains the image channels and removes the task identifier channels, forming normalized image data containing only pixel information, which is used as input to the perception module.

[0118] After the standardized image data is input into the perception module, multiple convolutional operations are performed sequentially according to the loaded convolutional kernel weights. Each convolutional layer is followed by a batch normalization operation. The normalization process performs feature value normalization and rescaling based on the loaded γ and β parameters. The convolutional operations can use standard convolution, depthwise separable convolution, or adjustable kernel convolution. The batch normalization operation can be configured with training or inference modes to control the parameter usage logic.

[0119] The high-dimensional feature data output by convolution and normalization operations contains redundant dimensions and unnecessary channel information. Channel dimension compression operations reduce the feature dimension based on the target channel setting, using 1x1 convolution, fully connected mapping, or channel selection strategies to generate first intermediate feature data with the dimensions required for downstream processing.

[0120] This embodiment achieves flexible selection and loading of perception modules by dynamically associating task identifiers with module parameters, enabling adaptive feature extraction without affecting the backbone network. Combining structural definition and weight configuration standardizes the perception module loading process and makes the structure controllable, improving the system's versatility and module reuse efficiency. The channel compression mechanism ensures uniform feature dimensions, facilitating structural compatibility and performance optimization in subsequent processing.

[0121] In one embodiment, step S40 above includes:

[0122] S401, initialize the weight parameters of the convolutional and fully connected layers in the image processing core network, and freeze the weight update function of the convolutional and fully connected layers in the image processing core network;

[0123] S402, Perform a cross-channel feature fusion operation on the first intermediate feature data to generate fused feature data;

[0124] S403, Perform multi-level feature map dimensionality reduction processing on the fused feature data to generate dimensionality-reduced feature data;

[0125] S404, Perform standardization processing on the dimensionality-reduced feature data to generate standardized feature data;

[0126] S405, the standardized feature data is input into the fully connected layer of the image processing core network, and a dimension mapping operation is performed through a linear transformation matrix to generate second intermediate feature data.

[0127] In this embodiment, the initialization process of the image processing core network includes setting the loading and update strategies for weight parameters. Convolutional layers and fully connected layers are common components in deep neural networks. Convolutional layers mainly handle the spatial extraction of local features, while fully connected layers are responsible for converting intermediate features into higher semantic representation vectors. The initialized weight parameters can be derived from a pre-trained model or randomly generated according to specific rules (such as Xavier or He initialization). The purpose of the freeze update function is to maintain the stability of the basic feature extraction capability and avoid parameter drift caused by gradient updates during training for specific tasks. This process can be achieved by setting an automatic differentiation flag or masking the parameter update path in the training graph.

[0128] Cross-channel feature fusion operations are used to integrate the semantics expressed by the first intermediate feature data across different channel dimensions, improving the global consistency and contextual relevance of features. Implementation methods include channel attention mechanisms (such as SE modules and ECA modules), multi-scale fusion modules (such as lateral connections in FPN structures), or feature channel cross-operations (such as convolution stacking, concatenation and remapping). This operation reconstructs the response intensity of each channel through weighted or structural combinations, enabling key channels to acquire more expressive power.

[0129] Even after fusion, features may still exhibit dimensionality redundancy. To reduce computational burden and preserve semantic hierarchy, multi-level feature map dimensionality reduction processing is required. This processing can be implemented using pyramid pooling, strided convolution, or sequential downsampling convolution to compress feature map sizes at different spatial scales, thereby preserving the hierarchical correspondence between local and global structures. The dimensionality-reduced feature maps should possess good scale adaptability and abstract representation capabilities, facilitating subsequent standardization and mapping operations.

[0130] Standardizing features after dimensionality reduction can eliminate inconsistencies in feature value distribution, enhancing the convergence stability and generalization ability of the network. Standardization operations include batch standardization, layer standardization, or image normalization (such as Z-score). This process normalizes the distribution of each feature channel to a uniform mean and variance, preventing certain channels from dominating subsequent mapping processes.

[0131] Standardized feature data is input into the fully connected layer of the core network. The fully connected layer uses a linear transformation matrix to map the input high-dimensional tensor to the intermediate semantic space required for the target task. The linear transformation is achieved through matrix multiplication, and the dimension of the transformation matrix is ​​determined by both the input feature length and the target dimension. This operation can be nested at multiple levels to progressively refine the feature representation. The resulting second intermediate feature data has a standardized structure and uniform dimensions, possessing universality for transmission to the output module.

[0132] This embodiment avoids gradient interference and redundant retraining during multi-task processing by freezing the basic feature extraction parameters, achieving modularity and high reusability. Feature fusion and dimensionality reduction operations enhance the expression compression capability while effectively preserving key semantics. Standardization and dimensionality mapping processes unify the format and numerical distribution of output features, improving the efficiency and accuracy of subsequent task adaptation and ensuring seamless collaboration between output modules.

[0133] In one embodiment, step S50 above includes:

[0134] S501, Based on the task identifier, generate an output module selection instruction;

[0135] S502, according to the output module selection instruction, query the module configuration index table in the task library to obtain the output module file path;

[0136] S503, Load the output module containing the deconvolution kernel parameter set and the upsampling ratio parameter according to the output module file path;

[0137] S504, Perform a channel dimension expansion operation on the second intermediate feature data to generate expanded feature data;

[0138] S505, input the extended feature data into the output module, execute the deconvolution operation corresponding to the deconvolution kernel parameter group, and generate deconvolution output data;

[0139] S506, Perform interpolation amplification processing on the deconvolution output data corresponding to the upsampling ratio parameter to generate amplified image data;

[0140] S507, the pixel values ​​of the magnified image data are normalized to generate target image data.

[0141] In this embodiment, the selection of the output module does not employ a static hard-coding method. Instead, an output module selection instruction is generated using the task identifier. This instruction can be generated by hash mapping, index query rules, or a task type classification model, ensuring consistency and scalability in calling different modules within a complex task system. The module configuration index table is a structured task-module mapping database containing the mapping relationship between task identifiers and corresponding module file paths. This structure is often implemented in the form of key-value pairs or relational indexes for efficient retrieval of target modules.

[0142] When loading the output module, the model structure containing the deconvolution kernel parameter set and upsampling ratio parameter must be obtained from the specified path. The deconvolution kernel parameter set includes multiple learnable kernel weight tensors, typically the inverse operation of a two-dimensional convolution kernel, used to achieve spatial reconstruction of the feature map. The upsampling ratio parameter defines the factor by which the output image size is expanded, often used to achieve low-resolution to high-resolution mapping. The loading process usually includes model structure initialization and parameter deserialization, with structure reconstruction and weight assignment completed through a tensor stream parser or inference engine.

[0143] The second intermediate feature data is the compressed deep representation, which often does not match the output module's requirements in terms of channel count or resolution. To meet the input conditions for deconvolution operations, channel dimension expansion is required. Expansion methods include channel duplication, zero padding, and 1×1 convolution mapping. Among these, 1×1 convolution can introduce lightweight trainable parameters, thereby improving accuracy.

[0144] After the extended feature data enters the output module, deconvolution is performed according to the deconvolution kernel parameter set. Deconvolution, also known as transposed convolution, performs interpolated parameter-weighted reconstruction on the input tensor through a reverse sliding window to restore the feature map in terms of spatial resolution. This operation not only reconstructs the macroscopic structure of the image but also combines semantic information from the extended channels to restore details such as edges and textures.

[0145] To improve the final image resolution and perceptual quality, the deconvolution output data needs to undergo interpolation and upsampling operations in conjunction with an upsampling ratio parameter. Interpolation methods can include bilinear interpolation, cubic interpolation, or deep learning-driven interpolation modules (such as PixelShuffle). The interpolation process adjusts the image size according to a preset ratio, fills in missing pixel values, and maintains continuity and smoothness.

[0146] The magnified image data does not yet have a pixel range that meets the output requirements, therefore pixel value normalization is necessary. Normalization can be achieved by using a linear transformation to map the values ​​to a preset range, such as [0, 255] or [0, 1], ensuring that the output image meets the standards for subsequent visualization or task deployment. In floating-point models, this step can also unify the pixel distribution to meet quantization requirements.

[0147] This embodiment implements task-driven output module selection and loading, avoiding the construction of independent network structures for each output effect and significantly reducing system redundancy. Channel expansion and deconvolution operations restore spatial information in compressed features, interpolation enhances image clarity and perceptible details, and pixel normalization ensures output uniformity and deployability. The overall workflow is modular, scalable, and adaptable to output formats, improving task processing flexibility and image generation quality.

[0148] In one embodiment, after step S50 above, the method further includes:

[0149] S601, extract the convolutional kernel weight reassembly and batch normalized layer parameters from the convolutional kernel weight storage area and the batch normalized layer parameter storage area of ​​the perception module, respectively.

[0150] S602, extract the deconvolution kernel weight reassembly from the deconvolution kernel parameter storage area of ​​the output module, and read the upsampling ratio parameter stored in the output module;

[0151] S603, Generate a module parameter set including the convolution kernel weight reorganization, batch normalization layer parameters, deconvolution kernel weight reorganization and upsampling ratio parameters;

[0152] S604, bind the task identifier to the module parameter set to generate module registration metadata;

[0153] S605, convert the module registration metadata into a preset binary serialization format to generate a standardized module configuration file;

[0154] S606, Update the module configuration index table of the task library according to the task identifier, and establish a mapping relationship between the standardized module configuration file and the task identifier;

[0155] S607, verify the integrity and loadability of the standardized module configuration file, and generate a module storage completion status signal.

[0156] In this embodiment, after the task is completed, to achieve long-term preservation and reuse of the model structure and task configuration, it is necessary to extract key parameters from the executed perception module and output module and manage their storage. The perception module, as the main body of feature extraction, has a convolutional layer kernel weight storage area containing the weight tensors of each convolutional layer. These weights are formed after training and are closely related to the feature structure of the input image. The batch normalization layer parameters include the scaling factor and offset factor in each normalization layer, which are used to control the mean and variance of the feature distribution under different tasks.

[0157] The output module, as the main body of image reconstruction, stores multiple weighted sets of deconvolution kernels in its deconvolution kernel parameter storage area. These parameters are trained to learn the spatial structure rules of image reconstruction. The upsampling ratio parameter is used to define the spatial expansion ratio of the feature map, usually expressed as an integer ratio (such as 2, 4, 8). After extracting the above parameters, they are packaged into a unified dataset called the module parameter set. This set constitutes the basic data carrier for subsequent module reconstruction and task reproduction.

[0158] The task identifier serves as a unique logical key throughout the entire task execution process. By binding it to the module parameter set, it generates module registration metadata. This metadata records the mapping relationship between the model structure, parameter configuration, and logical identifier used in the current task, and is a crucial foundation for the traceability and reusability of the task library.

[0159] To achieve unified management across platforms and processes, module registration metadata needs to be converted into a standard binary serialization format to generate a standardized module configuration file. This file encapsulates the model hierarchy, weight parameter tensors, parameter type identifiers, and index mapping rules. The serialization format can be a general-purpose protocol such as Protocol Buffers, MessagePack, or FlatBuffers, which offer structured processing, high compression rates, and fast decoding capabilities.

[0160] The module configuration index table is a mapping structure in the task library used to locate configuration files. It uses the task identifier as the primary key, and the configuration file path, parameter hash value, and timestamp as secondary fields. When the task library receives new module configuration data, it needs to check if the identifier already exists. If not, it adds a record; if it already exists, it updates the record and compares the hash value.

[0161] Integrity verification verifies the integrity of standardized module configuration files by calculating hash values ​​or CRC32 checksums to confirm their unaltered content. Loadability verification verifies the legality and actual operability of the data structure by loading the module structure and parameters in a sandbox environment and performing forward inference calculations. Finally, a module storage completion status signal is generated, indicating that the task configuration has been officially stored in the database, supporting subsequent reuse or version comparison.

[0162] Example Description: In offline business scenarios within the financial industry, account managers often need to collect images of customer-signed documents, ID cards, storefronts, and event venues using mobile devices such as smartphones and tablets. These images serve as risk control documentation and marketing materials. The quality of these raw images often varies due to inconsistent device models, arbitrary shooting angles, and uncontrollable weather conditions (such as rain or moisture interference). Some images may suffer from blurriness, low resolution, or fine-grained noise, severely impacting subsequent compliance review and image content recognition efficiency. To address these issues, the system first reads the raw images from the image acquisition device and determines whether the image format conforms to internal image processing specifications, such as unified resolution standards, color encoding spaces (e.g., RGB), and image compression formats (e.g., JPEG, PNG). If the format is incompatible, the system automatically performs image format conversion and resolution adjustment, and performs color standardization processing, converting the data into standardized image data for subsequent tasks. Based on the current image processing requirements, the system selects the "Image Clarification Enhancement" task from the task type library, generates a unique task identifier, encodes this identifier as a binary vector of a specific dimension, and binds it to the standardized image data through a data channel to form composite image data. The task identifier also serves as the logical key for all subsequent module selection and configuration. The system queries the task library for the perception module configuration record matching the current task identifier, loads the module structure containing convolutional kernel weights and batch normalization parameters, constructs the initial perception module, and inputs the bound composite image data. The image undergoes multi-layer convolution processing and batch normalization in the perception module to extract multi-scale features and perform channel compression, generating the first intermediate feature data. This first intermediate feature data is then input into the image processing core network with frozen weights. Through cross-channel feature fusion and multi-layer dimensionality reduction, it obtains semantic features with more abstract expressive power. Finally, the features are projected onto the task space in the fully connected layer, outputting the second intermediate feature data. The system further retrieves the output module associated with the task identifier from the task library, loads the structure containing deconvolution kernel parameters and upsampling ratio information, and performs dimensionality expansion and image-level reconstruction operations on the second intermediate feature data, including deconvolution processing and bilinear interpolation upscaling, ultimately outputting the target image with enhanced clarity. Upon completion of the task, the system automatically extracts all parameters used in the current task execution, including the convolution kernel and normalization parameters of the perceptual module, and the deconvolution kernel and upsampling ratio parameters of the output module, and encapsulates them into a module parameter set. This set is bound to the task identifier, generating module registration metadata, which is then serialized into a binary configuration file, updating the task library index table. After the task system verifies the integrity and loadability of the configuration file, it is successfully stored, ensuring automatic processing and task reuse for similar images in the future.

[0163] In medical scenarios such as pathological image processing, fundus screening image enhancement, or chest X-ray noise removal, medical institutions often need to acquire high-precision images for auxiliary diagnosis and AI model inference. However, due to differences in acquisition equipment, image compression mechanisms, and inconsistent image standardization processes among different hospitals, image data quality fluctuates greatly, affecting the generalization ability and stability of automated analysis systems. The image processing system performs unified format verification on image data uploaded from hospitals of different sources. If the image resolution is lower than the threshold required for AI diagnosis, resolution expansion is automatically performed. If the image encoding is inconsistent (e.g., grayscale and color images are mixed), color space standardization and format encapsulation are performed to generate standardized image data that meets the requirements of the diagnostic platform. Based on the specialty or task type (e.g., pathological image reconstruction or noise suppression) of the current input image, the system generates a corresponding task identifier from the task type library and encodes this identifier as a binary channel, embedding it into the image data to form composite image data. The system then determines the appropriate perception module configuration, loads a pre-trained feature extraction module, and completes multi-level feature encoding and compression of the standardized image input to obtain the first intermediate feature data. Subsequently, this intermediate feature data is input into the image processing core network with frozen parameters. This core network, constructed during the training phase by aggregating large-scale pathological image data, currently only performs forward inference to ensure the uniformity of output features. The output second intermediate feature data contains lesion information or structural edge information in the semantic space. The system selects the output module corresponding to the current task identifier to complete high-quality reconstruction processing. If the processing task is image enhancement, deconvolution is performed to restore details, combined with 4x upsampling interpolation to improve resolution, outputting a clear enhanced medical image. If it is a noise reduction task, parameter-guided filtering kernels are used to suppress unstructured noise during deconvolution, improving image diagnostics. After task execution, the system automatically extracts and saves all used model parameters and module configurations, performs serialization and integrity verification of the configuration files, and registers them in the task library index structure. Through this process, hospitals can quickly call the corresponding module configurations for different diseases and image types, improving diagnostic efficiency and image processing consistency, meeting data standardization and regulatory requirements, and cross-platform deployment needs.

[0164] This embodiment achieves unified extraction, standardized storage, and index management of the structural parameters of the perception and output modules after task execution, ensuring consistency and traceability of task configurations across multiple executions. Cross-platform compatibility is enhanced through binary serialization, and efficient retrieval and version replacement are supported through an indexing mechanism. File integrity and loadability verification mechanisms guarantee the security and operability of the task library data. The overall structure possesses high maintainability, low redundancy, and reproducible task configuration capabilities, providing support for the engineering deployment of multi-task models.

[0165] In one embodiment, a task identifier-based image processing apparatus is provided, which corresponds one-to-one with the task identifier-based image processing method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image processing device based on task identifiers of the present invention. The module includes an image preprocessing module 10, a task binding module 20, a perception module 30, a core feature network module 40, and an output module 50. Detailed descriptions of each functional module are as follows:

[0166] Image preprocessing module 10 is used to acquire the raw image data to be processed and convert the raw image data into a preset image format to generate standardized image data;

[0167] The task binding module 20 is used to bind the standardized image data with a task identifier that specifies the target task type for image processing, thereby generating composite image data.

[0168] The perception module 30 is used to select a corresponding perception module from the task library according to the task identifier, input the composite image data into the perception module, and process the composite image data through the perception module to generate first intermediate feature data.

[0169] The core feature network module 40 is used to input the first intermediate feature data into the image processing core network, and process the first intermediate feature data through the image processing core network to generate the second intermediate feature data.

[0170] The output module 50 is used to select a corresponding output module from the task library according to the task identifier, input the second intermediate feature data into the output module, and process the second intermediate feature data through the output module to generate target image data.

[0171] In one embodiment, the image preprocessing module 10 is specifically used for:

[0172] Read raw image data from an image acquisition device, the raw image data containing a first format type;

[0173] Verify whether the first format type of the original image data conforms to the preset image format requirements;

[0174] When the first format type does not meet the preset image format requirements, an image format conversion operation is performed to generate converted image data of the second format type;

[0175] The converted image data or the original image data conforming to the preset image format is input into the resolution adjustment module to generate adjusted image data with the preset resolution;

[0176] The adjusted image data is subjected to color space normalization processing to generate standardized color image data;

[0177] The standardized color image data is encapsulated into standardized image data containing format identifiers, resolution identifiers, and color identifiers.

[0178] In one embodiment, the task binding module 20 is specifically used for:

[0179] Based on a predefined task type library, generate a unique task code corresponding to the target task type;

[0180] The unique task code is converted into a binary code vector of a preset dimension;

[0181] Create a metadata tag containing the binary encoded vector;

[0182] The metadata tags are associated with and stored with standardized image data;

[0183] The binary encoded vector is superimposed as an independent data channel onto the channel dimension of the standardized image data to generate a composite image data with complete binding.

[0184] In one embodiment, the sensing module 30 is specifically used for:

[0185] Extract binary encoded vectors from the independent data channels of the composite image data;

[0186] Perform a decoding operation on the binary encoded vector to generate a target task type identifier;

[0187] Based on the target task type identifier, query the module registry in the task library and match the corresponding perception module configuration parameters. The perception module configuration parameters include convolution kernel weight recombination and batch normalization layer parameters.

[0188] Parse the network structure definition file in the configuration parameters of the perception module to generate the initialization structure of the perception module;

[0189] The convolutional kernel weights are recombined and injected into the kernel weight storage area of ​​each convolutional layer of the initialization perception module structure, and the batch normalized layer parameters are written into the batch normalized layer parameter storage area of ​​the initialization perception module structure to generate a perception module with loaded parameters.

[0190] Separate the standardized image data and binary encoded vector from the composite image data to generate the separated standardized image data;

[0191] The separated and standardized image data is input into the perception module, and multi-level convolution operations are performed through the convolution kernel weight reorganization. The output features are then batch-standardized using the batch standardization layer parameters to generate standardized feature data.

[0192] Perform channel dimension compression on the standardized feature data to generate first intermediate feature data.

[0193] In one embodiment, the core feature network module 40 is specifically used for:

[0194] Initialize the weight parameters of the convolutional and fully connected layers in the image processing core network, and freeze the weight update function of the convolutional and fully connected layers in the image processing core network;

[0195] Perform a cross-channel feature fusion operation on the first intermediate feature data to generate fused feature data;

[0196] The fused feature data is subjected to multi-level feature map dimensionality reduction processing to generate dimensionality-reduced feature data;

[0197] The dimensionality-reduced feature data is standardized to generate standardized feature data;

[0198] The standardized feature data is input into the fully connected layer of the image processing core network, and a dimension mapping operation is performed through a linear transformation matrix to generate second intermediate feature data.

[0199] In one embodiment, the output module 50 is specifically used for:

[0200] Based on the task identifier, an output module selection instruction is generated;

[0201] Based on the output module selection instruction, query the module configuration index table in the task library to obtain the output module file path;

[0202] Load the output module containing the deconvolution kernel parameter set and upsampling ratio parameter according to the output module file path;

[0203] Perform a channel dimension expansion operation on the second intermediate feature data to generate expanded feature data;

[0204] The extended feature data is input into the output module, and the deconvolution operation corresponding to the deconvolution kernel parameter set is executed to generate deconvolution output data.

[0205] The deconvolution output data is subjected to interpolation and amplification processing corresponding to the upsampling ratio parameter to generate amplified image data;

[0206] The magnified image data is normalized to generate target image data.

[0207] In one embodiment, the output module 50 is specifically used for:

[0208] The convolutional kernel weight reassembly and batch normalization layer parameters are extracted from the convolutional kernel weight storage area and the batch normalization layer parameter storage area of ​​the perception module, respectively.

[0209] Extract the deconvolution kernel weight reassembly from the deconvolution kernel parameter storage area of ​​the output module, and read the upsampling ratio parameter stored in the output module;

[0210] Generate a module parameter set that includes the convolution kernel weight reorganization, batch normalization layer parameters, deconvolution kernel weight reorganization, and upsampling ratio parameters;

[0211] Bind the task identifier to the module parameter set to generate module registration metadata;

[0212] The module registration metadata is converted into a preset binary serialization format to generate a standardized module configuration file;

[0213] Update the module configuration index table of the task library according to the task identifier, and establish a mapping relationship between the standardized module configuration file and the task identifier;

[0214] Verify the integrity and loadability of the standardized module configuration file, and generate a module storage completion status signal.

[0215] In one embodiment, a deterministic device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the deterministic machine device includes a processor, memory, network interface, and database connected via a system bus. The processor provides deterministic and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, the deterministic machine program, and the database. The internal memory provides an environment for the operation of the operating system and the deterministic machine program stored in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When executed by the processor, the deterministic machine program implements the functions or steps of a task identifier-based image processing method on the server side.

[0216] In one embodiment, a determining device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the deterministic device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides deterministic and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and the deterministic program. The internal memory provides an environment for the operation of the operating system and the deterministic program stored in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the deterministic program implements the user-side functions or steps of a task identifier-based image processing method.

[0217] In one embodiment, a deterministic machine device is provided, including a memory, a processor, and a deterministic machine program stored in the memory and executable on the processor. When the processor executes the deterministic machine program, it performs the following steps:

[0218] Acquire the raw image data to be processed, and convert the raw image data into a preset image format to generate standardized image data;

[0219] The standardized image data is bound to a task identifier that specifies the target task type for image processing to generate composite image data;

[0220] The corresponding perception module is selected from the task library according to the task identifier, and the composite image data is input into the perception module. The perception module processes the composite image data to generate the first intermediate feature data.

[0221] The first intermediate feature data is input into the image processing core network, and the image processing core network processes the first intermediate feature data to generate the second intermediate feature data.

[0222] The corresponding output module is selected from the task library according to the task identifier, and the second intermediate feature data is input into the output module. The output module processes the second intermediate feature data to generate target image data.

[0223] In one embodiment, a deterministic machine-readable storage medium is provided, on which a deterministic machine program is stored, the deterministic machine program performing the following steps when executed by a processor:

[0224] Acquire the raw image data to be processed, and convert the raw image data into a preset image format to generate standardized image data;

[0225] The standardized image data is bound to a task identifier that specifies the target task type for image processing to generate composite image data;

[0226] The corresponding perception module is selected from the task library according to the task identifier, and the composite image data is input into the perception module. The perception module processes the composite image data to generate the first intermediate feature data.

[0227] The first intermediate feature data is input into the image processing core network, and the image processing core network processes the first intermediate feature data to generate the second intermediate feature data.

[0228] The corresponding output module is selected from the task library according to the task identifier, and the second intermediate feature data is input into the output module. The output module processes the second intermediate feature data to generate target image data.

[0229] It should be noted that the functions or steps described above regarding determining the machine-readable storage medium or the machine device can be referred to in the relevant descriptions on the server side and user side in the aforementioned method embodiments. To avoid repetition, they will not be described one by one here.

[0230] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a deterministic machine program instructing related hardware. The deterministic machine program can be stored in a non-volatile deterministic machine-readable storage medium. When executed, the deterministic machine program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0231] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0232] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method of image processing based on a task identifier, characterized by, Includes the following steps: Acquire the raw image data to be processed, and convert the raw image data into a preset image format to generate standardized image data; The process of binding the standardized image data with a task identifier for a specified target task type in image processing to generate composite image data includes: generating a unique task code corresponding to the target task type according to a predefined task type library; converting the unique task code into a binary encoding vector of a preset dimension; creating a metadata tag containing the binary encoding vector; associating the metadata tag with the standardized image data; and superimposing the binary encoding vector as an independent data channel onto the channel dimension of the standardized image data to generate the bound composite image data. The corresponding perception module is selected from the task library according to the task identifier, and the composite image data is input into the perception module. The perception module processes the composite image data to generate the first intermediate feature data. The first intermediate feature data is input into the image processing core network, and the image processing core network processes the first intermediate feature data to generate the second intermediate feature data. The corresponding output module is selected from the task library according to the task identifier, and the second intermediate feature data is input into the output module. The output module processes the second intermediate feature data to generate target image data.

2. The image processing method based on task identifiers as described in claim 1, characterized in that, Acquire the raw image data to be processed, and convert the raw image data into a preset image format to generate standardized image data, including: Read raw image data from an image acquisition device, the raw image data containing a first format type; Verify whether the first format type of the original image data conforms to the preset image format requirements; When the first format type does not meet the preset image format requirements, an image format conversion operation is performed to generate converted image data of the second format type; The converted image data or the original image data conforming to the preset image format is input into the resolution adjustment module to generate adjusted image data with the preset resolution; The adjusted image data is subjected to color space normalization processing to generate standardized color image data; The standardized color image data is encapsulated into standardized image data containing format identifiers, resolution identifiers, and color identifiers.

3. The task identifier based image processing method of claim 1, wherein, The corresponding perception module is selected from the task library according to the task identifier, and the composite image data is input into the perception module. The perception module processes the composite image data to generate first intermediate feature data, including: Extract binary encoded vectors from the independent data channels of the composite image data; Perform a decoding operation on the binary encoded vector to generate a target task type identifier; Based on the target task type identifier, query the module registry in the task library and match the corresponding perception module configuration parameters. The perception module configuration parameters include convolution kernel weight recombination and batch normalization layer parameters. Parse the network structure definition file in the configuration parameters of the perception module to generate the initialization structure of the perception module; The convolutional kernel weights are recombined and injected into the kernel weight storage area of ​​each convolutional layer of the initialization perception module structure, and the batch normalized layer parameters are written into the batch normalized layer parameter storage area of ​​the initialization perception module structure to generate a perception module with loaded parameters. Separate the standardized image data and binary encoded vector from the composite image data to generate the separated standardized image data; The separated and standardized image data is input into the perception module, and multi-level convolution operations are performed through the convolution kernel weight reorganization. The output features are then batch-standardized using the batch standardization layer parameters to generate standardized feature data. Perform channel dimension compression on the standardized feature data to generate first intermediate feature data.

4. The task identifier based image processing method of claim 1, wherein, The first intermediate feature data is input into the image processing core network, and the image processing core network processes the first intermediate feature data to generate second intermediate feature data, including: Initialize the weight parameters of the convolutional and fully connected layers in the image processing core network, and freeze the weight update function of the convolutional and fully connected layers in the image processing core network; Perform a cross-channel feature fusion operation on the first intermediate feature data to generate fused feature data; The fused feature data is subjected to multi-level feature map dimensionality reduction processing to generate dimensionality-reduced feature data; The dimensionality-reduced feature data is standardized to generate standardized feature data; The standardized feature data is input into the fully connected layer of the image processing core network, and a dimension mapping operation is performed through a linear transformation matrix to generate second intermediate feature data.

5. The task identifier based image processing method of claim 1, wherein, The corresponding output module is selected from the task library according to the task identifier, and the second intermediate feature data is input into the output module. The output module processes the second intermediate feature data to generate target image data, including: Based on the task identifier, an output module selection instruction is generated; Based on the output module selection instruction, query the module configuration index table in the task library to obtain the output module file path; Load the output module containing the deconvolution kernel parameter set and upsampling ratio parameter according to the output module file path; Perform a channel dimension expansion operation on the second intermediate feature data to generate expanded feature data; The extended feature data is input into the output module, and the deconvolution operation corresponding to the deconvolution kernel parameter set is executed to generate deconvolution output data. The deconvolution output data is subjected to interpolation and amplification processing corresponding to the upsampling ratio parameter to generate amplified image data; The magnified image data is normalized to generate target image data.

6. The task identifier based image processing method of claim 1, wherein, After selecting a corresponding output module from the task library based on the task identifier, and inputting the second intermediate feature data into the output module, and processing the second intermediate feature data through the output module to generate target image data, the process further includes: The convolutional kernel weight reassembly and batch normalization layer parameters are extracted from the convolutional kernel weight storage area and the batch normalization layer parameter storage area of ​​the perception module, respectively. Extract the deconvolution kernel weight reassembly from the deconvolution kernel parameter storage area of ​​the output module, and read the upsampling ratio parameter stored in the output module; Generate a module parameter set that includes the convolution kernel weight reorganization, batch normalization layer parameters, deconvolution kernel weight reorganization, and upsampling ratio parameters; Bind the task identifier to the module parameter set to generate module registration metadata; The module registration metadata is converted into a preset binary serialization format to generate a standardized module configuration file; Update the module configuration index table of the task library according to the task identifier, and establish a mapping relationship between the standardized module configuration file and the task identifier; Verify the integrity and loadability of the standardized module configuration file, and generate a module storage completion status signal.

7. An image processing apparatus based on a task identifier, characterized by, The task identifier-based image processing device includes: The image preprocessing module is used to acquire the raw image data to be processed and convert the raw image data into a preset image format to generate standardized image data; The task binding module is used to bind the standardized image data with a task identifier for a specified target task type of image processing to generate composite image data. This includes: generating a unique task code corresponding to the target task type based on a predefined task type library; converting the unique task code into a binary encoding vector of a preset dimension; creating a metadata tag containing the binary encoding vector; associating the metadata tag with the standardized image data; and overlaying the binary encoding vector as an independent data channel onto the channel dimension of the standardized image data to generate the bound composite image data. The perception module is used to select the corresponding perception module from the task library according to the task identifier, input the composite image data into the perception module, and process the composite image data through the perception module to generate first intermediate feature data; The core feature network module is used to input the first intermediate feature data into the image processing core network, and process the first intermediate feature data through the image processing core network to generate the second intermediate feature data. The output module is used to select a corresponding output module from the task library according to the task identifier, input the second intermediate feature data into the output module, and process the second intermediate feature data through the output module to generate target image data.

8. A determination machine device, comprising: The determining device includes a memory, a processor, and a task identifier-based image processing program stored in the memory and executable on the processor. When executed by the processor, the task identifier-based image processing program implements the steps of the task identifier-based image processing method as described in any one of claims 1-6.

9. A computer readable storage medium, characterized in that, The storage medium stores an image processing program based on a task identifier, which, when executed by a processor, implements the steps of the image processing method based on a task identifier as described in any one of claims 1-6.

Citation Information

Patent Citations

  • High-resolution image transmission method and system based on optical communication

    CN118741055A

  • SAR (Synthetic Aperture Radar) target identification method based on multi-scale perception and reference attention

    CN118781483A