Image processing method, device, electronic device, storage medium and program product

CN119810472BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411896056.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-09-23
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

How to improve the presentation of detailed features while maintaining the overall structure of the image, making the detailed features clearer and more visible.

Method used

By extracting features from the original image of the target object, the edge features are extracted from the feature map using bidirectional strip convolution and contrastive learning, and an edge feature enhanced image is generated.

Benefits of technology

The enhanced image not only maintains the complete display of the overall structure of the target object, but also further refines the detailed features, making them clearer and more visible, especially the texture and grain features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810472B_ABST
    Figure CN119810472B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing method, apparatus, electronic device, storage medium, and program product, relating to the field of image processing technology. The method comprises: extracting features from an original image of a target object to obtain a feature map; performing a bidirectional strip convolution operation on the feature map based on a segmentation mask of the feature map to extract original edge features from the feature map; classifying the original edge features through contrastive learning to obtain classified feature information; and generating an edge feature-enhanced image of the target object based on the classified feature information and the original image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, specifically to the field of image acquisition and processing technology, and more particularly to an image processing method, device, electronic device, storage medium, and program product. Background Art

[0002] In the image field, the original image often includes two parts: overall structure and detailed features.

[0003] Accordingly, how to improve the presentation effect of image details while maintaining the overall structure of the image so that the details are more clearly visible is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] The embodiments of the present disclosure provide an image processing method, apparatus, electronic device, storage medium, and program product.

[0005] In a first aspect, an embodiment of the present disclosure proposes an image processing method, comprising: performing feature extraction on an original image of a target object to obtain a feature map; performing a bidirectional strip convolution operation on the feature map based on a segmentation mask of the feature map to extract original edge features from the feature map; classifying the original edge features through contrastive learning to obtain classification feature information; and generating an edge feature enhanced image of the target object based on the classification feature information and the original image.

[0006] In the second aspect, an embodiment of the present disclosure proposes an image processing device, including: a feature map extraction module, configured to extract features from an original image of a target object to obtain a feature map; an edge feature extraction module, configured to perform a bidirectional strip convolution operation on the feature map based on a segmentation mask of the feature map, and extract original edge features from the feature map; a classification feature generation module, configured to classify the original edge features through contrastive learning to obtain classification feature information; and an enhanced image generation module, configured to generate an edge feature enhanced image of the target object based on the classification feature information and the original image.

[0007] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the image processing method described in any implementation method in the first aspect when executing the instructions.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the image processing method described in any implementation manner in the first aspect when executed.

[0009] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the image processing method described in any implementation manner in the first aspect.

[0010] The disclosed embodiment, based on the original image of the target object, performs enhancement processing on the original image through a series of processing processes such as feature extraction, bidirectional strip convolution, and feature classification. The enhanced image after the enhancement processing not only maintains a complete display of the overall structure of the target object, but also further refines the detailed features that may appear in the target object, so that these detailed features are more clearly visible, and the scope of the target objects targeted by the image processing is increased, so that some target objects containing subtle texture features and line features can be displayed more clearly.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings:

[0013] Figure 1 is an exemplary system architecture in which the present disclosure may be applied;

[0014] Figure 2 A flowchart of an image processing method provided in an embodiment of the present disclosure;

[0015] Figure 3 A flowchart of another image processing method provided by an embodiment of the present disclosure;

[0016] Figure 4 A flowchart of another image processing method provided in an embodiment of the present disclosure;

[0017] Figure 5 A flowchart of another image processing method provided in an embodiment of the present disclosure;

[0018] Figure 6 A flowchart of an image processing method in an application scenario provided by an embodiment of the present disclosure;

[0019] Figure 7 A structural block diagram of an image processing device provided in an embodiment of the present disclosure;

[0020] Figure 8 A schematic structural diagram of an electronic device suitable for executing an image processing method provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other unless there is a conflict.

[0022] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0023] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the image processing method, apparatus, electronic device, and computer-readable storage medium of the present disclosure can be applied.

[0024] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0025] Users can use terminal devices 101, 102, 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, 103 and server 105 may be installed with various applications for enabling information communication between them, such as instant messaging applications.

[0026] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.

[0027] Server 105 can provide various services through various built-in applications. It should be noted that, in addition to being obtained from terminal devices 101, 102, and 103 via network 104, the original images can also be pre-stored locally on server 105 in various ways. Therefore, when server 105 detects that such data is already stored locally, it can choose to directly obtain the data locally. In this case, exemplary system architecture 100 may also not include terminal devices 101, 102, 103 and network 104.

[0028] Because optimization processes based on the original image may require significant computational resources and high computing power, the image processing methods provided in the subsequent embodiments of this disclosure are generally performed by a server 105 with significant computational power and resources. Accordingly, the image processing apparatus is generally located within the server 105. However, it should also be noted that, if the terminal devices 101, 102, and 103 also possess sufficient computational power and resources, the terminal devices 101, 102, and 103 may also utilize image processing applications installed thereon to perform the aforementioned computations delegated to the server 105, thereby outputting the same results as the server 105. In particular, in the presence of multiple terminal devices with varying computational power, if the image processing application determines that the terminal device it is in possession of possesses significant computational power and a significant amount of remaining computational resources, it may allow the terminal device to perform the aforementioned computations, thereby appropriately alleviating the computational burden on the server 105. Accordingly, the image processing apparatus may also be located within the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.

[0029] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0030] Please refer to Figure 2 , Figure 2 This is a flowchart of an image processing method provided by an embodiment of the present disclosure, wherein process 200 includes the following steps:

[0031] Step 201: extract features from the original image of the target object to obtain a feature map.

[0032] This step is intended to be performed by the execution subject of the image processing method (e.g. Figure 1 The server 105 shown in the figure obtains the original image of the target object and performs feature extraction on the original image to obtain a feature map. The target object refers to an object whose image can be obtained by taking a photo or other means, such as a person, document, book, cultural relic, etc. The original image of the target object can be a photo obtained by taking a photo with a camera or other photographic device. The execution subject can be controlling the camera to complete the photo taking, thereby obtaining the original image of the target object, or it can be obtaining a pre-stored photo from the cloud or other server, etc., and this embodiment is not limited to this. Feature extraction refers to extracting information or features that are useful for tasks (such as classification, recognition, detection, etc.) from raw data. In the field of image processing, feature extraction generally refers to extracting descriptors or feature vectors that can represent the content of the original image from the original image to form a feature map.

[0033] Step 202: Perform a bidirectional strip convolution operation on the feature map based on the segmentation mask of the feature map to extract the original edge features from the feature map.

[0034] This step aims to extract the original edge features from the feature map through a bidirectional striped convolution operation by the aforementioned execution entity. In this embodiment, the execution entity can use a classifier to classify each pixel or region based on the information in the feature map, assign a category label to each pixel or region based on the classification results, and combine them to generate a segmentation mask. A bidirectional striped convolution operation is then performed based on the segmentation mask. A striped convolution operation is a convolution operation in which the convolution kernel is striped. The convolution kernel of a striped convolution is much smaller in one dimension (usually width or height) than in another dimension, which enables it to better capture strip-like features in the data. In scenarios such as image segmentation and object detection, striped convolution can better capture slender objects or structures in the image (for example, slender objects in an image, character sequences in text, etc.). Bidirectional striped convolution refers to a striped convolution operation that considers both forward and backward information, thereby extracting more comprehensive feature information to obtain the original edge features. In practical applications, the above process can be used to highlight the texture features of one or more specific regions of the target object. Exemplarily, the execution entity may also generate an adaptive attention weight map based on the obtained original edge features, and adjust the attention weights for different areas and different features in the original image through the attention weight map, and use the attention weights to characterize and distinguish areas or features that need to be focused on.

[0035] Step 203: Classify the original edge features by contrastive learning to obtain classification feature information.

[0036] This step is intended to enable the aforementioned execution subject to classify the original edge features through contrastive learning to obtain classification feature information. Contrastive learning is a self-supervised learning method that learns data representation by comparing positive samples (e.g., similar but different samples after data augmentation or transformation) with negative samples (e.g., samples that are different from the positive samples), so that positive samples are closer in the representation space and negative samples are farther apart. This representation learning method can extract key features from the data and provide strong support for subsequent classification tasks. In this embodiment, the original edge features can be classified through contrastive learning to obtain classification feature information.

[0037] Step 204: Generate an edge feature enhanced image of the target object based on the classification feature information and the original image.

[0038] This step involves the aforementioned execution subject extracting classification feature information from the original image and combining it with the original image to generate an edge-feature-enhanced image of the target object. This process allows the edge-feature-enhanced image to present the overall structure of the target object while highlighting local details (i.e., the features corresponding to the edge features).

[0039] The image processing method provided by the embodiment of the present disclosure, based on the original image of the target object, enhances the original image through a series of processing processes such as feature extraction, bidirectional strip convolution, and feature classification. The method ensures that the enhanced image after the enhancement process not only maintains a complete display of the overall structure of the target object, but also further refines the detailed features that may appear in the target object, so that these detailed features are more clearly visible, and increases the scope of the target objects targeted by the image processing, and can more clearly display some target objects containing subtle texture features and line features.

[0040] In some optional implementations of the disclosed embodiments, step 203, classifying the original edge features using contrastive learning to obtain classification feature information, primarily includes: first, extracting similar edge features and dissimilar edge features from the edge feature information; then, adjusting the classification feature information by reducing the feature distance between similar edge features and / or increasing the feature distance between similar and dissimilar edge features. For example, the classification process using contrastive learning can begin by performing data augmentation. Common data augmentation techniques include cropping, flipping, rotation, random cropping, and color transformation. By generating diverse instances, data variability is increased. An encoder network is then trained. The encoder network takes the data augmented instances as input and maps them into a latent representation space, capturing meaningful features and similarities. The encoder network is typically a deep neural network architecture, such as a convolutional neural network (CNN) for image data or a recurrent neural network (RNN) for sequential data. The network learns to extract and encode high-level representations from the data augmented instances, thereby facilitating the distinction between similar and dissimilar instances in subsequent steps. A projection network can then be used to take the output of the encoder network and project it into a lower-dimensional space, often called a projection space or embedding space. This additional projection step helps enhance the discriminative power of the learned representations. By mapping the representations into a lower-dimensional space, the projection network reduces the complexity and redundancy of the data, helping to better separate similar and dissimilar instances. Once the enhanced instances are encoded and projected into the embedding space, contrastive learning is applied, aiming to maximize the consistency between positive pairs (instances from the same sample) and minimize the consistency between negative pairs (instances from different samples). This results in closer feature distances between similar edge features and greater distances between similar and dissimilar edge features, enabling classification and obtaining categorical feature information.

[0041] It should be noted that the above process is merely an example. In a specific implementation, steps such as data enhancement or projection network may be omitted as needed. Furthermore, for the distance change between similar edge features and non-similar edge features, only one type may be adjusted based on actual conditions. This embodiment is not limited to this.

[0042] In the embodiment of the present disclosure, the above process is a process of enhancing the original image through a series of processing processes such as feature extraction, bidirectional strip convolution, and feature classification. One of the main functions of this process is to focus on enhancing texture features. In practical applications, in addition to enhancing texture features, other features of the target object can also be enhanced. In some optional implementations of the embodiment of the present disclosure, such as Figure 3 As shown, the image processing method further includes:

[0043] Step 301: Based on a pre-trained image processing model, surface feature recognition is performed on the original image to obtain a feature recognition result.

[0044] In this embodiment, the training process of the image processing model mainly includes the following steps:

[0045] Step 1: Use the initial neural network model to extract high-level and low-level features from the sample image. In fields such as image processing, computer vision, and deep learning, high-level features, also known as high-level semantic features, typically refer to features extracted from deeper layers of deep learning models (such as convolutional neural networks). These features contain rich combinatorial information. Low-level features, also known as underlying features, typically refer to features extracted from shallower layers of deep learning models. These features are closer to the image's raw pixel data. Low-level image features typically include contours, edges, color, texture, and shape features. High-level and low-level features each have their own advantages and disadvantages in image understanding and analysis, and they are complementary. High-level features provide the overall content and meaning of the image, while low-level features provide local details and accurate object location information.

[0046] Step 2: Calculate the difference vector based on high-level features and low-level features.

[0047] After extracting the high-level and low-level features, the execution entity may calculate a difference vector based on the high-level and low-level features. For example, the obtained difference vector may undergo some post-processing, such as normalization or smoothing, to remove noise or improve the interpretability of the vector.

[0048] Step 3: Adjust the weight distribution of the neural network model based on the difference vector to obtain the image processing model.

[0049] During the training process of this initial neural network model, the difference vectors are used as a guide to calculate the gradients of the weights. These gradients indicate how the weights should be adjusted to reduce the differences between feature levels. Furthermore, an optimization algorithm (such as gradient descent) is used to update the neural network model's weights based on the calculated gradients. The goal of the weight update is to minimize the differences between high-level and low-level features, thereby improving model performance and ultimately achieving the image processing model.

[0050] It should be noted that, in this embodiment, the process of performing surface feature recognition on the original image based on a pre-trained image processing model to obtain feature recognition results can also be achieved through a neural network model such as CNN or RNN, and this embodiment is not limited to this.

[0051] Step 302: Generate a first optimized image based on the feature recognition result.

[0052] This step aims to generate a first optimized image of the original image based on the feature recognition results by the aforementioned execution entity. In this embodiment, through the aforementioned process and the use of a pre-trained image processing model, the surface features (e.g., bumps and undulations) detected in the original image of the target object can be more prominently detected. This not only emphasizes the major surface undulations of the target object, but also preserves the subtle gradations of bumps and undulations, thereby preserving the overall structure while also ensuring that fine local features are not neglected.

[0053] Step 303: Integrate the edge feature enhanced image and the first optimized image to obtain a first integrated image.

[0054] In this embodiment, through the above Figure 2 The edge-feature-enhanced image obtained in steps 201-204 of the illustrated embodiment and the first optimized image obtained in steps 301-302 can be further integrated to obtain a first integrated image. This first integrated image integrates detail features such as texture features highlighted by the edge-feature-enhanced image with surface features highlighted by the first optimized image, enabling a more comprehensive and specific representation of the features of the target object, allowing the first integrated image to more clearly and accurately depict the target object. The integration of the edge-feature-enhanced image and the first optimized image can be performed using an attention mechanism. For example, an attention mechanism is first applied to generate an attention weight map (attention distribution) that indicates which areas in the image are important for the integration task. The attention weight map (attention distribution) is then used to weight the extracted features to highlight features in important areas and suppress features in less important areas. Features from different input images are concatenated or weighted summed according to the attention weights to achieve feature integration, thereby obtaining the first integrated image.

[0055] It should be noted that, in this embodiment, the process of integrating different images is not limited to being achieved through the attention mechanism, but can also be achieved using more mature image fusion algorithms in related fields (such as Alpha fusion algorithm (alpha compositing), pyramid fusion algorithm, Poisson fusion algorithm, etc.). This embodiment is not limited to this.

[0056] In practical applications, in addition to enhancing texture features, other features of the target object can also be enhanced. In some optional implementations of the embodiments of the present disclosure, such as Figure 4 As shown, the image processing method further includes:

[0057] Step 401: The original image is optimized by adaptive histogram equalization and multi-scale processing to obtain a second optimized image. Adaptive histogram equalization is a computer image processing technology used for image enhancement, which aims to improve the contrast of the image. The image contrast can be changed by calculating the local histogram of the image and then redistributing the brightness. Multi-scale processing refers to the use of operations or representations of different scales to extract and analyze image information in image processing, which helps to capture multi-scale features in the image and improve the robustness and performance of image processing. In specific implementation, the execution entity can first perform multi-scale processing on the original image to extract feature information at different scales, and then perform adaptive histogram equalization on the image at each scale to enhance the local contrast of the image, and finally perform feature fusion to obtain the second optimized image. Through the above process, it is possible to simultaneously capture the multi-scale features in the original image and enhance the local contrast of the original image, thereby improving the performance and effect of image processing.

[0058] Step 402: Integrate the edge feature enhanced image, the first optimized image, and the second optimized image to obtain a second integrated image.

[0059] Furthermore, in this embodiment, through the above Figure 2In the illustrated embodiment, the edge feature-enhanced image obtained in steps 201-204, the first optimized image obtained in steps 301-302, and the second optimized image obtained in step 401 can be further integrated to obtain a second integrated image. This second integrated image integrates the texture features and other detail features highlighted by the edge feature-enhanced image, the surface features highlighted by the first optimized image, and the multi-scale features and contrast features highlighted by the second optimized image. This allows for a more comprehensive and specific representation of the features of the target object, enabling the second integrated image to more clearly and accurately depict the target object. This not only improves computational efficiency but also automatically adjusts the integration ratio based on the feature quality of different regions, ultimately outputting a balanced and complete integrated image. The integration process of the edge feature-enhanced image, the first optimized image, and the second optimized image can be performed using an attention mechanism. For example, the attention mechanism is first applied to generate an attention weight map (attention distribution) that indicates which regions in the image are important for the integration task. The attention weight map (attention distribution) is then used to weight the extracted features to highlight features in important regions and suppress features in unimportant regions. Features from different input images are concatenated or weighted-summed according to attention weights to achieve feature integration, thereby obtaining the second integrated image.

[0060] In some optional implementations of the disclosed embodiments, the executing entity may further integrate the edge feature-enhanced image and the second optimized image described in the above embodiments to obtain a third integrated image. This third integrated image integrates the detailed features, such as texture features, highlighted by the edge feature-enhanced image, with the multi-scale features and contrast features highlighted by the second optimized image. This can more comprehensively and specifically reflect the features of the target object, allowing the third integrated image to more clearly and accurately depict the target object. The integration of the edge feature-enhanced image and the second optimized image can be performed using an attention mechanism. For example, an attention mechanism is first applied to generate an attention weight map (attention distribution) that indicates which areas of the image are important for the integration task. The attention weight map (attention distribution) is then used to weight the extracted features to highlight features in important areas and suppress features in less important areas. Features from different input images are concatenated or weighted summed according to the attention weights to achieve feature integration, thereby obtaining the third integrated image.

[0061] Please refer to Figure 5 , Figure 5 This is a flowchart of another image processing method provided by an embodiment of the present disclosure, wherein process 500 includes the following steps:

[0062] Step 501: Acquire a test image obtained by photographing a target object with a photographing device in an arbitrary photographing posture.

[0063] In this embodiment, the original image of the target object can be obtained by photographing it using a camera. Furthermore, the camera can first photograph the target object in any shooting position using the camera to obtain a test image. The shooting position primarily refers to the position and angle of the camera during shooting. The executing entity can obtain the test image captured by the camera, as well as the shooting position used when capturing the test image.

[0064] Step 502: Determine the target shooting posture based on the actual shooting posture of the test image.

[0065] Based on the actual shooting posture of the test image captured by the camera, the execution entity can determine the target shooting posture to be used for the actual shooting of the target object. For example, to achieve fixed 4-view or 8-view shooting, the target object's outline and size information can be determined based on the test image. This outline and size information can then be combined to determine the specific shooting posture information (i.e., shooting angle and position, etc.) suitable for 4-view or 8-view shooting. For example, if the target object is large (such as a seismometer or other artifact, or an X-ray detector), the specific shooting posture determined can be within a larger capture area and have a wider range of shooting angles. On the other hand, if the target object is small (such as an oracle bone inscribed with oracle bone script or a miniature sculpture), the specific shooting posture determined can be within a smaller capture area, closer to the target alignment, and have a smaller range of shooting angles. The specific shooting posture can be adjusted based on the actual situation of the target object, and this embodiment is not limited to this.

[0066] Step 503: Acquire an original image obtained by photographing the target object by the photographing device in the target photographing posture.

[0067] After determining the target shooting posture for actual shooting, the shooting device can be used to actually shoot at the target shooting posture to obtain original images of the target object at different positions and angles, and the execution subject can obtain the original image.

[0068] Step 504: extract features from the original image of the target object to obtain a feature map.

[0069] Step 505: Perform a bidirectional strip convolution operation on the feature map based on the segmentation mask of the feature map to extract the original edge features from the feature map.

[0070] Step 506: Classify the original edge features by contrastive learning to obtain classification feature information.

[0071] Step 507: Generate an edge feature enhanced image of the target object based on the classification feature information and the original image.

[0072] The above steps 504-507 are similar to the following Figure 2 Steps 201-204 shown are consistent. For the same content, please refer to the corresponding part of the previous embodiment and will not be repeated here.

[0073] During implementation, the execution entity may also integrate edge feature images corresponding to original images captured at different shooting positions based on the target shooting posture. For example, for four-view shooting angles (including horizontal, downward, upward, and side views), original images of the target object at different angles can be obtained. Through steps 504-507 above, corresponding edge feature-enhanced images can be obtained. By combining the shooting postures corresponding to the four-view shooting angles, the execution entity can integrate edge feature-enhanced images of the target object at different angles.

[0074] Through the above process, on the one hand, automatic adjustment of shooting for target objects of different sizes and shapes can be achieved to improve shooting efficiency, and on the other hand, the characteristic information of all aspects of the target object can be reflected more comprehensively, accurately and clearly, so that the information of the target object presented in the obtained image is more in line with the actual situation of the target object.

[0075] The image processing methods described in any of the above embodiments, and the enhanced images and integrated images generated thereby, are all processes for generating images of a target object. In practical applications, in order to better present the target object, a three-dimensional modeling process may be further performed based on the enhanced images or integrated images obtained in any of the above embodiments to generate a three-dimensional model of the target object. In some optional implementations of the disclosed embodiments, the execution subject may further perform the following steps:

[0076] Step 508: Perform three-dimensional modeling on the target object based on the edge feature enhanced image and the target shooting posture to obtain a three-dimensional model of the target object.

[0077] In this embodiment, the process of the execution subject performing three-dimensional modeling of the target object based on the edge feature enhanced image and the target shooting posture can include multiple situations: the execution subject performs three-dimensional modeling directly based on the edge feature enhanced image and the target shooting posture, the execution subject performs three-dimensional modeling based on the first integrated image obtained by integrating the edge feature enhanced image and the first optimized image and the target shooting posture, the execution subject performs three-dimensional modeling based on the second integrated image obtained by integrating the edge feature enhanced image and the second optimized image and the target shooting posture, and the execution subject performs three-dimensional modeling based on the third integrated image obtained by integrating the edge feature enhanced image, the first optimized image and the second optimized image and the target shooting posture. The specific method used for three-dimensional modeling can be selected and configured according to actual needs, and this embodiment is not limited to this.

[0078] In some implementations of the present disclosure, step 508, performing three-dimensional modeling of the target object based on the edge feature enhanced image and the target shooting pose, to obtain the three-dimensional model of the target object, mainly includes:

[0079] Step 1: Group feature extraction is performed on the edge feature-enhanced image based on the target shooting pose to obtain modeling feature information. The target shooting pose includes the shooting position and posture, which determines the content and perspective of the original captured image, affecting the accuracy and authenticity of the 3D modeling. Therefore, the execution entity obtains the original image of the target object at different perspectives based on the target shooting pose, and processes it accordingly to obtain edge feature-enhanced images. Group feature extraction is performed on this edge feature-enhanced image (for example, group feature extraction is performed according to grid division) to obtain modeling feature information.

[0080] Step 2: Perform point cloud sampling based on modeling feature information to obtain point cloud data.

[0081] After acquiring the modeling feature information, the execution entity can perform point cloud sampling based on this information to obtain point cloud data. For example, for target areas with a high number of features within the target object, the point cloud sampling density can be increased, with sampling performed at a first sampling density. For non-target areas with fewer features (i.e., areas other than the target area), the point cloud sampling density can be reduced, with sampling performed at a second sampling density, where the first sampling density is greater than the second sampling density. This approach allows more data to be concentrated in the target area, ensuring that detailed features of the target object are sampled while also rationally allocating processing resources and improving processing efficiency.

[0082] Step 3: Perform mesh reconstruction and texture mapping based on point cloud data to obtain a three-dimensional model.

[0083] After acquiring the point cloud data through point cloud sampling, the execution subject can obtain a three-dimensional model of the target object through mesh reconstruction and texture mapping.

[0084] The image processing method of the disclosed embodiment performs three-dimensional modeling based on enhanced processing of the original image of the target object. This method can more accurately and clearly identify and display the detailed features and texture features of the target object. Furthermore, it can further incorporate these features into the three-dimensional model of the target object, enabling a more precise and comprehensive display of the content of the target object. This significantly improves the display quality of both the image and the model, significantly enhancing the modeling quality of small objects (such as oracle bones inscribed with oracle bone inscriptions) with rich surface details, and provides more reliable three-dimensional data support for subsequent digital research and maintenance work.

[0085] To deepen understanding, this disclosure also provides a specific implementation solution in combination with a specific application scenario, see Figure 6 Flow 600 is shown.

[0086] In this embodiment, the target object is a small cultural relic (such as an oracle bone with oracle bone inscriptions) as an example for description.

[0087] The process 600 mainly includes:

[0088] Step 601: Acquire a test image 3 of an oracle bone 2 captured by a photographing device 1 .

[0089] Step 602: Determine the target shooting posture based on the actual shooting posture of the oracle bone 2 captured by the trial image 3.

[0090] Step 603: Acquire the original image 8 captured by the shooting device 1 at the target shooting posture (shooting position 4-7).

[0091] Step 604: extract features from the original image 8 to obtain a feature map.

[0092] Step 605: Perform a bidirectional strip convolution operation on the feature map based on the segmentation mask of the feature map to extract the original edge features.

[0093] Step 606: Classify the original edge features through comparative learning to obtain classification feature information.

[0094] Step 607: Generate an edge feature enhanced image 9 of the oracle bone 2 based on the classification feature information and the original image 8.

[0095] Step 608 : Perform three-dimensional modeling of the oracle bone 2 based on the edge feature enhanced image 9 and the target shooting posture to obtain a three-dimensional model 10 of the oracle bone 2 .

[0096] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image processing device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0097] like Figure 7 As shown, the image processing device 700 of this embodiment may include: a feature map extraction module 701, an edge feature extraction module 702, a classification feature generation module 703 and an enhanced image generation module 704. Among them, the feature map extraction module 701 is configured to extract features from the original image of the target object to obtain a feature map. The edge feature extraction module 702 is configured to perform a bidirectional strip convolution operation on the feature map based on the segmentation mask of the feature map to extract the original edge features from the feature map. The classification feature generation module 703 is configured to classify the original edge features by contrastive learning to obtain classification feature information. The enhanced image generation module 704 is configured to generate an edge feature enhanced image of the target object based on the classification feature information and the original image.

[0098] In this embodiment, the specific processing of the feature map extraction module 701, the edge feature extraction module 702, the classification feature generation module 703 and the enhanced image generation module 704 and the technical effects thereof can be referred to in the respective Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiment are not repeated here.

[0099] In some optional implementations of this embodiment, the classification feature generation module 703 includes a feature extraction submodule configured to extract similar edge features and non-similar edge features from the edge feature information. The classification feature generation submodule is configured to adjust the classification feature information by reducing the feature distance between similar edge features and / or increasing the feature distance between similar edge features and non-similar edge features.

[0100] In some optional implementations of this embodiment, the image processing device further includes: a feature recognition module configured to perform surface feature recognition on the original image based on a pre-trained image processing model to obtain a feature recognition result. A first optimized image generation module configured to generate a first optimized image based on the feature recognition result. The training process of the image processing model includes the following steps: extracting high-level features and low-level features of the sample image respectively through an initial neural network model. Calculating a difference vector based on the high-level features and the low-level features. Adjusting the weight distribution of the neural network model based on the difference vector to obtain an image processing model.

[0101] In some optional implementations of this embodiment, the image processing device further includes: a first integrated image generation module configured to integrate the edge feature enhanced image and the first optimized image to obtain a first integrated image.

[0102] In some optional implementations of this embodiment, the image processing device further includes: a second optimized image generation module configured to optimize the original image through adaptive histogram equalization and multi-scale processing to obtain a second optimized image.

[0103] In some optional implementations of this embodiment, the image processing apparatus further includes: a second integrated image generation module configured to integrate the edge feature enhanced image, the first optimized image, and the second optimized image to obtain a second integrated image. The second integrated image generation module is further configured to fuse the edge feature enhanced image, the first optimized image, and the second optimized image based on their respective corresponding attention distributions to obtain the second integrated image.

[0104] In some optional implementations of this embodiment, the image processing apparatus further includes: a third integrated image generation module configured to integrate the edge feature enhanced image and the second optimized image to obtain a third integrated image.

[0105] In some optional implementations of this embodiment, the image processing device further includes: a test image acquisition module configured to acquire a test image of the target object captured by the camera in any shooting posture; a shooting posture determination module configured to determine the target shooting posture based on the actual shooting posture of the test image; and an original image acquisition module configured to acquire an original image of the target object captured by the camera in the target shooting posture.

[0106] The shooting posture determination module includes: a contour and size extraction module configured to extract contour information and size information of the target object based on the test image and the actual shooting posture; and a shooting posture determination submodule configured to determine the target shooting posture information based on the contour information and size information.

[0107] In some optional implementations of this embodiment, the image processing device further includes: an image integration module configured to integrate edge feature enhanced images corresponding to original images captured at different shooting positions based on the target shooting posture.

[0108] In some optional implementations of this embodiment, the image processing device further includes: a three-dimensional modeling module configured to perform three-dimensional modeling of the target object based on the edge feature enhanced image and the target shooting posture to obtain a three-dimensional model of the target object.

[0109] In some optional implementations of this embodiment, the three-dimensional modeling module includes: a modeling feature extraction submodule, configured to perform group feature extraction on the edge feature enhanced image based on the target shooting posture to obtain modeling feature information. A point cloud data acquisition submodule, configured to perform point cloud sampling based on the modeling feature information to obtain point cloud data. The three-dimensional modeling submodule is configured to perform mesh reconstruction and texture mapping based on the point cloud data to obtain a three-dimensional model. The point cloud data acquisition submodule is further configured to: based on the modeling feature information and the shooting posture information, collect first point cloud data of the target area with a first sampling density, and collect second point cloud data outside the target area with a second sampling density to obtain point cloud data. The first sampling density is greater than the second sampling density.

[0110] This embodiment exists as an apparatus embodiment corresponding to the above-mentioned method embodiment. The image processing apparatus provided by this embodiment, on the basis of acquiring the original image of the target object, performs enhancement processing on the original image through a series of processing processes such as feature extraction, bidirectional strip convolution, and feature classification. In this way, the enhanced image after the enhancement processing not only maintains a complete display of the overall structure of the target object, but also further refines the detail features that may appear in the target object, so that these detail features are more clearly visible, and the scope of the target objects targeted by the image processing is increased, so that some target objects containing subtle texture features and line features can be displayed more clearly.

[0111] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the image processing method described in any of the above embodiments can be implemented when the at least one processor executes them.

[0112] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the image processing method described in any of the above embodiments when executed.

[0113] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, which, when executed by a processor, can implement the image processing method described in any of the above embodiments.

[0114] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0115] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0116] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0117] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the image processing method. For example, in some embodiments, the image processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the image processing method by any other suitable means (e.g., via firmware).

[0118] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0122] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0123] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host. This is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.

[0124] According to the technical solution of the embodiment of the present disclosure, on the basis of obtaining the original image of the target object, the original image is enhanced through a series of processing processes such as feature extraction, bidirectional strip convolution, and feature classification, so that the enhanced image after the enhancement processing not only maintains a complete display of the overall structure of the target object, but also further refines the detailed features that may appear in the target object, so that these detailed features are more clearly visible, and the scope of the target objects targeted by the image processing is increased, and some target objects containing subtle texture features and line features can be displayed more clearly.

[0125] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0126] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An image processing method, comprising: Extract features from the original image of the target object to obtain a feature map; performing a bidirectional strip convolution operation on the feature map based on the segmentation mask of the feature map to extract original edge features from the feature map; Classifying the original edge features by contrastive learning to obtain classification feature information; generating an edge feature enhanced image of the target object based on the classification feature information and the original image; Based on a pre-trained image processing model, surface feature recognition is performed on the original image to obtain a feature recognition result; wherein the training process of the image processing model includes the following steps: extracting high-level features and low-level features of the sample image respectively through an initial neural network model; calculating a difference vector based on the high-level features and low-level features; adjusting the weight distribution of the neural network model based on the difference vector, so that the image processing model generates a first optimized image based on the feature recognition result; generating a first optimized image based on the feature recognition result; The edge feature enhanced image and the first optimized image are integrated to obtain a first integrated image.

2. The method according to claim 1, wherein The original edge features are classified by contrastive learning to obtain classification feature information, including: extracting similar edge features and non-similar edge features from the edge feature information; The classification feature information is obtained by adjusting in a manner of reducing the feature distance of the edge features of the same type and / or increasing the feature distance between the edge features of the same type and the edge features of different types.

3. The method according to claim 1, further comprising: The original image is optimized by adaptive histogram equalization and multi-scale processing to obtain a second optimized image.

4. The method according to claim 3, further comprising: The edge feature enhanced image, the first optimized image and the second optimized image are fused based on their respective corresponding attention distributions to obtain a second integrated image.

5. The method according to claim 1, further comprising: Acquire a test image of the target object captured by a shooting device in any shooting posture; Extracting the contour information and size information of the target object based on the test image and the actual shooting posture; Determine the target shooting posture based on the contour information and size information; The original image obtained by the shooting device shooting the target object in the target shooting posture is obtained.

6. The method according to claim 5, further comprising: Three-dimensional modeling is performed on the target object based on the edge feature enhanced image and the target shooting posture to obtain a three-dimensional model of the target object.

7. The method according to claim 6, wherein: The three-dimensional modeling is performed based on the edge feature enhanced image and the target shooting posture to obtain the three-dimensional model of the target object, including: performing group feature extraction on the edge feature enhanced image based on the target shooting posture to obtain modeling feature information; Performing point cloud sampling based on the modeling feature information to obtain point cloud data; Mesh reconstruction and texture mapping are performed based on the point cloud data to obtain the three-dimensional model.

8. The method according to claim 7, wherein: The performing point cloud sampling based on the modeling feature information to obtain point cloud data includes: Based on the modeling feature information and the shooting posture information, first point cloud data of the target area is collected with a first sampling density, and second point cloud data outside the target area is collected with a second sampling density to obtain the point cloud data; wherein the first sampling density is greater than the second sampling density.

9. An image processing device comprising: A feature map extraction module is configured to extract features from an original image of a target object to obtain a feature map; an edge feature extraction module configured to perform a bidirectional strip convolution operation on the feature map based on the segmentation mask of the feature map to extract original edge features from the feature map; a classification feature generation module, configured to classify the original edge features by contrastive learning to obtain classification feature information; an enhanced image generation module, configured to generate an edge feature enhanced image of the target object based on the classification feature information and the original image; The feature recognition module is configured to perform surface feature recognition on the original image based on a pre-trained image processing model to obtain a feature recognition result; wherein the training process of the image processing model includes the following steps: extracting high-level features and low-level features of the sample image respectively through an initial neural network model; calculating a difference vector based on the high-level features and low-level features; adjusting the weight distribution of the neural network model based on the difference vector, and obtaining the image processing model to generate a first optimized image based on the feature recognition result. A first optimized image generating module is configured to generate a first optimized image based on the feature recognition result; The first integrated image generation module is configured to integrate the edge feature enhanced image and the first optimized image to obtain a first integrated image.

10. The apparatus according to claim 9, further comprising: The second optimized image generation module is configured to optimize the original image through adaptive histogram equalization and multi-scale processing to obtain a second optimized image.

11. The apparatus according to claim 9, further comprising: The second integrated image generation module is configured to fuse the edge feature enhanced image, the first optimized image and the second optimized image based on their respective corresponding attention distributions to obtain a second integrated image.

12. The apparatus according to claim 9, further comprising: a test image acquisition module, configured to acquire a test image obtained by a shooting device shooting the target object in any shooting posture; A contour and size extraction module is configured to extract contour information and size information of the target object based on the test image and the actual shooting posture; a shooting posture determination module, configured to determine a target shooting posture based on the contour information and the size information; The original image acquisition module is configured to acquire the original image obtained by the shooting device shooting the target object in the target shooting posture.

13. The apparatus according to claim 12, further comprising: The three-dimensional modeling module is configured to perform three-dimensional modeling on the target object based on the edge feature enhanced image and the target shooting posture to obtain a three-dimensional model of the target object.

14. The device according to claim 13, wherein The three-dimensional modeling module includes: a modeling feature extraction submodule configured to perform group feature extraction on the edge feature enhanced image based on the target shooting posture to obtain modeling feature information; a point cloud data acquisition submodule, configured to perform point cloud sampling based on the modeling feature information to obtain point cloud data; The three-dimensional modeling submodule is configured to perform mesh reconstruction and texture mapping based on the point cloud data to obtain the three-dimensional model.

15. The device according to claim 14, wherein The point cloud data acquisition submodule is further configured to: Based on the modeling feature information and the shooting posture information, first point cloud data of the target area is collected with a first sampling density, and second point cloud data outside the target area is collected with a second sampling density to obtain the point cloud data; wherein the first sampling density is greater than the second sampling density.

16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the image processing method according to any one of claims 1 to 8. 17 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the image processing method according to claim 1 .

18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the image processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-scale aggregation cloud and cloud shadow identification method, system and device and storage medium

    CN115410081A

  • Photovoltaic cell defect segmentation method based on edge perception

    CN116645338A