Image processing method and device, electronic equipment and storage medium

By introducing a dense connection structure and a local texture feature extraction module into the game costume detection network, the problems of low efficiency and insufficient accuracy in traditional detection methods are solved, and efficient and accurate clipping detection is achieved.

CN121033451APending Publication Date: 2025-11-28NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510949729.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional game costume clipping detection relies on manual inspection, which is inefficient and prone to omissions. Existing target detection algorithms are not accurate enough when dealing with clothing texture details and texture abrupt changes, resulting in a high false alarm rate and making it difficult to meet the needs of high-quality game development.

Method used

An improved target detection network model is adopted, which enhances the transmission of clothing texture features and the ability to extract texture abrupt features in the wearing area by setting a dense connection structure in the backbone network and integrating a local texture feature extraction module in the neck network.

Benefits of technology

It significantly improves the accuracy and efficiency of penetration detection, reduces the workload of manual screening, lowers the false alarm rate, and enhances the model's adaptability to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033451A_ABST
    Figure CN121033451A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring target image data; preprocessing the target image data to obtain enhanced image data; an improved target detection network model is adopted to carry out cross-mode detection on the enhanced image data, and a detection result is generated; according to the improved target detection network model, a feature extraction module with a dense connection structure is arranged in a backbone network of the target detection network model; and accessing a local texture feature extraction module into the neck network of the target detection network model. A dense connection structure is arranged in a backbone network, so that the problem of information loss of clothing texture features in an interlayer transmission process is effectively solved, and the capturing capability of the network on fine texture changes is enhanced; the local texture feature extraction module is accessed to the neck network, and specific texture mutation features of the mold penetrating area are extracted, so that the recognition precision of the model on the mold penetrating area is remarkably improved. Through the synergistic effect of the two technical means, the accuracy of mold penetrating detection is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing, and more specifically, the embodiments of the present invention relate to an image processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] This section is intended to provide background or context for embodiments of the invention as set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.

[0003] In the field of game development, clothing clipping is a common problem affecting game quality. Traditional clipping detection mainly relies on manual visual inspection, which is inefficient and prone to omissions. The object detection algorithms used in related technologies have the following shortcomings when dealing with clipping: first, conventional convolutional neural networks easily lose detailed information about clothing textures during feature transfer; second, existing models struggle to effectively capture the unique texture abrupt changes in clipping areas; and third, the false positive rate is high, requiring extensive manual review. These problems necessitate a significant investment of human resources in the game testing phase, and the detection accuracy is insufficient to meet the demands of high-quality game development. Therefore, there is an urgent need for a method that can quickly and accurately detect clipping areas from massive amounts of game clothing images to improve testing efficiency and reduce labor costs. Summary of the Invention

[0004] In this context, embodiments of the present invention are intended to provide an image processing method, apparatus, electronic device, and storage medium to at least partially address the aforementioned problems existing in the related art.

[0005] In a first aspect of the present invention, an image processing method is provided, comprising: acquiring target image data; performing clipping detection on the target image data using an improved target detection network model, and generating a detection result; the improved target detection network model comprising: setting a feature extraction module with a dense connection structure in the backbone network of the target detection network model; and connecting a local texture feature extraction module in the neck network of the target detection network model.

[0006] In a second aspect of the present invention, an image processing apparatus is provided, comprising: an image acquisition module for acquiring target image data; and a clipping detection module for performing clipping detection on the target image data using an improved target detection network model and generating a detection result; the improved target detection network model comprising: a backbone network module including a feature extraction module having a densely connected structure; and a neck network module including an accessed local texture feature extraction module.

[0007] In a third aspect of the present invention, an electronic device is provided, comprising: a memory storing computer-executable instructions executable by a processor; and a processor for executing the computer-executable instructions to implement the steps of the image processing method described above.

[0008] In a fourth aspect of the present invention, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps of the image processing method described above.

[0009] This disclosure effectively solves the problem of information loss during the inter-layer transmission of clothing texture features by setting a dense connection structure in the backbone network, thus enhancing the network's ability to capture subtle texture changes. Furthermore, by integrating a local texture feature extraction module into the neck network, specifically extracting the unique texture abrupt changes in the clipping region, the model's recognition accuracy for the clipping region is significantly improved. The synergistic effect of these two techniques allows the overall solution to maintain high detection efficiency while significantly improving the accuracy of clipping detection. Attached Figure Description

[0010] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of the invention are illustrated in the drawings by way of example and not limitation, wherein: Figure 1 This is a schematic diagram illustrating the implementation environment of an image processing method provided in this embodiment of the disclosure; Figure 2 A flowchart of an image processing method provided in an embodiment of this disclosure; Figure 3 A flowchart illustrating the improved method for using the improved target detection network model in the image processing method provided in this disclosure embodiment; Figure 4 A schematic diagram of an improved target detection network model structure provided in an embodiment of this disclosure; Figure 5 A schematic diagram of training data including precisely labeled regions provided in an embodiment of this disclosure; Figure 6 A schematic diagram of the detection results output by the improved target detection network model provided in the embodiments of this disclosure; Figure 7 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.

[0011] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0012] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.

[0013] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0014] The accompanying drawings are schematic illustrations of this disclosure and are not necessarily drawn to scale. Some block diagrams shown in the drawings may be functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in hardware modules or integrated circuits, or in networks, processors, or microcontrollers. Implementations can be carried out in various forms and should not be construed as limited to the examples set forth herein. The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough description of embodiments of this disclosure. However, those skilled in the art will recognize that one or more specific details may be omitted when implementing the technical solutions of this disclosure, or other methods, components, apparatuses, steps, etc., may be used to replace one or more specific details.

[0015] Figure 1A system architecture diagram of the operating environment of this exemplary embodiment is shown. This system architecture may include a terminal device 110 and a server 120. The terminal device 110 may be a mobile phone, tablet computer, personal computer, smart wearable device, game console, etc., and has a display function capable of displaying a graphical user interface, which may include the operating system interface or the application interface. An application, such as a game program, is installed on the terminal device 110. The server 120 generally refers to the backend system providing application services in this exemplary embodiment; it may be a single server or a cluster of multiple servers. For example, a game server program is deployed on the server 120 to perform server-side game data processing. The terminal device 110 and the server 120 can be connected via a wired or wireless communication link for data transmission. The method in one exemplary embodiment of this disclosure can be executed by any one or more of the terminal device 110 and the server 120.

[0016] In one implementation, the above method can be implemented and executed based on a cloud interaction system. The cloud interaction system can be the system architecture described above. Various cloud applications, such as cloud gaming, can run under the cloud interaction system. Taking cloud gaming as an example, cloud gaming can be a game mode based on cloud computing. In the cloud gaming operation mode, the game program's execution entity and the game screen presentation entity are separated. The storage and execution of the game's control and interaction methods are completed on the cloud gaming server (such as the aforementioned server 120). The cloud gaming client (such as the aforementioned terminal device 110) is responsible for receiving and sending data and presenting the game screen. For example, the cloud gaming client can be a display device with data transmission capabilities located close to the user, such as a mobile terminal, television, computer, or PDA; while the cloud gaming server in the cloud performs information processing. When playing the game, the user operates the cloud gaming client to send operation commands to the cloud gaming server. The cloud gaming server runs the game according to the operation commands, encodes and compresses the game screen and other data, returns it to the cloud gaming client via the network, and finally, the cloud gaming client decodes and outputs the game screen.

[0017] In one implementation, the method described above can be implemented by the terminal device 110 alone. For example, without deploying the server 120, the terminal device 110 can run the application in a standalone environment to implement the game function and execute the method described above.

[0018] In game development, fashion is a crucial visual resource, and its quality directly impacts the player experience. However, during the design and import of fashion, the "clipping" problem often arises, where clothing appears unnaturally intertwined or overlapping with the character model. Traditional clipping detection methods rely primarily on manual frame-by-frame review, which is not only time-consuming and labor-intensive but also struggles to cover a large number of fashion resources, resulting in low testing efficiency and hindering product iteration speed.

[0019] While some image detection techniques exist for identifying abnormal regions in images, there is a lack of specifically optimized detection models for clipping issues, such as those involving game costumes with unique characteristics (e.g., texture breaks, discontinuous boundaries). Furthermore, due to the varied character poses and complex lighting in game scenes, conventional object detection algorithms suffer from insufficient accuracy and robustness, easily leading to false positives and false negatives.

[0020] Therefore, embodiments of the present invention aim to provide an image processing method to improve the accuracy and efficiency of game costume clipping detection.

[0021] See Figure 2 and Figure 3 According to one embodiment of this disclosure, an image processing method is provided, which may include the following steps: acquiring target image data; performing clipping detection on the target image data using an improved target detection network model, and generating detection results; the improved target detection network model includes: setting a feature extraction module with a dense connection structure in the backbone network of the target detection network model; and connecting a local texture feature extraction module in the neck network of the target detection network model. Thus, by optimizing the network structure and feature extraction method, the accuracy and efficiency of clipping detection are significantly improved, and the workload of manual screening is reduced.

[0022] In a specific application, when it is necessary to detect whether there is a clipping issue in the clothing of a game character, the system first collects multi-view images of the character as target image data. The improved detection model adopts a densely connected structure in the backbone network, which allows feature maps from each convolutional layer to be passed across layers. For example, the fabric texture features extracted in the second layer can directly participate in the edge detection calculation in the fourth layer. At the same time, the local texture module connected to the neck network analyzes the texture breakage features at the clipping boundary (such as the abrupt texture at the junction of the sleeve and the arm), and after fusing it with the global features of the backbone network, it outputs the accurate location of the clipping area.

[0023] The target image data refers to the image input that needs to be detected through clipping. It usually includes a complete image of a game character wearing clothing. This data is used for subsequent processing and analysis by the network model to ensure the accuracy of the detected object. Optionally, the target image data specifically includes clothing images of game characters in different poses and lighting conditions. During the data acquisition phase, the system obtains raw images through the game engine interface to ensure resolution and color fidelity. For dynamic poses, keyframe capture technology is used to record multiple motion frames to capture potential clipping issues in the clothing during movement. In addition, the image data needs to be standardized to RGB format and adjusted to the network input size to ensure efficient execution of subsequent preprocessing steps.

[0024] Optionally, the acquisition of target image data must meet specific quality requirements. For example, the character must maintain a frontal view to cover the main areas of the clothing, or side view data must be obtained through multi-view shooting to identify side occlusion or texture breakage issues. Lighting conditions must cover indoor, outdoor, and shadow scenes to ensure that the model is robust to material reflections and shadow occlusion under different lighting conditions. Optionally, the source of the target image data can be real-time data from the game engine or a pre-stored image dataset. Real-time data can be used for online detection to ensure that clothing clipping issues in the game can be detected and fixed in a timely manner, improving the gaming experience. Pre-stored image datasets can be used for offline detection, improving detection efficiency by batch processing large numbers of images. For real-time data, the generation of enhanced image data needs to be compatible with the real-time input stream, for example, by using a streaming pipeline to instantly enhance the fashion images acquired in real time, reducing latency. Among them, the improved target detection network model is a deep learning architecture optimized for garment pattern detection. Through the dense connection structure of the backbone network and the texture feature module of the neck network, it enhances the ability to identify garment pattern features such as broken clothing textures.

[0025] Dense connection paths are structures that establish cross-layer connections between multiple convolutional layers in the backbone network. Their purpose is to ensure that the output of each layer can be directly passed to all subsequent layers, avoiding information loss during inter-layer transmission.

[0026] The local texture feature extraction module is integrated into the neck network of the target detection network model. It extracts discontinuous texture features at the mold-through boundary in the target image through predefined operators.

[0027] Optionally, the design of densely connected paths can be extended according to actual task requirements. For example, in specific cases, information flow can be enhanced by increasing connection density, or the gradient vanishing problem can be alleviated by introducing residual connections. Furthermore, to improve computational efficiency, the connection strength can be appropriately reduced in some non-critical layers, thereby achieving a balance between performance and resource consumption. This flexibility allows the structure to be applied not only to the mold-detection task in this invention but also to other image understanding domains. Optionally, the local texture feature extraction module also needs to consider real-time performance and hardware compatibility during design. On the one hand, to avoid excessive computation leading to inference delays, the number of operators or the range of parameters can be reasonably limited; on the other hand, to adapt to the hardware conditions of different platforms, the internal structure of the module should adopt a general and efficient algorithm framework as much as possible to achieve efficient operation on various devices. This design approach not only improves the system's practicality but also provides a good foundation for subsequent functional expansion.

[0029] In an optional implementation, a feature extraction module with a densely connected structure is set in the backbone network of the object detection network model. This includes: establishing densely connected paths across multiple convolutional layers in the backbone network, so that the output of each layer can be directly passed to all subsequent layers; setting feature fusion units at the connection points of each layer, combining feature information from different levels through learnable weights; and setting a channel attention mechanism after feature fusion to adaptively adjust the contribution of each feature channel. In this way, through the densely connected structure and feature fusion mechanism, the transmission effect of clothing texture features between network layers is effectively enhanced, improving the model's ability to capture subtle texture changes.

[0030] For example, see Figure 4 In image processing, when detecting clipping regions in game costume images, the backbone network adopts a C3K2 structure with a dense connection mechanism. The output of each layer is not only passed to the next layer but also connected to all subsequent layers, achieving multi-level feature sharing. At each connection point, a feature fusion unit is deployed, which uses learnable parameters to weight and combine feature maps from different layers to form richer feature representations. Subsequently, the system applies a channel attention mechanism to dynamically adjust the weights of each feature channel according to the current task requirements, enhancing key features and suppressing non-key features. This process ensures that the model can capture subtle texture changes at different scales, effectively distinguishing between normal boundaries and clipping boundaries.

[0031] Among them, dense connection paths with cross-layer connections are a structural design that realizes direct connections between multiple layers in a neural network. Its role is to ensure that the output of each layer is not only passed to the next layer, but also directly passed to all subsequent layers, thereby reducing information loss and enhancing feature transfer efficiency.

[0032] Optionally, the dense connection path construction method is based on a modification of the C3K2 module of the standard YOLO network. By introducing connection channels between each convolutional layer, the current layer not only receives the output of its immediate predecessor layer, but also directly accesses the feature maps of all previous layers. This design ensures that the texture details extracted in shallow layers are fully preserved and passed to deeper layers. This structure allows each layer to access the output of all previous layers, thus retaining more original input information.

[0033] Optionally, cross-layer connections can be implemented using tensor concatenation operations within the neural network framework. For example, in the PyTorch framework, assuming the output of layer n is Tensor_n and the output of layer (n+1) is Tensor_n+1, then Tensor_n is concatenated with Tensor_n+1 using the `cat()` function and used as the input to layer (n+2). Simultaneously, Tensor_n also needs to be passed separately to layers (n+2, n+3, etc.) and all subsequent layers. The specific connection paths can be designed as follows: the output of layer 1 goes to layers 2, 3, and 4; the output of layer 2 goes to layers 3 and 4; and so on, forming a densely connected network structure. Furthermore, to avoid computational redundancy, a connection weight decay parameter can be set, applying a progressive weight decay coefficient to connections between layers that are far apart. For example, the connection weight between layers 1 and 4 could be 0.8 times its original value.

[0034] Furthermore, to ensure effective feature fusion across different layers, each connection path is equipped with a feature fusion unit to dynamically adjust the importance of features at different levels. These fusion units typically consist of learnable weight parameters that can be automatically optimized during training to achieve the best feature combination effect. Specifically, a feature fusion unit is a module set at each layer connection point, whose function is to combine feature information from different levels using learnable weights to achieve more efficient feature integration and utilization.

[0035] Optionally, the core of the feature fusion unit lies in constructing a mechanism that enables features from different levels to be effectively fused together to form a more representative feature representation. Traditional feature fusion methods typically employ simple concatenation or weighted summation, which may not fully utilize the relationships between features at different levels. Therefore, this disclosure introduces the concept of learnable weights into the feature fusion unit. These weight parameters can be automatically optimized during training using a backpropagation algorithm, thereby achieving adaptive selection and weighted combination of features at different levels. For example, before the output of each layer enters the next stage, it undergoes processing with a weight matrix, which dynamically adjusts the weight values ​​of each feature channel according to the needs of the current task. This method not only preserves important feature information but also suppresses irrelevant or redundant information, thereby improving the overall quality of the features.

[0036] Optionally, the channel attention mechanism is a structure used to enhance key feature channels. Its basic principle is to first perform global average pooling on each channel of the feature map to obtain a one-dimensional feature vector. Then, this vector is compressed and expanded through two fully connected layers to generate a set of weight values ​​equal to the number of original feature channels. These weight values ​​reflect the importance of each channel in the current task; higher values ​​indicate greater importance. Finally, these weight values ​​are multiplied back into the original feature map to complete the weighted enhancement of the features. In this disclosure, the channel attention mechanism is deployed after feature fusion, primarily to further enhance feature channels closely related to wear detection. For example, when detecting whether there is a break at the edge of clothing, some channels may mainly contain color contrast information, while others may mainly contain texture continuity information. The channel attention mechanism can automatically emphasize channels that better reflect the wear state according to the needs of the current task, thereby improving the accuracy of the detection results.

[0037] Optionally, in practical applications, combining feature fusion and channel attention mechanisms can significantly improve the model's adaptability to complex scenarios. For example, in scenarios with poor lighting conditions or special clothing materials, conventional feature extraction methods may be interfered with, leading to an increased false positive rate. However, through a densely connected structure, the model can retain more original information at multiple levels; through feature fusion, the model can integrate the advantages of features from different levels; and through channel attention, the model can dynamically adjust the importance distribution of features. The synergistic effect of these three mechanisms enables the model to maintain high detection performance under various complex conditions. Especially when dealing with large-scale datasets, this structural design can effectively reduce computational overhead, improve training and inference efficiency, and meet the dual requirements of real-time performance and accuracy in practical applications.

[0038] In an optional implementation, integrating a local texture feature extraction module into the neck network of the target detection network model includes: extracting texture features at the clipping boundary of the target image through local clipping texture features; and fusing the extracted texture features at the clipping boundary with other features of the neck network. In this way, by selectively extracting texture breakage features of the clipping region and fusing them with the overall features, the accuracy of the model in identifying the clipping region can be improved.

[0039] For example, see [link to previous article] Figure 4 The system adds a local feature extraction module based on Gabor and LBP operators to the neck part of the YOLO structure. This module specifically enhances the boundary texture anomalies of the clipping region and fuses them with the global features passed from the backbone network. After training, the model can more accurately distinguish the differences between normal human body edges and clipping regions, thereby significantly reducing the false positive rate and improving the recall rate.

[0040] Among them, the local texture feature extraction module is a component used to extract special texture features at the clipping boundary from the target image. Its role is to enhance the ability to distinguish clipping regions in the neck network.

[0041] Optionally, a local texture feature extraction module can be introduced as an additional feature extraction unit within the neck structure of the neural network. The core design of this module is to capture texture abrupt changes unique to the clipping region, such as discontinuities, breaks, or overlaps at the junction of clothing and body. These texture features are typically not effectively extracted through conventional convolution operations, thus requiring supplementary operators. For example, Gabor filters can be used to extract highly directional edge features, while LBP (Local Binary Pattern) is suitable for capturing subtle texture changes. These operators can generate high-dimensional feature vectors that reflect potential clipping cues in the image.

[0042] Optionally, to improve the performance of the local texture feature extraction module, a multi-scale analysis strategy can be adopted. Since images at different resolutions may exhibit different texture characteristics, multiple Gabor and LBP operators at different scales are designed to extract detailed small-scale features and well-defined medium-to-large-scale features, respectively. Subsequently, these multi-scale features are combined into a unified feature representation through feature concatenation or weighted fusion. This approach not only enhances the model's adaptability to complex scenes but also avoids information loss caused by a single scale. Furthermore, in practical deployment, operator parameters (such as the number of directions and scale range) can be adjusted to adapt to different types of game characters and clothing styles, improving versatility and robustness.

[0043] Optionally, feature fusion refers to merging feature maps from multiple different sources into a unified representation for use by subsequent processing modules. In this disclosure, feature fusion includes not only information from the local clipping texture feature extraction module but also multi-scale feature maps from the neck network. These feature maps typically contain semantic information at different levels, such as edges, shapes, and colors. By integrating this information with clipping-specific texture features, the system can achieve stronger semantic expressive power while maintaining high-resolution details. The specific method of feature fusion can be chosen according to actual needs, such as using a simple concatenation operation or dynamically assigning importance weights to each feature through an attention mechanism. Among them, the channel attention mechanism is a commonly used method, which adaptively emphasizes features that are important for clipping discrimination by weighting each feature channel. For example, in some cases, it may be necessary to emphasize the importance of local clipping texture features, while in other cases, it may be necessary to prioritize overall structural information. This method can significantly improve the accuracy and effectiveness of feature representation while keeping the model lightweight. In addition, cross-layer connection technology can be combined to allow local texture features to be directly passed to subsequent classification and regression layers, thereby enhancing the overall performance of the model.

[0044] In an optional implementation, the improved object detection network is trained as follows: a baseline image set containing clipping regions is acquired and initial labeled data containing pre-labeled regions is generated; based on the pre-labeled regions, a segmentation algorithm is used to generate basic training data containing precisely labeled regions; the basic training data is preprocessed to obtain augmented training data; and the improved object detection network is trained using the augmented training data. In this way, through multi-stage data processing and augmented training strategies, the model's generalization ability and detection accuracy are effectively improved, while reducing the workload of manual labeling.

[0045] In a specific application, the system first collects a large number of character images from a game engine as a baseline image set. These images include fashion effects for different characters, poses, and lighting conditions. Then, the clipping regions in some images are manually labeled with rectangular boxes, forming initial annotation data. Next, a segmentation algorithm (such as SAM) is used to expand the rectangular boxes into pixel-level mask annotations to more accurately describe the boundary contours of the clipping regions. Based on this, the system enhances these labeled data using two strategies: basic image enhancement and advanced hybrid enhancement. These strategies include operations such as image flipping, rotation, affine transformation, mixup, and cutmix, generating diverse training samples. Finally, the enhanced training data is input into an improved YOLO network for training, enabling the model to better recognize and distinguish clipping phenomena in various complex scenes.

[0046] The baseline image set refers to a set of original image data, with or without clipping regions, used to train the model. Its purpose is to provide sufficiently diverse samples so that the model can learn the behavior of clipping regions under different conditions.

[0047] Optionally, the composition of the benchmark image set can vary depending on the specific application scenario. For example, during game development, the system can extract images from multiple versions of game resources, covering different styles of clothing design, different character poses, and different ambient lighting conditions, thereby ensuring the model has strong generalization ability. Furthermore, to further improve training efficiency, the benchmark image set can be filtered and classified to remove duplicate or low-quality images, retaining representative samples. Simultaneously, the benchmark image set can be dynamically updated according to training needs; for example, when new fashion resources are added to a new version of the game, they can be promptly included in the training set to adapt to new detection tasks. The quality and diversity of the benchmark image set have a significant impact on model performance; therefore, in practical applications, it should be reasonably constructed and optimized based on specific task requirements.

[0048] Among them, the pre-annotated area refers to the location and range information of the through-line area formed by the preliminary annotation of the reference image. Its function is to provide a basis for subsequent fine annotation.

[0049] Optionally, the generation of pre-annotated regions usually relies on the experience and judgment of manual annotators, who manually draw rectangular boxes of possible clipping areas based on the image content.

[0050] Optionally, the generation of pre-labeled regions can go beyond simple rectangular bounding boxes; it can also include additional information such as the type, severity, and frequency of the clipping region. This information aids in subsequent data processing and analysis, providing more contextual information for model training.

[0051] In this context, the segmentation algorithm generating precise annotations refers to the process of extracting pixel-level contours from the initially labeled rectangular regions using a pre-defined image segmentation model (such as SAM) to generate accurate mask annotation data. Its purpose is to compensate for the coarseness of manual annotations, improve the detail accuracy of training data, and thus enhance the model's ability to recognize the boundaries of clipping regions. For example, see... Figure 5 The clipping area in the image is marked with a rectangle. The mask algorithm automatically identifies the clipping outline based on this rectangle and makes more accurate annotations on the clipping pockets (the light blue part in the rectangle).

[0052] Optionally, the segmentation algorithm can be implemented using a semi-automatic interactive method. For example, the system first inputs the coordinates of the initially labeled rectangle into the SAM model, and the model automatically generates an initial segmentation mask. Subsequently, the system optimizes the mask according to preset rules: if there is a deviation between the mask edge and the texture breakpoint of the overlay area, the segmentation boundary is fine-tuned through a gradient backpropagation mechanism. For example, for areas with abrupt changes in clothing material, the model can strengthen the gradient weights of edge pixels to ensure that the matching degree between the segmentation boundary and the actual overlay contour reaches more than 95%.

[0053] Preprocessing refers to the steps of enhancing and standardizing the basic training data. Through processes such as flipping and adding noise, a diverse set of enhanced samples that retains semantic features is generated. Its purpose is to expand the scale of the training data and enhance the model's adaptability to different lighting, viewpoints, and occluded scenes.

[0054] In an optional implementation, preprocessing the basic training data includes: performing geometric transformations on the basic training data using basic image enhancement algorithms; and performing linear blending or region replacement processing on the basic training data using advanced hybrid enhancement algorithms. This multi-layered image enhancement processing effectively expands the diversity of the training data, significantly improves the model's generalization ability, reduces overfitting, and thus enhances detection accuracy and robustness.

[0055] In a specific application, this preprocessing step can combine multiple enhancement strategies. For example, basic enhancement operations such as rotation, scaling, and translation can be performed simultaneously on the same image to simulate changes in image features from different perspectives. Furthermore, advanced enhancement methods such as mixup and cutmix can be used to linearly combine multiple sample images or replace local regions, generating more diverse training data. These enhancement techniques not only enrich the distribution of the input data but also enhance the model's ability to recognize clipping problems in complex scenes, providing high-quality data support for subsequent training.

[0056] Among them, basic image augmentation algorithms are techniques that expand a dataset by applying geometric transformations to the original image. Their role is to change the spatial properties of the image, enabling the model to learn more diverse visual patterns, thereby enhancing the model's generalization ability.

[0057] Optionally, basic image enhancement algorithms can be implemented in several ways. A common approach is image rotation, which involves rotating the image around its center point at a certain angle to simulate the viewing effect from different perspectives. Another approach is translation, which involves moving the image horizontally or vertically by a certain number of pixels, thereby altering the position of the target within the image. Furthermore, affine transformations, including scaling, cropping, and perspective transformations, can be used to further enhance the image's diversity. These enhancement methods primarily affect the overall structure of the image, rather than local texture details.

[0058] Optionally, in practical applications, basic image enhancement algorithms often incorporate random parameter settings to avoid the enhanced images exhibiting predictable characteristics. For example, the rotation angle, translation magnitude, or scaling ratio is randomly selected for each enhancement to ensure that each enhanced image possesses a certain degree of uniqueness. This approach helps break down potential commonalities in the dataset and prevents the model from over-relying on certain specific feature representations during training. Furthermore, the geometric transformation parameters of the basic image enhancement algorithm need to consider the reasonableness of clothing deformation. The rotation angle is typically controlled within ±30 degrees to avoid excessive distortion that could lead to fabric structural distortion; the flipping operation primarily uses horizontal flipping to maintain the physical integrity of the clothing; and the scaling ratio is set between 0.8 and 1.2 times to ensure that key features are not excessively compressed or stretched.

[0059] Among them, advanced hybrid augmentation algorithms are techniques that generate new samples by linearly combining or replacing regions of multiple images. Their role is to enhance the model's adaptability to complex scenes and promote the model's learning of shared features across samples by fusing semantic information from different images.

[0060] Optionally, advanced fusion enhancement algorithms mainly include two types: mixup and cutmix. The basic idea of ​​mixup is to linearly weight two images according to a certain ratio to generate a new synthetic image. This method allows the model to learn the transition relationships between different images during training, improving its ability to judge blurred boundary regions. On the other hand, cutmix generates a more realistic fusion image by replacing parts of one image with corresponding parts of another. Compared to mixup, cutmix preserves the spatial structure of the image, making it closer to real-world application scenarios and thus exhibiting superior performance in many tasks.

[0061] In an optional implementation, training the improved object detection network with training data includes: training a local texture feature extraction module to extract texture features at the boundaries of labeled regions using preset operators. This targeted training specifically for texture features at clipping boundaries effectively improves the model's ability to capture subtle clipping features, significantly reducing false positives and false negatives.

[0062] For example, during training, a combination of Gabor filters and LBP operators can be used to extract features from the edge locations of the labeled regions. For instance, Gabor filters capture abrupt changes in texture direction by adjusting spatial frequency and orientation parameters; while the LBP operator quantifies the degree of texture breakage at the boundary by statistically analyzing the binary pattern differences in the pixel neighborhood. The feature maps output by both are then weighted and fused, and used as local texture features input to the network model, thereby enhancing sensitivity to clipping regions.

[0063] Among them, the preset operator refers to a set of mathematical functions or algorithms used to extract local texture features in the target image. Its role is to obtain key information from the image that helps to identify the clipping region.

[0064] Optionally, the preset operators include, but are not limited to, Gabor filters and LBP operators, which are used to capture features such as edges, directionality, and local contrast in the image, respectively. Specifically, the Gabor filter bank consists of multiple sets of filters with different directions and scales, which separate the texture components of specific directions and frequencies in the image through convolution operations. For example, when a sudden change in texture direction is detected at the edge of clothing, the response value of the Gabor filter in the corresponding direction will produce a significant difference, thereby marking the potential wear-through boundary. The LBP operator generates a binary encoded feature map by calculating the difference in gray values ​​of each point in the pixel neighborhood. This encoding pattern exhibits a continuous distribution feature at normal human body boundaries, while exhibiting an abrupt distribution feature at wear-through breaks. The output features of the two operators are weighted and integrated through a feature fusion layer, and the fusion weight parameters can be dynamically optimized through a backpropagation algorithm.

[0065] Optionally, the selection and configuration of preset operators need to fully consider the characteristics of game fashion images. Since there are many types of clothing in the game, and each material has a different appearance, a single operator may not cover all situations. Therefore, this solution adopts a multi-operator parallel processing strategy in the local texture feature extraction module, that is, running multiple different types of operators simultaneously within the same module and fusing their outputs. This design not only improves the robustness of the model but also effectively addresses the challenges posed by different types of clothing. For example, some materials may be more suitable for processing with Gabor filters, while others are more suitable for analysis with LBP operators. By continuously adjusting the weights of each operator during training, the model can automatically select the most suitable feature extraction method for the current image.

[0066] Optionally, the training process of the local texture feature extraction module relies on high-quality labeled data. To improve the accuracy and generalization ability of feature extraction, this module is usually trained in conjunction with pre-labeled clipping region information. Specifically, the system uses existing labeled data to determine which regions belong to clipping boundaries and adjusts filtering parameters or operator weights accordingly, making the feature extraction process more focused on truly important regions. In practical applications, this data is usually provided by human annotators with initial bounding boxes, which are then used to generate pixel-level mask annotations through segmentation algorithms. With this high-precision data as support, the module can better learn how to distinguish between normal human body boundaries and abnormal clipping boundaries. During training, Gabor and LBP operators are first used to extract texture features from the labeled regions, and then these features are fused with features from the backbone network and the neck network. Through the backpropagation algorithm, the model can automatically adjust the parameters of each operator, gradually optimizing its ability to recognize clipping regions. At the same time, to avoid overfitting, the training data is also processed by various enhancement methods, such as rotation, scaling, and noise addition, to simulate various changes in real game scenarios. Ultimately, the well-trained local texture feature extraction module can maintain high detection accuracy when faced with clothing of different styles and materials.

[0067] In an optional implementation, the method further includes: using a pre-trained classifier to perform secondary verification on the detection results and outputting the final detection result. In this way, the secondary verification mechanism of the classifier effectively filters out false alarms in the initial detection, significantly improving the accuracy and reliability of the penetration detection.

[0068] For example, see Figure 6In the game costume clipping detection process, an initial detection box is first generated using an improved YOLO network. Then, a pre-trained EfficientNet classifier is used to perform a secondary judgment on each detection box, and only the areas that are confirmed by the classifier as true clipping are retained as the final output.

[0069] Among them, the classifier is a model used to perform secondary verification on the preliminary detection results, and its role is to improve the accuracy and stability of the detection results.

[0070] Optionally, the classifier can employ a lightweight and efficient architecture, such as the EfficientNet model, to balance processing speed and classification performance. This classifier is typically pre-trained on a large-scale labeled dataset and possesses strong generalization capabilities. In this disclosure, the classifier receives candidate detection boxes from the output of the object detection network as input, and determines whether they conform to typical clipping patterns by extracting image content features within these boxes. For example, the system can employ a convolutional neural network structure combined with an attention mechanism or other discriminative feature extraction modules to improve classification performance. Since the classifier is trained on an independent dataset, it can supplement the shortcomings of the detection model from different perspectives, reduce the false positive rate, and enhance the robustness of the overall system.

[0071] Optionally, the deployment method of the classifier can be flexibly adjusted according to the actual application scenario. In scenarios with high real-time requirements, the classifier can be integrated with the detection model in the same computation process to achieve end-to-end automated processing; while in the testing phase where higher accuracy is required, the system can also run the classifier as an independent sub-module to ensure that each judgment is based on complete contextual information. In addition, to adapt to constantly changing image features, the system can also periodically update the classifier parameters or retrain the model to maintain its continuous effectiveness and adaptability.

[0072] Optionally, the output of the secondary verification can form a feedback mechanism with the original detection results. For example, if the classifier continuously removes detection results for a certain type of scene, the system can trigger a retraining request for the detection network, using the removed cases as negative samples to supplement the training dataset of the detection network. This dynamic feedback mechanism enables the detection network and the classifier to form a closed-loop technology for co-evolution, continuously improving the robustness of the overall detection system.

[0073] In an optional implementation, the classifier is trained as follows: sample images containing and not containing clipping regions are acquired; clipping regions are cropped from the sample images containing clipping regions as positive samples; regions are randomly cropped from the sample images not containing clipping regions as negative samples; and the classifier is trained using the positive and negative samples. In this way, by constructing a balanced training dataset and implementing targeted sample processing, the classifier's ability to distinguish clipping regions is effectively improved, and the false positive rate is significantly reduced.

[0074] For example, when training a clipping detection classifier, clipping regions are cropped from labeled clipping images as positive samples, while the same number of regions are randomly cropped from normal images as negative samples. These samples are used to train an efficient and lightweight binary classification model (such as EfficientNet), enabling it to accurately distinguish between clipping and non-clipping regions. Thus, during the inference phase, this classifier is used to perform a secondary verification of the initial detection results, reducing false detections and improving the robustness and reliability of the detection system.

[0075] Positive samples refer to image regions cropped from sample images containing clipping areas and clearly labeled as clipping locations. Their purpose is to train the classifier to identify image features that actually exhibit clipping phenomena.

[0076] Optionally, positive samples can be cropped images of a single local region, or datasets consisting of multiple related clipping regions cropped from the same image. Furthermore, the selection of positive samples must ensure they possess clear clipping boundary features, such as texture breaks or abrupt contour changes. In addition, to enhance the generalization ability of the classifier, positive samples can also include clipping region samples under different lighting conditions, different character designs, and different poses. Moreover, positive samples should maintain consistency with their corresponding negative samples in terms of size, resolution, and color space to avoid introducing unnecessary bias factors. Simultaneously, to prevent overfitting, positive samples should undergo data augmentation processing, such as rotation, scaling, and brightness adjustment, to generate more diverse training data.

[0077] Optionally, the selection criteria for positive samples can be achieved through a combination of manual annotation and automatic segmentation algorithms. For example, manual annotation can first delineate the approximate range of the camouflage region in the original image, and then segmentation algorithms (such as the SAM model) can be used to perform pixel-level precise segmentation of these regions, thereby obtaining high-quality positive sample data. Furthermore, the distribution of positive samples also needs to be reasonably arranged, covering various typical camouflage scenarios while avoiding excessive concentration on a specific type of camouflage, so as not to cause the classifier to perform poorly when facing other types of camouflage. Therefore, in the process of constructing the positive sample dataset, it is also necessary to diversify the sample sources, including different styles of clothing, different material representations, and presentation effects under different lighting conditions. Through the above methods, the classifier can maintain high recognition accuracy and stability when facing complex and varied real-world detection tasks.

[0078] Negative samples refer to image regions randomly cropped from sample images that do not contain clipping regions. Their purpose is to train the classifier to recognize image features of non-clipping regions.

[0079] Optionally, negative samples can be randomly selected normal regions or non-clipping regions selected according to certain rules. Furthermore, the selection of negative samples should ensure that they do not possess the typical characteristics of clipping regions, such as texture breaks or abrupt contour changes. To enhance the classifier's discriminative ability, negative samples should cover as many normal game costume images as possible, including normal clothing for different characters, poses, and lighting conditions. In addition, to prevent the classifier from favoring certain types of normal regions, the distribution of negative samples should be as uniform as possible, avoiding excessive concentration on a small portion of the images. Furthermore, negative samples should maintain consistency with positive samples in terms of image size, resolution, and color space to reduce interference factors caused by data mismatch. Simultaneously, to improve the classifier's generalization ability, negative samples should undergo data augmentation processing, such as flipping, rotating, and brightness adjustment, to generate more diverse training data.

[0080] Optionally, the method for selecting negative samples can be designed based on statistical principles. For example, several regions can be randomly selected as negative samples from each sample image that does not contain clipping areas, ensuring that each image contributes a certain amount of training data. Simultaneously, to prevent some images from being oversampled and causing classifier bias, a certain sampling frequency limit can be set so that the probability of a region being selected in each image is approximately equal. Furthermore, the number of negative samples can be dynamically adjusted according to the complexity of the image. For example, for samples with simpler image content, the sampling density of their negative samples can be appropriately increased, while for samples with more complex image content, the sampling density of their negative samples can be appropriately decreased to maintain the overall balance of the data.

[0081] Optionally, the classifier training process typically includes multiple stages such as feature extraction, model building, and parameter optimization. In the feature extraction stage, the system can employ efficient convolutional neural network structures, such as EfficientNet and ResNet, to extract multi-level feature representations from images. In the model building stage, the system can select an appropriate classification architecture based on task requirements, such as fully connected layers or the Softmax activation function, to achieve the final class determination. In the parameter optimization stage, the system can measure the model's prediction error using the cross-entropy loss function and iteratively update the model parameters using the gradient descent algorithm. Furthermore, the system can introduce regularization techniques, such as L2 regularization and Dropout, to prevent overfitting and improve its performance on unknown data. Through these methods, the system can build an efficient, stable, and easily maintainable classifier, providing reliable technical support for subsequent detection and verification.

[0082] Through the above technical solution, this disclosure achieves rapid and efficient detection of clipping issues in game costumes. By introducing a feature extraction module with a densely connected structure, the transmission effect of clothing texture features between network layers is enhanced, feature loss is avoided, and the model's ability to perceive clothing edges in complex scenes is improved. Simultaneously, by adding a local clipping texture feature extraction module and combining it with operators such as Gabor and LBP to extract features at clipping boundaries, the model's accuracy in identifying clipping regions is further improved. Finally, by training a binary classifier to perform secondary verification of the initial detection results, the false positive rate is effectively reduced, thereby significantly improving the overall detection accuracy and reliability.

[0083] Corresponding to the above method embodiments, this invention provides an image processing apparatus, see [link to previous document]. Figure 7 The device includes: an image acquisition module for acquiring target image data; and a clipping detection module for performing clipping detection on the target image data using an improved target detection network model and generating detection results. The improved target detection network model includes: a backbone network module containing a feature extraction module with a densely connected structure; and a neck network module containing an accessed local texture feature extraction module.

[0084] In an optional implementation, the feature extraction module with a dense connection structure includes: a cross-layer connection unit, used to establish dense connection paths between multiple convolutional layers, so that the output of each layer is directly passed to all subsequent layers; a feature fusion unit, set at the connection points of each layer, used to combine feature information from different layers through learnable weights; and a channel attention unit, used to adaptively adjust the contribution of each feature channel after feature fusion.

[0085] In an optional implementation, the local texture feature extraction module includes: a texture extraction unit for extracting texture features at the mold-through boundary of the target image; and a feature fusion unit for fusing the extracted mold-through boundary texture features with other features of the neck network.

[0086] In an optional implementation, a training module is also included for training an improved object detection network model. The training module includes: a data acquisition unit for acquiring a set of baseline images containing clipping regions and generating initial labeled data; a data labeling unit for generating basic training data containing precisely labeled regions based on the pre-labeled regions using a segmentation algorithm; a data augmentation unit for preprocessing the basic training data to obtain augmented training data; and a network training unit for training the improved object detection network using the augmented training data.

[0087] In an optional implementation, the data augmentation unit includes: a basic augmentation subunit for performing geometric transformations on the training data using a basic image augmentation algorithm; and an advanced augmentation subunit for performing linear blending or region replacement processing on the training data using an advanced hybrid augmentation algorithm.

[0088] In an optional implementation, the network training unit includes a texture feature training subunit, which is used to train the local texture feature extraction module to extract texture features at the boundary of the labeled region using a preset operator through training data.

[0089] In an optional implementation, a verification module is also included, which is used to perform secondary verification on the detection results using a pre-trained classifier and output the final detection results.

[0090] In an optional implementation, a classifier training module is further included for training a classifier. The classifier training module includes: a sample acquisition unit for acquiring sample images containing the camouflage region and sample images not containing the camouflage region; a positive sample generation unit for cropping the camouflage region from the camouflage region sample image as a positive sample; a negative sample generation unit for randomly cropping regions from the non-camouflage region sample image as negative samples; and a classifier training unit for training the classifier using the positive and negative samples.

[0091] The image processing apparatus provided in this disclosure has the same implementation principle and technical effects as the aforementioned method embodiments. For the sake of brevity, any parts not mentioned in the apparatus embodiments can be referred to the corresponding content in the aforementioned method embodiments.

[0092] It should be noted that although several units / modules or sub-units / modules of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0093] This invention also provides an electronic device, such as... Figure 8 As shown, the electronic device includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement any image processing method of the present disclosure embodiments. For specific implementation methods and the resulting technical effects, please refer to the method embodiments, which will not be repeated here.

[0094] Figure 8 This is a schematic diagram of the structure of an electronic device. The electronic device 1100 includes a processor 1101 with one or more processing cores, a memory 1102 with one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 and the memory 1102 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0095] The processor 1101 is the control center of the electronic device 1100. It connects various parts of the electronic device 1100 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 1102, and calling data stored in the memory 1102, it executes various functions of the electronic device 1100 and processes data, thereby performing overall monitoring of the electronic device 1100.

[0096] Optionally, the electronic device 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. The processor 1101 is electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107. Those skilled in the art will understand that... Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0097] Exemplary media This invention also provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute any image processing method of this disclosure embodiment when run by a processor. For specific implementation methods and the resulting technical effects, please refer to the method embodiments, which will not be repeated here.

[0098] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0099] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0100] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, The method includes: Acquire target image data; An improved target detection network model is used to perform clipping detection on the target image data, and detection results are generated. The improved target detection network model includes: A feature extraction module with a dense connection structure is set in the backbone network of the target detection network model; A local texture feature extraction module is integrated into the neck network of the target detection network model.

2. The method according to claim 1, characterized in that, The feature extraction module with a dense connection structure set in the backbone network of the target detection network model includes: Dense connection paths are established between multiple convolutional layers of the backbone network to enable the output of each layer to be directly passed to all subsequent layers. Feature fusion units are set at the connection points of each layer, and feature information from different layers is combined through learnable weights. After feature fusion, a channel attention mechanism is set up to adaptively adjust the contribution of each feature channel.

3. The method according to claim 1, characterized in that, The step of integrating a local texture feature extraction module into the neck network of the target detection network model includes: The texture features at the clipping boundary of the target image are extracted using the local clipping texture features; The extracted texture features at the mold boundary are fused with other features of the neck network.

4. The method according to claim 1, characterized in that, The improved object detection network is trained in the following manner: Acquire a baseline image set containing the through-modeling region and generate initial annotation data containing the pre-annotated region; Based on the pre-labeled regions, a segmentation algorithm is used to generate basic training data containing precisely labeled regions; The basic training data is preprocessed to obtain enhanced training data; The improved target detection network is trained using the enhanced training data.

5. The method according to claim 4, characterized in that, The preprocessing of the basic training data includes: Geometric transformations are performed on the basic training data using a basic image enhancement algorithm. The basic training data is linearly mixed or region-replaced using an advanced hybrid enhancement algorithm.

6. The method according to claim 4, characterized in that, Training the improved object detection network using the training data includes: The local texture feature extraction module is trained using the training data to extract texture features at the boundary of the labeled region using a preset operator.

7. The method according to claim 1, characterized in that, The method further includes: The detection results are then validated a second time using a pre-trained classifier, and the final detection result is output.

8. The method according to claim 7, characterized in that, The classifier is trained in the following manner: Obtain sample images containing the clipping region and sample images not containing the clipping region; The clipping region is cropped from the sample image containing the clipping region and used as a positive sample; A region is randomly cropped from the sample image that does not contain the mold-through region and used as a negative sample; The classifier is trained using the positive and negative samples.

9. An image processing apparatus, characterized in that, The device includes: The image acquisition module is used to acquire target image data; The clipping detection module is used to perform clipping detection on the target image data using an improved target detection network model and generate detection results. The improved target detection network model includes: The backbone network module includes a feature extraction module with a densely connected structure. The neck network module includes an access local texture feature extraction module.

10. An electronic device, characterized in that, include: Memory stores computer-executable instructions that can be executed by a processor; A processor for executing the computer-executable instructions to implement the method as claimed in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.