A vamp color difference detection method and system based on twin feature coding
By using twin feature coding technology, the problems of low efficiency and poor objectivity in shoe upper color difference detection are solved, and high-precision, interpretable color difference detection is achieved in complex environments, making it suitable for industrial production lines.
Patent Information
- Application Number
- CN202610526763.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies for detecting color difference in shoe uppers suffer from problems such as low detection efficiency, lack of objectivity in results, susceptibility to interference from lighting and materials, and difficulty in quantification and localization. Furthermore, deep learning methods are not adaptable to complex environments.
A twin feature coding-based approach is adopted, which uses paired image acquisition and preprocessing, a feature coding network with shared weights, multi-scale feature aggregation, and explicit difference modeling to achieve color consistency discrimination and color difference localization of the left and right shoe uppers.
It improves the stability and consistency of test results, enhances the sensitivity to subtle color differences, enables interpretable localization and quantification of color difference areas, and adapts to complex industrial environments.
Smart Images

Figure CN122368211A_ABST
Abstract
Description
Technical Field
[0001] This invention is a method and system for detecting color difference in shoe uppers based on twin feature coding, belonging to the field of computer vision and industrial quality inspection technology. Background Technology
[0002] In the large-scale production and refined quality control system of the footwear industry, the consistency of shoe upper color is a core indicator for measuring the appearance quality of products and reflecting the brand's craftsmanship level. Even if there are slight color differences between the left and right shoes of the same pair that are difficult to distinguish with the naked eye, they will still disrupt the overall visual harmony of the finished shoes and be identified as appearance defects. This will have an irreversible negative impact on the end consumer's willingness to buy, the brand's market reputation and industry competitiveness, and the brand value.
[0003] Currently, color difference detection in the footwear industry still relies primarily on manual visual inspection. This method requires inspectors to compare the color of left and right shoe uppers or shoes from the same batch under a fixed light source. However, the inherent limitations of manual inspection make it unsuitable for the streamlined, high-volume production demands of modern footwear manufacturing: Firstly, it is inefficient; secondly, the results lack objectivity, heavily relying on the inspectors' visual sensitivity and experience, and are easily affected by subjective factors such as fatigue and emotions, leading to inconsistent results among different inspectors and even among the same inspector at different times; thirdly, the results lack quantitative standards, relying only on qualitative descriptions such as "color difference present" or "no color difference," making it impossible to accurately quantify the degree or range of color difference, or precisely pinpoint its location.
[0004] With the gradual penetration of machine vision and artificial intelligence technologies into the field of industrial quality inspection, automated color detection methods based on image analysis are beginning to be applied to the footwear inspection process. Existing technologies mainly fall into two categories: one is color analysis methods targeting single shoe upper images. These methods perform operations such as color space conversion, color histogram statistics, and single visual feature extraction on a single shoe upper image to achieve a quantitative description of the shoe upper's color distribution, and then determine whether the color meets the standard by using a preset threshold. The other is simple image comparison methods, which determine color differences between different shoe upper samples through direct image differencing and pixel-level feature comparison. However, both of these methods have significant technical limitations. Single-image analysis methods do not utilize the paired correspondence between left and right shoe uppers, making them susceptible to imaging errors in a single image, resulting in significant deviations in the judgment results. Simple image comparison methods only operate at the pixel or shallow feature level, failing to explore the deep visual features of the shoe upper and lacking the ability to identify subtle color differences.
[0005] Furthermore, the complex environment of actual shoe production sites further exacerbates the technical difficulty of automated color difference detection, becoming a key factor restricting the application of existing detection methods. On the one hand, the lighting conditions at the production site are uncontrollable. Variations in the intensity of natural light, aging and decay of production lamps, and differences in the layout of light sources at different workstations can all lead to inconsistencies in brightness and color temperature during shoe upper image acquisition, causing the same shoe upper to present different color visual effects at different times and workstations. On the other hand, the diversity of shoe upper materials brings complex imaging interference. Different materials such as leather, fabric, synthetic leather, and mesh have significant differences in reflective properties, texture structure, and color adsorption capacity, which can easily form highlight areas, shadow patches, or texture pseudo-differences in images. These non-essential visual differences caused by lighting and materials can seriously interfere with detection methods based on single features or shallow comparisons, leading to a large number of false detections and missed detections, failing to meet the quality control requirements of industrial production.
[0006] In recent years, deep learning technology has made breakthroughs in computer vision fields such as object detection, image classification, and industrial defect detection, providing new technical ideas for the automated detection of color difference in shoe uppers. However, directly applying it to shoe upper color difference detection still faces many technical challenges: First, the traditional single-input deep learning network structure cannot fully explore and utilize the inherent spatial correspondence and color association between the left and right shoe uppers, making it difficult to achieve accurate consistency discrimination of paired samples. Second, existing deep learning detection methods mostly focus on binary classification results of "with color difference / no color difference," lacking explicit modeling and visualization of color difference regions, resulting in poor interpretability of detection results and hindering quality inspectors from reviewing the results and tracing the source of color difference problems. Third, existing models are not adaptable enough to industrial environments, lacking targeted feature design for actual scenarios such as changes in lighting and material differences, making it difficult to balance robustness and detection sensitivity, leading to a significant decrease in detection accuracy in complex production environments. Summary of the Invention
[0007] To address the problems in the existing technology, this invention provides a method and system for detecting color difference in shoe uppers based on twin feature encoding.
[0008] The technical solution adopted by this invention to solve its technical problem is: a method for detecting color difference in shoe uppers based on twin feature coding, comprising the following steps: Step S1: Paired image acquisition and synchronization control step. Acquire images of the uppers of the left and right shoes of the same pair of shoes, and ensure consistency of the left and right shoe upper images in terms of time and imaging conditions through synchronous triggering, unified exposure parameters and unified light source control. Step S2, the shoe upper region extraction and paired consistent cropping step, detects or segments the shoe body region in the left and right shoe upper images, extracts the shoe upper region, and performs paired consistent cropping and scale normalization processing on the left and right shoe upper regions according to a unified rule. Step S3: Paired consistent geometric alignment step. Based on the shoe upper contour features, key point information or geometric center information, paired consistent geometric alignment processing is performed on the left and right shoe upper images to reduce the impact of differences in posture and placement angle on subsequent color difference analysis. Step S4: Pairwise consistent color normalization processing step. Color normalization, white balance correction or color constancy processing are performed simultaneously on the geometrically aligned left and right shoe images to reduce the interference of illumination changes and imaging condition differences on the color discrimination results. Step S5, Paired Feature Encoding and Local Feature Extraction: The processed left and right shoe upper images are input into a paired feature encoding network with shared weights to extract local texture features, edge structure features, and subtle color change features of the left and right shoe uppers. Step S6, Global Consistency Modeling and Multi-Scale Feature Aggregation, extracts multi-scale feature representations from different levels of the pairwise feature encoding network, and models the overall color distribution relationship of the left and right shoe uppers through the global consistency modeling mechanism to obtain aggregated features that integrate local details and global context information. Step S7, Explicit Difference Modeling and Difference Response Feature Generation: Based on the aggregated features of the left and right shoe uppers, explicit difference modeling is performed through feature difference, correlation operation or similarity measurement to generate difference response features that reflect the spatial distribution of color differences between the left and right shoe uppers. Step S8, Color Difference Judgment and Difference Region Location Output Step: Based on the difference response features, output the judgment result of whether there is a color difference between the left and right shoe surfaces, and generate the difference region location result to indicate the location of the color difference on the shoe surface.
[0009] Further, step S1 includes the following sub-steps: Step S11: Simultaneously image the uppers of the left and right shoes of the same pair of shoes using at least two synchronously triggered imaging devices to obtain the original images of the left and right shoe uppers. Step S12: During the image acquisition process, apply a uniform exposure time, gain parameter, and light source brightness and color temperature control strategy to the imaging process of the left and right shoe surfaces to ensure consistent imaging conditions. Step S13: Bind the acquired left and right shoe upper images in pairs according to the timestamp and spatial correspondence to form a one-to-one pair of shoe upper image input samples.
[0010] Further, step S2 includes the following sub-steps: Step S21: Use an object detection model or semantic segmentation model to detect or segment the shoe body region in the left and right shoe upper images to obtain candidate shoe upper regions; Step S22: The candidate area of the shoe upper is trimmed, and the size of the left and right shoe upper areas is normalized according to a preset scale; Step S23: Based on the spatial correspondence between the left and right shoe uppers in paired samples, perform paired consistent cropping on the cropped shoe upper area to ensure the consistency of the spatial structure of the left and right shoe upper areas.
[0011] Further, step S3 includes the following sub-steps: Step S31: Calculate the spatial transformation parameters between the left and right shoe upper images based on the shoe upper contour features, key point information, or geometric center information; Step S32: Based on the spatial transformation parameters, perform paired geometric alignment processing on the left and right shoe images to reduce spatial offset caused by differences in shooting angle and placement posture.
[0012] Further, step S4 includes the following sub-steps: Step S41: Simultaneously perform white balance correction, brightness normalization, or color constancy processing on the geometrically aligned left and right shoe images; Step S42: Map the left and right shoe images in a unified color space to reduce the impact of lighting changes and imaging condition differences on the color consistency judgment results.
[0013] Further, step S5 includes the following sub-steps: Step S51: Input the left and right shoe images processed by steps A to D into a pairwise feature encoding network with shared weights, so that the left and right shoe images are mapped to the same feature representation space. Step S52: Extract local texture features, edge structure features, and subtle color change features of the shoe upper through hierarchical convolution feature extraction structure; Step S53: Enhance the expression of local features related to color differences through cross-channel or cross-space feature enhancement mechanisms to improve the detection sensitivity of minute color differences.
[0014] Further, step S6 includes the following sub-steps: Step S61: Extract multi-scale feature representations corresponding to different receptive field scales from different network layers of the pairwise feature coding network; Step S62: Aggregate the multi-scale features by feature concatenation, weighted fusion, or fusion based on learnable parameters to form aggregated features that fuse local detail information and global context information; Step S63: Based on the aggregated features, model the overall color distribution consistency of the left and right shoe uppers to suppress the interference of local noise on the color difference discrimination results.
[0015] Further, step S7 includes the following sub-steps: Step S71: Based on the aggregated features of the left and right shoe uppers, construct explicit difference features between the left and right shoe uppers through at least one of the following methods: channel-wise difference, feature correlation operation, or similarity measurement. Step S72: Map the difference features in the spatial dimension to generate a difference response map that reflects the spatial distribution of the color difference between the left and right shoe uppers; Step S73: Weight the difference response map using a channel attention mechanism or a spatial weight adjustment mechanism to enhance the response of the real color difference region and suppress false differences caused by changes in illumination or material differences.
[0016] Further, step S8 includes the following sub-steps: Step S81: Perform feature compression and global convergence processing on the difference response features to form discriminative features for color difference discrimination; Step S82: Based on the discriminative features, output the discriminative result of whether there is a color difference between the left and right shoe uppers through the classifier; Step S83: Generate a difference heatmap based on the difference response map to indicate the spatial location of the color difference in the shoe upper.
[0017] A shoe upper color difference detection system based on twin feature coding includes an image acquisition module, a pairwise consistency preprocessing module, a pairwise feature coding module, a multi-scale feature aggregation module, an explicit difference modeling module, and a color difference discrimination and localization module. The modules work together to perform the above detection method. The image acquisition module is used to acquire images of the left and right shoe uppers of the same pair of shoes under the same time and imaging conditions. The paired consistency preprocessing module is used to perform shoe upper region extraction and paired consistency cropping, paired consistency geometric alignment and paired consistency color normalization on the acquired image. The paired feature encoding module is used to input the processed left and right shoe upper images into a paired feature encoding network with shared weights to extract local texture features, edge structure features, and subtle color change features of the left and right shoe uppers. The multi-scale feature aggregation module is used to extract multi-scale feature representations from different levels of the pairwise feature coding network and to model the overall color distribution relationship of the left and right shoe uppers to obtain aggregated features that integrate local details and global context information. The explicit difference modeling module is used to perform explicit difference modeling based on the aggregation features of the left and right shoe uppers, and generate difference response features that reflect the spatial distribution of color differences between the left and right shoe uppers. The color difference discrimination and positioning module is used to output a discrimination result on whether there is a color difference between the left and right shoe uppers based on the difference response features, and to generate a difference area positioning result to indicate the location of the color difference on the shoe upper.
[0018] An electronic device includes a processor and a memory, the memory storing a computer program that, when executed by the processor, causes the processor to perform the methods described above.
[0019] A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method described above.
[0020] The beneficial effects of this invention are: 1. This invention transforms the problem of shoe upper color difference detection into a problem of pairwise consistency discrimination by using a paired consistent image acquisition and preprocessing mechanism. This effectively reduces the impact of non-essential factors such as changes in illumination, shooting angle deviation, and material reflection on the detection results, thereby improving the stability and consistency of the detection results.
[0021] 2. This invention constructs a twin feature encoding structure with shared weights, enabling the left and right shoe images to be compared in the same feature space. This enhances the model's sensitivity to subtle color differences while avoiding the accumulation of biases caused by single-image discrimination methods.
[0022] 3. This invention adopts a combination of multi-scale feature collaborative modeling and explicit difference modeling, which can effectively distinguish between real color differences and pseudo differences caused by shadows, highlights or material textures while maintaining high sensitivity to local subtle color differences.
[0023] 4. By generating difference response maps and difference heat maps, this invention enables intuitive positioning of color difference areas on the shoe upper, significantly improving the interpretability of test results and facilitating verification and quality traceability by industrial quality inspectors. Attached Figure Description
[0024] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart of a method and system for detecting color difference in shoe uppers based on twin feature encoding, according to the present invention. Detailed Implementation
[0025] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0026] Example 1: A method for detecting color difference in shoe uppers based on twin feature coding This embodiment uses the consistency detection of left and right shoe uppers in an industrial production line as an application scenario to illustrate the overall process of the shoe upper color difference detection method of the present invention.
[0027] S01, Image Acquisition and Paired Consistency Preprocessing; In this embodiment, industrial cameras and standardized light sources are arranged on a conveyor belt or fixed inspection station to simultaneously image the uppers of the left and right shoes of the same pair. Preferably, the industrial camera uses a fixed focal length and fixed exposure parameters to reduce imaging differences.
[0028] After image acquisition, the system inputs the left and right shoe images as paired samples and performs paired preprocessing operations on them. Specifically, firstly, the system uses an object detection model or semantic segmentation model to detect or segment the shoe area in the image to obtain candidate areas for the shoe upper; then, the shoe upper areas are cropped and scaled; next, based on the shoe upper contour features or key point information, weak alignment processing is performed on the left and right shoe images to reduce spatial offset caused by differences in placement posture; finally, the left and right shoe images are color normalized using white balance or color constancy algorithms to reduce the impact of changes in lighting conditions on color discrimination.
[0029] During the model training phase, paired data augmentation operations, including geometric augmentation and color augmentation, can be applied simultaneously to the left and right shoe images to improve the model's adaptability to changes in the actual industrial environment.
[0030] S02, twin feature hybrid coding; The preprocessed left and right shoe images are input into a shared-weighted twin feature encoder for feature extraction. The twin feature encoder consists of a hierarchical convolutional feature extraction module and a lightweight self-attention feature extraction module. The hierarchical convolutional feature extraction module is used to extract local texture, edge, and color transition features of the shoe upper, while the self-attention feature extraction module is used to model the color consistency relationship across regions of the shoe upper.
[0031] By sharing weights, the left and right shoe images are mapped in the same feature space, providing a consistent feature basis for subsequent difference modeling.
[0032] S03, Multi-scale feature aggregation; Multi-scale features are extracted from different network layers of the Siamese feature encoder and fused using a multi-scale feature aggregation module. By fusing low-level local detail features and high-level global semantic features, the aggregated features can simultaneously reflect local color differences and overall tonal consistency information of the shoe upper.
[0033] S04, Explicit Difference Modeling and Difference Response Map Generation; After obtaining the aggregated features of the left and right shoes, a difference response map between the features of the left and right shoes is constructed by channel-wise correlation operation or similarity measurement. The difference features are then weighted by channel attention mechanism to highlight the feature responses related to the true color difference.
[0034] The generated difference response map can be further converted into a difference heatmap to visually indicate the spatial location where color difference in the shoe upper may occur.
[0035] S05, Color difference judgment and result output; The system processes the difference features input into a classification head and outputs a result indicating whether there is a color difference between the left and right shoes. In practical industrial applications, the system can simultaneously output the color difference determination result and the difference heatmap, and use the detection results for automatic sorting, alarms, or quality traceability.
[0036] Example 2: System Structure and Deployment Method In this embodiment, the shoe upper color difference detection system of the present invention is deployed in an industrial computer or edge computing device. The system includes an image acquisition and preprocessing module, a twin feature encoding module, a multi-scale feature aggregation module, a difference modeling module, and a color difference discrimination module. The modules are connected via a data bus or a high-speed communication interface.
[0037] The system can communicate with industrial cameras, conveyor belt control systems, and production management systems to achieve online real-time detection and automated control of color difference in shoe uppers.
[0038] Through the above embodiments, the present invention can achieve high-precision, stable and interpretable automatic detection of color difference in shoe uppers in complex industrial environments, verifying the effectiveness and feasibility of the method and system of the present invention in practical applications.
[0039] Example 3: Color Difference Detection under Different Lighting Conditions and with Different Shoe Upper Materials In this embodiment, the adaptability and robustness of the shoe upper color difference detection method of the present invention are explained in light of different lighting conditions and differences in shoe upper materials that may exist in the actual shoe manufacturing process.
[0040] Implementation methods under different lighting conditions: In real industrial environments, the acquisition of shoe upper images may be affected by factors such as changes in ambient light, aging of lighting fixtures, or partial obstruction, resulting in differences in image brightness and color temperature. In this embodiment, based on standard constant light source conditions, various different lighting conditions are introduced for shoe upper image acquisition, including different brightness intensities, different color temperatures, and locally uneven lighting conditions.
[0041] To address the above issues, after image acquisition, pairwise consistent preprocessing is used to simultaneously perform color normalization and brightness correction operations on the left and right shoe upper images to reduce the impact of lighting differences. Meanwhile, during the model training phase, pairwise consistent data augmentation is used to simulate various lighting change scenarios, enabling the model to learn to maintain its ability to distinguish the color difference of the real shoe upper under different lighting conditions.
[0042] Experimental results show that, under different lighting conditions, the method of the present invention can still stably distinguish between the false color difference caused by changes in lighting and the true color difference that exists in the shoe surface itself, ensuring the consistency and reliability of the detection results.
[0043] Implementation methods for different shoe upper materials: In the footwear industry, shoe uppers are made of a variety of materials, including but not limited to leather, fabrics, synthetic materials, and their combinations. Different materials exhibit significant differences in reflective properties, texture distribution, and color expression, which can easily interfere with color difference detection.
[0044] In this embodiment, images of left and right shoes with uppers made of various materials are acquired, and features are extracted using a twin feature encoder. A hierarchical convolutional feature extraction module is used to capture local texture and color detail features of shoe uppers made of different materials, while a lightweight self-attention feature extraction module is used to model the overall color consistency relationship across regions of the shoe upper, thereby reducing the interference of material texture differences on color difference discrimination.
[0045] In addition, through explicit difference modeling steps, the color difference between the left and right shoe uppers under the same material conditions is modeled in detail, so that the model can still accurately identify the real color difference area of the shoe upper even when there is material reflection or complex texture changes.
[0046] Through the above embodiments, the present invention demonstrates good adaptability and detection stability under different lighting conditions and different shoe upper materials, further verifying the versatility and practical value of the method and system of the present invention in actual complex industrial application environments.
[0047] The method of this invention has a clear overall structure and controllable computational complexity, making it suitable for deployment in industrial production lines or edge computing devices. It has good engineering application value and promotion prospects.
[0048] Although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for detecting color difference in shoe uppers based on twin feature encoding, characterized in that, include: Step S1: Acquire images of the uppers of the left and right shoes of the same pair, with consistent time and imaging conditions; Step S2: Process the shoe body area in the left and right shoe upper images, extract the shoe upper area, and perform paired consistent cropping and scale normalization on the left and right shoe upper areas according to uniform rules. Step S3: Based on the shoe upper contour features, key point information, or geometric center information, perform paired geometric alignment processing on the left and right shoe upper images; Step S4: Perform color normalization processing simultaneously on the geometrically aligned left and right shoe upper images; Step S5: Input the processed left and right shoe upper images into a pairwise feature encoding network with shared weights to extract local texture features, edge structure features, and subtle color change features of the left and right shoe uppers. Step S6: Extract multi-scale feature representations from different levels of the pairwise feature encoding network, and model the overall color distribution relationship of the left and right shoe uppers to obtain aggregated features that integrate local details and global context information. Step S7: Based on the aggregation features of the left and right shoe uppers, perform explicit difference modeling to generate difference response features that reflect the spatial distribution of color differences between the left and right shoe uppers; Step S8: Based on the difference response features, output the discrimination result of whether there is a color difference between the left and right shoe uppers, and generate the difference area positioning result to indicate the location of the color difference on the shoe upper.
2. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: In step S1, images of the left and right shoe surfaces of the same pair of shoes are acquired by synchronous triggering, unified exposure parameters and unified light source control, with consistent time and imaging conditions. Step S1 includes the following sub-steps: Step S11: Simultaneously image the uppers of the left and right shoes of the same pair of shoes using at least two synchronously triggered imaging devices to obtain the original images of the left and right shoe uppers. Step S12: During the image acquisition process, apply a uniform exposure time, gain parameter, and light source brightness and color temperature control strategy to the imaging process of the left and right shoe surfaces. Step S13: Bind the acquired left and right shoe upper images in pairs according to the timestamp and spatial correspondence to form a one-to-one pair of shoe upper image input samples.
3. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: In step S2, processing the shoe body region in the left and right shoe upper images includes at least one of the following: detecting and segmenting the shoe body region. Step S2 includes the following sub-steps: Step S21: Use an object detection model to detect the shoe body region in the left and right shoe upper images, or use a semantic segmentation model to segment the shoe body region in the left and right shoe upper images to obtain candidate shoe upper regions. Step S22: The candidate area of the shoe upper is trimmed, and the size of the left and right shoe upper areas is normalized according to a preset scale; Step S23: Based on the spatial correspondence between the left and right shoe uppers in paired samples, perform paired consistent cropping on the cropped shoe upper area.
4. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: Step S3 includes the following sub-steps: Step S31: Calculate the spatial transformation parameters between the left and right shoe upper images based on the shoe upper contour features, key point information, or geometric center information; Step S32: Based on the spatial transformation parameters, perform paired geometric alignment processing on the left and right shoe images to reduce spatial offset caused by differences in shooting angle and placement posture.
5. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: Step S4 includes the following sub-steps: Step S41: Perform white balance correction processing simultaneously on the left and right shoe images after geometric alignment. For the target image after white balance correction, select one of brightness normalization processing or color constancy processing to perform the process. Step S42: Map the left and right shoe images in a unified color space.
6. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: Step S5 includes the following sub-steps: Step S51: Input the left and right shoe images processed by steps S1 to S4 into a pairwise feature encoding network with shared weights, so that the left and right shoe images are mapped to the same feature representation space. Step S52: Extract local texture features, edge structure features, and subtle color change features of the shoe upper through hierarchical convolution feature extraction structure; Step S53: Enhance the expression of local features related to color differences by using one of the cross-channel feature enhancement mechanisms or the cross-space feature enhancement mechanisms.
7. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: In step S6, the overall color distribution relationship between the left and right shoe uppers is modeled using a global consistency modeling mechanism. Step S6 includes the following sub-steps: Step S61: Extract multi-scale feature representations corresponding to different receptive field scales from different network layers of the pairwise feature coding network; Step S62: Aggregate the multi-scale features by one of three fusion methods: feature concatenation, weighted fusion, and fusion based on learnable parameters, to form aggregated features that fuse local detail information and global context information. Step S63: Based on the aggregated features, model the overall color distribution consistency of the left and right shoe uppers.
8. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: In step S6, explicit difference modeling is performed using one of the following methods: feature difference, correlation operation, and similarity measurement. Step S6 includes the following sub-steps: Step S61: Based on the aggregated features of the left and right shoe uppers, construct explicit difference features between the left and right shoe uppers through at least one of the following methods: channel-wise difference, feature correlation operation, and similarity measurement. Step S62: Map the difference features in the spatial dimension to generate a difference response map that reflects the spatial distribution of the color difference between the left and right shoe uppers; Step S63: Weight the difference response map using one of the channel attention mechanism or the spatial weight adjustment mechanism.
9. The method for detecting color difference in shoe uppers based on twin feature encoding according to claim 1, characterized in that: Step S7 includes the following sub-steps: Step S71: Perform feature compression and global convergence processing on the difference response features to form discriminative features for color difference discrimination; Step S72: Based on the discriminative features, output the discriminative result of whether there is a color difference between the left and right shoe uppers through the classifier; Step S73: Generate a difference heatmap based on the difference response map to indicate the spatial location of the color difference in the shoe upper.
10. A shoe upper color difference detection system based on twin feature coding, characterized in that, It includes an image acquisition module, a pairwise consistency preprocessing module, a pairwise feature encoding module, a multi-scale feature aggregation module, an explicit difference modeling module, and a color difference discrimination and localization module, and the modules work together to perform the method described in any one of claims 1 to 9; The image acquisition module is used to acquire images of the left and right shoe uppers of the same pair of shoes under the same time and imaging conditions. The paired consistency preprocessing module is used to perform shoe upper region extraction and paired consistency cropping, paired consistency geometric alignment and paired consistency color normalization on the acquired image. The paired feature encoding module is used to input the processed left and right shoe upper images into a paired feature encoding network with shared weights to extract local texture features, edge structure features, and subtle color change features of the left and right shoe uppers. The multi-scale feature aggregation module is used to extract multi-scale feature representations from different levels of the pairwise feature coding network and to model the overall color distribution relationship of the left and right shoe uppers to obtain aggregated features that integrate local details and global context information. The explicit difference modeling module is used to perform explicit difference modeling based on the aggregation features of the left and right shoe uppers, and generate difference response features that reflect the spatial distribution of color differences between the left and right shoe uppers. The color difference discrimination and positioning module is used to output a discrimination result on whether there is a color difference between the left and right shoe uppers based on the difference response features, and to generate a difference area positioning result to indicate the location of the color difference on the shoe upper.