Intelligent identification and compliance management system based on standard dress

By acquiring the frontal and depth images of the target through a cosmetic mirror, constructing an orthogonal geometric benchmark, generating multi-view evidence sets and reconstructing relationships, this technology solves the problems of existing technologies such as the inability to identify small uniform parts, handle human body surface distortion, and lack of optical and physical verification. It achieves high-precision, interpretable, and verifiable dress compliance management, and improves the automation and reliability of work supervision.

CN120833494BActive Publication Date: 2025-11-18CHENGDU KANGTE NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511316655.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-11-18
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing technologies for personnel identification and compliance assessment have limitations, including the inability to accurately identify small uniform parts, difficulty in handling human body surface distortion, and lack of optical and physical verification. However, existing technologies, when applied to personnel dress identification and standardization requirements, cannot effectively identify small uniform parts, handle human body surface distortion, or address the limitations of existing methods for classifying and identifying uniform parts and clothing materials. Furthermore, existing technologies lack effective methods, equipment, processes, materials, and devices for classifying and identifying uniform parts and clothing materials.

Method used

Simultaneously acquire the target's frontal and depth images using a cosmetic mirror, construct an orthogonal geometric benchmark, and generate a frontal projection map, a chest and abdomen cylindrical unfolded map, and an upper arm conical unfolded map. Set slots for candidate region detection, introduce supplementary lighting differential and reflective consistency verification, establish a multi-view evidence group, generate a wearable token sequence, perform relationship reconstruction and two rounds of adjudication, output compliance results, and generate an evidence overlay map and comparison log.

Benefits of technology

This system achieves high-precision, interpretable, and verifiable dress code recognition and compliance management. It simultaneously acquires frontal and depth images of the target through a mirror and performs intelligent recognition of these images. Through orthogonal projection and relationship reconstruction modules, it achieves high-precision, interpretable, and verifiable dress code compliance management, improving the automation and reliability of work supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833494B_ABST
    Figure CN120833494B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent identification, and more particularly to an intelligent identification and compliance management system based on standard dress, which comprises an image acquisition module, an image structured analysis and relationship reconstruction module, and a clause comparison and compliance output module; wherein the image acquisition module acquires a target front view image and a depth image through a figure mirror and constructs an orthogonal geometric reference; the image structured analysis and relationship reconstruction module synchronously outputs an evidence superposition graph and a regional consistency result set under the constraint of the orthogonal geometric reference, with a wearing structure sequence graph connected by up-down coverage and adjacent continuity; and the clause comparison and compliance output module compares the wearing structure sequence graph and wearing token sequence with a preset dress clause comparison table item by item, and finally outputs the result of whether the staff's dress is compliant. The present application realizes high-precision, interpretable and reviewable dress compliance management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent recognition technology, specifically relating to an intelligent recognition and compliance management system based on standardized attire. Background Technology

[0002] With increasing demands for transparency and standardization in public safety, work supervision, and social governance, the compliance of staff attire during their work hours has received growing attention. Standardized attire not only involves personnel image management but also directly relates to the dignity of the work process and public trust in their conduct. To ensure that personnel dress in the workplace meets standards, several technical solutions based on computer vision and intelligent recognition have emerged both domestically and internationally. These solutions can be broadly categorized into three types: general clothing recognition methods based on image target detection; structured recognition methods based on key points and human pose estimation; and marker detection methods based on feature comparison. However, these existing technologies still have significant shortcomings when applied to identifying and assessing personnel attire compliance.

[0003] In the first category of methods, general clothing recognition primarily relies on trained supervised models to classify or detect clothing categories in images. For example, it involves detecting targets such as hats, uniforms, and shoes. This type of technology has some application in commercial scenarios, such as customer clothing analysis in the retail industry. However, when this technology is directly applied to the detection of personnel attire, it encounters significant difficulties. First, the main features of personnel uniforms are often reflected in details such as badges, epaulets, and name tags, and the area of ​​these details in the image is usually very small, making it easy for general object detection networks to miss or falsely detect items. Second, personnel uniforms are not significantly different in overall color and texture from other dark-colored clothing of similar materials, making it difficult for detection methods based on broad clothing categories to accurately distinguish them, thus failing to meet the requirements of precise normalization. Third, in practical deployment, this type of method often relies on a large number of training samples, but it is difficult to comprehensively cover samples corresponding to different positions, seasons, and environments of staff, resulting in a lack of robustness in the recognition results. Summary of the Invention

[0004] The main objective of this invention is to provide an intelligent identification and compliance management system based on standardized attire. It simultaneously acquires the target's frontal and depth images using a grooming mirror. Utilizing foreground extraction, geometric benchmark construction, and surface unfolding methods, it generates a frontal projection image, a chest and abdomen cylindrical unfolded image, and an upper arm conical unfolded image. Slots are set in each image to locate clothing elements such as cap badges, shoulder patches, name tags, and armbands. During candidate region detection, the system combines supplementary lighting differential and reflective consistency verification to effectively distinguish between genuine attire and printed materials. By establishing multi-view evidence groups across images, generating a wearable token sequence, performing relationship reconstruction, and conducting two rounds of adjudication, it ultimately obtains a wearable structure sequence diagram and a set of regional consistency results. After comparing each item with a preset clause comparison table, the system can output a conclusion on whether the staff's attire is compliant and generate a traceable evidence overlay image and a comparison log. This invention effectively overcomes the limitations of existing technologies, such as the inability to accurately identify small uniform parts, difficulty in handling human body surface distortion, and lack of optical and physical verification. It achieves high-precision, interpretable, and verifiable attire compliance management, improving the automation and reliability of work supervision.

[0005] To solve the above problems, the technical solution of the present invention is implemented as follows:

[0006] The intelligent recognition and compliance management system based on standardized attire includes: an image acquisition module, an image structured analysis and relationship reconstruction module, and a clause comparison and compliance output module. The image acquisition module acquires the target's frontal and depth images using a grooming mirror and constructs an orthogonal geometric datum. Under the constraints of the orthogonal geometric datum, the image structured analysis and relationship reconstruction module generates three projection images of the upper body point set: a frontal projection image, a chest and abdomen cylindrical development image, and an upper arm conical development image. Slots are divided on each image according to preset locations, and the source image, relative position, and adjacent slot numbers are recorded. Candidate regions are detected within each slot to obtain candidate regions, and preliminary elimination is completed using supplementary lighting difference and reflective consistency verification. The system establishes a correspondence between candidate regions among the three images to form a multi-view evidence group. For each multi-view evidence group, it generates a token containing region category, source image identifier, slot identifier, and information on adjacent relationships and boundary morphology. The system constructs a wearable token sequence according to the slot order. Using the wearable token sequence as input, it performs relationship reconstruction and conflict adjudication to obtain a wearable structure sequence diagram with slots as nodes and upper and lower coverage and adjacent continuity as lines. It simultaneously outputs an evidence overlay diagram and a region consistency result set. The clause comparison and compliance output module compares the wearable structure sequence diagram and the wearable token sequence with a preset dress code comparison table item by item, and finally outputs the result of whether the staff's dress is compliant. The wearable token sequence and the wearable structure sequence diagram are archived for review.

[0007] Furthermore, the image acquisition module obtains the frontal view image and depth image of the target through the dressing mirror, and obtains the upper body point set by foreground extraction and connected component filtering; the module uses the main direction of the point set to determine the trunk center line, uses the extreme point of the top of the head to determine the vertical reference, and uses the line connecting the extreme points on the outer sides of the two shoulders as the shoulder line reference; the trunk center line, shoulder line reference and vertical reference constitute an orthogonal geometric reference.

[0008] Furthermore, the image acquisition module simultaneously acquires the target frontal view image and depth image through the facial mirror, ensuring that both correspond to the same scene at the same time point. The depth image is then filled with dead pixels and filtered by median to obtain an effective depth image for segmentation. Based on the effective depth image, the depth histogram of the central region of the image is calculated, and the main peak closest to the camera is selected as the target depth range. Threshold segmentation is performed on the entire image within the target depth range to obtain a depth foreground mask. Adaptive threshold segmentation based on brightness and chromaticity is performed on the target frontal view image to obtain an appearance foreground mask. The intersection of the two masks is calculated, and opening and closing operations are performed to remove isolated noise and small holes, resulting in candidate foreground masks. Connected component labeling is performed on the candidate foreground masks, and the bounding rectangle, area, aspect ratio, and depth continuity index of each connected component are calculated. Connected components that meet the following conditions are selected as upper body connected components: the percentage of area to the total image area is within a preset area range, the aspect ratio is within a preset aspect ratio range, and the depth variation is within a preset depth range. The set of all pixels in the upper body connected components in the frontal view image and the effective depth image is defined as the upper body point set.

[0009] Furthermore, the image acquisition module calculates the geometric center and coordinate covariance features of the upper body point set to obtain the principal and secondary directions. From the principal and secondary directions, the one with the smaller angle to the vertical direction of the image is selected as the torso direction. Using the geometric center as the crossing point and the torso direction as the direction, a straight line passing through this crossing point is defined as the torso centerline. The head search zone is determined based on the circumscribed rectangle of the upper body point set, and this head search zone is located in the upper region of the circumscribed rectangle. The upper boundary contour is extracted within the head search zone, and the boundary point with the smallest ordinate and located in the neighborhood of the torso centerline is selected as the top extreme point. Boundary continuity is checked in the local neighborhood around the top extreme point. If the adjacent pixels above are all background and the left and right sides are still foreground within a set distance, then the point is confirmed as the final top extreme point. Otherwise, backtrack upwards along the center line of the torso to the nearest boundary point that meets the conditions and take it as the final extreme point of the head; define a shoulder search band based on the distance between the extreme point of the head and the lower boundary of the circumscribed rectangle of the upper body point set, wherein the shoulder search band is located in the upper half of the circumscribed rectangle and has a set shoulder height; count the foreground width row by row in the shoulder search band and select the row with the largest foreground width as the candidate row of the shoulder line; obtain the leftmost foreground pixel and the rightmost foreground pixel on the candidate row of the shoulder line as the initial values ​​of the extreme points of the two outer shoulders; starting from the initial values ​​of the extreme points of the two outer shoulders, perform short-distance outward expansion search and curvature extreme value retrieval along the outer boundary respectively to obtain two points with the outermost convex boundary and directly adjacent to the background, and define these two points as the extreme points of the two outer shoulders; use the line connecting the extreme points of the two outer shoulders as the shoulder line reference.

[0010] Furthermore, the image structured analysis and relational reconstruction module uses the circumscribed rectangle determined by the upper body point set as the initial cropping range, sets the torso centerline as the vertical reference, and sets the shoulder line reference as the horizontal reference. It calculates the planar projection transformation that remaps the original image plane to the orthogonal geometric reference plane formed by the vertical reference and the shoulder line reference. Subsequently, it completes cropping and size normalization using the line connecting the extreme points on the outer sides of the two shoulders as the horizontal positioning basis and the extreme point on the top of the head and the lower boundary of the circumscribed rectangle as the vertical positioning basis, to obtain the frontal projection image. The chest and abdomen cylindrical unfolding diagram uses the torso centerline as the cylindrical axis and rearranges the points around this axis into a horizontal sequence according to the angle order. Simultaneously, a longitudinal sequence is formed according to the relative height of the vertical reference; the upper arm conical development diagram takes the conical axis near each scapula as a reference, unfolds the upper arm point set along the circumferential direction into a transverse sequence, and forms a longitudinal sequence according to the arm length direction; based on the natural boundaries of the three images, the hat slot, shoulder slot, and chest slot are set sequentially from top to bottom on the orthographic projection diagram; the left chest slot and right chest slot are set around the central area on the chest and abdomen cylindrical development diagram, and side seam slots are set on both sides; the left arm slot and right arm slot are set on the upper arm conical development diagram respectively; each slot records its source image, relative position, and adjacent slot number as the sole basis for relationship determination.

[0011] Furthermore, the image structure analysis and relationship reconstruction module performs candidate region detection for each slot, including: edge extraction and closed contour filtering in the hat slot, retaining areas with concentric layers and specular reflection as candidate areas for the hat badge; line family detection, concentric ring structure and corner point counting in the shoulder slot, retaining quadrilateral areas enclosed by two approximately parallel boundaries as candidate areas for the epaulettes; rectangular border positioning and internal line scanning in the chest and two chest slots, retaining areas with rectangular borders and multiple horizontal short lines as candidate areas for the name badges; closed curved contour search and internal and external comparison verification in the left and right arm slots, retaining areas where the internal and external comparison difference exceeds a set threshold and the texture direction is independent of the fabric texture direction as candidate areas for the arm badges; and simultaneously performing supplementary lighting differential verification and reflective consistency verification on all candidate areas to eliminate false specular highlights that do not conform to the reflective characteristics of three-dimensional entities.

[0012] Furthermore, the image structured analysis and relationship reconstruction module, based on orthogonal geometric datum, maps the candidate region positions in the orthographic projection image to the corresponding thoracic and abdominal cylindrical unfolded image and upper arm conical unfolded image. It searches for similar candidate regions near the mapping point. If candidate regions exist in two or three images and their boundary orientation is consistent with adjacent slots, these candidate regions are merged into a multi-view evidence group of the same region. A unique token is generated for each multi-view evidence group. The token content includes: region category, source slot, adjacent slot relationship, boundary morphology features, and reflection consistency results. In addition, when the candidate region is an armband candidate region, the token content also includes texture independence results. If a candidate region exists only in a single image but its boundary morphology and slot orientation are stable and it does not overlap with adjacent slots, a token is still generated and a single-view mark is added.

[0013] Furthermore, the image structured parsing and relational reconstruction module concatenates the tokens into a wearable token sequence according to the predetermined order of the slots; wherein the predetermined order is as follows: hat slot token, shoulder slot token, left chest slot token, right chest slot token, left arm slot token, and right arm slot token.

[0014] Furthermore, the image structured analysis and relationship reconstruction module takes the wearing token sequence as input and performs two rounds of adjudication: checking whether the relationship between adjacent slots is coherent and reviewing the regional consistency of the multi-view evidence group.

[0015] The intelligent recognition and compliance management system based on standardized attire of this invention has the following beneficial effects: Firstly, the system simultaneously acquires frontal and depth images of the target through a dressing mirror and constructs an orthogonal geometric benchmark. This ensures that images of personnel from different installation angles, postures, and lighting conditions are unified into a stable coordinate system, significantly improving the accuracy and consistency of subsequent detection. Secondly, the system utilizes frontal projection diagrams, chest and abdomen cylindrical unfolded diagrams, and upper arm conical unfolded diagrams...Figure 3 The combined use of these images unfolds the curved surfaces of the chest, abdomen, and upper arms into a two-dimensional plane, effectively solving the problem of irregular shapes of identification objects caused by the curvature of the human body and perspective distortion. This allows standardized clothing elements such as name tags, epaulets, and armbands to be presented with stable boundaries and clear geometric relationships. Furthermore, by setting slots in each image and performing candidate region detection, the system introduces supplementary lighting differential and reflective consistency verification, enabling it to distinguish between genuinely worn three-dimensional identification objects and sticker patterns printed only on the surface of clothing based on physical reflection characteristics, reducing false detections common in traditional methods. The system also establishes candidate region correspondences across images and generates multi-view evidence sets, then uses token sequences and relationship reconstruction for two rounds of adjudication, effectively ensuring the dual reliability of detection results in terms of geometric coherence and optical consistency. Finally, in the clause comparison stage, the system can compare the wearing structure sequence diagram and the wearing token sequence item by item with the preset standard clause table, outputting compliance conclusions including occurrence, positional relationship, and material properties, and generating evidence overlay diagrams and comparison logs, achieving a complete closed loop from automatic identification to compliance auditing. Through the above design, the present invention not only improves the automation and accuracy of on-site supervision of standardized dress, but also provides a transparent, explainable and verifiable technical basis for subsequent work review, audit traceability and management decision-making, overcoming the limitations of existing technologies that cannot effectively identify small parts, lack the ability to unfold curved surfaces, and lack optical and physical verification methods. Attached Figure Description

[0016] Figure 1 A schematic diagram of the system structure of the intelligent recognition and compliance management system based on standardized dress provided in an embodiment of the present invention;

[0017] Figure 2 This is a schematic diagram of the wearable structure sequence provided in an embodiment of the present invention;

[0018] Figure 3 This is a schematic diagram of the texture direction angle distribution curve provided in an embodiment of the present invention. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0020] refer to Figure 1 The intelligent recognition and compliance management system based on standardized dress code includes: an image acquisition module, an image structured analysis and relationship reconstruction module, and a clause comparison and compliance output module.

[0021] The image acquisition module uses a grooming mirror to acquire the target's frontal and depth images and constructs an orthogonal geometric reference. Its principle lies in combining imaging geometry with the stable morphological features of the upper human body. First, temporal and spatial consistency is established at the acquisition end, and then a reproducible geometric reference is completed within the imaging domain. This reference is used for subsequent structured analysis and relationship reconstruction in a smart recognition and compliance management system based on standardized attire. The grooming mirror synchronously acquires the target's frontal and depth images via hardware triggering, allowing the two sensing channels to share the same timestamp and visual axis, thus achieving spatial correspondence at the pixel level. The camera's internal and external parameters are fixed in the storage medium after factory calibration. During acquisition, lens distortion correction and principal point correction are performed on the target's frontal image, while parallax-to-distance mapping and coordinate alignment are performed on the depth image. This ensures that the pixel coordinate systems of the two channels are consistent. This consistency guarantees that the upper body point set subsequently generated by the two channels has a single coordinate meaning, avoiding human contour misalignment caused by temporal displacement.

[0022] In real-world scenarios, depth images contain random dead pixels and small-scale noise. The image acquisition module first performs dead pixel filling and median filtering. Dead pixels typically occur on reflective materials, dark fabrics, or in areas with skewed incident angles, directly causing depth holes and foreground fragmentation. Dead pixel filling restores isolated holes to continuous depth by interpolating within a small neighborhood using effective pixels as weights. Median filtering suppresses impulse noise while preserving the step characteristics of human contour edges, making the foreground boundary closer to the real physical boundary during subsequent thresholding. To distinguish between people and the background, the image acquisition module calculates a depth histogram in the central region of the effective depth image and selects the range corresponding to the main peak closer to the camera side as the target depth range. This is based on the fact that in entrance / exit and guard booth scenes, people are usually located in the foreground activity area, while background walls and environmental components are located further away and form secondary peaks. Selecting the main peak closer to the camera side satisfies two effects simultaneously: first, it utilizes the physical characteristics of depth ranging, which has high accuracy and small quantization error at close range, to obtain a more reliable contour; second, it naturally eliminates multipath echoes caused by distant backgrounds and glass reflections, decoupling the foreground definition from the environment. A depth foreground mask is obtained by thresholding the entire image based on the target depth range, while an appearance foreground mask is generated on the target's frontal image based on brightness and chromaticity contrast. The intersection of the two masks is used to complement sensing defects: the depth foreground mask maintains separability even when clothing colors are close to the background, while the appearance foreground mask provides visual evidence at depth hole locations. The intersection result is then subjected to opening and closing operations to remove isolated noise and holes, resulting in a well-shaped foreground candidate mask.

[0023] Foreground candidate masks are used for connected component labeling, resulting in several candidate body regions. The image acquisition module calculates the bounding rectangle, area, aspect ratio, and depth continuity index for each connected component. Based on the prior knowledge that people in entrance / exit layouts are typically located near the image center and occupy a moderate proportion of the area, connected components with a reasonable area percentage, an aspect ratio close to the shape of the upper human body, smooth depth variations, and a bounding rectangle center falling within the center's neighborhood are selected. The set of all pixels in these components within the target's frontal view image and the effective depth image is defined as the upper body point set. This selection strategy leverages the geometric regularity of the upper human body's stable width-to-height relationship and continuous depth distribution on the observation plane, effectively excluding non-human targets such as small personal items, roadside billboards, or background edges, thus providing clean samples for subsequent geometric benchmark calculations.

[0024] After the upper body point set is determined, the image acquisition module determines the torso centerline based on the statistical shape principal axis. Specifically, it calculates the geometric center and coordinate covariance features of the upper body point set to obtain the principal and secondary directions. When a person is standing or walking slowly, the major axis of the upper body contour is nearly parallel to the ground normal. Therefore, the direction with the smaller angle to the vertical direction of the image is selected from the principal and secondary directions as the torso direction. The torso centerline is defined with the geometric center as the crossing point and this direction as the pointing direction. The principle of this method is to treat the human body as a two-dimensional projection of an approximately rigid body, using the second-order statistics of the point set to obtain a principal axis estimate robust to local occlusion and local noise. Compared to the simple edge peak method, the principal axis estimate is insensitive to local perturbations such as clothing wrinkles and backpack contours, and can remain stable under monitored posture changes.

[0025] To determine the vertical positioning anchor point, the image acquisition module constructs a head search band on the upper part of the bounding rectangle of the upper body point set. Within this search band, the upper boundary contour is extracted, and the boundary point with the smallest ordinate in the neighborhood of the torso centerline is searched as a candidate for the top-head extremum point. The physical meaning of the top-head extremum point is the uppermost boundary of the human body and background in the vertical direction, typically corresponding to the highest point of the hat brim or hair outline. To avoid false detections due to background gaps or light spots, the module performs a boundary continuity check in the neighborhood of the top-head extremum point, confirming that the area above the point is background, while the short distances to the left and right are still foreground, thus ensuring that the point is at the upper end of the true human body contour. This strategy provides a unified starting point for vertical positioning in the orthographic projection image, facilitating size normalization for different individuals and heights.

[0026] The shoulder line reference is determined based on the existence of a locally maximum width position of the upper body in the image along the vertical direction, which typically corresponds to the acromion line. The image acquisition module sets a shoulder search band based on the distance between the extreme point at the top of the head and the lower boundary of the circumscribed rectangle. This shoulder search band is located in the upper half of the circumscribed rectangle and has a fixed shoulder height range. The foreground width is statistically analyzed row by row within the shoulder search band, and the row with the largest foreground width is selected as the candidate shoulder line row. The leftmost and rightmost foreground pixels in this row are then located as the initial values ​​for the extreme points on the outer sides of the two shoulders. Subsequently, a short-distance outward expansion search and curvature extreme value retrieval are performed along the outer boundaries on both sides to find the position with the most convex boundary and directly adjacent to the background as the extreme points on the outer sides of the two shoulders. This processing utilizes the geometric features of the shoulder in a two-dimensional projection: the boundary curvature near the acromion exhibits an extreme value and is convex. Combined with the maximum width criterion, it can recover the horizontal reference that conforms to the bony landmarks of the human body from the interference of capes, bag straps, or chest objects. The shoulder line reference is defined by connecting the extreme points on the outer sides of the two shoulders, achieving alignment with the natural horizontal reference of the human body. After the torso centerline, the extreme point at the top of the head, and the shoulder line reference are determined, the image acquisition module constructs a straight line strictly parallel to the torso centerline, using the extreme point at the top of the head as the crossing point, as a vertical reference. Rigid fine-tuning is then performed on the shoulder line reference to ensure it remains orthogonal to the vertical reference. These three elements together constitute an orthogonal geometric reference. This orthogonal geometric reference is based on the standardization principles of imaging geometry: the monitoring installation angle and slight ground tilt introduce global pitch and roll deviations. An orthogonal geometric reference established through local human body stability features can absorb these deviations during the geometric alignment process. Subsequent slot division and position comparison are scene-independent and device-independent under this reference.

[0027] Based on orthogonal geometric references, the image acquisition module remaps the target's frontal image onto a geometric reference plane with vertical and shoulder line references as coordinate axes, obtaining a frontal projection image. The purpose of this remapping is to eliminate nonlinear scale variations and slight rotational deviations caused by perspective, ensuring that the vertical and horizontal directions are aligned with the human body reference axes. Cropping and size normalization use the line connecting the outermost extreme points of the two shoulders as the horizontal positioning reference, and the extreme point at the top of the head and the lower boundary of the circumscribed rectangle as the vertical positioning reference, ensuring that the frontal projection images obtained for different individuals and at different passing speeds have comparable scale definitions and uniform start and end boundaries. The depth image is cropped using the same mapping, making the target's frontal image, depth image, and frontal projection image spatially identical. Orthogonal geometric datum and orthographic projection provide a unified coordinate system for image structured analysis and relational reconstruction, enabling the chest and abdomen cylindrical unfolding and the upper arm conical unfolding to unfold with the trunk centerline and the extreme points on the outer sides of the shoulders as stable anchor points. The position and number of the slots remain consistent across different devices and locations, thereby directly supporting the consistency and reproducibility of the intelligent recognition and compliance management system based on standardized attire in candidate region detection, multi-view evidence group establishment, and wearable structure sequence diagram generation.

[0028] The image structured analysis and relationship reconstruction module operates under the constraints of orthogonal geometric benchmarks. It transforms the pixel distribution of parts related to standardized attire from viewpoint-dependent to viewpoint-independent part and relationship representations, enabling the same wearable element to be located and compared with consistent coordinates and topology under different poses, camera installation angles, and lighting conditions. This module uses orthographic projection maps, chest and abdomen cylindrical development maps, and upper arm conical development maps as geometric carriers. It defines the search domain and structural constraints through slots and adjacency relationships, performs deterministic screening based on the physical reflection characteristics and boundary morphology of candidate regions, and solidifies image evidence into verifiable serialized data through multi-view evidence groups and tokens. Finally, it reconstructs a wearable structure sequence map and a set of region consistency results, providing stable input for clause comparison in a smart recognition and compliance management system based on standardized attire.

[0029] The orthographic projection is used to establish a planar reference that aligns with the natural vertical and horizontal orientation of the human body. The chest and abdomen cylindrical development is used to unfold the curved surfaces of the torso onto a plane to eliminate circumferential occlusion and perspective scaling. The upper arm conical development is used to flatten the conical surface of the upper arm to separate the sleeve side from the front and back areas of the arm. The cylinder and cone are chosen as the development models because the torso and upper arm approximate cylindrical and conical surfaces respectively in common clothing and standing passage scenarios. This choice allows for a monotonic, continuous, and geometrically reversible mapping of surface points to a two-dimensional plane, avoiding metric distortion caused by free deformation. During development, the torso centerline is used as the cylindrical axis, and the extreme points on the outer sides of the shoulders and the shoulder line reference are used to determine the upper arm conical axis, ensuring a calculable correspondence between the three projections for the same physical part. The advantages of doing this are twofold: firstly, the nonlinear deformation of name badges, shoulder patches, armbands, etc., caused by surface curvature in the original image is transformed into approximately regular boundaries in the unfolded domain, which facilitates stable detection; secondly, any camouflage patches that exist only in a single projection and lack surface continuity will lose consistent evidence when corresponding across images, and will thus be eliminated by subsequent processes.

[0030] Defining slots on three types of maps transforms the search problem into determining the existence and relationships of targets within a small, semantically fixed area. Slots are arranged along the natural partitions of human anatomy and standard clothing, including hat slots, shoulder slots, chest slots, left chest slots, right chest slots, side seam slots, left arm slots, and right arm slots. Each slot records its source map, relative position, and adjacent slot numbers, used to constrain vertical overlap and horizontal coherence within the same map, and to establish one-to-one or one-to-many mapping windows between different maps. The principle behind using slots is to transform open-world search into boundary and texture determination within a priori correct geometric window, significantly reducing the probability of false detections and binding the results to coordinate positions with fixed meanings, facilitating automated clause comparison.

[0031] Candidate region detection primarily relies on deterministic image processing, directly designing rules based on the physical and geometric origins of the wearable elements. In the hat slot, the cap badge typically exhibits a highlight and concentric layers of metal or enamel. Edge extraction and closed contour filtering yield a compact set of contours, and the synchronicity of the highlight response with switching on / off is observed using supplementary lighting difference to confirm it as a three-dimensional solid rather than a printed highlight. In the shoulder slot, the epaulette often consists of two approximately parallel boundaries and corner points. Line family detection and corner point counting output quadrilateral regions that conform to length and shape relationships, and the stability of the boundary lines distinguishes pseudo-lines generated by folds. In the chest and two chest slots, the name badge generally includes a regular border and line engravings. Rectangular border positioning and internal line scanning can form stable candidates within the unfolded domain. If the line maintains intensity before and after supplementary lighting difference while the border highlight increases with supplementary lighting, it indicates a combination of solid engraving and a transparent cover, consistent with the physical reflective characteristics of a solid name badge. In the slots of the left and right arms, the arm patch has a closed curved outline, and its internal texture direction is independent of the fabric texture direction. By searching the closed curved edges and verifying the internal and external comparisons, candidates can be stably output. If the internal texture is periodically repeating and the highlights of the supplementary light difference do not change with the switch, it can be identified as a printed film and subsequently rejected. The introduction of supplementary light difference and reflective consistency is based on the principle of optical reflection: when the incident light is enhanced, the specular component of the metal or glaze increases while the diffuse reflection ratio remains relatively stable. The brightness change of the printed pattern mainly comes from the overall exposure rather than local specular reflection. The response modes of the two in the difference map are different, which can be used as evidence for physical authenticity separation.

[0032] Cross-image correspondence uses orthogonal geometric datums as a bridge, projecting candidate regions from the orthographic projection onto the corresponding angles and heights of the chest and abdomen cylindrical unfolded diagram and the upper arm conical unfolded diagram, and searching for similar candidates in slots near the mapping landing points. The principle of establishing correspondence is that real wearable elements attached to curved surfaces must provide mutually corroborating boundary segments in multiple images under the continuous mapping from 3D curved surfaces to 2D unfolded diagrams; if it only appears in one image and cannot find a consistent boundary in the corresponding unfolded domain, then the 3D consistency of the candidate is insufficient and does not meet the requirement of multi-view common evidence. After completing the correspondence, the candidate sets of the same physical element in different images are combined into a multi-view evidence group, and its source image identifier, slot identifier, adjacency relationship and boundary morphology are recorded with a unified identifier.

[0033] Tokens are used to solidify multi-view evidence sets into structured data, facilitating subsequent serialization and audit review. Each token contains region category, source image identifier, slot identifier and adjacency relationship, boundary morphology information, reflectivity consistency result, and texture independence result in the armband scene. If there is only a single image evidence but the boundary morphology and slot orientation are stable and do not cover adjacent slots, a token annotated with a single-view mark is generated to retain weak evidence. Tokens are connected in a fixed slot order to form a wearable token sequence: hat slot token, shoulder slot token, left chest slot token, right chest slot token, left arm slot token, and right arm slot token. The purpose of serialization is to determine the spatial structure carried by the sequence, eliminate the randomness of parallel detection, and make the adjacency relationship and the vertical coverage relationship have discernible constraints in the sequence position.

[0034] The relationship reconstruction and conflict resolution process takes the wearable token sequence as input. First, it checks the continuity of adjacent slot relationships within the same image. Then, it verifies the regional consistency of the multi-view evidence group across different images. Adjacent slot relationship continuity ensures that the shoulder slot token is positioned above the two chest slot tokens and that their boundaries do not intersect, and that the chest slot and the two chest slots maintain a prescribed order and do not overlap laterally. Regional consistency verification checks whether the boundary orientation of the orthographic projection map is consistent with the boundary orientation of the unfolded domain, and whether the supplementary lighting difference and reflection consistency maintain the same conclusion across the three images. If any check fails, the conflict backup at that location replaces the current token; if the replacement still fails, the slot location is marked as missing, and the reason is recorded in the regional consistency result set. This two-level resolution effectively solves the problems of false boundaries caused by duplicate candidates and occlusion, and the temporary validity of film camouflage in single-view vision, ensuring that the final output wearable structure sequence map contains only geometrically and optically consistent elements. The wearable structure sequence diagram uses slots as nodes and upper / lower coverage and adjacent connections as lines, visually showing the appearance, relative position, and coverage relationship of cap badges, shoulder patches, name tags, and armbands in a unified coordinate system. The regional consistency result set records the evidence source, reflectivity consistency result, and texture independence result for each node. This output directly serves the clause comparison and compliance output module of the intelligent recognition and compliance management system based on standardized attire. Clause comparison can be performed item by item based on slot positions and connection relationships, and the evidence overlay diagram can replay the boundaries and differences of each comparison item on three images, forming a traceable chain for management auditing.

[0035] The clause comparison and compliance output module converts the clothing structure sequence diagram and clothing token sequence into three types of evidence—presence of parts, positional relationship, and material properties—that can be directly compared item by item with the clothing clause comparison table. Under orthogonal geometric datum, it completes clause matching, conflict resolution, and result solidification using unified coordinates. This embodiment first loads the clothing clause comparison table into layers based on season, job position, location, and equipment status. Layered loading is necessary because different layers have different requirements for element occurrence and relative relationships. Determining the applicable layers first reduces the involvement of irrelevant clauses, improves the reliability of matching judgments, and avoids multiple mutually exclusive requirements on the same input. Layer selection is guided by the basic elements of the clothing token sequence. For example, when there are hat slot tokens and shoulder slot tokens, the regular field layer is prioritized; when there are signs of outer garment coverage and the chest slot boundary is obscured, the winter layer is prioritized. The advantage of this selection strategy is that the starting point of the clause comparison is consistent with the on-site clothing status, reducing discrepancies in subsequent position and coverage judgments.

[0036] The clause comparison process employs a combination of sequential and retrospective judgment. Sequential judgment begins with the occurrence clause, using the presence status of each slot node in the wearable structure sequence diagram as direct evidence. The cap slot, shoulder slot, left chest slot, right chest slot, left arm slot, and right arm slot are compared item by item to check if they meet the layering requirements. Prioritizing the occurrence clause is because occurrence is a prerequisite for position and coverage judgments; confirming the existence of elements first reduces invalid calculations of positional relationships in cases of missing elements. Retrospective judgment uses the regional consistency result set as a trigger condition. Nodes with weak evidence or single-view markings in the occurrence judgment undergo cross-image re-verification, using the consistency of boundary orientation and reflection between the three images as reinforcing evidence to ensure stable occurrence conclusions even under complex lighting and slight posture deviations.

[0037] After the occurrence of an item is confirmed, the verification phase of the positional relationship clause begins. This phase involves determining relative height, left-right attribution, adjacency continuity, and vertical coverage based on orthogonal geometric references. Relative height is determined in the coordinate system of the orthographic projection diagram, using the shoulder slot node as the lateral reference line to verify that the hat slot is above it and the two chest slots are below it. Left-right attribution is determined using the torso centerline as the boundary, verifying that the left chest slot and left arm slot are to the left of the torso centerline, and the right chest slot and right arm slot are to the right of the torso centerline. Adjacency continuity is determined based on the connection relationships in the wearing structure sequence diagram, checking whether the continuous arrangement of the shoulder slots and two chest slots in the lateral direction meets the hierarchical order. Vertical coverage is confirmed by the boundary overlap of the evidence overlay diagram, verifying whether the coverage relationship between the name tag boundary and the clothing front boundary is consistent with the clause. The positional relationship clause is placed in the second phase because it relies on the confirmed occurrence and has the advantages of scale independence and viewpoint independence under orthogonal geometric references, enabling the restoration of images of different individuals under different camera installation angles to a unified interpretation coordinate system, reducing errors caused by cross-scene differences.

[0038] The third stage, material properties and authenticity determination, involves reading the consistency of reflectivity for tokens in the hat, shoulder, and chest slots, and the texture independence for tokens in the left and right arm slots. This stage is based on the physical differences between optical reflection and fabric texture: metal or enamel surfaces produce enhanced specular reflection during differential lighting, while fabric backgrounds are predominantly diffuse. These two types of textures are distinguishable in their differential response and spatial distribution of boundary highlights. The internal texture direction of the armband is independent of the fabric texture, while the periodic texture of the printed film is often related to the fabric texture or exhibits a uniform brightness shift during differential lighting. By reading these results and comparing them item by item with the material property requirements in the layered clauses, the final confirmation of authenticity and material can be achieved when both occurrence and positional relationships are satisfied. This stage is placed after positional relationships because material properties often depend on the final boundaries of the candidates. Completing geometric constraints first provides a stable extraction window for material judgment, reducing the probability of misjudgment caused by lighting fluctuations.

[0039] The conflict resolution operates using the same mechanism across all three phases. When inconsistencies exist among multiple pieces of evidence relied upon for the same clause, the clause comparison and compliance output module replaces and backtracks evidence according to its source priority. The priorities, from highest to lowest, are: multi-view evidence groups, strong boundary of the orthographic projection, consistent boundary of the unfolded domain, and single-view marked evidence. This priority is chosen because multi-view evidence groups simultaneously satisfy geometric and optical cross-validation, resulting in the highest reliability; strong boundary of the orthographic projection has a stable outline under frontal human conditions; consistent boundary of the unfolded domain can compensate for boundary gaps caused by slight rotation; and single-view marked evidence is only retained for weak existence and is not used for the final compliance conclusion. If inconsistencies still exist after replacement, the clause comparison and compliance output module marks the corresponding node of that clause as pending review and provides the source of evidence, conflict location, and suggested re-collection conditions in the output, facilitating on-site re-collection and review with minimal cost.

[0040] The output consists of a compliance conclusion, a cause list, an evidence overlay diagram, and a comparison log. The compliance conclusion is generated at the hierarchical and clause / item levels. A compliance result is given if all three types of evidence (occurrence, positional relationship, and material attributes) are matched; a non-compliance reason is given if any type of evidence is not matched. The cause list is designed for management processes, listing each unmatched slot node, unmet relationship type, and corresponding token identifier for quick identification of rectification points. The evidence overlay diagram overlays boundaries, lines, and differential responses onto each clause / item in the frontal projection, chest / abdomen cylindrical development, and upper arm conical development diagrams, forming visual verification materials. The comparison log records the clause comparison table version, hierarchical selection, token number, judgment order, and conflict replacement as immutable items, along with the wearing token sequence and wearing structure sequence. Figure 1 The data is archived to ensure the traceability of conclusions during post-audit and cross-device reuse. Through the above principles and execution process, the clause comparison and compliance output module transforms the wearable structure sequence diagram and wearable token sequence into an interpretable judgment chain oriented towards standard clauses, enabling the intelligent recognition and compliance management system based on standardized attire to achieve consistent clause comparison and compliance output across different devices and scenarios.

[0041] Furthermore, the image structured analysis and relational reconstruction module, after receiving the upper body point set, uses the circumscribed rectangle determined by the upper body point set as the initial cropping range. This aims to limit the processing domain to areas directly related to standardized attire, reducing interference from background textures and edge components, while confining the computational load of subsequent mapping and unfolding within a stable pixel set. The torso centerline is set as the vertical reference, and the shoulder line reference as the horizontal reference. Subsequently, a planar projection transformation is calculated to remap the original image plane to the orthogonal geometric reference plane formed by the vertical and shoulder line references. This process incorporates the overall deflection caused by the monitoring installation angle, slight roll, and tilt into geometric alignment, ensuring that the coordinate axes are naturally aligned with the human body's vertical and horizontal orientation. Subsequent positioning and comparison of the hat, shoulders, chest, and upper arms are performed under a unified interpreted coordinate system. The line connecting the extreme points on the outer sides of the two shoulders is used as the horizontal positioning basis, and the extreme point on the top of the head and the lower boundary of the circumscribed rectangle are used as the vertical positioning basis to complete cropping and size normalization, resulting in a frontal projection image. The reason for choosing the line connecting the extreme points on the outer sides of the two shoulders as the horizontal reference is that the shoulder width fluctuates little in the passage scenario and is not sensitive to slight changes in posture, making it suitable as an anchor point for the horizontal scale; choosing the extreme point at the top of the head and the lower boundary of the circumscribed rectangle as the vertical start and end reference can ensure that the chest and abdomen area is fully included and maintains a comparable longitudinal coverage range between different people, which is convenient for judging the vertical relationship based on relative height in the subsequent clause comparison stage.

[0042] The cylindrical unfolded diagram of the chest and abdomen uses the torso centerline as the cylindrical axis. Points around this axis are rearranged into a horizontal sequence according to their angles, while a vertical sequence is formed according to their relative heights. Under actual wearing conditions, the torso's shape approximates a cylindrical surface. Directly observing the badge or pocket boundaries in the original diagram is easily affected by circumferential curvature and perspective scaling. After unfolding into the cylindrical domain, circumferential compression is transformed into monotonic horizontal coordinates. Elements on the chest are locally approximated as rectangles, resulting in better boundary continuity and facilitating stable detection and cross-time alignment. Organizing points into the horizontal sequence according to their angles while preserving relative height maintains the distinguishability of the front, middle, and rear regions even with slight lateral rotation. For left and right chest elements appearing in the front and middle regions, the unfolded diagram has a consistent metric meaning with the lateral offset from the centerline, facilitating direct matching with pre-defined slots.

[0043] The upper arm conical development diagram uses the conical axis near each shoulder blade as a reference, unfolding the upper arm point set into a horizontal sequence along the circumference and forming a vertical sequence along the arm length. The upper arm shape gradually tapers as the distance from the shoulder joint increases, becoming closer to a conical surface. In the original diagram, the arm patch boundary often appears elliptical or irregular due to slight arm rotation. After unfolding to the conical domain, the arm patch outline tends to be closed and its size is stable. The sleeve seam and the front and back boundaries of the arm form reproducible positioning lines in the horizontal direction, thus fixing the relative relationship between the arm patch and the fabric texture to a two-dimensional plane, which is beneficial for performing texture independence and boundary consistency judgment.

[0044] Setting slots based on natural boundaries in the three projections transforms the open search of the entire map into an existence and relationship determination within a limited window. The frontal projection map sequentially sets hat slots, shoulder slots, and chest slots from top to bottom. This is because this field of view directly corresponds to the "identification appearing on the hat," "structural components appearing on the shoulder," and "name badge appearing on the chest" described in the clauses, and the vertical order is unaffected by camera posture under orthogonal geometric references. The chest and abdomen cylindrical unfolded map sets left and right chest slots around the central area, with side seam slots on both sides. This is because the center of the unfolded area corresponds to the front-middle region of the human body, suitable for holding elements such as name badges and number plates; the side seam positions form stable circumferential coordinate constants after flattening, serving as boundary lines to prevent confusion and avoid misjudging side pockets, zippers, or placket folds as chest elements. The upper arm conical unfolded map sets left and right arm slots, covering the common wearing area from the front to the side of the arm. The vertical range extends from the shoulder line reference to the upper boundary of the forearm, accommodating differences in posture, whether the arm is raised or naturally drooping. By using slots, candidate region detection is reduced from the entire image without boundaries to a semantically fixed window, reducing false triggers and providing clear spatial labels for subsequent relationship determination.

[0045] Each slot records its source image, relative position, and adjacent slot numbers, serving as the sole basis for relationship determination. This aims to maintain consistent topological relationships across images, time periods, and devices. The source image identifier establishes a one-to-one or one-to-many correspondence window between the orthographic projection, the chest and abdomen cylindrical development, and the upper arm conical development, ensuring that evidence of the same wearable element in different images can be aggregated. The relative position is expressed in coordinates based on orthogonal geometric references, representing the area of ​​belonging and offset within each image. This facilitates direct judgment based on relative height, left-right belonging, and top-bottom coverage during the clause comparison phase. Adjacent slot numbers solidify the top-bottom and left-right adjacency relationships, allowing constraints such as "the shoulder slot is higher than both chest slots" and "the left chest slot and right chest slot are in the same longitudinal band" to be directly verified by the machine through numbered connections. Using this set of records as the sole basis for relationship determination facilitates reproducible auditing, maintains isomorphic data structures across different implementation versions, and reduces ambiguity caused by differences in implementation details.

[0046] Furthermore, when the image structure analysis and relationship reconstruction module performs candidate region detection for each slot, it first performs edge extraction and closed contour screening in the hat slot. Positions with stable gradient changes and light-dark transitions are used as edge sources. A set of closed contours is formed through connectivity reconstruction. Within the closed contours, the spatial distribution of concentric layers and specular reflection is calculated. Regions with multi-layered annular transitions and local specular enhancement when supplementary lighting is activated are preferentially retained as candidate areas for the hat badge. The rationale behind this approach is that metal or enamel materials typically possess both concentric layers and specular reflection signals simultaneously. These signals overlap spatially, distinguishing them from the diffuse reflection and single-layer bright spots of fabrics or printed patterns, thus maintaining stable detection results even under slight changes in posture such as tilting the head up or down.

[0047] In the shoulder slot, line family detection is performed to identify concentric ring structures and count corner points. First, line family detection is used to obtain two boundaries with consistent directions and stable spacing. Then, corner point counts are performed at the ends of these two boundaries, and the quadrilateral region enclosed by the two approximately parallel boundaries is retained as the candidate area for the epaulette. The combination of line family and corner points can simultaneously constrain length relationships and endpoint structures, avoiding mistaking short lines caused by folds as epaulette boundaries. The quadrilateral enclosed by the two approximately parallel boundaries maintains geometric stability in the frontal projection and the chest and abdomen cylinder development, facilitating subsequent verification of the vertical and horizontal relationships with adjacent slots.

[0048] The rectangular borders and internal lines are located and scanned within the chest area and the two chest slots. First, the rectangular borders are located based on boundary closure and aspect ratio. Then, the distribution of multiple horizontal short lines is scanned horizontally within the borders, and the area with both rectangular borders and multiple horizontal short lines is selected as the candidate area for the name badge. The manufacturing characteristics of name badges are usually manifested by the coexistence of borders and internal lines. The rectangular borders provide shape constraints, while the multiple horizontal short lines provide content hierarchy constraints. The superposition of the two can exclude interference such as placket creases and pocket folds from the candidates. When supplementary lighting is turned on, the borders often show highlights along the edges, while the brightness of the internal lines changes little. This difference also provides a basis for subsequent reflection consistency verification.

[0049] Closed curved contours were searched and verified through internal and external comparisons in the slots of the left and right arms. First, closed curved contours were searched within the conical unfolded image of the upper arm. Then, the color, brightness, and texture direction inside and outside the contour were compared in pairs. Areas where the difference between internal and external contrast exceeded a set threshold and the texture direction was independent of the fabric grain direction were retained as candidate areas for the armband. The conical unfolded image of the upper arm flattened the circumferential deformation of the arm, and the closed curved contours tended to be stable. If the internal texture was decoupled from the fabric grain direction and the contrast was significant, it indicated that the area belonged to an independent object attached to the fabric surface, which was more in line with the characteristics of armband production and wearing. Supplemental lighting differential verification and reflective consistency verification were performed on all candidate areas simultaneously. The difference in the images before and after supplemental lighting was used to verify the proportional change of local specular enhancement and diffuse reflection, eliminating false highlights that did not conform to the reflective characteristics of three-dimensional entities, and distinguishing printed films from physical accessories from an optical behavior perspective.

[0050] The image structured analysis and relationship reconstruction module, based on orthogonal geometric benchmarks, maps candidate region positions in the orthographic projection image to the corresponding chest and abdomen cylindrical unfolded image and upper arm conical unfolded image. The mapping uses a unified human body reference axis and unified pixel coordinates, binding the position points of the same physical part in different images to the same semantic window. It searches for similar candidate regions near the mapped points, checking whether the boundary orientation is consistent with the longitudinal and transverse directions of the unfolded coordinates and whether it is consistent with adjacent slots. The actual worn element is attached to the curved surface, leaving mutually corroborating boundary fragments and highlight distributions in the three images. If candidate regions exist in two or all three images and their boundary orientation is consistent with adjacent slots, these candidate regions are merged into a multi-view evidence group for the same region. This fusion can filter out occasional candidates caused by shadows and fabric creases from a single viewpoint, retaining only geometrically and optically consistent evidence.

[0051] For each multi-view evidence group, a unique token is generated, and the region category, source slot, adjacent slot relationship, boundary morphology features, and reflectivity consistency results are written into a structured record. When the candidate region is an armband candidate region, texture independence results are appended to the token to illustrate the decoupling of internal texture and fabric grain. The unique token transforms image evidence into verifiable data items, facilitating relation reconstruction and clause comparison processing in a fixed order and with fixed key values, avoiding semantic drift caused by implementation differences. If a candidate region exists only in a single image but its boundary morphology and slot orientation are stable and it does not overlap with adjacent slots, a token is still generated and labeled with a single-view mark to preserve weak evidence in cases of occlusion or insufficient viewpoint. This preservation strategy allows subsequent conflict adjudication to retrospectively improve the evidence level when more frames or more viewpoints are obtained, without affecting the current confirmation of a clearly established multi-view evidence group. Through the above process, candidate region detection, cross-image correspondence, and unique token generation form a continuous evidence solidification chain, providing stable, interpretable, and traceable input for relation reconstruction, conflict adjudication, and clause comparison.

[0052] Furthermore, the image structured analysis and relationship reconstruction module concatenates the tokens into a wearable token sequence according to the predetermined order of the slots. The fixed order is: hat slot token, shoulder slot token, left chest slot token, right chest slot token, left arm slot token, and right arm slot token. This fixed order is used to convert the spatial structure into a linear structure, facilitating judgment and recording within the same data channel. The top-to-bottom, center-to-side arrangement aligns with the natural positions under orthogonal geometric references, directly mapping to the up / down and left / right logic in the clauses, reducing semantic conversion across modules. When generating the wearable token sequence, if multiple candidate tokens exist for each slot, the token generated by the multi-view evidence group is selected first; if there are still ties, the one with the more complete boundary shape is selected. If no token is generated for a slot, a placeholder for the missing token is written at that position along with the reason for the missing token. Each item in the sequence carries the source slot, adjacent slot relationships, boundary shape features, reflection consistency results, and texture independence results in the armband scene. Subsequent adjudication can complete most judgments without returning to the image domain, while ensuring traceability of records.

[0053] Using the token sequence as input, the module performs two rounds of adjudication. The first round checks the continuity of adjacent slot relationships by establishing pairwise checks for adjacent items in the sequence. The check between the hat slot token and the shoulder slot token verifies their relative height and horizontal alignment, requiring the hat slot token to be above the shoulder slot token and both to maintain vertical continuity near the torso centerline. The check between the shoulder slot token and the left and right chest slot tokens verifies their vertical overlap and horizontal affiliation, requiring the left and right chest slot tokens to be simultaneously below the shoulder slot tokens and respectively located on the left and right sides of the torso centerline. The check between the left and right chest slot tokens and the left and right arm slot tokens verifies their horizontal continuity and non-intersecting boundaries, requiring the arm-related tokens to maintain horizontal continuity with the corresponding chest side, and their boundaries not to overlap the chest boundaries. This first round is performed because the continuity of adjacent slot relationships directly reflects the basic correctness of human anatomy and the placement of the tokens. Once the vertical order or left-right affiliation is disrupted, even if a single candidate appears valid in the image, it should not be included in the final structure.

[0054] When inconsistencies occur in the first round, the module performs backtracking and replacement according to a unified strategy. If there is a conflicting backup for the current token, it is replaced with the backup and re-verified. If only a token with a single-view marker causes inconsistency, it first backtracks to the previous position with a multi-view evidence group and tightens the boundary morphology features of that position to avoid coverage judgment errors caused by loose boundaries. If inconsistency still occurs after replacement, a missing placeholder is reserved at that slot position, and the reason for the missing position is recorded as whether the relationship between adjacent slots is not consistent. This process ensures that the overall topology of the sequence remains reliable even in the case of partial occlusion or transient reflective interference.

[0055] The second round verifies the regional consistency of the multiview evidence group by performing cross-image consistency comparisons on tokens originating from the multiview evidence group in the sequence. For the hat slot token and shoulder slot token, it checks whether the boundary orientation is consistent in the orthographic projection and the chest and abdomen cylindrical unfolded image, and verifies whether the reflection consistency results are consistent in the two images. For the left chest slot token and right chest slot token, it checks whether the boundary of the central region in the chest and abdomen cylindrical unfolded image corresponds to the rectangular border of the orthographic projection image. If the orthographic projection image shows an outgoing line but the unfolded domain does not show a corresponding boundary, the multiview evidence is deemed insufficient. For the left arm slot token and right arm slot token, it checks whether the closed curved contour in the upper arm conical unfolded image is consistent with the arc-shaped boundary in the orthographic projection image, and verifies whether the texture independence results remain decoupled in the two images. The reason for placing this round after checking the continuity of adjacent slot relationships is that the spatial location of multiview evidence depends on the local topology determined in the previous round. Ensuring topological continuity first can reduce ambiguity in the cross-image search window, thereby improving the accuracy of consistency judgment.

[0056] If insufficient regional consistency is found during the second round of review, the module performs three actions in sequence. First, if the token has both reflectivity consistency and texture independence results and the only difference is in the boundary orientation, the reflectivity consistency result or texture independence result is used as the primary evidence, requiring a re-comparison after tightening the boundary morphology features in another image. Second, if the evidence for the token comes entirely from a single image and no similar candidate can be found in the corresponding image, the token is downgraded to a single-view token and retained in the sequence, but it is not used for forced matching during the clause comparison stage. Third, if contradictory reflectivity consistency results or texture independence results appear among multi-view evidence, the token is cleared and a missing placeholder is written in the corresponding slot. At the same time, the source of the conflict is recorded in the regional consistency result set, awaiting supplementary evidence from subsequent frames or subsequent passages. Through such review and processing, false candidates caused by shadows, folds, and films can be cleaned up at the cross-image level, and the final retained wearable token sequence takes into account both geometric and optical consistency.

[0057] After two rounds of adjudication, the module generates a wearable structure sequence diagram based on the adjudicated wearable token sequence and simultaneously outputs a regional consistency result set. The wearable structure sequence diagram reconstructs slot nodes based on sequence positions and expresses top-to-bottom overlap and adjacent continuity with lines. The regional consistency result set stores the source slot, adjacent slot relationships, boundary morphological features, reflectivity consistency results, and texture independence results for each token. The "serialization followed by adjudication" process is adopted to complete the maximum range of conflict resolution with minimal state switching, ensuring that the intelligent recognition and compliance management system based on standardized attire outputs a consistent data structure under different devices and postures, providing stable input and traceable evidence for clause comparison and compliance output.

[0058] A grooming mirror with a resolution of The target frontal view image and the depth image in the same field of view are acquired collaboratively. Camera intrinsic parameters are denoted as... ( (where the intrinsic parameter matrix is ​​denoted as ) and the extrinsic parameter is denoted as . ( Let be a rotation matrix. (The translation vector is denoted as ). The acquired front view image of the target is denoted as . Depth image is denoted as ( For pixels The distance (in meters). Obtained through factory calibration. and And apply distortion correction and coordinate alignment.

[0059] Construct a histogram in the depth domain to separate the foreground. Let The depth value represents Pixel count, histogram bin count At the center of the image Regional statistics The main near-peak position was detected as follows: ,in Take the depth window , Take the window's half width as the reference. The reason for choosing the main peak closer to the camera side is... On the one hand, close-range measurements result in smaller quantification errors; on the other hand, the background at the entrance / exit is relatively far away, with the distant peak corresponding to the wall and environment, allowing for better cropping. It can reduce interference from distant views. It generates a depth foreground mask. ,in This is an indicator function. For The appearance foreground mask is obtained by using an adaptive threshold. Find the intersection of the two and perform morphological opening and closing operations (the radius of the structuring element is taken as...). (pixels) to obtain the foreground candidate mask .

[0060] Connectivity labeling yields a set of regions. For each Calculate the circumscribed rectangle ,area Aspect Ratio Depth standard deviation The filtering criteria are: The largest connected component that satisfies the conditions. The outer rectangle is Define the upper body point set. .

[0061] The pixel coordinates of the upper body point set are as follows Computational geometric center and covariance matrix .right The principal direction vector is obtained by eigenvalue decomposition. and secondary direction vector Choose the direction with the smaller angle to the vertical axis of the image as the torso direction, denoted as . Let the center line of the torso be... .exist Above The height range defines the head search band, extracts the upper boundary, and... The point with the smallest ordinate in the neighborhood is taken as the extreme point at the top. .according to The vertical distance determines the shoulder search band; the maximum foreground width is statistically analyzed along the row. The extreme points on the outer sides of the shoulders are obtained by refining the curvature extreme values ​​at both ends of the line. , Define shoulder line reference. .

[0062] by For the crossing point, parallel to Define vertical reference .calculate Angle with the vertical axis of the image Construct rotation and translation transformations make vertical, The horizontal plane, along with the other two, forms an orthogonal geometric datum. Through homography ( (A matrix is ​​used to map the original image onto an orthogonal geometric reference plane) to obtain the orthographic projection image. .by As a horizontal positioning, and The bottom edge is used for vertical positioning; the cropped and normalized edges are then used. Scale. Numerically, the center of the cropping frame. ,Width ,high Then scale it to the target scale.

[0063] The upper body point set is back-projected into 3D using camera intrinsics. (Pixel...) 3D coordinates ,in These are camera coordinates. The cylindrical coordinates, with the torso axis as the cylinder axis, are defined as follows: ,in To pass The axis center point obtained by back projection, The scaling factor from angle to height to pixel (in this embodiment) , (Obtain the development diagram of the thoracic and abdominal cylinder). The size is .

[0064] by and The left and right upper arm conical axes are obtained by neighborhood back projection and joint axis estimation. The conical unfolding coordinates are defined as follows: ,in The three-dimensional coordinates of the shoulder point , , The scaling factor (take) , The cone development diagrams of the left and right upper arms were obtained. All dimensions are .

[0065] exist Set the cap slot as a rectangular area. Shoulder slots are The chest slot is .exist Set the center of the left chest slot at , The center of the right breast slot is at , Side seam slot along .exist Set the left arm slot and right arm slot to cover. , Each slot records the source image identifier, relative position, and adjacent slot number.

[0066] The frame with the fill light off is denoted as The frame is recorded as the fill light is turned on. Canny edges were created on the hat slot, and the closed contour was filtered to obtain three closed candidates, among which the candidate with the largest area was selected. Area is Calculate the mirror response ratio. ,in For the current candidate region, To prevent constants with a denominator of zero. In this example... Exceeding the empirical threshold These areas, arranged in concentric layers, are marked as candidate areas for the cap badge. Probability Hough lines are drawn in the shoulder slots to obtain two main lines. , angle difference Spacing between the two lines The endpoint corner count is Forming a quadrilateral, marked as the epaulette candidate area.

[0067] A minimum bounding rectangle fit and internal horizontal line scan were performed within the chest and both chest slots. A rectangle was detected in the left chest slot. (Width ,high ), Internal continuous short line count The complementary light differential enhances the average intensity along the rectangular edge. Internal line enhancement Regions meeting the characteristics of "stronger edges and weaker interior" for physical name badges are marked as candidate regions. The right chest slot did not reach the row count threshold and is therefore marked as missing. Perform a closed-edge search to obtain candidate results. Calculate the difference in brightness between the inside and outside. ,in These are the mean values ​​inside and outside the candidate region, respectively. Calculate the texture orientation independence index. ,in This represents the direction vector of the main texture inside and outside the region. (Example) Above the threshold Specular enhancement of fill light difference Areas lower than metal components but still showing robust differences in contrast to fabric are marked as candidate areas for armbands. No closed curved edge candidates were detected in the right arm and are marked as missing.

[0068] Will Left chest rectangle Mapped to of Landing point The window detects rectangular segments with consistent boundary orientation. .Will Left arm candidate mapping to window Closed curved edge segments were detected. The same law applies to cap badges and shoulder boards. The corresponding fragment was detected in the front and middle bands. Therefore, a multi-view evidence group was constructed. The right chest and right arm do not correspond and remain missing.

[0069] Generate a unique token for each multi-view evidence group The token content includes the region category. Source slot Adjacent slot relationships Boundary morphology characteristics Consistency of reflectivity and texture independence results in the armband scene. For example, the token on the left breast badge. The token on the left arm patch. Construct a sequence of wearing tokens in a predetermined order. .

[0070] The continuity of adjacent slots is checked based on orthogonal geometric reference coordinates. Verification is made that the cap is positioned above the shoulder. Verify that both breasts are positioned below the shoulders and correctly aligned left and right: .in Projecting tokens to Horizontal and vertical positions, For the center line of the torso Calculate the horizontal coordinates. Verify that the left arm and left chest are horizontally continuous and their boundaries do not intersect, and calculate the horizontal distance. Falling into the permitted zone First round passed; missing items remain as placeholders. Review the regional consistency of the multi-view evidence group. For name tags, compare... and Boundary direction difference Consistency of supplemental lighting The conclusions are consistent in both images; for the left arm patch, the degree of closure is... for ,exist for , Consistent The second round passed; no evidence of multiple vision was found in the right chest and right arm, maintaining the unit's absence.

[0071] Based on the summer field duty stratification criteria: cap badge, shoulder insignia, name tag (at least one on the left or right chest), and armband (at least one on the left or right arm) must be present, and their positions must meet the vertical and horizontal requirements. (Hit set) All conditions must be met for occurrence and positional relationship; the right chest and right arm are acceptable options, and their absence does not affect compliance. The final output conclusion is compliance, and an evidence overlay diagram and a comparison log are generated. The comparison log records the version number. Layered selection Each token And the results of the two rounds of rulings, archived with the number 2025-08-23-001. Replacing the key threshold in this embodiment with a self-learning or configuration value will not change the process: near-peak window Ensures coverage of human body thickness at common passage distances of 1.0–2.0 meters; mirror response ratio threshold. Differential response capable of distinguishing between metal / enamel and ordinary fabric under 8-bit quantization; scaling factor for cylinder and cone expansions. It only affects the pixel scale reaching the two-dimensional plane, and does not affect the logic of cross-image correspondence and two rounds of adjudication.

[0072] like Figure 2As shown, the wearable structure sequence diagram illustrates the core data structure generated after relationship reconstruction and conflict resolution in this invention. This sequence diagram employs a graph theory model, using slots as nodes and connecting them with top-to-bottom coverage relationships and adjacent coherence relationships to construct a hierarchical spatial topology of wearable items. In specific implementation, the system takes the wearable token sequence as input and outputs a structured sequence diagram after two rounds of resolution. The first round of resolution primarily checks the coherence of adjacent slot relationships, including spatial coherence and logical coherence verification. Spatial coherence ensures that the positional relationships between adjacent slots conform to human geometry constraints, while logical coherence verifies the rationality of the dressing order. The second round of resolution reviews the regional consistency of the multi-view evidence group, resolving potential conflicts through cross-image consistency verification and evidence strength assessment. As can be seen from the node distribution in the sequence diagram, the hat slot is located at the top of the structure diagram, establishing a hierarchical connection with the shoulder slots through top-to-bottom coverage relationships (connected by solid lines). The shoulder slots, as intermediate-level nodes, connect downwards to the left and right chest slots, reflecting the vertical coverage logic of the clothing. The left and right chest slots are interconnected through adjacency (connected by dashed lines), reflecting the lateral spatial layout of the chest area. Furthermore, the left and right chest slots are also connected to the left and right arm slots, respectively, forming a spatial topology extending from the torso to the limbs. The weight parameter of each connection reflects the reliability of the corresponding relationship; the system filters low-reliability connections using weight thresholds to ensure the stability of the structural diagram. This invention's wearable structure sequence diagram simultaneously outputs an evidence overlay diagram and a set of regional consistency results, providing a structured data foundation for subsequent clause comparison and compliance determination. This data structure not only records the spatial relationships of each garment item but also retains confidence information during the detection process, supporting verification and traceability functions, significantly improving the system's reliability and practicality.

[0073] like Figure 3As shown, the texture direction angle distribution curve illustrates the core technical principle of texture independence verification in this invention. This technology effectively distinguishes the texture features of the fabric texture and the armband texture by statistically analyzing their direction angles. Specifically, the system uses the Sobel gradient operator to perform edge detection on the region to be detected, with an angle resolution of 5° and a window size of 15×15 pixels. After smoothing the image using a Gaussian filter (σ=1.0), the gradient direction angle of each pixel is calculated. Within the angle range of 0° to 180°, the gradient intensity distribution in each angle interval is statistically analyzed to form a texture direction histogram. From the distribution curve, it can be observed that the fabric texture (blue curve) exhibits a clear unimodal distribution characteristic, with the main direction θ1 concentrated near 0°, corresponding to the horizontal texture direction, which conforms to the regular arrangement characteristics of the warp and weft threads of the fabric. Its gradient intensity reaches a peak of 1.2 in the main direction, while the intensity value is relatively low in other angle intervals, showing obvious directionality. In contrast, the main direction θ2 of the armband texture (red curve) is located at 45°, exhibiting a diagonal distribution characteristic. The peak intensity of the texture is approximately 1.4, and its distribution is relatively concentrated, reflecting the unique directional characteristics of the patch's surface pattern. The angle difference between the principal directions of the two textures is |θ2-θ1|=|45°-0°|=45°, significantly exceeding the preset independence threshold of 30°, thus meeting the technical requirement of texture independence. This invention establishes an angle difference calculation model, determining texture independence when the angle between the principal directions of the two textures is ≥30°. This determination method effectively avoids the technical problem of misidentifying non-target areas such as fabric wrinkles and stains as patches, significantly improving the accuracy and reliability of candidate region detection.

[0074] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent recognition and compliance management system based on standardized dress code, characterized in that, The system includes: an image acquisition module, an image structured analysis and relationship reconstruction module, and a clause comparison and compliance output module. The image acquisition module acquires the target's frontal and depth images using a dressing mirror and constructs an orthogonal geometric datum. Under the constraints of the orthogonal geometric datum, the image structured analysis and relationship reconstruction module generates three projection images of the upper body point set: a frontal projection image, a chest and abdomen cylindrical development image, and an upper arm conical development image. Slots are divided on each image according to preset locations, and the source image, relative position, and adjacent slot numbers are recorded. Candidate regions are detected within each slot to obtain candidate regions, and preliminary elimination is completed using supplementary lighting difference and reflection consistency verification. A candidate region is established between the three images. The corresponding relationships of selected regions are used to form multi-view evidence groups. For each multi-view evidence group, tokens containing region category, source image identifier, slot identifier, and information on adjacent relationships and boundary morphology are generated, and a wearable token sequence is constructed according to the slot order. The relationship reconstruction and conflict adjudication are performed with the wearable token sequence as input to obtain a wearable structure sequence diagram with slots as nodes and upper and lower coverage and adjacent continuity as lines. The evidence overlay diagram and region consistency result set are output simultaneously. The clause comparison and compliance output module compares the wearable structure sequence diagram and wearable token sequence with the preset dress code comparison table item by item, and finally outputs the result of whether the staff's dress is compliant. The wearable token sequence and wearable structure sequence diagram are archived for review.

2. The intelligent recognition and compliance management system based on standardized attire as described in claim 1, characterized in that, The image acquisition module obtains the frontal view image and depth image of the target through the dressing mirror, and obtains the upper body point set by foreground extraction and connected component filtering; the module uses the main direction of the point set to determine the trunk center line, uses the extreme point of the top of the head to determine the vertical reference, and uses the line connecting the extreme points on the outer sides of the two shoulders as the shoulder line reference; the trunk center line, shoulder line reference and vertical reference constitute an orthogonal geometric reference.

3. The intelligent recognition and compliance management system based on standardized attire as described in claim 2, characterized in that, The image acquisition module simultaneously acquires the target frontal view image and depth image through the mirror, ensuring that the two correspond to the same scene at the same time point. It also performs bad pixel filling and median filtering on the depth image to obtain an effective depth image for segmentation. Based on the effective depth image, the depth histogram of the central region of the image is calculated, and the main peak near the camera side is selected as the target depth range. Thresholding segmentation is performed on the entire image within the target depth range to obtain a depth foreground mask. An adaptive thresholding based on brightness and chromaticity is performed on the target front view image to obtain the appearance foreground mask; The intersection of the two masks is calculated and opening and closing operations are performed to remove isolated noise and small holes, resulting in a candidate foreground mask. Connected components are labeled on the candidate foreground mask, and the bounding rectangle, area, aspect ratio, and depth continuity index of each connected component are calculated. Select connected components that meet the following conditions as upper body connected components: the percentage of area to the total area of ​​the image is within a preset area range, the aspect ratio is within a preset aspect ratio range, and the depth variation is within a preset depth range; define the set of all pixels in the upper body connected component in the front view image and the effective depth image as the upper body point set.

4. The intelligent recognition and compliance management system based on standardized attire as described in claim 3, characterized in that, The image acquisition module calculates the geometric center and coordinate covariance features of the upper body point set to obtain the principal and secondary directions. From these two directions, the one with the smaller angle to the vertical direction of the image is selected as the torso direction. Using the geometric center as the crossing point and the torso direction as the pointing direction, a straight line passing through this crossing point is defined as the torso centerline. A head search zone is determined based on the circumscribed rectangle of the upper body point set, located in the upper region of the circumscribed rectangle. The upper boundary contour is extracted within the head search zone, and the boundary point with the smallest ordinate and located in the neighborhood of the torso centerline is selected as the top-of-head extreme point. Boundary continuity is checked in the local neighborhood around the top-of-head extreme point. If the adjacent pixels above are all background and the left and right sides are still foreground within a set distance, then this point is confirmed as the final top-of-head extreme point; otherwise... Then, trace back upwards along the center line of the torso to the nearest boundary point that meets the conditions and take it as the final extreme point of the head. Based on the distance between the extreme point of the head and the lower boundary of the circumscribed rectangle of the upper body point set, define a shoulder search band, where the shoulder search band is located in the upper half of the circumscribed rectangle and has a set shoulder height. Count the foreground width row by row within the shoulder search band and select the row with the largest foreground width as the candidate row for the shoulder line. Obtain the leftmost and rightmost foreground pixels on the candidate row for the shoulder line as the initial values ​​of the extreme points on the outer sides of the two shoulders. Starting from the initial values ​​of the extreme points on the outer sides of the two shoulders, perform short-distance outward expansion search and curvature extreme value retrieval along the outer boundary to obtain two points with the outermost convex boundary and directly adjacent to the background. Define these two points as the extreme points on the outer sides of the two shoulders. Use the line connecting the extreme points on the outer sides of the two shoulders as the shoulder line reference.

5. The intelligent recognition and compliance management system based on standardized attire as described in claim 4, characterized in that, The image structured analysis and relation reconstruction module uses the circumscribed rectangle determined by the upper body point set as the initial cropping range, sets the torso centerline as the vertical reference, and sets the shoulder line reference as the horizontal reference. It calculates the planar projection transformation that remaps the original image plane to the orthogonal geometric reference plane formed by the vertical and shoulder line references. Subsequently, using the line connecting the extreme points on the outer sides of the shoulders as the horizontal positioning basis and the extreme point on the top of the head and the lower boundary of the circumscribed rectangle as the vertical positioning basis, cropping and size normalization are completed to obtain the frontal projection image. The chest and abdomen cylindrical unfolding diagram uses the torso centerline as the cylindrical axis, rearranging the points around this axis into a horizontal sequence according to angular order. The upper arm conical development diagram is formed by taking the conical axis near each scapula as a reference, unfolding the upper arm point set into a transverse sequence along the circumferential direction and forming a longitudinal sequence along the arm length direction; based on the natural boundaries of the three images, the cap slot, shoulder slot, and chest slot are set sequentially from top to bottom on the orthographic projection diagram; the left chest slot and right chest slot are set around the central area on the chest and abdomen cylindrical development diagram, and side seam slots are set on both sides; the left arm slot and right arm slot are set on the upper arm conical development diagram respectively; each slot records its source image, relative position, and adjacent slot number as the sole basis for relationship determination.

6. The intelligent recognition and compliance management system based on standardized attire as described in claim 5, characterized in that, The image structured analysis and relationship reconstruction module performs candidate region detection for each slot, including: edge extraction and closed contour filtering in the hat slot, retaining areas with concentric layers and specular reflection as candidate areas for the cap badge; line family detection, concentric ring structure and corner point counting in the shoulder slot, retaining quadrilateral areas enclosed by two approximately parallel boundaries as candidate areas for the epaulettes; rectangular border positioning and internal line scanning in the chest and two chest slots, retaining areas with rectangular borders and multiple horizontal short lines as candidate areas for the name badges; closed curved contour search and internal and external comparison verification in the left and right arm slots, retaining areas where the internal and external comparison difference exceeds a set threshold and the texture direction is independent of the fabric texture direction as candidate areas for the arm badges; and simultaneously performing supplementary lighting differential verification and reflective consistency verification on all candidate areas to eliminate false specular highlights that do not conform to the reflective characteristics of three-dimensional entities.

7. The intelligent recognition and compliance management system based on standardized attire as described in claim 6, characterized in that, The image structured analysis and relationship reconstruction module is based on orthogonal geometric reference. It maps the candidate region positions in the orthogonal projection map to the corresponding thoracic and abdominal cylindrical unfolded map and upper arm conical unfolded map. It searches for similar candidate regions near the mapping landing point. If there are candidate regions in two or three images and the boundary direction is consistent with the adjacent slot, these candidate regions are merged into a multi-view evidence group of the same region. A unique token is generated for each multi-view evidence group. The token includes: region category, source slot, adjacent slot relationship, boundary morphology features, and reflection consistency results. Furthermore, when the candidate region is an armband candidate region, the token content also includes the texture independence result; if the candidate region exists only in a single image but its boundary shape and slot direction are stable and do not overlap with adjacent slots, a token is still generated and a single-view mark is added.

8. The intelligent recognition and compliance management system based on standardized attire as described in claim 7, characterized in that, The image structured parsing and relation reconstruction module concatenates the tokens into a wearable token sequence according to the predetermined order of the slots; wherein the predetermined order is as follows: hat slot token, shoulder slot token, left chest slot token, right chest slot token, left arm slot token and right arm slot token.

9. The intelligent recognition and compliance management system based on standardized attire as described in claim 8, characterized in that, The image structured analysis and relationship reconstruction module takes the wearing token sequence as input and performs two rounds of adjudication: checking whether the relationship between adjacent slots is coherent and reviewing the regional consistency of the multi-view evidence group.

Citation Information

Patent Citations

  • Police dressing specification detection method and device based on image intelligent analysis, electronic equipment and storage medium

    CN114283445A

  • System and method to detect proper seatbelt usage and distance

    WO2022260793A1