A standardized image acquisition method, system and device
Patent Information
- Application Number
- CN202610931744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-25
AI Technical Summary
现有方式通常仅依赖通用拍照APP完成图像记录,未引入针对采集过程的实时图像识别与质量判定功能,导致拍摄距离、角度及光照条件完全依赖用户经验,难以保证不同采集主体之间的一致性
[0009]本发明有益效果在于:通过在采集区域引入标准化参照贴纸,并将其与目标采集对象进行同场景约束,使采集过程具备统一的空间基准;进一步通过对实时图像数据执行逐帧结构轮廓识别及特征信息提取,结合标尺类结构的空间分布关系与比色区域的成像响应特征,对画面的位置偏移、结构完整性及清晰度进行联合判定,从而在源头上保证采集画面的规范性与一致性;在此基础上,以判定结果反向控制图像采集触发条件,仅在满足预设质量约束时执行采集,有效避免无效图像与低质量数据的产生;同时,通过将目标图像与采集过程中的空间分布状态、结构特征信息等基础数据进行关联绑定,并以结构化数据包形式统一封装与加密存储,使图像数据具备可追溯的上下文信息;由此,彻底改变了以往临床研究中需制定复杂SOP、反复培训医护、依赖微信等私人渠道零散催缴数据且常面临信息缺失的传统模式,实现了仅需下载APP即可完成标准化采集,研究人员在服务器端即可全局掌握各地采集进度与数据质量,以判定结果自动触发归档的闭环控制机制大幅简化了操作流程,显著降低了对人员专业技能的依赖及批量采集成本,使大规模数据收集在基层医疗机构得以高效、便捷地实施,不仅提升了图像采集的标准化程度与数据可靠性,而且增强了后续分析处理的准确性与可溯源能力,显著提高整体数据采集与管理的技术效果。
Smart Images

Figure CN122824974A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of standardized image acquisition technology, and in particular to a standardized image acquisition method, system and device. Background Technology
[0002] While image-based data acquisition methods have been applied in various scenarios, the field of medical macroscopic image acquisition, centered on smart terminal apps, still lacks standardized software acquisition control mechanisms and unified algorithmic constraint processes. Existing methods typically rely solely on general-purpose photo apps for image recording, lacking real-time image recognition and quality assessment functions for the acquisition process. This results in shooting distance, angle, and lighting conditions depending entirely on user experience, making it difficult to ensure consistency across different acquisition subjects. Furthermore, current mobile acquisition processes lack automatic recognition and software-side constraints for standardized reference information, failing to achieve unified control of size and color benchmarks at the app level, leading to significant fluctuations in proportional consistency and color stability of acquired images. Due to the lack of app-based frame-by-frame analysis and real-time validity assessment mechanisms, the shooting process typically involves simple storage processing only after completion, failing to remove non-compliant images during the acquisition phase, resulting in inconsistent data quality. Therefore, there is an urgent need for an acquisition scheme that uses a smart terminal app as the execution entity, integrating real-time image recognition algorithms and standardized reference constraint mechanisms to achieve process-oriented control and automated standardization of the shooting behavior, thereby improving the standardization and comparability of image data. Summary of the Invention
[0003] Therefore, it is necessary to provide a standardized image acquisition method, system, and device to solve at least one of the aforementioned technical problems.
[0004] To achieve the above objectives, a standardized image acquisition method, applied to a smart terminal application, includes the following steps: Step S1: Paste a standardized reference sticker on the collection area of the target object, so that the target object and the standardized reference sticker are within the same shooting angle. Use the application to call the camera of the smart terminal to obtain real-time image data containing the target object and the reference sticker, and generate the basic data set of the image acquisition process. Step S2: Analyze the real-time image data frame by frame to identify the structural outline of the standardized reference sticker; extract the feature information from the standardized reference sticker based on the structural outline, and combine the spatial distribution of the feature information in the image and the image clarity to perform a validity judgment on the current image and output the judgment result; Step S3: Based on the judgment result, control the image acquisition trigger condition. When the judgment result meets the preset condition, control the camera of the smart terminal to acquire the target image and obtain the target image. Step S4: Associate and bind the target image with the basic data set to generate a structured data packet containing image data and associated information; encrypt the structured data packet and send it to the local machine and cloud for archiving.
[0005] Between steps S3 and S4, a no-reference deep learning image quality assessment is performed on the target image, outputting a continuous quality score in the range of 0 to 1. The quality score is compared with a preset quality threshold. Only when the quality score meets the threshold constraint will step S4 be performed to generate structured data packets and perform encrypted archiving; otherwise, a resampling prompt will be triggered.
[0006] The present invention also provides a standardized image acquisition system for performing the standardized image acquisition method described above, the standardized image acquisition system comprising: The standardized reference sticker setting module is used to affix standardized reference stickers to the acquisition area of the target object, ensuring that the target object and the standardized reference sticker are within the same shooting angle. The application then uses the smart terminal's camera to acquire real-time image data containing both the target object and the reference sticker, generating the basic data set for the image acquisition process. The validity determination module is used to analyze real-time image data frame by frame, identify the structural outline of the standardized reference sticker; extract feature information from the standardized reference sticker based on the structural outline, and combine the spatial distribution of feature information in the image and image clarity to perform validity determination on the current image and output the determination result. The image acquisition module is used to control the image acquisition triggering conditions based on the judgment result. When the judgment result meets the preset conditions, it controls the camera of the smart terminal to acquire the target image and obtain the target image. The storage and archiving module is used to associate and bind target images with basic data sets to generate structured data packets containing image data and associated information; it also encrypts the structured data packets and sends them to local storage and the cloud for archiving.
[0007] The image quality assessment module is used to perform a no-reference image quality assessment on the target image and output a quality score; compare the quality score with a preset threshold to generate a quality pass / fail judgment result, and transmit the judgment result to the storage and archiving module as one of the bases for whether to perform encrypted archiving.
[0008] A standardized image acquisition device includes: The intelligent terminal module is equipped with an application that performs image acquisition control, reference sticker feature recognition, validity determination and interactive prompts, and interacts with the remote server through a communication interface. The reference sticker module, located in the adjacent area of the target object, is a flexible attachment structure. The surface of the reference sticker module is divided into a colorimetric card array area and a scale area along its length. The colorimetric card array area integrates multiple standard color blocks arranged in a matrix to provide a reference for color reproduction and white balance calibration during image acquisition. The scale area integrates continuous and equally spaced scale markings to provide a physical reference for size measurement. The reference sticker module also includes a black and white grid area located between the colorimetric card array area and the scale area. The black and white grid area is composed of alternating black and white squares and is used to assist the algorithm in identifying image distortion and edge focus status during image acquisition. The surface integrates rulers, color charts, and black and white grids to provide size, color references, and deformation calibration standards; An image acquisition module, located on a mobile terminal, includes a camera and is used to acquire image data containing the target object and a reference sticker module, and transmit it to the intelligent control module. The data encapsulation module, integrated into the intelligent control module, is used to associate image data, validity judgment results, and acquisition time and location information to generate structured data packets; The communication storage module is electrically connected to the intelligent control module. It includes a communication interface and a storage unit, and is used to encrypt structured data packets and upload them to an external storage terminal for archiving via a communication network.
[0009] The beneficial effects of this invention are as follows: By introducing standardized reference stickers into the acquisition area and constraining them with the target acquisition object within the same scene, a unified spatial benchmark is provided for the acquisition process; furthermore, by performing frame-by-frame structural contour recognition and feature information extraction on real-time image data, and combining the spatial distribution relationship of scale-like structures with the imaging response characteristics of the colorimetric region, the positional offset, structural integrity, and sharpness of the image are jointly determined, thereby ensuring the standardization and consistency of the acquired image from the source; based on this, the determination results are used to control the image acquisition triggering conditions, and acquisition is only performed when preset quality constraints are met, effectively avoiding the generation of invalid images and low-quality data; simultaneously, by associating and binding the target image with basic data such as the spatial distribution state and structural feature information during the acquisition process, and presenting it in the form of structured data packets... Unified packaging and encrypted storage endow image data with traceable contextual information. This completely changes the traditional model of clinical research, which required complex SOPs, repeated training of medical staff, reliance on private channels such as WeChat for sporadic data collection, and often faced information gaps. It enables standardized data collection by simply downloading an APP. Researchers can monitor the collection progress and data quality in various locations from the server side. The closed-loop control mechanism that automatically triggers archiving based on the results greatly simplifies the operation process, significantly reduces the dependence on personnel professional skills and the cost of batch collection, and enables large-scale data collection to be implemented efficiently and conveniently in primary healthcare institutions. It not only improves the standardization and reliability of image collection, but also enhances the accuracy and traceability of subsequent analysis and processing, significantly improving the overall technical effect of data collection and management.
[0010] Furthermore, by introducing no-reference deep learning image quality assessment, the target image is comprehensively scored based on structural compliance judgment, enabling automatic quality review and batch screening after acquisition. This allows researchers to have a global grasp of data quality distribution on the server side, further strengthening the closed-loop control effect of quality judgment triggering archiving. Attached Figure Description
[0011] Figure 1 A schematic diagram illustrating the steps of a standardized image acquisition method; Figure 2 This is a schematic diagram of a standardized image acquisition system; Figure 3 A schematic diagram of the standardized image of the ruler; Figure 4 A flowchart for standardizing the compliance determination and triggering of image acquisition for reference stickers; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0012] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0013] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0014] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0015] To achieve the above objectives, please refer to Figures 1 to 4 A standardized image acquisition method, applied to a smart terminal application, includes the following steps: Step S1: Paste a standardized reference sticker on the collection area of the target object, so that the target object and the standardized reference sticker are within the same shooting angle. Use the application to call the camera of the smart terminal to obtain real-time image data containing the target object and the reference sticker, and generate the basic data set of the image acquisition process. Step S2: Analyze the real-time image data frame by frame to identify the structural outline of the standardized reference sticker; extract the feature information from the standardized reference sticker based on the structural outline, and combine the spatial distribution of the feature information in the image and the image clarity to perform a validity judgment on the current image and output the judgment result; Step S3: Based on the judgment result, control the image acquisition trigger condition. When the judgment result meets the preset condition, control the camera of the smart terminal to acquire the target image and obtain the target image. Step S4: Associate and bind the target image with the basic data set to generate a structured data packet containing image data and associated information; encrypt the structured data packet and send it to the local machine and cloud for archiving.
[0016] This embodiment provides a standardized image acquisition implementation method based on a smart terminal application (APP), suitable for home or personal use scenarios. Users install a dedicated image acquisition APP on a mobile smart terminal (including but not limited to smartphones and tablets) and use it in conjunction with disposable flexible standardized reference stickers to achieve standardized image acquisition of the target object (e.g., skin surface area).
[0017] At the start of data acquisition, the user affixes a standardized reference sticker to an area adjacent to the target object, ensuring that both the target object and the standardized reference sticker are within the same field of view of the smart terminal's camera. The app acquires a real-time image data stream containing the target object and the standardized reference sticker by accessing the smart terminal's camera, and simultaneously constructs a basic data set during the acquisition process. This basic data set includes at least: real-time image frame sequence data; the initial spatial distribution of the standardized reference sticker in the image; and the device acquisition timestamp and acquisition session identifier. The spatial distribution of the standardized reference sticker in the image is characterized as follows: locating the imaging area of the standardized reference sticker in the real-time image data and extracting its pixel distribution range in the image; expressing the spatial positional relationship of the standardized reference sticker in the image based on the pixel distribution range to form the corresponding spatial distribution state.
[0018] The app performs frame-by-frame analysis of real-time image data. Its core employs a multi-level integrated visual algorithm framework to perform structural recognition and compliance assessment of standardized reference stickers. For each frame, object detection is performed using the following algorithm: an optimized deep learning object detection model based on MobileNetSSD or YOLO-tiny, or a pre-trained feature-specific classifier, is applied to extract features and detect objects in each frame. The model identifies the geometric boundaries of the standardized reference stickers, generating candidate regions for structural contours.
[0019] Under the constraints of mobile computing power, a rapid initial screening is further performed: the Sobel operator is used to extract the edge map; the edge density is calculated to determine whether the sticker appears completely; the Tenengrad gradient function is used to calculate the sharpness index; the LBP texture operator is used to identify the scale area; the HSV color space is used to segment and identify the colorimetric area; if the edge density, sharpness, or color features do not meet the threshold, the current frame is directly determined to be unsuitable for acquisition.
[0020] Input the multidimensional features into the Lightweight Gradient Boosting Model (LightGBM) and output the initial screening probability value P: if P < threshold T1 → the current frame is invalid; if P ≥ threshold T1 → enter the precise judgment stage.
[0021] For image frames that pass the initial screening, a second stage of precise analysis is performed: A multi-task learning network is constructed using ShuffleNetV2+ Multi-Scale Feature Pyramid (FPN-lite) structure combined with a self-attention mechanism. This includes: precise localization based on YOLOv5s or SSD models of the standardized reference sticker bounding box, scale marking area, and colorimetric area; outputting overall image sharpness score, scale area sharpness, and colorimetric area sharpness through a fully connected regression network; and achieving pixel-level segmentation based on MobileNetV3+DeepLabV3+Lite to determine whether the sticker body is complete; whether the scale area coverage is ≥ a preset threshold; and whether the colorimetric area is fully displayed. The following constraints are used for joint judgment: spatial offset (based on the difference between the actual coordinates of the scale / color chart and the reference coordinates), edge sharpness index, segmentation integrity index, and color response consistency index, to form a unified judgment result: valid or invalid, and the result is output in real time through the APP interface.
[0022] The app sends a shooting trigger command to the smart terminal to control the camera to perform high-resolution image acquisition only when the current frame meets the validity criteria. At the moment of acquisition, the current frame state is locked to prevent data distortion caused by jitter or offset. Simultaneously, automatic optimization processing is performed, including: white balance correction based on a colorimeter; scale normalization based on a scale; and lightweight image enhancement (sharpening / noise reduction).
[0023] After data acquisition, the app automatically executes a data encapsulation process, binding the following information to the target image data: basic data set (frame sequence + spatial distribution status); validity determination result; device information (device ID, system version); timestamp and location information; and generating a data packet according to a structured format, such as JSON or EXIF extended structure. Encryption processing (such as AES encryption or a hybrid encryption mechanism) is performed, and the data is synchronously uploaded via the communication module to: the local storage module; and the cloud server archiving system, forming a complete standardized image acquisition closed-loop data chain.
[0024] Please refer to [link / reference needed] for further information. Figure 3 This is an example of standardized image acquisition for this case. The scratches on the forearm, the vertically attached ruler, and the color chart are all within the same frame. The ruler provides a precise length reference, while the color chart is used for color calibration, ensuring the standardization and comparability of the images in terms of size and color. The appearance size and color chart specifications can be flexibly adjusted and adapted according to different clinical acquisition needs. Furthermore, this standardized reference sticker has been mass-produced and standardized. The small size shown on the right is the mainstream form for current mass applications, aiming to further improve portability and applicability to primary care through compact design.
[0025] Please refer to [link / reference needed] for further information. Figure 4 It begins by detecting four elements: frame border, ruler, color chart, and focus, and determines whether they are satisfied simultaneously. If not, the interface prompts for adjustment. After capturing real-time images, it performs real-time compliance checks in a loop. If the conditions are met, the interface displays compliance confirmation. After the user triggers shooting (manual / automatic), it performs shooting, binds metadata, encrypts and uploads the data, and returns to the acquisition interface to prepare for the next shooting, thus achieving compliance control and automated processing of the entire acquisition process.
[0026] Preferably, step S1 includes: Standardized reference stickers are set up within the collection area, and the relative positional relationship between the standardized reference stickers and the target collection object is defined, so that the standardized reference stickers form a spatial constraint benchmark within the field of view; Based on spatial constraints, the structural layout of the standardized reference sticker is preset, the spatial orientation characteristics of the standardized reference sticker are determined, and the target acquisition object and the standardized reference sticker are placed in the same field of view. The application calls the camera of the smart terminal to acquire real-time image data containing the target object and standardized reference stickers, and retains the spatial distribution of the standardized reference stickers in the image. The basic dataset is constructed by associating real-time image data with spatial distribution status.
[0027] In this embodiment, the execution process of "setting standardized reference stickers and constructing spatial constraint benchmarks within the acquisition area" in step S1 is achieved collaboratively by the smart terminal APP and the disposable flexible standardized reference stickers. The user pastes the standardized reference stickers in the vicinity of the target acquisition object in the home environment, so that the reference stickers and the target acquisition object enter the same field of view of the smart terminal camera, thereby establishing an initial spatial constraint relationship at the physical level. After starting the acquisition session, the APP calls the camera to acquire a real-time image data stream containing the target acquisition object and the standardized reference stickers, and performs preliminary localization processing based on a pre-trained visual model in each frame of the image. Specifically, it uses "application of a deep learning object detection model optimized based on MobileNetSSD or YOLO-tiny, or a pre-trained specific feature classifier to extract features and detect objects in each frame of the image" to identify the initial geometric position and boundary area of the standardized reference stickers in the image. On this basis, the sticker contour pixel set is further extracted through edge detection and connected component analysis, and the bounding rectangle of the contour and the coordinates of its center point in the image coordinate system are used as spatial position representation quantities, thereby constructing the spatial distribution state of the standardized reference stickers in the image.
[0028] Based on the pre-defined structural layout of the scale and colorimetric areas within the sticker, its spatial orientation characteristics are pre-modeled. The relative geometric relationship between the scale and colorimetric areas in the standard template is used as a baseline structural constraint and mapped to the current frame image coordinate system. This constraint governs the relative positional relationship between the target object and the sticker, ensuring the target object is within the effective framing area defined by the sticker. During continuous framing, the app performs stability tracking on consecutive frames. By analyzing the time series of changes in the sticker's center point coordinates, rotation angle, and scale, a dynamic and stable expression of the spatial constraint baseline is formed. The app also determines in real-time whether the target object and the sticker are simultaneously within the effective framing range. Real-time image data is bound and encapsulated with the spatial distribution state, composed of "sticker boundary information, spatial coordinate distribution state, and structural orientation relationship parameters output by the detection model," forming a basic data set. This basic data set is synchronously associated with timestamps, device identifiers, and acquisition session IDs during generation, providing a unified data foundation for subsequent frame-by-frame compliance judgment and structured image encapsulation.
[0029] Preferably, the process involves using an application to access the camera of a smart terminal to acquire real-time image data containing the target object and standardized reference stickers, while preserving the spatial distribution of the standardized reference stickers within the image. The application calls the camera of the smart terminal to acquire real-time image data containing the target object and standardized reference stickers; Locate the imaging area of the standardized reference sticker in real-time image data and extract its pixel distribution range in the image; The spatial positional relationship of the standardized reference sticker in the image is expressed based on the pixel distribution range, forming the corresponding spatial distribution state; The spatial distribution state is recorded and maintained in real-time image data, so that the spatial distribution state of the standardized reference sticker is stored synchronously with the image data.
[0030] In this embodiment, the execution process of "acquiring real-time image data containing the target object and standardized reference stickers through the application calling the camera of the smart terminal, and retaining the spatial distribution state of the standardized reference stickers in the image" is completed by the real-time vision processing module of the smart terminal APP. After the user starts the acquisition, the APP acquires a continuous video frame data stream containing the target object and standardized reference stickers by calling the camera of the smart terminal, and performs synchronous processing on each frame image. After the image enters the processing pipeline, the image is initially located and identified based on a pre-trained target detection model. Specifically, "the application of a deep learning target detection model optimized based on MobileNetSSD or YOLO-tiny, or a pre-trained specific feature classifier, is used to extract features and detect targets in each frame image" to quickly lock the candidate imaging area of the standardized reference sticker in the image and output its initial bounding box coordinate information. Building upon this foundation, the app further refines the candidate regions with pixel-level processing. By performing edge detection (such as using the Sobel or Canny operators) and connected component analysis on the images within the bounding boxes, it extracts the pixel-level distribution range of the standardized reference sticker, including its outer contour pixel set, aspect ratio features, and geometric center point coordinates. This enables precise localization of the sticker's imaging area. Based on the extracted pixel distribution range, the spatial positional relationship of the standardized reference sticker in the image is expressed in a structured manner. Specifically, the coordinates of the bounding rectangle, the center point coordinates, and the scale ratio parameters are uniformly mapped to the image coordinate system to form a spatial distribution state vector. This spatial distribution state vector includes at least position coordinate information, scale information, and rotation angle information, used to characterize the spatial pose of the sticker in the current field of view. After completing the spatial representation, the app binds and stores this spatial distribution state with the corresponding real-time image frame. That is, when each frame of image data is generated, the corresponding spatial distribution state data is synchronously attached, and its temporal series consistency is maintained through a memory caching mechanism, so that the spatial distribution state of the standardized reference sticker can be continuously updated and synchronously recorded with the real-time image data stream. Simultaneously, consistency verification of spatial distribution status is performed between consecutive frames. By conducting time series analysis on changes in center point trajectory, scale fluctuation range, and angle offset, the stability and traceability of spatial distribution status are ensured, thereby forming a synchronous structured data stream containing "image frame data + sticker spatial distribution status", providing basic input data for subsequent validity determination and standardized acquisition control.
[0031] Preferably, step S2 includes: Frame-by-frame scanning is performed on real-time image data, structural contours are extracted based on the geometric boundary features of standardized reference stickers, and consistency comparison of structural contours is performed in consecutive frames to determine the existence status of standardized reference stickers. Under the constraint of structural contour consistency results, the feature information in the standardized reference sticker is analyzed within the area defined by the structural contour. The feature information includes the imaging response features of the scale marking area and the colorimetric area. Based on the relative positional relationship of the feature information within the structural contour, the structural integrity judgment result is output. Under the condition that the structural integrity judgment result is valid, the distribution offset degree is calculated based on the spatial distribution of feature information in the image, and the image quality judgment parameters are constructed by combining the edge sharpness of the structural contour and the imaging response characteristics of the colorimetric region. Joint constraint analysis is performed based on the distribution offset degree and image quality judgment parameters, and the validity judgment result corresponding to the current frame is output.
[0032] In this embodiment, the process of performing frame-by-frame scanning and validity determination of real-time image data in step S2 is completed by the multi-level visual analysis algorithm module built into the smart terminal APP. After the user completes the acquisition preparation in step S1, the APP analyzes the real-time video stream output by the camera frame by frame. In each frame, structural contour extraction is performed based on the geometric boundary features of the standardized reference sticker. The gradient calculation of the image is performed using edge detection operators (such as Sobel operator or Canny operator). A candidate contour set is generated by combining connected component analysis and contour closure verification. The candidate contours are screened according to the geometric constraints preset by the standardized reference sticker (including aspect ratio range, rectangularity threshold and boundary continuity constraints) to obtain stable structural contours. The APP performs consistency comparison processing on the stable structural contours between consecutive frames. By performing time series matching analysis on the size ratio change of the contour, the center point drift and the boundary shape similarity, it determines whether the structural contour remains stable between multiple frames, thereby determining the existence state of the standardized reference sticker in the current acquisition process. Under the constraint that the structural contour consistency result is met, the analysis area is further limited to the inside of the structural contour. Feature information analysis is performed on this area. The feature information includes at least the imaging response features of the scale marking area and the colorimetric area. Specifically, local region segmentation is performed inside the contour, and the scale area and colorimetric area are identified and extracted by combining HSV color space distribution features and local texture features (such as LBP operator). The relative spatial position relationship of the scale area and the colorimetric area within the structural contour is calculated, including the vertical distribution relationship, the center offset distance and the proportional relationship matching degree. The relative position relationship is then compared with the preset standard layout template for consistency.
[0033] When both the scale marking area and the colorimetric area satisfy the preset structural topology relationship, the output structural integrity judgment result is valid; otherwise, the output is invalid. Under the premise that the structural integrity judgment result is valid, the APP further calculates the distribution offset degree based on the spatial distribution of feature information in the image. Specifically, it calculates the spatial offset by calculating the difference between the actual position of the scale area and the colorimetric area in the image coordinate system and the standard template position, and weights and fuses multiple offsets to form a unified distribution offset degree index. At the same time, it calculates the edge sharpness index by combining the gradient response intensity of the structural contour edge, and performs quantitative analysis on the color response consistency of the colorimetric area (based on RGB / HSV distribution variance) to construct a comprehensive judgment parameter that characterizes the current image quality. The distribution offset degree and image quality judgment parameters are input into the joint constraint analysis module. Consistency verification is performed by setting multi-dimensional threshold constraint rules (including offset level threshold, sharpness threshold and color consistency threshold). When the spatial offset degree is lower than the preset threshold and the image quality judgment parameters meet the sharpness and color stability constraints, the current frame is output as a valid acquisition frame. The validity judgment result is fed back to the APP front end in real time to control the shooting process. Otherwise, the current frame is determined to be invalid and a readjustment prompt is triggered. This realizes a frame-by-frame closed-loop control acquisition mechanism based on "structural contour consistency + feature area spatial relationship + multi-dimensional image quality evaluation".
[0034] Preferably, the real-time image data is scanned frame by frame, the structural contour is extracted based on the geometric boundary features of the standardized reference sticker, and the consistency of the structural contour is compared in consecutive frames to determine the existence state of the standardized reference sticker, including: Candidate boundaries are extracted from real-time image data and filtered according to the geometric constraints of standardized reference stickers to output the initial structural contour. Based on the initial structural contour, perform closure and boundary continuity checks, and eliminate contour segments that do not meet the constraints to obtain stable structural contours. The morphological parameters of the stable structural profile are matched in consecutive frames, and structural profile consistency data is output based on the size ratio and the degree of consistency of boundary morphology. The validity of stable structural profiles is confirmed based on structural profile consistency data, and the existence status of standardized reference stickers is output when consistency constraints are met.
[0035] In this embodiment, the process of "performing frame-by-frame scanning of real-time image data and determining the existence status of standardized reference stickers" is executed by the frame-by-frame visual recognition and structural stability analysis module built into the smart terminal APP. After the user starts the acquisition, the APP continuously receives the real-time video stream output by the camera and performs frame-by-frame scanning processing on each frame of the image. For the current frame image, a candidate boundary extraction operation is performed. After grayscale processing of the image, gradient calculation is performed using the Sobel operator or the Canny edge detection operator to obtain the set of all potential edge contours in the image, forming a candidate boundary dataset. The candidate boundaries are filtered according to the preset geometric constraints of the standardized reference sticker. The geometric constraints include a preset aspect ratio range, a rectangularity threshold, edge linearity constraints, and an area ratio range. By comprehensively filtering the circumscribed rectangle parameters of the candidate boundaries, the contour polygon approximation results, and the edge line fitting errors, an initial structural contour that meets the constraints is output.
[0036] After obtaining the initial structural contour, the APP further performs closure and boundary continuity verification. By performing connectivity analysis and closure detection on the contour point set, it determines whether the contour forms a complete closed curve. At the same time, it performs breakpoint detection and curvature continuity analysis on the contour boundary, eliminating contour segments with obvious breaks, noise interference, or discontinuous boundaries. Contours that meet the closure and continuity constraints are identified as stable structural contours. In the continuous frame processing stage, cross-frame consistency matching analysis is performed on the stable structural contours. By extracting the contour's morphological parameters, including but not limited to contour area, circumscribed rectangle size ratio, center point coordinates, and boundary curvature features, and calculating the parameter change rate and similarity index between adjacent frames on the time series, structural contour consistency data is output based on the consistency of size ratio and boundary morphology.
[0037] The consistency data obtained from consecutive frames is input into the consistency determination module. By setting a stability threshold condition, the persistence and stability of the contour in multiple frames are comprehensively evaluated. When the structural contour meets the size fluctuation range constraint and the morphological similarity is higher than the preset threshold in consecutive frames, the stable structural contour is determined to be a valid contour, and the existence status of the standardized reference sticker is output as valid. This realizes a progressive visual recognition process based on candidate boundary extraction, geometric constraint screening, closure verification, and cross-frame consistency verification. The current sticker recognition status is fed back in real time through the APP for subsequent image acquisition control.
[0038] Preferably, under the constraint of structural contour consistency results, the feature information in the standardized reference sticker is parsed within the area defined by the structural contour, and the structural integrity judgment result is output based on the relative positional relationship of the feature information within the structural contour, including: Feature information is extracted from the imaging area of the standardized reference sticker within the area defined by the structural contour. The feature information includes the imaging response features of the scale marking area and the colorimetric area. The scale marking area is identified based on the imaging response characteristics of the scale marking area, and the colorimetric area is identified based on the imaging response characteristics of the colorimetric area, generating the identification result. Based on the recognition results of the scale marking area and the colorimetric area, the relative positional relationship between the scale marking area and the colorimetric area within the structural contour is extracted; The relative positional relationship is matched with the preset positional relationship, and the existence status and positional relationship of the scale marking area and colorimetric area are determined based on the matching result; When both the scale marking area and the colorimetric area meet the preset positional relationship, the output structural integrity judgment result is valid; otherwise, the output structural integrity judgment result is invalid.
[0039] In this embodiment, the process of parsing the feature information of the standardized reference sticker and outputting the structural integrity judgment result under the constraint of structural contour consistency is executed by the structural semantic parsing and region consistency judgment module in the smart terminal APP. After the user confirms the existence status of the sticker, the APP only enters the fine parsing process for image frames that meet the structural contour consistency constraint. The ROI region extraction operation is performed within the determined structural contour defined area. That is, based on the stable structural contour coordinates output by the previous steps, the corresponding effective area of the sticker is extracted in the image coordinate system, and multi-channel feature information extraction processing is performed in this area. The feature information includes the imaging response features of the scale marking area and the imaging response features of the colorimetric area. By performing gray-level gradient analysis, LBP texture feature extraction and HSV color space distribution modeling on the ROI areas respectively, the scale marking area is enhanced by high-frequency texture density features and periodic edge response features, and the colorimetric area is identified by color clustering analysis and color block stability detection, thereby generating candidate recognition results corresponding to the scale marking area and the colorimetric area respectively.
[0040] After feature extraction, the app fuses and verifies the results of deep learning detection (e.g., region detection boxes output by YOLO-tiny or MobileNetSSD optimized models) with traditional image feature results. It then confirms and identifies the scale marking region and the colorimetric region separately, outputting the identification results for both types of regions, including their precise position coordinates within the structural contour, the size of the circumscribed rectangle, and the coordinates of the center point. Based on the identification results, the app extracts the relative positional relationship between the scale marking region and the colorimetric region within the structural contour. Specifically, it calculates the relative distance between their center points, the vertical and horizontal offsets, and the normalized proportional relationship in the preset template coordinate system, thereby forming a spatial topological relationship expression within the structure.
[0041] Next, the APP matches and verifies the relative positional relationship with the pre-defined standardized reference sticker template positional relationship. The template positional relationship is constructed based on fixed layout parameters calibrated at the factory, including the standard vertical distribution relationship of the scale area and the colorimetric area, the standard spacing ratio, and the standard alignment deviation range. By calculating the difference between the actual detection result and the template parameters, it is determined whether the scale marking area and the colorimetric area simultaneously meet the positional constraint conditions. When both the scale marking area and the colorimetric area meet the preset threshold range in terms of spatial position, size ratio, and relative layout relationship, the structural integrity judgment result is output as valid, and the sticker structure of the current frame is marked as complete and usable. Otherwise, the structural integrity judgment result is output as invalid, and the APP interface prompts the user in real time to adjust the shooting angle or sticker position, thereby realizing the automatic structural integrity judgment process based on "multimodal feature recognition + spatial topology calculation + template matching constraint".
[0042] In another embodiment, an automatic mask generation algorithm based on SAM (Segment Anything Model) is used to segment the image at the pixel level to generate a set of candidate masks. Each candidate mask is then subjected to multi-level filtering to accurately determine whether it is a valid color chart region.
[0043] The input image is preprocessed as follows: the longer side of the image is scaled to a preset size, CLAHE contrast enhancement is performed on the L channel in the Lab color space, and bilateral filtering is used to remove noise while preserving edges. The preprocessed enhanced image is used for segmentation, while the scaled-only unenhanced image is retained for subsequent color analysis to avoid introducing tone shifts during enhancement. The SAM ViT-B model's automatic mask generator (SamAutomaticMaskGenerator) performs zero-shot full-image segmentation on the preprocessed image. A uniform grid of points (default 32×32=1024 cue points) is laid on the image, and the mask is predicted point by point. A multi-scale cropping strategy is employed to generate a full-image cropping layer and four 2×2 sub-cropping blocks, improving recall of targets at different scales. Built-in quality filtering includes: prediction IoU threshold ≥0.86, stability score threshold ≥0.92, bounding box NMS threshold 0.7, and a minimum mask area of 100 pixels, filtering out low-quality and duplicate masks.
[0044] Shape filtering is performed on the masks generated by SAM: the outer contour of the mask is extracted, and the ratio of the contour area to the area of the minimum bounding rectangle (cv2.minAreaRect) is calculated, defined as the rectangularity. A rectangularity ≥ 0.8 and a contour area ≥ 10 pixels are required to ensure that the candidate masks have an approximately rectangular geometric shape and to exclude irregular fragments.
[0045] Color verification is performed on the mask that has passed shape filtering: Pixels within the mask area are analyzed in the HSV color space on the unenhanced scaled image. Only pixels with saturation ≥ 80 and lightness > 40 are included in the color determination, excluding gray undersaturated areas and dark shadows. The three color components—red (hue < 10 or > 170), green (hue ∈ [35, 85]), and blue (hue ∈ [100, 130])—are detected separately. The mask must contain all three colors (red, green, and blue) to pass color verification, ensuring that the detected object is a standard color correction chart and not some other rectangular object.
[0046] Spatial relationship analysis is performed on candidate masks selected through both shape and color filtering: the bounding boxes of candidate masks are compared pairwise. If the bounding box of mask A completely contains the bounding box of mask B, and the area of A is ≥ 95% of the area of B, then A is determined to be external occlusion interference (such as feet, hands, etc.) and is discarded, retaining the smaller internal mask B as the valid color swatch area. This logic ensures that when the color swatch is partially covered by occlusions, the system can still output the accurate segmentation area of the color swatch itself.
[0047] The final output detection results are as follows: the brightness of the area outside the mask is reduced to 30%, while the masked area retains its original brightness, generating an overlay visualization image; simultaneously, JSON format metadata is output, recording the ID, bounding box coordinates (xywh), area, SAM prediction IoU, stability score, and detection results of each color component (has_red / has_green / has_blue) for each detected color swatch. Detection is considered successful when at least one mask passes all filtering conditions; otherwise, no valid color swatch is detected, and the results are output to the _not_found directory.
[0048] In another embodiment, the accurate color card segmentation results obtained from the above-mentioned SAM automatic mask generation and multi-level filtering can be further used to construct a large-scale training dataset to optimize the performance of subsequent YOLO series object detection models or other lightweight detection models, thereby eliminating the dependence on the computationally intensive SAM module.
[0049] Specifically, the effective color swatch mask output by SAM (including bounding box coordinates, mask region, and JSON metadata) is applied to unlabeled or weakly labeled original acquired images to generate synthetic training data in the following manner: First, the color swatch mask obtained by SAM segmentation is post-processed to extract its transparent mask image or binary mask; then, the color swatch region (including standard color blocks) corresponding to the mask is cropped from the original image, or the mask overlay information is directly retained; next, batch processing is performed on a large number of unlabeled target acquisition scene images (e.g., skin / wound images under different lighting, angles, and distances). Quantity processing: Using enhancement strategies such as random affine transformation (rotation, scaling, translation, slight perspective distortion), color perturbation, and lighting simulation, the extracted standardized color chart regions are superimposed onto appropriate positions on the unlabeled image according to a preset relative positional relationship (referencing the fixed layout of the scale-color chart in the patent), forming a synthetic image containing a complete standardized reference sticker; at the same time, based on the metadata (boundary box, IoU score, color verification result) output by SAM, corresponding annotation information is automatically generated, including the boundary box annotation of the color chart region, pixel-level mask annotation, structural integrity label, and offset degree label; Using the above method, diverse and high-quality training datasets ranging from tens of thousands to hundreds of thousands of data points can be automatically generated from a small amount of real-world labeled data in a short time, significantly reducing the cost of manual annotation and improving the scene coverage of the dataset. Based on the generated synthetic dataset, YOLOv5s, YOLOv8n, or other lightweight object detection models are fine-tuned end-to-end, further improving the model's detection accuracy and robustness in real-time mobile environments for the structural contours, scale markings, and colorimetric areas of standardized reference stickers. After training, the optimized YOLO model is integrated into a smart terminal application to replace or assist the initial object detection step in step S2 above.
[0050] During dataset generation, SAM segmentation results from real-world images can be combined with synthetic images for mixed training (using Mixup or CutMix strategies) to further mitigate the neighborhood offset problem. This implementation achieves a technical link from "detection" to "data closed-loop self-enhancement" through SAM zero-shot segmentation capabilities, which not only reduces the dependence of model training on large-scale manual annotation but also significantly improves the algorithm generalization ability and deployment efficiency of the entire standardized image acquisition system.
[0051] Preferably, under the condition that the structural integrity determination result is valid, the distribution offset degree is calculated based on the spatial distribution of feature information in the image, and the image quality determination parameters are constructed by combining the edge sharpness of the structural contour and the imaging response characteristics of the colorimetric region, including: Under the condition that the structural integrity judgment result is valid, the spatial position of the scale marking area and colorimetric area in the image is extracted, the distribution coordinates of them relative to the structural contour are obtained, and the corresponding benchmark distribution is established based on the preset standard distribution relationship. Based on the matching of the distribution coordinates with the reference distribution, the spatial offset of the scale marking area and the colorimetric area is calculated, and the spatial offset is integrated to obtain the degree of distribution offset. Edge response information of the boundary region is extracted based on the structural contour, and edge sharpness is calculated. Based on a joint analysis of distribution offset and edge sharpness, a unified set of image quality judgment parameters is formed to characterize the image state.
[0052] In this embodiment, the process of "calculating the distribution offset and constructing image quality judgment parameters under the condition that the structural integrity judgment result is met" is executed by the spatial geometric analysis and image quality assessment module in the smart terminal APP. After the user confirms the structural integrity of the standardized reference sticker in the previous steps, the quality assessment stage is only entered for structurally complete frames. Based on the identified structural contour coordinate system, the APP extracts the precise spatial positions of the scale marking area and the colorimetric area in the image coordinate system, including their center point coordinates, circumscribed rectangle coordinates, and region centroid coordinates, and maps the spatial positions uniformly to the structural contour normalized coordinate system, thereby obtaining the actual distribution coordinates of the scale marking area and the colorimetric area within the structural contour. At the same time, a benchmark distribution model is established based on the preset standardized reference sticker template. This model defines the standard relative positional relationship, standard spacing ratio, and standard center alignment coordinates of the scale marking area and the colorimetric area, which are used as a spatial comparison benchmark.
[0053] The app performs spatial offset calculations, matching the actual distribution coordinates point-by-point with the baseline distribution model. It calculates the center point offset, scale deviation, and angular offset between the scale marking area and the colorimetric area, and weights and fuses these offset components to form a unified spatial distribution offset index. Simultaneously, it calculates the local offset contribution value for both the scale and colorimetric areas, summarizing these to obtain a distribution offset degree parameter, which characterizes the degree of spatial stability deviation of the current sticker in the image. After completing the spatial offset calculation, the app further performs edge response analysis on the structural contour boundary area. By calculating the gradient of contour edge pixels (e.g., based on the Sobel or Laplacian operator), it extracts the edge intensity distribution curve and calculates the average edge gradient and high-frequency response energy, thus forming an edge sharpness index to characterize the current image's focus quality and boundary sharpness. Simultaneously, it performs imaging response analysis on the colorimetric area, calculating the color mean, variance, and color offset of each standard color block in the colorimetric area in the HSV color space to evaluate its color reproduction consistency and stability, thereby obtaining the imaging response characteristic value of the colorimetric area.
[0054] The app inputs distribution offset degree, edge sharpness, and colorimetric region imaging response characteristics into a joint evaluation model. By setting a multi-dimensional weighted fusion function, it comprehensively calculates the three types of indicators to generate a unified image quality judgment parameter. This judgment parameter reflects spatial stability, imaging sharpness, and color accuracy, and is compared with a preset quality threshold. When the image quality judgment parameter meets the preset constraints, the current frame is marked as a high-quality acquisition frame; otherwise, it is judged as a low-quality frame and feedback is sent to the user to prompt them to adjust the shooting posture. This realizes a multi-dimensional image quality evaluation mechanism based on "spatial offset calculation + edge response analysis + color consistency assessment".
[0055] Preferably, the procedure further includes the following steps after step S3 and before step S4: The target image is subjected to a no-reference image quality assessment. The target image is converted into a preset color space representation and scaled according to a preset size rule. The result is input into a pre-trained dual-branch convolutional neural network quality assessment model and outputs a quality score that represents the overall visual quality of the image. The quality score is compared with a preset quality threshold. If the quality score is greater than or equal to the preset quality threshold, the target image quality is determined to be qualified and step S4 is executed. Otherwise, the target image quality is determined to be unqualified and a re-sampling prompt is triggered, and archiving is not performed.
[0056] In one embodiment, a no-reference image quality assessment is performed on the target image. Specifically, the target image is converted from its original RGB representation to the CIELab color space to separate the luminance and chrominance components. Based on this, the target image is scale-normalized according to a preset size rule to eliminate the impact of resolution differences on subsequent evaluation results. The normalized image is input into a pre-trained, offline-trained, two-branch convolutional neural network quality assessment model. One branch extracts spatial domain-based texture sharpness features, while the other extracts frequency domain-based structure-preserving features. By fusing the features from the two branches, a comprehensive visual quality score for the corresponding target image is output. This quality score is a continuous numerical value used to characterize the overall quality status of the image in terms of sharpness, noise level, and structural distortion.
[0057] The quality score is compared with a preset quality threshold. The preset quality threshold can be set based on historical sample statistics. When the quality score is greater than or equal to the threshold, the target image is determined to meet the quality requirements and is passed to the subsequent step S4 for archiving. When the quality score is less than the threshold, the target image is determined to have problems such as blurring, abnormal exposure, or noise interference.
[0058] If an image is deemed unqualified, the system triggers a re-acquisition prompt mechanism, outputting a prompt message to the acquisition terminal through the human-computer interaction interface to guide the user to re-acquire the image; at the same time, the current image archiving process is terminated to prevent low-quality images from entering the subsequent processing link, thereby ensuring the quality consistency of the overall dataset.
[0059] Most importantly, a joint constraint analysis is performed on the distribution offset degree and the judgment parameters to form the validity judgment result of the current image, including: The degree of distribution offset is mapped to intervals to determine its corresponding offset level, and the offset level is used as a spatial stability constraint input to the decision framework. The decision parameters are analyzed hierarchically, and the effective components that characterize the boundary imaging quality are extracted. The decision parameters are then constrained and reconstructed based on the structural contour correlation. Under the constraints of spatial stability and decision parameters, a joint decision condition is constructed, and the consistency of the matching relationship between the two is verified to form a unified constraint decision state. The validity determination result of the current screen is output based on the constraint determination state.
[0060] In this embodiment, the process of "jointly constraining and analyzing the distribution offset degree and judgment parameters and outputting the image validity judgment result" is executed by the multi-constraint fusion judgment module in the smart terminal APP. After the user completes the image quality assessment, the current frame enters the validity judgment stage. The APP performs interval mapping processing on the distribution offset degree obtained in the previous steps, that is, divides the continuous numerical offset degree parameter into multiple offset level intervals according to the preset segmentation threshold, such as low offset level, medium offset level and high offset level, and converts the offset degree into a spatial stability constraint label through a lookup table mapping method. The spatial stability constraint is used to characterize the spatial jitter degree and positional stability of the current standardized reference sticker in the image, and serves as the input parameter for subsequent judgment.
[0061] The image quality assessment parameters are subjected to hierarchical analysis. By decomposing their internal components, effective components for characterizing boundary imaging quality are extracted, including edge sharpness, structural contour contrast, and color consistency of the colorimetric region. Based on the structural contour correlation, the components are reconstructed. That is, by combining the spatial correlation between the scale marking area and the colorimetric region in the structural contour, different quality indicators are normalized and weighted to form a standardized quality evaluation vector under structural constraints, thereby avoiding the impact of fluctuations in a single indicator on the overall judgment result.
[0062] The app constructs a joint judgment framework for spatial stability constraints and image quality constraints. It simultaneously inputs the spatial stability constraints corresponding to the offset level and the reconstructed quality evaluation vector into the joint judgment model. By setting multi-dimensional consistency verification rules, it performs consistency analysis on the matching relationship between the two. The consistency verification includes at least the judgment of the matching relationship between spatial stability and sharpness, the judgment of the collaborative constraint between spatial offset and color consistency, and the judgment of the linkage constraint between structural integrity and edge quality. When all constraint relationships meet the preset threshold conditions, a unified constraint judgment state is formed, and the current image is judged to be in a stable and high-quality effective acquisition state. Otherwise, it is judged to be invalid or in an adjustment state.
[0063] The app outputs the validity judgment result of the current image based on the constraint judgment state, and provides real-time feedback to the user through the interface, including prompts such as "can be shot", "need to adjust position" or "need to improve clarity", thereby realizing a joint judgment mechanism based on "offset level mapping + mass component reconstruction + multi-constraint consistency verification", and providing a unique judgment basis for subsequent automatic triggering of shooting control.
[0064] Of particular importance, step S3 includes: Based on the judgment result, the current image acquisition status is analyzed, and an acquisition trigger condition model corresponding to the judgment result is constructed. Under the constraints of the acquisition trigger condition model, the trigger control state of the image acquisition device is determined, and a trigger enable signal is generated when the determination result meets the preset conditions. When the trigger enable signal is activated, the image acquisition device is controlled to perform an image capture operation to obtain the corresponding target image; After image capture is completed, the target image is associated with the judgment result corresponding to the trigger enable signal.
[0065] In one embodiment, the validity determination result output by the preceding steps is analyzed in terms of state, and divided into condition-satisfied state, critical transition state, and condition-unsatisfied state. Based on different states, corresponding acquisition trigger condition models are established. Among them, the condition-satisfied state corresponds to the combination of conditions that the standardized reference sticker structure is complete, the spatial distribution is stable, and the imaging is clear. The critical transition state is used to characterize the situation where some conditions are close to the threshold boundary but have not been fully satisfied. The condition-unsatisfied state is used to characterize the situation where there are obvious deficiencies or deviations.
[0066] Under the constraints of the acquisition trigger condition model, when determining the trigger control status of the image acquisition device in real time, the validity determination result of the current frame can be jointly compared with the historical determination results of multiple consecutive frames. When several consecutive frames are in a state that meets the conditions, and the spatial position and sharpness variation of the standardized reference sticker are within a stable range, the current acquisition state is determined to have met the trigger condition, and a trigger enable signal is generated. When in a critical transition state, the trigger lock state is maintained and guidance information is output to prompt adjustment of the shooting angle or distance.
[0067] After generating the trigger enable signal, the trigger control of the image acquisition device is constrained, allowing image capture operations to be performed only during the effective period of the trigger enable signal; at the same time, a hold time window is set for the trigger enable signal. If the state is determined to fluctuate within this time window, the trigger enable signal is automatically revoked to avoid accidental triggering of shooting in an unstable state.
[0068] When the image acquisition device is controlled to perform image capture operation under the action of the trigger enable signal, the image acquisition device can be directly called to obtain one or more high-resolution images, and the image with the highest clarity and complete standardized reference sticker structure can be selected as the target image.
[0069] After image capture is completed, the target image is associated with the judgment result corresponding to the trigger enable signal, and the state information at the trigger time is recorded, including the spatial distribution state of the standardized reference sticker and the imaging quality parameters, so that the target image and the judgment basis at the time of its generation are correlated for subsequent data encapsulation and traceability processing.
[0070] In another embodiment, after completing step S3 to acquire the target image and before performing step S4 to encapsulate the data, the smart terminal application or cloud server performs a no-reference image quality assessment on the target image as a post-acquisition quality verification step.
[0071] Specifically, the target image is preprocessed by converting it into an RGB three-channel representation; scaling it proportionally to a fixed height preset value; adaptively calculating the width based on the original aspect ratio; and using an interpolation algorithm to maintain the geometric proportions.
[0072] The preprocessed image is input into a pre-trained dual-branch convolutional neural network quality assessment model. The model includes a sub-convolutional network branch for extracting local distortion-sensitive features, and a quality regression branch connected to the pre-trained feature extraction network. The model is trained on a no-reference image quality dataset and can output a continuous quality score in the range of 0 to 1 for a single image without relying on a reference image. The higher the score, the better the overall visual quality of the image.
[0073] The quality score output by the model is compared with a preset quality threshold: when the quality score is greater than or equal to the preset quality threshold, the target image quality is deemed acceptable, and the process of generating and encrypting the structured data package in step S4 is allowed; when the quality score is lower than the preset quality threshold, the target image quality is deemed unacceptable, a resampling prompt is output on the application interface, and the acquisition is marked as pending resampling, without triggering encryption archiving.
[0074] For batch acquisition scenarios, the system can perform the above evaluation on all target images in sequence, generate evaluation records containing image identifiers, quality scores, pass / fail markers and timestamps, and summarize the average score, highest score, lowest score and pass / fail number statistics, so that researchers can have a global grasp of the data quality collected from various locations on the server.
[0075] This implementation combines referenceless deep learning quality assessment with structural compliance determination under standardized reference sticker constraints, forming a dual quality assurance mechanism for structural compliance and overall image quality, further reducing the probability of low-quality, blurry, or distorted images entering the archiving system.
[0076] After the quality assessment is completed and the data is deemed qualified, the structured data package also includes a quality score and a qualification mark field, which are stored together with the image data, the validity judgment result, the collection time, and the location information to support subsequent data quality traceability and batch management.
[0077] Of particular importance, step S4 includes: Based on the acquired target image, the acquisition process data corresponding to the target image is retrieved from the basic dataset, and the association between the target image and the acquisition process data is established. The target image and acquisition process data under the constraint of the association relationship are uniformly organized and integrated according to the preset data structure to generate a structured data package containing image data and association information; The structured data packets are subjected to integrity processing to generate corresponding data identification information, and the structured data packets are encrypted based on preset encryption rules. The encrypted structured data packets are sent to the preset storage terminal for storage, and the storage feedback results are received to complete the archiving process.
[0078] In one embodiment, after the target image is acquired, the acquisition process data corresponding to the target image is retrieved from the local basic data set using the target image as an index. The acquisition process data includes at least the user business information entered by the intelligent control terminal in step S100, as well as the automatically acquired timestamp information, geographical location information, and device unique identification information. It also includes the unique identification information of the standardized reference sticker and its spatial distribution status data in the image, thereby forming a one-to-one correspondence between the target image and the acquisition process data. When establishing the association, the acquisition trigger event of the target image is used as the core node. The acquisition process data is aligned with the time axis. User input information, image acquisition time, GPS positioning information and device ID information are uniformly mapped to the same acquisition event identifier. A data association index structure is built based on the acquisition event identifier, so that the target image and the corresponding acquisition context information form an inseparable data link relationship logically. After completing the association relationship construction, the target image and acquisition process data under the association relationship constraints are uniformly organized and processed. According to the preset structured data template, various types of data are integrated into fields. Among them, image data is used as the main data unit, acquisition process data is used as the auxiliary metadata unit, and the scale information of the standardized reference sticker, color card information and its spatial distribution parameters are embedded as key verification fields into the structured data structure, thereby generating a structured data package containing image data, size calibration information, color reference information and acquisition process information. After generating the structured data packet, an integrity verification process is performed on the structured data packet. Specifically, this includes consistency checks on the integrity of image data, metadata fields, and standardized reference sticker information. Based on the check results, corresponding data identification information is generated. The data identification information is used to uniquely represent the collection event and its associated data status to ensure uniqueness and non-confusion in subsequent storage and tracing processes. After integrity processing is completed, the structured data packets are encrypted based on preset encryption rules. These rules include segmenting and encoding the structured data packets, generating a dynamic encryption key by combining device identification information and acquisition event identifiers, and encapsulating the image data and metadata as a whole to form an irreversible encrypted data carrier. The encrypted structured data packets are automatically sent to a preset cloud storage terminal for archiving through the intelligent control terminal. After receiving the data packets, the cloud storage terminal verifies and matches the data identification information and returns a storage confirmation feedback result. When the storage feedback result is received, it is written into the acquisition event record, thus completing a complete standardized image acquisition and archiving process.
[0079] The present invention also provides a standardized image acquisition system for performing the standardized image acquisition method described above, the standardized image acquisition system comprising: The standardized reference sticker setting module 101 is used to paste a standardized reference sticker on the acquisition area of the target acquisition object, so that the target acquisition object and the standardized reference sticker are within the same shooting angle. The application calls the camera of the smart terminal to acquire real-time image data containing the target acquisition object and the reference sticker, and generates the basic data set for the image acquisition process. The validity determination module 102 is used to perform frame-by-frame analysis of real-time image data, identify the structural contour of the standardized reference sticker; extract feature information from the standardized reference sticker based on the structural contour, and perform validity determination on the current image in combination with the spatial distribution of feature information in the image and image clarity, and output the determination result. The image acquisition module 103 is used to control the image acquisition triggering conditions based on the judgment result. When the judgment result meets the preset conditions, it controls the camera of the smart terminal to acquire the target image and obtain the target image. The storage and archiving module 104 is used to associate and bind the target image with the basic data set to generate a structured data packet containing image data and associated information; to encrypt the structured data packet and send it to the local machine and the cloud for archiving.
[0080] The image quality assessment module 105 is used to perform a no-reference image quality assessment on the target image and output a quality score; compare the quality score with a preset threshold to generate a quality pass / fail judgment result, and transmit the judgment result to the storage and archiving module 104 as one of the bases for whether to perform encrypted archiving.
[0081] A standardized image acquisition device includes: The intelligent terminal module is equipped with an application that performs image acquisition control, reference sticker feature recognition, validity determination and interactive prompts, and interacts with the remote server through a communication interface. The reference sticker module, located in the adjacent area of the target object, is a flexible attachment structure. The surface of the reference sticker module is divided into a colorimetric card array area and a scale area along its length. The colorimetric card array area integrates multiple standard color blocks arranged in a matrix to provide a reference for color reproduction and white balance calibration during image acquisition. The scale area integrates continuous and equally spaced scale markings to provide a physical reference for size measurement. The reference sticker module also includes a black and white grid area located between the colorimetric card array area and the scale area. The black and white grid area is composed of alternating black and white squares and is used to assist the algorithm in identifying image distortion and edge focus status during image acquisition. The surface integrates rulers, color charts, and black and white grids to provide size, color references, and deformation calibration standards; An image acquisition module, located on a mobile terminal, includes a camera and is used to acquire image data containing the target object and a reference sticker module, and transmit it to the intelligent control module. The data encapsulation module, integrated into the intelligent control module, is used to associate image data, validity judgment results, and acquisition time and location information to generate structured data packets; The communication storage module is electrically connected to the intelligent control module. It includes a communication interface and a storage unit, and is used to encrypt structured data packets and upload them to an external storage terminal for archiving via a communication network.
[0082] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. The invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A standardized image acquisition method, characterized in that, Applications used on smart terminals include the following steps: Step S1: Paste a standardized reference sticker on the collection area of the target object, so that the target object and the standardized reference sticker are within the same shooting angle. Use the application to call the camera of the smart terminal to obtain real-time image data containing the target object and the reference sticker, and generate the basic data set of the image acquisition process. Step S2: Analyze the real-time image data frame by frame to identify the structural outline of the standardized reference sticker; extract the feature information from the standardized reference sticker based on the structural outline, and combine the spatial distribution of the feature information in the image and the image clarity to perform a validity judgment on the current image and output the judgment result; Step S3: Based on the judgment result, control the image acquisition trigger condition. When the judgment result meets the preset condition, control the camera of the smart terminal to acquire the target image and obtain the target image. Step S4: Associate and bind the target image with the basic data set to generate a structured data packet containing image data and associated information; encrypt the structured data packet and send it to the local machine and cloud for archiving.
2. The standardized image acquisition method according to claim 1, characterized in that, Step S1 includes: Standardized reference stickers are set up within the collection area, and the relative positional relationship between the standardized reference stickers and the target collection object is defined, so that the standardized reference stickers form a spatial constraint benchmark within the field of view; Based on spatial constraints, the structural layout of the standardized reference sticker is preset, the spatial orientation characteristics of the standardized reference sticker are determined, and the target acquisition object and the standardized reference sticker are placed in the same field of view. The application calls the camera of the smart terminal to acquire real-time image data containing the target object and standardized reference stickers, and retains the spatial distribution of the standardized reference stickers in the image. The basic dataset is constructed by associating real-time image data with spatial distribution status.
3. The standardized image acquisition method according to claim 2, characterized in that, The application accesses the camera of the smart terminal to acquire real-time image data containing the target object and standardized reference stickers, and preserves the spatial distribution of the standardized reference stickers in the image, including: The application calls the camera of the smart terminal to acquire real-time image data containing the target object and standardized reference stickers; Locate the imaging area of the standardized reference sticker in real-time image data and extract its pixel distribution range in the image; The spatial positional relationship of the standardized reference sticker in the image is expressed based on the pixel distribution range, forming the corresponding spatial distribution state; The spatial distribution state is recorded and maintained in real-time image data, so that the spatial distribution state of the standardized reference sticker is stored synchronously with the image data.
4. The standardized image acquisition method according to claim 1, characterized in that, Step S2 includes: Frame-by-frame scanning is performed on real-time image data, structural contours are extracted based on the geometric boundary features of standardized reference stickers, and consistency comparison of structural contours is performed in consecutive frames to determine the existence status of standardized reference stickers. Under the constraint of structural contour consistency results, the feature information in the standardized reference sticker is analyzed within the area defined by the structural contour. The feature information includes the imaging response features of the scale marking area and the colorimetric area. Based on the relative positional relationship of the feature information within the structural contour, the structural integrity judgment result is output. Under the condition that the structural integrity judgment result is valid, the distribution offset degree is calculated based on the spatial distribution of feature information in the image, and the image quality judgment parameters are constructed by combining the edge sharpness of the structural contour and the imaging response characteristics of the colorimetric region. Joint constraint analysis is performed based on the distribution offset degree and image quality judgment parameters, and the validity judgment result corresponding to the current frame is output.
5. The standardized image acquisition method according to claim 4, characterized in that, Frame-by-frame scanning of real-time image data is performed, structural contours are extracted based on the geometric boundary features of standardized reference stickers, and consistency comparisons of structural contours are conducted in consecutive frames to determine the existence status of standardized reference stickers, including: Candidate boundaries are extracted from real-time image data and filtered according to the geometric constraints of standardized reference stickers to output the initial structural contour. Based on the initial structural contour, perform closure and boundary continuity checks, and eliminate contour segments that do not meet the constraints to obtain stable structural contours. The morphological parameters of the stable structural profile are matched in consecutive frames, and structural profile consistency data is output based on the size ratio and the degree of consistency of boundary morphology. The validity of stable structural profiles is confirmed based on structural profile consistency data, and the existence status of standardized reference stickers is output when consistency constraints are met.
6. The standardized image acquisition method according to claim 4, characterized in that, Under the constraint of structural contour consistency, the feature information in the standardized reference sticker is parsed within the area defined by the structural contour, and the structural integrity judgment result is output based on the relative positional relationship of the feature information within the structural contour, including: Feature information is extracted from the imaging area of the standardized reference sticker within the area defined by the structural contour. The feature information includes the imaging response features of the scale marking area and the colorimetric area. The scale marking area is identified based on the imaging response characteristics of the scale marking area, and the colorimetric area is identified based on the imaging response characteristics of the colorimetric area, generating the identification result. Based on the recognition results of the scale marking area and the colorimetric area, the relative positional relationship between the scale marking area and the colorimetric area within the structural contour is extracted; The relative positional relationship is matched with the preset positional relationship, and the existence status and positional relationship of the scale marking area and colorimetric area are determined based on the matching result; When both the scale marking area and the colorimetric area meet the preset positional relationship, the output structural integrity judgment result is valid; otherwise, the output structural integrity judgment result is invalid.
7. The standardized image acquisition method according to claim 4, characterized in that, Under the condition that the structural integrity judgment result is valid, the distribution offset degree is calculated based on the spatial distribution of feature information in the image, and the image quality judgment parameters are constructed by combining the edge sharpness of the structural contour and the imaging response characteristics of the colorimetric region, including: Under the condition that the structural integrity judgment result is valid, the spatial position of the scale marking area and colorimetric area in the image is extracted, the distribution coordinates of them relative to the structural contour are obtained, and the corresponding benchmark distribution is established based on the preset standard distribution relationship. Based on the matching of the distribution coordinates with the reference distribution, the spatial offset of the scale marking area and the colorimetric area is calculated, and the spatial offset is integrated to obtain the degree of distribution offset. Edge response information of the boundary region is extracted based on the structural contour, and edge sharpness is calculated. Based on a joint analysis of distribution offset and edge sharpness, a unified set of image quality judgment parameters is formed to characterize the image state.
8. The standardized image acquisition method according to claim 4, characterized in that, The steps following step S3 and before step S4 include: The target image is subjected to a no-reference image quality assessment. The target image is converted into a preset color space representation and scaled according to a preset size rule. The result is input into a pre-trained dual-branch convolutional neural network quality assessment model and outputs a quality score that represents the overall visual quality of the image. The quality score is compared with a preset quality threshold. If the quality score is greater than or equal to the preset quality threshold, the target image quality is determined to be qualified and step S4 is executed. Otherwise, the target image quality is determined to be unqualified and a re-sampling prompt is triggered, and archiving is not performed.
9. A standardized image acquisition system, characterized in that, For performing the standardized image acquisition method as described in claim 1, the standardized image acquisition system includes: The standardized reference sticker setting module is used to affix standardized reference stickers to the acquisition area of the target object, ensuring that the target object and the standardized reference sticker are within the same shooting angle. The application then uses the smart terminal's camera to acquire real-time image data containing both the target object and the reference sticker, generating the basic data set for the image acquisition process. The validity determination module is used to analyze real-time image data frame by frame, identify the structural outline of the standardized reference sticker; extract feature information from the standardized reference sticker based on the structural outline, and combine the spatial distribution of feature information in the image and image clarity to perform validity determination on the current image and output the determination result. The image acquisition module is used to control the image acquisition triggering conditions based on the judgment result. When the judgment result meets the preset conditions, it controls the camera of the smart terminal to acquire the target image and obtain the target image. The storage and archiving module is used to associate and bind target images with basic data sets to generate structured data packets containing image data and associated information; it also encrypts the structured data packets and sends them to local storage and the cloud for archiving. The image quality assessment module is used to perform a no-reference image quality assessment on the target image and output a quality score; compare the quality score with a preset threshold to generate a quality pass / fail judgment result, and transmit the judgment result to the storage and archiving module as one of the bases for whether to perform encrypted archiving.
10. A standardized image acquisition device, characterized in that, include: The intelligent terminal module is equipped with an application that performs image acquisition control, reference sticker feature recognition, validity determination and interactive prompts, and interacts with the remote server through a communication interface. The reference sticker module, located in the adjacent area of the target object, is a flexible attachment structure. The surface of the reference sticker module is divided into a colorimetric card array area and a scale area along its length. The colorimetric card array area integrates multiple standard color blocks arranged in a matrix to provide a reference for color reproduction and white balance calibration during image acquisition. The scale area integrates continuous and equally spaced scale markings to provide a physical reference for size measurement. The reference sticker module also includes a black and white grid area located between the colorimetric card array area and the scale area. The black and white grid area is composed of alternating black and white squares and is used to assist the algorithm in identifying image distortion and edge focus status during image acquisition. The surface integrates rulers, color charts, and black and white grids to provide size, color references, and deformation calibration standards; An image acquisition module, located on a mobile terminal, includes a camera and is used to acquire image data containing the target object and a reference sticker module, and transmit it to the intelligent control module. The data encapsulation module, integrated into the intelligent control module, is used to associate image data, validity judgment results, and acquisition time and location information to generate structured data packets; The communication storage module is electrically connected to the intelligent control module. It includes a communication interface and a storage unit, and is used to encrypt structured data packets and upload them to an external storage terminal for archiving via a communication network.