An intelligent and transparent management and control system and method for a detection center based on an internet of things

By using work order-driven task orchestration and improving the Mask2Former model, the problem of inconsistent data collection and supervision in the testing center was solved, realizing automatic collection and centralized management of testing data, and improving the transparency and traceability of the testing process.

CN122135370APending Publication Date: 2026-06-02HUBEI DAZE INTELLIGENT SENSING TECHNOLOGY CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI DAZE INTELLIGENT SENSING TECHNOLOGY CO LTD
Filing Date
2026-02-27
Publication Date
2026-06-02

Smart Images

  • Figure CN122135370A_ABST
    Figure CN122135370A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent and transparent management and control system and method for testing centers based on the Internet of Things (IoT). The system includes: constructing a testing center and platform architecture, establishing a correspondence between resource identifiers; parsing work order elements and distributing them to corresponding testing equipment; completing testing data acquisition and preprocessing to obtain standard testing data; constructing an improved Mask2Former model and outputting masking results; outputting semantic interpretation information based on zero-shot semantic segmentation; generating traceability records and testing reports, and completing security protection. This invention, by integrating the improved Mask2Former model and zero-shot semantic segmentation, achieves transparent supervision of the entire testing process, a closed-loop traceability system for data and video, and safe and reliable operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) management technology, and in particular to an intelligent and transparent management system and method for a testing center based on IoT. Background Technology

[0002] To implement the requirements for building a green, modern, digital, and intelligent supply chain and transparent laboratories, existing power material testing centers generally suffer from fragmented data acquisition links, heavy reliance on manual labor, and difficulties in data aggregation. The equipment at testing sites is complex and undergoes rapid updates, including integrated testing instruments with interface capabilities, intelligent instruments relying on host computer software for output, standalone digital instruments communicating via serial ports, and numerous analog instruments and gauges still requiring manual reading or data entry. Due to inconsistencies in equipment communication methods, output formats, and data field definitions, testing data is often scattered across various media, including equipment control processor interfaces, host computer databases, host computer file directories, gateway forwarding logs, and manual spreadsheets, resulting in a multi-source, heterogeneous, incomparable, and unmanageable situation. In actual operations, personnel need to switch between different systems, search for result files, copy and paste or manually transcribe key indicators, and then manually verify and summarize them. This easily leads to duplicate entries, omissions, inconsistencies in units and precision, and missing timestamps, making statistical analysis and quality traceability difficult to conduct. Due to the lack of a unified data model and centralized aggregation path based on work orders, the flow of test data is difficult to close. Report generation often relies on manual sorting and subjective judgment, which is not only inefficient but also amplifies human error, making it difficult to meet the testing center's needs for large-scale processing of high-frequency, multi-batch tasks.

[0003] Existing testing centers still have shortcomings in terms of transparent process supervision and integrated platform management. Video surveillance varies in terms of workstation coverage, clarity, viewing angle, and unified management. Some workstations have blind spots or video streams are not associated with testing tasks. As a result, although the testing operation process can be recorded, it lacks the ability to retrieve by work order, locate by step, and trace by event at the task level, making it difficult to achieve full-process supervision and evidence solidification. Equipment access lacks unified standards and adaptation strategies. Different protocols and different network forms coexist, making it difficult for the platform to carry out unified identity management, unified resource identification management, and unified data governance, forming information silos. Video streams and testing data are also difficult to align in the same spatiotemporal coordinates, resulting in a disconnect between data results and operation processes, making it difficult to achieve a traceable closed loop from task issuance, testing execution to result output.

[0004] Therefore, how to provide an intelligent and transparent management and control system and method for testing centers based on the Internet of Things is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose an intelligent and transparent management system and method for testing centers based on the Internet of Things (IoT). This invention comprehensively utilizes work order-driven task orchestration, IoT access to heterogeneous testing equipment and aggregation of standard testing data, video and data traceability linked by work order number, and terminal and boundary security protection technologies to form a closed-loop process from work order reception and issuance, automatic data acquisition and preprocessing, centralized storage, process visualization and monitoring to report generation. At the algorithmic level, it innovatively introduces an improved Mask2Former model and zero-shot semantic segmentation to achieve task-oriented mask query generation, cross-frame mask trajectory association, structured mask output, open-vocabulary semantic assignment, and parallel semantic interpretation output. This unifies the visualization and traceability of the testing process with structured analysis results and standard testing data. Compared with existing technologies, this invention can achieve unified access and automatic data acquisition and aggregation even under conditions of inconsistent device interface protocols and data formats, reducing errors from manual transcription and improving the transparency, traceability, and safe and reliable operation of the testing process.

[0006] According to an embodiment of the present invention, an intelligent and transparent management and control method for a testing center based on the Internet of Things includes:

[0007] Build the testing center and platform architecture, deploy wireless access points and wired networks, complete the access of video monitoring terminals for each testing station, and establish the corresponding relationship of resource identifiers;

[0008] Receive the inspection work order issued by the superior platform, parse the work order elements, and issue the work order task to the corresponding inspection equipment and video acquisition task based on the work station number;

[0009] The testing equipment is connected to the Internet of Things according to the interface type to complete the collection and preprocessing of testing data, obtain standard testing data, and centrally collect and store it according to the work order number;

[0010] An improved Mask2Former model is constructed. The mask query construction module generates task-oriented mask queries based on standard detection data. A cross-frame mask consistency module is introduced to perform cross-frame association to form mask trajectories. The structured expression mask decoder performs structured masking and outputs the mask results.

[0011] Based on zero-shot semantic segmentation, open-vocabulary semantic assignment is performed on the mask results. The category information in the mask results is explicitly decoupled and reconstructed to obtain the category semantic representation. Semantic interpretation information corresponding to the mask is generated in parallel output header through semantic interpretation information.

[0012] The structured mask results and semantic interpretation information are associated and summarized by work order number to generate a visual traceability record and test report of the detection process, thus completing the security protection of the terminal and data transmission.

[0013] Optionally, the step of completing the access of video monitoring terminals at each testing station and establishing the correspondence of resource identifiers includes:

[0014] Wireless access points and wired switching networks are deployed. The testing terminals, testing equipment network nodes, IoT gateways and video monitoring terminals at each testing station are all connected to the testing center network. A management platform server is deployed on the platform side to establish a communication connection with the testing center network.

[0015] The video monitoring terminals at each testing station are registered and network parameters are configured. The access information of each video stream is obtained and the corresponding video stream identifier is generated. The network nodes of each testing device and the IoT gateway are registered and access information is obtained and the corresponding device identifier is generated. At the same time, the station identifier corresponding to the testing station to which the device belongs is generated.

[0016] Establish a resource identifier mapping table, bind the workstation identifier with the device identifier and video stream identifier under the workstation and write them into the resource identifier mapping table, and establish an index relationship based on the resource identifier mapping table to support locating the device identifier and video stream identifier by workstation identifier and locating the workstation identifier by device identifier.

[0017] Optionally, the parsing process obtains the work order elements, and based on the workstation number, the work order tasks are distributed to the corresponding testing equipment and video acquisition tasks, including:

[0018] Receive the testing work order issued by the superior platform and parse it to obtain the work order elements, including the work order number, testing item identifier, sample identifier, workstation number, target equipment set, and work order time parameters;

[0019] Based on the workstation number, the set of device identifiers and video stream identifiers corresponding to the workstation number are retrieved from the resource identifier mapping table, and the target device set is matched with the set of device identifiers to determine the target device identifier;

[0020] The system generates a task, issues an instruction, and sends it to the network node or IoT gateway of the detection device corresponding to the target device identifier to trigger the detection task. It also generates a video acquisition instruction and sends it to the video monitoring terminal corresponding to the video stream identifier to start video acquisition and archiving. On the platform side, it writes a work order execution record and establishes the association between the work order number and the target device identifier, video stream identifier, and workstation number.

[0021] Optionally, the detection data includes detection index result data, process data, judgment and conclusion data, and equipment operation and status data.

[0022] Optionally, obtaining standard test data and centrally aggregating and storing it according to work order number includes:

[0023] Establish a list of testing equipment access based on equipment identification and determine the data acquisition channels according to interface type. The data acquisition channels include HTTP interface acquisition channel, RESTful interface acquisition channel, host computer database reading channel, host computer file system reading channel, RS-485 communication and IoT gateway protocol conversion channel, Bluetooth transmission channel, and mobile service terminal input channel.

[0024] Based on the work order number, the data acquisition channel corresponding to the target device identifier is triggered to perform data acquisition and generate raw data records, which contain detection data.

[0025] The raw data records are preprocessed to generate standard test data. The preprocessing includes data field extraction, field naming mapping, unit conversion, data type conversion, timestamp completion, missing value marking, outlier marking, and writing of data source channel identifier. The standard test data is centrally aggregated and stored according to work order number, and an association index is established between work order number and sample identifier, work station number, equipment identifier, and timestamp.

[0026] Optionally, the output mask result includes:

[0027] An improved Mask2Former model is constructed, including a mask query construction module, a cross-frame mask consistency module, and a structured expression mask decoder;

[0028] Input the work order elements and standard test data into the mask query construction module to generate a task-oriented mask query set. Perform mask candidate segmentation processing on the task-oriented mask query set and video frames, and output the mask candidate set.

[0029] The candidate mask set is input into the cross-frame mask consistency module. The cross-frame mask consistency module calculates the regional overlap, position displacement and appearance feature similarity of the candidate masks between adjacent video frames, and performs time alignment of the candidate masks matched across frames by combining the timestamps in the standard detection data, and establishes cross-frame association to generate mask trajectories.

[0030] The mask trajectory is input into the structure representation mask decoder, which generates a structured mask result from the mask trajectory. The structured mask result includes a binary mask, a set of contour points, bounding box coordinates, area parameters, and standard detection data index information corresponding to the structured mask result, and outputs the mask result.

[0031] Optionally, generating semantic interpretation information corresponding to the mask includes:

[0032] Receive the masking results and complete the zero-shot semantic segmentation initialization process, read the open vocabulary category set, and establish the configuration of the field category set and the field prompt text set;

[0033] A category information explicit decoupling and reconstruction module is constructed, including a field splitting and hint construction unit, a field semantic encoding unit, a region feature extraction unit, and a field matching and category semantic generation unit, wherein:

[0034] The field splitting and prompt building unit splits the open vocabulary category set into object field, location field, defect type field, morphology field, and severity field to form a field category set. For each field category set, a field prompt text set is generated according to the prompt template.

[0035] The field semantic encoding unit encodes the field prompt text set separately to obtain the field semantic vector set;

[0036] The region feature extraction unit extracts region appearance features, region location features, and work order context features from each structured mask in the masking result and generates a set of region feature vectors respectively.

[0037] The field matching and category semantic generation unit matches the set of regional feature vectors with the set of field semantic vectors to calculate the field matching score of each field, and determines the value of each field based on the field matching score and combines them to generate a category semantic representation of the structured mask;

[0038] The semantic interpretation information is generated by the parallel output header based on the category semantic representation and the structured mask result. The semantic interpretation information includes the category name, field splitting result and the contour point set of the structured mask. The semantic interpretation information is then bound to the structured mask result for output.

[0039] Optionally, the security protection for the terminal and data transmission includes:

[0040] Receive structured masking results and semantic interpretation information, read standard detection data aggregated and stored by work order number, establish the association relationship between structured masking results, semantic interpretation information and standard detection data based on work order number, and establish the mapping relationship between masking results and standard detection data records and video frames based on standard detection data index information and timestamps in structured masking results;

[0041] A visual traceability record of the detection process is generated by work order number and written into the traceability dataset. The traceability record includes the video stream identifier corresponding to the work station number, time alignment relationship, structured mask result, semantic interpretation information and event time period information. The standard detection data, structured mask result and semantic interpretation information are summarized by work order number to generate a detection report dataset and output the detection report.

[0042] Deploy security probes on the terminal side to perform terminal identification and operational status collection, and control access to mobile storage media. On the boundary side, enable firewalls or isolation devices, configure communication whitelists, and block communication outside the whitelist. Create security audit records for the transmission and storage of standard test data, structured mask results, semantic interpretation information, traceability datasets, and test reports.

[0043] An intelligent and transparent management and control system for a testing center based on the Internet of Things, according to an embodiment of the present invention, includes the following modules:

[0044] The architecture access module is used to deploy network and video access and establish resource identifier mapping relationships;

[0045] The work order parsing and distribution module is used to receive inspection work orders, parse work order elements, and distribute equipment tasks and video acquisition tasks.

[0046] The IoT data acquisition and preprocessing module is used for heterogeneous device access and detection data acquisition and preprocessing, generating standard detection data and aggregating and storing it.

[0047] Improve the Mask2Former processing module to generate task-oriented mask queries and cross-frame associations based on standard detection data, form mask trajectories, and output mask results;

[0048] The zero-sample semantic assignment module is used to perform open-vocabulary semantic assignment and explicit decoupling reconstruction of category information on the mask results, and output semantic explanation information.

[0049] The association summary and traceability module is used to associate and summarize structured mask and semantic interpretation information by work order number, and generate traceability records and detection reports;

[0050] The security protection module is used for terminal protection and data transmission security protection, and forms a security audit log.

[0051] The beneficial effects of this invention are:

[0052] This invention transforms the collection and processing of inspection data from decentralized and manual to automated and centralized, through a work order-driven task orchestration system and an IoT access system for heterogeneous inspection equipment. After receiving and parsing the inspection work order, the platform automatically distributes the task to the corresponding workstation and target equipment, simultaneously triggering the video acquisition task bound to the work order. This ensures that inspection tasks, equipment execution, and process images are interconnected under the same work order key. Addressing the challenges of varying equipment interface capabilities, inconsistent protocols, and diverse data formats, this invention performs access and preprocessing based on interface type. It standardizes inspection data (results, process data, judgment and conclusion data, equipment operation and status data, and traceability metadata) into standardized inspection data, centrally aggregating and storing it by work order number. This reduces omissions, errors, and inconsistencies caused by manual transcription and secondary entry. Furthermore, through resource identifier mapping and indexing, it achieves unified positioning and retrieval of workstations, equipment, and video streams, improving data aggregation efficiency and traceability under multi-workstation concurrent inspection.

[0053] This invention achieves a transformation from visualization to judgment and explanation in terms of transparent supervision and intelligent analysis. It improves the Mask2Former model by generating task-oriented mask queries through a work order-driven mask query construction module, forming mask trajectories using a cross-frame mask consistency module, and outputting structured mask results from a structured expression mask decoder. This provides statistically significant and searchable structured representations of abnormal areas and key processes. Zero-sample semantic segmentation assigns open-vocabulary semantic values ​​to the mask results, obtaining category semantic representations through explicit decoupling and reconstruction of category information. Semantic explanation information is then generated in parallel from the output header, generating explanation information bound to the mask, improving adaptability and review efficiency in the face of new defects and task switching scenarios. Combined with terminal security probes, boundary firewalls or isolation devices, and communication whitelist security protection measures, it ensures secure and reliable operation of terminals and data transmission, strengthening the transparency, traceability, and closed-loop management of the detection process. Attached Figure Description

[0054] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0055] Figure 1 This is a flowchart of an intelligent and transparent management and control method for a testing center based on the Internet of Things proposed in this invention;

[0056] Figure 2 This is a structural block diagram of the improved Mask2Former model for an intelligent and transparent management and control method for testing centers based on the Internet of Things proposed in this invention.

[0057] Figure 3This is a functional diagram of an intelligent and transparent management and control system for a testing center based on the Internet of Things proposed in this invention. Detailed Implementation

[0058] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0059] refer to Figure 1 and Figure 2 A method for intelligent and transparent management and control of testing centers based on the Internet of Things, comprising:

[0060] Build the testing center and platform architecture, deploy wireless access points and wired networks, complete the access of video monitoring terminals for each testing station, and establish the corresponding relationship of resource identifiers;

[0061] Receive the inspection work order issued by the superior platform, parse the work order elements, and issue the work order task to the corresponding inspection equipment and video acquisition task based on the work station number;

[0062] The testing equipment is connected to the Internet of Things according to the interface type to complete the collection and preprocessing of testing data, obtain standard testing data, and centrally collect and store it according to the work order number;

[0063] An improved Mask2Former model is constructed. The mask query construction module generates task-oriented mask queries based on standard detection data. A cross-frame mask consistency module is introduced to perform cross-frame association to form mask trajectories. The structured expression mask decoder performs structured masking and outputs the mask results.

[0064] Based on zero-shot semantic segmentation, open-vocabulary semantic assignment is performed on the mask results. The category information in the mask results is explicitly decoupled and reconstructed to obtain the category semantic representation. Semantic interpretation information corresponding to the mask is generated in parallel output header through semantic interpretation information.

[0065] The structured mask results and semantic interpretation information are associated and summarized by work order number to generate a visual traceability record and test report of the detection process, thus completing the security protection of the terminal and data transmission.

[0066] In this embodiment, the step of completing the access of video monitoring terminals at each testing station and establishing the correspondence of resource identifiers includes:

[0067] Wireless access points and wired switching networks are deployed. The testing terminals, testing equipment network nodes, IoT gateways and video monitoring terminals at each testing station are all connected to the testing center network. A management platform server is deployed on the platform side to establish a communication connection with the testing center network.

[0068] The video monitoring terminals at each testing station are registered and configured with network parameters. Access information for each video stream is obtained, and a corresponding video stream identifier is generated. Network nodes for each testing device are registered with the IoT gateway. Device access information is obtained, and a corresponding device identifier is generated. Simultaneously, a station identifier corresponding to the testing station to which the device belongs is generated. The network parameter configuration specifically includes:

[0069] Network parameter configuration includes setting static IP addresses, subnet masks, default gateways, and DNS servers for the video surveillance terminals and detection equipment network nodes respectively; configuring a time synchronization address consistent with the management and control platform and completing clock calibration; setting access port numbers, bitstream parameters, and transmission protocol parameters; configuring device names, workstation numbers, and resource identifiers; and submitting registration after completing connectivity tests and authentication information verification.

[0070] A resource identifier mapping table is established to bind workstation identifiers with device identifiers and video stream identifiers under the workstation and write them into the resource identifier mapping table. An index relationship is established based on the resource identifier mapping table to support locating device identifiers and video stream identifiers by workstation identifiers, and locating workstation identifiers by device identifiers. Specifically, the resource identifier mapping table is as follows:

[0071] The resource identifier mapping table is a structured data table maintained on the platform side. It is used to record the unique identifiers and affiliations of resources in the testing center. It includes workstation identifier fields, device identifier fields, video stream identifier fields, and corresponding access information fields. Each record represents the binding relationship between the workstation and the devices and video streams under the workstation, and includes the device access address and port, video stream access address and channel number, registration time and status information, so as to quickly retrieve the associated device identifier and video stream identifier by workstation identifier, and to retrieve the workstation identifier by device identifier.

[0072] In this embodiment, the step of parsing the work order elements and issuing the work order task to the corresponding detection equipment and video acquisition task based on the workstation number includes:

[0073] The system receives and parses inspection work orders from the superior platform to obtain work order elements. These elements include the work order number, inspection item identifier, sample identifier, workstation number, target equipment set, and work order time parameter. Specifically, the process of parsing and obtaining these work order elements involves:

[0074] The process of parsing and obtaining work order elements includes format recognition and field parsing of the received work order data packet, extracting the work order number, inspection item identifier, sample identifier, workstation number, target equipment set, and work order time parameter according to field mapping, and performing field integrity verification, type verification, and value validity verification on the extracted results. If the work order contains nested lists, the equipment list and inspection item list are traversed and parsed to generate the target equipment set. Finally, the parsed work order elements are written into the work order element record and the corresponding work order index is generated.

[0075] Based on the workstation number, the set of device identifiers and video stream identifiers corresponding to the workstation number are retrieved from the resource identifier mapping table, and the target device set is matched with the set of device identifiers to determine the target device identifier;

[0076] The task generation process involves issuing instructions and sending them to the network node or IoT gateway of the detection device corresponding to the target device identifier to trigger the detection task. A video acquisition instruction is generated and sent to the video monitoring terminal corresponding to the video stream identifier to initiate video acquisition and archiving. On the platform side, a work order execution record is written, and the association between the work order number and the target device identifier, video stream identifier, and workstation number is established.

[0077] The task issuance instruction is generated by the platform based on the work order element and resource identifier mapping table. First, the target device identifier and access channel type are determined. Then, the work order number, target device identifier, test item parameters, sample identifier, workstation number, execution time parameters and return address are filled in according to the unified instruction template. A unique instruction number and timestamp are generated for the instruction. The instruction is serialized into the message format corresponding to the device channel and written into the queue to be issued, forming a task issuance instruction for the network node or IoT gateway of the test device.

[0078] The video acquisition instruction is generated by the platform based on the work order elements, video stream identifier, and work order time parameters. The work order number, video stream identifier, workstation number, archive location identifier, and return address are filled in according to the video acquisition instruction template. A unique instruction number and timestamp are generated for the instruction. The instruction is serialized into a control command format supported by the video monitoring terminal and written into the acquisition scheduling queue to form a video acquisition instruction for starting acquisition and archiving.

[0079] In this embodiment, the detection data includes detection index result data, process data, judgment and conclusion data, and equipment operation and status data.

[0080] In this embodiment, obtaining standard test data and centrally aggregating and storing it according to work order number includes:

[0081] Establish a list of testing equipment access based on equipment identification and determine the data acquisition channels according to interface type. The data acquisition channels include HTTP interface acquisition channel, RESTful interface acquisition channel, host computer database reading channel, host computer file system reading channel, RS-485 communication and IoT gateway protocol conversion channel, Bluetooth transmission channel, and mobile service terminal input channel.

[0082] Based on the work order number, the data acquisition channel corresponding to the target device identifier is triggered to perform data acquisition and generate raw data records, which contain detection data.

[0083] The raw data records are preprocessed to generate standard test data. The preprocessing includes data field extraction, field naming mapping, unit conversion, data type conversion, timestamp completion, missing value marking, outlier marking, and writing of data source channel identifier. The standard test data is centrally aggregated and stored according to work order number, and an association index is established between work order number and sample identifier, work station number, equipment identifier, and timestamp.

[0084] In this embodiment, the output mask result includes:

[0085] An improved Mask2Former model is constructed, including a mask query construction module, a cross-frame mask consistency module, and a structured representation mask decoder, wherein:

[0086] Based on the original Mask2Former's backbone feature extraction, pixel decoding, and query decoding framework, the learnable query input end is replaced by a mask query construction module that receives work order elements and standard detection data and outputs a task-oriented mask query set. A cross-frame mask consistency module is added to the mask prediction output end to match and associate the mask candidates of consecutive video frames according to the timestamp and generate a mask trajectory as a time-series input.

[0087] The original mask output header is replaced or expanded into a structured representation mask decoder, which outputs the contour point set, bounding box, area and orientation structured fields while outputting the binary mask, and carries standard detection data index information, forming an improved Mask2Former model that includes task-oriented query, cross-frame trajectory and structured output.

[0088] The work order elements and standard detection data are input into the mask query construction module to generate a task-oriented mask query set. The task-oriented mask query set and video frames are then processed for mask candidate segmentation, and a mask candidate set is output, in which:

[0089] The generated task-oriented mask query set is specifically as follows:

[0090] After receiving the work order elements and standard test data, the mask query construction module extracts and encodes the fields of test item identifier, sample identifier, workstation number, target equipment set, and test index results, judgments and conclusions, equipment operation and status in the standard test data to generate a task context vector. Based on the task context vector, it determines the number of queries and query category slots, expands the task context vector according to the slots to generate query seed vectors, and concatenates or replaces the corresponding positions with the preset query template vectors. It outputs a task-oriented mask query set containing each query number, the corresponding task field pointer, and the query vector.

[0091] The preset query template is a set of query slots and initialization vectors defined on the model side. It is used to specify the organization and semantics of task-oriented mask queries. Each template includes a template number, slot type identifier, number of slots, initial query vector for each slot, and binding relationship between the slot and work order field or standard detection data field.

[0092] The mask candidate segmentation process is as follows: feature extraction is performed on video frames to obtain pixel-level feature maps, and the task-oriented mask query is used as the input query decoder to interact with the pixel features. The output is the mask embedding and class embedding corresponding to each query. The mask prediction head maps the mask embedding to a pixel-level binary mask and calculates the confidence score to complete the generation of candidate masks for each query. The candidate masks are filtered according to the confidence score threshold, and overlapping candidates are deduplicated and merged to obtain a mask candidate set containing the mask region, confidence score, corresponding query number and timestamp. The confidence score is obtained by combining the class prediction probability of the mask candidate and the average probability that the pixel in the mask region belongs to the foreground. The confidence score threshold is set to 0.5.

[0093] The candidate mask set is input into the cross-frame mask consistency module. This module calculates the region overlap, positional displacement, and appearance feature similarity of the candidate masks between adjacent video frames. It also combines the timestamps from the standard detection data to perform time alignment on the candidate masks matched across frames, establishing cross-frame associations and generating mask trajectories.

[0094] The calculation of the region overlap, position displacement, and appearance feature similarity of the candidate masks is as follows: For any pair of candidate masks in two adjacent frames, the binary masks are first aligned at the pixel level and the region overlap is calculated. The region overlap is the ratio of the number of pixels in the intersection of the two mask regions to the number of pixels in the union. Then, the center point coordinates of the two candidate masks are calculated. The center point coordinates are the average of the coordinates of all foreground pixels in the mask region. The position displacement is the Euclidean distance of the difference between the two center point coordinates. The appearance feature vector is obtained by feature aggregation according to the mask region. The appearance feature similarity is the cosine similarity of the two appearance feature vectors. A matching feature consisting of overlap, displacement, and similarity is formed for each pair of candidate masks.

[0095] The process of establishing cross-frame associations to generate mask trajectories is as follows: Based on the timestamps of standard detection data, adjacent frame pairing windows are determined. The matching features of candidate masks of two frames within the same window are calculated pairwise, and pairs with an overlap of less than a threshold are filtered out. For the remaining pairs, a comprehensive matching score is calculated and sorted from high to low scores. A one-to-one matching method is used to select the candidate mask of the next frame with the highest score and which is not occupied for each candidate mask of the previous frame to establish cross-frame associations. The continuously established associations are concatenated in chronological order to form mask trajectories. The comprehensive matching score is obtained by averaging the regional overlap and appearance feature similarity, and then multiplying it by the complementary value after the position displacement is normalized. The overlap threshold is set to 0.3.

[0096] The mask trajectory is input into the structure representation mask decoder. The structure representation mask decoder generates a structured mask result from the mask trajectory. The structured mask result includes a binary mask, a set of contour points, bounding box coordinates, area parameters, and standard detection data index information corresponding to the structured mask result. The mask result is then output. Specifically, generating a structured mask result from the mask trajectory involves:

[0097] Traverse each frame mask according to the trajectory identifier and use the corresponding binary mask as the basic output. At the same time, extract the boundary of the binary mask to obtain the contour point set, find the minimum bounding rectangle of the contour point set to obtain the bounding box coordinates, and count the number of foreground pixels of the binary mask and combine it with the pixel physical calibration coefficient to obtain the area parameter. Write the trajectory identifier, frame number, timestamp and standard detection data index information carried by the mask trajectory into the structured mask result record.

[0098] In this embodiment, generating semantic interpretation information corresponding to the mask includes:

[0099] Receive the masking results and complete the zero-shot semantic segmentation initialization process, read the open vocabulary category set, and establish the configuration of the field category set and the field prompt text set;

[0100] A category information explicit decoupling and reconstruction module is constructed, including a field splitting and hint construction unit, a field semantic encoding unit, a region feature extraction unit, and a field matching and category semantic generation unit, wherein:

[0101] The field splitting and prompt construction unit splits the open vocabulary category set into object fields, location fields, defect type fields, morphology fields, and severity fields to form field category sets. For each field category set, a field prompt text set is generated according to the prompt template. The specific process of generating the field prompt text set is as follows:

[0102] After splitting the open vocabulary category set into object field, location field, defect type field, morphology field, and severity field, a field vocabulary list is created for each field. For each field value in the field vocabulary list, multiple prompt texts are generated according to the prompt template. The prompt template includes field value and area class template, and field value class template that appears in the workstation scenario. The multiple prompt texts corresponding to the same field value are summarized to form a field prompt text set.

[0103] The field semantic encoding unit encodes the field prompt text set separately to obtain the field semantic vector set;

[0104] The region feature extraction unit extracts region appearance features, region location features, and work order context features from each structured mask in the masking result and generates a set of region feature vectors for each. Specifically, the extraction of region appearance features, region location features, and work order context features involves:

[0105] The appearance features of the region are obtained by mask pooling the binary mask corresponding to the structured mask on the feature map of the video frame. The location features of the region are formed by normalizing the bounding box coordinates, contour point set and area parameters to form the location feature vector. The work order context features are obtained by encoding the detection conclusion and equipment status field corresponding to the detection item identifier, sample identifier, work station number, target equipment set and standard detection data index information in the work order elements.

[0106] The field matching and category semantic generation unit matches the set of region feature vectors with the set of field semantic vectors to calculate the field matching score for each field. Based on the field matching score, it determines the value of each field and combines them to generate a category semantic representation of the structured mask, where:

[0107] The process of obtaining the field matching score for each field is as follows: calculate the cosine similarity between the region appearance feature vector and the set of field semantic vectors of the current field to obtain a set of similarity scores, and take the maximum similarity of multiple prompt texts corresponding to the same field value as the matching score of the field value.

[0108] The generation of the category semantic representation of the structured mask is specifically as follows: the field value with the highest score in the field matching score table of each field is selected as the predicted value of the field, and the values ​​of the object field, part field, defect type field, morphology field and severity field are concatenated in order to generate category labels. At the same time, the highest score of each field and the corresponding field value are written into the category semantic record to obtain the category semantic representation of the structured mask.

[0109] When constructing the explicit decoupling and reconstruction module for category information, the module is divided into a field splitting and prompting construction unit, a field semantic encoding unit, a region feature extraction unit, and a field matching and category semantic generation unit. The field splitting and prompting construction unit splits the open vocabulary category set into a field category set and generates a corresponding field prompt text set. The field semantic encoding unit encodes the field prompt text set to obtain a field semantic vector set. The region feature extraction unit extracts the region appearance features, region location features, and work order context features for each structured mask and generates a region feature vector set. The field matching and category semantic generation unit matches the region feature vector with the field semantic vector to obtain a field matching score and determines the value of each field accordingly and combines them to form a category semantic representation.

[0110] Semantic explanation information is generated through a parallel output header based on category semantic representation and structured mask results. This semantic explanation information includes category names, field splitting results, and the contour point set of the structured mask. The semantic explanation information is then bound to the structured mask results for output. Specifically, the generation of semantic explanation information involves:

[0111] The system reads the category name and extracts the values ​​of the object field, location field, defect type field, morphology field, and severity field to form a field splitting result. It reads the contour point set and the corresponding trajectory identifier, frame number, and timestamp from the structured mask result. It serializes the category name, field splitting result, and contour point set into an interpretation record and generates an interpretation identifier for the interpretation record. It writes the binding relationship between the interpretation identifier and the corresponding structured mask identifier into the output data structure to realize the binding output of semantic interpretation information and structured mask result.

[0112] In this embodiment, the security protection for terminal and data transmission includes:

[0113] Receive structured masking results and semantic interpretation information, read standard detection data aggregated and stored by work order number, establish the association relationship between structured masking results, semantic interpretation information and standard detection data based on work order number, and establish the mapping relationship between masking results and standard detection data records and video frames based on standard detection data index information and timestamps in structured masking results;

[0114] A visual traceability record of the detection process is generated by work order number and written into the traceability dataset. The traceability record includes the video stream identifier corresponding to the work station number, time alignment relationship, structured mask result, semantic interpretation information and event time period information. The standard detection data, structured mask result and semantic interpretation information are summarized by work order number to generate a detection report dataset and output the detection report.

[0115] Deploy security probes on the terminal side to perform terminal identification and operational status collection, and control access to mobile storage media. On the boundary side, enable firewalls or isolation devices, configure communication whitelists, and block communication outside the whitelist. Create security audit records for the transmission and storage of standard test data, structured mask results, semantic interpretation information, traceability datasets, and test reports.

[0116] refer to Figure 3 An intelligent and transparent management and control system for testing centers based on the Internet of Things includes the following modules:

[0117] The architecture access module is used to deploy network and video access and establish resource identifier mapping relationships;

[0118] The work order parsing and distribution module is used to receive inspection work orders, parse work order elements, and distribute equipment tasks and video acquisition tasks.

[0119] The IoT data acquisition and preprocessing module is used for heterogeneous device access and detection data acquisition and preprocessing, generating standard detection data and aggregating and storing it.

[0120] Improve the Mask2Former processing module to generate task-oriented mask queries and cross-frame associations based on standard detection data, form mask trajectories, and output mask results;

[0121] The zero-sample semantic assignment module is used to perform open-vocabulary semantic assignment and explicit decoupling reconstruction of category information on the mask results, and output semantic explanation information.

[0122] The association summary and traceability module is used to associate and summarize structured mask and semantic interpretation information by work order number, and generate traceability records and detection reports;

[0123] The security protection module is used for terminal protection and data transmission security protection, and forms a security audit log.

[0124] Example 1:

[0125] To verify the feasibility of this invention in practice, it was applied to a power material quality inspection center, which undertakes the incoming inspection and sampling of various materials, including high-voltage electrical appliances, hardware fasteners, and insulating materials. The center simultaneously operates new intelligent testing instruments and multiple older, stand-alone digital instruments and analog measuring tools. Two prominent problems were encountered during the upgrade process, which aimed to improve work order-driven, data-centralized, transparent, and traceable processes: First, the equipment interfaces and data formats were inconsistent. Test results were scattered across instrument interfaces, host computer databases, file directories, and manual forms, requiring operators to repeatedly copy and re-enter data, leading to omissions, errors, inconsistent unit precision, and version inconsistencies within the same work order. Second, although the testing process was video-enabled, it was impossible to quickly locate key time periods according to the work order. In case of disputes, it was difficult to align the operation process with the test results for evidence collection. Third, the outdated terminal system introduced risks of weak passwords, unauthorized access, and the introduction of data via mobile media, affecting safe and reliable operation and audit traceability.

[0126] In the pilot application, the testing center first completed the deployment of wireless access points and wired networks and connected them to cameras at each workstation, establishing a resource identification mapping for workstations, equipment, and video streams, enabling the platform to locate equipment and video resources by workstation. After the upper-level platform issues a work order, the management platform parses the work order elements and automatically issues tasks according to the workstation number, while simultaneously triggering video acquisition and archiving. On the data side, IoT access and acquisition preprocessing are completed according to interface type. Instruments that can be directly connected acquire data through the interface, host computer devices read data from the database or file directory, serial port devices acquire data through IoT gateway protocol conversion, and analog measuring instruments input data through mobile terminals. Finally, standard test data is uniformly generated and centrally aggregated and stored according to the work order number. On the transparent supervision side, key frames from workstation videos are fed into the algorithm service. The improved Mask2Former model combines work order elements and standard detection data to generate task-oriented mask queries, establish mask trajectories between consecutive frames, and output structured masks. Zero-shot semantic segmentation assigns open-vocabulary semantic values ​​to the structured masks, decomposes them into category semantic representations through category fields, and outputs verifiable semantic explanation information in parallel. Finally, the structured masks and explanation information are written into the traceability record by work order number and participate in report generation.

[0127] After pilot operation, the platform processed hundreds of work orders, connecting to various devices including smart interface instruments, host computer instruments, serial port devices, and analog measuring tools. Standard testing data records reached millions, and were aligned with and indexed to timestamps in the video streams. On-site statistics showed a significant decrease in manual intervention in data collection and entry. The time spent by personnel searching for, copying, verifying, and backfilling data was significantly reduced, shortening the overall report generation cycle. In dispute review scenarios, key segments of the workstation video could be directly located using the work order number, and structured masks and semantic interpretation information could be overlaid to complete evidence collection. Review response time was reduced from lengthy document searches to quick location and verification. On the security side, terminal security probes and boundary whitelist policies intercepted abnormal logins, unauthorized access, and unauthorized writing to removable media, and retained audit records. No data leakage or business interruption occurred, meeting the testing center's requirements for secure and reliable operation and traceable closed-loop management.

[0128] Table 1 Comparison of Intelligent and Transparent Management Solutions for Testing Centers

[0129] Comparison indicators Manual process Single Interface Acquisition No video captured Video without intelligence Baseline model Invention Solution Access duration (min / unit) 180 95 70 72 68 42 Missed entry rate (%) 4.8 2.9 1.7 1.8 1.5 0.6 Error rate (%) 3.6 2.1 1.2 1.3 1.1 0.4 Closed-loop delay (min) 210 135 92 98 85 46 Video retrieval(s) 420 360 340 110 95 28 Abnormal detection rate (%) 52 61 66 70 78 89 Explanation of availability (%) 15 22 25 38 54 83 Blocking rate (%) 58 67 71 73 76 92

[0130] As shown in Table 1, in terms of data access and data quality indicators, the proposed solution is optimal overall in terms of access time, omission rate, and error rate. The access time per device is 42 minutes, which is approximately 76.7% lower than the manual process of 180 minutes per device. It is also significantly better than single-interface acquisition (95 minutes per device), acquisition without video (70 minutes per device), video without intelligence (72 minutes per device), and the baseline model (68 minutes per device). Regarding data integrity and accuracy, the proposed solution has an omission rate of 0.6% and an error rate of 0.4%, which are 87.5% and 88.9% lower than the manual process (4.8% and 3.6%, respectively). These rates are also better than the baseline model (1.5% and 1.1%), indicating that after the heterogeneous device access and standardized preprocessing links are improved, the omissions and errors caused by manual transcription and secondary entry are significantly suppressed.

[0131] From the perspective of business closed-loop and transparent retrieval indicators, the solution of this invention shows the most significant improvement in closed-loop latency and video retrieval. The closed-loop latency is 46 minutes, which is about 78.1% shorter than the manual process of 210 minutes. It is also better than the single-interface collection of 135 minutes, collection without video of 92 minutes, video without intelligence of 98 minutes, and baseline model of 85 minutes. This reflects the improvement in process efficiency brought about by the automatic distribution, centralized aggregation, and correlation summary capabilities under the work order drive. The video retrieval time is 28 seconds, which is about 93.3% shorter than the manual process of 420 seconds. It is also significantly better than the video without intelligence of 110 seconds and baseline model of 95 seconds. This indicates that after establishing a video index by work order number and superimposing structured results, it is possible to locate key time periods more quickly and support traceability and review.

[0132] From the perspective of intelligent recognition and security performance indicators, the solution of this invention forms a combined advantage in anomaly detection rate, interpretability, and blocking rate. The anomaly detection rate reaches 89%, which is higher than that of manual process (52%), single interface acquisition (61%), acquisition without video (66%), video without intelligence (70%), and baseline model (78%). This indicates that after introducing the improved Mask2Former trajectory-based structured segmentation and zero-sample semantic segmentation field decoupling semantic assignment, the detection capability of abnormal areas and key events is stronger. The interpretability rate is 83%, which is 29 percentage points higher than that of the baseline model (54%) and also much higher than that of manual process (15%), reflecting the effectiveness of parallel output semantic interpretation information in review and reporting scenarios. The security blocking rate is 92%, which is higher than that of manual process (58%) and baseline model (76%), indicating that the terminal-side and boundary-side security protection strategies can effectively improve risk interception capabilities in engineering implementation.

[0133] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent and transparent management and control of a testing center based on the Internet of Things, characterized in that, include: Build the testing center and platform architecture, deploy wireless access points and wired networks, complete the access of video monitoring terminals for each testing station, and establish the corresponding relationship of resource identifiers; Receive the inspection work order issued by the superior platform, parse the work order elements, and issue the work order task to the corresponding inspection equipment and video acquisition task based on the work station number; The testing equipment is connected to the Internet of Things according to the interface type to complete the collection and preprocessing of testing data, obtain standard testing data, and centrally collect and store it according to the work order number; An improved Mask2Former model is constructed. The mask query construction module generates task-oriented mask queries based on standard detection data. A cross-frame mask consistency module is introduced to perform cross-frame association to form mask trajectories. The structured expression mask decoder performs structured masking and outputs the mask results. Based on zero-shot semantic segmentation, open-vocabulary semantic assignment is performed on the mask results. The category information in the mask results is explicitly decoupled and reconstructed to obtain the category semantic representation. Semantic interpretation information corresponding to the mask is generated in parallel output header through semantic interpretation information. The structured mask results and semantic interpretation information are associated and summarized by work order number to generate a visual traceability record and test report of the detection process, thus completing the security protection of the terminal and data transmission.

2. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The process of completing the access of video monitoring terminals at each testing station and establishing the corresponding relationship of resource identifiers includes: Wireless access points and wired switching networks are deployed. The testing terminals, testing equipment network nodes, IoT gateways and video monitoring terminals at each testing station are all connected to the testing center network. A management platform server is deployed on the platform side to establish a communication connection with the testing center network. The video monitoring terminals at each testing station are registered and network parameters are configured. The access information of each video stream is obtained and the corresponding video stream identifier is generated. The network nodes of each testing device and the IoT gateway are registered and access information is obtained and the corresponding device identifier is generated. At the same time, the station identifier corresponding to the testing station to which the device belongs is generated. Establish a resource identifier mapping table, bind the workstation identifier with the device identifier and video stream identifier under the workstation and write them into the resource identifier mapping table, and establish an index relationship based on the resource identifier mapping table to support locating the device identifier and video stream identifier by workstation identifier and locating the workstation identifier by device identifier.

3. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The parsing process yields the work order elements, and based on the workstation number, the work order tasks are distributed to the corresponding testing equipment and video acquisition tasks, including: Receive the testing work order issued by the superior platform and parse it to obtain the work order elements, including the work order number, testing item identifier, sample identifier, workstation number, target equipment set, and work order time parameters; Based on the workstation number, the set of device identifiers and video stream identifiers corresponding to the workstation number are retrieved from the resource identifier mapping table, and the target device set is matched with the set of device identifiers to determine the target device identifier; The system generates a task, issues an instruction, and sends it to the network node or IoT gateway of the detection device corresponding to the target device identifier to trigger the detection task. It also generates a video acquisition instruction and sends it to the video monitoring terminal corresponding to the video stream identifier to start video acquisition and archiving. On the platform side, it writes a work order execution record and establishes the association between the work order number and the target device identifier, video stream identifier, and workstation number.

4. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The detection data includes detection index result data, process data, judgment and conclusion data, and equipment operation and status data.

5. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The process of obtaining standard test data and centrally aggregating and storing it according to work order number includes: Establish a list of testing equipment access based on equipment identification and determine the data acquisition channels according to interface type. The data acquisition channels include HTTP interface acquisition channel, RESTful interface acquisition channel, host computer database reading channel, host computer file system reading channel, RS-485 communication and IoT gateway protocol conversion channel, Bluetooth transmission channel, and mobile service terminal input channel. Based on the work order number, the data acquisition channel corresponding to the target device identifier is triggered to perform data acquisition and generate raw data records, which contain detection data. The raw data records are preprocessed to generate standard test data. The preprocessing includes data field extraction, field naming mapping, unit conversion, data type conversion, timestamp completion, missing value marking, outlier marking, and writing of data source channel identifier. The standard test data is centrally aggregated and stored according to work order number, and an association index is established between work order number and sample identifier, work station number, equipment identifier, and timestamp.

6. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The output mask result includes: An improved Mask2Former model is constructed, including a mask query construction module, a cross-frame mask consistency module, and a structured expression mask decoder; Input the work order elements and standard test data into the mask query construction module to generate a task-oriented mask query set. Perform mask candidate segmentation processing on the task-oriented mask query set and video frames, and output the mask candidate set. The candidate mask set is input into the cross-frame mask consistency module. The cross-frame mask consistency module calculates the regional overlap, position displacement and appearance feature similarity of the candidate masks between adjacent video frames, and performs time alignment of the candidate masks matched across frames by combining the timestamps in the standard detection data, and establishes cross-frame association to generate mask trajectories. The mask trajectory is input into the structure representation mask decoder, which generates a structured mask result from the mask trajectory. The structured mask result includes a binary mask, a set of contour points, bounding box coordinates, area parameters, and standard detection data index information corresponding to the structured mask result, and outputs the mask result.

7. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The generation of semantic interpretation information corresponding to the mask includes: Receive the masking results and complete the zero-shot semantic segmentation initialization process, read the open vocabulary category set, and establish the configuration of the field category set and the field prompt text set; A category information explicit decoupling and reconstruction module is constructed, including a field splitting and hint construction unit, a field semantic encoding unit, a region feature extraction unit, and a field matching and category semantic generation unit, wherein: The field splitting and prompt building unit splits the open vocabulary category set into object field, location field, defect type field, morphology field, and severity field to form a field category set. For each field category set, a field prompt text set is generated according to the prompt template. The field semantic encoding unit encodes the field prompt text set separately to obtain the field semantic vector set; The region feature extraction unit extracts region appearance features, region location features, and work order context features from each structured mask in the masking result and generates a set of region feature vectors respectively. The field matching and category semantic generation unit matches the set of regional feature vectors with the set of field semantic vectors to calculate the field matching score of each field, and determines the value of each field based on the field matching score and combines them to generate a category semantic representation of the structured mask; The semantic interpretation information is generated by the parallel output header based on the category semantic representation and the structured mask result. The semantic interpretation information includes the category name, field splitting result and the contour point set of the structured mask. The semantic interpretation information is then bound to the structured mask result for output.

8. The intelligent and transparent management and control method for a testing center based on the Internet of Things according to claim 1, characterized in that, The security protection for the terminal and data transmission includes: Receive structured masking results and semantic interpretation information, read standard detection data aggregated and stored by work order number, establish the association relationship between structured masking results, semantic interpretation information and standard detection data based on work order number, and establish the mapping relationship between masking results and standard detection data records and video frames based on standard detection data index information and timestamps in structured masking results; A visual traceability record of the detection process is generated by work order number and written into the traceability dataset. The traceability record includes the video stream identifier corresponding to the work station number, time alignment relationship, structured mask result, semantic interpretation information and event time period information. The standard detection data, structured mask result and semantic interpretation information are summarized by work order number to generate a detection report dataset and output the detection report. Deploy security probes on the terminal side to perform terminal identification and operational status collection, and control access to mobile storage media. On the boundary side, enable firewalls or isolation devices, configure communication whitelists, and block communication outside the whitelist. Create security audit records for the transmission and storage of standard test data, structured mask results, semantic interpretation information, traceability datasets, and test reports.

9. An intelligent and transparent management and control system for a testing center based on the Internet of Things, comprising the intelligent and transparent management and control method for a testing center based on the Internet of Things as described in any one of claims 1 to 8, characterized in that, Includes the following modules: The architecture access module is used to deploy network and video access and establish resource identifier mapping relationships; The work order parsing and distribution module is used to receive inspection work orders, parse work order elements, and distribute equipment tasks and video acquisition tasks. The IoT data acquisition and preprocessing module is used for heterogeneous device access and detection data acquisition and preprocessing, generating standard detection data and aggregating and storing it. Improve the Mask2Former processing module to generate task-oriented mask queries and cross-frame associations based on standard detection data, form mask trajectories, and output mask results; The zero-sample semantic assignment module is used to perform open-vocabulary semantic assignment and explicit decoupling reconstruction of category information on the mask results, and output semantic explanation information. The association summary and traceability module is used to associate and summarize structured mask and semantic interpretation information by work order number, and generate traceability records and detection reports; The security protection module is used for terminal protection and data transmission security protection, and forms a security audit log.