Autonomous characterization processing methods, systems, devices, and storage media

By introducing a large language model and a thought chain reasoning mechanism, a structured task process is automatically generated, enabling efficient automated analysis and cross-device collaborative operation of micro-nano scale material samples. This solves the problems of manual intervention and unstructured results in existing technologies, and realizes the automated transformation from image to physical operation.

CN121903339BActive Publication Date: 2026-06-30PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIV SHENZHEN GRADUATE SCHOOL
Filing Date
2026-03-24
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing technologies for the study of micro- and nano-scale materials and structural samples suffer from a speed asymmetry problem between high-throughput data acquisition and inefficient manual intervention. This makes it difficult to achieve stable and reliable individual-level instance identification and cross-device automatic addressing. Furthermore, the lack of understanding of complex imaging scenarios results in unstructured analysis results that cannot directly drive physical operations.

Method used

By introducing a pre-trained large language model and a pre-defined thought chain reasoning mechanism, combined with a human-computer interaction interface, a structured task flow is automatically generated. Through instance-level visual analysis and pixel-to-physical mapping, the target object is automatically located, filtered, measured, and statistically analyzed, and physical execution instructions are generated.

Benefits of technology

It achieves efficient and reliable individual-level instance recognition and cross-device collaborative operation, reduces manual intervention, improves the targeting of data collection and the reliability of analysis, and opens up an automatic addressing channel from image perception to physical operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903339B_ABST
    Figure CN121903339B_ABST
Patent Text Reader

Abstract

This application discloses an autonomous representation processing method, system, device, and storage medium, relating to the field of artificial intelligence technology. The method includes: receiving user-input representation requirements; analyzing and decomposing the representation requirements using a large language model and a pre-set thought chain reasoning mechanism to obtain a task flow; performing imaging and positioning operations according to the task flow to acquire imaging data; extracting target object instances and calculating attributes from the imaging data to generate instance-level structured objects; and forming representation conclusions and / or decision information or physical execution instructions based on the instance-level structured objects. This method utilizes a large model to achieve autonomous task planning, and parallel visual analysis supports automatic positioning, filtering, measurement, and statistics, significantly reducing manual costs. Dynamic process decomposition improves individual recognition accuracy in complex scenarios. Executable instructions are generated based on pixel-physical mapping, establishing a closed loop from image perception to physical operation, effectively solving the problem of cross-device collaboration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to autonomous representation processing methods, systems, devices, and storage media. Background Technology

[0002] Currently, the research and preparation of micro- and nano-scale materials and structural samples are rapidly developing towards automated and intelligent experimental platforms. Existing technologies are mainly divided into two categories: one is a hardware automation platform integrating imaging / scanning equipment, precision motion mechanisms, and control units, capable of automatic sample loading, field-of-view scanning, and data acquisition; the other is a software analysis system based on traditional computer vision or task-specific deep learning models, used to process images and output target detection or segmentation results.

[0003] However, both types of systems still face significant bottlenecks in practical applications. On the one hand, although high throughput has been initially achieved in data acquisition, the crucial characterization, screening, and evaluation stages still heavily rely on manual intervention. Researchers need to spend a significant amount of time browsing the field of view and making experience-based judgments, resulting in a speed asymmetry of "fast acquisition, slow analysis," making it difficult to form a sustainable closed loop. On the other hand, existing analysis tools are mostly fixed scripts or static pipelines, lacking the ability to understand high-level scientific research intentions. They cannot automatically break down ambiguous or complex task objectives into multi-step executable operations, nor can they dynamically adjust strategies based on intermediate results. In particular, in complex imaging scenarios with high density, overlap, or strong interference, traditional methods struggle to achieve stable and reliable individual-level instance identification and quantification. Furthermore, their output is usually limited to pixel coordinates and does not establish a mapping relationship with the coordinate system of the physical actuator, resulting in "seeing but not finding," and automatic cross-device addressing still requires manual intervention.

[0004] Therefore, there is an urgent need for an autonomous representation processing method that can achieve high representation efficiency and low labor costs, while supporting stable and reliable individual-level instance identification, as well as cross-device collaborative automatic addressing and physical operation capabilities.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main objective of this application is to provide an autonomous characterization processing method, system, device, and storage medium, aiming to solve the technical problem of how to achieve high characterization efficiency and low labor costs while supporting stable and reliable individual-level instance identification, as well as cross-device collaborative automatic addressing and physical operation capabilities.

[0007] To achieve the above objectives, this application proposes an autonomous representation processing method, which includes:

[0008] Receive representation requirements from users through the human-computer interaction interface;

[0009] By combining a pre-trained large language model and a pre-defined thought chain reasoning mechanism, the representation requirements are analyzed and decomposed to obtain the task flow.

[0010] Perform imaging and positioning operations according to the task flow to acquire imaging data;

[0011] Target object instance extraction and attribute calculation are performed on the imaging data to generate instance-level structured objects;

[0012] Based on the instance-level structured object, characterization conclusions and / or decision information or physical execution instructions are formed.

[0013] In one embodiment, the step of combining a pre-trained large language model and a preset thought chain reasoning mechanism to analyze and decompose the representation requirements to obtain the task flow includes:

[0014] The representation requirements are input into the pre-trained large language model, and semantic understanding and logical reasoning are performed based on preset system prompts and constraint rules to generate a structured task description.

[0015] The task description is broken down into multiple interrelated subtasks through the preset thought chain reasoning mechanism and output in the form of a task flow.

[0016] In one embodiment, the step of performing imaging and positioning operations according to the task flow to acquire imaging data includes:

[0017] Based on the imaging range, field of view position, and acquisition strategy in the task flow, plan the global scanning path;

[0018] According to the global scanning path, imaging and positioning operations are performed at different spatial locations within the target area to obtain corresponding imaging data.

[0019] In one embodiment, the step of performing target object instance extraction and attribute calculation on the imaging data to generate instance-level structured objects includes:

[0020] The imaging data is processed using a pre-trained instance segmentation model to identify and segment multiple target object instances, wherein each target object instance corresponds to an independent instance region;

[0021] Generate corresponding instance-level results based on the target object instance in each instance region;

[0022] Based on the instance-level results, attribute calculations are performed on the target object instance to obtain the attribute results;

[0023] The attribute results are written into the structured fields corresponding to the target object instance, and the extended analysis tool is called to generate the final instance-level structured object.

[0024] In one embodiment, writing the attribute result into a structured data field associated with the target object instance and invoking an extended analysis tool to generate the final instance-level structured object includes:

[0025] The attribute results are written into the structured data field associated with the target object instance to form an initial instance-level structured object;

[0026] The initial instance-level structured object is input into the corresponding extended analysis tool, and the extended analysis tool analyzes the initial instance-level structured object to generate extended attributes.

[0027] The extended attributes are written into the initial instance-level structured object to form the final instance-level structured object.

[0028] In one embodiment, the step of forming characterization conclusions and / or decision information or physical execution instructions based on the instance-level structured object includes:

[0029] The instance-level structured objects are aggregated and filtered to generate structured output results that meet the task objectives, and representational conclusions and / or decision information are formed based on the structured output results; or

[0030] By using a pixel-physical mapping mechanism, the pixel coordinates in the instance-level structured object are converted into physical coordinates, and physical execution instructions are generated.

[0031] In one embodiment, the step of converting the pixel coordinates of the instance-level structured object into physical coordinates using a pixel-physical mapping mechanism and generating physical execution instructions includes:

[0032] The positioning actuator is controlled to move to the target physical location, and imaging data is collected at the target physical location. The pixel coordinates and physical coordinates of the target reference point at the target physical location are obtained to form pixel-physical correspondence point data.

[0033] Based on the pixel-physical correspondence point data, a mapping model is constructed;

[0034] Extract the corresponding pixel coordinates from the instance-level structured object, and substitute the pixel coordinates into the mapping model to calculate the corresponding physical coordinates;

[0035] The physical coordinates are converted into physical execution commands for driving the positioning actuator.

[0036] Furthermore, to achieve the above objectives, this application also proposes an autonomous characterization processing system, which includes:

[0037] The receiving module is used to receive the representation requirements input by the user through the human-computer interaction interface;

[0038] The cognitive module is used to combine a pre-trained large language model and a preset thought chain reasoning mechanism to analyze and decompose the representation requirements to obtain the task flow.

[0039] The imaging control module is used to perform imaging and positioning operations according to the task flow and acquire imaging data;

[0040] The visual perception and analysis module is used to perform target object instance extraction and attribute calculation on the imaging data to generate instance-level structured objects;

[0041] The decision-making and execution module is used to form characterization conclusions and / or decision information or physical execution instructions based on the instance-level structured objects.

[0042] In addition, to achieve the above objectives, this application also proposes an autonomous characterization processing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the autonomous characterization processing method as described above.

[0043] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the autonomous representation processing method described above.

[0044] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the autonomous representation processing method described above.

[0045] This application proposes an autonomous representation processing method, system, device, and storage medium. The method includes: receiving representation requirements input by a user through a human-computer interaction interface; parsing and decomposing the representation requirements by combining a pre-trained large language model and a preset thought chain reasoning mechanism to obtain a task flow; performing imaging and positioning operations according to the task flow to acquire imaging data; extracting target object instances and calculating attributes from the imaging data to generate instance-level structured objects; and forming representation conclusions and / or decision information or physical execution instructions based on the instance-level structured objects. This solution introduces a large language model for autonomous planning and process orchestration of the representation task, combined with parallel visual analysis and attribute calculation, to achieve automatic positioning, filtering, measurement, and statistical analysis of target objects, significantly reducing manual intervention and lowering labor costs. Simultaneously, by combining a large language model and a thought chain reasoning mechanism during task execution, the analysis process is dynamically decomposed, improving the accuracy of individual instance recognition in complex or uncertain scenarios. Furthermore, physical execution instructions are generated based on instance-level structured objects, achieving end-to-end automatic addressing and linkage from image perception to physical operation, effectively solving the cross-device collaboration problem of "seeing but not finding". Attached Figure Description

[0046] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart illustrating an embodiment of the autonomous characterization processing method of this application.

[0049] Figure 2 This is a flowchart illustrating Embodiment 2 of the autonomous characterization processing method of this application;

[0050] Figure 3 This is a flowchart illustrating Embodiment 3 of the autonomous characterization processing method of this application;

[0051] Figure 4 This is a diagram illustrating the overall hardware and software architecture of the system provided in Embodiment 1 of this application.

[0052] Figure 5 A simplified flowchart illustrating the autonomous characterization processing method provided in Embodiment 1 of this application;

[0053] Figure 6This is a schematic diagram of the module structure of the autonomous characterization processing system according to an embodiment of this application;

[0054] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the autonomous characterization processing method in the embodiments of this application.

[0055] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0056] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0057] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0058] The main solution of this application embodiment is as follows: receiving the representation requirements input by the user through the human-computer interaction interface; combining the pre-trained large language model and the preset thought chain reasoning mechanism to parse and decompose the representation requirements to obtain the task flow; performing imaging and positioning operations according to the task flow to obtain imaging data; performing target object instance extraction and attribute calculation on the imaging data to generate instance-level structured objects; and forming representation conclusions and / or decision information or physical execution instructions based on the instance-level structured objects.

[0059] In this embodiment, for ease of description, the autonomous representation processing system will be used as the execution subject in the following description.

[0060] Because existing automated characterization platforms typically rely on preset scripts or manually set parameters, they lack the ability to understand the user's high-level intent and struggle to transform vague or semantic research goals into executable operation sequences. Furthermore, while their visual analysis modules can output pixel-level detection results, these results are often limited to the image coordinate system and lack a reliable mapping relationship with the coordinate system of the physical actuator, resulting in "seeing but not finding," and requiring significant manual intervention for cross-device automatic addressing. In addition, in complex imaging scenarios such as high density, overlap, or low contrast, existing methods struggle to achieve stable and reliable individual-level instance recognition and quantification, and the resulting unstructured image data is typically stored as isolated files, failing to effectively accumulate into traceable and reusable knowledge assets.

[0061] This application provides a solution. First, it introduces a pre-trained large language model and combines it with a pre-defined thought chain reasoning mechanism to perform semantic parsing and task decomposition on the natural language representation requirements input by the user. This automatically generates a structured task flow that includes imaging range planning, target selection conditions, analysis steps, and operation instructions, thereby giving the system the ability to understand high-level scientific research intentions and autonomously plan execution paths. Second, it adopts a highly robust instance-level visual analysis method to accurately segment and calculate multi-dimensional attributes of target objects in complex imaging scenarios, generating independent instance-level structured results for each target object, significantly improving the reliability and quantification accuracy of individual recognition. Third, it establishes a learnable mapping model between pixel coordinates and physical execution coordinates. By obtaining corresponding point data through system calibration, and based on this model, it converts the key pixel positions of the target objects identified in the image into absolute coordinates in physical space in real time, thereby generating physical execution instructions that can directly drive the positioning actuator, completely opening up a cross-device collaborative channel from "seeing" to "finding and doing".

[0062] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions. The following description uses a personal computer as an example to illustrate this embodiment and the subsequent embodiments.

[0063] Based on this, embodiments of this application provide an autonomous characterization processing method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the autonomous characterization processing method of this application.

[0064] In this embodiment, the autonomous characterization processing method includes steps S10 to S50:

[0065] Step S10: Receive the representation requirements input by the user through the human-computer interaction interface;

[0066] In this embodiment, the system receives natural language instructions or high-level representation targets input by the user or upper-layer application through a human-computer interaction interface. The natural language instructions may include descriptions of selection constraints, analysis indicators, output formats, and subsequent operation requirements for the target object.

[0067] Step S20: Combining the pre-trained large language model and the preset thought chain reasoning mechanism, the representation requirements are analyzed and decomposed to obtain the task flow.

[0068] It should be noted that the large language model adopts a Transformer model based on an Encoder-Decoder or Decoder-only architecture (such as Qwen, LLaMA, ChatGLM, etc.), and in this embodiment is configured to include the following three functional sub-modules:

[0069] Semantic understanding module: Receives natural language representation requirements input by users through the human-computer interaction structure, and extracts key semantic elements such as target object, observation scale, imaging modality, quantitative indicators and constraints by using pre-trained word embedding layer and multi-head attention mechanism;

[0070] Task planning module: Integrates a pre-set structured thought chain reasoning template, which is embedded in the model's context window as learnable hints or soft rules. Based on semantic understanding results, the model activates the matching thought chain path and outputs a step-by-step, ordered, and executable task flow sequence through autoregressive generation.

[0071] Domain Adaptation Module: Building upon general pre-training, the model has undergone domain-adaptive fine-tuning on high-quality instruction-procedure pair datasets covering fields such as materials science, micro / nano fabrication, electron microscopy, and bioimaging. This dataset is expert-annotated, and each sample includes: (a) natural language representation requirements; (b) the corresponding standard operating procedure (SOP); and (c) equipment parameter constraints (such as accelerating voltage, scan step size, and signal-to-noise ratio requirements). During fine-tuning, an instruction-tuning strategy was employed, enabling the model not only to understand terminology but also to generalize to unseen complex tasks.

[0072] During training, the large language model employs a two-stage training strategy. First, self-supervised pre-training is performed on a general corpus, learning basic language representation capabilities through masked language modeling or causal language modeling tasks. Then, domain-adaptive fine-tuning is conducted on a high-quality instruction-procedure dataset encompassing fields such as materials science, micro / nano fabrication, electron microscopy, and biological imaging. This dataset is annotated by domain experts, with each sample including natural language representation requirements, corresponding standard operating procedures (SOPs), and equipment parameter constraints. The fine-tuning stage utilizes an instruction-tuning method, simultaneously optimizing semantic understanding, task planning, and constraint compliance to achieve three sub-goals: the model accurately extracts key elements from input instructions, such as the target object, observation scale, and imaging modality, and generates logically sound, complete, and physically executable task flow sequences based on a pre-defined structured thought chain reasoning template; simultaneously, by introducing a loss term or post-processing verification mechanism that is aware of equipment constraints, the output flow is ensured to conform to instrument operating specifications. The final trained model can reliably transform the user's fuzzy, high-level representation requirements into structured control instructions for automated experimental platforms, providing precise guidance for subsequent imaging, analysis, and decision-making.

[0073] The Chain-of-Thought (CoT) reasoning mechanism refers to a prompting engineering strategy that guides large language models through step-by-step reasoning. This strategy includes, but is not limited to, embedding structured intermediate reasoning examples, domain-specific logical rule templates, triggerable task decomposition paths, and dynamically context-enhanced soft prompts into input prompts. The structured intermediate reasoning examples are pre-constructed by experts and contain multi-step reasoning chains from natural language representation requirements to the final executable task flow. The logical rule templates encode dependencies of common representation tasks in the form of if-then rules or state transition diagrams and inject them as prior knowledge into the model context. The dynamically context-enhanced soft prompts are generated in real-time based on the current experimental platform's capability list, sample status, and historical operation records, constraining the reasoning direction and ensuring the physical feasibility of the output flow. Through this strategy, large language models can explicitly unfold multi-step logical deductions when parsing user representation requirements, rather than directly generating end-to-end instructions, thereby significantly improving the accuracy, robustness, and interpretability of task decomposition.

[0074] Understandably, since the system's automated platform can only respond to fixed parameters or scripted instructions, it cannot handle fuzzy, open, or semantically rich research objectives. If the semantic parsing and task planning stages are skipped directly, subsequent imaging and analysis will lack clear guidance or even perform invalid operations. Therefore, executing step S20 avoids task deviations or process breaks caused by a lack of intent understanding, thereby realizing the automatic transformation from the user's high-level representation needs to a schedulable, executable, and verifiable machine-level task process, laying the decision-making foundation for the entire autonomous representation processing closed loop.

[0075] In one feasible embodiment, step S20 may include steps S21-S22:

[0076] Step S21: Input the representation requirements into the pre-trained large language model, and perform semantic understanding and logical reasoning based on preset system prompts and constraint rules to generate a structured task description;

[0077] In this embodiment, the representation requirements are input into the pre-trained large language model. The large language model performs semantic parsing and logical reasoning on the input under preset system prompts and preset constraint rules. The large language model can identify the target type, constraints, and expected output format of the representation task, and generate a structured task description accordingly.

[0078] The system prompts are used to define the model’s role, task objectives, and output format, including but not limited to instruction templates (such as “You are an autonomous characterization processing system, please generate an executable task flow according to user requirements”); domain knowledge context (such as common target types and typical parameter ranges in micro-nano material characterization); and output structure specifications (such as requiring task fields to be returned in JSON or YAML format).

[0079] The constraint rules are used to limit the legal boundaries and operational feasibility of the task, including but not limited to: physical equipment capability limitations (such as the maximum size of the imaging field of view and the lower limit of positioning accuracy); safe operation specifications (such as prohibiting the triggering of high-energy lasers in unfocused areas); and logical consistency verification rules (such as "if high-resolution imaging is required, coarse positioning must be completed first").

[0080] Step S22: Through the preset thought chain reasoning mechanism, the task description is broken down into multiple interrelated sub-tasks and output in the form of a task flow.

[0081] Furthermore, the system utilizes the aforementioned thought chain reasoning mechanism to progressively decompose high-level representation requirements into a set of subtasks with clear execution order, parameter dependencies, and operational constraints, forming a schedulable execution plan. These subtasks include, but are not limited to, operational steps such as imaging range planning, target object analysis, attribute calculation, result filtering, and physical positioning.

[0082] In this embodiment, the thought chain reasoning mechanism guides the large language model to generate intermediate reasoning processes in a logical order by embedding step-by-step reasoning instructions in system prompts. For example, the prompts may include guiding statements such as: "Please think about the following steps: First, determine the imaging area; second, perform target detection; then, filter objects that meet the conditions based on constraints; finally, generate physical instructions to drive the positioning mechanism." Thus, the model not only outputs the final task flow but also explicitly presents the reasoning path, significantly improving the accuracy, interpretability, and executability of task decomposition.

[0083] Through the above steps, the system uses a large language model as its core cognitive engine. By combining structured semantic understanding with step-by-step reasoning guidance, it achieves automatic transformation from high-level representation needs to specific execution actions. This effectively avoids process breaks or operational deviations caused by ambiguous intentions or task complexity, thereby supporting the system to complete end-to-end autonomous representation tasks without human intervention.

[0084] Step S30: Perform imaging and positioning operations according to the task flow to acquire imaging data;

[0085] It should be noted that the imaging data refers to a set of images or scan data collected by the imaging device under system control, covering the target area specified in the task. Its forms include, but are not limited to, optical microscopic images, electron microscopic images, laser confocal scans, atomic force microscopy (AFM) topographic images, or spectral / energy spectral spatial mapping data. The imaging data is stored in digital format and associated with corresponding physical spatial location information, serving as the input source for subsequent visual perception and instance analysis.

[0086] Understandably, existing automated platforms typically employ fixed fields of view or preset sampling strategies, making it difficult to dynamically adjust the imaging range, resolution, or acquisition density based on task semantics. This leads to over-acquisition in sparse target areas, insufficient coverage in critical areas, or an inability to perform targeted high-resolution imaging of specific objects. Furthermore, if imaging and positioning operations are not coordinated with the task flow, it will result in data redundancy, inefficiency, and even the omission of critical targets. Therefore, executing step S30 avoids the waste of resources and the risk of target omission caused by blind acquisition or static scanning, thereby achieving task-driven adaptive data acquisition—that is, imaging is performed only in the spatial area required by the task, with the required accuracy and strategy, significantly improving the targeting, completeness, and overall system representation efficiency of data acquisition.

[0087] In one feasible embodiment, step S30 may include steps S31-S32:

[0088] Step S31: Based on the imaging range, field of view position, and acquisition strategy in the task flow, plan the global scanning path;

[0089] In this embodiment, the system analyzes the spatial coverage requirements (such as target area boundaries and regions of interest), field of view parameters (such as single-shot imaging field of view size and overlap rate), and acquisition strategies (such as resolution level, sampling density, and whether to enable high-resolution local supplementary sampling) contained in the task flow, and comprehensively generates a global scanning path that covers the entire target area and meets the task accuracy requirements. This path consists of a series of ordered spatial coordinate points or trajectory segments, which are used to guide subsequent positioning and imaging operations.

[0090] Step S32: According to the global scanning path, perform imaging and positioning operations at different spatial locations within the target area to obtain corresponding imaging data.

[0091] Specifically, the system controls the positioning actuator to move sequentially to each spatial position specified by the global scanning path. After achieving stable positioning at each position, the imaging device is triggered to acquire images or scan data of the corresponding field of view. The acquired imaging data is automatically associated with the physical coordinate information of the current position and is summarized into an imaging data set covering the entire target area, serving as the input basis for subsequent visual perception and instance analysis.

[0092] Through the above steps, the task flow of large model generation is directly mapped into an executable scanning path. Guided by this path, orderly positioning and imaging operations are completed, avoiding blind scanning and invalid data acquisition. This significantly improves imaging efficiency, data correlation, and the reliability of subsequent analysis while ensuring full coverage of the target area. Furthermore, the acquired imaging data is strictly aligned with its physical location, laying the foundation for high-precision mapping from pixel coordinates to physical coordinates and effectively supporting the "what you see is what you find" cross-device automatic addressing capability.

[0093] Step S40: Perform target object instance extraction and attribute calculation on the imaging data to generate instance-level structured objects;

[0094] Understandably, traditional image analysis methods typically only output category-level detection results or pixel-level semantic segmentation masks, making it difficult to accurately distinguish between individuals in complex scenes where target objects are densely packed, overlapping, or have blurred boundaries. Furthermore, their output is mostly unstructured data, lacking quantitative descriptions of each individual object, which prevents subsequent precise screening, statistical analysis, or physical operations. Therefore, executing step S40 avoids analytical biases and reliance on manual verification caused by instance confusion or missing attributes, thereby achieving individual-level recognition, multi-dimensional quantification, and structured representation of target objects, providing a reliable data foundation for high-level decision-making and automated execution.

[0095] In one feasible embodiment, step S40 may include steps S41 to S44:

[0096] Step S41: The imaging data is processed using a pre-trained instance segmentation model to identify and segment multiple target object instances.

[0097] To achieve reliable differentiation and quantitative analysis of multiple target objects in imaging data, this embodiment introduces an instance-level target object analysis mechanism, which is used to extract mutually independent target object instances from the original imaging data and establish an instance-level structured representation containing spatial location and attribute information for each target object instance.

[0098] Specifically, the system first performs preprocessing operations on the received raw imaging data. These preprocessing operations include, but are not limited to, noise reduction, contrast enhancement, brightness normalization, background correction, or other processing steps that enhance the salience of the target object.

[0099] The pre-trained instance segmentation model (such as the instance segmentation network Mask R-CNN, YOLO-World, or a self-developed architecture) is used to process the pre-processed imaging data, so that each target object corresponds to an independent instance region (i.e., a pixel-level mask), and different instances are assigned a unique identifier even if they are spatially adjacent or partially overlapping to ensure distinguishability.

[0100] The instance segmentation model can be trained and optimized in the following ways:

[0101] First, a training dataset containing imaging data and corresponding instance annotation information is constructed. The instance annotation information is used to identify the independent instance region of each target object in the imaging data.

[0102] Subsequently, the instance segmentation model is trained based on the training dataset to enable the model to predict the target object instance region from the input imaging data.

[0103] After the model is trained, its performance in terms of segmentation accuracy, recall, and generalization is evaluated using an independent validation dataset. Based on the evaluation results, the model structure, loss function, or hyperparameters are iteratively optimized.

[0104] It should be noted that the training data sources, annotation methods, model structures, and optimization strategies described are all illustrative examples and do not constitute a limitation of the present invention.

[0105] Through the above steps, the system can effectively distinguish different individuals into independent instance regions in complex scenarios where target objects are in contact, overlap, or have unclear boundaries, avoiding the mistaken merging of multiple target objects into a single connected region, and significantly improving the accuracy of subsequent attribute calculations and decision analysis.

[0106] Step S42: Generate corresponding instance-level results based on the target object instance in each instance region;

[0107] In this embodiment, for each target object instance identified through instance segmentation, the system generates an initial instance-level result. This instance-level result includes, but is not limited to:

[0108] Instance identifier: Used to distinguish different target objects;

[0109] Instance region identifier: used to describe the spatial extent of the target object instance in the imaging data;

[0110] Instance location description: Used to describe the location of the target object instance in pixel coordinate space.

[0111] Step S43: Based on the instance-level results, perform attribute calculations on the target object instance to obtain attribute results;

[0112] After obtaining instance-level results, the system further quantifies the multi-dimensional attributes of each target object instance. Specifically, based on its instance region, it calculates its geometric, morphological, and photometric features to generate attribute results. These attributes include, but are not limited to: area, perimeter, equivalent diameter, roundness, convex hull ratio, principal orientation angle, texture features (such as gray-level co-occurrence matrix energy and entropy), brightness / gray-level statistics (mean, variance, histogram distribution), and physical estimates that can be derived under specific modalities (such as calibrated thickness, layer, or composition confidence). All attributes are output in numerical or structured vector form and bound to the corresponding instance identifier.

[0113] Step S44: Write the attribute results into the structured field corresponding to the target object instance, and call the extended analysis tool to generate the final instance-level structured object.

[0114] The system writes the attribute results obtained in step S43 into the structured data container of the target object instance, forming an intermediate object containing basic attributes. Subsequently, according to the analysis requirements specified in the task flow, one or more extended analysis tools are dynamically invoked to perform specialized enhancement analysis on the intermediate object. The output of the extended analysis tools (such as the number of internal holes, crack branch structure, nearest neighbor distance with other instances, etc.) is written back to the original object as new fields, finally generating an enhanced instance-level structured object containing multidimensional basic attributes and task-related extended attributes. The instance-level structured object includes the instance identifier, pixel domain identifier, position coordinates, confidence level, and multidimensional attribute fields of the target object instance.

[0115] Through the above steps, the system not only achieves individualized modeling and standardized expression of each target object, but also flexibly supports diverse high-level analysis tasks without modifying the core architecture through a modular extension mechanism. This avoids the difficulties in reuse and reliance on manual intervention caused by attribute missingness, unstructured data, or rigid analysis capabilities in traditional methods, thus providing a complete, reliable, and traceable data foundation for subsequent quantitative screening, intelligent decision-making, and physical execution.

[0116] Step S50: Based on the instance-level structured object, form a representation conclusion and / or decision information or physical execution instructions.

[0117] Understandably, existing automated analysis systems typically stop at image detection or simple statistics, and their outputs are mostly unstructured reports or isolated data files, lacking semantic alignment with task objectives and unable to directly drive downstream physical devices. Without high-level semantic aggregation and executable transformation of instance-level structured objects, the analysis results will be "understandable but unusable," still requiring manual interpretation, screening, and operational planning. Therefore, step S50 avoids the disconnect between analysis results and task objectives, as well as the problem of separation between perception and execution, thereby achieving a leap from individual quantification to task closure—that is, automatically generating understandable scientific conclusions, actionable decision suggestions, or executable physical control commands, truly supporting autonomous representation without human intervention.

[0118] In one feasible embodiment, step S50 may include steps S51-S52:

[0119] Step S51: Summarize and filter the instance-level structured objects to generate structured output results that meet the task objectives, and form representation conclusions and / or decision information based on the structured output results;

[0120] In this embodiment, the system summarizes and filters the set of instance-level structured objects, generates structured output results that meet the task objectives, and forms characterization conclusions and / or decision information based on the structured output results.

[0121] The task objectives include, but are not limited to, target counting, size screening, morphology evaluation, defect identification, distribution statistics, or operation candidate recommendation.

[0122] The characterization conclusions include quantitative statistical results such as the number of target objects, distribution density, size mean and variance, morphology pass rate, and defect detection rate.

[0123] The decision information includes judgmental outputs such as whether the process acceptance criteria are met, whether a high-resolution re-inspection is triggered, and whether a specific target is recommended to proceed to the next process.

[0124] Step S52: Using a pixel-physical mapping mechanism, the pixel coordinates in the instance-level structured object are converted into physical coordinates, and physical execution instructions are generated.

[0125] In this embodiment, a pre-calibrated pixel-physical coordinate mapping model is invoked to convert the key pixel coordinates recorded in the instance-level structured object into absolute coordinates in the physical coordinate system of the positioning actuator. Subsequently, standardized physical execution instructions are generated based on these physical coordinates. These physical execution instructions can be directly parsed and executed by the hardware control unit to achieve automatic addressing, high-resolution imaging, marking, or processing operations on specific target objects. The system can also store the structured results and related metadata in the data management module to support subsequent task reuse, comparative analysis, and model iteration, thereby completing a complete autonomous representation closed-loop process.

[0126] Through the above steps, the system not only completes the extraction of "data" into "knowledge" (representation of conclusions and decisions), but also opens up the path from "knowledge" to "action" (physical execution instructions), effectively solving the problem of the separation of "analysis is analysis, operation is manual" in traditional platforms, and truly realizing the autonomous representation processing closed loop of "what you see is what you get, and what you get is controllable".

[0127] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Further elaborating on step S44, the autonomous representation processing method also includes steps A1 to A3:

[0128] Step A1: Write the attribute results into the structured data field associated with the target object instance to form an initial instance-level structured object;

[0129] In this embodiment, the attribute results calculated in step S43 are written into the data container of the corresponding target object instance in the form of key-value pairs or structured records, generating an initial instance-level structured object containing basic geometric and appearance attributes, which serves as the input carrier for subsequent extended analysis.

[0130] Step A2: Input the initial instance-level structured object into the corresponding extended analysis tool, and use the extended analysis tool to analyze the initial instance-level structured object to generate extended attributes;

[0131] To meet the diverse needs of different representation tasks in terms of analytical objectives and dimensions, this embodiment introduces an extended analysis tool mechanism. This mechanism takes an initial instance-level structured object as input and further processes and quantifies specific structural features or analytical indicators of the target object in a modular manner, thereby achieving flexible support for various analytical tasks without changing the overall system architecture.

[0132] The extended analysis tools are uniformly encapsulated as optional analysis modules and invoked as needed by the cognitive control layer through standardized interfaces. Different extended analysis tools are functionally independent, and their specific implementations can be configured or replaced according to the analysis objectives. The number and type of these tools do not constitute a limitation on the present invention.

[0133] During system operation, the initial instance-level structured object is used as input data for the extended analysis tools. Each extended analysis tool performs further analysis operations on the target object based on the instance region, pixel coordinates, and calculated attribute results in the initial instance-level structured object to obtain extended attributes. The extended attributes include, but are not limited to, higher-order quantitative indicators such as the number of internal defects, principal axis direction angle, contour curvature distribution, neighboring instance density, and void filling rate.

[0134] The extended analysis tools may include, but are not limited to, the following types:

[0135] (1) Geometric structure feature analysis tool: used to further analyze the geometric morphology of target object instances. Based on the instance region corresponding to the target object instance, the geometric structure feature analysis calculates the boundary contour, corner distribution, main direction information or shape description parameters of the target object, and generates quantitative results reflecting the geometric structure features of the target object accordingly. The analysis results can be stored as instance-level extended attributes for subsequent filtering, classification or statistical analysis.

[0136] (2) Local structure or defect feature analysis tool: used to identify and quantify local structural features within a target object instance. The local structure or defect feature analysis tool can perform secondary segmentation, morphological processing, or feature detection on the instance region to identify local regions or specific structures within the instance and calculate the corresponding quantity, distribution location, or scale parameters. The analysis process can be implemented based on traditional image processing algorithms, machine learning models, or a combination of both.

[0137] (3) Multi-instance relationship analysis tool: used to analyze the spatial relationships or relative attributes between multiple target object instances. The analysis content may include the relative distance, directional relationship, arrangement characteristics, or statistical distribution characteristics between target objects. Through the multi-instance relationship analysis tool, the system can be extended from single-instance analysis to multi-instance collaborative analysis to support more complex representation tasks.

[0138] Step A3: Write the extended attributes into the initial instance-level structured object to form the final instance-level structured object.

[0139] In this embodiment, the system writes back the extended attributes output by each extended analysis tool as new fields to the corresponding initial instance-level structured object, ultimately generating an enhanced instance-level structured object containing basic attributes and task-related extended attributes. This object has complete individual description capabilities and can be directly used for subsequent filtering, decision-making, or physical execution.

[0140] The above-described method involves writing the attribute results into a structured data field associated with the target object instance to form an initial instance-level structured object. This initial instance-level structured object is then input into a corresponding extended analysis tool, which performs analysis based on the initial instance-level structured object to generate extended attributes. Finally, these extended attributes are written into the initial instance-level structured object to form a final instance-level structured object. This approach not only achieves standardized, scalable, and traceable structured representation of the target object but also significantly improves analytical flexibility and task adaptability in complex research scenarios through a modular extended analysis mechanism. It avoids the functional limitations and repetitive development costs caused by the fixed analytical capabilities of traditional systems, thus providing key technical support for building an evolvable and reusable autonomous representation processing platform.

[0141] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Further elaborating on step S52, the autonomous representation processing method also includes steps B1 to B4:

[0142] Step B1: Control the positioning actuator to move to the target physical location, collect imaging data at the target physical location, and obtain the pixel coordinates and physical coordinates of the target reference point at the target physical location to form pixel-physical correspondence point data;

[0143] It should be noted that the positioning actuator refers to a programmable motion platform used to precisely control the position of the imaging device or sample in physical space, and its form includes, but is not limited to, an electric displacement stage, a robotic arm, or a coordinate system of a subsequent precision instrument.

[0144] To enable the imaging analysis results to directly drive the physical actuator, this embodiment introduces a pixel-physical mapping mechanism to establish the correspondence between the pixel coordinates of the target object in the imaging data and the physical coordinates of the actual physical actuator. This enables the instance-level structured object to be executable and supports automatic addressing, resetting and subsequent operations of the target object.

[0145] In this embodiment, the imaging data obtained in step S30 is defined in pixel coordinate space, which is represented by pixel position parameters and used to describe the relative positional relationship of the target object in the imaging field of view; the motion position of the positioning actuator is defined in physical coordinate space, which is used to describe the displacement or pose state of the actuator in actual space. The pixel-physical mapping mechanism is used to establish a stable and computable mapping relationship between the above two coordinate spaces.

[0146] Specifically, firstly, the position of any point in the imaging data is expressed in pixel coordinates. It means that among them and These represent the horizontal and vertical pixel positions of the point in the imaging field of view, respectively; correspondingly, the position of the positioning actuator in actual space is represented by physical coordinates. It means that among them and This represents the position parameters of the actuator in the physical coordinate space.

[0147] Then, by controlling the positioning actuator to stop at multiple known physical locations and acquiring imaging data at the corresponding locations, the pixel coordinates of the target reference point at the locations are recorded. With physical coordinates This allows for the acquisition of multiple sets of pixel-physical correspondence point data. The calibration process can be performed periodically during system initialization or operation.

[0148] Step B2: Construct a mapping model based on the pixel-physical correspondence point data;

[0149] In this embodiment, a mapping model from pixel coordinates to physical coordinates is constructed based on the pixel-physical correspondence point data system. In one embodiment, the mapping model may take the form of a linear or affine mapping, for example:

[0150]

[0151] in, These are the mapping parameters obtained through a calibration process. In other embodiments, the mapping model may also employ polynomial mapping, piecewise mapping, or other nonlinear mapping forms, and the specific model form can be selected according to the actual characteristics of the imaging system and the actuator.

[0152] Step B3: Extract the corresponding pixel coordinates from the instance-level structured object, and substitute the pixel coordinates into the mapping model to calculate the corresponding physical coordinates;

[0153] In this embodiment, the system extracts the key pixel coordinates of the target object from the instance-level structured object. These key pixel coordinates can be the center point coordinates, feature point coordinates, or other representative pixel positions of the target object. The system then substitutes these pixel coordinates into an established mapping model to calculate the corresponding physical coordinates. .

[0154] Step B4: Convert the physical coordinates into physical execution commands for driving the positioning actuator.

[0155] In this embodiment, the system converts the calculated physical coordinates into physical execution commands that the positioning actuator can recognize, and sends them from the cognitive control layer to the hardware support layer to drive the positioning actuator to complete the automatic addressing, resetting or alignment operation of the target object, thereby realizing a closed-loop connection from pixel-level analysis results to physical action execution.

[0156] The above-described method controls the positioning actuator to move to the target physical location, collects imaging data at the target physical location, and obtains the pixel coordinates and physical coordinates of the target reference point at the target physical location to form pixel-physical correspondence point data. Based on the pixel-physical correspondence point data, a mapping model is constructed. The corresponding pixel coordinates are extracted from the instance-level structured object and substituted into the mapping model to calculate the corresponding physical coordinates. The physical coordinates are converted into physical execution commands to drive the positioning actuator. This effectively solves the core bottlenecks of "disconnect between perception and execution" and "seeing but not finding" in the prior art, avoids manual intervention or operation failures caused by inconsistencies in coordinate systems, and thus achieves fully automatic target addressing and precise operation across devices and modalities, significantly improving the intelligence level and task reliability of the autonomous representation system.

[0157] For example, to help understand the implementation flow of the autonomous characterization processing method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 4 , Figure 4 A system hardware and software architecture diagram for an autonomous representation processing method is provided, specifically:

[0158] In this embodiment, the system adopts a layered architecture design, which is functionally divided into a hardware support layer, a cognitive control layer, and a visual perception and analysis tool layer. Each layer collaborates through standardized data and control interfaces, thereby achieving a closed-loop process from high-level representation of intent input to target object analysis result output and physical execution instruction generation.

[0159] The hardware support layer, serving as the physical foundation for system operation, is used to perform operations such as carrying the object to be displayed, acquiring imaging data, and adjusting its spatial position. It provides a stable data source and executable physical operation capabilities for the upper-level cognitive control and visual analysis. In one embodiment, the hardware support layer includes at least an imaging device, a positioning execution mechanism, and a connected computing and control unit.

[0160] The imaging device is used to acquire images or scan data of the representative object. The imaging device can be an optical microscope, electron microscope, or other imaging equipment with two-dimensional or multi-dimensional imaging capabilities. Preferably, the imaging device is equipped with an image sensor to acquire imaging data of the target object at different field-of-view positions, and transmits the imaging data to a computing and control unit via a data interface.

[0161] The positioning actuator is used to adjust the relative position between the object to be characterized and the imaging device. It may include an electrically driven displacement stage, a linear motion module, a rotary platform, or other controllable motion mechanisms. In one embodiment, the positioning actuator has at least two-dimensional movement capability along a planar direction and optionally has focusing or height adjustment capability in the vertical direction. The positioning actuator is communicatively connected to a computing and control unit to receive displacement or positioning commands and provide feedback on the current pose information.

[0162] The computing and control unit is used for unified management of the imaging device and the positioning actuator. It can be a server, workstation, or embedded computing device, and internally deploys system control programs, communication interfaces, and data processing modules. The computing and control unit can receive control commands from the cognitive control layer and convert them into corresponding imaging trigger signals or physical motion control signals. It is also responsible for receiving the raw data collected by the imaging device and providing it to the visual perception and analysis tool layer for processing.

[0163] During system operation, the hardware support layer executes imaging and positioning operations according to the task flow generated by the cognitive control layer. The positioning execution mechanism drives the imaging device to acquire imaging data at different spatial locations, thereby forming an imaging data set covering the area of ​​the object to be characterized.

[0164] The cognitive control layer, serving as the system's decision-making and scheduling hub, is used to understand the intent of the representational tasks, generate task flows, and coordinate and control lower-level functional modules and hardware units. Using a large language model as its core cognitive engine, the cognitive control layer, in collaboration with the visual perception and analysis tool layer and the hardware support layer, achieves the automatic transformation from high-level representational needs to specific execution actions.

[0165] In one embodiment, the cognitive control layer is configured to receive natural language instructions or high-level representation targets from a user or upper-layer application, and perform semantic parsing and logical reasoning on the input. Through preset system prompts and constraint rules, the large language model can identify the target type, constraints, and expected output format of the representation task, and generate a structured task description accordingly.

[0166] Furthermore, the cognitive control layer, through a thought chain reasoning mechanism, decomposes the high-level representation requirements into multiple interrelated sub-tasks, forming a set of task flows or execution plans with a clear execution order and parameter constraints. These sub-tasks include, but are not limited to, operational steps such as imaging range planning, target object analysis, attribute calculation, result filtering, and physical positioning.

[0167] During the execution phase, the cognitive control layer invokes the analysis modules in the visual perception and analysis tool layer through standardized interfaces, and adjusts or supplements subsequent task flows based on intermediate analysis results, thereby achieving dynamic task orchestration with feedback capabilities. Simultaneously, the cognitive control layer can convert analysis results into control commands for the hardware support layer, driving the imaging device and positioning execution mechanism to complete corresponding physical operations.

[0168] It should be noted that the cognitive control layer focuses on understanding the intent, logical planning, and process control of the representation task, without relying on a specific material system, sample type, or imaging device type. By decoupling the cognitive control logic from the specific analysis algorithm and hardware implementation, the cognitive control layer of this invention can adapt to different application scenarios and support the expansion and replacement of subsequent functional modules.

[0169] The visual perception and analysis tool layer, as the collection of perception and analysis capabilities of the system in this embodiment, is used to perform target object recognition, instance-level segmentation, attribute calculation, and structured output of the imaging data acquired by the hardware support layer, and provides callable analysis capabilities to the cognitive control layer through a standardized interface. The visual perception and analysis tool layer consists of multiple relatively independent analysis modules. Each module is encapsulated in a modular way and selected, combined, and scheduled by the cognitive control layer as needed through a unified calling interface, thereby forming an expandable "skill library".

[0170] In one embodiment, the visual perception and analysis tool layer includes at least a target object instance extraction module and an attribute calculation module. The target object instance extraction module performs object-level recognition and separation on the input imaging data, outputting an instance-level result for each target object. The result includes at least the instance identifier of the target object, a spatial location description, and a region representation associated with that instance. Preferably, the target object instance extraction module can be implemented using instance segmentation to obtain independent target object instances even in scenes where target objects are densely packed, in contact, or overlapping.

[0171] The attribute calculation module is used to perform quantitative analysis of the target object based on instance-level results and output geometric and / or physical attributes related to the target object. The attribute calculation module can further convert pixel-domain features into physically meaningful measurement results based on imaging calibration parameters or other scale conversion parameters.

[0172] In one embodiment, the visual perception and analysis tool layer may further include one or more task-specific analysis modules for detecting and quantifying specific structural features or phenomena of a target object. These task-specific analysis modules can be implemented using traditional image processing algorithms, machine learning models, or a combination of both, and their applicable objects and analysis targets can be expanded according to actual needs.

[0173] In addition, the visual perception and analysis tool layer preferably outputs a unified format of instance-level structured object collections as the basis for the cognitive control layer to perform filtering, sorting, statistical summarization and decision generation.

[0174] Based on the functional descriptions of the modules above, the hardware support layer, cognitive control layer, and visual perception and analysis tool layer are functionally independent but closely coordinated during operation, jointly completing the autonomous representation task through bidirectional interaction of data flow and control flow.

[0175] Specifically, the user or upper-layer application first inputs natural language commands or high-level representation targets into the system through a human-computer interaction interface. This input is then passed to the cognitive control layer. The cognitive control layer performs semantic parsing and logical reasoning on the input, generates a structured task flow or execution plan, and schedules lower-layer functional modules according to the execution plan.

[0176] During execution, the cognitive control layer sends imaging and positioning-related control commands to the hardware support layer to drive the imaging device and positioning actuator to acquire imaging data at different spatial locations. The hardware support layer then transmits the acquired raw imaging data to the visual perception and analysis tool layer in real time or in batches as input for subsequent analysis and processing.

[0177] The visual perception and analysis tool layer performs analysis operations such as target object instance extraction and attribute calculation on the imaging data, generating an instance-level structured object set, and returns the analysis results to the cognitive control layer. Based on the instance-level structured object set, the cognitive control layer judges the task execution status and can adjust, supplement, or terminate subsequent task processes according to the analysis results, thereby forming a closed-loop execution process with feedback capabilities.

[0178] When physical addressing or subsequent operations are required, the cognitive control layer can further convert the analyzed target object location information into physical execution commands for the hardware support layer, driving the positioning actuator to complete the automatic addressing, reset, or related operations of the target object. Finally, the system can output the analysis results, decision information, or execution status to the user or store them in the data management module for subsequent task retrieval and reuse.

[0179] Through the above approach, this embodiment realizes a system operation mechanism with the cognitive control layer as the core, which unifies and coordinates imaging data acquisition, instance-level analysis, and physical execution. This enables each functional layer to collaboratively complete complex representation tasks while maintaining mutual decoupling, and provides a good foundation for subsequent functional expansion and application migration.

[0180] For example, to help understand the implementation flow of the autonomous characterization processing method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 5 , Figure 5 A simplified flowchart of an autonomous representation processing method is provided, specifically:

[0181] First, the system receives natural language commands or high-level representation targets input by users or upper-layer applications, and feeds the input into a large language model. Under preset system prompts and tool call constraints, the model performs semantic understanding and logical reasoning, and uses a thought chain reasoning mechanism to parse the abstract requirements into a set of structured subtasks. These subtasks are output in the form of a task flow, which includes at least: the type of analysis module to be called, the calling order, key parameters or constraints, and the expected intermediate / final output fields.

[0182] Then, according to the task flow, the system plans a global scanning path for the imaging range, field of view, and acquisition strategy, and controls the hardware support layer to perform imaging data acquisition according to the global scanning path. In one embodiment, the system drives the positioning actuator to move according to preset or dynamically adjusted step parameters, and triggers the imaging device to acquire images or scan data at the corresponding positions, thereby obtaining an imaging data set covering the target area. The acquired imaging data is transmitted to the computing and control unit as input for subsequent visual analysis.

[0183] Subsequently, the system performs object-level identification and separation on the imaging data, generating instance-level results for each target object. For each target object instance, the system outputs the region representation, location description, and necessary confidence information related to that instance, thereby achieving the transformation from image-level perception to instance-level representation. In scenarios where target objects are dense, in contact, or overlapping, this instance-level analysis can improve individual discrimination ability and result stability.

[0184] After obtaining instance-level results, the system further performs quantitative calculations on the geometric and / or physical properties of each target object instance and writes the calculation results into the corresponding instance data structure, forming a unified format of instance-level structured object collections. In one embodiment, the system can convert pixel-domain measurement results into physically meaningful metrics based on imaging calibration parameters and scale conversion parameters, and retain the key location fields of the target object for subsequent addressing or operations.

[0185] Finally, the system summarizes and filters the set of instance-level structured objects, generating structured output results that meet the task objectives, and feeds these results back to the cognitive control layer. Based on the structured results, the cognitive control layer forms analytical conclusions and / or decision information, and can generate executable physical execution instructions when needed to drive the hardware support layer to complete automatic addressing, resetting, or subsequent operations of the target object. The system can also store the structured results and related metadata in the data management module to support subsequent task reuse, comparative analysis, and model iteration, thereby completing a full autonomous representation closed-loop process.

[0186] Through the methods described in the above embodiments, this application realizes a closed-loop operation mechanism from high-level intent input, imaging data acquisition, instance-level analysis, attribute extraction to result output and physical execution, enabling the system to complete autonomous representation and analysis in a unified manner under different task objectives and different imaging scenarios, and possessing good scalability and applicability.

[0187] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the autonomous characterization processing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0188] This application also provides an autonomous characterization processing system; please refer to [reference needed]. Figure 6 The autonomous characterization processing system includes:

[0189] The receiving module 10 is used to receive the representation requirements input by the user through the human-computer interaction interface;

[0190] The cognitive module 20 is used to combine a pre-trained large language model and a preset thought chain reasoning mechanism to analyze and decompose the representation requirements to obtain the task flow.

[0191] Imaging control module 30 is used to perform imaging and positioning operations according to the task flow and acquire imaging data;

[0192] The visual perception and analysis module 40 is used to perform target object instance extraction and attribute calculation on the imaging data to generate instance-level structured objects;

[0193] The decision-making and execution module 50 is used to form characterization conclusions and / or decision information or physical execution instructions based on the instance-level structured object.

[0194] The autonomous representation processing system provided in this application, employing the autonomous representation processing method in the above embodiments, can solve the technical problem of how to achieve high representation efficiency and low labor costs while supporting stable and reliable individual-level instance identification, as well as cross-device collaborative automatic addressing and physical operation capabilities. Compared with the prior art, the autonomous representation processing system provided in this application has the same beneficial effects as the autonomous representation processing method provided in the above embodiments, and other technical features of the autonomous representation processing system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0195] This application provides an autonomous characterization processing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the autonomous characterization processing method in the above embodiment 1.

[0196] The following is for reference. Figure 7 The diagram illustrates a structural schematic suitable for implementing the autonomous characterization processing device of the embodiments of this application. The autonomous characterization processing device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The autonomous characterization processing device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0197] like Figure 7 As shown, the autonomous characterization processing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the autonomous characterization processing device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the autonomous characterization processing device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows an autonomous characterization processing device with various systems, it should be understood that implementing or possessing all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0198] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0199] The autonomous characterization processing device provided in this application, employing the autonomous characterization processing method in the above embodiments, can solve the technical problem of how to achieve high characterization efficiency and low labor costs while supporting stable and reliable individual-level instance identification, as well as cross-device collaborative automatic addressing and physical operation capabilities. Compared with the prior art, the beneficial effects of the autonomous characterization processing device provided in this application are the same as those of the autonomous characterization processing method provided in the above embodiments, and other technical features in this autonomous characterization processing device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0200] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0201] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0202] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the autonomous characterization processing method in the above embodiments.

[0203] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0204] The aforementioned computer-readable storage medium may be included in the autonomous characterization processing device; or it may exist independently and not assembled into the autonomous characterization processing device.

[0205] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the autonomous representation processing device, the autonomous representation processing device: receives representation requirements input by the user through a human-computer interaction interface; combines a pre-trained large language model and a preset thought chain reasoning mechanism to parse and decompose the representation requirements to obtain a task flow; performs imaging and positioning operations according to the task flow to acquire imaging data; performs target object instance extraction and attribute calculation on the imaging data to generate instance-level structured objects; and forms representation conclusions and / or decision information or physical execution instructions based on the instance-level structured objects.

[0206] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0207] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0208] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0209] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described autonomous representation processing method, thereby solving the technical problem of autonomous representation processing. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the autonomous representation processing method provided in the above embodiments, and will not be repeated here.

[0210] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the autonomous representation processing method described above.

[0211] The computer program product provided in this application solves the technical problem of how to achieve high representation efficiency and low labor costs while supporting stable and reliable individual-level instance identification, as well as cross-device collaborative automatic addressing and physical operation capabilities. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the autonomous representation processing method provided in the above embodiments, and will not be repeated here.

[0212] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A self-representation processing method, characterized in that, The autonomous characterization processing method includes: Receive representation requirements from users through the human-computer interaction interface; The representation requirements are input into a pre-trained large language model, and semantic understanding and logical reasoning are performed based on preset system prompts and constraint rules to generate a structured task description. The task description is broken down into multiple interrelated sub-tasks through the preset thought chain reasoning mechanism and output in the form of a task flow. Based on the imaging range, field of view position, and acquisition strategy in the task flow, plan the global scanning path; According to the global scanning path, imaging and positioning operations are performed at different spatial locations within the target area to obtain corresponding imaging data; The imaging data is processed using a pre-trained instance segmentation model to identify and segment multiple target object instances, wherein each target object instance corresponds to an independent instance region; Generate corresponding instance-level results based on the target object instance in each instance region; Based on the instance-level results, attribute calculations are performed on the target object instance to obtain the attribute results; The attribute results are written into the structured fields corresponding to the target object instance, and the extended analysis tool is called to generate the final instance-level structured object. Based on the instance-level structured object, characterization conclusions and / or decision information or physical execution instructions are formed.

2. The autonomous characterization processing method as described in claim 1, characterized in that, The step of writing the attribute results into the structured data field associated with the target object instance and calling the extended analysis tool to generate the final instance-level structured object includes: The attribute results are written into the structured data field associated with the target object instance to form an initial instance-level structured object; The initial instance-level structured object is input into the corresponding extended analysis tool, and the extended analysis tool analyzes the initial instance-level structured object to generate extended attributes. The extended attributes are written into the initial instance-level structured object to form the final instance-level structured object.

3. The autonomous characterization processing method as described in claim 1, characterized in that, The step of forming representational conclusions and / or decision information or physical execution instructions based on the instance-level structured object includes: The instance-level structured objects are aggregated and filtered to generate structured output results that meet the task objectives, and representational conclusions and / or decision information are formed based on the structured output results; or By using a pixel-physical mapping mechanism, the pixel coordinates in the instance-level structured object are converted into physical coordinates, and physical execution instructions are generated.

4. The autonomous characterization processing method as described in claim 3, characterized in that, The step of converting the pixel coordinates of the instance-level structured object into physical coordinates and generating physical execution instructions using a pixel-physical mapping mechanism includes: The positioning actuator is controlled to move to the target physical location, and imaging data is collected at the target physical location. The pixel coordinates and physical coordinates of the target reference point at the target physical location are obtained to form pixel-physical correspondence point data. Based on the pixel-physical correspondence point data, a mapping model is constructed; Extract the corresponding pixel coordinates from the instance-level structured object, and substitute the pixel coordinates into the mapping model to calculate the corresponding physical coordinates; The physical coordinates are converted into physical execution commands for driving the positioning actuator.

5. An autonomous characterization processing system, characterized in that, The autonomous characterization processing system includes: The receiving module is used to receive the representation requirements input by the user through the human-computer interaction interface; The cognitive module is used to input the representation requirements into a pre-trained large language model, and perform semantic understanding and logical reasoning based on preset system prompts and constraint rules to generate a structured task description; through the preset thought chain reasoning mechanism, the task description is decomposed into multiple interrelated sub-tasks and output in the form of a task flow. The imaging control module is used to plan the global scanning path according to the imaging range, field of view position and acquisition strategy in the task flow; According to the global scanning path, imaging and positioning operations are performed at different spatial locations within the target area to obtain corresponding imaging data; The visual perception and analysis module is used to process the imaging data using a pre-trained instance segmentation model, identify and segment multiple target object instances, wherein each target object instance corresponds to an independent instance region; generate corresponding instance-level results based on the target object instances in each instance region; perform attribute calculations on the target object instances based on the instance-level results to obtain attribute results; write the attribute results into the structured fields corresponding to the target object instances, and call extended analysis tools to generate the final instance-level structured object; The decision-making and execution module is used to form characterization conclusions and / or decision information or physical execution instructions based on the instance-level structured objects.

6. An autonomous characterization processing device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the autonomous characterization processing method as described in any one of claims 1 to 4.

7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the autonomous representation processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Robot multi-modal fusion autonomous decision-making method and system based on large language model

    CN121351007A

  • Android system user interface interaction method and device based on large language model

    CN121455596A