Image annotation method for spacecraft target type and attribute classification
By building a method combining multi-dimensional attribute model and generative large-scale models, the problems of multi-dimensional attribute modeling and low-sample robustness in spacecraft image annotation are solved, and efficient and accurate spacecraft image annotation are achieved.
Patent Information
- Application Number
- CN202510396261.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to take into account the multi-dimensional attribute refinement modeling, robustness of learning in small samples and multi-path results in spacecraft image annotation, resulting in low labeling efficiency, low accuracy and poor robustness.
By constructing a multi-dimensional attribute model, combining benchmark prompt templates and few-sample learning, a generative large model is used for preliminary annotation, and the output results are optimized through thinking chains and multi-path fusion mechanisms, and finally consistency analysis is performed to generate accurate image annotation results.
It improves the labeling accuracy and robustness of spacecraft images in small sample scenarios, reduces manual intervention, improves labeling efficiency and accuracy, and is suitable for complex and changeable spacecraft image recognition and classification.
Smart Images

Figure CN120259814A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to an image annotation method, system, device and medium for classifying spacecraft target types and attributes. Background Art
[0002] When performing fine-grained classification (FGIC) and multi-attribute labeling of spacecraft images, current research mainly faces the following technical difficulties: ① Insufficient multi-dimensional feature representation: Space targets (such as spacecraft) are highly complex in terms of appearance, function, orbital environment, etc. Existing studies often focus on a single attribute, feature or task, and it is difficult to take into account multiple dimensions such as key components, geometric features, lighting conditions, and attitude changes; therefore, they perform poorly in fine-grained, multi-attribute spacecraft labeling. ② Lack of data sets and expensive labeling costs: Compared with vehicles or general targets, the number of publicly available spacecraft images is limited, and professional labeling is difficult. Most existing data sets only have simple or single-dimensional labels, which is difficult to support refined, multi-attribute labeling needs. ③ Poor robustness of few-sample labeling strategies: Traditional deep learning methods often rely on fixed labels or single-path reasoning, which cannot adapt to the variability of spacecraft lighting, distance, and attitude; in few-sample scenarios, overfitting or reasoning bias is more likely to occur, resulting in insufficient recognition accuracy and stability.
[0003] ④ Existing research cannot directly cover the core requirements of this invention: CN118446314A mainly focuses on "intention recognition" and safety warning during the movement of spacecraft, and relies on textual or numerical orbital information; if the multi-dimensional visual attributes of the spacecraft are annotated in a graphic mode, the solution is difficult to directly migrate and cannot deeply explore fine-grained components and geometric features.
[0004] CN118172531A focuses on “event information extraction” and multimodal causal relationships, and is suitable for multimodal event element analysis; however, it has no direct means for the classification and semi-supervised labeling of fine visual elements such as key components of spacecraft, lighting, geometric shapes, etc., making it difficult to meet the robustness requirements in scenarios with few samples.
[0005] CN118364801A mainly discusses the automated iteration and template matching of large model text prompts, and the need for graphic and text annotation of fine-grained attributes such as spacecraft appearance, key components, and lighting. However, there is a lack of corresponding multi-attribute modeling and few-sample fusion strategies, making it difficult to balance accuracy and robustness.
[0006] Even if one attempts to combine the ideas of the above three patents, it can only provide some reference in parts such as text reasoning, causal analysis, or prompt templates, and it is difficult to form a complete technical chain that "takes into account visual fine-grainedness, multi-attribute annotation, and few-shot robustness". They lack comprehensive modeling of multi-dimensional visual features such as key components, geometric shapes, and lighting of spacecraft, and have not established a highly robust and accurate annotation process, so they cannot solve the problems that the present invention attempts to crack.
[0007] In summary, there is still a lack of an effective overall solution for image annotation of spacecraft types and attributes, which needs to simultaneously meet the following requirements: refined modeling of a multi-dimensional attribute system; few-shot learning with context in large language models; result fusion and consistency verification of multi-path (multiple inferences). To maintain annotation efficiency, stability, and high accuracy in highly complex space scenes with variable lighting and attitudes. Summary of the Invention
[0008] Aiming at the problem in the prior art that there is a lack of an overall solution for spacecraft image annotation that takes into account refined modeling of multi-dimensional attributes, few-shot learning robustness, and multi-path result fusion. The present invention provides an image annotation method for spacecraft target type and attribute classification, which combines a semi-supervised annotation method of multi-attribute modeling and context learning, and uses a generative large model to improve the annotation accuracy and robustness of spacecraft images in few-shot scenarios.
[0009] To achieve the above object, the present invention provides the following technical solutions.
[0010] In a first aspect, the present invention provides an image annotation method for spacecraft target type and attribute classification, including: Constructing a multi-dimensional attribute model based on key elements of a spacecraft image; Based on the multi-dimensional attribute model, constructing a benchmark prompt template, obtaining a preliminary output result through the benchmark prompt template, and optimizing the benchmark prompt template and the preliminary output result through few-shot learning to obtain a first output result; Outputting and supervising the thinking process of the first output result through a chain of thought, and obtaining a second output result through an optimized prompt template; Generating an output result dataset according to the first output result, the second output result, and several output results; Performing consistency analysis on the output result dataset to obtain an image annotation result; Performing performance verification on the image annotation result to obtain an image annotation result for spacecraft target type and attribute classification.
[0011] As a further improvement of the present invention, the constructing a multi-dimensional attribute model based on key elements of a spacecraft image includes: Establish a multi-dimensional attribute label system based on the appearance features, key components, lighting conditions, and mission types of spacecraft in the spacecraft images; Construct a multi-dimensional attribute model according to the multi-dimensional attribute label system.
[0012] As a further improvement of the present invention, based on the multi-dimensional attribute model, construct a benchmark prompt template, obtain a preliminary output result through the benchmark prompt template, and optimize the benchmark prompt template and the preliminary output result through few-shot learning to obtain a first output result, including: Obtain a benchmark prompt template according to the multi-dimensional attribute model; Based on the spacecraft target image and the benchmark prompt template, and using a large language model to perform preliminary annotation on the space target image to obtain a preliminary output result; Specify the output format and optional values in the benchmark prompt template, and guide the generative large model to understand the annotation requirements through few-shot examples, and optimize the preliminary output result to obtain a first output result.
[0013] As a further improvement of the present invention, output supervision is performed on the thinking process of the first output result through the chain of thought, and a second output result is obtained by optimizing the prompt template, including: Obtain the reasoning path of the thinking process based on the benchmark prompt template during the generation of the first output result, that is, the reasoning path for generating the first output result; Use the prompt words of the chain of thought to optimize the prompt template, and perform output supervision on the reasoning path for generating the first output result to obtain a second output result.
[0014] As a further improvement of the present invention, generate an output result dataset according to the first output result, the second output result, and several output results, including: Use several recognition models to perform recognition processing on the spacecraft target image, obtain the intermediate reasoning process and reasoning results when several different recognition models recognize the spacecraft target image, and then, through a multi-path fusion mechanism, obtain several reasoning paths to form several output results; Generate an output result dataset according to the first output result, the second output result, and several output results.
[0015] As a further improvement of the present invention, perform consistency analysis on the output result dataset to obtain an image annotation result, including: According to the self-consistent method, fuse the multi-attribute annotation results of the reasoning paths of the output results in the output result dataset to obtain an image annotation result.
[0016] As a further improvement of the present invention, perform performance verification on the image annotation result to obtain an image annotation result for spacecraft target type and attribute classification, including: The performance of the image annotation results is verified according to the average accuracy of a single image and the overall perfect accuracy. The data that passes the verification is the hint template for the classification of spacecraft target types and attributes and the image annotation results.
[0017] In a second aspect, the present invention provides an image annotation system for classifying spacecraft target types and attributes, including: A multi-attribute model construction module: used to construct a multi-dimensional attribute model based on the key elements of spacecraft images; A first output result acquisition module: used to construct a benchmark hint template based on the multi-dimensional attribute model, obtain a preliminary output result through the benchmark hint template, and optimize the benchmark hint template and the preliminary output result through few-shot learning to obtain a first output result; A second output result acquisition module: used to perform output supervision on the thinking process of the first output result through a chain of thought, and obtain a second output result through an optimized hint template; An output result set generation module: used to generate an output result data set according to the first output result, the second output result, and several output results; An image annotation result module: used to perform consistency analysis on the output result data set to obtain an image annotation result; A verification annotation structure module: used to perform performance verification on the image annotation result to obtain an image annotation result for the classification of spacecraft target types and attributes.
[0018] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the image annotation method for classifying spacecraft target types and attributes are implemented.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the image annotation method for classifying spacecraft target types and attributes are implemented.
[0020] In a fifth aspect, the present invention provides a computer program product including computer instructions, and when the computer instructions are executed by a processor, the steps of the image annotation method for classifying spacecraft target types and attributes are implemented.
[0021] Compared with the prior art, the present invention has the following beneficial effects: The present invention realizes the refined description and modeling of spacecraft image features by constructing a multi-dimensional attribute model based on the key elements of spacecraft, overcoming the limitation that traditional methods are difficult to take into account the detailed expression of multi-dimensional attributes when dealing with complex spacecraft images. The introduction of the multi-dimensional attribute model enables image annotation to capture more comprehensively the key information such as the shape, structure, and function of the spacecraft, providing a richer and more accurate feature basis for subsequent image recognition and classification. On this basis, the present invention integrates few-shot learning techniques to optimize the preliminary output results, greatly improving the learning efficiency and annotation accuracy under the condition of limited samples. It effectively solves the problem of scarce samples faced in spacecraft image annotation. Even in scenarios with limited data resources, it can maintain high annotation robustness and generalization ability, opening up new possibilities for spacecraft image analysis and application.
[0022] Furthermore, the present invention uses the chain of thought technique to identify and output the deep thinking process of the preliminarily optimized output results, enhancing the logic and coherence of the annotation results. It also realizes the mining and utilization of implicit information in the image by simulating the reasoning process of humans, improving the precision and reliability of the annotation. The application of the chain of thought makes the annotation process no longer a simple feature matching and classification, but incorporates a higher level of cognition and understanding, which is particularly important for dealing with complex and variable spacecraft images. The second output result generated through this process complements the first output result, jointly constituting a more comprehensive and in-depth annotation information set.
[0023] Furthermore, the present invention also sets up a consistency analysis mechanism for the output result data set. This mechanism can automatically verify and integrate multiple output results to ensure the consistency and accuracy of the final annotation results. This step not only reduces the need for human intervention and improves the annotation efficiency, but also effectively avoids the accumulation of errors and contradictions in the annotation process through algorithm-level optimization, providing a solid guarantee for the accurate annotation of spacecraft images. In addition, the consistency analysis of the output result data set helps to discover and correct potential problems in the annotation process, providing valuable data support for subsequent algorithm optimization and model improvement. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. In the drawings: Figure 1 is a schematic flow chart of an image annotation method for spacecraft target type and attribute classification according to the present invention; Figure 2 is a schematic diagram of the overall process and functional structure of an image annotation method for spacecraft target type and attribute classification according to the present invention Figure 3Schematic diagram of a multi-attribute labeling system for an image annotation method for spacecraft target type and attribute classification according to the present invention Figure 4 The following are the few sample example images and corresponding annotations of the present invention; wherein: Figure (a) is Example 1; Figure (b) is Example 2; Figure 5 A schematic diagram of a process in an embodiment of the present invention; Figure 6 A schematic diagram of the structure of an image annotation system for classifying spacecraft target types and attributes according to the present invention; Figure 7 Schematic diagram of the structure of the embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. The described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present invention.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0027] Aiming at the problem that the existing spacecraft image annotation technology lacks an overall solution that takes into account the multi-dimensional attribute refinement modeling, few-sample learning robustness and multi-path result fusion. The present invention provides an image annotation method for spacecraft target type and attribute classification, such as Figure 1 As shown, the method includes: S100: Constructing a multi-dimensional attribute model based on key elements of spacecraft images; S200: constructing a benchmark prompt template based on the multi-dimensional attribute model, obtaining a preliminary output result through the benchmark prompt template, and optimizing the benchmark prompt template and the preliminary output result through few-sample learning to obtain a first output result; S300: Output supervision of the thinking process of the first output result is performed through the thinking chain, and the second output result is obtained by optimizing the prompt template; S400: Generate an output result data set according to the first output result, the second output result and the plurality of output results; S500: Perform consistency analysis on the output result dataset to obtain the image annotation result; S600: Perform performance verification on the image annotation result to obtain the image annotation result for spacecraft target type and attribute classification.
[0028] The present invention combines a semi-supervised annotation method of multi-attribute modeling and context learning, and utilizes a generative large model to improve the annotation accuracy and robustness of spacecraft images in few-shot scenarios.
[0029] The following further explains the present invention with reference to specific drawings.
[0030] An image annotation method for spacecraft target type and attribute classification according to the present invention, as Figure 2 shown, specifically includes the following steps: S1: Based on the key elements of spacecraft images, such as type, components, lighting, geometric shape, etc., establish a multi-dimensional attribute model; According to elements such as spacecraft appearance features, key components, lighting conditions, mission types, etc., establish a multi-dimensional attribute label system, and define optional values and default placeholders such as "others" for each attribute. Among them, spacecraft appearance features include but are not limited to model categories and geometric shapes.
[0031] When researching the "image annotation for spacecraft target type and attribute classification" task, based on multi-faceted characteristics such as spacecraft appearance structure, key components, lighting environment, functional uses, etc., form multi-dimensional modeling to obtain a multi-dimensional attribute model.
[0032] To enhance the completeness and operability of the annotation result, the following division principles and label systems are adopted in the multi-attribute modeling stage.
[0033] (1) Consider both the local and the whole. Spacecraft are often composed of multiple separable or relatively independent functional components (such as solar panels, antennas, robotic arms, etc.), and at the same time, the overall shape (such as cylinder, capsule, modular combination) also has important reference value for task type judgment. When performing multi-attribute modeling in the present invention, both the type, quantity, distribution position, etc. of "Key Components" are defined, and the overall geometric shape (Body Shape) is separately described. This "local + whole" strategy helps to balance details and global consistency in subsequent annotation.
[0034] (2) Optional fields and knowledge integration. Spacecraft of different models may have special components or different shapes. To avoid omission or confusion, the present invention adopts an optional field dictionary and combines the experience of domain experts to provide predefined values for common component names, overall shapes, or mission types. For attributes that cannot be judged or are out of range, "others" is used as a placeholder, which not only unifies the output format but also retains room for expansion.
[0035] (3) Gradual granularity division. When "key component = robotic arm" is recognized, its position (top, bottom, symmetric distribution, etc.) can be further judged to form a finer-grained annotation. At the same time, the attributes of the main body and components can be split into multiple sub-fields as needed (such as component shape, quantity, size, distribution position) to adapt to different levels of annotation requirements.
[0036] Based on the above principles, the main attributes required for spacecraft image annotation are finally divided into six dimensions, including: target type (type), mission type (mission_types), key components (key_components), overall geometric characteristics (positional_scale_and_geometric_characteristics), illumination conditions (illumination_conditions), and confidence level (confidence_level). Table 1 and Figure 3 gives an exemplary label system: Table 1 Exemplary label system
[0037] If there are multiple spacecraft targets in the image, multiple sets of structured fields (such as multiple spacecraft_targets entries) can be generated accordingly in the subsequent annotation output, and records are made separately at the levels of key components, illumination, etc.
[0038] In this step, based on the knowledge of spacecraft domain experts, a multi-dimensional attribute label covering satellite model, mission type, key components, geometric form, illumination conditions, confidence level, etc. is constructed to form a multi-dimensional attribute model, and a benchmark prompt template is formed according to the multi-dimensional attribute model.
[0039] Combining the spacecraft target image and the benchmark prompt template, the large language model (LLM) can directly perform preliminary annotation on the space target image to obtain a preliminary output result, reducing the dependence on large-scale annotation data. Compared with the existing technologies that only focus on single attributes or orbital information, this multi-attribute system can model the shape characteristics and application functions of spacecraft more comprehensively.
[0040] S2: Specify the output format and optional values in the baseline prompt template, and use few-shot examples to guide the generative large model to understand the annotation requirements; Based on the predefined output format, such as the JSON structure, compile the baseline prompt template and add few-shot examples to guide the generative large model to understand and perform multi-attribute annotation.
[0041] Specifically, the design of the graphic-modal prompt template is as follows: After completing multi-attribute modeling, the present invention combines the baseline prompt template with in-context learning (ICL) to guide the generative large model to perform multi-attribute annotation of spacecraft images, obtaining a first output result. Further, the chain of thought (CoT) and multi-path self-consistent fusion (SC) mechanisms are used to process the first output result to improve the accuracy and stability, obtaining a second output result and several output results.
[0042] Specifically, the baseline prompt template is as follows: (1) JSON output structure and optional fields. This application uses a fixed-format JSON as the annotation output and strictly limits the optional values in fields such as "type", "mission_types", "key_components", etc. If it cannot be determined, use "others" as a placeholder to constrain the output normativity of the large model.
[0043] (2) Task description and constraints. At the beginning of the prompt template, instruct the large model: "Please perform spacecraft attribute recognition, output in JSON format; only select from the given list of values, do not add extra fields; keep the result concise and do not write explanatory sentences." Such constraints can reduce the situation where the model generates redundant descriptions.
[0044] (3) Exemplary baseline prompt template. In simplified form (the JSON structure is omitted), it includes fields such as target type, mission type, key components, overall geometric features, lighting conditions, and confidence, and is accompanied by descriptions of several optional values. The present invention is not limited to using this format and can be adjusted according to application requirements in actual implementation.
[0045] Specifically, the few-shot examples are as follows: To enable the large model to fully understand the annotation requirements, the present invention prepares several few-shot examples based on the baseline prompt template. Such as Figure 4 shown, each example includes the target image and the corresponding attribute annotation description. Through in-context learning, the accuracy and robustness of the large model during testing can be significantly improved.
[0046] Example 1: Image name: "Cargo Spaceship.png"; ① Target object: "Model": "Cygnus Spaceship", "Quantity": 1; ② Mission type: "Cargo"; ③ Key components: i) "Component type": "Sub-body", "Quantity": 1, "Distribution location": "Other", "Component shape": "Cylindrical", ii) "Component type": "Solar panel", "Quantity": 2, "Distribution location": "Symmetrically on both sides", "Component shape": "Circular"; ④ Physical dimensions and overall ratio: "Main body shape": "Cylindrical", "Overall ratio": "Normal"; ⑤ Lighting condition: "Good"; ⑥ Confidence level: "High".
[0047] Example 2: Image name: "Hubble Telescope.png"; ① Target object: "Type": "Space Telescope", "Quantity": 1; ② Mission type: "Astronomical observation"; ③ Key components: i) "Component type": "Sub-body", "Quantity": 1, "Distribution location": "Middle", "Component shape": "Cylindrical", ii) "Component type": "Solar panel", "Quantity": 2, "Distribution location": "Symmetrically on both sides", "Component shape": "Rectangular"; ④ Physical dimensions and overall ratio: "Main body shape": "Cylindrical", "Overall ratio": "Normal"; ⑤ Lighting condition: "Good"; ⑥ Confidence level: "High".
[0048] S3: Adopt a mechanism that combines chain of thought and multi-path to improve the adaptability and annotation accuracy for different image scenarios; During the annotation process, first let several different recognition models output their intermediate reasoning processes, that is, the chain of thought, and then obtain multiple independent reasoning paths by calling the same or different models multiple times to form richer candidate annotation results and generate several output results; among them, several different recognition models include but are not limited to generative models, as well as the same model or different models used several times.
[0049] When optimizing the prompt template in the present invention, a chain of thought (CoT) is introduced, allowing the large model to explicitly give the intermediate reasoning process before outputting the final JSON. The "Chain-of-thought:" part can be used by the system or manually to check the logical rationality. Once an error is found, it can be corrected before generating the JSON, thereby greatly improving the credibility and interpretability of the results.
[0050] (1)Step-by-step reasoning output: In the method of the present invention, before generating an answer, the large model will, according to the newly added prompt words, first output a concise reasoning process, mark it as "Chain-of-thought: ……", and then clearly end at the end of the paragraph; subsequently, the large model will output the words "Final JSON:" and the legal JSON content as the final available annotation result of the system. At this time, the chain-of-thought text will not be incorporated into the final JSON file, but will be recorded in the log or intermediate debugging information for analysis and backtracking; through this way of "outputting the reasoning process first and then the final result", the present invention can verify the reasoning process of the large model. Once obvious errors or self-contradictory logics are found in the chain of thought, corrections or regenerations can be made before generating the final JSON.
[0051] (2)Sub-task division: Identify the main objects in the target image (such as satellites, spacecraft, or space stations), and point out the basis for this identification judgment in the "Chain-of-thought:" paragraph; analyze the task type of the target (such as manned, observation, cargo, etc.), and briefly explain the inference reason in the chain of thought; label the key components of the target (such as solar panels, antennas, robotic arms, etc.) and their distribution characteristics, and record the reasoning logic or identification steps in the chain-of-thought text; judge the lighting conditions (such as good or low light) and confidence level of the target, and if there is no conflict, then enter the link of generating "Final JSON:".
[0052] For each input image, the present invention proposes a generation strategy of multi-path reasoning. The so-called "multi-path" not only refers to multiple reasoning calls on the same model, but also includes calling multiple different generative large models (such as ChatGPT, Claude, Gemini, etc.) for the same prompt template and input data, or using the same large model but different models (such as ChatGPT-4o, o1-preview, Gemini-2.0-Flash, Gemini-1.5-Flash, Claude 3.5 Sonnet, Claude 3.5 Haiku, etc.) to generate reasoning results respectively. In this way, more diverse and independent annotation reasoning results can be obtained, providing more sufficient candidate paths for subsequent consistency analysis.
[0053] The specific process includes: Generating multiple reasoning paths: Call the generative large model to generate multiple reasoning paths using the same input multiple times. Each reasoning path is independently generated and contains six-dimensional annotation information of the type and attributes of the target image.
[0054] Therefore, based on few-shot examples and chain-of-thought supervision, this step decomposes complex fine-grained annotation tasks into multi-step reasoning. Subsequently, the Self-Consistency method is introduced to fuse and perform consistency verification on the reasoning results of multiple paths: the multiple paths can come from multiple independent calls of the same large model, or from different large models and multiple versions of the same model; by comparing and fusing multiple reasoning chains, the single-path reasoning error is reduced, and the annotation stability and accuracy are improved. Compared with existing technologies, this link greatly enhances the adaptability to uncertainties such as illumination and pose changes in few-shot scenarios.
[0055] S4: Finally, with the assistance of a small amount of manual verification or prior information, an efficient semi-supervised image annotation process is achieved.
[0056] Specifically, consistency analysis and voting are performed on the attribute annotation results of multiple paths to automatically generate the final merged annotation; if obvious conflicts occur, they can be corrected by combining a small amount of manual verification, thus completing the semi-supervised annotation process.
[0057] After completing the prompt template design and multi-path fusion, in order to reduce the cost of manual annotation and ensure high accuracy under few-shot conditions, a "semi-supervised" strategy is adopted to annotate the attributes of spacecraft images, and various evaluation methods are combined to quantify its effect.
[0058] S41: Self-Consistency When the large model is called multiple times on the same image or different large model versions are used, multiple "reasoning paths" are generated. The present invention introduces a self-consistent fusion mechanism to perform attribute-level voting and merging on these paths: when analyzing multiple reasoning paths, the Self-Consistency (SC) method is introduced, aiming to fuse the multi-attribute annotation results of each path to obtain a more robust and accurate final output. This process is automatically completed by the large model after referring to the attribute results of different paths, and specifically includes the following steps: For each attribute field output by each reasoning path (such as target type, task type, number and location of key components, etc.), its rationality and similarity to other paths are evaluated. The large model will be guided in the prompt to compare the key attributes of each path. If most paths give the same or similar values for a certain attribute, it means that the attribute has higher consistency.
[0059] For each attribute field (such as "whether manned", "number of key components", "illumination condition", etc.), the large model "votes" according to the frequency of the same or similar outputs in the paths. Based on the voting results, the large model will regard the attribute value with the highest frequency or confidence as the "candidate result", and further comprehensively consider other prompt information and semantic context to automatically judge and correct conflicting or divergent items.
[0060] After completing the evaluation and voting of attributes, the large model comprehensively processes all attribute fields and automatically generates a "fusion result" as the final annotation output for the spatial target. This result is not a direct selection of the overall output of a certain path, but the optimal consistent annotation obtained based on the attribute-level analysis and voting combination of multiple paths. For example, if most paths consider the target task type to be "manned" and the large model does not find reasons for conflicts, then the "task type" in the final result adopts the attribute value of "manned". This final fusion result can also be accompanied by a "reason" field to briefly explain the voting and fusion process of each attribute.
[0061] S42: Semi-supervised Annotation (1) Loading and Initialization, load or initialize the generative large model (multiple versions are optional) through the API interface, and prepare the prompt template and optional field dictionary; set few-shot examples and the thought chain prompting method.
[0062] (2) Image Processing and Upstream Detection, if there are multiple spacecraft targets in the image, target detection or cropping can be performed first; generate an encoded or Base64 representation for each image as the input for the subsequent large model.
[0063] (3) Annotation and Fusion, call the large model using the benchmark prompt template to output the preliminary JSON result (including multiple inference paths); perform self-consistent (SC) fusion on the results of multiple paths to form a single candidate annotation; if there are obvious conflicts or indecision in some fields, a small amount of manual verification can be entered, and domain experts or annotators can conduct manual review or supplementation, which is the main manifestation of the "semi-supervised" mechanism.
[0064] (4) Saving and Iteration, save the final annotation result in a unified format (JSON file or database), and can compare it with the prior annotation data to iteratively update the model prompt or few-shot examples to gradually improve the overall performance.
[0065] S43: Performance Verification To more comprehensively evaluate the effectiveness of the context learning method described in the present invention in multi-attribute annotation tasks, the present invention introduces the following two measurement indicators to characterize the correct recognition rate of the model in the single-image dimension and the overall dimension.
[0066] ① Per-Image Average Accuracy (PIAA), for any image , let its ground truth attribute set be , and the model prediction attribute set be After splitting the six-dimensional attributes of the image into discrete "key-value pairs" or tags, the ratio of the number of correct matches to the total number of true values is measured by set operations and defined as follows: (1) Furthermore, on the entire dataset, the arithmetic mean of the single-image accuracies of images is obtained: where is the number of attributes correctly predicted for the th image, is the number of true attributes, and this metric can measure the situation of "partially correct".
[0067] ② All-Attributes-Correct Accuracy (AACA). Under a more strict determination, if all the attributes of the th image match the true values, it is recorded as correct; otherwise, it is recorded as incorrect. Then the overall all-attributes-correct rate on the entire dataset is: (3) This metric requires that all the attributes of an image be correctly recognized by the model to be counted as 1.
[0068] Finally, based on the above metrics, the present invention verified the performance of the "Zero-shot" and "Few-shot" schemes on a public dataset (such as URSO). The results of the ablation experiment are shown in Table 2, and the test set selected contains 450 images randomly selected from 9 spacecraft categories.
[0069] Table 2 Results of the ablation experiment
[0070] Through this method, the effectiveness and superiority of the method of the present invention in the task of spacecraft target type and attribute classification and annotation are verified, demonstrating a significant performance improvement.
[0071] Compared with the prior art, in the cases of few samples, variable lighting environments, etc., the present embodiment can still maintain high annotation robustness and can be rapidly iterated and extended in scenarios such as on-orbit service and fault monitoring. Moreover, the modules in the present application work together to ensure both high precision and robustness, and can significantly reduce the human input, which has good applicability to the spacecraft attribute annotation in application scenarios such as on-orbit service and space situation monitoring. A multi-attribute label system is constructed for fine-grained visual elements such as target type, key components, and lighting conditions, taking into account the spacecraft's shape and mission characteristics, and enhancing the refined understanding of spacecraft images. Compared with the technical solutions based only on orbital features or text prompts, it can capture the diverse appearance and functional attributes of spacecraft more comprehensively. Secondly, through few-sample examples and chain-of-thought supervision, the complex attribute annotation is decomposed into multi-step reasoning, and multiple reasoning paths are fused and consistency verified to reduce the deviation of single-path reasoning; at the same time, the manual correction link is retained to ensure the practicality in the case of few samples. Even in the case of few samples or under harsh conditions such as large changes in lighting, attitude, and altitude, high precision can still be maintained, and it is adaptable to different types of large models and various application scenarios; it not only makes up for the dependence on large-scale data but also can be rapidly extended to other fine-grained space target recognition tasks.
[0072] In summary, the present invention starts from the perspective of spacecraft multi-attribute modeling and in-context learning (ICL), combines few-sample examples, chain-of-thought supervision, and self-consistent multi-path fusion means to form a semi-supervised and high-precision annotation algorithm; in this process, existing studies, whether used alone or in simple combination, cannot provide a complete multi-dimensional visual attribute modeling and few-sample self-consistent fusion mechanism. Based on the comprehensive method proposed by the present invention, the accuracy and efficiency of spacecraft image recognition can be significantly improved in large-scale or complex orbital / lighting environments, and the human annotation cost can be reduced; it has broad application value in diverse space applications such as on-orbit service, fault monitoring, and space situation awareness. The present invention proposes a semi-supervised annotation method combining multi-attribute modeling and in-context learning (ICL) for the multi-attribute classification annotation requirements of spacecraft target images, such as type, key components, lighting conditions, geometric shape, etc., and significantly improves the accuracy and robustness in few-sample scenarios through multi-path fusion. This solution makes full use of the capabilities of generative large models in text-image modality understanding and multi-stage reasoning, which not only reduces the manual annotation cost but also ensures high annotation efficiency and stability in the cases of variable lighting, attitude, and limited data.
[0073] The following further explains and illustrates the present invention with specific embodiments. An image annotation method for spacecraft target type and attribute classification, the specific steps are as follows: S1: Template design. Refer to the baseline prompt template and few-shot examples to guide the large model to understand the typical features of the Hubble Telescope in terms of task type (astronomical observation), key components (sub-body + solar panels), etc.
[0074] S2: Template optimization. Split the annotation task into sub-steps such as identifying objects, task types, key components, etc. through Chain of Thought (CoT); and combine with Self-Consistency (SC) to perform consistency analysis and voting selection on the results of multiple reasoning paths.
[0075] S3: Result fusion and verification. Finally, select the one with the highest consistency among the multi-path reasoning results as the annotation output; as Figure 5 shown, annotation information such as "Type = space telescope", "Key components = solar panel + sub-body", "Task type = astronomical observation", "Lighting conditions = good", "Confidence = high" can be obtained. Aggregating and comparing the generated multiple annotation results can effectively reduce the errors caused by randomness or few-shot samples.
[0076] This method is experimented on the publicly available spacecraft image dataset (such as URSO, etc.), and the accuracy, reasoning efficiency, and robustness of multi-attribute recognition are quantitatively evaluated; preliminary tests show that in the fine-grained spacecraft attribute classification, compared with the baseline prompt template, this invention can improve the annotation accuracy by about 15%, and the semi-supervised annotation also shortens the annotation time.
[0077] It can be seen that this invention first uses the baseline prompt template and few-shot examples to obtain preliminary annotations, then performs consistency processing on multi-path reasoning through Chain of Thought and Self-Consistency, and finally realizes high-accuracy automatic annotation of the Hubble Telescope and even other types of spacecraft images. In practical engineering applications, this invention can also be extended to other space target scenarios with complex appearances and multi-attribute requirements, having good generality and robustness. Through multi-attribute modeling, context prompting, Chain of Thought, and Self-Consistency means, high-accuracy semi-supervised image annotation is completed on the basis of a small amount of manual verification. This method can be widely applied to various fields such as on-orbit servicing of spacecraft, space situation monitoring, and fault diagnosis, while improving the annotation efficiency and accuracy, significantly reducing the labor cost.
[0078] In summary, the present method first performs preliminary annotation of space target images based on Prompt Engineering, combined with Few shot, on the basis of a multi-dimensional attribute template. Subsequently, the complex annotation task is decomposed into multi-stage reasoning through the Chain of Thought (CoT), and multiple independent reasoning paths are generated when calling multiple different large models or different versions of the same large model. For these independent reasoning results, the present invention further utilizes the Self-Consistency (SC) method for automated voting and fusion, integrating multi-path consistency to generate a robust and highly accurate final annotation output. Experimental results show that the present invention can achieve higher accuracy and stability at a relatively low labor cost on public datasets (such as URSO), and is applicable to the automatic annotation requirements of complex space target images in multiple scenarios and tasks.
[0079] The present invention conducts strict performance verification on the image annotation results to ensure that the obtained annotation results not only meet the actual requirements of spacecraft target type and attribute classification, but also maintain stable performance in various application scenarios. The implementation of this link is not only a direct test of the effectiveness of the annotation method, but also an important proof of its feasibility in actual operation. Through performance verification, the method proposed by the present invention has demonstrated excellent application potential in the field of spacecraft image annotation, not only improving the accuracy and efficiency of annotation, but also promoting the overall progress of spacecraft image analysis technology. Therefore, the present invention provides a comprehensive and efficient solution for spacecraft image annotation by combining advanced technologies such as multi-attribute modeling, few-shot learning, chain of thought reasoning, and consistency analysis. This method not only solves the problems of low annotation accuracy, poor robustness, and strong sample dependence in the prior art, but also significantly improves the intelligent level and automation degree of annotation, laying a solid foundation for in-depth research and wide application in the fields of spacecraft image recognition, classification, monitoring, and management. Its innovative technical ideas and remarkable practical effects will undoubtedly have a profound impact on the development of spacecraft image annotation technology, promoting the field to a higher level.
[0080] The second object of the present invention is to propose an image annotation system for spacecraft target type and attribute classification, such as Figure 6 shown, including: A multi-attribute model construction module 100: used to construct a multi-dimensional attribute model based on the key elements of spacecraft images; A first output result acquisition module 200: used to construct a benchmark prompt template based on the multi-dimensional attribute model, obtain a preliminary output result through the benchmark prompt template, and optimize the benchmark prompt template and the preliminary output result through few-shot learning to obtain a first output result; Second Output Result Obtaining Module 300: It is used to perform output supervision on the thinking process of the first output result through the chain of thought, and obtain the second output result by optimizing the prompt template; Output Result Set Generating Module 400: It is used to generate an output result data set according to the first output result, the second output result, and several output results; Image Annotation Result Module 500: It is used to perform consistency analysis on the output result data set to obtain an image annotation result; Annotation Structure Verification Module 600: It is used to perform performance verification on the image annotation result to obtain the image annotation result of the spacecraft target type and attribute classification.
[0081] As Figure 7 shown, the third object of the present invention is to provide an electronic device, which includes: a processor 701, a memory 702, and a display screen 703. Among them, the memory 702 and the display screen 703 are both connected to the processor 701, such as through a bus 704. Optionally, the electronic device may further include a transceiver 705. It should be noted that in practical applications, the transceiver 705 is not limited to one, and the structure of this electronic device does not constitute a limitation to the embodiments of the present application.
[0082] The processor 701 may be a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 701 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0083] The bus 704 may include a path for transmitting information between the above components. The bus 704 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 704 may be divided into an address bus, a data bus, a control bus, etc.
[0084] The memory 702 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0085] The memory 702 is used to store the application program code for implementing the solution of this application and is controlled by the processor 701 for execution. The processor 701 is used to execute the application program code stored in the memory 702 to implement the content shown in the foregoing method embodiments.
[0086] Figure 7 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0087] The fourth object of the present invention is to provide a computer-readable storage medium storing a computer program, on which a computer program is stored. When the program is executed by a processor, it implements each process of the foregoing Figure 1 shown method embodiments. For example, a memory including instructions, and the above instructions can be executed by the processor of the electronic device to complete the above method.
[0088] A computer-readable storage medium can be a tangible device that holds and stores instructions used by an instruction execution device. A computer-readable storage medium can be, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above. Specifically, a computer-readable storage medium can be a portable computer disk, a hard disk, a USB flash drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, an optical disc, a magnetic disk, a mechanical coding device, and any combination of the above.
[0089] The fifth object of the present invention is to provide a computer program product, including computer instructions, which when executed by a processor, implement each process of the method embodiment shown above Figure 1 and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0090] Many embodiments and many applications other than the examples provided will be apparent to those skilled in the art upon reading the above description. Therefore, the scope of this teaching should not be determined with reference to the above description, but should be determined with reference to the full scope of the foregoing claims and the equivalents thereof. For the sake of completeness, all articles and references, including patent applications and published announcements, are incorporated herein by reference. The omission of any aspect of the subject matter disclosed herein in the foregoing claims is not intended to abandon such subject matter, nor should it be considered that the applicant has not considered such subject matter as part of the disclosed inventive subject matter.
[0091] The above is a further detailed description of the present invention. It cannot be determined that the specific implementation of the present invention is limited thereto. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope determined by the claims submitted for the present invention.
Claims
1. An image annotation method for classifying the types and attributes of spacecraft targets, characterized in that, Including: Construct a multi-dimensional attribute model based on the key elements of spacecraft images; Construct a benchmark prompt template based on the multi-dimensional attribute model, obtain a preliminary output result through the benchmark prompt template, and optimize the benchmark prompt template and the preliminary output result through few-shot learning to obtain a first output result; Output supervise the thinking process of the first output result through the chain of thought, and obtain a second output result by optimizing the prompt template; Generate an output result dataset according to the first output result, the second output result and several output results; Conduct consistency analysis on the output result dataset to obtain an image annotation result; Conduct performance verification on the image annotation result to obtain an image annotation result for spacecraft target type and attribute classification.
2. The image annotation method for classifying the target type and attributes of a spacecraft according to claim 1, characterized in that, The constructing a multi-dimensional attribute model based on the key elements of spacecraft images includes: Establish a multi-dimensional attribute label system based on the appearance features, key components, lighting conditions, and mission types of spacecraft in spacecraft images; Construct a multi-dimensional attribute model according to the multi-dimensional attribute label system.
3. The image annotation method for classifying the target type and attributes of a spacecraft according to claim 1, wherein The constructing a benchmark prompt template based on the multi-dimensional attribute model, obtaining a preliminary output result through the benchmark prompt template, and optimizing the benchmark prompt template and the preliminary output result through few-shot learning to obtain a first output result includes: Obtain a benchmark prompt template according to the multi-dimensional attribute model; Based on the spacecraft target image and the benchmark prompt template, and use a large language model to preliminarily annotate the space target image to obtain a preliminary output result; Specify the output format and optional values in the benchmark prompt template, and guide the generative large model to understand the annotation requirements through few-shot examples, and optimize the preliminary output result to obtain a first output result.
4. An image annotation method for classifying the target type and attributes of a spacecraft, as claimed in claim 1, wherein The output supervising the thinking process of the first output result through the chain of thought and obtaining a second output result by optimizing the prompt template includes: Obtain the reasoning path of generating the first output result, which is the thinking process based on the benchmark prompt template during the process of generating the first output result; Use the prompt words of the chain of thought to optimize the prompt template, and output supervise the reasoning path of generating the first output result to obtain a second output result.
5. The image annotation method for classifying the target type and attributes of a spacecraft according to claim 1, characterized in that, The generating an output result dataset according to the first output result, the second output result and several output results includes: Use several recognition models to perform recognition processing on the spacecraft target image, obtain the intermediate reasoning process and reasoning results when several different recognition models recognize the spacecraft target image, and then, through a multi-path fusion mechanism, obtain several reasoning paths to form several output results; Generate an output result dataset according to the first output result, the second output result and several output results.
6. A method for image annotation for classifying the target type and attributes of a spacecraft, as claimed in claim 1, wherein The conducting consistency analysis on the output result dataset to obtain an image annotation result includes: According to the self-consistency method, fuse the multi-attribute annotation results of the reasoning paths of the output results in the output result dataset to obtain an image annotation result.
7. A method for image annotation for classifying the target type and attributes of a spacecraft, as claimed in claim 1, wherein The conducting performance verification on the image annotation result to obtain an image annotation result for spacecraft target type and attribute classification includes: The performance of the image annotation results is verified according to the average accuracy of a single image and the overall perfect accuracy. The data passing the verification are the hint templates for the classification of spacecraft target types and attributes and the image annotation results.
8. An image annotation system for classifying the target types and attributes of spacecraft, characterized in that, Including: Construct a multi-attribute model module: used to construct a multi-dimensional attribute model based on the key elements of spacecraft images; Obtain the first output result module: used to construct a benchmark hint template based on the multi-dimensional attribute model, obtain a preliminary output result through the benchmark hint template, and optimize the benchmark hint template and the preliminary output result through few-shot learning to obtain the first output result; Obtain the second output result module: used to output and supervise the thinking process of the first output result through the chain of thought, and obtain the second output result through an optimized hint template; Generate an output result set module: used to generate an output result data set according to the first output result, the second output result, and several output results; Image annotation result module: used to perform consistency analysis on the output result data set to obtain the image annotation result; Verify the annotation structure module: used to perform performance verification on the image annotation result to obtain the image annotation result for the classification of spacecraft target types and attributes.
9. An electronic device, characterized in that, Including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for image annotation for the classification of spacecraft target types and attributes according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the method for image annotation for the classification of spacecraft target types and attributes according to any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Multi-modal event intelligent sensing method based on large model prompt project
CN118172531A
Target spacecraft intention recognition method and system based on large language model
CN118446314A
Cited By
Method for identifying aircraft type in satellite image, electronic equipment, storage medium and program product
CN120894704A
Low-altitude remote sensing image-oriented ground feature fine-grained attribute extraction method
CN121661531A
Image multi-attribute automatic pre-labeling method based on multi-modal large language model
CN121661646A