Secondary individual identification method and system driven by dynamic knowledge
Through the dynamic knowledge-driven secondary individual recognition method, combined with pre-trained and customized training object detection model and large language model, the problem of individual recognition in dynamic scenarios is solved, the semantic cognition and recognition of new individuals is realized, the model training process is simplified, and the accuracy and real-timeness of recognition are improved.
Patent Information
- Application Number
- CN202510471443.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to achieve efficient and accurate individual recognition in dynamic scenarios, especially under the influence of factors such as changes in lighting conditions, viewing angle changes and occlusion. Traditional methods lack semantic information and cannot distinguish individuals. The data set preparation period is long, making it difficult to identify new individuals.
Using a dynamic knowledge-driven secondary individual recognition method, through incremental feature updates and object detection model custom training, combined with pre-trained and customized training object detection models, a large language model or search engine is used to generate semantic categories and probability values, automatically generate data sets and update models.
The semantic cognition and recognition of new individuals in dynamic scenarios is realized, the training process of the object detection model is simplified, the training difficulty and cycle are reduced, and the accuracy and real-time nature of individual recognition are improved.
Smart Images

Figure CN120374952A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robotics, and in particular, relates to a secondary individual recognition method and system driven by dynamic knowledge. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] The ability of individual recognition is crucial for a robot to build a map of a dynamic scene. During the process of topological mapping of a dynamic scene, the robot needs to observe the environment, find key target individuals that can distinguish the environment and describe the scene, and establish a scene topological structure with independent individual attributes. The robot needs to discover objects from different perspectives to achieve position discovery and object search.
[0004] For the recognition of individuals, traditional feature matching methods such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), ORB, XFeat, etc., mainly rely on significant feature points in images. Factors such as lighting conditions, perspective changes, and occlusions will affect the detection of feature points and the generation of descriptors, thus affecting the accuracy of object recognition; traditional methods lack semantic information and cannot achieve the recognition of semantic attributes through the robot itself. Feature-based recognition methods lack the ability to distinguish individuals and can only recognize the position point information with special light and shadow features in the scene, and objects need to be marked manually.
[0005] With the development of deep learning, new methods such as SuperPoint and D2-Net that have emerged in recent years combine key point detection and feature descriptor calculation through data learning to improve efficiency and performance, but it is still difficult to achieve high levels of real-time performance and accuracy simultaneously. For popular object recognition model frameworks such as YOLO, they can recognize object categories but still cannot distinguish at the individual level. And the method of establishing an individual dataset and training the model has many disadvantages, such as a long preparation period for the dataset, inability to handle new individuals that may be added at any time, and after new individuals appear, retraining the model requires a lot of resources and time. Summary of the Invention
[0006] In order to solve the technical problems existing in the above background art, the present invention provides a secondary individual recognition method and system driven by dynamic knowledge, which can realize the semantic cognition of new individuals and the recognition of new individuals for individuals that are not recognized or misrecognized by the target detection model in the early stage through incremental feature update, individual screening, and customized training of the target detection model.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] The first aspect of the present invention provides a secondary individual recognition method driven by dynamic knowledge, which includes:
[0009] Obtain a scene image and segment it into several screening regions;
[0010] For each screening region, input it into a pre-trained object detection model and a custom-trained object detection model respectively, obtain semantic categories and probability values, and select the semantic category corresponding to the high probability value to obtain semantic nodes. Feature sub-extraction is performed on each semantic node. For a screening region, if the probability values obtained by the pre-trained object detection model and the custom-trained object detection model are both less than the threshold, the screening region is classified as an unknown class;
[0011] For a certain semantic node in the current observation, if there already exists a semantic node with the same semantic category, match the semantic nodes based on the feature sub. If the match is successful, merge the feature subs of the two semantic nodes and update the incremental feature of the semantic node;
[0012] As the acquisition perspective of the scene image moves, for a certain semantic category, if the intersection over union of two semantic nodes in consecutive observations is greater than the threshold, determine whether to add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category;
[0013] For the unknown class, generate semantic categories and probability values through a large language model or a search engine to train the object detection model and update the custom-trained object detection model.
[0014] Further, the method for determining whether to add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category is: calculate the ratio of the hit amount of the feature sub of the semantic node in the current observation in the incremental feature to the number of feature subs in the incremental feature. If the ratio is less than a certain value, add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category.
[0015] Further, the segmentation includes: segment the scene image using a segmentation model to divide the scene into several screening regions; for the regions not segmented by the segmentation model, cluster them into multiple screening regions.
[0016] Further, the step of matching semantic nodes based on feature subs includes: calculate the number of matching feature subs of two semantic nodes with the same semantic category, calculate the ratio of the number of matching feature subs to the number of feature subs of the semantic node of this semantic category in the current observation to obtain the feature density. If the feature density meets the set conditions, the two semantic nodes are successfully matched.
[0017] The second aspect of the present invention provides a secondary individual recognition system driven by dynamic knowledge, which includes:
[0018] An image segmentation module, which is configured to: obtain a scene image and segment it into several screened regions;
[0019] A semantic node generation module, which is configured to: for each screened region, input a pre-trained object detection model and a custom-trained object detection model respectively, obtain a semantic category and a probability value, and select the semantic category corresponding to the high probability value to obtain a semantic node, perform feature sub-extraction on each semantic node. For a screened region, if the probability values obtained by the pre-trained object detection model and the custom-trained object detection model are both less than the threshold, the screened region is classified as an unknown class;
[0020] An individual node matching module, which is configured to: for a certain semantic node in the current observation, if there already exists a semantic node with the same semantic category, match the semantic nodes based on the feature sub, and if the match is successful, merge the feature subs of the two semantic nodes and update the incremental feature of the semantic node;
[0021] An incremental feature update module, which is configured to: as the acquisition perspective of the scene image moves, for a certain semantic category, if the intersection over union of two semantic nodes in consecutive observations is greater than the threshold, determine whether to add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category;
[0022] A custom training module, which is configured to: for unknown classes, generate semantic categories and probability values through a large language model or a search engine to train the object detection model and update the custom-trained object detection model.
[0023] Further, the method for determining whether to add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category is: calculate the ratio of the hit quantity of the feature sub of the semantic node in the current observation in the incremental feature to the number of feature subs in the incremental feature. If the ratio is less than a certain value, add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category.
[0024] Further, the segmentation includes: segmenting the scene image using a segmentation model to divide the scene into several screened regions; clustering the regions not segmented by the segmentation model into multiple screened regions.
[0025] Further, the step of matching the semantic nodes based on the feature sub includes: calculating the number of matching feature subs of two semantic nodes with the same semantic category, calculating the ratio of the number of matching feature subs to the number of feature subs of the semantic nodes with this semantic category in the current observation to obtain a feature density. If the feature density meets the set conditions, the two semantic nodes are successfully matched.
[0026] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a secondary individual recognition method driven by dynamic knowledge as described above are implemented.
[0027] The fourth aspect of the present invention provides a computer device, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor. When the processor executes the program, the steps in a secondary individual recognition method driven by dynamic knowledge as described above are implemented.
[0028] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0029] For individuals that are not recognized or misrecognized by the target detection model in the early stage, the present invention can realize new individual semantic cognition and new individual recognition through incremental feature update, individual screening, and customized training of the target detection model.
[0030] Through the recognition of the target area by the large language model or the search engine, the present invention can automatically generate a data set and annotation information, greatly simplifying the training process of the target detection model.
[0031] Through the parallel target detection models, the present invention reduces the training difficulty and training cycle of the target detection model, providing a method for unique individual recognition for the construction of the scene topology map. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0033] Figure 1 is a flowchart of a secondary individual recognition method driven by dynamic knowledge according to Embodiment 1 of the present invention;
[0034] Figure 2 is a schematic structural diagram of a computer device according to Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0036] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0037] Example 1
[0038] This embodiment provides a secondary individual recognition method driven by dynamic knowledge.
[0039] Regarding how to enable a robot to recognize different individuals in the construction of a dynamic environment scene and continuously learn and cognize unknown individuals, a secondary individual recognition method driven by dynamic knowledge provided in this embodiment is used to achieve a more refined and accurate environmental cognition ability of the robot.
[0040] The secondary individual recognition method driven by dynamic knowledge provided in this embodiment first uses a segmentation model to segment the scene to form candidate regions, and then clusters the remaining regions through depth data; secondly, uses object detection algorithms and feature vector similarity for preliminary entity recognition and semantic segmentation to obtain the mask of the object, and at the same time divides the regions into two categories: known semantics and unknown semantics; again, uses image feature extraction algorithms alone or in combination to extract texture features from the masked part and continuously observes and records individual features; then inputs the unknown candidate regions to a large language model, online query or database query to obtain alternative answers for the recognition results of the unknown regions, corrects the individual semantics manually, and at the same time stores the generated annotation information and the image as a data set; finally, continuously loops the above process to achieve the cognition of all individuals in the scene.
[0041] The secondary individual recognition method driven by dynamic knowledge provided in this embodiment, as Figure 1 shown, includes the following steps:
[0042] Step 1, Scene segmentation and screening domain.
[0043] Step 101, Use a segmentation model including but not limited to SAM, FastSAM, MASK-RNN, etc. to segment the input scene image, and segment the scene into several screening regions (abbreviation, screening domain).
[0044] In step 102, there are still cases of pixel omission after segmentation. Cluster the pixel regions of the region to be screened (referring to the region in the scene image except for the segmented screening regions) through depth data; if the scene image does not have depth information, cluster the remaining pixels through visual features such as color, texture, and shape to form several screening domains.
[0045] In this embodiment, the clustering can adopt but not limited to clustering algorithms such as K-means, GMM or DBSCAN.
[0046] Among them, the depth data refers to the distance of each pixel point relative to the camera, that is, the actual distance from this pixel point to the center of the camera lens. Using an RGB-D camera, the depth data of each pixel point in the scene image can be obtained.
[0047] Step 2, Secondary Feature Recognition.
[0048] Step 201, Creation of Individual Nodes
[0049] (1) For each screening domain, input the pre-trained object detection model and the custom-trained object detection model respectively to obtain the pixel-level semantic category and probability value (confidence level); since the pre-trained object detection model and the custom-trained object detection model generate their respective probability values for the semantic category recognition results, the semantic category corresponding to the high probability value is selected to obtain the semantic mask region.
[0050] Let the input image be I ∈ R H×W×3 , and K screening regions (segmentation masks) are generated by algorithms such as FastSAM where indicates that the pixel (i, j) belongs to the screening domain k, indicates not belonging, H k and W k are the height and width of the original image block region corresponding to the screening domain respectively. The segmented screening region is represented as a set of pixel coordinates Use algorithms such as YOLO to obtain the semantic category of the segmented region R k to obtain the semantic category label c k , where c k belongs to the predefined category set C = {c1, c2,..., c N}, and N is the total number of categories. Combine the segmentation mask and the semantic category for the screening domain k to obtain the semantic mask region with the semantic category c k represented as a set containing pixel coordinates and semantic categories
[0051] For a certain screening domain, if the probability values obtained by the pre-trained object detection model and the custom-trained object detection model are both less than the threshold, then this screening domain is classified as an unknown class, and the unknown class does not participate in the following steps and directly enters Step 204.
[0052] In this embodiment, the object detection model can be but is not limited to YOLO.
[0053] Among them, the pre-trained object detection model refers to the object detection model trained based on but not limited to 80 categories of the COCO dataset.
[0054] (2) Create a mapping (or dictionary), where the key is the region identifier k (screening domain number), and the value is the semantic mask region corresponding to this screening domain
[0055] (3) Extract features including but not limited to XFeat, ORB, and SIFT feature descriptors from the object images corresponding to each semantic mask region.
[0056] Among them, the feature descriptor refers to a feature descriptor, which is a representation of a picture or a picture block. The feature descriptor converts a picture with a size of width×height×3 (number of channels) into a feature vector with a length of n.
[0057] (4) Record the semantic category corresponding to a semantic mask region obtained in (1) and the feature descriptor as an individual node, and at the same time assign it a unique node number. The specific format is shown in Table 1.
[0058] Step 202: Query individual nodes.
[0059] Table 1: Node record data format
[0060]
[0061]
[0062] The semantics recognized by YOLO objects may be repeated. Taking the situation shown in Table 1 as an example, the items recognized as computers will have two individual nodes Node_id with the same semantic category. At this time, calculate the matching situation of the feature descriptors of the two semantic individuals of the same category, and calculate the number of feature hits des match Feature density under the current observation:
[0063]
[0064] Among them, des view represents the number of feature descriptors extracted from the mask region corresponding to the semantic category of the current image frame.
[0065] Specifically, traverse each semantic mask region in the current observation and perform the following steps:
[0066] (1) Determine whether the node table (for example, Table 1) is empty. If it is, go to step (2); otherwise, go to step (3);
[0067] (2) Number the nodes, and add the semantic category and feature descriptor corresponding to the semantic mask region to the node table;
[0068] (3) Determine whether the semantic category corresponding to the semantic mask region already exists in the node table. If not, number the nodes and add the semantic category and feature descriptor corresponding to the semantic mask region to the node table. For example, if "chair" has not appeared in Table 1, then directly record the node semantics and feature descriptors; otherwise, go to step (4);
[0069] (4) Match the semantic node features. If the match is successful, merge the feature subsets: Match the currently observed node (such as the node corresponding to node_id4) with the nodes of the same semantic category in the node table (such as the node corresponding to node_id2), calculate the matching situation of the feature subsets of two individuals of the same category, that is, calculate the feature density of the number of feature hits under the current observation. If the feature density meets the set conditions, add the feature subset of the currently observed node (such as the node corresponding to node_id4) to the feature subset of the node of the same semantic category in the node table (such as the node corresponding to node_id2) (the added ones are the feature subsets that did not match), and update the incremental features of the semantic node. If the feature density does not meet the set conditions, corresponding to the node number of the current observation, add the semantic category and feature subset corresponding to the current observation node to the node table (as shown in Table 1, a new node_id4 is opened to record a new computer node).
[0070] It should be noted that all the feature subsets of the semantic nodes recorded in the node table are the incremental features of the semantic nodes.
[0071] Step 203: During the process of moving with the acquisition perspective of the scene image, calculate the IoU (Intersection over Union) value for two semantic mask regions of the same object category (semantic category) in consecutive observations:
[0072]
[0073] Among them, S nj represents the semantic mask region belonging to object category j in the nth observation, and S (n+1)j represents the semantic mask region belonging to object category j in the (n + 1)th observation.
[0074] When the IoU value is greater than the threshold, lock the two semantic mask regions as the same individual object. While locking the observed object, calculate the ratio of the number of feature subsets and the total number of feature subsets matched by the object corresponding to S (n+1)j :
[0075]
[0076] Among them, des match represents the number of successfully matched feature points of the semantic mask region S (n+1)j of the latest observation of object category j and all the feature points currently contained in object category j (that is, the hit amount of the feature subsets extracted under the semantic mask region in the current observation in the incremental features), des allis the number of all feature points of the object (semantic category) observed up to the end of the last observation (i.e., the number of feature points in the incremental features of object category j). The total feature subset is obtained by incrementally updating (taking the union) the feature vector when the number of matching feature subsets in the feature vector is lower than a certain ratio (i.e., new features appear) during the continuous change of the observation angle; in other words, the total feature subset refers to the set of all feature descriptors extracted for a specific object or entity at all observation angles; when observing an object for the first time and extracting its features, these features constitute the initial "total feature subset"; as more observations are made from different angles and new extracted features are added, this set continues to grow to form a more complete feature description.
[0077] When the content of the feature subset in the object semantic mask is less than a certain ratio, new observed values (feature subsets) are added to the incremental features to update the incremental features, forming feature records for the same individual under different observation angles.
[0078] That is, as the acquisition perspective of the scene image moves, for a certain semantic category, if the intersection over union of two semantic mask regions in consecutive observations is greater than the threshold, then calculate the ratio of the hit quantity of the feature subset in the current observed semantic mask region to the number of feature subsets in the incremental features. If the ratio is less than a certain value, add the feature subset of the current observed semantic mask to the incremental features corresponding to the node of this semantic category, that is, add it to the node table.
[0079] It should be noted that in step 202, the feature density is calculated using the currently observed features and the incremental features of the same semantic node to determine the specific node corresponding to the current observation. Here, the incremental features are used to record the features of each individual at different observation angles, so as to distinguish different individuals with the same semantics and make up for the defect that the object detection algorithm cannot distinguish individuals of the same category.
[0080] By analogy with the process of 3D modeling, an object is observed from different perspectives, and the final model is not obtained from the result of a single observation. After the final model is established, it is matched through a single observation to determine which object model it is. The role of the incremental features is similar, recording the features of the object from different perspectives so that it can be distinguished whether it is the same individual from different perspectives. Thus, for objects with the same semantics, it can be distinguished whether they are the same individual in the case of multiple perspectives. In step 202, if the features are not incremental, after the process continues for a period of time, although it is the same individual, because the observations from different perspectives are inconsistent with the initial features, it will be determined as a new node.
[0081] Step 204, learning and updating of unknown individuals.
[0082] For the areas of unknown classes after the above process, they should enter the learning session. Generate images from the data in these areas, and call the APIs of general large language models including but not limited to ChatGPT or search engines including but not limited to Bing to generate candidate results (semantic classes and probability values) of possible image information. Check the candidate results through manual screening and record them in the node table shown in Table 1. At the same time, generate a dataset from the image data, mask data, and annotation data of this query, and train a target detection model with this dataset to generate a proprietary model (a custom-trained target detection model). This proprietary model will generate the result with the highest confidence through discrimination with the pre-trained model in the next use.
[0083] A two-level individual recognition method driven by dynamic knowledge provided in this embodiment solves the problem that the object detection model cannot distinguish individuals. At the same time, for individuals that were not labeled in the early stage of YOLO and those that are not recognized, it can still achieve individual semantic recognition and individual identification.
[0084] A two-level individual recognition method driven by dynamic knowledge provided in this embodiment can automatically generate a dataset and annotation information through the recognition of the target area by a large language model or a search engine, greatly simplifying the training process of the target detection model.
[0085] A two-level individual recognition method driven by dynamic knowledge provided in this embodiment reduces the training difficulty and training cycle of the target detection model through parallel target detection models, providing a method for unique individual recognition for the construction of the scene topology map.
[0086] Embodiment 2
[0087] This embodiment provides a two-level individual recognition system driven by dynamic knowledge, which specifically includes:
[0088] An image segmentation module, which is configured to: obtain a scene image and segment it into several screened areas;
[0089] A semantic node generation module, which is configured to: for each screened area, input it into a pre-trained target detection model and a custom-trained target detection model respectively to obtain a semantic class and a probability value, and select the semantic class corresponding to the high probability value to obtain a semantic node. Extract feature sub-vectors for each semantic node. For the screened area, if the probability values obtained by the pre-trained target detection model and the custom-trained target detection model are both less than the threshold, then divide the screened area into the unknown class;
[0090] An individual node matching module, which is configured to: for a semantic node in the current observation, if there already exists a semantic node of the same semantic category, match the semantic nodes based on the feature sub-pairs, and if the match is successful, merge the feature sub-pairs of the two semantic nodes and update the incremental features of the semantic nodes;
[0091] An incremental feature update module, which is configured to: as the acquisition perspective of the scene image moves, for a certain semantic category, if the intersection over union of two semantic nodes in consecutive observations is greater than a threshold, determine whether to add the feature sub-pairs of the semantic node in the current observation to the incremental features of this semantic category;
[0092] A customized training module, which is configured to: for unknown classes, generate semantic categories and probability values through a large language model or a search engine to train an object detection model and update the customized trained object detection model.
[0093] Further, the method for determining whether to add the feature sub-pairs of the semantic node in the current observation to the incremental features of this semantic category is: calculate the ratio of the hit quantity of the feature sub-pairs of the semantic node in the current observation in the incremental features to the number of feature sub-pairs in the incremental features. If the ratio is less than a certain value, add the feature sub-pairs of the semantic node in the current observation to the incremental features of this semantic category.
[0094] Further, the segmentation includes: segmenting the scene image using a segmentation model and dividing the scene into several screening regions; clustering the regions not segmented by the segmentation model into multiple screening regions.
[0095] Further, the steps of matching the semantic nodes based on the feature sub-pairs include: calculating the number of matching feature sub-pairs of two semantic nodes with the same semantic category, calculating the ratio of the number of matching feature sub-pairs to the number of feature sub-pairs of the semantic node of this semantic category in the current observation to obtain the feature density. If the feature density meets the set conditions, the two semantic nodes are successfully matched.
[0096] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1 one by one, and the specific implementation process is the same, so it will not be repeated here.
[0097] Embodiment 3
[0098] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a two-level individual recognition method driven by dynamic knowledge as described in Embodiment 1 above.
[0099] Embodiment 4
[0100] This embodiment provides a computer device, such as Figure 2As shown, it includes a display device, an input device, a computer-readable storage medium (volatile memory and non-volatile storage medium), a processor, a communication interface (i.e., a network interface), and a computer program stored on the computer-readable storage medium and executable on the processor. Among them, the processor, the communication interface, and the computer-readable storage medium can be connected through a bus or other means. Among them, the communication interface is used to receive and send data, and when the processor executes the program, it implements the steps in a secondary individual recognition method driven by dynamic knowledge as described in Embodiment 1 above.
[0101] Among them, any reference to memory, storage, database, or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more of the processes or Figure 1 blocks or multiple blocks.
[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more of the processes orFigure 1 The functions specified in one or more boxes.
[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 or more boxes.
[0105] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A secondary individual recognition method driven by dynamic knowledge, characterized in that, Including: Obtain a scene image and segment it to obtain a number of screening regions; For each screening region, input it into a pre-trained object detection model and a custom-trained object detection model respectively to obtain semantic categories and probability values, and select the semantic category corresponding to the high probability value to obtain semantic nodes. Feature sub-extraction is performed on each semantic node. For the screening region, if the probability values obtained by the pre-trained object detection model and the custom-trained object detection model are both less than the threshold, the screening region is classified as an unknown class; For a certain semantic node in the current observation, if there already exists a semantic node with the same semantic category, match the semantic nodes based on the feature sub. If the match is successful, merge the feature subs of the two semantic nodes and update the incremental feature of the semantic node; As the acquisition perspective of the scene image moves, for a certain semantic category, if the intersection over union of two semantic nodes in consecutive observations is greater than the threshold, determine whether to add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category; For the unknown class, generate semantic categories and probability values through a large language model or a search engine to train the object detection model and update the custom-trained object detection model.
2. The secondary individual recognition method driven by dynamic knowledge as described in claim 1, wherein The method for determining whether to add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category is: calculate the ratio of the hit quantity of the feature sub of the semantic node in the current observation in the incremental feature to the number of feature subs in the incremental feature. If the ratio is less than a certain value, add the feature sub of the semantic node in the current observation to the incremental feature of this semantic category.
3. A secondary individual recognition method driven by dynamic knowledge as described in claim 1, characterized in that, The segmentation includes: segment the scene image using a segmentation model to divide the scene into several screening regions; cluster the regions not segmented by the segmentation model into multiple screening regions.
4. A secondary individual recognition method driven by dynamic knowledge as described in claim 1, characterized in that, The steps of matching semantic nodes based on feature subs include: calculate the number of matching feature subs of two semantic nodes with the same semantic category, calculate the ratio of the number of matching feature subs to the number of feature subs of the semantic node of this semantic category in the current observation to obtain the feature density. If the feature density meets the set conditions, the two semantic nodes are successfully matched.
5. A secondary individual recognition system driven by dynamic knowledge, characterized in that Including: An image segmentation module configured to: obtain a scene image and segment it to obtain a number of screening regions; A semantic node generation module configured to: for each screening region, input it into a pre-trained object detection model and a custom-trained object detection model respectively to obtain semantic categories and probability values, and select the semantic category corresponding to the high probability value to obtain semantic nodes. Feature sub-extraction is performed on each semantic node. For the screening region, if the probability values obtained by the pre-trained object detection model and the custom-trained object detection model are both less than the threshold, the screening region is classified as an unknown class; An individual node matching module configured to: for a certain semantic node in the current observation, if there already exists a semantic node with the same semantic category, match the semantic nodes based on the feature sub. If the match is successful, merge the feature subs of the two semantic nodes and update the incremental feature of the semantic node; An incremental feature update module, which is configured to: as the acquisition perspective of the scene image moves, for a certain semantic category, if the intersection over union of two semantic nodes in consecutive observations is greater than a threshold, determine whether to add the feature subset of the semantic node in the current observation to the incremental features of the semantic category; A customized training module, which is configured to: for unknown classes, generate semantic categories and probability values through a large language model or a search engine to train an object detection model and update the customized trained object detection model.
6. The secondary individual recognition system driven by dynamic knowledge as claimed in claim 5, wherein The method for determining whether to add the feature subset of the semantic node in the current observation to the incremental features of the semantic category is: calculate the ratio of the hit quantity of the feature subset of the semantic node in the current observation in the incremental features to the number of feature subsets in the incremental features. If the ratio is less than a certain value, add the feature subset of the semantic node in the current observation to the incremental features of the semantic category.
7. The secondary individual recognition system driven by dynamic knowledge as claimed in claim 5, wherein The segmentation includes: using a segmentation model to segment the scene image and dividing the scene into several screening regions; clustering the regions not segmented by the segmentation model into multiple screening regions.
8. A secondary individual recognition system driven by dynamic knowledge as claimed in claim 5, wherein, The step of matching semantic nodes based on feature subsets includes: calculating the number of matching feature subsets of two semantic nodes with the same semantic category, calculating the ratio of the number of matching feature subsets to the number of feature subsets of the semantic nodes of the semantic category in the current observation to obtain a feature density. If the feature density meets the set conditions, the two semantic nodes are successfully matched.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a secondary individual recognition method driven by dynamic knowledge as described in any one of claims 1-4.
10. A computer device, comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a secondary individual recognition method driven by dynamic knowledge as described in any one of claims 1-4.