An object classification method, device, equipment and medium

By constructing node graphs and utilizing knowledge graph databases, combined with pre-saved models, the problem of common sense errors in object classification is solved, and more accurate object classification is achieved.

CN114332614BActive Publication Date: 2025-07-22CHINA ORDNANCE SCI INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111622073.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-07-22
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

The prior art is prone to common sense errors when classifying objects, resulting in inaccurate classification results.

Method used

By constructing node graphs and utilizing knowledge graph databases, combining the pre-saved first model and second model, the object categories in each detection box are determined, and common sense errors are reduced and classification accuracy is improved.

Benefits of technology

It significantly improves the accuracy of object classification and reduces common sense errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332614B_ABST
    Figure CN114332614B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an object classification method, apparatus, device, and medium. The electronic device determines a feature vector and a classification vector corresponding to each detection box in the image to be classified according to the first model, constructs a node corresponding to each feature vector, and determines a node graph. According to the target preset category corresponding to each node in the node graph and the classification included in each knowledge graph in the saved knowledge graph database, it is determined that the edge connecting every two nodes in the node graph is the target number of times that the target preset categories corresponding to the two nodes appear in the same knowledge graph. The second model determines the target category of each object according to each node graph and the image to be classified, which can greatly reduce common sense errors and improve the accuracy of object classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an object classification method, apparatus, device, and medium. Background Art

[0002] With the development of technology, the research and development of software products have put forward higher requirements for user experience. In the field of object classification, rough classification of objects, that is, classifying the major categories to which the objects belong, can no longer support today's intelligent application scenarios. For example, in the case of pet classification, the ability to classify pets into major categories such as cats and dogs can no longer meet the requirements of the application scenario for user experience. What is needed in the interaction scenario is a more detailed classification result, such as outputting specific categories such as corgi dogs and Persian cats.

[0003] In order to meet the need for specific classification of objects to be classified, in the prior art, a trained neural network is usually used to determine the category of the object to be classified based on the feature vector of the object to be classified in the image. However, this classification method may have common sense errors for similar objects, such as classifying a basketball as a football and an egg as a table tennis ball, that is, the final classification result is inaccurate. Summary of the Invention

[0004] This application provides an object classification method, apparatus, device, and medium to solve the problem of inaccurate classification results when classifying objects in an image in the prior art.

[0005] An embodiment of this application provides an object classification method, and the method includes:

[0006] Input a to-be-classified image pre-annotated with each detection box into a pre-saved first model, and obtain a classification vector and a feature vector corresponding to each detection box output by the first model, where each value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension;

[0007] For each detection box, obtain a set number of target preset categories in the classification vector corresponding to this detection box; construct a node corresponding to the feature vector of this detection box, and save the target preset category corresponding to this node;

[0008] Combine any target preset category corresponding to each node to determine a node graph corresponding to each category combination. Any two nodes in this node graph are connected to each other. Among them, for any two connected nodes in this node graph, according to the target preset categories corresponding to these two nodes in this category combination and the number of times each two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to this category combination, and save the target number for the edge connecting these two nodes.

[0009] According to each of the node graphs, the image to be classified, and the second model saved in advance, determine the target category corresponding to the object included in each detection box.

[0010] Further, the inputting the image to be classified with each detection box pre-annotated into the first model to obtain the classification vector and feature vector corresponding to each detection box output by the first model includes:

[0011] The first model determines the feature vector corresponding to each detection box; according to the feature vector and each preset category configured in advance, determine the probability that the object included in this detection box is classified into each preset category, and according to the corresponding relationship between each preset category and the vector dimension specified in advance, output the classification vector and feature vector corresponding to this detection box.

[0012] Further, the obtaining the set number of target preset categories in the classification vector corresponding to this detection box includes:

[0013] Sort the values corresponding to each dimension in the classification vector in descending order, select the preset number of target values ranked at the front, and determine the preset classification corresponding to each target value as the target preset category.

[0014] Further, the determining the target category corresponding to the object included in each detection box according to each of the node graphs, the image to be classified, and the second model saved in advance includes:

[0015] The second model determines the spatial distance between any two detection boxes in the to-be-classified image according to the position information of each detection box in the input to-be-classified image, and determines the distance matrix corresponding to the spatial distance; saves the target number for the edge connecting any two nodes in each node graph, and determines the relationship matrix corresponding to the target number; determines the first feature matrix corresponding to the node graph according to the feature vector corresponding to each node in the node graph; for the feature vector corresponding to each node in the node graph, updates the feature vector according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each pre-configured category, determines the probability that the object contained in the detection box corresponding to the updated feature vector is classified into each preset category, and determines the category corresponding to the maximum probability as the category corresponding to the updated feature vector; for any two connected nodes in the node graph, determines the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, and determines the loss value corresponding to the node graph according to the relationship probability corresponding to every two connected nodes in the node graph and a preset first function; determines the target category corresponding to the object contained in each detection box as the category corresponding to each node in the node graph with the minimum loss value.

[0016] Further, the updating of the feature vector corresponding to each node in the node graph according to the relationship matrix, the distance matrix, and the feature matrix includes:

[0017] For each node in the node graph, perform a preset number of updates, where each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, and the relationship matrix and the distance matrix between this node and each other node in the node graph, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in this node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached.

[0018] Further, the determining of the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database includes:

[0019] For any two connected nodes in the node graph, obtain the number of times that the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times that the two nodes appear simultaneously in the knowledge graph database.

[0020] An embodiment of the present application further provides an object classification device, the device includes:

[0021] A classification module, configured to input a to-be-classified image pre-annotated with each detection box into a pre-stored first model, and obtain a classification vector and a feature vector corresponding to each detection box output by the first model, where each value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension;

[0022] A processing module, configured to, for each detection box, obtain a set number of target preset categories in the classification vector corresponding to this detection box; construct a node corresponding to the feature vector of this detection box, and save the target preset category corresponding to this node; combine any one of the target preset categories corresponding to each node to determine a node graph corresponding to each category combination, where any two nodes in this node graph are connected to each other, and where, for any two connected nodes in this node graph, according to the target preset categories corresponding to these two nodes in this category combination, and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to this category combination, and save the target number for the edge connected by these two nodes;

[0023] The classification module is further configured to determine the target category corresponding to the object included in each detection box according to each of the node graphs, the to-be-classified image, and a pre-stored second model.

[0024] Further, the classification module is specifically configured to, for each detection box, the first model determines the feature vector corresponding to this detection box; according to the feature vector and each preset category pre-configured, determine the probability that the object included in this detection box is classified into each preset category, and according to the corresponding relationship between each preset category and the vector dimension pre-specified, output the classification vector and the feature vector corresponding to this detection box.

[0025] Further, the processing module is specifically configured to sort the values corresponding to each dimension in the classification vector in descending order, select a preset number of target values ranked at the front, and determine the preset classification corresponding to each target value as the target preset category.

[0026] Further, the classification module is specifically configured to: the second model determines the spatial distance between any two detection frames in the to-be-classified image according to the position information of each detection frame in the input to-be-classified image, and determines a distance matrix corresponding to the spatial distance; for each target graph, save the target number of times for the edge connecting any two nodes, and determine a relationship matrix corresponding to the target number of times; determine a first feature matrix corresponding to the target graph according to the feature vector corresponding to each node in the target graph; for the feature vector corresponding to each node in the target graph, update the feature vector according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each pre-configured category, determine the probability that the object included in the detection frame corresponding to the updated feature vector is classified into each preset category, and determine the category corresponding to the maximum probability as the category corresponding to the updated feature vector; for any two connected nodes in the target graph, determine the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, and determine the loss value corresponding to the target graph according to the relationship probability corresponding to every two connected nodes in the target graph and a preset first function; determine the category corresponding to each node in the target graph with the minimum loss value as the target category corresponding to the object included in each detection frame.

[0027] Further, the classification module is specifically configured to, for each node in the target graph, perform an update for a preset number of times, where each update process includes: determining the sum of the product of the feature matrix, a preset parameter, and the relationship matrix and the distance matrix between the node and each other node in the target graph, and determining the updated feature vector of the node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in the node, and performing the update of the next feature vector according to the updated feature matrix until the preset number of times is reached.

[0028] Further, the classification module is specifically configured to, for any two connected nodes in the target graph, obtain the number of times that the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times that the two nodes appear simultaneously in the knowledge graph database.

[0029] An embodiment of the present application further provides an electronic device, which at least includes a processor and a memory. When the processor executes a computer program stored in the memory, it implements the steps of any one of the above object classification methods.

[0030] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any one of the above object classification methods are implemented.

[0031] In the embodiment of the present application, an image to be classified pre-labeled with each detection box is input into a pre-stored first model, and a classification vector and a feature vector corresponding to each detection box output by the first model are obtained. Each dimension value in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension. For each detection box, a set number of target preset categories in the classification vector corresponding to the detection box are obtained; nodes corresponding to the feature vector of the detection box are constructed, and the target preset category is stored corresponding to the node. Any one of the target preset categories corresponding to each node is combined to determine a node graph corresponding to each category combination. Any two nodes in the node graph are connected to each other. Among them, for any two connected nodes in the node graph, according to the target preset categories corresponding to the two nodes in the category combination and the number of times that every two preset categories appear simultaneously in the stored knowledge graph database, the target number corresponding to the target preset category corresponding to the category combination is determined, and the target number is stored for the edge connecting the two nodes. According to each node graph, the image to be classified, and a pre-stored second model, the target category of the object included in each detection box is determined. That is, in the embodiment of the present application, according to the pre-stored knowledge graph database, as well as the pre-configured first model and second model, the target category of the object corresponding to each detection box in the image to be classified is determined. Since the knowledge graph database is constructed according to the actual scenario, when determining the target category of each object based on the knowledge graph database, common sense errors can be greatly reduced, and the accuracy of object classification is improved. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions of the present application, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic diagram of an object classification process provided by an embodiment of the present application;

[0034] Figure 2 It is a schematic diagram of object classification provided by an embodiment of the present application;

[0035] Figure 3 It is a schematic structural diagram of an object classification device provided by an embodiment of the present application;

[0036] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0037] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0038] In order to improve the accuracy of object classification, an embodiment of the present application provides an object classification method, device, equipment, and medium.

[0039] Embodiment 1:

[0040] Figure 1 A schematic diagram of an object classification process provided by an embodiment of the present application, the process including:

[0041] S101: Input a to-be-classified image pre-annotated with each detection box into a pre-stored first model, and obtain a classification vector and a feature vector corresponding to each detection box output by the first model, where each value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to the dimension.

[0042] An object classification method provided by an embodiment of the present application is applied to an electronic device, and the electronic device may be a PC or the like.

[0043] In an embodiment of the present application, when an electronic device performs object classification on a to-be-classified image, at least one detection box is annotated in the to-be-classified image, and each detection box includes an object. The electronic device classifies the object included in each detection box.

[0044] Wherein, in an embodiment of the present application, the electronic device may identify an object in the to-be-classified image based on a pre-configured target detection algorithm, determine a detection box corresponding to each object, and specifically, according to the target detection algorithm, for each object in the to-be-classified image, determine the area where the object is located, and identify the area with a detection box.

[0045] The electronic device includes a first model that has been pre-trained to classify the objects contained in each detection box in the image to be classified. When the electronic device obtains the image to be classified with each detection box labeled, it inputs the image to be classified into the first model, and the first model outputs a classification vector and a feature vector corresponding to each detection box. Among them, both the classification vector and the feature vector contain multiple dimensions, and each dimension corresponds to a numerical value. Specifically, the component vector corresponding to the detection box is used to represent the probability that the object contained in the detection box is classified into the preset category corresponding to each dimension, and the feature vector corresponding to the detection box is used to represent the vector that can reflect the characteristics of the object contained in the detection box in the image to be classified.

[0046] It should be noted that, in order to improve the accuracy of object classification, in the embodiments of the present application, the first model can recognize objects of multiple preset categories pre-stored therein, that is, recognize objects of preset categories. Therefore, the first model can determine the probability that the object contained in each detection box belongs to each preset category, and based on the corresponding relationship between each preset category and each dimension in the component vector pre-stored, and the probability that the object belongs to each preset category, output the classification vector corresponding to the detection box. For each dimension in the classification vector of the detection box, the numerical value corresponding to this dimension is the probability that the object contained in the detection box is classified into the preset category corresponding to this dimension. That is to say, the component vector is used to represent the probability that the object in its corresponding detection box belongs to each preset category.

[0047] Among them, the training process of the first model includes:

[0048] Input the sample image to be classified with detection boxes labeled into the first model, and receive the sample classification vector and the sample feature vector corresponding to each detection box in the image to be classified output by the first model, where the image to be classified is also marked with the original classification vector and the original feature vector corresponding to each detection box;

[0049] Adjust the parameters of the first model according to the loss values corresponding to the sample classification vector and the original classification vector, and the loss values corresponding to the sample feature vector and the original feature vector;

[0050] If the sample quantity for which both of these two loss values are less than the threshold meets the requirement or the number of iterations of the first model reaches the maximum value, it is considered that the training of the first model is completed.

[0051] S102: For each detection box, obtain a set number of target preset categories in the classification vector corresponding to the detection box; construct nodes corresponding to the feature vector of the detection box, and save the target preset categories corresponding to the nodes.

[0052] In the embodiment of the present application, after the first model outputs the classification vector corresponding to each detection box, the electronic device can further determine the target category of the object included in each detection box according to the classification vector corresponding to each detection box. However, since there are many types of preset categories included in the classification vector, and the probabilities of some of the preset categories in the classification vector are very small, it can be directly determined that these preset categories with very small probabilities are not the target categories of the object corresponding to the classification vector. Based on this, in order to reduce the load pressure on the electronic device, in the embodiment of the present application, for each classification vector, the electronic device obtains a set number of target preset categories in the classification vector, where the set number of target preset categories can be the set number of preset categories with relatively large probabilities.

[0053] In order to further determine the target category of the object included in each detection box, in the embodiment of the present application, the detection is based on a knowledge graph. Specifically, for each detection box, the electronic device can construct a node corresponding to the feature vector of the detection box and save the target preset category determined for the detection box corresponding to the node.

[0054] S103: Combine any one target preset category corresponding to each node to determine a node graph corresponding to each category combination. Any two nodes in the node graph are connected to each other. Among them, for any two connected nodes in the node graph, according to the target preset categories corresponding to the two nodes in the category combination and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to the category combination, and save the target number for the edge connecting the two nodes.

[0055] In order to further determine the target category of the object included in each detection box, in the embodiment of the present application, after constructing a node corresponding to the feature vector for each detection box, because the image to be classified contains multiple detection boxes, the feature vector of each detection box corresponds to a corresponding node, and each target preset category determined for the corresponding detection box is saved corresponding to each node. Therefore, multiple node graphs of the image to be detected can be constructed. Each node graph includes a node corresponding to the feature vector of each detection box and one target preset category corresponding to each node. Among every two node graphs, at least one node with different target preset categories is included.

[0056] Specifically, in the embodiment of the present application, the electronic device combines any one target preset category corresponding to each node to determine a node graph corresponding to each category combination. Any two nodes in the node graph are connected to each other.

[0057] Among them, in the embodiment of the present application, a knowledge graph database is pre-stored in the electronic device. The knowledge graph database contains multiple knowledge graphs, and each knowledge graph contains at least two classifications. All the classifications included in each knowledge graph are the classifications corresponding to the objects included in a sample image. Then, according to each knowledge graph in the knowledge graph database, determine the target number of times that the two target preset categories co-occur in one knowledge graph, and save the target number for the edge connected by the two nodes.

[0058] S104: According to each of the node graphs, the image to be classified, and the pre-stored second model, determine the target category corresponding to the object included in each detection frame.

[0059] In the embodiment of the present application, a pre-trained second model is stored in the electronic device. After constructing multiple node graphs, according to each node graph and the image to be classified, determine the target category corresponding to the object included in each detection frame.

[0060] Specifically, input each node graph and the image to be classified into the second model. The second model determines the node graph with the most accurate classification in the node graph according to the information carried in the image to be classified, and outputs the node graph with the most accurate classification. The electronic device determines the target preset classification corresponding to each node in the node graph output by the second model as the target classification of the object corresponding to each node.

[0061] In the embodiment of the present application, the electronic device determines the feature vector and classification vector corresponding to each detection frame in the image to be classified according to the first model, constructs a node corresponding to each feature vector, and determines a node graph. According to the target preset category corresponding to each node in the node graph and the classifications included in each knowledge graph in the saved knowledge graph database, it is determined that the edge connecting every two nodes in the node graph is the target number of times that the target preset categories corresponding to the two nodes appear in the same knowledge graph. The second model determines the target category of each object according to each node graph and the image to be classified, which can greatly reduce common sense errors and improve the accuracy of object classification.

[0062] Embodiment 2:

[0063] In order to perform an initial classification on the objects in the image to be classified, on the basis of the above embodiment, in the embodiment of the present application, the step of inputting the image to be classified with each detection frame pre-annotated into the first model and obtaining the classification vector and feature vector corresponding to each detection frame output by the first model includes:

[0064] For each detection box, the first model determines the feature vector corresponding to the detection box; according to the feature vector and each pre-configured category, it determines the probability that the object contained in the detection box is classified into each preset category, and according to the correspondence between each preset category and the vector dimension stipulated in advance, outputs the classification vector and the feature vector corresponding to the detection box.

[0065] In the embodiment of the present application, when the first model stored in the electronic device processes the image to be classified, for each detection box, it determines the feature vector corresponding to the detection box, and then according to the feature vector, determines the classification vector corresponding to the detection box, and finally outputs the classification vector and the feature vector corresponding to each detection box.

[0066] Specifically, in the embodiment of the present application, for each detection box, the first model processes the image content corresponding to the detection box, obtains a vector that can reflect the characteristics of the object contained in the corresponding detection box of the image content, and determines this vector as the feature vector corresponding to the detection box.

[0067] A plurality of preset categories are pre-configured in the first model. For each detection box, the first model determines the probability that the object contained in the detection box is each preset category according to the feature vector corresponding to the detection box, and determines the classification vector corresponding to the detection box according to the correspondence between each preset category and the vector dimension stipulated in advance.

[0068] Among them, the first model can be a ResNet-50 convolutional neural network model.

[0069] Embodiment 3:

[0070] In order to reduce the load pressure of the electronic device and improve the efficiency of object classification, on the basis of the above embodiments, in the embodiment of the present application, the set number of target preset categories in obtaining the classification vector corresponding to the detection box includes:

[0071] Sort the values corresponding to each dimension of the classification vector in descending order, select the preset number of target values in the front of the sorting, and determine the preset classification corresponding to each target value as the target preset category.

[0072] In the embodiment of the present application, after determining the classification vector corresponding to each detection box, since the value of each dimension of the component vector is the probability that the object contained in the detection box is the preset category corresponding to this dimension, if this probability is too small, this preset category can be excluded.

[0073] In order to reduce the load pressure of the electronic device and improve the efficiency of object classification, in the embodiments of the present application, for the classification vector corresponding to each detection box, the electronic device will select the target preset category with a relatively high probability from the preset categories corresponding to the classification vector, and then further classify the object according to the target preset category.

[0074] Specifically, in the embodiments of the present application, for each classification vector, the electronic device arranges the values corresponding to each dimension in the classification vector in descending order, selects the preset number of target values ranked at the front, and determines the preset classification corresponding to the target value as the target preset classification.

[0075] Embodiment 4:

[0076] In order to further classify the objects contained in the detection box, based on the above embodiments, in the embodiments of the present application, determining the target category corresponding to the object contained in each detection box according to each node graph, the image to be classified, and the pre-stored second model includes:

[0077] The second model determines the spatial distance between any two detection boxes in the image to be classified according to the position information of each detection box in the input image to be classified, and determines the distance matrix corresponding to the spatial distance; for each edge connecting any two nodes in each node graph, the target number of times is saved, and the relationship matrix corresponding to the target number of times is determined; according to the feature vector corresponding to each node in the node graph, the first feature matrix corresponding to the node graph is determined; for the feature vector corresponding to each node in the node graph, the feature vector is updated according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each preset category, the probability that the object contained in the detection box corresponding to the updated feature vector is classified into each preset category is determined, and the category corresponding to the maximum probability is determined as the category corresponding to the updated feature vector; for any two connected nodes in the node graph, according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, the relationship probability between the categories corresponding to the two nodes is determined, and according to the relationship probability between each two connected nodes in the node graph and the preset first function, the loss value corresponding to the node graph is determined; the category corresponding to each node in the node graph with the minimum loss value is determined as the target category corresponding to the object contained in each detection box.

[0078] In the embodiments of the present application, the second model saved in the electronic device will further determine the category of the object contained in each detection box, and the electronic device determines the category of the object contained in each detection box determined by the second model as the target category corresponding to the object.

[0079] Among them, the second model is the target category corresponding to the object included in each detection box determined based on the image to be classified, each constructed node graph, and the saved knowledge graph data.

[0080] Specifically, the second model determines the spatial distance between any two detection boxes in the image to be classified according to the position information of each detection box in the image to be classified. Since the detection box is a region, the spatial distance between the two detection boxes can be determined according to the center points of the two detection boxes or the pixel points at other preset positions. For example, in the image to be classified, a plane coordinate system is constructed, and the position information of each detection box is represented by the coordinates corresponding to the center point or the upper left vertex of each detection box. When calculating the spatial distance between any two detection boxes, the following formula can be used:

[0081]

[0082] Among them, ρ is used to represent the spatial distance between any two detection boxes, and (x1, y1) and (x2, y2) respectively represent the coordinates corresponding to the center points of the two detection boxes in the image to be classified.

[0083] After the second model determines the spatial distance between any two detection boxes, it determines the product of the spatial distance and the preset identity matrix, and determines this product as the distance matrix corresponding to the spatial distance. Similarly, for each node graph, the second model determines the product of the target number corresponding to the edge connecting any two nodes in the node graph and the preset identity matrix, and determines this product as the relationship matrix corresponding to the target number. And according to the eigenvector corresponding to each node in the node graph, according to the pre-configured rules, each eigenvector is used as a row or a column to determine the first eigenmatrix corresponding to the node graph.

[0084] For each node in each node graph, the second model updates the eigenvector corresponding to each node in the node graph according to the first eigenmatrix, and the distance matrix and relationship matrix between this node and other nodes. Specifically, the second model determines the distance matrix and relationship matrix between this node and any other node, calculates the product of the distance matrix, relationship matrix and the first eigenmatrix, and determines the sum value of the products corresponding to this node and each other node as the updated eigenvector of this node. And according to the updated eigenvector and each preset category pre-configured, it determines the probability that the object included in the detection box corresponding to each updated eigenvector is classified into each preset category, and determines the category corresponding to the maximum probability as the category corresponding to the updated eigenvector. That is, the second model reclassifies this node according to the updated eigenvector of each node, and determines the unique category corresponding to this node in the node graph.

[0085] After reclassifying each node in each node graph, for each node graph, the second model determines the relationship probability corresponding to any two connected nodes in the node graph according to the categories corresponding to the two nodes after reclassification and the number of times each two preset categories appear simultaneously in the saved knowledge graph database. Among them, the more times the categories corresponding to the two nodes after reclassification appear in the knowledge graph database, the greater the relationship probability corresponding to the two nodes.

[0086] For each node graph, the second model determines the loss value corresponding to the node graph according to the relationship probability corresponding to each two connected nodes in the node graph and a preset first function; the category corresponding to each node in the node graph with the smallest loss value is determined as the target category corresponding to the object included in each detection box.

[0087] Specifically, in the embodiment of the present application, for each node in each node graph, the second model determines the sum value of the relationship probabilities between the node and each other node, and determines the total sum value of the sum values corresponding to each node. The negative number corresponding to the ratio of the total sum value to the number of preset categories included in the knowledge graph database is determined as the loss value corresponding to the node graph. Among them, in the embodiment of the present application, in order to reduce errors, when determining the loss value corresponding to each node graph, the cross-entropy loss value corresponding to the node graph is determined.

[0088] Among them, when calculating the cross-entropy loss value corresponding to each node graph, the following formula can be used for calculation:

[0089]

[0090] Among them, L represents the cross-entropy loss value, N represents the number of preset categories included in the knowledge graph database, τ(y i =s) is the activation function, that is, the preset function. When the equation in the parentheses holds, the value of the activation function is 1, otherwise it is 0. Among them, in y i =j, y i is the serial number corresponding to the category corresponding to node i in the knowledge graph database, s is an arbitrary number, and the value of s is the serial number corresponding to each preset category in the knowledge graph database, p i j is the relationship probability corresponding to the two nodes. Since the relationship probability is calculated by exponentiation to avoid a relatively small calculated relationship probability value, when using it, the logarithm of the relationship probability needs to be calculated.

[0091] Among them, in the embodiment of the present application, the second model can be a graph convolutional neural network model.

[0092] Embodiment 5:

[0093] In order to update the feature vectors corresponding to each detection box, based on the above embodiments, in the embodiments of the present application, for the feature vectors corresponding to each node in the node graph, updating the feature vectors according to the relationship matrix, the distance matrix, and the feature matrix includes:

[0094] For each node in the node graph, perform a preset number of updates, where each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, the relationship matrix between this node and each other node in the node graph, and the distance matrix, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vectors of each node in this node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached.

[0095] In the embodiments of the present application, when the first model performs an initial classification on the objects included in each detection box in the image to be classified, for each detection box, only the feature vector corresponding to this detection box determined based on the image content corresponding to this detection box is used for the initial classification, and whether the classification conforms to common sense and the relationship between the objects included in the detection box are not considered.

[0096] Therefore, for each node of each node graph, the second model updates the feature vector corresponding to this node according to the feature matrix corresponding to this node graph, the distance matrix between this node and each other node in the node graph, and the relationship matrix determined based on the knowledge graph database, so that the classification of this node is more in line with common sense and the accuracy of the classification is improved.

[0097] Specifically, in the embodiments of the present application, for each node in the node graph, perform a preset number of updates. Each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, the relationship matrix and the distance matrix between this node and each other node in the node graph, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vectors of each node in this node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached. Among them, the preset parameters used in each update process can be the same or different.

[0098] And for each node in each node graph, when updating the feature information corresponding to this node, in order to avoid the influence of nodes with a weak relationship with this node on the update, nodes with a strong relationship with this node can also be selected to update the feature vector of this node.

[0099] Specifically, for each node in each node graph, determine the first preset number of nodes with a large target frequency corresponding to the edges connected to the node in the node graph, and then select the second preset number of nodes that are closer to the node from the first preset number of nodes. Calculate the spatial distance between each node in the second preset number of nodes and the node, as well as the target frequency corresponding to the two nodes, perform normalization processing on each spatial distance and each target frequency respectively to obtain a distance matrix and a relationship matrix, determine the product of the distance matrix, the relationship matrix, the feature matrix, and the preset parameters, determine the sum value of the product corresponding to each node in the second preset number of nodes, use the sum value as the input of the activation function, and determine the output of the activation function as the updated feature vector of the node.

[0100] Among them, the process of updating the feature vector corresponding to the nodes in each node graph can be implemented by the following formula:

[0101]

[0102] Among them, represents the updated feature vector of the node, represents the distance matrix obtained after normalizing the spatial distance between two nodes, where A′ represents the relationship matrix obtained after normalizing the target frequency of two nodes, and I N represents the identity matrix, X (l) represents the feature matrix, W (l) represents the preset parameter, f(*) represents the activation function, r i represents the first preset number of nodes with a large target frequency corresponding to the edges connected to the node in the node graph, and d ri represents the second preset number of nodes that are closer to the node among the first preset number of nodes.

[0103] Embodiment 6:

[0104] In order to determine the relationship probability between any two nodes in each node graph, based on the above embodiments, in the embodiment of the present application, the determining the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the stored knowledge graph database includes:

[0105] For any two connected nodes in the node graph, obtain the number of times that the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times that the two nodes appear simultaneously in the knowledge graph database.

[0106] In the embodiment of the present application, for any two connected nodes in each node graph, when determining the relationship probability corresponding to the two nodes, the number of times any one of the two nodes co-occurs with each other category in the knowledge graph database is obtained, and the target number of times corresponding to the edge between the two nodes in the node graph is obtained. The ratio of the sum of the target number of times and the number of times any one node co-occurs with each other category is calculated, and this ratio is determined as the relationship probability between the two nodes.

[0107] In addition, in the embodiment of the present application, for any two connected nodes in each node graph, when determining the relationship probability corresponding to the two nodes, if the ratio of the sum of the target number of times and the number of times any one node co-occurs with each other category is directly calculated, the value of the finally determined ratio may be small, that is, the relationship probability corresponding to the two nodes is small. Therefore, in the embodiment of the present application, the relationship probability corresponding to the two nodes can be calculated in the form of an exponential function. Specifically, it can be calculated by the following formula:

[0108]

[0109] Wherein, represents the relationship probability between node i and node j, exp represents the indicator function, represents the target number of times corresponding to the edge between node i and node j in the node graph, m represents other nodes in the knowledge graph database, represents the number of times node j co-occurs with each other category.

[0110] Figure 2 is a schematic diagram of object classification provided by the embodiment of the present application. As shown in this Figure 2 figure, the process includes inputting the to-be-classified image marked with a detection frame into the first model pre-stored in the electronic device. The first model outputs a classification vector and a feature vector corresponding to each detection frame; for each detection frame, the electronic device obtains a set number of target preset categories in the classification vector corresponding to the detection frame, constructs a node corresponding to the feature vector of the detection frame, and saves the target preset category corresponding to the node. Any one of the target preset categories corresponding to each node is combined to determine a node graph corresponding to each category combination. Any two nodes in the node graph are connected to each other; the second model stored in the electronic device determines the target category corresponding to the object included in each detection frame according to each node graph, the to-be-classified image, and the stored knowledge graph database.

[0111] Embodiment 7:

[0112] Figure 3 is a schematic structural diagram of an object classification device provided by the embodiment of the present application. The device includes:

[0113] The classification module 301 is configured to input a classification image with each detection box pre-annotated into a pre-stored first model, and obtain a classification vector and a feature vector corresponding to each detection box output by the first model, where the value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension;

[0114] The processing module 302 is configured to, for each detection box, obtain a set number of target preset categories in the classification vector corresponding to this detection box; construct a node corresponding to the feature vector of this detection box, and save the target preset category corresponding to this node; combine any one of the target preset categories corresponding to each node to determine a node graph corresponding to each category combination, where any two nodes in this node graph are connected to each other. Among them, for any two connected nodes in this node graph, according to the target preset categories corresponding to these two nodes in this category combination, and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to this category combination, and save the target number for the edge connecting these two nodes;

[0115] The classification module 301 is further configured to determine the target category corresponding to the object included in each detection box according to each node graph, the classification image, and a pre-stored second model.

[0116] In a possible implementation manner, the classification module 301 is specifically configured to, for each detection box, the first model determines the feature vector corresponding to this detection box; according to the feature vector and each preset category configured in advance, determine the probability that the object included in this detection box is classified into each preset category, and according to the corresponding relationship between each preset category and the vector dimension stipulated in advance, output the classification vector and the feature vector corresponding to this detection box.

[0117] In a possible implementation manner, the processing module 302 is specifically configured to sort the values corresponding to each dimension in the classification vector in descending order, select the preset number of target values ranked at the front, and determine the preset classification corresponding to each target value as the target preset category.

[0118] In a possible implementation manner, the classification module 301 is specifically configured to: based on the position information of each detection box in the to-be-classified image input by the second model, determine the spatial distance between any two detection boxes in the to-be-classified image, and determine a distance matrix corresponding to the spatial distance; save the target number of times for the edge connecting any two nodes in each node graph, and determine a relationship matrix corresponding to the target number of times; determine a first feature matrix corresponding to the node graph according to the feature vector corresponding to each node in the node graph; for the feature vector corresponding to each node in the node graph, update the feature vector according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each pre-configured category, determine the probability that the object included in the detection box corresponding to the updated feature vector is classified into each preset category, and determine the category corresponding to the maximum probability as the category corresponding to the updated feature vector; for any two connected nodes in the node graph, determine the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, and determine the loss value corresponding to the node graph according to the relationship probability corresponding to every two connected nodes in the node graph and a preset first function; determine the target category corresponding to the object included in each detection box as the category corresponding to each node in the node graph with the minimum loss value.

[0119] In a possible implementation manner, the classification module 301 is specifically configured to perform a preset number of updates on each node in the node graph, where each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, and the relationship matrix and the distance matrix between this node and each other node in the node graph, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in this node, and performing the update of the next feature vector according to the updated feature matrix until the preset number of times is reached.

[0120] In a possible implementation manner, the classification module 301 is specifically configured to, for any two connected nodes in the node graph, obtain the number of times that the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times that the two nodes appear simultaneously in the knowledge graph database.

[0121] Embodiment 8:

[0122] Based on the above embodiments, an embodiment of the present invention further provides an electronic device. Figure 4 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, asFigure 4 As shown in the figure, it includes: a processor 41, a communication interface 42, a memory 43, and a communication bus 44. Among them, the processor 41, the communication interface 42, and the memory 43 complete mutual communication through the communication bus 44;

[0123] A computer program is stored in the memory 43. When the program is executed by the processor 41, the processor 41 is caused to execute the following steps:

[0124] Input the to-be-classified image pre-annotated with each detection box into the pre-saved first model, and obtain the classification vector and feature vector corresponding to each detection box output by the first model, where the value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension;

[0125] For each detection box, obtain a set number of target preset categories in the classification vector corresponding to this detection box; construct a node corresponding to the feature vector of this detection box, and save the target preset category corresponding to this node;

[0126] Combine any target preset category corresponding to each node, determine the node graph corresponding to each category combination, and any two nodes in this node graph are connected to each other. Among them, for any two connected nodes in this node graph, according to the target preset categories corresponding to these two nodes in this category combination, and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to this category combination, and save the target number for the edge connecting these two nodes;

[0127] According to each of the node graphs, the to-be-classified image, and the pre-saved second model, determine the target category corresponding to the object included in each detection box.

[0128] In a possible implementation manner, the step of inputting the to-be-classified image pre-annotated with each detection box into the first model and obtaining the classification vector and feature vector corresponding to each detection box output by the first model includes:

[0129] The first model determines the feature vector corresponding to each detection box; according to the feature vector and each preset category pre-configured, determine the probability that the object included in this detection box is classified into each preset category, and according to the pre-specified corresponding relationship between each preset category and the vector dimension, output the classification vector and feature vector corresponding to this detection box.

[0130] In a possible implementation manner, the step of obtaining a set number of target preset categories in the classification vector corresponding to this detection box includes:

[0131] Sort the values corresponding to each dimension in the classification vector in descending order, select a preset number of target values at the front of the sorting, and determine the preset classification corresponding to each target value as the target preset category.

[0132] In a possible implementation manner, the determining, according to each of the node graphs, the image to be classified, and a pre-stored second model, the target category corresponding to the object included in each detection frame includes:

[0133] The second model determines the spatial distance between any two detection frames in the image to be classified according to the position information of each detection frame in the input image to be classified, and determines a distance matrix corresponding to the spatial distance; saves the target number of times for the edge connecting any two nodes in each node graph, and determines a relationship matrix corresponding to the target number of times; determines a first feature matrix corresponding to the node graph according to the feature vector corresponding to each node in the node graph; for the feature vector corresponding to each node in the node graph, updates the feature vector according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each preset category, determines the probability that the object included in the detection frame corresponding to the updated feature vector is classified into each preset category, and determines the category corresponding to the maximum probability as the category corresponding to the updated feature vector; for any two connected nodes in the node graph, determines the relationship probability between the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the pre-stored knowledge graph database, and determines the loss value corresponding to the node graph according to the relationship probability between every two connected nodes in the node graph and a preset first function; determines the category corresponding to each node in the node graph with the minimum loss value as the target category corresponding to the object included in each detection frame.

[0134] In a possible implementation manner, the updating the feature vector according to the relationship matrix, the distance matrix, and the feature matrix for the feature vector corresponding to each node in the node graph includes:

[0135] For each node in the node graph, perform a preset number of updates, where each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, and the relationship matrix and the distance matrix between this node and each other node in the node graph, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in this node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached.

[0136] In a possible implementation manner, determining the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the stored knowledge graph database includes:

[0137] For any two connected nodes in the node graph, obtain the number of times that the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times that the two nodes appear simultaneously in the knowledge graph database.

[0138] Since the principle of the above electronic device for solving problems is similar to that of the object classification method, the implementation of the above electronic device can refer to the embodiments of the method, and the repeated parts will not be described again.

[0139] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 42 is used for communication between the above electronic device and other devices. The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk storage. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0140] The above processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0141] Embodiment 9:

[0142] Based on the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to execute the following steps when executing:

[0143] Input the classification image pre - labeled with each detection box into the pre - saved first model, and obtain the classification vector and feature vector corresponding to each detection box output by the first model, where the value corresponding to each dimension in the classification vector is the probability that the object contained in the corresponding detection box is classified into the preset category corresponding to this dimension;

[0144] For each detection box, obtain a set number of target preset categories in the classification vector corresponding to this detection box; construct a node corresponding to the feature vector of this detection box, and save the target preset category corresponding to this node;

[0145] Combine any target preset category corresponding to each node, and determine the node graph corresponding to each category combination. Any two nodes in this node graph are connected to each other. Among them, for any two connected nodes in this node graph, according to the target preset categories corresponding to these two nodes in this category combination, and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to this category combination, and save the target number for the edge connecting these two nodes;

[0146] According to each of the node graphs, the classification image, and the pre - saved second model, determine the target category corresponding to the object contained in each detection box.

[0147] In a possible implementation manner, the step of inputting the classification image pre - labeled with each detection box into the first model and obtaining the classification vector and feature vector corresponding to each detection box output by the first model includes:

[0148] For each detection box, the first model determines the feature vector corresponding to this detection box; according to the feature vector and each preset category configured in advance, determine the probability that the object contained in this detection box is classified into each preset category, and according to the corresponding relationship between each preset category and the vector dimension stipulated in advance, output the classification vector and feature vector corresponding to this detection box.

[0149] In a possible implementation manner, the step of obtaining a set number of target preset categories in the classification vector corresponding to this detection box includes:

[0150] Sort the values corresponding to each dimension in the classification vector in descending order, select the preset number of target values ranked in the front, and determine the preset classification corresponding to each target value as the target preset category.

[0151] In a possible implementation manner, the step of determining the target category corresponding to the object contained in each detection box according to each of the node graphs, the classification image, and the pre - saved second model includes:

[0152] The second model determines the spatial distance between any two detection boxes in the to-be-classified image according to the position information of each detection box in the input to-be-classified image, and determines the distance matrix corresponding to the spatial distance; saves the target number of times for the edges connecting any two nodes in each node graph, and determines the relationship matrix corresponding to the target number of times; determines the first feature matrix corresponding to the node graph according to the feature vector corresponding to each node in the node graph; for the feature vector corresponding to each node in the node graph, updates the feature vector according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each pre-configured category, determines the probability that the object contained in the detection box corresponding to the updated feature vector is classified into each preset category, and determines the category corresponding to the maximum probability as the category corresponding to the updated feature vector; for any two connected nodes in the node graph, determines the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, and determines the loss value corresponding to the node graph according to the relationship probability corresponding to every two connected nodes in the node graph and a preset first function; determines the target category corresponding to the object contained in each detection box as the category corresponding to each node in the node graph with the minimum loss value.

[0153] In a possible implementation manner, the updating of the feature vector corresponding to each node in the node graph according to the relationship matrix, the distance matrix, and the feature matrix includes:

[0154] For each node in the node graph, perform an update for a preset number of times. Each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, and the relationship matrix and the distance matrix between this node and each other node in the node graph, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in this node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached.

[0155] In a possible implementation manner, the determining the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database includes:

[0156] For any two connected nodes in the node graph, obtain the number of times the category corresponding to any one of the nodes co-occurs with each of the other categories in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times the two nodes co-occur in the knowledge graph database.

[0157] Since the principle of the above computer-readable storage medium for solving problems is similar to the object classification method, the implementation of the above computer-readable storage medium can refer to the embodiments of the method, and the repeated parts will not be elaborated.

[0158] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0159] The present application is described with reference to the flowcharts and / or block diagrams of the method, device (system), and computer program product according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the specified functions in Figure 1 one process or multiple processes and / or blocksFigure 1 Steps of functions specified in one or more boxes.

[0162] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

Claims

1. A method for object classification, characterized in that, The method includes: Input the classification image pre-labeled with each detection box into the pre-saved first model, and obtain the classification vector and feature vector corresponding to each detection box output by the first model, where the value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension; wherein, the first model is used to classify the objects included in each detection box in the classification image; For each detection box, obtain a set number of target preset categories in the classification vector corresponding to the detection box; construct a node corresponding to the feature vector of the detection box, and save the target preset categories corresponding to the node; Combine any one of the target preset categories corresponding to each node, determine the node graph corresponding to each category combination, and any two nodes in the node graph are connected to each other. Among them, for any two connected nodes in the node graph, according to the target preset categories corresponding to the two nodes in the category combination, and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, determine the target number corresponding to the target preset category corresponding to the category combination, and save the target number for the edge connected by the two nodes; According to each of the node graphs, the classification image, and the pre-saved second model, determine the target category corresponding to the object included in each detection box; the second model is used to determine the node graph with the most accurate classification in each of the node graphs according to the information carried in the classification image, and output the node graph with the most accurate classification; The determining the target category corresponding to the object included in each detection box includes: Determine the target classification corresponding to the object corresponding to each node in the node graph output by the second model as the target classification of the object corresponding to each node.

2. The method according to claim 1, characterized in that, The inputting the classification image pre-labeled with each detection box into the first model and obtaining the classification vector and feature vector corresponding to each detection box output by the first model includes: The first model determines the feature vector corresponding to each detection box; according to the feature vector and each preset category configured in advance, determine the probability that the object included in the detection box is classified into each preset category, and according to the corresponding relationship between each preset category and the vector dimension specified in advance, output the classification vector and feature vector corresponding to the detection box.

3. The method according to claim 1, characterized in that, The obtaining a set number of target preset categories in the classification vector corresponding to the detection box includes: Sort the values corresponding to each dimension in the classification vector in descending order, select the preset number of target values ranked at the front, and determine the preset classification corresponding to each target value as the target preset category.

4. The method according to claim 1, characterized in that, The determining the target category corresponding to the object included in each detection box according to each of the node graphs, the classification image, and the pre-saved second model includes: The second model determines the spatial distance between any two detection boxes in the image to be classified according to the position information of each detection box in the input image to be classified, and determines the distance matrix corresponding to the spatial distance; for each node graph, the target number of times is saved for the edge connecting any two nodes, and the relationship matrix corresponding to the target number of times is determined; according to the feature vector corresponding to each node in the node graph, the first feature matrix corresponding to the node graph is determined; for the feature vector corresponding to each node in the node graph, according to the relationship matrix, the distance matrix, and the feature matrix, the feature vector is updated; according to the updated feature vector and each preset category, the probability that the object included in the detection box corresponding to the updated feature vector is classified into each preset category is determined, and the category corresponding to the maximum probability is determined as the category corresponding to the updated feature vector; for any two connected nodes in the node graph, according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database, the relationship probability of the categories corresponding to the two nodes is determined, and according to the relationship probability corresponding to every two connected nodes in the node graph and a preset first function, the loss value corresponding to the node graph is determined; the category corresponding to each node in the node graph with the minimum loss value is determined as the target category corresponding to the object included in each detection box.

5. The method according to claim 4, wherein The updating of the feature vector corresponding to each node in the node graph according to the relationship matrix, the distance matrix, and the feature matrix includes: For each node in the node graph, update is performed for a preset number of times. Each update process includes: determining the sum value of the product of the feature matrix, a preset parameter, and the relationship matrix and the distance matrix between this node and each other node in the node graph, and determining the updated feature vector of this node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in this node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached.

6. The method according to claim 4, characterized in that, The determining of the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times that every two preset categories appear simultaneously in the saved knowledge graph database includes: For any two connected nodes in the node graph, obtain the number of times that the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; according to each number of times and the number of times that the two nodes appear simultaneously in the knowledge graph database, determine the relationship probability of the categories corresponding to the two nodes.

7. An object classification device, characterized in that, The device includes: A classification module, configured to input an image to be classified pre-annotated with each detection box into a pre-saved first model, and obtain the classification vector and the feature vector corresponding to each detection box output by the first model, where each value corresponding to each dimension in the classification vector is the probability that the object included in the corresponding detection box is classified into the preset category corresponding to this dimension; wherein, the first model is used to classify the object included in each detection box in the image to be classified. A processing module, which is used to, for each detection box, obtain a set number of target preset categories in the classification vector corresponding to the detection box; construct a node corresponding to the feature vector of the detection box, and save the target preset categories corresponding to the node; combine any one of the target preset categories corresponding to each node to determine a node graph corresponding to each category combination, where any two nodes in the node graph are connected to each other. Among them, for any two connected nodes in the node graph, according to the target preset categories corresponding to the two nodes in the category combination and the number of times that every two preset categories appear simultaneously in the pre-saved knowledge graph database, determine the target number of times corresponding to the target preset category corresponding to the category combination, and save the target number of times for the edge connecting the two nodes; The classification module is further used to, according to each of the node graphs, the image to be classified, and a pre-saved second model, determine the target category corresponding to the object included in each detection box; the second model is used to, according to the information carried in the image to be classified, determine the node graph with the most accurate classification in each of the node graphs, and output the node graph with the most accurate classification; Specifically, the classification module is used to determine the target classification corresponding to the object of each node in the node graph output by the second model as the target classification of the object corresponding to each node.

8. The device according to claim 7, wherein Specifically, the classification module is used to, for each detection box, the first model determines the feature vector corresponding to the detection box; according to the feature vector and each pre-configured category, determine the probability that the object included in the detection box is classified into each preset category, and according to the corresponding relationship between each preset category and the vector dimension stipulated in advance, output the classification vector and the feature vector corresponding to the detection box.

9. The device according to claim 7, characterized in that, Specifically, the processing module is used to sort the values corresponding to each dimension in the classification vector in descending order, select a preset number of target values ranked in the front, and determine the preset classification corresponding to each target value as the target preset category.

10. The device according to claim 7, characterized in that, Specifically, the classification module is used to, according to the position information of each detection box in the input image to be classified, the second model determines the spatial distance between any two detection boxes in the image to be classified, and determines the distance matrix corresponding to the spatial distance; Save the target number of times for the edge connecting any two nodes in each node graph, and determine the relationship matrix corresponding to the target number of times; According to the feature vector corresponding to each node in the node graph, determine the first feature matrix corresponding to the node graph; For the feature vector corresponding to each node in the node graph, update the feature vector according to the relationship matrix, the distance matrix, and the feature matrix; according to the updated feature vector and each pre-configured category, determine the probability that the object included in the detection box corresponding to the updated feature vector is classified into each preset category, and determine the category corresponding to the maximum probability as the category corresponding to the updated feature vector; For any two connected nodes in the node graph, determine the relationship probability of the categories corresponding to the two nodes according to the categories corresponding to the two nodes and the number of times each two preset categories appear simultaneously in the stored knowledge graph database. According to the relationship probability corresponding to each two connected nodes in the node graph and a preset first function, determine the loss value corresponding to the node graph; determine the target category corresponding to the object included in each detection box as the category corresponding to each node in the node graph with the minimum loss value.

11. The device according to claim 10, characterized in that, The classification module is specifically configured to update each node in the node graph a preset number of times. Each update process includes: determining the sum value of the product of the feature matrix, the preset parameters, and the relationship matrix and the distance matrix between the node and each other node in the node graph, and determining the updated feature vector of the node according to the sum value and a second preset function; updating the feature matrix according to the updated feature vector of each node in the node, and performing the next update of the feature vector according to the updated feature matrix until the preset number of times is reached.

12. The device according to claim 9, characterized in that, The classification module is specifically configured to, for any two connected nodes in the node graph, obtain the number of times the category corresponding to any one node appears simultaneously with each other category in the knowledge graph database; determine the relationship probability of the categories corresponding to the two nodes according to each number of times and the number of times the two nodes appear simultaneously in the knowledge graph database.

13. An electronic device, characterized in that, The electronic device at least includes a processor and a memory. The processor is configured to implement the steps of the object classification method according to any one of claims 1-6 when executing the computer program stored in the memory.

14. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements the steps of the object classification method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image classification method and device

    CN110163301A

  • YOLO-based image target recognition method and apparatus, electronic device, and storage medium

    WO2020164282A1