Method and apparatus for updating an object recognition model
By combining visual and speech information to update the object recognition model, the problem that existing models cannot recognize untrained categories is solved, and a higher recognition rate and recognition results that are more in line with user needs are achieved.
Patent Information
- Application Number
- CN202010215064.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-03-24
AI Technical Summary
Existing object recognition models can only identify the object categories involved in the training process, cannot identify the object categories not involved, and the recognition results may not meet the user's needs.
By obtaining the target image collected by the camera device and the voice information collected by the voice device, the object recognition model is updated to include the characteristics and category tags of the target object, and combining voice indications and visual information to improve the recognition rate and accuracy of the model.
It enhances the recognition ability of the object recognition model, can identify more categories of objects, improves the recognition rate and the convenience of users to update the model, and ensures that the recognition results are more in line with user needs.
Smart Images

Figure CN113449548B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence (AI), and more specifically, to methods and apparatuses for updating an object recognition model. Background Art
[0002] Artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Research in the field of artificial intelligence includes robots, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theories, etc.
[0003] Object detection is one of the classic problems in computer vision, and its task is to use a bounding box to mark the position of an object in an image and give the category of the object. Among them, giving the category of an object in object detection is achieved through object recognition. Object recognition, which can also be understood as object classification, is a method of distinguishing different categories of objects according to the characteristics of the objects.
[0004] With the development of artificial intelligence technology, object recognition is no longer only achieved by relying on traditional manual recognition, but can also be achieved using an object recognition model based on deep learning.
[0005] An object recognition model based on deep learning usually uses a large number of object images of known categories to train the object recognition model, so that the object recognition model can learn the unique characteristics of different categories of objects and record the corresponding relationship between the characteristics of different categories of objects and the category labels. In this way, when an object image is input into the trained object recognition model in actual business applications, the category of the object can be inferred based on the object image, so as to achieve the purpose of object recognition.
[0006] For example, when a user uses a terminal device such as a mobile phone for object detection, a trained object recognition model can be used to identify the category of the object in the image captured by the terminal such as the mobile phone.
[0007] The above method of enabling an object recognition model to recognize the category of an object through training has the following problems: The object recognition model can only recognize the object categories involved in the training process. If the category of the target object does not belong to the object categories involved in the training process, the object recognition model cannot recognize the target object.
[0008] One way to solve this problem is that after training the object recognition model, during the use of the object recognition model, the object recognition model can be updated so that the object recognition model can recognize more types of objects.
[0009] Therefore, how to update the object recognition model has become a technical problem to be solved urgently. Summary of the Invention
[0010] This application provides a method and device for updating an object recognition model, enabling the object recognition model to recognize more objects and improving the recognition rate of the object recognition model.
[0011] In a first aspect, this application provides a method for updating an object recognition model. The method includes: obtaining a target image collected by a camera device; obtaining first voice information collected by a voice device, where the first voice information is used to indicate a first category of a target object in the target image; and updating a first object recognition model according to the target image and the first voice information. After the update, the first object recognition model includes features of the target object and a first label, and there is a corresponding relationship between the features of the target object and the first label, and the first label is used to represent the first category.
[0012] In the method of this application, since object features and the category label corresponding to the features can be added to the object recognition model, the object recognition model can recognize objects of this category, thereby improving the recognition rate of the object recognition model and further enhancing the intelligence of the object recognition model. In addition, the method of this application enables the user to indicate the category of the object to be updated through voice, which can improve the convenience of the user to update the object recognition model.
[0013] In combination with the first aspect, in a first possible implementation manner, the updating the first object recognition model according to the target image and the first voice information includes: determining, according to the similarity between the first label and each label in at least one category of labels, that the first label is a first category of labels in the at least one category of labels, and the similarity between the first label and the first category of labels is greater than the similarity between the first label and other labels in the at least one category of labels; using a second object recognition model to determine, according to the target image, a first probability that the category label of the target object is the first category of labels; and when the first probability is greater than or equal to a preset probability threshold, adding the features of the target object and the first label to a feature library of the first object recognition model.
[0014] In this implementation manner, when it is determined that the probability that the category indicated by the user's voice is the true category of the target object is relatively high, the first object recognition model is updated, which helps to improve the accuracy of the category label corresponding to the object feature updated to the first object recognition model, thereby improving the recognition accuracy of the first object recognition model.
[0015] Combined with the first possible implementation manner, in the second possible implementation manner, determining that the first label is the first type of label among the at least one type of label according to the similarity between the first label and each type of label in the at least one type of label includes: determining that the first label is the first type of label among the at least one type of label according to the similarity between the semantic features of the first label and the semantic features of each type of label in the at least one type of label; wherein, the similarity between the first label and the first type of label is greater than the similarity between the first label and other types of labels in the at least one type of label, including: the distance between the semantic features of the first label and the semantic features of the first type of label is less than the distance between the semantic features of the first label and the semantic features of the other types of labels.
[0016] Combined with the first aspect or any of the above possible implementation manners, in the third possible implementation manner, the target image includes a first object, and the target object is the object closest to the first object in the direction indicated by the first object in the target image, and the first object includes an eyeball or a finger.
[0017] In this implementation manner, an eyeball or a finger in the target image can be specified in advance as the pointing object, and the object in the direction indicated by the pointing object is determined as the target object, which helps to accurately mark the first category indicated by the user's voice to the target object specified by the user, thereby improving the recognition accuracy of the updated first object recognition model.
[0018] Combined with the third possible implementation manner, in the fourth possible implementation manner, updating the first object recognition model according to the target image and the first voice information includes: determining a bounding box of the first object in the target image according to the target image; determining the direction indicated by the first object according to the image in the bounding box; performing visual saliency detection on the target image to obtain a plurality of salient regions in the target image; determining a target salient region from the plurality of salient regions according to the direction indicated by the first object, where the target salient region is the salient region closest to the first object in the direction indicated by the first object among the plurality of salient regions; and updating the first object recognition model according to the target salient region, where the object in the target salient region includes the target object.
[0019] Combined with the fourth possible implementation, in the fifth possible implementation, determining the direction indicated by the first object based on the image in the bounding box includes: using a classification model to classify the image in the bounding box to obtain the target category of the first object; and determining the direction indicated by the first object according to the target category of the first object.
[0020] In a second aspect, the present application provides an apparatus for updating an object recognition model. The apparatus includes: an acquisition module, configured to acquire a target image collected by an imaging device; the acquisition module is further configured to acquire first voice information collected by a voice device, where the first voice information is used to indicate a first category of a target object in the target image; and an update module, configured to update a first object recognition model according to the target image and the first voice information. In the updated first object recognition model, features of the target object and the first label are included, and there is a corresponding relationship between the features of the target object and the first label, where the first label is used to represent the first category.
[0021] In the apparatus of the present application, since object features and the category label corresponding to the features can be added to the object recognition model, the object recognition model can recognize objects of this category, thereby improving the recognition rate of the object recognition model and further enhancing the intelligence of the object recognition model. In addition, the method of the present application enables the user to indicate the category of the object to be updated by voice, which can improve the convenience of the user to update the object recognition model.
[0022] Combined with the second aspect, in the first possible implementation, the update module is specifically configured to: determine that the first label belongs to a first category label among the at least one category of labels according to the similarity between the first label and each category of labels in the at least one category of labels, where the similarity between the first label and the first category label is greater than the similarity between the first label and other category labels in the at least one category of labels; use a second object recognition model to determine, according to the target image, a first probability that the category label of the target object is the first category label; and when the first probability is greater than or equal to a preset probability threshold, add the features of the target object and the first label to the first object recognition model.
[0023] In this implementation, the first object recognition model is updated only when the probability that the category indicated by the user is the true category of the target object is relatively high, which helps to improve the accuracy of the category label corresponding to the object features updated to the first object recognition model, thereby improving the recognition accuracy of the first object recognition model.
[0024] Combined with the first possible implementation manner, in the second possible implementation manner, the updating module is specifically configured to: determine that the first tag is the first type of tag among the at least one type of tags according to the similarity between the semantic features of the first tag and the semantic features of each type of tag in the at least one type of tags; wherein the similarity between the first tag and the first type of tag is greater than the similarity between the first tag and other types of tags in the at least one type of tags, including: the distance between the semantic features of the first tag and the semantic features of the first type of tag is less than the distance between the semantic features of the first tag and the semantic features of the other types of tags.
[0025] Combined with the second aspect or any of the above possible implementation manners, in the third possible implementation manner, the target image includes a first object, and the target object is the object closest to the first object in the direction indicated by the first object in the target image, and the first object includes an eyeball or a finger.
[0026] In this implementation manner, the eyeball or finger in the target image can be specified in advance as the pointing object, and the object in the direction indicated by the pointing object is determined as the target object, which helps to accurately mark the first tag indicated by the user through voice to the target object specified by the user, thereby improving the recognition accuracy of the updated first object recognition model.
[0027] Combined with the third possible implementation manner, in the fourth possible implementation manner, the updating module is specifically configured to: determine the bounding box of the first object in the target image according to the target image; determine the direction indicated by the first object according to the image in the bounding box; perform visual saliency detection on the target image to obtain a plurality of salient regions in the target image; determine a target salient region from the plurality of salient regions according to the direction indicated by the first object, where the target salient region is the salient region closest to the first object in the direction indicated by the first object among the plurality of salient regions; update the first object recognition model according to the target salient region, where the object in the target salient region includes the target object.
[0028] Combined with the fourth possible implementation manner, in the fifth possible implementation manner, the updating module is specifically configured to: use a classification model to classify the image in the bounding box to obtain the target category of the first object; determine the direction indicated by the first object according to the target category of the first object.
[0029] In a third aspect, the present application provides a method for updating an object recognition model, the method comprising: obtaining a target image collected by a camera device; obtaining first indication information for indicating a first category of a target object in the target image; when the target confidence of the first category being the actual category of the target object is greater than or equal to a preset confidence threshold, updating a first object recognition model according to the target image and the first indication information, wherein a feature library of the updated first object recognition model includes features of the target object and a first label, and there is a corresponding relationship between the features of the target object and the first label, and the first label is used to represent the first category; wherein, the target confidence is determined according to a first probability, the first probability being the probability that a second object recognition model recognizes the target object and the first label is a first type of label, the second object recognition model being used to recognize an image to obtain the probability that the category label of an object in the image is each type of label in at least one type of label, the at least one type of label being obtained by clustering the category labels corresponding to the features in the feature library of the first object recognition model, the at least one type of label including the first type of label, and the similarity between the first label and the first type of label being greater than the similarity between the first label and other types of labels in the at least one type of label.
[0030] In this method, the first object recognition model is updated only when the confidence that the first label is the true label of the target object is relatively high. Wherein, this confidence is determined according to the probability that the second object recognition model infers from the target image that the category label of the object to be detected is the label class to which the first label belongs. In this way, it helps to improve the accuracy of the category label corresponding to the object feature updated to the first object recognition model, and thus the recognition accuracy of the first object recognition model can be improved.
[0031] In combination with the third aspect, in a first possible implementation manner, the updating the first object recognition model according to the target image and the first indication information includes: determining that the first label is the first type of label according to the similarity between the first label and each type of label in the at least one type of label; inputting the target image into the second object recognition model to obtain the first probability; determining the target confidence according to the first probability; and when the target confidence is greater than or equal to the confidence threshold, adding the features of the target object and the first label to the first object recognition model.
[0032] Combined with the third aspect or the first possible implementation manner, in the second possible implementation manner, the similarity between the first tag and the first type of tags is greater than the similarity between the first tag and other types of tags in the at least one type of tags, including: the distance between the semantic features of the first tag and the semantic features of the first type of tags is less than the distance between the semantic features of the first tag and the semantic features of the other types of tags.
[0033] Combined with the second possible implementation manner, in the third possible implementation manner, the confidence level is the first probability.
[0034] Combined with the third aspect or any of the above possible implementation manners, in the fourth possible implementation manner, the first indication information includes voice information collected by a voice device or text information collected by a touch device.
[0035] In a fourth aspect, the present application provides a device for updating an object recognition model, and the device includes corresponding modules for implementing the method in the third aspect or any one of the implementation manners.
[0036] In a fifth aspect, the present application provides a device for updating an object recognition model, and the device includes: a memory for storing instructions; a processor for executing the instructions stored in the memory, and when the instructions stored in the memory are executed, the processor is used to execute the method in the first aspect or any one of the possible implementation manners.
[0037] In a sixth aspect, the present application provides a device for updating an object recognition model, and the device includes: a memory for storing instructions; a processor for executing the instructions stored in the memory, and when the instructions stored in the memory are executed, the processor is used to execute the method in the third aspect or any one of the possible implementation manners.
[0038] In a seventh aspect, the present application provides a computer-readable medium, and the computer-readable medium stores instructions for a device to execute, and the instructions are used to implement the method in the first aspect or any one of the possible implementation manners.
[0039] In an eighth aspect, the present application provides a computer-readable medium, and the computer-readable medium stores instructions for a device to execute, and the instructions are used to implement the method in the third aspect or any one of the possible implementation manners.
[0040] In a ninth aspect, the present application provides a computer program product containing instructions, and when the computer program product runs on a computer, the computer is caused to execute the method in the first aspect or any one of the possible implementation manners.
[0041] In a tenth aspect, the present application provides a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the method in the third aspect or any one of its possible implementation manners.
[0042] In an eleventh aspect, the present application provides a chip. The chip includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in the first aspect or any one of its possible implementation manners.
[0043] Optionally, as an implementation manner, the chip may further include a memory. Instructions are stored in the memory. The processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the method in the first aspect or any one of its possible implementation manners.
[0044] In a twelfth aspect, the present application provides a chip. The chip includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in the third aspect or any one of its possible implementation manners.
[0045] Optionally, as an implementation manner, the chip may further include a memory. Instructions are stored in the memory. The processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the method in the third aspect or any one of its possible implementation manners.
[0046] In a thirteenth aspect, the present application provides a computing device. The computing device includes a processor and a memory, where: computer instructions are stored in the memory, and the processor executes the computer instructions to implement the method in the first aspect or any one of its possible implementation manners.
[0047] In a fourteenth aspect, the present application provides a computing device. The computing device includes a processor and a memory, where: computer instructions are stored in the memory, and the processor executes the computer instructions to implement the method in the third aspect or any one of its possible implementation manners.
[0048] In a fifteenth aspect, the present application provides an object recognition method. The method includes: obtaining a to-be-recognized image collected by a camera device; using a first object recognition model to perform category recognition on the to-be-recognized image, where the first object recognition model is obtained by adding the features of a target object and a first label to a target image collected by the camera device and first voice information collected by a voice device, and there is a corresponding relationship between the features of the target object and the first label. The first voice information is used to indicate a first category of the target object in the target image, and the first label is used to represent the first category.
[0049] In some possible implementation manners, the features of the target object and the first label are added to the first object recognition model when a first probability is greater than or equal to a preset probability threshold. The first probability is obtained by using a second object recognition model based on the target image, and the first probability is the probability that the class label of the target object is the first class label. The first class label is determined based on the first label and at least one class of labels, and the similarity between the first label and the first class label is greater than the similarity between the first label and other class labels in the at least one class of labels.
[0050] In some possible implementation manners, that the similarity between the first label and the first class label is greater than the similarity between the first label and other class labels in the at least one class of labels includes: the distance between the semantic features of the first label and the semantic features of the first class label is less than the distance between the semantic features of the first label and the semantic features of the other class labels.
[0051] In some possible implementation manners, the target image includes a first object, and the target object is the object closest to the first object in the direction indicated by the first object in the target image. The first object includes an eyeball or a finger.
[0052] In a sixteenth aspect, the present application provides an object recognition device, and the device includes corresponding modules for implementing the method in the fifteenth aspect or any one of the implementation manners.
[0053] In a seventeenth aspect, the present application provides an object recognition device, and the device includes: a memory for storing instructions; a processor for executing the instructions stored in the memory. When the instructions stored in the memory are executed, the processor is used to execute the method in the fifteenth aspect or any one of the possible implementation manners.
[0054] In an eighteenth aspect, the present application provides a computer-readable medium, and the computer-readable medium stores instructions for a device to execute. The instructions are used to implement the method in the fifteenth aspect or any one of the possible implementation manners.
[0055] In a nineteenth aspect, the present application provides a computer program product including instructions. When the computer program product runs on a computer, the computer is caused to execute the method in the fifteenth aspect or any one of the possible implementation manners.
[0056] In a twentieth aspect, the present application provides a chip. The chip includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in the fifteenth aspect or any one of the possible implementation manners.
[0057] Optionally, as an implementation, the chip may further include a memory, in which instructions are stored, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the method in the fifteenth aspect or any one of its possible implementations.
[0058] In a twenty-first aspect, the present application provides a computing device, which includes a processor and a memory, where: computer instructions are stored in the memory, and the processor executes the computer instructions to implement the method in the fifteenth aspect or any one of its possible implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic diagram of an artificial intelligence main framework according to the present application.
[0060] Figure 2 It is a schematic structural diagram of a system architecture provided by an embodiment of the present application.
[0061] Figure 3 It is a schematic structural diagram of a convolutional neural network provided by an embodiment of the present application.
[0062] Figure 4 It is a schematic structural diagram of another convolutional neural network provided by an embodiment of the present application.
[0063] Figure 5 It is a schematic hardware structure diagram of a chip provided by an embodiment of the present application.
[0064] Figure 6 It is a schematic diagram of a system architecture provided by an embodiment of the present application.
[0065] Figure 7 It is a schematic flowchart of a method for updating an object recognition model according to an embodiment of the present application.
[0066] Figure 8 It is a schematic diagram of clustering labels based on label features to obtain at least one class of labels.
[0067] Figure 9 It is a schematic flowchart of a method for updating an object recognition model according to another embodiment of the present application.
[0068] Figure 10 It is a schematic flowchart of a method for updating an object recognition model according to another embodiment of the present application.
[0069] Figure 11 It is a schematic diagram of a method for determining a main body area in the gesture indication direction.
[0070] Figure 12This is a schematic flowchart for updating an object recognition model based on a user instruction in this application.
[0071] Figure 13 This is another schematic flowchart for updating an object recognition model based on a user instruction in this application.
[0072] Figure 14 This is an exemplary structural diagram of an apparatus for updating an object recognition model in this application.
[0073] Figure 15 This is another exemplary structural diagram of an apparatus for updating an object recognition model in this application. Detailed implementation manners
[0074] Figure 1 This shows a schematic diagram of an artificial intelligence main framework, which describes the overall working process of an artificial intelligence system and is applicable to the general requirements in the field of artificial intelligence.
[0075] The above artificial intelligence main framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis).
[0076] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom".
[0077] The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecosystem of the system.
[0078] (1) Infrastructure:
[0079] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through a basic platform. It communicates with the external through sensors; the computing power is provided by intelligent chips (such as CPU, NPU, GPU, ASIC, FPGA and other hardware acceleration chips); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0080] (2) Data
[0081] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.
[0082] (3) Data processing
[0083] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0084] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0085] Reasoning refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, and using formalized information for machine thinking and problem-solving according to the reasoning control strategy. The typical function is search and matching.
[0086] Decision-making refers to the process of making decisions after the intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.
[0087] (4) General capabilities
[0088] After the data undergoes the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0089] (5) Intelligent products and industry applications
[0090] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields. It is the encapsulation of the overall artificial intelligence solution, productizing intelligent information decision-making and realizing landing applications. Its application fields mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, safe city, intelligent terminal, etc.
[0091] For example, a large number of images containing objects can be obtained as data; then the data is processed, that is, deep learning is performed on the correlation between the object categories and object features in the images; after data processing, an object recognition model with general capabilities can be obtained; deploying the object recognition model to the infrastructure, such as deploying it to devices such as robots, an intelligent product with object recognition capabilities can be obtained; after the intelligent product collects an image, the object recognition model deployed on it can be used to recognize the image to obtain the category of the object in the image, thus realizing the industry application of object recognition.
[0092] Object recognition in this application can also be referred to as image recognition, which refers to the technology of using a computer to process, analyze, and understand images to identify various different types of objects in the images. The model used to implement object recognition is called an object recognition model, and can also be called an image recognition model.
[0093] The object recognition model can be obtained through training. The following will introduce an exemplary method for training the object recognition model in conjunction with Figure 2 an introduction to an exemplary method for training the object recognition model.
[0094] In Figure 2 it, the data acquisition device 260 is used to acquire training data. For example, the training data can include training images and the corresponding categories of the training images, where the results of the training images can be pre-annotated manually.
[0095] After the training data is acquired, the data acquisition device 260 stores these training data in the database 230, and the training device 220 trains to obtain the object recognition model 201 based on the training data maintained in the database 230.
[0096] The following will describe how the training device 220 obtains the object recognition model 201 based on the training data. The training device 220 processes the input original image, compares the output image category with the category annotated for the original image, until the difference between the image category output by the training device 120 and the category annotated for the original image is less than a certain threshold, thereby completing the training of the object recognition model 201.
[0097] The above object recognition model 201 can be used for object recognition. The object recognition model 201 in the embodiments of this application can specifically be a neural network. It should be noted that in actual applications, the training data maintained in the database 230 does not necessarily all come from the acquisition of the data acquisition device 260, and it may also be received from other devices. Additionally, it should be noted that the training device 220 does not necessarily train the object recognition model completely based on the training data maintained in the database 230, and it may also obtain training data from the cloud or other places for model training. The above description should not be used as a limitation to the embodiments of this application.
[0098] The object recognition model 201 trained according to the training device 220 can be applied to different systems or devices, such as being applied to the execution device 210 shown in Figure 2 where the execution device 210 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle-mounted terminal, etc., and can also be a server or the cloud, etc. In Figure 2In it, the execution device 210 configures an input / output (I / O) interface 212 for data interaction with external devices. A user can input data to the I / O interface 212 through a client device 240. The input data in the embodiments of the present application may include: an image to be recognized input by the client device.
[0099] A preprocessing module 213 is used to preprocess the input data (such as an image to be processed) received by the I / O interface 212. In the embodiments of the present application, the preprocessing module 213 may also be absent.
[0100] When the execution device 210 preprocesses the input data, or when a processing module 211 of the execution device 210 performs related processes such as calculations, the execution device 210 can call data, code, etc. in the data storage system 250 for corresponding processing, or can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 250.
[0101] Finally, the I / O interface 212 returns the processing result, such as the obtained image category, to the client device 240, and thus provides it to the user.
[0102] In Figure 2 In the case shown, the user can manually give the input data, and this manual giving can be operated through the interface provided by the I / O interface 212. In another case, the client device 240 can automatically send the image to be recognized to the I / O interface 212. If the client device 240 is required to automatically send the image to be recognized and user authorization is needed, the user can set the corresponding permissions in the client device 240. The user can view the category recognition result output by the execution device 210 in the client device 240, and the specific presentation form can be specific ways such as display, sound, action, etc. The client device 240 can also be used as a data acquisition end to acquire the image to be recognized input to the I / O interface 212 and the category recognition result output from the I / O interface 212 as new sample data, and store them in the database 230. Of course, it can also be acquired without passing through the client device 240, but directly store the input data of the I / O interface 212 and the output result of the I / O interface 212 as new sample data in the database 230 by the I / O interface 212.
[0103] It should be noted that Figure 2 is only a schematic diagram of a method for training an object recognition model provided by the embodiments of the present application. Figure 2 The positional relationship between the devices, components, modules, etc. shown in Figure 2 does not constitute any limitation. For example, in Figure 2 the data storage system 250 is an external memory relative to the execution device 210. In other cases, the data storage system 250 can also be placed in the execution device 210.
[0104] The object recognition model can be implemented through a neural network. Further, it can be implemented through a deep neural network. Even further, it can be implemented through a convolutional neural network.
[0105] To better understand the solution of the embodiments of the present application, the following first introduces the relevant terms of the neural network and other related concepts that may be involved in the embodiments of the present application.
[0106] (1) Neural network
[0107] A neural network (neural networks, NN) is a complex network system formed by a large number of simple processing units (called neurons) widely interconnected. It reflects many basic characteristics of the human brain function and is a highly complex non-linear dynamic learning system.
[0108] A neural network can be composed of neural units. A neural unit can refer to an operation unit with x s and intercept 1 as inputs. The output of this operation unit can be as shown in formula (1-1):
[0109]
[0110] where s = 1, 2,... n, n is a natural number greater than 1, W s is the weight of x s and b is the bias of the neural unit. f is the activation function of the neural unit (activation functions), which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.
[0111] (2) Deep neural network
[0112] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. When dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is to say, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer.
[0113] Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the bias vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs such a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and bias vectors is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer.
[0114] In summary, the coefficient from the k-th neuron in the L - 1-th layer to the j-th neuron in the L-th layer is defined as
[0115] It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically speaking, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0116] (3) Convolutional neural network (CNN)
[0117] A convolutional neural network is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature plane. A convolutional layer refers to the layer of neurons in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can only be connected to some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weights here are the convolutional kernels. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of an image is the same as that of other parts. That is to say, the image information learned in a certain part can also be used in another part. So for all positions on the image, we can use the same learned image information. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more convolutional kernels there are, the richer the image information reflected by the convolution operation.
[0118] The convolutional kernels can be initialized in the form of matrices of random sizes, and during the training process of the convolutional neural network, the convolutional kernels can learn to obtain reasonable weights. Additionally, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network while reducing the risk of overfitting.
[0119] The structure of the convolutional neural network in the embodiments of this application can be as Figure 3 shown. In Figure 3 , the convolutional neural network (CNN) 300 can include an input layer 310, a convolutional layer / pooling layer 320 (where the pooling layer is optional), and a neural network layer 330.
[0120] Among them, the input layer 310 can obtain the image to be recognized and hand over the obtained image to be recognized to the convolutional layer / pooling layer 320 and the subsequent neural network layer 330 for processing, and the category recognition result of the image can be obtained.
[0121] Next, a detailed introduction will be given to the internal layer structure of the CNN 300 in Figure 3 .
[0122] Convolutional layer / pooling layer 320:
[0123] Convolutional layer:
[0124] As Figure 3The convolutional layer / pooling layer 320 shown may include layers such as examples 321 - 326. For example: In one implementation, layer 321 is a convolutional layer, layer 322 is a pooling layer, layer 323 is a convolutional layer, layer 324 is a pooling layer, layer 325 is a convolutional layer, and layer 326 is a pooling layer; in another implementation, 321 and 322 are convolutional layers, 323 is a pooling layer, 324 and 325 are convolutional layers, and 326 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.
[0125] Taking the convolutional layer 321 as an example, the internal working principle of one convolutional layer will be introduced below.
[0126] The convolutional layer 321 may include many convolutional operators, also known as kernels, which act as a filter for extracting specific information from the input image matrix in image recognition. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolutional operation on the image, the weight matrix usually processes the input image pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction, thus completing the work of extracting specific features from the image. The size of this weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image, and during the convolutional operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices with the same size (row × column), that is, multiple matrices of the same type, are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image, and here the dimension can be understood as being determined by the above-mentioned "multiple". Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors in the image, and yet another weight matrix is used to blur the unwanted noise in the image, etc. The sizes (row × column) of these multiple weight matrices are the same, and the sizes of the convolutional feature maps extracted by these multiple weight matrices of the same size are also the same. Then, the multiple convolutional feature maps of the same size that are extracted are combined to form the output of the convolutional operation.
[0127] The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. The individual weight matrices formed by the weight values obtained through training can be used to extract information from the input image, so that the convolutional neural network 200 can make correct predictions.
[0128] When the convolutional neural network 300 has multiple convolutional layers, the initial convolutional layer (for example, 321) often extracts more general features, which can also be called low-level features. As the depth of the convolutional neural network 300 increases, the features extracted by the later convolutional layers (for example, 326) become more and more complex, such as high-level semantic features. Features with higher semantics are more suitable for the problem to be solved.
[0129] Pooling layer:
[0130] Since it is often necessary to reduce the number of training parameters, it is often necessary to periodically introduce a pooling layer after the convolution layer. Figure 3 The layers 321-326 illustrated in 320 may be a convolution layer followed by a pooling layer, or may be multiple convolution layers followed by one or more pooling layers. For example, in the image recognition process, the only purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer may include an average pooling operator and / or a maximum pooling operator for sampling the input image to obtain an image of smaller size. The average pooling operator may calculate the pixel values in the image within a specific range to generate an average value as the result of average pooling. The maximum pooling operator may take the pixel with the largest value within a specific range as the result of maximum pooling. In addition, just as the size of the weight matrix used in the convolution layer should be related to the image size, the operator in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer may be smaller than the size of the image input to the pooling layer, and each pixel in the image output by the pooling layer represents the average value or maximum value of the corresponding sub-region of the image input to the pooling layer.
[0131] Neural network layer 330:
[0132] After being processed by the convolution layer / pooling layer 320, the convolution neural network 300 is not sufficient to output the required output information. As mentioned above, the convolution layer / pooling layer 320 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolution neural network 300 needs to use the neural network layer 330 to generate one or a group of outputs of the required number of classes. Therefore, the neural network layer 330 may include multiple hidden layers (such as Figure 3 331, 332 to 33n) and the output layer 340 shown, the parameters contained in the multi-layer hidden layers can be pre-trained according to relevant training data for image recognition.
[0133] After the multi-layer hidden layer in the neural network layer 330, that is, the last layer of the entire convolutional neural network 300 is the output layer 340. The output layer 340 has a loss function similar to categorical cross-entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 300 (such as Figure 3 The propagation from 310 to 340 is forward propagation) is completed, the backpropagation (such as Figure 3 The propagation from 340 to 310 is backpropagation) will start to update the weights and biases of the previously mentioned layers to reduce the loss of the convolutional neural network 300, that is, the error between the class recognition result output by the convolutional neural network 300 and the ideal class.
[0134] The structure of the convolutional neural network in the embodiments of the present application can be as Figure 4 shown. In Figure 4 , the convolutional neural network (CNN) 400 may include an input layer 410, a convolutional layer / pooling layer 420 (where the pooling layer is optional), and a neural network layer 430. Compared with Figure 3 , Figure 4 In the convolutional layer / pooling layer 420 in , multiple convolutional layers / pooling layers (421 to 426) are parallel, and the respectively extracted features are all input to the full neural network layer 430 for processing. The neural network layer 430 may include multiple hidden layers, that is, hidden layer 1 to hidden layer n, which can be denoted as 431 to 43n.
[0135] It should be noted that Figure 3 and Figure 4 The convolutional neural networks shown are only examples of two possible convolutional neural networks in the embodiments of the present application. In specific applications, the convolutional neural network in the embodiments of the present application may also exist in the form of other network models.
[0136] Figure 5 This is the hardware structure of a chip for running or training an object recognition model provided by the embodiments of the present application. The chip includes a neural network processor 50. The chip can be set in an execution device 210 as shown in Figure 2 to complete the computing work of the processing module 211. The chip can also be set in a training device 220 as shown in Figure 2 to complete the training work of the training device 220 and output the object recognition model 201. The algorithms of each layer in the convolutional neural networks shown in Figure 3 and Figure 4 can all be implemented in the chip as shown in Figure 5 .
[0137] The neural network processor NPU 50 is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and tasks are assigned by the main CPU. The core part of the NPU is the arithmetic circuit 503, and the controller 504 controls the arithmetic circuit 503 to extract data from the memory (weight memory or input memory) and perform operations.
[0138] In some implementations, the arithmetic circuit 503 includes multiple processing engines (PEs) inside. In some implementations, the arithmetic circuit 503 is a two-dimensional systolic array. The arithmetic circuit 503 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 503 is a general matrix processor.
[0139] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 501 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 508.
[0140] The vector calculation unit 507 can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. For example, the vector calculation unit 507 can be used for network calculations in non-convolutional / non-FC layers of the neural network, such as pooling, batch normalization, local response normalization, etc.
[0141] In some implementations, the vector calculation unit 507 can store the processed output vector into the unified buffer 506. For example, the vector calculation unit 507 can apply a non-linear function to the output of the arithmetic circuit 503, such as a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 507 generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 503, such as for use in subsequent layers in the neural network.
[0142] The unified memory 506 is used to store input data and output data.
[0143] The weight data directly transfers the input data in the external memory to the input memory 501 and / or the unified memory 506 through the direct memory access controller (DMAC) 505, stores the weight data in the external memory into the weight memory 502, and stores the data in the unified memory 506 into the external memory.
[0144] The bus interface unit (BIU) 510 is used to interact between the main CPU, the DMAC, and the instruction fetch buffer 509 through the bus.
[0145] The instruction fetch buffer 509 connected to the controller 504 is used to store the instructions used by the controller 504.
[0146] The controller 504 is used to call the instructions cached in the instruction fetch memory 509 to control the working process of the arithmetic accelerator.
[0147] Generally, the unified memory 506, the input memory 501, the weight memory 502, and the instruction fetch memory 509 are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.
[0148] Among them, Figure 3 and Figure 4 The operations of each layer in the convolutional neural network shown can be executed by the arithmetic circuit 503 or the vector calculation unit 507.
[0149] Such as Figure 6 As shown, the embodiment of the present application provides a system architecture 600. The system architecture includes local devices 601, local devices 602, an execution device 610, and a data storage system 650. Among them, the local devices 601 and 602 are connected to the execution device 610 through a communication network.
[0150] The execution device 610 can be implemented by one or more servers. Optionally, the execution device 610 can be used in cooperation with other computing devices, such as data storage devices, routers, load balancers, and other devices. The execution device 610 can be arranged at a single physical site or distributed across multiple physical sites. The execution device 610 can use the data in the data storage system 650 or call the program code in the data storage system 650 to implement the object recognition method of the embodiments of the present application.
[0151] Users can operate their respective user devices (such as local device 601 and local device 602) to interact with the execution device 610. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart vehicle, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc.
[0152] The local device of each user can interact with the execution device 610 through a communication network using any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, or any combination thereof.
[0153] In one implementation, the local device 601 and the local device 602 collect the images to be recognized and send the images to be recognized to the execution device 610. The execution device 610 uses the object recognition model deployed thereon to recognize the images to be recognized and returns the recognition result to the local device 601 or the local device 602.
[0154] In another implementation, the object recognition model can be directly deployed on the local device 601 or the local device 602. In this way, after the local device 601 or the local device 602 captures the images to be recognized through the imaging device, it can use the object recognition model to recognize the images to be recognized by itself.
[0155] Currently, for the object recognition model, generally, a large number of object images of known categories are used to train the object recognition model, so that the object recognition model can learn the unique features of different categories of objects and record the corresponding relationship between the features of different categories of objects and the category labels. In this way, when the trained object recognition model inputs an object image in the actual business application, it can infer the category of the object based on the object image, so as to achieve the purpose of object recognition.
[0156] For example, when a user uses a terminal device such as a mobile phone for object detection, the trained object recognition model can be used to recognize the category of the object in the image captured by the terminal such as the mobile phone.
[0157] The method of training the object recognition model to recognize the categories of objects has the following problems: the object recognition model can only recognize the object categories involved in the training process. If the category of the target object does not belong to the object categories involved in the training process, the object recognition model cannot recognize the target object; or, after the object recognition model recognizes the object category, the recognition result does not meet the user's needs. For example, the object recognition model recognizes the object in an image as a dog, but what the user needs is the recognition result of "Husky".
[0158] To address the above problems, after training the object recognition model, during the use of the object recognition model, the user can run to update the object recognition model so that the recognition result of the object recognition model can meet the user's needs. Therefore, this application proposes a method for updating the object recognition model.
[0159] First, several exemplary application scenarios of the method for updating the object recognition model in this application will be introduced below.
[0160] Application scenario 1:
[0161] When the user uses a robot, they hope that the robot can recognize the objects they are interested in.
[0162] For example, if the user hopes that the robot can recognize the Doraemon in their hand, the user can point to the doll and say to the robot, "This is Doraemon". The robot simultaneously obtains visual information and the user's voice command, then obtains the knowledge "This is Doraemon" based on the user's voice command, obtains the image features corresponding to "Doraemon" based on the visual information, and generates the correspondence between the category label "Doraemon" and the corresponding image, and inputs the category label "Doraemon" and the corresponding image into the robot's model. In this way, the updated model has the ability to recognize "Doraemon".
[0163] Application scenario 2:
[0164] When the user uses a robot to recognize the objects they are interested in, they are not satisfied with the object category output by the robot and inform the robot of the more accurate category of the object.
[0165] For example, the user points to Peppa Pig, and the robot obtains visual information and outputs the category "pig" based on the visual information. The user is not satisfied and says to the robot, "Wrong, this is Peppa Pig". The robot simultaneously obtains visual information and the user's voice command, generates the correspondence between the category label "Peppa Pig" and the corresponding image, and inputs the category label "Peppa Pig" and the corresponding image into the robot's model. The updated model of the robot has the ability to recognize "Peppa Pig".
[0166] Application scenario 3:
[0167] When the user uses the robot to identify an object of interest, if the object category output by the robot is incorrect, the user tells the robot the correct category of the object.
[0168] For example, a child says to a children's robot, "Xiaoyi Xiaoyi, help me recognize an object." The robot says to the child, "Please put the object in front of me." The child puts the object in front of the robot. The robot identifies the name of the object and says to the child, "I guess this is an apple. Am I right?" The child says, "No, this is an orange." The robot adds the features corresponding to the orange and the category label "orange" to the model and says, "I remember. Next time I see it, I'll be able to recognize it."
[0169] The following introduces a schematic flowchart of the method for updating the object recognition model of the present application.
[0170] Figure 7 It is a schematic flowchart of the method for updating the object recognition model according to an embodiment of the present application. As Figure 4 shown, the method may include S710, S720, and S730. The method may be executed by the aforementioned execution device or local device.
[0171] S710, obtain a target image.
[0172] The target image may be an image obtained by a camera device on a smart device for visual information collection. The target image may include one or more objects.
[0173] S720, obtain first voice information, where the first voice information is used to indicate a first category of a target object in the target image.
[0174] The target object refers to an object of interest to the user in the target image, that is, an object whose category the user hopes to know. The first voice information may be voice information collected by a voice device on the smart device, for example, voice information collected by a voice device such as a microphone.
[0175] After the smart device collects a voice command input by the user, it can obtain the knowledge in the voice command, thereby obtaining the first voice information. For example, the method of natural language understanding may be used to obtain the knowledge in the voice command, thereby obtaining the first voice information.
[0176] As an example, the first voice information may include the following content "This is A", where A is the category label of the object. For example, this is Peppa Pig, this is an orange, this is Doraemon, where Peppa Pig, orange, and Doraemon are the first categories of the target objects in the target image.
[0177] S730. Update the first object recognition model according to the target image and the first voice information. In the updated first object recognition model, the features of the target object and the first label are included, and there is a corresponding relationship between the features of the target object and the first label. The first label is used to represent the first category.
[0178] In other words, a first label for representing the first category can be generated, the features of the target object and the first label, as well as the corresponding relationship between the features of the target object and the first label, are added to the first object recognition model.
[0179] The first object recognition model is used to identify the category of the object in the image. In some examples, after the first object recognition model inputs an image, it can output the category label of the object in the image.
[0180] The first object recognition model can be a neural network model. In some examples, the first object recognition model can be obtained through training with a training set, which can include a large number of images. These images can include objects of different categories, and the categories of these objects are known, that is, the labels corresponding to the images in the training set are known. The method of obtaining the first object recognition model through training with the training set can refer to the prior art and will not be elaborated here.
[0181] For example, the training set of the first object recognition model can include the ImageNet dataset and the corresponding labels. The ImageNet dataset refers to the publicly available dataset used in the ImageNet large scale visual recognition challenge (ILSVRC) competition.
[0182] Another example is that the training set of the first object recognition model can include the OpenImages dataset and the corresponding labels.
[0183] The features of the target object can be obtained by the first object recognition model extracting features from the target image. For example, the first object recognition model can include a feature extraction sub-model, and use this feature extraction sub-model to extract the features in the target image. The feature extraction sub-model can be a dense convolutional neural network, a dilated neural network, a residual neural network, etc.
[0184] The method of the embodiment of the present application improves the recognition rate of the object recognition model and further enhances the intelligence of the object recognition model by adding object features and the category label corresponding to the features to the object recognition model, so that the object recognition model can recognize the objects of this category.
[0185] In addition, the method of the embodiment of the present application enables the user to indicate the category of the target object through voice, which can improve the convenience for the user to update the object recognition model.
[0186] In some possible implementation manners, S730, that is, updating the first object recognition model according to the target image and the first voice information, may include: determining, according to the target image and the first voice information, a target confidence that the first category indicated by the first voice information is the actual category of the target object; when the target confidence is greater than or equal to a preset confidence threshold, adding the features of the target object, the first label, and the correspondence between the features and the first label to the first object recognition model.
[0187] In these implementation manners, when it is determined that the confidence of the first category indicated by the user for the target object through voice is relatively high, the features of the target object, the first label, and the correspondence are updated to the first object recognition model.
[0188] Or rather, when it is determined that the confidence of the first category indicated by the user for the target object through voice is relatively low, for example, less than or less than the preset confidence threshold, the first object recognition model may not be updated. For example, when the first label in the first voice information input by the user is incorrect, or when an error occurs in obtaining the first voice information in the user's voice, or when the obtained target object is not the object that the user specifies to be recognized, the first object recognition model may not be updated.
[0189] These implementation manners can improve the recognition accuracy of the updated first object recognition model, and further improve the intelligence of the first object recognition model.
[0190] In some possible implementation manners, the target confidence may be determined according to a first probability, the first probability may be inferred by a second object recognition model for the category of the target object, and indicates the probability that the category label of the target object is a first type of label. The second object recognition model is used to recognize an image to obtain the probability that the category label of the object in the image is each category label in at least one category of labels. The at least one category of labels is obtained by clustering the category labels in the first object recognition model. The at least one category of labels includes the first type of label, and the similarity between the first label and the first type of label is greater than the similarity between the first label and other category labels in the at least one category of labels.
[0191] In this implementation manner, the second object recognition model may be a neural network, for example, a convolutional neural network.
[0192] An exemplary method for obtaining the at least one category of labels is introduced below.
[0193] In an exemplary method, the class labels in the first object recognition model can be obtained first, and the semantic features of each class label in the first object recognition model can be extracted; then, based on the semantic features of all class labels, all class labels in the first object recognition model can be clustered to obtain the at least one class of labels, where each class of labels in the at least one class of labels can include one or more class labels among all class labels of the first object recognition model.
[0194] For example, the BERT model can be used to extract the semantic features of each class label in the first object recognition model. The full name of BERT is bidirectional encoder representation from transformers.
[0195] For example, the k-means method can be used to cluster the class labels of the first object recognition model based on the semantic features of the class labels of the first object recognition model to obtain at least one class of labels.
[0196] Figure 8 FIG. is a schematic diagram for clustering labels based on label features to obtain at least one class of labels. Figure 8 In, one point represents one label feature, and one oval box represents one class of labels.
[0197] Taking the training set of the first object recognition model including the ImageNet dataset as an example, the BERT model can be used to extract features of 1000 class labels corresponding to the ImageNet dataset; using K-means based on the features of these 1000 class labels, these 1000 class labels can be clustered into 200 class labels. These 200 class labels are the above-mentioned at least one class of labels.
[0198] An exemplary method for obtaining the second object recognition model is introduced below.
[0199] In an exemplary method, the training set of the first object recognition model can be obtained, and the class label of each training data in the training set can be modified to the class label corresponding to it, so as to obtain a new training set; then, the new training set can be used to train a classification model, so as to obtain the second object recognition model. When the trained second object recognition model inputs an image, it can infer which class of label the object in the image corresponds to.
[0200] Taking the training set of the first object recognition model as including the ImageNet dataset and the above 200 class labels obtained by clustering as an example, the mobilenet model is trained using the ImageNet dataset and these 200 class labels to obtain a second object recognition model. Among them, the label corresponding to each image in the ImageNet dataset is mapped from the original 1000 class labels to the corresponding label in these 200 class labels.
[0201] The following combines Figure 9 to introduce an exemplary implementation manner of obtaining the target confidence and updating the first object recognition model according to the target confidence. As Figure 9 shown, the method may include S910 to S940.
[0202] S910, determine that the first label is the first class label among the at least one class of labels according to the similarity between the first label and each class label in the at least one class of labels, where the first label is used to represent the first category of the target object.
[0203] In an exemplary method, the similarity between the first label and each class label in the at least one class of labels may be obtained, and the class label with the largest similarity may be determined as the first class label.
[0204] When obtaining the similarity between the first label and each class label in the at least one class of labels, in an exemplary implementation manner, the similarity between the first label and the central label of each class label may be obtained, and this similarity may be used as the similarity between the first label and each class label.
[0205] The similarity between the first label and the central label of each class label may be measured by the distance between the semantic feature of the first label and the semantic feature of the central label of each class label. Among them, the smaller the distance, the greater the similarity.
[0206] One calculation method for the distance between the semantic feature of the first label and the semantic feature of the central label of each class label is as follows: extract the semantic feature of the first label, extract the semantic feature of the central label, and calculate the distance between the semantic feature of the first label and the semantic feature of the central label.
[0207] For example, using the BERT model to extract the semantic feature of the first label, a feature vector can be obtained; using the BERT model to extract the semantic feature of the central label, another feature vector can be obtained; calculate the distance between these two feature vectors, such as cosine distance, Euclidean distance, etc.
[0208] When obtaining the similarity between the first label and each type of label in the at least one type of label, in another exemplary implementation, the similarity between the first label and each label in each type of label may be obtained, and the average similarity may be calculated, and the average similarity may be used as the similarity between the first label and each type of label.
[0209] S920, use the second object recognition model to infer the category of the target object, and obtain a first probability that the category label of the target object is the first type of label.
[0210] For example, input the target image into the second object recognition model. After the second object recognition model makes an inference, it outputs the label of the category to which the target object belongs and the probability that the target object belongs to each category therein, and these labels include the first type of label. That is to say, the second object recognition model can output a first probability that the category label of the target object is the first type of label.
[0211] S930, determine a target confidence level that the first category is the true category of the target object according to the first probability.
[0212] That is to say, according to the first probability that the category label of the target object is the first type of label, determine a target confidence level that the first label indicated by the user is the true category label of the target object.
[0213] For example, the first probability may be used as the target confidence level. Of course, other operations may also be performed according to the first probability to obtain the target confidence level, and this embodiment does not limit this.
[0214] S940, when the target confidence level is greater than or equal to the confidence level threshold, add the features of the target object and the first label to the first object recognition model.
[0215] In this embodiment, it is determined which type of label the first label specified by the user belongs to, and the probability that the trained classification model infers that the target object belongs to the category identified by this type of label is used, and the confidence level that the first label specified by the user is the true category label of the target object is determined according to this probability. And when this confidence level exceeds the preset confidence level threshold, it is considered that the first label specified by the user is reliable, and then the first object recognition model is updated. This can improve the recognition accuracy of the updated first object recognition model.
[0216] Such as Figure 10As shown, the first label is "Russian Blue Cat"; the BERT model is used to extract semantic features of the first label, obtaining the first semantic feature; according to this first semantic feature, it is determined that "Russian Blue Cat" belongs to the "cat" class label, that is, the first class label is the "cat" class label; the target object in the target image is a dog; the classification model is used to infer the target image, and it is known that the first probability that the label of the target object in the target image is the "cat" class label is 0.04, and this first probability is used as the target confidence; since the target confidence 0.04 is less than the preset confidence threshold 0.06, therefore, the first object recognition model is not updated.
[0217] As can be seen from Figure 10 the example shown, in the present application, the first object recognition model is updated according to the target confidence, which can avoid updating the wrong label to the first object recognition model in the case where the user indicates the wrong label for the target object, thereby avoiding reducing the misrecognition of the first object recognition model.
[0218] It can be understood that Figure 9 the first label in the method shown is not limited to being indicated by the user through the first voice message, and it can be indicated by the user in any way. For example, it can be indicated by the user through a text message.
[0219] In some possible scenarios, the target image may include multiple objects, and the user is only interested in one of them, that is, the object that the user hopes to be recognized is one of these multiple objects, or rather, the first label indicated by the user through voice is the expected label of one of these multiple objects, and the object that the user is interested in is the target object.
[0220] In this scenario, it is necessary to determine which object in the target image is the target object expected by the user, so as to correctly know which target object the first label indicated by the user corresponds to, thereby helping to improve the recognition accuracy of the updated first object recognition model.
[0221] For the above scenario, the present application further proposes a method for determining the target object from the target image. In one implementation manner of determining the target object from the target image, one or more object categories can be specified in advance, and it is stipulated that the object pointed to by the object of this one or more categories in the target image is the target object, that is, the object located in the direction indicated by this one or more categories is the target object. In the embodiments of the present application, for the sake of description convenience, these pre-specified objects are referred to as the first objects.
[0222] The pre-specified objects described herein may include hands and / or eyeballs. Of course, it can also be other categories, and the present application does not limit this.
[0223] The following describes how to determine a target object from a target image based on a first object.
[0224] In one implementation of determining a target object from a target image based on a first object, the following steps 1 to 5 may be included.
[0225] Step 1: Perform target detection on the target image to obtain the position of the first object and the bounding box of the first object.
[0226] Taking the first object as a hand for example, the single shot multi-box detector (SSD) can be used to detect the position where the hand is located and the bounding box of the hand.
[0227] Taking the first object as an eyeball for example, the SSD detection method can be used to detect the position where the face is located and the bounding box of the face.
[0228] Step 2: Determine the direction indicated by the first object according to the image in the bounding box.
[0229] Taking the first object as a hand for example, a trained classification model can be used to classify the hand within the bounding box, so as to obtain the direction indicated by the hand. For example, finger images can be divided into 36 categories, and the directions indicated by adjacent two categories of finger images are spaced 10 degrees apart; correspondingly, hand images are also divided into 36 categories, and each category of finger images corresponds to a direction indicated by the hand.
[0230] Taking the first object as an eyeball for example, a trained classification model can be used to classify the eyeballs within the face bounding box, so as to obtain the direction indicated by the eyeballs. Eyeball images can be divided into 36 categories, and the directions indicated by adjacent two categories of eyeball images are spaced 10 degrees apart.
[0231] Step 3: Perform visual saliency detection on the target image to obtain multiple salient regions in the target image.
[0232] One implementation of obtaining the salient regions in the target image may include: calculating the saliency probability map of the target image; inputting the saliency probability map into the main region candidate box generation model to obtain the main region candidate box; segmenting the image saliency probability through the main region into the main pixel set and the non-main pixel set; respectively calculating the average saliency probability value of the main pixel set and the average saliency probability value of the non-main pixel set, and finding the ratio of these two probability values, and taking this ratio as the saliency score of the main region; the main region with a saliency score greater than the pre-configured score threshold is the aforementioned salient region, and if there are multiple main regions with saliency scores greater than the score threshold, multiple salient regions are obtained.
[0233] In some implementations, the target image is input into a saliency probability map generation model, and a saliency probability map of the target image is generated based on the output of the saliency probability map generation model. The saliency probability map of the target image may include probability values corresponding one-to-one to the pixel values in the target image, and each probability value represents the saliency probability of the position where its corresponding pixel value is located. Among them, the candidate bounding box generation module for the main region can be trained using a saliency detection dataset.
[0234] As an example, the saliency probability map generation model can be a binary classification segmentation model. After the target image is input into the binary classification segmentation model, the segmentation model can segment two types of things in the target image, one is the salient object and the other is the background. And the segmented image can output the probability that each pixel in the target image belongs to the corresponding category, and these probabilities constitute the saliency probability map.
[0235] In some implementations, the candidate bounding box generation model for the main region can use methods such as selective search or connected region analysis to obtain the candidate bounding box of the main region of the target image.
[0236] Step four, according to the direction indicated by the first object and the position of the bounding box, determine the target salient region from the multiple salient regions. The target salient region is the salient region among the multiple salient regions that is closest to the bounding box in the direction indicated by the first object.
[0237] That is to say, the salient region closest to the first object in the direction indicated by the first object is determined as the target salient region.
[0238] Taking the first object including a hand as an example, obtain the salient regions in the direction indicated by the hand, calculate the distances between these salient regions and the hand, and finally determine the salient region with the smallest distance as the target salient region.
[0239] As Figure 11 shown, the salient region 1 in the direction indicated by the finger is closer to the finger than the salient region 2 in the direction indicated by the finger. Therefore, the salient region 1 is the target salient region.
[0240] Taking the first object including an eyeball as an example, obtain the salient regions in the direction indicated by the eyeball, calculate the distances between these salient regions and the eyeball, and finally determine the salient region with the smallest distance as the target salient region. The object in the target salient region is the target object.
[0241] Step five, update the first object recognition model according to the target salient region.
[0242] For example, obtain the features of the object in the target salient region and add the features and the first label to the first object recognition model.
[0243] Figure 12 This is a schematic flowchart for updating an object recognition model based on a user instruction in this application.
[0244] S1201, Receive a user instruction and obtain the indication information "This is A" in the user instruction, where A is the first label.
[0245] S1202, Obtain a target image, perform multi-subject saliency detection on the target image, and obtain multiple saliency regions.
[0246] S1203, Based on the direction indicated by the user (the direction indicated by a gesture or the direction indicated by eye movement), obtain the target saliency region located in the direction indicated by the user from the multiple saliency regions.
[0247] S1204, Judge the confidence level that the first label A is the class label of the object in the target saliency region.
[0248] Use the 1000-class labels corresponding to the ImageNet dataset, extract features for these 1000-class labels using the BERT model; use K-means to cluster the 1000 classes into 200 classes and generate 200 cluster centers; map the 1000-class object labels to 200 superclasses; use the ImageNet dataset and the corresponding 200-class superclass labels to train a 200-class classification model using mobilenetv2; extract features for label A using BERT; calculate the distance between the BERT features of label A and the BERT feature centers corresponding to the 200 superclasses, and select the superclass H corresponding to the closest distance as the superclass of the label; input the saliency region into the mobilenetv2 model to generate the probability that this region belongs to superclass H, and use it as the confidence level of the user feedback.
[0249] S1205, If the confidence level is greater than the threshold, update the model.
[0250] For example, if the confidence level is less than the threshold, do not perform model update; if the confidence level is greater than the threshold, update the model.
[0251] Figure 13 This is a schematic flowchart for correcting an object recognition model based on a user instruction in this application
[0252] S1301, Perform multi-subject saliency detection on the target image to obtain multiple saliency regions.
[0253] S1302, For the OpenImage dataset, change the labels corresponding to the target boxes in the images to "object".
[0254] S1303. Train the Fast RCNN model using the modified OpenImages dataset.
[0255] S1304. Input the target image into the Fast RCNN model to generate N saliency regions, where N is a positive integer.
[0256] S1305. Obtain the saliency region indicated by the user based on the direction indicated by the user in the target image (the direction indicated by the gesture or the direction indicated by the eyeball).
[0257] For example, the following steps can be performed to obtain the saliency region indicated by the user: a) Train the Fast RCNN model using the hand dataset to obtain the hand detection Fast RCNN model; b) Train the finger direction classification model using the finger direction dataset. The finger direction dataset annotates the direction indicated by the finger, with each class interval of 10 degrees, for a total of 36 classes; c) Input the target image into the hand detection Fast RCNN model to obtain the hand position region; d) Input the hand position region into the finger direction classification model to obtain the finger direction; e) Obtain the nearest main region in the gesture direction from the N saliency regions and calculate the distance d1 between the nearest main region and the finger; f) Train the SSD model using the face detection dataset; g) Train the eyeball direction classification model using the eyeball direction dataset. The eyeball direction dataset annotates the direction indicated by the eyeball, with each class interval of 10 degrees, for a total of 36 classes; h) Input the target image into the face detection SSD model to obtain the face position region; i) Input the face position region into the eyeball direction classification model to obtain the direction indicated by the eyeball; j) Obtain the nearest main region in the direction indicated by the eyeball and calculate the distance d2 between the main region and the eyeball; k) If d1 is less than d2, use the nearest main region in the direction indicated by the gesture as the main region pointed to by the user; if d1 is greater than d2, use the nearest main region in the direction indicated by the eyeball as the main region pointed to by the user.
[0258] S1306. Perform class recognition on the saliency region pointed to by the user to obtain label A * .
[0259] S1307. Collect the user instruction and obtain the content "This is A" in the user instruction. If A and A * are inconsistent, judge the confidence level of label A. If the confidence level is greater than or equal to the threshold, update the model.
[0260] In an object recognition method in this application, after obtaining the to-be-recognized image collected by the imaging device, the first object recognition model updated by any of the above methods can be used for recognition. This object recognition method can be executed by the aforementioned execution device or local device.
[0261] For example, after adding the image features of "Doraemon" and the first label containing the semantic features of "Doraemon", as well as the corresponding relationship between the image features and the first label, to the first object recognition model, when using the first object recognition model for object recognition, if the imaging device captures an image containing "Doraemon", the first object recognition model can first extract the features of the image and calculate the similarity between the features and the features in the feature library of the first object recognition model, so as to determine that the category of the image is "Doraemon".
[0262] Figure 14 FIG. is an exemplary structural diagram of an apparatus for updating an object recognition model according to the present application. The apparatus 1400 includes an acquisition module 1410 and an update module 1420. In some implementation manners, the apparatus 1400 may be the foregoing execution device or local device.
[0263] The apparatus 1400 may implement any of the foregoing methods. For example, the acquisition module 1410 is used to execute S710 and S720, and the update module 1420 is used to execute S730.
[0264] The present application further provides an apparatus 1500 as shown in Figure 15 FIG. Figure 15 . The apparatus 1500 includes a processor 1502, a communication interface 1503, and a memory 1504. An example of the apparatus 1500 is a chip. Another example of the apparatus 1500 is a computing device. Another example of the apparatus 1500 is a server.
[0265] The processor 1502, the memory 1504, and the communication interface 1503 may communicate through a bus. Executable code is stored in the memory 1504, and the processor 1502 reads the executable code in the memory 704 to execute the corresponding method. Other software modules required for other running processes, such as an operating system, may also be included in the memory 1504. The operating system may be LINUX TM , UNIX TM , WINDOWS TM and the like.
[0266] For example, the executable code in the memory 1504 is used for any of the foregoing methods, and the processor 1502 reads the executable code in the memory 1504 to execute any of the foregoing methods.
[0267] Among them, the processor 1502 may be a central processing unit (CPU). The memory 1504 may include volatile memory, such as random access memory (RAM). The memory 1504 may also include non-volatile memory (2NVM), such as read-only memory (2ROM), flash memory, hard disk drive (HDD), or solid state disk (SSD).
[0268] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0269] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0270] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0271] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0272] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, can exist physically alone for each unit, or two or more units can be integrated into one unit.
[0273] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory, magnetic disks, or optical discs that can store program codes.
[0274] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for updating an object recognition model, characterized in that Including: Obtaining a target image collected by an imaging device; Obtaining first voice information collected by a voice device, where the first voice information is used to indicate a first category of a target object in the target image; Updating a first object recognition model according to the target image and the first voice information. In the updated first object recognition model, features of the target object and a first label are included, and there is a corresponding relationship between the features of the target object and the first label, where the first label is used to represent the first category; The updating the first object recognition model according to the target image and the first voice information includes: Determining that the first label is a first type of label among at least one type of label according to the similarity between the first label and each type of label in the at least one type of label, where the similarity between the first label and the first type of label is greater than the similarity between the first label and other types of labels in the at least one type of label, and the at least one type of label is determined according to the category labels in the first object recognition model; Using a second object recognition model to determine a first probability that the category label of the target object is the first type of label according to the target image; When the first probability is greater than or equal to a preset probability threshold, adding the features of the target object and the first label to the first object recognition model.
2. The method according to claim 1, characterized in that, The determining that the first label is a first type of label among at least one type of label according to the similarity between the first label and each type of label in the at least one type of label includes: Determining that the first label is a first type of label among at least one type of label according to the similarity between the semantic features of the first label and the semantic features of each type of label in the at least one type of label; Where the similarity between the first label and the first type of label being greater than the similarity between the first label and other types of labels in the at least one type of label includes: The distance between the semantic features of the first label and the semantic features of the first type of label is less than the distance between the semantic features of the first label and the semantic features of the other types of labels.
3. The method according to claim 1 or 2, characterized in that, The target image includes a first object, and the target object is the object in the target image that is located in the direction indicated by the first object and is closest to the first object. The first object includes an eyeball or a finger.
4. The method according to claim 3, wherein The updating the first object recognition model according to the target image and the first voice information includes: Determining a bounding box of the first object in the target image; Determining the direction indicated by the first object according to the image in the bounding box; Performing visual saliency detection on the target image to obtain a plurality of salient regions in the target image; Determining a target salient region from the plurality of salient regions according to the direction indicated by the first object, where the target salient region is the salient region among the plurality of salient regions that is closest to the bounding box of the first object in the direction indicated by the first object; Updating the first object recognition model according to the target salient region, where the object in the target salient region includes the target object.
5. The method according to claim 4, wherein Determining the direction indicated by the first object based on the image in the bounding box includes: Using a classification model to classify the image in the bounding box to obtain the target category of the first object; Determining the direction indicated by the first object according to the target category of the first object.
6. An apparatus for updating an object recognition model, characterized in that, Includes: An acquisition module for acquiring a target image collected by an imaging device; The acquisition module is further configured to acquire first voice information collected by a semantic device, where the first voice information is used to indicate a first category of a target object in the target image; An update module for updating a first object recognition model according to the target image and the first voice information. After the update, the first object recognition model includes features of the target object and a first label, and there is a corresponding relationship between the features of the target object and the first label, and the first label is used to represent the first category; Specifically, the update module is configured to: Determine that the first label is a first type of label among the at least one type of label according to the similarity between the first label and each type of label in the at least one type of label, where the similarity between the first label and the first type of label is greater than the similarity between the first label and other types of labels in the at least one type of label, and the at least one type of label is determined according to the category labels in the first object recognition model; Using a second object recognition model to determine a first probability that the category label of the target object is the first type of label according to the target image; When the first probability is greater than or equal to a preset probability threshold, adding the features of the target object, the first label, and the corresponding relationship between the features and the first label to the first object recognition model.
7. The device according to claim 6, characterized in that, Specifically, the update module is configured to: Determine that the first label is a first type of label among the at least one type of label according to the similarity between the semantic features of the first label and the semantic features of each type of label in the at least one type of label; Among them, the similarity between the first label and the first type of label being greater than the similarity between the first label and other types of labels in the at least one type of label includes: The distance between the semantic features of the first label and the semantic features of the first type of label is less than the distance between the semantic features of the first label and the semantic features of the other types of labels.
8. The device according to claim 6 or 7, characterized in that, The target image includes a first object, and the target object is the object in the target image that is located in the direction indicated by the first object and is closest to the first object, and the first object includes an eyeball or a finger.
9. The device according to claim 8, wherein Specifically, the update module is configured to: Determine a bounding box of the first object in the target image; Determine the direction indicated by the first object according to the image in the bounding box; Performing visual saliency detection on the target image to obtain multiple salient regions in the target image; Determining a target salient region from the multiple salient regions according to the direction indicated by the first object, where the target salient region is the salient region among the multiple salient regions that is closest to the bounding box of the first object in the direction indicated by the first object; Update the first object recognition model according to the target salient region, wherein the objects within the target salient region include the target object.
10. The device according to claim 9, characterized in that, The updating module is specifically configured to: Use a classification model to classify the image in the bounding box to obtain the target category of the first object; Determine the direction indicated by the first object according to the target category of the first object.
11. An apparatus for updating an object recognition model, characterized in that, Comprising: A processor, the processor being coupled to a memory; The memory is used for storing instructions; The processor is configured to execute the instructions stored in the memory so that the device executes the method according to any one of claims 1 to 5.
12. A computer-readable medium, characterized in that, Comprising instructions, when the instructions are run on a processor, causing the processor to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method for manipulating image based on tracked eye movements
CN101943982A
Image annotating method, apparatus and electronic device
CN107223246A
Display object determination method and device based on image content, medium and equipment
CN108288208A
N image processing system based on artificial intelligence and computer equipment
CN109800805A