Object recognition method and device, storage medium and electronic device
By using twin neural network architecture in the image recognition model to train the Regnet model and adjust the shared weight, the hardware dependence and economic cost problems caused by the complex structure of the image recognition model in the prior art are solved, and the goals of high recognition accuracy and low training cost are achieved.
Patent Information
- Application Number
- CN202311794775.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
The image recognition model in the prior art has complex structure, which leads to an increase in dependence on hardware and an increase in economic costs, making it difficult to achieve the goal of high recognition accuracy.
The Regnet model is trained based on the twin neural network architecture, and the shared weight is adjusted by adjusting the loss value of the cross entropy loss function and the comparison loss function to achieve optimization of the image recognition model.
The cost of model training is reduced, the recognition accuracy of the image recognition model is improved, and the feature learning ability of the network is improved without increasing the number of parameters.
Smart Images

Figure CN120198780A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of smart homes, and in particular, to an object recognition method, device, storage medium, and electronic device. Background Art
[0002] With the progress of science and technology and the development of artificial intelligence, various algorithms are increasingly applied to people's daily lives. Especially for household appliances equipped with cameras, as one of the daily household appliances with high usage frequency, its intelligent development is crucial. The most critical purpose of the development of intelligent technology is to provide convenience for daily life, facilitating users to quickly recognize faces, actions, and objects through cameras. For example, recognizing the user's outfits and providing guiding suggestions and management for related purchases, etc.
[0003] However, most of the image recognition models in related technologies have complex structures, and complex network structures will increase the dependence on hardware and economic costs. Summary of the Invention
[0004] The purpose of the present application is to provide an object recognition method, device, storage medium, and electronic device, which are used to train a model with a relatively low training cost, and the obtained image recognition model can also have a relatively high recognition accuracy.
[0005] The present application provides an object recognition method, including:
[0006] Obtain a target image including a target object to be recognized, and input the target image into a target recognition model to obtain a prediction result of the target object; based on the target probability value indicated by the prediction result, determine whether the object type of the target object is the same as a preset type; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
[0007] Optionally, the determining whether the object type of the target object is the same as the preset type based on the probability value indicated by the prediction result includes: when the target probability value is greater than a preset probability threshold, determining that the object type of the target object is the same as the preset type; or, when the target probability value is less than or equal to the preset probability threshold, determining that the object type of the target object is different from the preset type.
[0008] Optionally, before inputting the target object into the target recognition model to obtain the prediction result of the target object, the method further includes: obtaining a first image and a second image to be input; the first image includes a first object; the second image includes a second object; generating a label value based on a comparison result between the object type of the first object and the object type of the second object; wherein, when the object type of the first object is the same as the object type of the second object, the label value is a first preset value; when the object type of the first object is different from the object type of the second object, the label value is a second preset value.
[0009] Optionally, after generating the label value based on the comparison result between the object type of the first object and the object type of the second object, the method further includes: respectively inputting the first image and the second image into a first Regnet model and a second Regnet model in a siamese neural network architecture, and obtaining a first prediction value output by the first Regnet model and a second prediction value output by the second Regnet model; the model weights of the first Regnet model and the second Regnet model are both target weights; calculating a cross-entropy loss value based on the first prediction value and the second prediction value, and calculating a contrast loss value based on the first prediction value, the second prediction value, and the label value; adjusting the target weight based on the sum of the cross-entropy loss value and the contrast loss value.
[0010] Optionally, adjusting the target weight based on the sum of the cross-entropy loss value and the contrast loss value includes: when the label value is the first preset value, adjusting the target weight to reduce the difference between the first prediction value and the second prediction value; or, when the label value is the second preset value and the target difference is greater than or equal to a preset adjustment threshold, keeping the target weight unchanged; or, when the label value is the second preset value and the target difference is less than the preset adjustment threshold, adjusting the target weight to increase the difference between the first prediction value and the second prediction value; wherein, the target difference is the difference between the first prediction value and the second prediction value.
[0011] Optionally, calculating the contrast loss value based on the first prediction value and the second prediction value includes:
[0012] Calculating the contrast loss value based on the following formula one:
[0013]
[0014] Wherein, W is the target weight; Y is the label value; X1 is the first prediction value; X2 is the second prediction value; when Y = 0, it indicates that the label value is the first preset value; when Y = 1, it indicates that the label value is the second preset value; D w = ||X1 - X2|| 2 ; m is the target difference.
[0015] Optionally, after adjusting the target weight based on the sum of the cross-entropy loss value and the contrast loss value, the method further includes: adjusting the initialized Regnet model based on the adjusted target weight to obtain the target recognition model.
[0016] This application also provides an object recognition device, including:
[0017] An acquisition module, configured to acquire a target image including a target object to be recognized; a recognition module, configured to input the target image into a target recognition model to obtain a prediction result of the target object; a judgment module, configured to judge whether the object type of the target object is the same as a preset type based on the target probability value indicated by the prediction result; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and the shared weight of the siamese Regnet model is adjusted based on a target loss value during the training process; the target loss value is determined based on the loss value of the cross-entropy loss function and the loss value of the contrast loss function.
[0018] Optionally, the judgment module is specifically configured to determine that the object type of the target object is the same as the preset type when the target probability value is greater than a preset probability threshold; the judgment module is specifically further configured to determine that the object type of the target object is different from the preset type when the target probability value is less than or equal to the preset probability threshold.
[0019] Optionally, the device further includes: a generation module; the acquisition module is further configured to acquire a first image and a second image to be input; the first image includes a first object; the second image includes a second object; the generation module is configured to generate a label value based on a comparison result of the object type of the first object and the object type of the second object; wherein, when the object type of the first object is the same as the object type of the second object, the label value is the first preset value; when the object type of the first object is different from the object type of the second object, the label value is the second preset value.
[0020] Optionally, the device further includes: a training module; the training module is configured to input the first image and the second image into a first Regnet model and a second Regnet model in a siamese neural network architecture respectively, and obtain a first predicted value output by the first Regnet model and a second predicted value output by the second Regnet model; the model weights of the first Regnet model and the second Regnet model are both target weights; the training module is further configured to calculate a cross-entropy loss value based on the first predicted value and the second predicted value, and calculate a contrast loss value based on the first predicted value, the second predicted value and the label value; the training module is further configured to adjust the target weight based on the sum of the cross-entropy loss value and the contrast loss value.
[0021] Optionally, the training module is specifically configured to adjust the target weight to reduce the difference between the first predicted value and the second predicted value when the label value is the first preset value; the training module is further specifically configured to keep the target weight unchanged when the label value is the second preset value and the target difference is greater than or equal to a preset adjustment threshold; the training module is further specifically configured to adjust the target weight to increase the difference between the first predicted value and the second predicted value when the label value is the second preset value and the target difference is less than the preset adjustment threshold; wherein, the target difference is the difference between the first predicted value and the second predicted value.
[0022] Optionally, the training module is specifically configured to calculate the contrast loss value based on the following formula one:
[0023]
[0024] wherein, W is the target weight; Y is the label value; X1 is the first predicted value; X2 is the second predicted value; when Y = 0, it means the label value is the first preset value; when Y = 1, it means the label value is the second preset value; D w = ||X1 - X2|| 2 ; m is the target difference.
[0025] Optionally, the generating module is further configured to adjust the initialized Regnet model based on the adjusted target weight to obtain the target recognition model.
[0026] The present application also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the steps of implementing the object recognition method as described in any one of the above through the computer program.
[0027] The present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, and when the program runs, it implements the steps of any one of the above object recognition methods.
[0028] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the above object recognition methods.
[0029] For the object recognition method, device, storage medium, and electronic device provided by the present application, first, a target image including a target object to be recognized is obtained, and the target image is input into a target recognition model to obtain a prediction result of the target object; then, based on the target probability value indicated by the prediction result, it is determined whether the object type of the target object is the same as a preset type; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function. In this way, the model can be trained at a relatively low training cost, and the obtained image recognition model can also have a relatively high recognition accuracy. Description of the Drawings
[0030] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0031] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0032] Figure 1 It is a schematic diagram of the hardware environment of an interaction method of an intelligent device according to an embodiment of the present application;
[0033] Figure 2 It is a schematic flowchart of the object recognition method provided by the present application;
[0034] Figure 3 It is a schematic flowchart of the model training provided by the present application;
[0035] Figure 4 It is a schematic structural diagram of the object recognition device provided by the present application;
[0036] Figure 5It is a schematic structural diagram of the electronic device provided by this application. Detailed implementation manners
[0037] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0038] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0039] According to one aspect of the embodiments of this application, an object recognition method is provided. This object recognition method is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and IntelligenceHouse ecosystem. Optionally, in this embodiment, the above object recognition method can be applied to, for example Figure 1 the hardware environment composed of a terminal device 102 and a server 104 as shown. As Figure 1 shown, the server 104 is connected to the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.
[0040] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart sound box, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart floor sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steam box, a smart microwave oven, a smart kitchen water heater, a smart purifier, a smart water dispenser, a smart door lock, etc.
[0041] The following describes the professional terms involved in the embodiments of the present application:
[0042] RegNet model: It is a neural network model based on differentiable architecture search. It can automatically generate the optimal network structure and parameters according to different tasks and hardware platforms. The RegNet model is characterized by being able to flexibly adjust the depth, width, resolution, and grouping of the network to achieve the best performance and efficiency. The 2RegNet model has achieved excellent results in multiple computer vision tasks, such as image classification, object detection, semantic segmentation, etc.
[0043] Cross-entropy loss function: It is a commonly used loss function for classification problems. It measures the difference between the predicted probability distribution of the model and the probability distribution of the true labels. The smaller the difference, the smaller the loss, and the more accurate the model.
[0044] Embedding layer: It is a method of converting discrete data (such as words or categories) into continuous vector representations. It can capture the semantic relationships and similarities between data, reduce the dimension of the data, and improve the efficiency and performance of the model.
[0045] Backpropagation algorithm: It is a method for optimizing the parameters of a neural network model. Its principle is to take the derivative of each parameter according to the loss function, and then update the parameters in the opposite direction of the gradient to minimize the loss function.
[0046] In view of the above technical problems existing in the related art, the embodiments of the present application provide a method for identifying similar categories based on a twin Regnet network structure framework and a regular term of a contrastive loss function, which is used to improve the network selection effect of the original Regnet model. Without increasing the original network parameters, the improved network optimizes the design space according to the improved loss function, searches for and selects a better network structure, and better identifies the target in the image.
[0047] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the object recognition method provided by the embodiments of the present application.
[0048] As Figure 2 shown, an object recognition method provided by the embodiments of the present application may include the following steps 201 and 202:
[0049] Step 201: Obtain a target image containing a target object to be recognized, and input the target image into a target recognition model to obtain a prediction result of the target object.
[0050] Among them, the target recognition model is obtained by training the Regnet model based on a twin neural network architecture, and during the training process, the shared weights of the twin Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
[0051] Exemplarily, the above target image contains a target object to be recognized, and the above target recognition model is used to classify the target object and output a binary classification result of the target object, that is, whether the object type of the target object is a preset type.
[0052] For example, taking the above target object as a skirt, the above target recognition model is used to identify whether the target object is a skirt and outputs a probability value that the target object is a skirt.
[0053] Exemplarily, the above target recognition model is obtained by training the Regnet model based on a twin neural network architecture, and during the training process, the shared weights of the twin Regnet model are adjusted based on the sum of the loss values of the cross-entropy loss function and the contrastive loss function.
[0054] It should be noted that since most of the image recognition models in the related art have complex structures and a large number of parameters, they are not conducive to application in actual scenarios. Based on this, the embodiments of the present application are based on the original lightweight network structure, and by changing the learning rules, relevant parameter adjustments are completed, so that the network can obtain better learning parameters and have a better feature learning effect.
[0055] Step 202: Based on the target probability value indicated by the prediction result, determine whether the object type of the target object is the same as the preset type.
[0056] Exemplarily, after obtaining the target probability value output by the target recognition model, it is possible to determine whether the object type of the target object is the same as the preset type based on this target probability value.
[0057] Specifically, the above step 202 may further include the following step 202a or step 202b:
[0058] Step 202a: When the target probability value is greater than the preset probability threshold, determine that the object type of the target object is the same as the preset type.
[0059] Step 202b: When the target probability value is less than or equal to the preset probability threshold, determine that the object type of the target object is different from the preset type.
[0060] Exemplarily, in the embodiments of the present application, it is possible to determine that the object type of the target object is the same as the preset type by setting a probability threshold.
[0061] Optionally, in the embodiments of the present application, the Regnet model can be trained based on the training method of the siamese neural network architecture.
[0062] Exemplarily, before the above step 201, the object recognition method provided in the embodiments of the present application may further include the following steps 203 and 204:
[0063] Step 203: Obtain a first image and a second image to be input.
[0064] Exemplarily, the first image includes a first object; the second image includes a second object. The above first image and second image are two randomly selected images from the sample set.
[0065] Step 204: Generate a label value based on the comparison result between the object type of the first object and the object type of the second object.
[0066] Wherein, when the object type of the first object is the same as the object type of the second object, the label value is a first preset value; when the object type of the first object is different from the object type of the second object, the label value is a second preset value.
[0067] Exemplarily, after obtaining the above first image and second image, it is also necessary to generate a label value according to the object types of the objects included in the first image and the second image. The label value can be manually labeled or automatically generated according to the object type markers carried by each image.
[0068] For example, when the object type of the first object is the same as the object type of the second object, the label value can be 0, that is, the above first preset value; when the object type of the first object is different from the object type of the second object, the label value is 1, that is, the above second preset value.
[0069] It can be understood that based on the above label value, the prediction result of the model can be judged, and then the weight of the model can be adjusted.
[0070] Exemplarily, after the above step 204, the object recognition method provided by the embodiments of the present application may further include the following steps 205 to step 207:
[0071] Step 205: Input the first image and the second image into the first Regnet model and the second Regnet model in the siamese neural network architecture respectively, and obtain the first prediction value output by the first Regnet model and the second prediction value output by the second Regnet model.
[0072] Wherein, the model weights of the first Regnet model and the second Regnet model are both target weights.
[0073] Exemplarily, in the embodiments of the present application, an initialized Regnet_x_400mf model, that is, the above first Regnet model and second Regnet model, can be constructed through the Torchvision library and a pre-trained model that has been trained by IMAGNET. And select the model weight at this time as the initial shared weight.
[0074] Exemplarily, Regnet_x_400mf is obtained by optimizing the design space of the Regnet model. By directly introducing the Regnet_x_400mf in the torchvision dataset and the weight file of Regnet_x_400mf obtained after pre-training on the Imagenet dataset, the initial basic weight network Regnet_x_400mf in this article is constructed.
[0075] Step 206: Calculate the cross-entropy loss value based on the first prediction value and the second prediction value, and calculate the contrast loss value based on the first prediction value, the second prediction value and the label value.
[0076] Step 207: Adjust the target weight based on the sum of the cross-entropy loss value and the contrastive loss value.
[0077] Exemplarily, as Figure 3 shown, it is a schematic diagram of the model training process provided by an embodiment of the present application. First, randomly select two images from the same sample set, including Image 1 and Image 2; Second, perform an image preprocessing method with a unified shape and size on the images with different pixels; Third, Image 1 passes through Regnet Model 1 to obtain Prediction Result 1, and Image 2 passes through Regnet Model 2 to obtain Prediction Result 2; Fourth, calculate the cross-entropy loss function according to the prediction results, and introduce the embedding layer to calculate the Contrastive Loss function value. According to the sum of the two obtained loss function values, use the backpropagation algorithm to adjust the shared weight, that is, the above-mentioned target weight; Fifth, if the loss function value is less than the threshold, end this training. If the loss function is greater than or equal to the threshold, loop through the above steps.
[0078] Specifically, the step of calculating the contrastive loss value based on the first prediction value, the second prediction value, and the label value in the above step 206 may further include the following step 206a:
[0079] Step 206a: Calculate the contrastive loss value based on the following formula (1):
[0080]
[0081] where W is the target weight; Y is the label value; X1 is the first prediction value; X2 is the second prediction value; when Y = 0, it means the label value is the first preset value; when Y = 1, it means the label value is the second preset value; D w = ||X1 - X2|| 2 ; m is the target difference.
[0082] Exemplarily, when Y = 0, adjust the parameters to minimize the distance between X1 and X2. When Y = 1, if the distance between X1 and X2 is greater than m, no optimization is performed (time-saving and labor-saving); if the distance between X1 and X2 is less than m, increase the distance between the two to m.
[0083] Specifically, the adjustment of the target weight in the above step 207 may include any one of the following steps 207a1 to 207a3:
[0084] Step 207a1: In the case where the label value is the first preset value, adjust the target weight to reduce the difference between the first prediction value and the second prediction value.
[0085] Step 207a2, when the tag value is the second preset value and the target difference is greater than or equal to the preset adjustment threshold, keep the target weight unchanged.
[0086] Step 207a3, when the tag value is the second preset value and the target difference is less than the preset adjustment threshold, adjust the target weight to increase the difference between the first prediction value and the second prediction value.
[0087] Wherein, the target difference is the difference between the first prediction value and the second prediction value.
[0088] Exemplarily, after completing the training of the model based on the above steps, the final weight can be obtained, that is, the adjusted target weight above. Based on this target weight, a target recognition model can be constructed.
[0089] Exemplarily, after the above step 207, the object recognition method provided by the embodiments of the present application may further include the following step 208:
[0090] Step 208, adjust the initialized Regnet model based on the adjusted target weight to obtain the target recognition model.
[0091] It can be understood that the training process of the model is mainly used to obtain the weight of the model, that is, the adjusted target weight above. After that, based on this weight, it can be combined with the Regnet model to obtain the target recognition model.
[0092] The object recognition method provided by the embodiments of the present application can improve the classification accuracy of the classification method, improve the feature learning ability of the network without increasing the number of parameters, and is more suitable for similar heterogeneous classification tasks.
[0093] The object recognition method provided by the embodiments of the present application, first, obtains a target image including a target object to be recognized, and inputs the target image into the target recognition model to obtain a prediction result of the target object; then, based on the target probability value indicated by the prediction result, determines whether the object type of the target object is the same as the preset type; wherein, the target recognition model is obtained by training the Regnet model based on the twin neural network architecture, and the shared weight of the twin Regnet model is adjusted based on the target loss value during the training process; the target loss value is determined based on the loss value of the cross-entropy loss function and the loss value of the contrast loss function. In this way, the model can be trained at a lower training cost, and the obtained image recognition model can also have a high recognition accuracy.
[0094] It should be noted that for the object recognition method provided in the embodiments of the present application, the execution subject may be an object recognition device, or a control module in the object recognition device for executing the object recognition method. In the embodiments of the present application, taking the object recognition device as an example to execute the object recognition method, the object recognition device provided in the embodiments of the present application will be described.
[0095] It should be noted that in the embodiments of the present application, the object recognition methods shown in the above-mentioned various method drawings are all exemplarily described by taking one of the drawings in the embodiments of the present application as an example. Specifically, when implemented, the object recognition methods shown in the above-mentioned various method drawings can also be implemented in combination with any other drawings that can be combined as shown in the above embodiments, which will not be elaborated here.
[0096] The object recognition device provided by the present application will be described below, and the description below can be mutually corresponded and referred to with the object recognition method described above.
[0097] Figure 4 is a schematic structural diagram of an object recognition device provided in an embodiment of the present application, as Figure 4 shown, and specifically includes:
[0098] An acquisition module 401, configured to acquire a target image including a target object to be recognized; a recognition module 402, configured to input the target image into a target recognition model to obtain a prediction result of the target object; a judgment module 403, configured to judge whether the object type of the target object is the same as a preset type based on the target probability value indicated by the prediction result; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
[0099] Optionally, the judgment module 403 is specifically configured to determine that the object type of the target object is the same as the preset type when the target probability value is greater than a preset probability threshold; the judgment module 403 is specifically further configured to determine that the object type of the target object is different from the preset type when the target probability value is less than or equal to the preset probability threshold.
[0100] Optionally, the device further includes: a generation module; the acquisition module 401 is further configured to acquire a first image and a second image to be input; the first image includes a first object; the second image includes a second object; the generation module is configured to generate a label value based on a comparison result between the object type of the first object and the object type of the second object; wherein, when the object type of the first object is the same as the object type of the second object, the label value is a first preset value; when the object type of the first object is different from the object type of the second object, the label value is a second preset value.
[0101] Optionally, the device further includes: a training module; the training module is configured to respectively input the first image and the second image into a first Regnet model and a second Regnet model in a siamese neural network architecture, and obtain a first predicted value output by the first Regnet model and a second predicted value output by the second Regnet model; the model weights of the first Regnet model and the second Regnet model are both target weights; the training module is further configured to calculate a cross-entropy loss value based on the first predicted value and the second predicted value, and calculate a contrast loss value based on the first predicted value, the second predicted value, and the label value; the training module is further configured to adjust the target weight based on the sum of the cross-entropy loss value and the contrast loss value.
[0102] Optionally, the training module is specifically configured to adjust the target weight to reduce the difference between the first predicted value and the second predicted value when the label value is the first preset value; the training module is specifically further configured to keep the target weight unchanged when the label value is the second preset value and the target difference is greater than or equal to a preset adjustment threshold; the training module is specifically further configured to adjust the target weight to increase the difference between the first predicted value and the second predicted value when the label value is the second preset value and the target difference is less than the preset adjustment threshold; wherein, the target difference is the difference between the first predicted value and the second predicted value.
[0103] Optionally, the training module is specifically configured to calculate the contrast loss value based on the following formula one:
[0104]
[0105] wherein, W is the target weight; Y is the label value; X1 is the first predicted value; X2 is the second predicted value; when Y = 0, it means the label value is the first preset value; when Y = 1, it means the label value is the second preset value; D w= ||X1 - X2|| 2 ; m is the target difference value.
[0106] Optionally, the generating module is further configured to adjust the initialized Regnet model based on the adjusted target weight to obtain the target recognition model.
[0107] The object recognition device provided by this application first obtains a target image containing a target object to be recognized, and inputs the target image into a target recognition model to obtain a prediction result of the target object; then, based on the target probability value indicated by the prediction result, it determines whether the object type of the target object is the same as a preset type; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function. In this way, the model can be trained at a relatively low training cost, and the obtained image recognition model can also have a relatively high recognition accuracy.
[0108] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute an object recognition method, which includes: obtaining a target image containing a target object to be recognized, and inputting the target image into a target recognition model to obtain a prediction result of the target object; based on the target probability value indicated by the prediction result, determining whether the object type of the target object is the same as a preset type; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
[0109] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0110] On the other hand, this application also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the object recognition method provided by the above-mentioned various methods. The method includes: obtaining a target image containing a target object to be recognized, and inputting the target image into a target recognition model to obtain a prediction result of the target object; based on the target probability value indicated by the prediction result, determining whether the object type of the target object is the same as a preset type; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
[0111] On another aspect, this application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it executes the object recognition method provided by the above-mentioned various methods. The method includes: obtaining a target image containing a target object to be recognized, and inputting the target image into a target recognition model to obtain a prediction result of the target object; based on the target probability value indicated by the prediction result, determining whether the object type of the target object is the same as a preset type; wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. An object recognition method, characterized in that, Including: Obtain a target image including a target object to be recognized, and input the target image into a target recognition model to obtain a prediction result of the target object; Based on the target probability value indicated by the prediction result, determine whether the object type of the target object is the same as a preset type; Wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weights of the siamese Regnet model are adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrastive loss function.
2. The object recognition method according to claim 1, characterized in that, The determining whether the type of the target object is the same as the preset type based on the probability value indicated by the prediction result includes: When the target probability value is greater than a preset probability threshold, determine that the object type of the target object is the same as the preset type; Or, When the target probability value is less than or equal to the preset probability threshold, determine that the object type of the target object is different from the preset type.
3. The object recognition method according to claim 1 or 2, characterized in that, Before inputting the target object into the target recognition model to obtain the prediction result of the target object, the method further includes: Obtain a first image and a second image to be input; the first image includes a first object; the second image includes a second object; Generate a label value based on the comparison result of the object type of the first object and the object type of the second object; Wherein, when the object type of the first object is the same as the object type of the second object, the label value is a first preset value; when the object type of the first object is different from the object type of the second object, the label value is a second preset value.
4. The object recognition method according to claim 3, characterized in that, After generating the label value based on the comparison result of the object type of the first object and the object type of the second object, the method further includes: Input the first image and the second image into a first Regnet model and a second Regnet model in a siamese neural network architecture respectively, and obtain a first prediction value output by the first Regnet model and a second prediction value output by the second Regnet model; the model weights of the first Regnet model and the second Regnet model are both target weights; Calculate a cross-entropy loss value based on the first prediction value and the second prediction value, and calculate a contrastive loss value based on the first prediction value, the second prediction value and the label value; Adjust the target weights based on the sum of the cross-entropy loss value and the contrastive loss value.
5. The object recognition method according to claim 4, characterized in that The adjusting the target weights based on the sum of the cross-entropy loss value and the contrastive loss value includes: When the label value is the first preset value, adjust the target weights to reduce the difference between the first prediction value and the second prediction value; Or, When the label value is the second preset value and the target difference is greater than or equal to a preset adjustment threshold, keep the target weights unchanged; Or, When the label value is the second preset value and the target difference is less than the preset adjustment threshold, adjust the target weight to increase the difference between the first prediction value and the second prediction value; Wherein, the target difference is the difference between the first prediction value and the second prediction value.
6. The object recognition method according to claim 5, characterized in that Calculating the contrast loss value based on the first prediction value and the second prediction value includes: Calculating the contrast loss value based on the following formula (1): Where, W is the target weight; Y is the label value; X1 is the first predicted value; X2 is the second predicted value; when Y = 0, it means the label value is the first preset value; when Y = 1, it means the label value is the second preset value; D w = ||X1 - X2|| 2 ; m is the target difference.
7. The object recognition method according to any one of claims 4 to 6, characterized in that After adjusting the target weight based on the sum of the cross-entropy loss value and the contrast loss value, the method further includes: Adjusting the initialized Regnet model based on the adjusted target weight to obtain the target recognition model.
8. An object recognition device, characterized in that, The device includes: An acquisition module, configured to acquire a target image including a target object to be recognized; A recognition module, configured to input the target image into a target recognition model to obtain a prediction result of the target object; A judgment module, configured to judge whether the object type of the target object is the same as a preset type based on the target probability value indicated by the prediction result; Wherein, the target recognition model is obtained by training a Regnet model based on a siamese neural network architecture, and during the training process, the shared weight of the siamese Regnet model is adjusted based on a target loss value; the target loss value is determined based on the loss value of a cross-entropy loss function and the loss value of a contrast loss function.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the object recognition method according to any one of claims 1 to 7.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the object recognition method according to any one of claims 1 to 7 through the computer program.