An object recognition method, device and equipment
By using convolutional neural networks in object recognition technology for feature extraction, the problem of insufficient feature extraction capability in the prior art is solved, and the accuracy of object recognition is improved.
Patent Information
- Application Number
- CN202011558197.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-12-24
AI Technical Summary
In the existing object recognition technology, the ability to extract features is limited and it is impossible to extract features at a higher level, resulting in a low accuracy of object recognition.
Convolutional neural network is used to extract feature of image frames, combining preset segmentation algorithms and machine learning algorithms to achieve deeper abstraction and recognition of object features.
By extracting deeper features, the accuracy of object recognition is improved and the ability to recognize object categories is enhanced.
Smart Images

Figure CN112613508B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of object recognition, and particularly relates to an object recognition method, device and equipment. Background Art
[0002] Currently, traditional object recognition technology mainly obtains a frame of image, then divides the picture into candidate box images of different sizes, then extracts features from the candidate box images to generate high-dimensional vectors, and classifies the candidate box images through machine learning algorithms such as Adaboot and SVM, and finally obtains the category information of the candidate box images. However, in the above method, the feature extraction ability is limited and cannot perform higher-level feature extraction, resulting in a low accuracy rate when performing object recognition. Summary of the Invention
[0003] Embodiments of this application provide an object recognition method, device and equipment, which can solve the problem that in the existing object recognition method, the feature extraction ability is limited and higher-level feature extraction cannot be performed, resulting in a low accuracy rate when performing object recognition.
[0004] In a first aspect, embodiments of this application provide an object recognition method, including:
[0005] Obtain an image to be recognized;
[0006] Perform segmentation processing on the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame;
[0007] Extract features from the image frame through a convolutional neural network to obtain object features corresponding to the image frame;
[0008] Recognize the object features through a machine learning algorithm to obtain the object category of the object included in the image to be recognized.
[0009] Further, the convolutional neural network is a mobilenet-ssd neural network, a yolov4-tiny neural network or a nanodet neural network.
[0010] Further, the preset segmentation algorithm is a brute-force search algorithm, a Selective Search algorithm or an edge detection-based image segmentation algorithm.
[0011] Further, the machine learning algorithm is a support vector machine algorithm or an Adaboot algorithm.
[0012] Further, the training process of the convolutional neural network includes:
[0013] Obtain a preset convolutional neural network and obtain a training sample set; the training sample set includes sample image frames and their corresponding object feature labels;
[0014] Input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object feature corresponding to the sample image frame;
[0015] Obtain the difference information between the sample object feature corresponding to the sample image frame and the object feature label corresponding to the sample image frame according to a preset loss function;
[0016] If the difference information meets the preset training stop condition, stop training and use the current preset convolutional neural network lightweight model as the convolutional neural network for outputting the object feature corresponding to the image frame;
[0017] If the difference information does not meet the first preset training stop condition, adjust the preset parameters and return to input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object feature corresponding to the sample image frame.
[0018] Further, the obtaining of the training sample set includes:
[0019] Obtain a sample image frame and its corresponding object feature label, and obtain the object category of the object in the sample image frame;
[0020] If the number of sample image frames corresponding to the object category is less than a preset number, perform transformation processing on the sample image frame to obtain a transformed sample image frame;
[0021] Determine the training sample set according to the sample image frame and its corresponding object feature label, and the transformed sample image frame and its corresponding object feature label.
[0022] In a second aspect, an embodiment of the present application provides an object recognition device, including:
[0023] An obtaining unit, configured to obtain an image to be recognized;
[0024] A segmentation unit, configured to perform segmentation processing on the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame;
[0025] An extraction unit, configured to extract features of the image frame through a convolutional neural network to obtain the object feature corresponding to the image frame;
[0026] A recognition unit, configured to recognize the object feature through a machine learning algorithm to obtain the object category of the object included in the image to be recognized.
[0027] Further, the convolutional neural network is a mobilenet-ssd neural network, a yolov4-tiny neural network, or a nanodet neural network.
[0028] Further, the preset segmentation algorithm is a brute-force search algorithm, a Selective Search algorithm, or an edge-detection-based image segmentation algorithm.
[0029] Further, the machine learning algorithm is a support vector machine algorithm or an Adaboot algorithm.
[0030] Further, the object recognition device further includes a training unit, specifically for:
[0031] Obtain a preset convolutional neural network and obtain a training sample set; the training sample set includes sample image frames and their corresponding object feature labels;
[0032] Input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object feature corresponding to the sample image frame;
[0033] Obtain the difference information between the sample object feature corresponding to the sample image frame and the object feature label corresponding to the sample image frame according to a preset loss function;
[0034] If the difference information meets the preset training stop condition, stop training and use the current preset convolutional neural network lightweight model as the convolutional neural network for outputting the object feature corresponding to the image frame;
[0035] If the difference information does not meet the first preset training stop condition, adjust the preset parameters and return to input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object feature corresponding to the sample image frame.
[0036] Further, the training unit is specifically for:
[0037] Obtain a sample image frame and its corresponding object feature label, and obtain the object category of the object in the sample image frame;
[0038] If the number of sample image frames corresponding to the object category is less than a preset number, perform transformation processing on the sample image frame to obtain a transformed sample image frame;
[0039] Determine the training sample set according to the sample image frame and its corresponding object feature label, and the transformed sample image frame and its corresponding object feature label.
[0040] In a third aspect, an embodiment of the present application provides an object recognition device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the object recognition method described in the first aspect above is implemented.
[0041] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the object recognition method described in the first aspect above is implemented.
[0042] In an embodiment of the present application, a to-be-recognized image is obtained; the to-be-recognized image is segmented according to a preset segmentation algorithm to obtain at least one image frame; the image frame is subjected to feature extraction through a convolutional neural network to obtain an object feature corresponding to the image frame; and the object feature is recognized through a machine learning algorithm to obtain an object category of an object included in the to-be-recognized image. In the above method, feature extraction is performed on the image frame through a convolutional neural network, and deeper features of the image frame can be extracted, improving the accuracy of object recognition. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0044] Figure 1 is a schematic flowchart of an object recognition method provided in the first embodiment of the present application;
[0045] Figure 2 is a schematic diagram of an object recognition device provided in the second embodiment of the present application;
[0046] Figure 3 is a schematic diagram of an object recognition device provided in the third embodiment of the present application. Detailed Embodiments
[0047] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0048] It should be understood that, as used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0049] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0050] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0051] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0052] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0053] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an object recognition method provided by the first embodiment of this application. In this embodiment, the execution subject of an object recognition method is a device with object recognition function, such as a server, a personal computer, a mobile phone, a robot, etc. As Figure 1 shown, the object recognition method may include:
[0054] S101: Obtain the image to be recognized.
[0055] When the device detects an image recognition instruction, it acquires the image to be recognized. The method for the device to acquire the image to be recognized is not limited here. It can be sent to this device after being captured by a device with a camera function, or this device itself has a capture function to obtain the image to be recognized. For example, a robot equipped with a camera module can directly obtain the image to be recognized through its own photographing function.
[0056] S102: Perform segmentation processing on the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame.
[0057] The preset segmentation algorithm is pre-stored in the device. The device performs segmentation processing on the image to be recognized according to the preset segmentation algorithm to obtain at least one image frame. When the device performs segmentation on the image to be recognized, multiple image frames can be obtained, and the sizes of the multiple image frames can be inconsistent. Each image frame may include one kind of object.
[0058] Among them, the preset segmentation algorithm can be a brute-force search algorithm, a Selective Search algorithm, or an edge detection-based image segmentation algorithm, or other image segmentation algorithms can be adopted, which is not limited here.
[0059] S103: Extract features from the image frame through a convolutional neural network to obtain the object features corresponding to the image frame.
[0060] The device inputs the image frame into the convolutional neural network, and the convolutional neural network extracts features from the image frame to obtain the object features corresponding to the image frame.
[0061] In this embodiment, using a convolutional neural network for feature extraction can more effectively extract features from the image frame, can perform higher-level abstraction on the features, so that subsequent object classification is also more accurate.
[0062] Among them, if this method is applied to a robot embedded platform, in order to meet the real-time requirement of object recognition, some lightweight networks can be adopted for the convolutional neural network. Therefore, the convolutional neural network can be a mobilenet-ssd neural network, a yolov4-tiny neural network, or a nanodet neural network.
[0063] Specifically, the training process of the convolutional neural network can include:
[0064] Acquire a preset convolutional neural network and a training sample set; the training sample set includes sample image frames and their corresponding object feature labels.
[0065] It can be understood that the more abundant the quantity and categories of the sample data in the training sample set are, the higher the accuracy of the trained convolutional neural network will be. For example, the object feature labels of the collected sample image frames can be shoes, socks, electric wires, bar stool bases, pets, feces, etc. When collecting, it is necessary to ensure the diversity and richness of each object feature label, ensure that there are corresponding collected data for each object feature label in different scenarios, and ensure that there are more than 30,000 images for each object feature label.
[0066] Specifically, when obtaining the training sample set, obtain the sample image frames and their corresponding object feature labels, and obtain the object categories of the objects in the sample image frames; the device can perform data preprocessing on the obtained sample image frames, and the data preprocessing can include data cleaning, data augmentation, and data balancing.
[0067] Data cleaning mainly processes the sample image frames with label disorders in each object feature label, and picks them out and places them into the correct object feature labels.
[0068] Data augmentation mainly enhances some object feature labels with relatively small quantities. If the number of sample image frames corresponding to the object category is less than the preset quantity, then perform transformation processing on the sample image frames to obtain transformed sample image frames. The transformation processing is not limited here. For example, the sample image frames can be randomly flipped left and right, randomly cropped, randomly added Gaussian noise, randomly adjusted in brightness, and transformed in color space, etc.
[0069] Data balancing is mainly achieved through data augmentation. Enhance some object feature labels with relatively small quantities, and then ensure that the quantity of each object feature label is at the same order of magnitude finally.
[0070] Determine the training sample set according to the sample image frames and their corresponding object feature labels, the transformed sample image frames and their corresponding object feature labels.
[0071] After the device obtains the training sample set, it starts training. The training process is mainly to first load the sample image frames into the memory, then input them to a preset convolutional neural network, the preset convolutional neural network performs forward inference to calculate the loss, the loss is used to backward update the weights of the preset convolutional neural network, and as the weights are continuously updated, the training loss is also continuously decreasing until the stop training condition is met.
[0072] Specifically, input the sample image frames into a preset lightweight model of the convolutional neural network for processing to obtain the sample object features corresponding to the sample image frames; obtain the difference information between the sample object features corresponding to the sample image frames and the object feature labels corresponding to the sample image frames according to the preset loss function.
[0073] If the difference information meets the preset training stop condition, stop the training, and use the current preset lightweight convolutional neural network model as the convolutional neural network for outputting the object features corresponding to the image frame; if the difference information does not meet the first preset training stop condition, adjust the preset parameters, and return to input the sample image frame into the preset lightweight convolutional neural network model for processing to obtain the sample object features corresponding to the sample image frame.
[0074] In one implementation, a validation set can also be set. During the training process, the accuracy of the validation set is also continuously increasing until the training loss and the accuracy of the validation set reach a stable value, then stop the training. Obtain a sample validation set; the sample validation set includes sample validation pictures and their corresponding sample validation result labels; obtain the recognition result corresponding to the sample validation picture according to the current preset convolutional neural network, and determine the accuracy corresponding to the sample validation set according to the recognition result corresponding to the sample validation picture and the sample validation result label; if the difference information and the accuracy meet the second preset training stop condition, stop the training, and use the current preset lightweight convolutional neural network model as the object recognition model; if the difference information and the accuracy do not meet the second preset training stop condition, adjust the preset parameters, and return to input the sample image into the preset lightweight convolutional neural network model for processing to obtain the sample image recognition result corresponding to the sample image.
[0075] During the training process, the preset parameters to be adjusted can include one or more of the learning rate, the learning rate schedule, the number of the sample images input into the preset lightweight convolutional neural network model each time, and the iteration training period.
[0076] The purpose of adjusting the preset parameters is mainly to find an optimal accuracy on the validation set. After several rounds of fine-tuning of the preset parameters, finally obtain the optimal accuracy on the validation set, that is, the corresponding optimal model. If the accuracy of the validation set meets the requirements, the model can be sent to an embedded platform such as a robot for on-device deployment, otherwise return to re-determine the sample training set to ensure that the accuracy of the validation set meets the requirements.
[0077] S104: Identify the object features through a machine learning algorithm to obtain the object category of the object included in the image to be recognized.
[0078] The device identifies the object features through a machine learning algorithm to obtain the object category of the object included in the image to be recognized. Since the object features are extracted by a convolutional neural network and the extracted object features have a higher level, the object category of the object included in the image to be recognized can be obtained more accurately according to the preset machine learning algorithm.
[0079] Among them, the machine learning algorithm can be a support vector machine algorithm or an Adaboot algorithm.
[0080] In the embodiment of the present application, an image to be recognized is obtained; the image to be recognized is segmented according to a preset segmentation algorithm to obtain at least one image frame; the image frame is subjected to feature extraction through a convolutional neural network to obtain the object feature corresponding to the image frame; and the object feature is recognized through a machine learning algorithm to obtain the object category of the object included in the image to be recognized. In the above method, feature extraction is performed on the image frame through a convolutional neural network, which can extract deeper features of the image frame and improve the accuracy of object recognition.
[0081] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0082] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an object recognition device provided in the second embodiment of the present application. Each unit included is used to execute Figure 1 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figure 1 the relevant descriptions in the corresponding embodiment. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 2 , the object recognition device 2 includes:
[0083] An acquisition unit 210, configured to acquire an image to be recognized;
[0084] A segmentation unit 220, configured to segment the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame;
[0085] An extraction unit 230, configured to perform feature extraction on the image frame through a convolutional neural network to obtain the object feature corresponding to the image frame;
[0086] A recognition unit 240, configured to recognize the object feature through a machine learning algorithm to obtain the object category of the object included in the image to be recognized.
[0087] Further, the convolutional neural network is a mobilenet-ssd neural network, a yolov4-tiny neural network, or a nanodet neural network.
[0088] Further, the preset segmentation algorithm is a brute-force search algorithm, a Selective Search algorithm, or an edge-detection-based image segmentation algorithm.
[0089] Further, the machine learning algorithm is a support vector machine algorithm or an Adaboot algorithm.
[0090] Further, the object recognition device further includes a training unit, specifically configured to:
[0091] Obtain a preset convolutional neural network and a training sample set; the training sample set includes sample image frames and their corresponding object feature labels;
[0092] Input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object feature corresponding to the sample image frame;
[0093] Obtain the difference information between the sample object feature corresponding to the sample image frame and the object feature label corresponding to the sample image frame according to a preset loss function;
[0094] If the difference information meets the preset training stop condition, stop training and use the current preset convolutional neural network lightweight model as the convolutional neural network for outputting the object feature corresponding to the image frame;
[0095] If the difference information does not meet the first preset training stop condition, adjust the preset parameters and return to input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object feature corresponding to the sample image frame.
[0096] Further, the training unit is specifically configured to:
[0097] Obtain a sample image frame and its corresponding object feature label, and obtain the object category of the object in the sample image frame;
[0098] If the number of sample image frames corresponding to the object category is less than a preset number, perform transformation processing on the sample image frame to obtain a transformed sample image frame;
[0099] Determine the training sample set according to the sample image frame and its corresponding object feature label, and the transformed sample image frame and its corresponding object feature label.
[0100] Figure 3 It is a schematic diagram of the object recognition device provided in the third embodiment of the present application. As Figure 3 shown, the object recognition device 3 of this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as an object recognition program. When the processor 30 executes the computer program 32, it implements the steps in the above-mentioned various object recognition method embodiments, such as Figure 1 the steps 101 to 104 shown. Or, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-mentioned device embodiments, such asFigure 2 The functions of the modules 210 to 240 shown.
[0101] Exemplarily, the computer program 32 may be divided into one or more modules / units, which are stored in the memory 31 and executed by the processor 30 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the object recognition device 3. For example, the computer program 32 may be divided into an acquisition unit, a segmentation unit, an extraction unit, and an identification unit, and the specific functions of each unit are as follows:
[0102] The acquisition unit is used to acquire the image to be recognized;
[0103] The segmentation unit is used to perform segmentation processing on the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame;
[0104] The extraction unit is used to extract features of the image frame through a convolutional neural network to obtain the object features corresponding to the image frame;
[0105] The identification unit is used to identify the object features through a machine learning algorithm to obtain the object category of the object included in the image to be recognized.
[0106] The object recognition device may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 This is only an example of the object recognition device 3 and does not constitute a limitation on the object recognition device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the object recognition device may further include an input / output device, a network access device, a bus, etc.
[0107] The so-called processor 30 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0108] The memory 31 may be an internal storage unit of the object recognition device 3, such as a hard disk or memory of the object recognition device 3. The memory 31 may also be an external storage device of the object recognition device 3, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the object recognition device 3. Further, the object recognition device 3 may also include both an internal storage unit and an external storage device of the object recognition device 3. The memory 31 is used to store the computer program and other programs and data required by the object recognition device. The memory 31 may also be used to temporarily store the data that has been output or will be output.
[0109] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned device / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference may be specifically made to the method embodiments section, and details will not be elaborated here.
[0110] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details will not be elaborated here.
[0111] The embodiments of the present application also provide a timing device for a virtual timer. The timing device for the virtual timer includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the above method embodiments are implemented.
[0112] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments can be implemented.
[0113] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, it enables the mobile terminal to execute the steps in the above-mentioned method embodiments when executed.
[0114] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps in the above-mentioned method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0115] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0116] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0117] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the apparatus or unit can be in electrical, mechanical or other forms.
[0118] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0119] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An object recognition method, characterized in that, it includes: Obtain the image to be recognized; Perform segmentation processing on the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame; each image frame includes one object; Extract features of the image frame through a convolutional neural network to obtain the object features corresponding to the image frame; Identify the object features through a machine learning algorithm to obtain the object category of the object included in the image to be recognized; The training process of the convolutional neural network includes: Obtain a preset convolutional neural network and obtain a training sample set; the training sample set includes sample image frames and their corresponding object feature labels; Input the sample image frame into a preset convolutional neural network lightweight model for processing to obtain the sample object features corresponding to the sample image frame; Obtain the difference information between the sample object features corresponding to the sample image frame and the object feature labels corresponding to the sample image frame according to a preset loss function; If the difference information meets the first preset stop training condition, stop training, and use the current preset convolutional neural network lightweight model as the convolutional neural network for outputting the object features corresponding to the image frame; If the difference information does not meet the first preset stop training condition, adjust the preset parameters, and return to input the sample image frame into the preset convolutional neural network lightweight model for processing to obtain the sample object features corresponding to the sample image frame; Obtain a sample validation set; the sample validation set includes sample validation pictures and their corresponding sample validation result labels; Obtain the recognition result corresponding to the sample validation picture according to the current preset convolutional neural network, and determine the accuracy rate corresponding to the sample validation set according to the recognition result corresponding to the sample validation picture and the sample validation result label; If the difference information and the accuracy rate meet the second preset stop training condition, stop training, and use the current preset convolutional neural network lightweight model as the object recognition model; If the difference information and the accuracy rate do not meet the second preset stop training condition, adjust the preset parameters, and return to input the sample image into the preset convolutional neural network lightweight model for processing to obtain the sample image recognition result corresponding to the sample image; Wherein, during the training process, the adjusted preset parameters include one or more of the learning rate, the learning rate schedule, the number of the sample images input into the preset convolutional neural network lightweight model each time, and the iteration training period.
2. The object recognition method according to claim 1, characterized in that, The convolutional neural network is a mobilenet-ssd neural network, a yolov4-tiny neural network or a nanodet neural network.
3. The object recognition method according to claim 1, characterized in that, The preset segmentation algorithm is a brute-force search algorithm, a Selective Search algorithm or an edge detection-based image segmentation algorithm.
4. The object recognition method according to claim 1, characterized in that, The machine learning algorithm is a support vector machine algorithm or an Adaboot algorithm.
5. The object recognition method according to claim 1, It is characterized in that The obtaining of the training sample set includes: Obtaining a sample image frame and its corresponding object feature label, and obtaining the object category of the object in the sample image frame; If the number of sample image frames corresponding to the object category is less than a preset number, performing transformation processing on the sample image frame to obtain a transformed sample image frame; Determining the training sample set according to the sample image frame and its corresponding object feature label, the transformed sample image frame and its corresponding object feature label.
6. An object recognition device It is characterized in that Applied to the method according to any one of claims 1 to 5, including: An obtaining unit, configured to obtain an image to be recognized; A segmentation unit, configured to perform segmentation processing on the image to be recognized according to a preset segmentation algorithm to obtain at least one image frame; An extraction unit, configured to extract object features corresponding to the image frame by means of a convolutional neural network; A recognition unit, configured to recognize the object features by means of a machine learning algorithm to obtain the object category of the object included in the image to be recognized.
7. The object recognition device according to claim 6 It is characterized in that The convolutional neural network is a mobilenet-ssd neural network, a yolov4-tiny neural network or a nanodet neural network.
8. An object recognition device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
9. A computer-readable storage medium storing a computer program It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Image recognition and neural network model training method, device and system
CN110163369A
Image-based multi-target segmentation and recognition method and system
CN111444773A