Object recognition method and device
The object recognition method in video frames addresses the challenge of accurately detecting coal self-heating hazards by using image processing algorithms to track smoke or fire regions, improving detection precision and reducing false positives in coal mining.
Patent Information
- Application Number
- CN202210550506.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-20
AI Technical Summary
In the prior art, fire detection in open-pit coal mines is inaccurate and costly, making it difficult to effectively prevent fires caused by spontaneous combustion of coal.
By obtaining multi-frame images from videos, using image-based smoke detection algorithms and edge detection algorithms, combining object recognition models, filtering and identifying smoke or flame regions with dynamic characteristics, and generating accurate identification results.
It improves the identification accuracy and efficiency of coal spontaneous combustion fires, reduces detection costs, and ensures the accuracy and reliability of identification results.
Smart Images

Figure CN114973082B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to an object recognition method. Background Art
[0002] As a major energy source, coal plays an important role in the process of economic development. Coal spontaneous combustion is one of the main disasters in the coal mining process. Due to the large number of fires caused by coal spontaneous combustion currently, coal spontaneous combustion has become one of the main disasters affecting safe production. The spontaneous combustion of coal not only brings huge economic and energy losses, but also generates a large amount of harmful gases during the process of coal spontaneous combustion, causing serious pollution to the environment.
[0003] In addition, since it is not easy to detect fires in the coalfields of open-pit coal mines, how to prevent the occurrence of coalfield fires is an urgent problem to be solved. Although there are currently examples of detecting large-area fires through satellite remote sensing technology, the detection areas obtained by satellite detection are inaccurate and the cost is relatively high. Therefore, an effective method is urgently needed to solve such problems. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide an object recognition method. One or more embodiments of this specification also relate to an object recognition device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, an object recognition method is provided, including:
[0006] Obtain at least two frames of images to be recognized and at least two frames of reference images from the video to be processed;
[0007] Determine the initial image to be recognized that contains the initial recognition area among the at least two frames of images to be recognized;
[0008] Determine the reference image to be recognized adjacent to the initial image to be recognized among the at least two frames of reference images;
[0009] Perform type recognition on the object to be recognized in the initial recognition area of the initial image to be recognized according to the reference image to be recognized, and generate a recognition result.
[0010] Optionally, the obtaining at least two frames of images to be recognized from the video to be processed includes:
[0011] Determine the video format of the video to be processed, and decode the video to be processed through the decoding protocol corresponding to the video format;
[0012] Extract video frames from the generated decoded results at preset time intervals to obtain at least two images to be recognized.
[0013] Optionally, determining the initial image to be recognized that contains the initial recognition region among the at least two images to be recognized includes:
[0014] Perform object recognition processing on the at least two images to be recognized using a preset object recognition algorithm, and determine the initial image to be recognized that contains the initial recognition region among the at least two images to be recognized according to the recognition results.
[0015] Optionally, the object to be recognized includes smoke, and the initial recognition region includes an initial smoke recognition region;
[0016] Correspondingly, performing object recognition processing on the at least two images to be recognized using a preset object recognition algorithm includes:
[0017] Perform object recognition processing on the at least two images to be recognized using an image-based smoke detection algorithm to obtain the recognition results of the initial smoke recognition region in the at least two images to be recognized.
[0018] Optionally, performing type recognition on the object to be recognized in the initial recognition region of the initial image to be recognized according to the reference image to be recognized includes:
[0019] Determine whether the reference image to be recognized contains an initial recognition region;
[0020] If so, filter the initial recognition region included in the initial image to be recognized according to the initial recognition region included in the reference image to be recognized to generate a target recognition region;
[0021] Perform type recognition on the object to be recognized in the target recognition region.
[0022] Optionally, filtering the initial recognition region included in the initial image to be recognized according to the initial recognition region included in the reference image to be recognized includes:
[0023] Determine the first position information of the initial recognition region included in the reference image to be recognized in the reference image to be recognized;
[0024] Determine the second position information of the initial recognition region included in the initial image to be recognized in the initial image to be recognized;
[0025] Determine the offset between the first position information and the second position information, and filter the initial recognition region included in the initial image to be recognized according to the offset.
[0026] Optionally, screening the initial recognition regions included in the initial image to be recognized according to the initial recognition regions included in the reference image to be recognized includes:
[0027] Using an edge detection algorithm to detect the first contour information of the object to be recognized in the initial recognition regions included in the reference image to be recognized;
[0028] Using an edge detection algorithm to detect the second contour information of the object to be recognized in the initial recognition regions included in the initial image to be recognized;
[0029] Comparing the first contour information with the second contour information, and screening the initial recognition regions included in the initial image to be recognized according to the comparison result.
[0030] Optionally, type recognition of the object to be recognized in the target recognition region includes:
[0031] Performing type recognition on the object to be recognized in the target recognition region through an object recognition model.
[0032] Optionally, the object recognition method further includes:
[0033] Obtaining training data, where the training data includes training images containing objects to be recognized;
[0034] Training the object recognition model to be trained according to the training data to generate the object recognition model.
[0035] Optionally, the object recognition method further includes:
[0036] Judging whether the recognition result is accurate;
[0037] If not, adjusting the recognition result, and optimizing the parameters of the object recognition model based on the initial image to be recognized and the adjustment result.
[0038] Optionally, obtaining the training data includes:
[0039] Obtaining training images containing objects to be recognized;
[0040] Cropping the training images, and using the cropping results and the training images as training data.
[0041] According to the second aspect of the embodiments of the present specification, an object recognition device is provided, including:
[0042] An acquisition module, configured to acquire at least two images to be recognized and at least two reference images from a video to be processed;
[0043] A first determination module, configured to determine an initial image to be recognized that contains an initial recognition region among the at least two images to be recognized.
[0044] A second determination module, configured to determine a reference image to be recognized that is adjacent to the initial image to be recognized among the at least two reference images.
[0045] A recognition module, configured to perform type recognition on the object to be recognized in the initial recognition region of the initial image to be recognized according to the reference image to be recognized, and generate a recognition result.
[0046] According to a third aspect of the embodiments of the present specification, there is provided a computing device, including:
[0047] A memory and a processor;
[0048] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of any one of the object recognition methods.
[0049] According to a fourth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of any one of the object recognition methods are implemented.
[0050] According to a fifth aspect of the embodiments of the present specification, there is provided a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above object recognition method.
[0051] In an embodiment of the present specification, at least two images to be recognized and at least two reference images are obtained from a video to be processed, an initial image to be recognized that contains an initial recognition region among the at least two images to be recognized is determined, a reference image to be recognized that is adjacent to the initial image to be recognized among the at least two reference images is determined, and type recognition is performed on the object to be recognized in the initial recognition region of the initial image to be recognized according to the reference image to be recognized, and a recognition result is generated.
[0052] After an initial image to be recognized that contains an initial recognition region is recognized in an embodiment of the present specification, the dynamic characteristics of the object to be recognized can also be used to determine a reference image to be recognized adjacent to each initial image to be recognized from at least two reference images, and then, according to whether the reference image to be recognized contains a recognition region, type recognition is further performed on the object to be recognized in the recognition region of each initial image to be recognized to ensure the accuracy of the recognition result. Description of the Drawings
[0053] Figure 1 is a flowchart of an object recognition method provided by an embodiment of the present specification;
[0054] Figure 2 It is an architecture diagram of an object recognition system provided by an embodiment of this specification;
[0055] Figure 3 It is a process flow chart of an object recognition method provided by an embodiment of this specification;
[0056] Figure 4 It is a schematic structural diagram of an object recognition device provided by an embodiment of this specification;
[0057] Figure 5 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0058] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0059] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0060] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0061] First, the noun terms related to one or more embodiments of this specification are explained.
[0062] FCOS: Fully Convolutional One-Stage Object Detection, a first-order fully convolutional object detection.
[0063] Faster RCNN: A three-dimensional convolutional neural network.
[0064] In this specification, an object recognition method is provided. This specification also relates to an object recognition device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail one by one in the following embodiments.
[0065] Figure 1 The flowchart of an object recognition method provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0066] Step 102, obtain at least two frames of images to be recognized and at least two frames of reference images from the video to be processed.
[0067] Specifically, the video to be processed is a video generated by video capturing of an actual scene containing the object to be detected. This video to be processed can be obtained by shooting with an image capturing device, such as a smartphone or a camera equipped with a camera. The image to be recognized is an image containing the object to be recognized, and this image to be recognized can be obtained by extracting video frames from the video to be processed. The reference image is also an image containing the object to be recognized, and the reference image can be obtained by extracting video frames from other video frames in the video to be processed except the images to be recognized.
[0068] Among them, when extracting video frames from the video to be processed, at least two frames of images to be recognized can be uniformly extracted according to the frame rate of the video to be processed and at a certain time interval. For example, if the frame rate of the video to be processed is 50fps and 4 frames of images to be recognized need to be extracted within 1s, the extraction time interval of the video frames can be determined as 250ms. If the frame rate of the video to be processed is 100fps and 4 frames of images to be recognized need to be extracted within 1s, the extraction time interval of the video frames can be determined as 500ms. The reference image can also be uniformly extracted from the video to be processed according to the frame rate of the video to be processed and at a certain time interval. The time interval can be different from the time interval used for extracting the images to be recognized, or the position of the first frame of the reference image in the video to be processed can be different from the position of the first frame of the images to be recognized in the video to be processed.
[0069] Alternatively, at least two frames of images to be recognized or at least two frames of reference images can also be randomly extracted from the video to be processed. The specific extraction method can be determined according to actual needs and is not limited here.
[0070] In practical applications, the object to be detected can be coal, and the video to be processed can be a video obtained by shooting coal. The object to be recognized can be smoke or flame, and this image to be recognized can be used to detect and recognize smoke and flame of coal to determine whether the coal is self-igniting.
[0071] Among them, the video stream of coal (the video to be processed) can be collected by an ordinary camera, and there are no special requirements for the camera in this specification. Usually, the high-mounted method can be adopted to cover a large open-air detection area.
[0072] In specific implementation, since after obtaining the image to be recognized, it is necessary to determine the initial image to be recognized that contains the initial recognition area in the image to be recognized, and this determination process can be implemented by the algorithm detection module. The algorithm detection module can perform object recognition processing on the image to be recognized through a preset object recognition algorithm. Therefore, at least two images to be recognized are obtained from the video to be processed. Specifically, the video format of the video to be processed can be determined, and the video to be processed is decoded through the decoding protocol corresponding to the video format.
[0073] Video frames are extracted from the generated decoding result at preset time intervals to obtain at least two images to be recognized.
[0074] Specifically, when using the device to shoot the video to be processed of the target detection object, the device usually encodes the shooting result into a specific format. For example, a general camera uses the H264 format. Therefore, after obtaining the video to be processed, the video format of the video to be processed can be determined first, and then the video to be processed is decoded through the decoding protocol corresponding to the video format. Then, video frames are extracted from the generated decoding result at preset time intervals to obtain the corresponding video frame extraction result, that is, at least two images to be recognized are obtained.
[0075] In the case where the video format is the H264 format, since the video to be processed is compressed, specifically, several frames of images in the video to be processed are first divided into a group (GOP, that is, a sequence), and then each frame image in each group is defined as three types, namely I frame, B frame, and P frame. Then, the I frame is used as the base frame, the P frame is predicted by the I frame, and then the B frame is predicted by the I frame and the P frame. Finally, the prediction difference between the I frame data and the B frame image is stored and transmitted.
[0076] When decoding it, that is, it is necessary to use the data of the I frame to reconstruct the complete image of the I frame, sum the predicted value and the prediction error in the I frame to reconstruct the complete P frame image. For the B frame image, the B frame uses the previous I or P frame and the subsequent P frame as reference frames to determine the predicted value and two motion vectors of each point of the B frame image, and the prediction difference and the motion vector are transmitted. The receiving end determines the predicted value in the reference frame of the B frame image according to the motion vector and sums it with the difference to obtain the sample value of each point of the B frame image, so as to obtain the complete B frame image.
[0077] Among them, when extracting video frames from the video to be processed, at least two images to be recognized can be evenly extracted at a certain time interval according to the frame rate of the video to be processed. For example, if the frame rate of the video to be processed is 50fps and 4 images to be recognized need to be extracted within 1s, the extraction time interval of the video frames can be determined to be 250ms.
[0078] In the embodiments of this specification, by pulling the video stream collected by the camera and performing decoding and frame extraction processing, the generated images to be recognized meet the requirements of the preset object recognition algorithm for input, which is conducive to ensuring the recognition efficiency and the accuracy of the recognition result in the object recognition process.
[0079] Step 104: Determine the initial image to be recognized that contains the initial recognition region among the at least two images to be recognized.
[0080] Specifically, the initial image to be recognized is one frame or at least two frames among the at least two images to be recognized; the initial recognition region is the recognition region that contains the object to be recognized.
[0081] After the embodiments of this specification extract at least two images to be recognized, object recognition can be performed on each image to be recognized to obtain the initial image to be recognized that contains the initial recognition region. Since the initial recognition region is the recognition region that contains the object to be recognized, the obtained initial image to be recognized is the image to be recognized that is initially recognized from the at least two images to be recognized and may contain the object to be recognized.
[0082] In specific implementation, the embodiments of this specification can use an algorithm detection module to process the video to be processed to determine the initial image to be recognized that contains the initial recognition region among the at least two images to be recognized. When the algorithm detection module processes the video to be processed, the preset object recognition algorithm can be used. Therefore, to determine the initial image to be recognized that contains the initial recognition region among the at least two images to be recognized, specifically, the preset object recognition algorithm can be used to perform object recognition processing on the at least two images to be recognized, and the initial image to be recognized that contains the initial recognition region among the at least two images to be recognized can be determined according to the recognition result.
[0083] Further, when the target detection object is coal, the object to be recognized includes smoke, and the initial recognition region includes an initial smoke recognition region;
[0084] Correspondingly, when using the preset object recognition algorithm to perform object recognition processing on the at least two images to be recognized, specifically, the image-based smoke detection algorithm is used to perform object recognition processing on the at least two images to be recognized, and the recognition result of the initial smoke recognition region in the at least two images to be recognized is obtained.
[0085] Specifically, when the target detection object is coal, the object to be recognized can be smoke or fire, and the initial recognition area can be the recognition area containing smoke or fire. Therefore, the preset object recognition algorithm can be the coal smoke and fire recognition algorithm based on the single-stage detection network FCOS, specifically, the smoke detection algorithm based on a single-frame image. This algorithm can be used to perform preliminary smoke recognition on each of the at least two frames of images to be recognized, and obtain the corresponding recognition result, which can be the initial smoke recognition area containing smoke or fire in the image to be recognized.
[0086] In the embodiment of this specification, the smoke and fire detection algorithm based on a single-frame picture is used to perform preliminary recognition of smoke and open fire on each image to be recognized. Compared with the three-dimensional convolutional neural network Faster RCNN, it has advantages such as being anchor-free, Proposal free, and having the Centerness idea, and has more advantages in terms of performance and computational resource occupancy. Although the smoke and fire detection based on pictures has more advantages in terms of speed and flexibility, for smoke and fire, the smoke and fire detection method based on pictures ignores the dynamic characteristics of smoke or fire, including its shape, light changes, etc. In addition, for low-resolution images to be recognized, it is difficult to distinguish static targets from other similar objects in the image, and false alarms are likely to occur. Therefore, after the preliminary recognition of the images to be recognized, in order to ensure the accuracy of the recognition result, the initial area to be recognized in the recognition result can be further screened.
[0087] Step 106: Determine the reference image to be recognized adjacent to the initial image to be recognized among the at least two frames of reference images.
[0088] Specifically, as mentioned above, the initial image to be recognized is the image to be recognized determined by the initial recognition of at least two frames of images to be recognized and may contain the object to be recognized. Although the smoke and fire detection based on pictures has more advantages in terms of speed and flexibility, for smoke and fire, the smoke and fire detection method based on pictures ignores the dynamic characteristics of smoke or fire, including its shape, light changes, etc. In addition, for low-resolution images to be recognized, it is difficult to distinguish static targets from other similar objects in the image, and false alarms are likely to occur. Therefore, after the preliminary recognition of the images to be recognized, in order to ensure the accuracy of the recognition result, the initial area to be recognized in the recognition result can be further screened.
[0089] Therefore, to further ensure the accuracy of the recognition result of the initial recognition, after determining the initial image to be recognized in the embodiments of this specification, it is possible to further determine the reference images to be recognized adjacent to each initial image to be recognized among the at least two extracted reference images, so as to further recognize the object to be recognized in the initial recognition area of the initial image to be recognized according to the situation of the object to be recognized included in the reference images to be recognized, thereby ensuring the accuracy of the recognition result.
[0090] For example, there are 5 initial images to be recognized, namely P1, P2, P3, P4, and P5. For P1, the reference image to be recognized adjacent to P1 can be determined in the reference images first, and then the type of P1 can be recognized according to the reference image to be recognized. For P2, the reference image to be recognized adjacent to P2 can be determined in the reference images first, and then the type of P2 can be recognized according to the reference image to be recognized, and so on. The reference images to be recognized adjacent to each initial image to be recognized are determined respectively, so as to recognize the type of each initial image to be recognized according to the reference images to be recognized.
[0091] Step 108: Perform type recognition on the object to be recognized in the initial recognition area of the initial image to be recognized according to the reference image to be recognized, and generate a recognition result.
[0092] Specifically, in the case where the object to be recognized has dynamic characteristics, after determining the reference image to be recognized adjacent to the initial image to be recognized, the initial recognition area can be screened according to whether the reference object contains the object to be recognized, so as to utilize its dynamic characteristics to ensure the accuracy of the recognition result.
[0093] When specifically implemented, performing type recognition on the object to be recognized in the initial recognition area of the initial image to be recognized according to the reference image to be recognized includes:
[0094] Judge whether the reference image to be recognized contains the initial recognition area;
[0095] If so, screen the initial recognition area included in the initial image to be recognized according to the initial recognition area included in the reference image to be recognized, and generate a target recognition area;
[0096] Perform type recognition on the object to be recognized in the target recognition area.
[0097] Specifically, an object with dynamic characteristics, such as smoke or flame, can exist for a certain period of time, and its shape is continuously and dynamically changing. Therefore, if the object to be recognized has dynamic characteristics and there is an initial recognition area of the object to be recognized in the initial image to be recognized, then it can be determined whether the initial recognition area in the initial image to be recognized contains the image to be recognized, that is, to determine the accuracy of the recognition result of the initial recognition area, according to whether the adjacent reference image to be recognized in the extracted reference images contains the initial recognition area of the object to be recognized. For example, if it is determined that the initial image to be recognized contains the initial recognition area of the object to be recognized, and the initial recognition areas of the object to be recognized are also included in the two adjacent reference images to be recognized before and after the initial image to be recognized, then the initial recognition area included in the initial image to be recognized can be further screened according to the initial recognition area included in the reference image to be recognized, so as to further determine the object to be recognized included in the initial image to be recognized according to the screening result; while if it is determined that the initial image to be recognized contains the initial recognition area of the object to be recognized, but the initial recognition areas of the object to be recognized are not included in the two adjacent reference images to be recognized before and after the initial image to be recognized, then it can be determined that the accuracy of the recognition result of the initial recognition area in the initial image to be recognized is relatively low, and the initial image to be recognized can be deleted without subsequent screening and recognition processes.
[0098] In practical applications, the two adjacent reference images to be recognized before and after the initial image to be recognized can be two images among at least two extracted reference images. To ensure the accuracy of the recognition result, in the embodiments of this specification, the object recognition process can be first performed on the reference image to be recognized, and then it can be determined whether the reference image to be recognized contains the initial recognition area of the object to be recognized according to the recognition result, so as to further determine the accuracy of the recognition result of the initial recognition area in the initial image to be recognized.
[0099] Furthermore, in the case where it is determined that the reference image to be recognized contains the initial recognition area, since the object to be recognized has dynamic characteristics, the position of the object to be recognized will change to a certain extent in different frames of the image to be recognized, that is, there will be a certain offset in the position of the object to be recognized in different frames of the image to be recognized. Therefore, in the embodiments of this specification, the initial image to be recognized can be further screened according to the magnitude of the offset between the positions of the object to be recognized in different frames of the image to be recognized, or the magnitude of the offset between the positions of the initial recognition areas.
[0100] Based on this, screening the initial recognition area included in the initial image to be recognized according to the initial recognition area included in the reference image to be recognized includes:
[0101] Determining the first position information of the initial recognition area included in the reference image to be recognized in the reference image to be recognized;
[0102] Determine the second position information of the initial region to be recognized included in the initial image to be recognized in the initial image to be recognized;
[0103] Determine the offset between the first position information and the second position information, and filter the initial recognition regions included in the initial image to be recognized according to the offset.
[0104] Alternatively, filtering the initial recognition regions included in the initial image to be recognized according to the initial recognition regions included in the reference image to be recognized includes:
[0105] Use an edge detection algorithm to detect the first contour information of the object to be recognized in the initial recognition region included in the reference image to be recognized;
[0106] Use an edge detection algorithm to detect the second contour information of the object to be recognized in the initial recognition region included in the initial image to be recognized;
[0107] Compare the first contour information with the second contour information, and filter the initial recognition regions included in the initial image to be recognized according to the comparison result.
[0108] Specifically, it can be determined by determining the first position information of the initial recognition region included in the reference image to be recognized in the reference image to be recognized, and then determining the second position information of the initial region to be recognized included in the initial image to be recognized in the initial image to be recognized. Then, through the first position information and the second position information, the first center point coordinates of the initial recognition region included in the reference image to be recognized and the second center point coordinates of the initial region to be recognized included in the initial image to be recognized can be determined respectively.
[0109] Among them, since the initial image to be recognized may include one or at least two initial recognition regions, similarly, the reference image to be recognized may include one or at least two initial recognition regions. Therefore, after determining the center coordinates of each initial recognition region, the distance between the second center point coordinates of the initial region to be recognized included in the initial image to be recognized and the first center point coordinates of each initial recognition region included in the reference image to be recognized can be calculated, that is, the offset of the first center point coordinates relative to the second center point coordinates. For any initial recognition region in the initial image to be recognized, if it is determined that the offset between the first center point coordinates and the second center point coordinates of the initial recognition region is greater than 0 and less than or equal to the preset offset threshold, then the initial recognition region can be determined as the target recognition region.
[0110] Alternatively, the first contour information of the object to be recognized in the initial recognition region included in the reference image to be recognized can be detected through an edge detection algorithm, and then the second contour information of the object to be recognized in the initial recognition region included in the initial image to be recognized can be detected. Then, the initial recognition region included in the initial image to be recognized can be screened by comparing the first contour information and the second contour information.
[0111] Among them, since the initial image to be recognized may include one or at least two initial recognition regions, and similarly, the reference image to be recognized may include one or at least two initial recognition regions. Therefore, after determining the contour information of the object to be recognized in each initial recognition region, the second contour information of the object to be recognized in the initial recognition region included in the initial image to be recognized can be compared with the first contour information of the object to be recognized in each initial recognition region included in the reference image to be recognized, so as to determine whether the contour of the object to be recognized in adjacent frames of images to be recognized has changed. For any initial recognition region in the initial image to be recognized, if there is no second contour information that is consistent with the first contour information of the object to be recognized in this initial recognition region, this initial recognition region can be determined as the target recognition region.
[0112] After screening and determining the target recognition region, due to the complexity of the actual object recognition scenario. For example, the object to be recognized is smoke or fire, but there are various objects similar to smoke and fire in the actual scenario. Only using the reference image to be recognized to screen the initial recognition region may still result in inaccurate screening results and thus false alarms. Therefore, an object recognition model can be added to further classify the screening results, that is, the type of the object to be recognized in the target recognition region can be recognized to generate the corresponding recognition result.
[0113] In specific implementation, to recognize the type of the object to be recognized in the target recognition region, specifically, the object recognition model can be used to recognize the type of the object to be recognized in the target recognition region.
[0114] Specifically, after determining the target recognition region in the initial image to be recognized, the initial image to be recognized including the target recognition region can be input into the object recognition model for processing, and the type recognition results corresponding to the objects to be recognized included in each target recognition region in this initial image to be recognized output by the object recognition model can be obtained. Among them, the type recognition result can include types such as smoke or fire.
[0115] In practical applications, to ensure the accuracy of the model output results, before using the object recognition model for type recognition, the object recognition model to be trained needs to be trained first, which is specifically implemented through the following method:
[0116] Obtain training data, where the training data includes training images containing objects to be recognized;
[0117] Train an object recognition model to be trained based on the training data to generate the object recognition model.
[0118] Specifically, training the object recognition model is equivalent to adjusting the model parameters of the model. Therefore, after obtaining the training data, that is, the training images containing the objects to be recognized, the training images can be input into the object recognition model for processing. In addition, there may be recognition regions containing the objects to be recognized in the training images, enabling the object recognition model to perform type recognition on the objects to be recognized in the recognition regions and adjust the model parameters based on the recognition results, thereby generating the trained object recognition model.
[0119] In addition, the embodiment of this specification can also re - input the images to be recognized that the object recognition model fails to recognize or has inaccurate recognition results into the object recognition model for learning to optimize the object recognition model, specifically, by judging whether the recognition results are accurate;
[0120] If not, adjust the recognition results, and optimize the parameters of the object recognition model based on the initial images to be recognized and the adjustment results.
[0121] Specifically, after the object recognition model outputs the type recognition results of the objects to be recognized in the images to be recognized, the user can judge whether the type recognition results are accurate. When the type recognition results are consistent with the actual types of the objects to be recognized, it can be determined that the recognition results are accurate; when the type recognition results are inconsistent with the actual types of the objects to be recognized, or the type recognition results do not include the types of the objects to be recognized, that is, the object recognition model fails to successfully recognize the types of the objects to be recognized, it is determined that the recognition results are inaccurate.
[0122] In this case, the images to be recognized can be directly used as training data and re - input into the object recognition model for learning, or the user can label the types of the objects to be recognized in the images to be recognized, and then input the images to be recognized and the labeling results into the object recognition model for learning to adjust and optimize the model parameters of the object recognition model.
[0123] The embodiment of this specification optimizes the model parameters of the object recognition model by means of data reflux of difficult example samples, which is beneficial to improving the model accuracy of the object recognition model, and thus beneficial to improving the accuracy of its recognition results.
[0124] When specifically implemented, obtaining the training data includes:
[0125] Obtain training images containing objects to be recognized;
[0126] Crop the training image, and use the cropping result and the training image as training data.
[0127] Specifically, since most of the training images used to train the object recognition model usually come from pictures collected from the network and are not obtained by taking pictures of the real scenes of customers. However, since object recognition using the object recognition model can be applied to a variety of different object recognition scenarios, that is, the object recognition model is pursued to have generality in multiple scenarios and a high algorithm recognition recall rate. However, in the real scenes of customers, there are often various objects similar to the objects to be recognized, such as smoke and fire. In addition, due to the large differences in camera configurations and scenes among different customers, and the large differences in data distributions, the accuracy of the results obtained by simply using the preset object recognition algorithm and edge detection algorithm to process the image to be recognized is relatively low. To solve this problem, in the embodiments of this specification, on the basis of data feedback, an object recognition model is added to perform type recognition on the results generated by using the preset object recognition algorithm and edge detection algorithm to process the image to be recognized. This object recognition model not only uses the data collected from the network for training, but also uses the customer feedback data to improve the accuracy of the model.
[0128] In practical applications, when training the object recognition model, the pictures collected from the network and the pictures collected from the real scenes of customers can be processed as follows:
[0129] Determine the correctly recognized regions and false alarm regions in the target recognition region generated by using the preset object recognition algorithm and edge detection algorithm to recognize the image to be recognized in the pictures collected from the network. At the same time, randomly crop the pictures of the real scenes of customers, and use the correctly recognized regions, false alarm regions, and the cropped pictures together as training data to train the object recognition model, so that the object recognition model has stronger adaptability to different scenes of customers. Moreover, the object recognition model is simple to train, has a shorter model iteration cycle, and can also be specially customized and optimized for customers, supporting more functions.
[0130] In the embodiments of this specification, by cropping the training image and using the cropped image and the training image together to train the object recognition model, it is beneficial to improve the model accuracy of the object recognition model, and thus beneficial to improve the accuracy of its recognition results.
[0131] In one embodiment of this specification, at least two to-be-recognized images and at least two reference images are obtained from a video to be processed, an initial to-be-recognized image containing an initial recognition region is determined from the at least two to-be-recognized images, a reference to-be-recognized image adjacent to the initial to-be-recognized image is determined from the at least two reference images, and the type of the object to be recognized in the initial recognition region of the initial to-be-recognized image is recognized according to the reference to-be-recognized image, and a recognition result is generated.
[0132] After an initial to-be-recognized image containing an initial recognition region is recognized in an embodiment of this specification, the dynamic characteristics of the object to be recognized can also be utilized to determine a reference to-be-recognized image adjacent to each initial to-be-recognized image from at least two reference images, and then, according to whether the reference to-be-recognized image contains a recognition region, the type of the object to be recognized in the to-be-recognized region of each initial to-be-recognized image is further recognized to ensure the accuracy of the recognition result.
[0133] Figure 2 The architecture diagram of an object recognition system provided according to one embodiment of this specification is shown, which specifically includes:
[0134] A video acquisition device, a video processing module, an algorithm detection module, an event warning module, an evidence storage module, and an event notification module. Among them, the video processing module, the algorithm detection module, the event warning module, and the evidence storage module together constitute a real-time early warning system for coal spontaneous combustion.
[0135] Specifically, the video acquisition device can be a camera. In an embodiment of this specification, a video stream related to coal is acquired by an ordinary camera, and there is no special requirement for the camera. Generally, a high-position installation method is adopted to cover more open-air detection areas.
[0136] After the video acquisition device acquires a video stream, the video processing module of the real-time early warning system for coal spontaneous combustion can perform operations such as pulling the stream and decoding and extracting frames on the video stream so that the processed video frames meet the requirements of the algorithm detection module for input.
[0137] After the algorithm detection module receives the video frames extracted by the video processing module, it uses a smoke and flame detection algorithm based on single-frame images to preliminarily identify smoke or flame in the received video frames. Compared with the three-dimensional convolutional neural network Faster RCNN, it has advantages such as being anchor-free, Proposal free, and having the Centerness idea, and is more advantageous in terms of performance and computational resource occupancy. Although smoke and flame detection based on images is more advantageous in terms of speed and flexibility, for smoke and flame, identifying smoke or flame based on single-frame images ignores the dynamic characteristics of smoke and flame, including changes in shape and light. In addition, it is difficult to distinguish static targets in low-resolution images from similar objects, and false alarms are likely to occur. Therefore, further processing is required after preliminary identification.
[0138] Based on this, the feature of the event warning module is aimed at the dynamic characteristics of smoke and flame. For example, in the video frames of the few frames before and after the target frame within a short period of time, the recognition area of the same smoke or flame will not change significantly, and natural smoke and flame will be accompanied by shape changes during the appearance process. For these characteristics, the following will be adopted:
[0139] a. Compare whether the position of the recognition area of the same smoke or flame in the video frames output by the algorithm detection module and the adjacent video frames shifts;
[0140] b. Use a morphological smoke and flame edge contour detection algorithm for inter-frame comparison, and further screening and filtering are carried out by excluding target recognition areas with no change at all.
[0141] In addition, due to the complex actual scene, there are various objects similar to smoke and flame. Just using the event warning module to screen according to the contour and position may still have problems of false alarms due to inaccurate recognition results. Therefore, an object recognition model can be added to further identify the screening results of the event warning module.
[0142] After screening and recognition, a final event trigger event notification is generated.
[0143] For the training of the object recognition model, in addition to training data, return data in the actual scene will also be used to improve the accuracy of the model training results. At the same time, images in the actual scene can be randomly cropped as training data, so that the trained model has stronger adaptability in the actual scene. Moreover, the training process of the object recognition model is relatively simple, the model iteration cycle is shorter, and it can also be customized and optimized for special scenes, supporting more functions.
[0144] In addition, if the event warning module issues an event warning for a certain video frame, the evidence storage module will store the video frame, the position information corresponding to the target recognition area in the video frame, and its corresponding warning time according to the output result of the event warning module. Alternatively, the evidence storage module can also perform search slicing on the original video stream to integrate the video frames with warnings in the order in which they appear in the original video stream, generate a continuous video and store it locally, and also provide a network synchronization upload function for users handling the event to query and view.
[0145] After identifying the initial image to be recognized that contains the initial recognition area in the embodiments of this specification, the dynamic characteristics of the object to be recognized can also be used to determine the adjacent reference images to be recognized of each initial image to be recognized from at least two images to be recognized, and then, according to whether the reference images to be recognized contain a recognition area, further perform type recognition on the objects to be recognized in the recognition areas of each initial image to be recognized to ensure the accuracy of the recognition result.
[0146] The following combines the attached Figure 3 , taking the application of the object recognition method provided in this specification in the coal smoke detection scenario as an example, to further illustrate the object recognition method. Among them, Figure 3 FIG. shows a flowchart of the processing procedure of an object recognition method provided by an embodiment of this specification, which specifically includes the following steps.
[0147] Step 302, obtain the video to be processed through the video processing module, determine the video format of the video to be processed, and decode the video to be processed through the decoding protocol corresponding to the video format.
[0148] Step 304, extract video frames from the generated decoding result at a first preset time interval to obtain at least two images to be recognized, and extract video frames from the generated decoding result at a second preset time interval to obtain at least two reference images.
[0149] Step 306, use the image-based smoke detection algorithm through the algorithm detection module to perform smoke recognition processing on at least two images to be recognized, and obtain the recognition results of the initial smoke recognition areas in at least two images to be recognized.
[0150] Step 308, determine the initial images to be recognized that contain the initial recognition area among at least two images to be recognized according to the recognition results.
[0151] Step 310, determine the reference images to be recognized adjacent to the initial images to be recognized among at least two reference images through the event warning module.
[0152] Step 312, determine whether the reference images to be recognized contain the initial recognition area.
[0153] If so, execute step 314.
[0154] Step 314: Use an edge detection algorithm to detect the first contour information of the object to be recognized in the initial recognition region included in the reference image to be recognized, and detect the second contour information of the object to be recognized in the initial recognition region included in the initial image to be recognized.
[0155] Step 316: Compare the first contour information with the second contour information, and screen the initial recognition region included in the initial image to be recognized according to the comparison result to generate a first target recognition region.
[0156] Step 318: Use an edge detection algorithm to detect the first contour information of the object to be recognized in the initial recognition region included in the reference image to be recognized, and use a detection algorithm to detect the second contour information of the object to be recognized in the initial recognition region included in the initial image to be recognized.
[0157] Step 320: Compare the first contour information with the second contour information, and screen the first target recognition region included in the initial image to be recognized according to the comparison result to generate a second target recognition region.
[0158] Step 322: Use an object recognition model to recognize the type of the object in the second target recognition region to generate a corresponding smoke recognition result.
[0159] After the embodiment of this specification recognizes the initial image to be recognized including the initial recognition region, it can also use the dynamic characteristics of smoke to determine the reference images to be recognized adjacent to each initial image to be recognized from at least two images to be recognized, and then further perform smoke recognition on the regions to be recognized in each initial image to be recognized according to whether the reference images to be recognized contain recognition regions, so as to ensure the accuracy of the recognition result.
[0160] Corresponding to the above method embodiment, this specification also provides an embodiment of an object recognition device. Figure 4 The following shows a schematic structural diagram of an object recognition device provided by an embodiment of this specification. As Figure 4 shown, the device includes:
[0161] An acquisition module 402, configured to acquire at least two images to be recognized and at least two reference images from a video to be processed;
[0162] A first determination module 404, configured to determine the initial images to be recognized that include an initial recognition region among the at least two images to be recognized;
[0163] A second determination module 406, configured to determine the reference images to be recognized adjacent to the initial images to be recognized among the at least two reference images;
[0164] An identification module 408, configured to perform type identification on an object to be identified in an initial identification region of the initial image to be identified according to the reference image to be identified, and generate an identification result.
[0165] Optionally, the obtaining module 402 is configured to:
[0166] Determine the video format of the video to be processed, and decode the video to be processed through a decoding protocol corresponding to the video format;
[0167] Extract video frames from the generated decoding result at preset time intervals to obtain at least two images to be identified.
[0168] Optionally, the first determination module 404 is further configured to:
[0169] Perform object recognition processing on the at least two images to be identified by using a preset object recognition algorithm, and determine, according to the recognition result, the initial image to be identified that contains the initial recognition region among the at least two images to be identified.
[0170] Optionally, the object to be identified includes smoke, and the initial recognition region includes an initial smoke recognition region;
[0171] Correspondingly, the first determination module 404 is further configured to:
[0172] Perform object recognition processing on the at least two images to be identified by using an image-based smoke detection algorithm to obtain the recognition result of the initial smoke recognition region in the at least two images to be identified.
[0173] Optionally, the identification module 408 is further configured to:
[0174] Determine whether the reference image to be identified contains an initial recognition region;
[0175] If so, filter the initial recognition region included in the initial image to be identified according to the initial recognition region included in the reference image to be identified, and generate a target recognition region;
[0176] Perform type identification on the object to be identified in the target recognition region.
[0177] Optionally, the identification module 408 is further configured to:
[0178] Determine first position information of the initial recognition region included in the reference image to be identified in the reference image to be identified;
[0179] Determine the second position information of the initial region to be recognized included in the initial image to be recognized in the initial image to be recognized;
[0180] Determine the offset between the first position information and the second position information, and screen the initial recognition region included in the initial image to be recognized according to the offset.
[0181] Optionally, the recognition module 408 is further configured to:
[0182] Use an edge detection algorithm to detect the first contour information of the object to be recognized in the initial recognition region included in the reference image to be recognized;
[0183] Use an edge detection algorithm to detect the second contour information of the object to be recognized in the initial recognition region included in the initial image to be recognized;
[0184] Compare the first contour information with the second contour information, and screen the initial recognition region included in the initial image to be recognized according to the comparison result.
[0185] Optionally, the recognition module 408 is further configured to:
[0186] Perform type recognition on the object to be recognized in the target recognition region through an object recognition model.
[0187] Optionally, the object recognition device further includes a training module, which is configured to:
[0188] Obtain training data, where the training data includes training images containing objects to be recognized;
[0189] Train the object recognition model to be trained according to the training data to generate the object recognition model.
[0190] Optionally, the object recognition device further includes a judgment module, which is configured to:
[0191] Judge whether the recognition result is accurate;
[0192] If not, adjust the recognition result, and optimize the parameters of the object recognition model based on the initial image to be recognized and the adjustment result.
[0193] Optionally, the training module is further configured to:
[0194] Obtain training images containing objects to be recognized;
[0195] Crop the training images, and use the cropping results and the training images as training data.
[0196] The above is a schematic solution of an object recognition device according to this embodiment. It should be noted that the technical solution of this object recognition device and the technical solution of the above object recognition method belong to the same concept. For the details not described in the technical solution of the object recognition device, reference can be made to the description of the technical solution of the above object recognition method.
[0197] Figure 5 FIG. 4 shows a block diagram of a computing device 500 according to an embodiment of the present specification. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0198] The computing device 500 further includes an access device 540, and the access device 540 enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Card (NIC)), such as IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.
[0199] In an embodiment of the present specification, the above components of the computing device 500 and Figure 5 other components not shown in FIG. 4 may also be connected to each other, for example, via a bus. It should be understood that Figure 5 the block diagram of the computing device shown in FIG. 4 is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0200] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 500 can also be a mobile or stationary server.
[0201] Among them, the processor 520 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above object recognition method are implemented.
[0202] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above object recognition method belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above object recognition method.
[0203] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above object recognition method are implemented.
[0204] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above object recognition method belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above object recognition method.
[0205] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above object recognition method.
[0206] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above object recognition method belong to the same concept. For the detailed content not described in the technical solution of the computer program, reference can be made to the description of the technical solution of the above object recognition method.
[0207] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0208] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0209] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0210] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0211] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. An object recognition method, comprising: Obtaining at least two frames of images to be recognized and at least two frames of reference images from a video to be processed; Determining an initial image to be recognized that contains an initial recognition region among the at least two frames of images to be recognized; Determining a reference image to be recognized adjacent to the initial image to be recognized among the at least two frames of reference images; Performing type recognition on the object to be recognized in the initial recognition region of the initial image to be recognized according to the reference image to be recognized, and generating a recognition result, wherein the performing type recognition on the object to be recognized in the initial recognition region of the initial image to be recognized according to the reference image to be recognized includes: in the case of determining that the reference image to be recognized contains an initial recognition region, screening the initial recognition region according to the offset magnitude between the positions of the objects to be recognized or the offset magnitude between the positions of the initial recognition regions in different frames of images to be recognized, generating a target recognition region, and performing type recognition on the object to be recognized in the target recognition region.
2. The object recognition method according to claim 1, wherein the obtaining at least two frames of images to be recognized from the video to be processed includes: Determining the video format of the video to be processed, and decoding the video to be processed through a decoding protocol corresponding to the video format; Performing video frame extraction on the generated decoding result at a preset time interval to obtain at least two frames of images to be recognized.
3. The object recognition method according to claim 1, wherein the determining an initial image to be recognized that contains an initial recognition region among the at least two frames of images to be recognized includes: Performing object recognition processing on the at least two frames of images to be recognized by using a preset object recognition algorithm, and determining an initial image to be recognized that contains an initial recognition region among the at least two frames of images to be recognized according to the recognition result.
4. The object recognition method according to claim 3, wherein the object to be recognized includes smoke, and the initial recognition region includes an initial smoke recognition region; Correspondingly, the performing object recognition processing on the at least two frames of images to be recognized by using a preset object recognition algorithm includes: Performing object recognition processing on the at least two frames of images to be recognized by using an image-based smoke detection algorithm to obtain the recognition result of the initial smoke recognition region in the at least two frames of images to be recognized.
5. The object recognition method according to claim 1, wherein the performing type recognition on the object to be recognized in the initial recognition region of the initial image to be recognized according to the reference image to be recognized includes: Judging whether the reference image to be recognized contains an initial recognition region; If so, screening the initial recognition region included in the initial image to be recognized according to the initial recognition region included in the reference image to be recognized, and generating a target recognition region, including: screening the initial recognition region according to the offset magnitude between the positions of the objects to be recognized or the offset magnitude between the positions of the initial recognition regions in different frames of images to be recognized, and generating a target recognition region; Performing type recognition on the object to be recognized in the target recognition region.
6. The object recognition method according to claim 5, wherein screening the initial recognition regions included in the initial image to be recognized according to the initial recognition regions included in the reference image to be recognized comprises: Determining first position information of the initial recognition regions included in the reference image to be recognized in the reference image to be recognized; Determining second position information of the initial recognition regions included in the initial image to be recognized in the initial image to be recognized; Determining an offset between the first position information and the second position information, and screening the initial recognition regions included in the initial image to be recognized according to the offset.
7. The object recognition method according to claim 5, wherein screening the initial recognition regions included in the initial image to be recognized according to the initial recognition regions included in the reference image to be recognized comprises: Detecting first contour information of an object to be recognized in the initial recognition regions included in the reference image to be recognized by using an edge detection algorithm; Detecting second contour information of an object to be recognized in the initial recognition regions included in the initial image to be recognized by using an edge detection algorithm; Comparing the first contour information with the second contour information, and screening the initial recognition regions included in the initial image to be recognized according to the comparison result.
8. The object recognition method according to claim 5, wherein recognizing the type of the object to be recognized in the target recognition region comprises: Recognizing the type of the object to be recognized in the target recognition region through an object recognition model.
9. The object recognition method according to claim 8 further comprises: Obtaining training data, wherein the training data includes training images containing objects to be recognized; Training a to-be-trained object recognition model according to the training data to generate the object recognition model.
10. The object recognition method according to claim 9 further comprises: Judging whether the recognition result is accurate; If not, adjusting the recognition result, and optimizing parameters of the object recognition model based on the initial image to be recognized and the adjustment result.
11. The object recognition method according to claim 9, wherein obtaining the training data comprises: Obtaining training images containing objects to be recognized; Cropping the training images, and using the cropping results and the training images as training data.
12. An object recognition device, comprising: An obtaining module configured to obtain at least two images to be recognized and at least two reference images from a video to be processed; A first determining module configured to determine an initial image to be recognized including initial recognition regions from the at least two images to be recognized; A second determining module configured to determine a reference image to be recognized adjacent to the initial image to be recognized from the at least two reference images; An identification module, configured to perform type identification on an object to be identified in an initial identification region of the initial image to be identified according to the reference image to be identified, and generate an identification result, where the performing type identification on the object to be identified in the initial identification region of the initial image to be identified according to the reference image to be identified includes: when it is determined that the reference image to be identified contains the initial identification region, screening the initial identification region according to the magnitude of the offset between the positions of the objects to be identified in different frames of images to be identified or the magnitude of the offset between the positions of the initial identification regions, generating a target identification region, and performing type identification on the object to be identified in the target identification region.
13. A computing device, comprising: a memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the object identification method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the object identification method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Video character recognition method and device, equipment and storage medium
CN114332902A