Processing system, method and storage medium for moving objects in images to be detected

By combining machine learning models and geometric calculation methods, the moving objects in the image to be detected are identified and removed, and the problems of low recognition accuracy and efficiency in the prior art are solved, thereby achieving higher recognition accuracy and navigation and positioning accuracy.

CN114118188BActive Publication Date: 2025-05-06AUDI AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010894972.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-31
Publication Date
2025-05-06
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

In the prior art, when identifying and removing moving objects in the image to be detected, it is difficult to accurately identify semi-static objects and static objects, and the recognition accuracy of geometric methods is difficult to adjust, resulting in misidentification.

Method used

Using a processing system combining preset machine learning models and geometric calculation methods, dynamic objects are identified through machine learning models, and semi-static objects with changing positions are identified using geometric calculation methods, identifying and removing areas of moving objects.

Benefits of technology

It improves the recognition accuracy and efficiency of moving objects in the detected image, reduces the amount and complexity of the training data of the machine learning model, and enhances the accuracy of navigation positioning and the effect of 3D reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114118188B_ABST
    Figure CN114118188B_ABST
Patent Text Reader

Abstract

The present invention claims protection for a processing system for mobile objects in an image to be detected, a corresponding method, a computer device, and a computer-readable storage medium. The processing system includes: an image recognition unit, configured to: identify mobile objects in the image to be detected based on a preset machine learning model and a geometric calculation method; a mobile object identification unit, configured to: for mobile objects identified based on one or both of the preset machine learning model and the geometric calculation method, use an identification layer to mark the identified mobile objects in the image to be detected; an image processing unit, configured to: remove the area where the mobile objects marked by the identification layer are located in the image to be detected. Utilizing the technical solution of the present invention, mobile objects in the image to be detected can be accurately and efficiently identified and removed, and the images identified and processed by the solution of the present invention can improve the accuracy of navigation positioning based on high-precision maps, and can also be beneficial to 3D reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more specifically, to a processing system for a moving object in an image to be detected, a corresponding method, a computer device and a computer-readable storage medium. Background Art

[0002] When the vehicle is in autonomous driving mode and uses a high-precision map for navigation, or in an application scenario where 3D map reconstruction is required, it will be disturbed by moving objects in the target map area or the reconstructed scene, thereby affecting the positioning accuracy of the navigation system or affecting the 3D map reconstruction. In the prior art, machine learning methods or geometric methods are usually used alone to identify the moving objects in the image, but the use of machine learning methods alone may result in the inability to identify all moving objects (for example, semi-static objects and static objects that are not dynamic objects but have changed positions), and the use of geometric methods alone may lead to the erroneous identification of non-moving objects as moving objects due to the inability to accurately adjust the recognition accuracy.

[0003] Therefore, it is necessary to design a technical solution to accurately and efficiently identify and remove moving objects in scene images, so as to facilitate navigation positioning or 3D reconstruction based on the scene images after recognition processing. Summary of the invention

[0004] At least to solve the above technical problems, the present invention proposes a processing solution for moving objects in an image to be detected, aiming to accurately and efficiently identify and remove the moving objects in the image to be detected.

[0005] As a first aspect of the present invention, a processing system for a moving object in an image to be detected is provided, wherein the processing system comprises:

[0006] An image recognition unit is configured to: recognize a moving object in the image to be detected based on a preset machine learning model and a geometric calculation method;

[0007] A mobile object identification unit is configured to: for a mobile object identified based on one or both of a preset machine learning model and a geometric calculation method, mark the identified mobile object using an identification layer in the image to be detected;

[0008] The image processing unit is configured to remove the area where the moving object marked by the identification layer is located in the image to be detected.

[0009] In one embodiment, the predefined object classification attributes include static objects, semi-static objects, and dynamic objects, and the method of identifying the moving object in the image to be detected based on the preset machine learning model and geometric calculation method includes:

[0010] Using the machine learning model to identify dynamic objects in the image to be detected; and / or

[0011] The geometric calculation method is used to identify semi-static objects and / or dynamic objects whose positions have changed in the image to be detected.

[0012] In one embodiment, the machine learning model includes a classification model pre-trained based on predefined object classification attributes and a deep convolutional neural network, wherein:

[0013] The using the machine learning model to identify the dynamic object in the image to be detected includes:

[0014] Using the classification model to identify and classify the objects in the image to be detected according to the object classification attributes to determine the dynamic object; and / or

[0015] The method of using the geometric calculation method to identify the semi-static object and / or dynamic object whose position has changed in the image to be detected includes:

[0016] Constructing a reference image set of reference images associated with the image to be detected for comparison;

[0017] Selecting a reference image with the largest overlap with the image to be detected from the reference image set;

[0018] Determine a projection point of at least one detection point in the reference image in the image to be detected, calculate the Hamming distance between any pair of the detection points and the projection point, and determine whether the Hamming distance is greater than a preset threshold;

[0019] In response to the Hamming distance being greater than a preset threshold, it is determined that a position of an object corresponding to the projection point in the image to be detected has changed.

[0020] In one embodiment, the image processing unit is further configured to:

[0021] Pixel information of a position corresponding to the removed area in the image to be detected is obtained from a reference image associated with the image to be detected, and the removed area in the image to be detected is filled based on the pixel information.

[0022] As a second aspect of the present invention, a processing method for a moving object in an image to be detected is provided, wherein the processing method comprises:

[0023] Identify the moving object in the image to be detected based on a preset machine learning model and a geometric calculation method;

[0024] For a moving object identified based on one or both of a preset machine learning model and a geometric calculation method, marking the identified moving object using an identification layer in the image to be detected;

[0025] The area where the moving object marked by the identification layer is located in the image to be detected is removed.

[0026] In one embodiment, the predefined object classification attributes include static objects, semi-static objects, and dynamic objects, and the method of identifying the moving object in the image to be detected based on the preset machine learning model and geometric calculation method includes:

[0027] Using the machine learning model to identify dynamic objects in the image to be detected; and / or

[0028] The geometric calculation method is used to identify semi-static objects and / or dynamic objects whose positions have changed in the image to be detected.

[0029] In one embodiment, the machine learning model includes a classification model pre-trained based on predefined object classification attributes and a deep convolutional neural network, wherein:

[0030] The using the machine learning model to identify the dynamic object in the image to be detected includes:

[0031] Using the classification model to identify and classify the objects in the image to be detected according to the object classification attributes to determine the dynamic object; and / or

[0032] The method of using the geometric calculation method to identify the semi-static object and / or dynamic object whose position has changed in the image to be detected includes:

[0033] Constructing a reference image set of reference images associated with the image to be detected for comparison;

[0034] Selecting a reference image with the largest overlap with the image to be detected from the reference image set;

[0035] Determine a projection point of at least one detection point in the reference image in the image to be detected, calculate the Hamming distance between any pair of the detection points and the projection point, and determine whether the Hamming distance is greater than a preset threshold;

[0036] In response to the Hamming distance being greater than a preset threshold, it is determined that a position of an object corresponding to the projection point in the image to be detected has changed.

[0037] In one embodiment, the processing method further includes:

[0038] Pixel information of a position corresponding to the removed area in the image to be detected is obtained from a reference image associated with the image to be detected, and the removed area in the image to be detected is filled based on the pixel information.

[0039] As a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program implements the processing method of the present invention when executed by a processor.

[0040] As a fourth aspect of the present invention, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, wherein the processor implements the processing method of the present invention when executing the computer program.

[0041] By using the technical solution of the present invention, by combining the preset machine learning model with the geometric calculation method, it is possible to ensure the accuracy of recognition of the moving object in the image to be detected, improve the recognition efficiency, and reduce the amount and complexity of the training data of the preset machine learning model. Furthermore, the image processed by the solution of the present invention can not only improve the accuracy of navigation positioning based on high-precision maps, but also facilitate 3D reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Non-limiting and non-exhaustive embodiments of the present invention are described by way of example with reference to the following drawings, in which:

[0043] Figure 1 A schematic diagram showing a processing system for a moving object in an image to be detected according to an embodiment of the present invention;

[0044] Figure 2 A flow chart showing a method for processing a moving object in an image to be detected according to one embodiment of the present invention;

[0045] Figure 3a-3c A schematic diagram of an application of a processing system and / or method for a moving object in an image to be detected according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and implementation schemes. It should be understood that the specific implementation schemes described herein are only used to explain this application and are not used to limit this application.

[0047] As a first aspect of the present invention, a processing system 100 for detecting moving objects in an image to be detected is provided. Specifically, the processing system 100 includes an image recognition unit 110, a moving object identification unit 120, and an image processing unit 130, and the units are communicatively coupled.

[0048] As can be understood, in order to identify the moving objects in the image to be detected, it is necessary to predefine the object classification attributes in the image to be detected. For example, the predefined object classification attributes may include static objects, semi-static objects, and dynamic objects. Specifically, taking the road intersection scene as an example, the static objects may include traffic signs, buildings on the side of the lane, and zebra crossings, etc.; the semi-static objects may include doors of shops along the street, vehicles parked on the roadside, etc.; the dynamic objects may include moving pedestrians, moving vehicles, etc. The predefined object classification attributes may be pre-set by those skilled in the art as needed when the processing system 100 is initialized, and may also be adjusted accordingly according to application requirements later.

[0049] The image recognition unit 110 may be configured to recognize a moving object in the image to be detected based on a preset machine learning model and a geometric calculation method.

[0050] Depending on the situation, the image recognition unit 110 can first use the machine learning model to identify dynamic objects in the image to be detected, and then use the geometric calculation method to identify semi-static objects and / or dynamic objects whose positions have changed in the image to be detected; or, it can also execute the above two methods at the same time, and then identify all moving objects in the image based on the common recognition results of the two methods.

[0051] In one embodiment, the preset machine learning model may be a classification model pre-trained based on predefined object classification attributes, a pre-acquired training data set, and a deep convolutional neural network, wherein the training data set may be an image data set, wherein each image contains an object corresponding to the predefined object classification attribute. The image recognition unit 110 may use the classification model to identify and classify the object in the image to be detected according to the object classification attribute to determine the dynamic object.

[0052] In another embodiment, the image recognition unit 110 can identify semi-static objects and / or dynamic objects whose positions in the image to be detected have changed through the following steps: Step 1, constructing a reference image set of reference images associated with the image to be detected for comparison; Step 2, selecting a reference image with the greatest overlap with the image to be detected from the reference image set; Step 3, determining the projection point of at least one detection point in the reference image in the image to be detected, calculating the Hamming distance between any pair of the detection points and the projection points, and judging whether the Hamming distance is greater than a preset threshold; Step 4, in response to the Hamming distance being greater than the preset threshold, determining that the position of the object corresponding to the projection point in the image to be detected has changed.

[0053] Specifically, in step 1, a collection of multiple groups of environmental images associated with the image to be detected (wherein the environmental images do not include any moving objects) can be used as a reference image set; if there are no environmental images that do not include any moving objects or the number of such environmental images is insufficient, the image area in the environmental image that does not include moving objects can also be used as a reference image in the reference image set.

[0054] In step 2, the image recognition unit 110 can extract point features of any frame of reference image in the reference image set, and then select a frame of image with the largest overlap with the current image to be detected from the reference image set based on the point features for any frame of image to be detected input into the processing system 100 as a reference image. Specifically, the image recognition unit 110 can extract multiple point features of any frame of reference image in the reference image set using a fast feature point extraction and description algorithm, and determine the corresponding projections of the multiple point features in the image to be detected, respectively calculate the sum of the absolute values ​​of the Hamming distances between the multiple point features of each frame of reference image and the corresponding projections on the image to be detected, and select the group with the smallest sum of absolute values, so as to determine that the reference image corresponding to the group with the smallest sum of absolute values ​​has the largest overlap with the current image to be detected. The fast feature point extraction and description algorithm can be, for example, an oriented rapid rotation strategy (ORB).

[0055] In step 3 and step 4, the image recognition unit 110 may determine at least one detection point x in the reference image based on the reference image selected in step 2 that has the greatest overlap with the image to be detected. For example, the detection point x may be determined by using a fast feature point extraction and description algorithm to determine the point feature in step 2, and then the projection point x' of the detection point x in the reference image in the current image to be detected is calculated; based on multiple detection points x and corresponding projection points x', the viewing angle difference α between the reference image and the current image to be detected is calculated to determine whether the viewing angle difference α is greater than a preset viewing angle difference threshold; in response to the viewing angle difference α being less than the preset viewing angle difference threshold, the feature descriptors B of the detection points x in the reference image are calculated respectively. x and the feature descriptor B of the projection point x' in the current image to be detected x' Based on the identification of the feature descriptor, the detection point x and the projection point x' are determined in pairs, and the Hamming distance H between each pair of the detection point x and the projection point x' is calculated. dist , determine the Hamming distance H dist Is it greater than the preset Hamming distance threshold τ B ; In response to the Hamming distance H dist Greater than the preset Hamming distance threshold τ B, it is determined that the object corresponding to the projection point x' in the current environment image is a semi-static object and / or a dynamic object whose position has changed.

[0056] For example, the viewing angle difference threshold is preferably 20° to 30°. In addition, the preset Hamming distance threshold τB can be adjusted according to different needs or scenarios. There are technical difficulties in setting a suitable preset Hamming distance threshold τB. Specifically, if the preset Hamming distance threshold τB is set very small, static objects will be misidentified as dynamic objects when the image to be detected is identified by a geometric method; and if the preset Hamming distance threshold τB is set very large, dynamic objects will be misidentified as static objects when the image to be detected is identified by a geometric method. Therefore, in order to overcome this technical difficulty, preferably, in the solution of the present invention, it is not necessary to set the preset Hamming distance threshold τB particularly accurately, as long as the preset Hamming distance threshold τB can make a part of the dynamic objects be identified as static objects. Because even if the setting of the preset Hamming distance threshold τB makes a part of the dynamic objects be identified as static objects, it can be corrected by the machine learning model. For example, if the preset Hamming distance threshold τB is set in the geometric calculation method, a part of the dynamic objects will be ignored, but the recognition efficiency of the geometric calculation method can be improved, and the machine learning model (artificial intelligence classification model) can be used to identify the ignored dynamic objects. As for the machine learning model itself, it is only necessary to consider the dynamic objects that frequently appear in the application scenarios (for example, in the scene of road intersections, only pedestrians and vehicles need to be considered) for learning and training, thereby reducing the amount and complexity of training data. Overall, by combining the recognition results of the machine learning model and the geometric calculation method, it is possible to ensure the accuracy of recognition of moving objects in the image to be detected and improve the recognition efficiency.

[0057] The mobile object identification unit 120 may be configured to: for a mobile object identified based on one or both of a preset machine learning model and a geometric calculation method, mark the identified mobile object using an identification layer in the image to be detected.

[0058] As can be understood, in the stage of identifying the moving object in the image to be detected, the reference image and the image to be detected are both RGB images. When the mobile object identification unit 120 identifies the identified moving object, it is necessary to convert the image to be detected from an RGB image to a depth image. However, due to the time difference between the RGB and depth images and the depth discontinuity of the moving object itself, there is an error between the position of the moving object in the depth image and the RGB image. Therefore, in order to eliminate the error, the segmented area of ​​the moving object can be enlarged by a preset or available image segmentation algorithm, and then the identification layer mark is used based on the enlarged segmented area to ensure that the moving object is marked. The mark can highlight or otherwise highlight the area where the moving object is located.

[0059] The image processing unit 130 may be configured to remove the area where the moving object marked by the identification layer is located in the image to be detected.

[0060] As can be understood, any available image processing algorithm can be used to remove the area where the moving object marked by the identification layer in the image to be detected is located. For example, the image processing algorithm can be a region-based segmentation method (FCN) or an edge-based segmentation method (Mask-RCNN).

[0061] Preferably, the image processing unit 130 may be further configured to: obtain pixel information of a position corresponding to the removed area in the image to be detected from a reference image associated with the image to be detected, and fill the removed area in the image to be detected based on the pixel information.

[0062] Depending on the situation, if the image to be detected is used for 3D reconstruction after removing the moving object, the "blanks" left in the removed area of ​​the image to be detected will affect the 3D reconstruction due to the lack of pixel information. For example, when using 3D reconstruction to build a high-precision map, if the "blanks" are not filled, the map cannot be reconstructed for the "blank" areas during 3D reconstruction due to the lack of pixel information.

[0063] To solve this problem, as described above, pixel information of the position corresponding to the removed area can be obtained from a series of reference image frames including the image to be detected (for example, from the previous frame and the next frame), and then the removed area in the image to be detected can be filled based on the pixel information.

[0064] According to a second aspect of the present invention, a processing method 200 for a moving object in an image to be detected is provided. The processing method 200 comprises:

[0065] S210, identifying a moving object in the image to be detected based on a preset machine learning model and a geometric calculation method;

[0066] S220, for a moving object identified based on one or both of a preset machine learning model and a geometric calculation method, marking the identified moving object using an identification layer in the image to be detected;

[0067] S230: Remove the area where the moving object marked by the identification layer is located in the image to be detected.

[0068] In one embodiment, the predefined object classification attributes include static objects, semi-static objects and dynamic objects, and the identifying of moving objects in the image to be detected based on a preset machine learning model and a geometric calculation method includes: using the machine learning model to identify dynamic objects in the image to be detected; and / or using the geometric calculation method to identify semi-static objects and / or dynamic objects whose positions have changed in the image to be detected.

[0069] In one embodiment, the machine learning model includes a classification model pre-trained based on predefined object classification attributes and a deep convolutional neural network, wherein the use of the machine learning model to identify dynamic objects in the image to be detected includes: using the classification model to identify and classify objects in the image to be detected according to the object classification attributes to determine the dynamic objects; and / or the use of the geometric calculation method to identify semi-static objects and / or dynamic objects whose positions have changed in the image to be detected includes: constructing a reference image set of reference images associated with the image to be detected for comparison; selecting a reference image with the greatest overlap with the image to be detected from the reference image set; determining the projection point of at least one detection point in the reference image in the image to be detected, calculating the Hamming distance between any pair of the detection points and the projection points, and determining whether the Hamming distance is greater than a preset threshold; in response to the Hamming distance being greater than the preset threshold, determining that the position of the object corresponding to the projection point in the image to be detected has changed.

[0070] In one embodiment, the processing method further includes: acquiring pixel information of a position corresponding to the removed area in the image to be detected from a reference image associated with the image to be detected, and filling the removed area in the image to be detected based on the pixel information.

[0071] It should be understood that the specific features described in the first aspect of this document regarding the processing system for moving objects in an image to be detected can also be similarly applied to the second aspect of the processing method for moving objects in an image to be detected for similar extension. For the sake of simplicity, it is not described in detail.

[0072] Combine the following Figure 3a-3c The application of the processing system and / or method for a moving object in an image to be detected according to the present invention is schematically described.

[0073] If you can understand, Figure 3a-3c The application scenario shown is that a vehicle uses a high-precision map to navigate on the road. As shown in the figure, for example, the scene (presented in grayscale) may include: the current vehicle O, traffic sign A, lane side building B, zebra crossing C, moving pedestrians D, the vehicle in front E, a disabled person in a wheelchair F, a pet dog G, and a movable roadside fruit stand H. It should be noted that the moving objects in the scene are shown in the form of rectangular boxes, where the rectangular boxes corresponding to the identifiable moving objects are filled with oblique lines, and the rectangular boxes corresponding to the moving objects that cannot be identified are filled with gray.

[0074] Reference Figure 3a If only the preset machine learning model is used to identify the image of the scene, it may only be able to identify the vehicle E in front and the moving pedestrian D as moving objects due to limited training data, and other untrained moving object types (such as a disabled person F in a wheelchair) cannot be identified. Since unlearned object types cannot be identified by the preset machine learning model, it also means that massive data training is required to achieve recognition in complex road conditions.

[0075] refer to Figure 3b If only the geometric calculation method is used to identify the image of the scene, it is impossible to determine the preset Hamming distance threshold τ B The reasonable value of may result in that the moving objects in the local area of ​​the scene (especially the objects at a long distance) may not be recognized due to the small difference, or on the contrary, it may cause the static objects to be mistakenly recognized as moving objects. For example, it may result in the failure to recognize the moving pedestrian D (gray) far away from the current vehicle O, or the traffic sign A may be mistakenly recognized as a moving object due to the change of light, while only the moving pedestrian D (slash) close to the current vehicle O, the disabled person F in a wheelchair, the movable fruit stand H on the roadside, and the vehicle E in front are recognized as moving objects.

[0076] Reference Figure 3c In the present invention, the preset machine learning model is used in combination with the geometric calculation method. When identifying based on the geometric calculation method, the preset Hamming distance threshold τ BThe setting is larger. Although this setting will ignore some dynamic objects (such as those dynamic objects at a long distance mentioned above), it can improve the recognition efficiency of the geometric calculation method. Then the machine learning model (classification model) can be used to identify the ignored dynamic objects. As for the machine learning model itself, it only needs to consider the dynamic objects that often appear in the application scenario (for example, in the scene of a road intersection, only pedestrians and vehicles need to be considered) for learning and training, thereby reducing the amount and complexity of training data. Overall, by combining the recognition results of the machine learning model and the geometric calculation method, it can not only ensure the recognition accuracy of the moving objects in the image to be detected, but also improve the recognition efficiency. For example, Figure 3c As shown, most of the moving objects in the scene can be identified, including a moving pedestrian D (diagonal line), a pet dog G, a disabled person F in a wheelchair, a movable fruit stand H on the roadside, and a vehicle E in front.

[0077] Based on the above Figure 3a-3c From the description of different situations, we can see that the current vehicle O is based on Figure 3c The recognition results shown have the highest navigation positioning accuracy when used for high-precision map navigation, because after the identified moving objects are removed, the interference of the moving objects is filtered out, and the only things left in the scene are the traffic sign A, the building B beside the lane, and the zebra crossing C, all of which are static objects, which are consistent with the data in the high-precision map and are easy to achieve accurate positioning.

[0078] For 3D reconstruction, for example, when the navigation system of the current vehicle O needs to update the high-precision map, the vehicle camera device can also be used to collect road environment images. Figure 3c The recognition results shown in the figure remove all moving objects, and then build a high-precision map based on only the traffic signs A, the buildings B beside the lane, and the zebra crossing C that remain in the scene, filtering out the interference of moving objects and making the map more accurate.

[0079] It should be understood that the various units of the processing system 100 for the mobile object in the image to be detected of the present invention can be implemented in whole or in part by software, hardware, firmware or a combination thereof. Each of the units can be embedded in the processor of the computer device in the form of hardware or firmware or independent of the processor, or can be stored in the memory of the computer device in the form of software for the processor to call to perform the operation of each unit. Each of the units can be implemented as an independent component or module, or two or more units can be implemented as a single component or module.

[0080] It should be understood by those skilled in the art that Figure 1The schematic diagram of the processing system 100 for the moving object in the image to be detected is only an exemplary block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device, processor or computer program embodying the solution of the present invention. The specific computer device, processor or computer program may include more or fewer components or modules than shown in the figure, or combine or split some components or modules, or have different arrangements of components or modules.

[0081] As a third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program implements the steps of the method of the second aspect of the present invention when executed by a processor. In one embodiment, the computer program is distributed on a plurality of computer devices or processors coupled to a network so that the computer program is stored, accessed and executed in a distributed manner by one or more computer devices or processors. A single method step / operation, or two or more method steps / operations, may be performed by a single computer device or processor or by two or more computer devices or processors. One or more method steps / operations may be performed by one or more computer devices or processors, and one or more other method steps / operations may be performed by one or more other computer devices or processors. One or more computer devices or processors may perform a single method step / operation, or perform two or more method steps / operations.

[0082] As a fourth aspect of the present invention, a computer device is provided, which includes a memory and a processor, wherein the memory stores computer instructions executable by the processor, and when the computer instructions are executed by the processor, the processor is instructed to execute the steps of the processing method for the mobile object in the image to be detected according to the second aspect of the present invention. The computer device can be a server, a vehicle-mounted terminal, or any other electronic device with necessary computing and / or processing capabilities in a broad sense. In one embodiment, the computer device may include a processor, a memory, a network interface, a communication interface, etc. connected through a system bus. The processor of the computer device can be used to provide necessary computing, processing and / or control capabilities. The memory of the computer device may include a non-volatile storage medium and an internal memory. An operating system, a computer program, etc. may be stored in or on the non-volatile storage medium. The internal memory can provide an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface and the communication interface of the computer device can be used to connect and communicate with external devices through a network. When the computer program is executed by the processor, the steps of the method of the present invention are executed.

[0083] It will be appreciated by those skilled in the art that all or part of the steps of the processing method for a mobile object in an image to be detected of the present invention can be completed by instructing related hardware such as a computer device or a processor through a computer program, and the computer program can be stored in a non-temporary computer-readable storage medium, and the steps of the processing method for a mobile object in an image to be detected of the present invention are implemented when the computer program is executed. Depending on the circumstances, any reference to a memory, storage, database or other medium herein may include non-volatile and / or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0084] The various technical features described above can be combined arbitrarily. Although all possible combinations of these technical features are not described, any combination of these technical features should be considered to be covered by this specification as long as there is no contradiction in such combination.

[0085] Although the present invention has been described in conjunction with the embodiments, it should be understood by those skilled in the art that the above description and the accompanying drawings are only exemplary and non-restrictive, and the present invention is not limited to the disclosed embodiments. Various modifications and variations are possible without departing from the spirit of the present invention.

Claims

1. A processing system for a moving object in an image to be detected, characterized in that: The processing system comprises: An image recognition unit is configured to: identify the moving object in the image to be detected based on a preset machine learning model and a geometric calculation method, wherein the predefined object classification attributes include semi-static objects and dynamic objects, and the identifying the moving object in the image to be detected based on the preset machine learning model and the geometric calculation method at least includes: using the geometric calculation method to identify the semi-static object and / or dynamic object whose position has changed in the image to be detected; A mobile object identification unit is configured to: for a mobile object identified based on one or both of a preset machine learning model and a geometric calculation method, mark the identified mobile object using an identification layer in the image to be detected; The image processing unit is configured to: remove the area where the moving object marked by the identification layer is located in the image to be detected, Wherein, the step of using the geometric calculation method to identify the semi-static object and / or dynamic object whose position has changed in the image to be detected includes: Based on at least one detection point in the reference image and a projection point of the at least one detection point in the image to be detected, it is determined whether a position of an object corresponding to the projection point in the image to be detected has changed.

2. The processing system according to claim 1, characterized in that The predefined object classification attributes include static objects, semi-static objects and dynamic objects, and the method of identifying the moving objects in the image to be detected based on the preset machine learning model and geometric calculation method further includes: The machine learning model is used to identify dynamic objects in the image to be detected.

3. The processing system according to claim 2, characterized in that The machine learning model includes a classification model pre-trained based on predefined object classification attributes and a deep convolutional neural network, wherein: The using the machine learning model to identify the dynamic object in the image to be detected includes: Using the classification model to identify and classify the object in the image to be detected according to the object classification attribute to determine the dynamic object; and The determining, based on at least one detection point in the reference image and a projection point of the at least one detection point in the image to be detected, whether a position of an object corresponding to the projection point in the image to be detected has changed comprises: Constructing a reference image set of reference images associated with the image to be detected for comparison; Selecting a reference image with the largest overlap with the image to be detected from the reference image set; Determine a projection point of at least one detection point in the reference image in the image to be detected, calculate the Hamming distance between any pair of the detection points and the projection point, and determine whether the Hamming distance is greater than a preset threshold; In response to the Hamming distance being greater than a preset threshold, it is determined that a position of an object corresponding to the projection point in the image to be detected has changed.

4. The processing system according to claim 3, characterized in that The image processing unit is further configured to: Pixel information of a position corresponding to the removed area in the image to be detected is obtained from a reference image associated with the image to be detected, and the removed area in the image to be detected is filled based on the pixel information.

5. A method for processing a moving object in an image to be detected, characterized in that: The processing method comprises: Identify the moving object in the image to be detected based on a preset machine learning model and a geometric calculation method, wherein the predefined object classification attributes include semi-static objects and dynamic objects, and the identifying the moving object in the image to be detected based on the preset machine learning model and the geometric calculation method at least includes: using the geometric calculation method to identify the semi-static object and / or dynamic object whose position in the image to be detected has changed; For a moving object identified based on one or both of a preset machine learning model and a geometric calculation method, marking the identified moving object using an identification layer in the image to be detected; removing the area where the moving object marked by the identification layer is located in the image to be detected, Wherein, the step of using the geometric calculation method to identify the semi-static object and / or dynamic object whose position has changed in the image to be detected includes: Based on at least one detection point in the reference image and a projection point of the at least one detection point in the image to be detected, it is determined whether a position of an object corresponding to the projection point in the image to be detected has changed.

6. The processing method according to claim 5, characterized in that: The predefined object classification attributes include static objects, semi-static objects and dynamic objects, and the method of identifying the moving objects in the image to be detected based on the preset machine learning model and geometric calculation method further includes: The machine learning model is used to identify dynamic objects in the image to be detected.

7. The processing method according to claim 6, characterized in that: The machine learning model includes a classification model pre-trained based on predefined object classification attributes and a deep convolutional neural network, wherein: The using the machine learning model to identify the dynamic object in the image to be detected includes: Using the classification model to identify and classify the object in the image to be detected according to the object classification attribute to determine the dynamic object; and The determining, based on at least one detection point in the reference image and a projection point of the at least one detection point in the image to be detected, whether a position of an object corresponding to the projection point in the image to be detected has changed comprises: Constructing a reference image set of reference images associated with the image to be detected for comparison; Selecting a reference image with the largest overlap with the image to be detected from the reference image set; Determine a projection point of at least one detection point in the reference image in the image to be detected, calculate the Hamming distance between any pair of the detection points and the projection point, and determine whether the Hamming distance is greater than a preset threshold; In response to the Hamming distance being greater than a preset threshold, it is determined that a position of an object corresponding to the projection point in the image to be detected has changed.

8. The processing method according to claim 7, characterized in that: The processing method also includes: Pixel information of a position corresponding to the removed area in the image to be detected is obtained from a reference image associated with the image to be detected, and the removed area in the image to be detected is filled based on the pixel information.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the processing method according to any one of claims 5 to 8.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the processing method according to any one of claims 5 to 8 is implemented.

Citation Information

Patent Citations

  • Street front order event video detection method based on deep learning and motion consistency

    CN108304798A

  • Point cloud map construction method and device, equipment and storage medium

    CN111402414A