Stereo Image Analysis and Recognition Method and System Based on Cloud Computer

By using the image data captured by multi-photographing devices and the stereoscopic image mapping relationship, the training data of the image recognition model is automatically checked and the training data of the image recognition model is solved, and efficient training data generation is achieved.

CN116883959BActive Publication Date: 2025-06-24V & G INFORMATION SYSTEM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311014692.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-14
Publication Date
2025-06-24
Estimated Expiration
2043-08-14

AI Technical Summary

Technical Problem

Existing image recognition technology relies on manual labeling of data, making it inconvenient to obtain training data.

Method used

By acquiring the image data captured by at least two imaging devices, the coordinate information of the target object is determined using the image analysis model, and the trustworthiness of the coordinate information is verified through the stereoscopic image mapping relationship to form training data.

Benefits of technology

Without manual labeling of data, high-quality training data can be easily generated and the training efficiency of image recognition models can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883959B_ABST
    Figure CN116883959B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for three-dimensional image analysis and recognition based on a cloud computer. The method includes: acquiring image data captured by at least two imaging devices for photographing a target object; inputting the first image data and the second image data into an image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data; determining the three-dimensional image mapping relationship between the first imaging device and the second imaging device; verifying the first coordinate information and the second coordinate information according to the three-dimensional image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information; when the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, using the first coordinate information as the annotation data of the first image data and using the second coordinate information as the annotation data of the second image data to form training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly to a three-dimensional image analysis and recognition method and system based on a cloud computer. Background Art

[0002] Image recognition refers to the technology of using a computer to process, analyze, and understand images to identify various different patterns of targets and objects, and is a practical application of applying deep learning algorithms.

[0003] Existing image recognition uses a pre-trained model to recognize the input image to determine the information of the target object in the image. For example, in the scenario of recognizing an image of a vehicle on a road, the coordinates of the target object (vehicle) (including position and distance, and the distance can also be referred to as depth) can be recognized based on the input image, and subsequent analysis can be performed. For example, the vehicle speed can be judged based on the coordinates of the vehicle at different times.

[0004] Existing image recognition models are usually trained using training data. The training data usually includes image data for input into the model and annotation data corresponding to the image data (such as the coordinates in the above example are annotations). The existing annotation scheme is usually manually annotated, and it is very inconvenient to obtain the training data. Summary of the Invention

[0005] The present invention provides a three-dimensional image analysis and recognition method and system based on a cloud computer to facilitate the acquisition of training data.

[0006] To solve the above technical problems, the present invention is implemented as follows:

[0007] In a first aspect, the present application provides a three-dimensional image analysis and recognition method based on a cloud computer. The method includes: obtaining image data captured by at least two imaging devices for photographing a target object, where the at least two imaging devices include a first imaging device and a second imaging device, and the image data includes first image data captured by the first imaging device and second image data captured by the second imaging device; inputting the first image data and the second image data into an image analysis model to determine first coordinate information of the target object in the first image data and second coordinate information of the target object in the second image data, where both the first coordinate information and the second coordinate information include position information and depth information; obtaining the spatial position relationship between the first imaging device and the second imaging device, and determining the three-dimensional image mapping relationship between the first imaging device and the second imaging device based on the spatial position relationship; verifying the first coordinate information and the second coordinate information based on the three-dimensional image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information; when the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, using the first coordinate information as the annotation data for the first image data and the second coordinate information as the annotation data for the second image data to form training data, so as to perform model training based on the training data.

[0008] Further, the imaging device includes a first optical component for long-distance shooting, a second optical component for close-range shooting, and an image sensor for receiving information from the first optical component and the second optical component; the obtaining image data captured by at least two imaging devices for photographing a target object includes: determining the first image data based on the acquisition data of the image sensor of the first imaging device; determining the second image data based on the acquisition data of the image sensor of the second imaging device.

[0009] Further, the determining the first image data based on the acquisition data of the image sensor of the first imaging device, or the determining the second image data based on the acquisition data of the image sensor of the second imaging device includes: performing data analysis on the acquisition data of the image sensor of the imaging device to determine whether the target object is within the shooting areas corresponding to the first optical component and the second optical component; when it is recognized that the target object is within the shooting areas corresponding to the first optical component and the second optical component, using the corresponding acquisition data as the image data.

[0010] Further, the verifying the first coordinate information and the second coordinate information based on the three-dimensional image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information includes: adjusting the first coordinate information based on the three-dimensional image mapping relationship to obtain mapped coordinate information; comparing the mapped coordinate information with the second coordinate information to determine the credibility of the first coordinate information and the second coordinate information based on the comparison result.

[0011] Further, the first coordinate information includes a first sub - coordinate and a second sub - coordinate. The first sub - coordinate corresponds to the first optical component of the first imaging device, and the second sub - coordinate corresponds to the second optical component of the first imaging device. The second coordinate information includes a third sub - coordinate and a fourth sub - coordinate. The third sub - coordinate corresponds to the first optical component of the second imaging device, and the fourth sub - coordinate corresponds to the second optical component of the second imaging device.

[0012] Further, the process of adjusting the first coordinate information according to the stereo image mapping relationship to obtain the mapped coordinate information can be determined according to the following formula: A = G * X, where A is the mapped coordinate information, G is the mapping matrix representing the stereo image mapping relationship, and X is the first coordinate information; the mapped coordinate information includes a first mapped coordinate corresponding to the first sub - coordinate and a second mapped coordinate corresponding to the second sub - coordinate. Comparing the mapped coordinate information with the second coordinate information to determine the credibility of the first coordinate information and the second coordinate information according to the comparison result includes: comparing the first mapped coordinate with the third sub - coordinate to determine the first comparison result; comparing the second mapped coordinate with the fourth sub - coordinate to determine the second comparison result; and determining the credibility of the first coordinate information and the second coordinate information according to the first comparison result and the second comparison result.

[0013] Further, the method further includes: obtaining the first fixed position relationship and the first focal length relationship between the first optical component and the second optical component of the first imaging device to determine the first fixed mapping relationship; determining the third mapped coordinate according to the first sub - coordinate and the first fixed mapping relationship; comparing the third mapped coordinate with the second sub - coordinate to determine the credibility of the first sub - coordinate and the second sub - coordinate, so as to determine the credibility of the first coordinate information; and discarding the first coordinate information when the credibility of the first coordinate information is lower than the first credibility threshold.

[0014] Further, the method further includes: obtaining the second fixed position relationship and the second focal length relationship between the first optical component and the second optical component of the second imaging device to determine the second fixed mapping relationship; determining the fourth mapped coordinate according to the third sub - coordinate and the second fixed mapping relationship; comparing the fourth mapped coordinate with the fourth sub - coordinate to determine the credibility of the third sub - coordinate and the fourth sub - coordinate, so as to determine the credibility of the second coordinate information; and discarding the second coordinate information when the credibility of the second coordinate information is lower than the second credibility threshold.

[0015] Second aspect, the present application provides a three-dimensional image analysis and recognition system based on a cloud computer. The system includes: an image data acquisition module, configured to acquire image data captured by at least two imaging devices for photographing a target object. The at least two imaging devices include a first imaging device and a second imaging device. The image data includes first image data captured by the first imaging device and second image data captured by the second imaging device; a coordinate information acquisition module, configured to input the first image data and the second image data into an image analysis model to determine first coordinate information of the target object in the first image data and second coordinate information of the target object in the second image data. The first coordinate information and the second coordinate information both include position information and depth information; a mapping relationship acquisition module, configured to acquire the spatial position relationship between the first imaging device and the second imaging device, and determine the three-dimensional image mapping relationship between the first imaging device and the second imaging device according to the spatial position relationship; a coordinate reliability verification module, configured to verify the first coordinate information and the second coordinate information according to the three-dimensional image mapping relationship to determine the reliability of the first coordinate information and the second coordinate information; a training data generation module, configured to, when the reliability of the first coordinate information and the second coordinate information exceeds a preset reliability threshold, use the first coordinate information as the annotation data of the first image data and the second coordinate information as the annotation data of the second image data to form training data for model training according to the training data.

[0016] Third aspect, the present application provides an electronic device, including: a memory and at least one processor; the memory is configured to store computer execution instructions; the at least one processor is configured to execute the computer execution instructions stored in the memory, so that the at least one processor executes the method as described in the first aspect.

[0017] The present application provides a three-dimensional image analysis and recognition method based on a cloud computer. The method includes: obtaining image data captured by at least two imaging devices for photographing a target object, where the at least two imaging devices include a first imaging device and a second imaging device, and the image data includes first image data captured by the first imaging device and second image data captured by the second imaging device; inputting the first image data and the second image data into an image analysis model to determine first coordinate information of the target object in the first image data and second coordinate information of the target object in the second image data, where both the first coordinate information and the second coordinate information include position information and depth information; obtaining the spatial position relationship between the first imaging device and the second imaging device, and determining the three-dimensional image mapping relationship between the first imaging device and the second imaging device based on the spatial position relationship; verifying the first coordinate information and the second coordinate information according to the three-dimensional image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information; when the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, using the first coordinate information as the annotation data of the first image data and the second coordinate information as the annotation data of the second image data to form training data for model training based on the training data.

[0018] The solution of the present application can be applied to a cloud computer, which can also be referred to as cloud computing, a server, etc. The cloud computer can interact with the imaging device to obtain the image data collected by the imaging device. The cloud computer can obtain the images collected by the imaging device, analyze them to determine the coordinates of the target object in each image, and then use the three-dimensional image mapping relationship between the images of different imaging devices to verify the coordinates determined by different imaging devices, determine the credibility of the coordinates obtained by the image analysis model, and when the credibility is high, use the corresponding coordinates as the annotation of the image data to form training data. It should be noted that in this solution, an image analysis model can be used to screen the collected image data and the obtained coordinates (remove low-quality analysis results), so as to use more accurate coordinates as annotation data to form high-quality training data.

[0019] Specifically, the cloud computer in this solution can obtain image data captured by at least two imaging devices for photographing a target object. The at least two imaging devices include a first imaging device and a second imaging device. The image data includes first image data captured by the first imaging device and second image data captured by the second imaging device. After obtaining the first image data and the second image data, the first image data and the second image data can be input into an image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data. Both the first coordinate information and the second coordinate information include position information and depth information. Subsequently, the spatial position relationship between the first imaging device and the second imaging device can be obtained, and the stereoscopic image mapping relationship between the first imaging device and the second imaging device can be determined based on the spatial position relationship, so as to verify the first coordinate information and the second coordinate information according to the stereoscopic image mapping relationship and determine the credibility of the first coordinate information and the second coordinate information. When the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data for model training based on the training data. The image analysis model in this solution can be trained based on the training data, and other models can also be trained.

[0020] Among them, in the first image data and the second image data captured at the same time, the position of the target object in the three-dimensional space should be fixed. Therefore, the analysis results of the target object in the two images should be consistent. Thus, this solution can use the spatial position relationship between the first imaging device and the second imaging device to determine the stereoscopic image mapping relationship between the first imaging device and the second imaging device, so as to perform mapping conversion on the first coordinate information according to the stereoscopic image mapping relationship, determine the mapped coordinate information, and analyze whether the mapped coordinate information is the same as the second coordinate information. If they are the same or the difference is small, it can be determined that the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold. Thus, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data. The first imaging device and the second imaging device in this solution can be imaging devices corresponding to a certain road section, so as to collect images of the road and vehicles for analysis to determine training data. This solution does not require users to manually annotate images, which is convenient for generating training data.

[0021] In addition, in this solution, the imaging devices (the first imaging device and the second imaging device) may include a first optical component for long-distance shooting, a second optical component for close-up shooting, and an image sensor for receiving information from the first optical component and the second optical component. By projecting the information of the two optical components onto the same image sensor, when performing conversion, information conversion can be carried out based on the same stereo image mapping relationship, making data processing simpler and more convenient. Moreover, the data corresponding to the first optical component of the first imaging device can be matched with the data corresponding to the first optical component of the second imaging and acquisition device, and the data corresponding to the second optical component of the first imaging device can be matched with the data corresponding to the second optical component of the second imaging and acquisition device, thereby performing two verifications to more accurately determine the credibility of coordinate analysis. It is also possible to match the data corresponding to the first optical component of the first imaging device with the data corresponding to the second optical component of the first imaging device, and further verify the analyzed coordinates to improve the credibility of the training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0023] Figure 1 is a flowchart of a method for stereo image analysis and recognition based on a cloud computer according to an embodiment of the present application;

[0024] Figure 2 is a schematic diagram of the steps of a method for stereo image analysis and recognition based on a cloud computer according to an embodiment of the present application;

[0025] Figure 3 is a schematic structural diagram of a system for stereo image analysis and recognition based on a cloud computer according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] The solution of the present application can be applied to a cloud computer, such as Figure 1As shown, a cloud computer can also be referred to as cloud computing, a server, etc. The cloud computer can interact with a camera device to obtain image data collected by the camera device. The camera device includes a first camera device and a second camera device. The image data includes first image data captured by the first camera device and second image data captured by the second camera device. The first camera device and the second camera device can upload the collected image data to the cloud computer for cloud computing. The camera device (the first camera device and the second camera device) can include a first optical component for long-distance shooting, a second optical component for close-up shooting, and an image sensor for receiving information from the first optical component and the second optical component. The shooting angles of the first optical component and the second optical component can be the same or different, and can be specifically set according to requirements. The focal lengths of the first optical component and the second optical component can be different, and the focal length relationship between the first optical component and the second optical component is preset.

[0028] The cloud computer can obtain images collected by different camera devices, analyze them, and determine the coordinates of the target object in each image. Then, using the stereo image mapping relationship between the images of different camera devices, the coordinates determined by different camera devices are verified to determine the credibility of the coordinates obtained by the image analysis model. When the credibility is relatively high, the corresponding coordinates are used as the annotation of the image data to form training data. When the credibility is relatively low, the corresponding data is discarded. It should be noted that in this solution, an image analysis model can be used to screen the collected image data and the analyzed coordinates (remove low-quality analysis results), so as to use more accurate coordinates as annotation data to form high-quality training data.

[0029] Specifically, the cloud computer in this solution can obtain image data captured by at least two imaging devices for photographing a target object. The at least two imaging devices include a first imaging device and a second imaging device. The image data includes first image data captured by the first imaging device and second image data captured by the second imaging device. After obtaining the first image data and the second image data, the first image data and the second image data can be input into an image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data. Both the first coordinate information and the second coordinate information include position information and depth information. Then, the spatial position relationship between the first imaging device and the second imaging device can be obtained, and the stereoscopic image mapping relationship between the first imaging device and the second imaging device can be determined based on the spatial position relationship, so as to verify the first coordinate information and the second coordinate information according to the stereoscopic image mapping relationship and determine the credibility of the first coordinate information and the second coordinate information. When the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data for model training based on the training data. The image analysis model in this solution can be trained based on the training data, and other models can also be trained.

[0030] Among them, in the first image data and the second image data captured at the same time, the position of the target object in the three-dimensional space should be fixed. Therefore, the analysis results of the target object in the two images should be consistent. Thus, this solution can use the spatial position relationship between the first imaging device and the second imaging device to determine the stereoscopic image mapping relationship between the first imaging device and the second imaging device, so as to perform mapping conversion on the first coordinate information according to the stereoscopic image mapping relationship, determine the mapped coordinate information, and analyze whether the mapped coordinate information is the same as the second coordinate information. If they are the same or the difference is small, it can be determined that the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold, so that the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data. The first imaging device and the second imaging device in this solution can be imaging devices corresponding to a certain road section, so as to collect images of the road and vehicles for analysis and determine training data. This solution does not require users to manually annotate images, which is convenient for generating training data.

[0031] In addition, in this solution, by projecting the information of two optical components onto the same image sensor, when performing conversion, information conversion can be carried out based on the same stereoscopic image mapping relationship, making data processing simpler and more convenient. Moreover, the data corresponding to the first optical component of the first imaging device can be matched with the data corresponding to the first optical component of the second imaging acquisition device, and the data corresponding to the second optical component of the first imaging device can be matched with the data corresponding to the second optical component of the second imaging acquisition device, thereby performing two verifications to more accurately determine the credibility of coordinate analysis. It is also possible to match the data corresponding to the first optical component of the first imaging device with the data corresponding to the second optical component of the first imaging device, and further verify the analyzed coordinates to improve the credibility of the training data.

[0032] Specifically, the embodiment of the present application provides a method for stereoscopic image analysis and recognition based on a cloud computer, as Figure 2 shown, the method includes:

[0033] Step 202, obtain image data captured by at least two imaging devices for capturing a target object, the at least two imaging devices including a first imaging device and a second imaging device, and the image data including first image data captured by the first imaging device and second image data captured by the second imaging device.

[0034] Step 204, input the first image data and the second image data into an image analysis model, and determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data, where both the first coordinate information and the second coordinate information include position information and depth information.

[0035] Step 206, obtain the spatial position relationship between the first imaging device and the second imaging device, and determine the stereoscopic image mapping relationship between the first imaging device and the second imaging device based on the spatial position relationship.

[0036] Step 208, verify the first coordinate information and the second coordinate information based on the stereoscopic image mapping relationship, and determine the credibility of the first coordinate information and the second coordinate information.

[0037] Step 210, when the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, use the first coordinate information as the annotation data for the first image data and the second coordinate information as the annotation data for the second image data to form training data for model training based on the training data.

[0038] The solution of this application can be applied to cloud computers, which can also be referred to as cloud computing, servers, etc. The cloud computer can interact with a camera device to obtain the image data collected by the camera device. The cloud computer can acquire the images captured by the camera device and analyze them to determine the coordinates of the target object in each image. Then, using the stereo image mapping relationship between the images of different camera devices, the coordinates determined by different camera devices are verified to determine the credibility of the coordinates obtained by the image analysis model. When the credibility is relatively high, the corresponding coordinates are used as the annotations of the image data to form training data. It should be noted that in this solution, an image analysis model can be used to screen the collected image data and the analyzed coordinates (remove low-quality analysis results), so as to use more accurate coordinates as annotation data to form high-quality training data.

[0039] Specifically, the cloud computer in this solution can obtain the image data captured by at least two camera devices for photographing the target object. The at least two camera devices include a first camera device and a second camera device. The image data includes the first image data captured by the first camera device and the second image data captured by the second camera device. After obtaining the first image data and the second image data, the first image data and the second image data can be input into the image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data. Both the first coordinate information and the second coordinate information include position information and depth information. Then, the spatial position relationship between the first camera device and the second camera device can be obtained, and the stereo image mapping relationship between the first camera device and the second camera device can be determined based on the spatial position relationship. Based on the stereo image mapping relationship, the first coordinate information and the second coordinate information are verified to determine the credibility of the first coordinate information and the second coordinate information. When the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data for model training based on the training data. The image analysis model in this solution can be trained based on the training data, or other models can be trained.

[0040] Among them, in the first image data and the second image data captured at the same time, the position of the target object in the three-dimensional space should be fixed. Therefore, the analysis results of the target object in the two images should be consistent. Thus, this solution can use the spatial position relationship between the first imaging device and the second imaging device to determine the stereoscopic image mapping relationship between the first imaging device and the second imaging device, so as to perform mapping conversion on the first coordinate information according to the stereoscopic image mapping relationship, determine the mapped coordinate information, and analyze whether the mapped coordinate information is the same as the second coordinate information. If they are the same or the difference is small, it can be determined that the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold. Thus, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data. The first imaging device and the second imaging device in this solution can be imaging devices corresponding to a certain road section, so as to collect images of the road and vehicles, and then analyze to determine the training data. This solution does not require users to manually annotate the images, which is convenient for generating training data.

[0041] In this solution, the imaging device (the first imaging device and the second imaging device) can be composed of two optical components. Specifically, as an optional embodiment, the imaging device includes a first optical component for long-distance shooting, a second optical component for close-up shooting, and an image sensor for receiving the information of the first optical component and the second optical component; the obtaining of the image data captured by at least two imaging devices for shooting the target object includes: determining the first image data according to the acquisition data of the image sensor of the first imaging device; determining the second image data according to the acquisition data of the image sensor of the second imaging device. By projecting the information of the two optical components onto the same image sensor, when performing conversion, information conversion can be performed with the same stereoscopic image mapping relationship, and data processing is simpler and more convenient. The positional relationship between the first optical component and the second optical component can be distributed along the same vertical plane, the position of the second optical component is below the first optical component, and the orientations of the first optical component and the second optical component are both at an angle towards the lower oblique direction for shooting. The vertical planes where the first imaging device and the second imaging device are located can be different, and the captured images of the first imaging device and the second imaging device have an overlapping part. When the target object appears in the overlapping part, the corresponding image is used as the image data to be analyzed.

[0042] Moreover, the images captured by the first optical component and the second optical component (of the first imaging device or the second imaging device) also have an overlapping part, so that when the target object appears in the corresponding area, both the first optical component and the second optical component can collect information. Specifically, as an optional embodiment, determining the first image data based on the acquisition data of the image sensor of the first imaging device, or determining the second image data based on the acquisition data of the image sensor of the second imaging device includes: performing data analysis on the acquisition data of the image sensor of the imaging device to determine whether the target object is within the shooting areas corresponding to the first optical component and the second optical component; when it is recognized that the target object is within the shooting areas corresponding to the first optical component and the second optical component, using the corresponding acquisition data as the image data. It can be determined whether the target object is within the shooting areas corresponding to the first optical component and the second optical component by analyzing whether the target object is included in the images corresponding to the first optical component and the second optical component.

[0043] After determining the coordinates of the target object in the first image data and the second image data, mapping can be performed based on the spatial position relationship between the two imaging devices, so as to determine the credibility of the first coordinate information and the second coordinate information according to whether the coordinate analysis results are consistent (the same or similar). Specifically, as an optional embodiment, verifying the first coordinate information and the second coordinate information based on the stereo image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information includes: adjusting the first coordinate information based on the stereo image mapping relationship to obtain the mapped coordinate information; comparing the mapped coordinate information with the second coordinate information, so as to determine the credibility of the first coordinate information and the second coordinate information according to the comparison result. Among them, in the first image data and the second image data captured at the same time, the position of the target object in the three-dimensional space should be fixed. Therefore, the analysis results of the target object in the two images should be consistent. Thus, this solution can use the spatial position relationship between the first imaging device and the second imaging device to determine the stereo image mapping relationship between the first imaging device and the second imaging device, perform mapping conversion on the first coordinate information based on the stereo image mapping relationship to determine the mapped coordinate information, and analyze whether the mapped coordinate information is the same as the second coordinate information. If they are the same or the difference is small, it can be determined that the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold. Thus, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form the training data.

[0044] The present solution can also split the image data corresponding to a camera device (according to different optical components) to obtain two parts of images, and the target object is located in both parts of the images. Therefore, two coordinates can be determined, and the reliability of the first coordinate information and the second coordinate information can be determined through secondary verification based on the two coordinates of different camera devices. Specifically, as an optional embodiment, the first coordinate information includes a first sub-coordinate and a second sub-coordinate. The first sub-coordinate corresponds to the first optical component of the first camera device, and the second sub-coordinate corresponds to the second optical component of the first camera device. The second coordinate information includes a third sub-coordinate and a fourth sub-coordinate. The third sub-coordinate corresponds to the first optical component of the second camera device, and the fourth sub-coordinate corresponds to the second optical component of the second camera device. Specifically, as an optional embodiment, the process of adjusting the first coordinate information according to the stereo image mapping relationship to obtain the mapped coordinate information can be determined according to the following formula: A = G * X, where A is the mapped coordinate information, G is the mapping matrix representing the stereo image mapping relationship, and X is the first coordinate information; the mapped coordinate information includes a first mapped coordinate corresponding to the first sub-coordinate and a second mapped coordinate corresponding to the second sub-coordinate. Comparing the mapped coordinate information with the second coordinate information to determine the reliability of the first coordinate information and the second coordinate information according to the comparison result includes: comparing the first mapped coordinate with the third sub-coordinate to determine the first comparison result; comparing the second mapped coordinate with the fourth sub-coordinate to determine the second comparison result; and determining the reliability of the first coordinate information and the second coordinate information according to the first comparison result and the second comparison result. The coordinates of the images corresponding to the first optical components of different camera devices can be compared, and the coordinates of the images corresponding to the second optical components of different camera devices can be compared, so as to determine the reliability of the first coordinate information and the second coordinate information according to the two comparison results, improving the quality of the coordinate information.

[0045] The information such as the angles and focal lengths of the first optical component and the second optical component is different. Therefore, the first optical component and the second optical component can be regarded as two different imaging devices, so as to perform the verification of two coordinates of the same imaging device. Specifically, as an optional embodiment, the method further includes: obtaining the first fixed position relationship and the first focal length relationship between the first optical component and the second optical component of the first imaging device to determine the first fixed mapping relationship; determining the third mapping coordinate according to the first sub-coordinate and the first fixed mapping relationship; comparing the third mapping coordinate with the second sub-coordinate to determine the credibility of the first sub-coordinate and the second sub-coordinate, so as to determine the credibility of the first coordinate information; when the credibility of the first coordinate information is lower than the first credibility threshold, discarding the first coordinate information. This solution can compare the first sub-coordinate and the second sub-coordinate to determine whether the results are consistent. If the results are consistent, the next analysis can be carried out. If the results are inconsistent, the corresponding information can be discarded to reduce low-quality data.

[0046] The processing strategy of the second imaging device for coordinate information is the same as that of the first imaging device. Specifically, as an optional embodiment, the method further includes: obtaining the second fixed position relationship and the second focal length relationship between the first optical component and the second optical component of the second imaging device to determine the second fixed mapping relationship; determining the fourth mapping coordinate according to the third sub-coordinate and the second fixed mapping relationship; comparing the fourth mapping coordinate with the fourth sub-coordinate to determine the credibility of the third sub-coordinate and the fourth sub-coordinate, so as to determine the credibility of the second coordinate information; when the credibility of the second coordinate information is lower than the second credibility threshold, discarding the second coordinate information.

[0047] Based on the above embodiments, the embodiment of the present application further provides a stereoscopic image analysis and recognition system based on a cloud computer, as Figure 3 shown, the system includes:

[0048] An image data acquisition module 302, configured to acquire image data captured by at least two imaging devices for photographing a target object. The at least two imaging devices include a first imaging device and a second imaging device, and the image data includes first image data captured by the first imaging device and second image data captured by the second imaging device.

[0049] A coordinate information acquisition module 304, configured to input the first image data and the second image data into an image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data. The first coordinate information and the second coordinate information both include position information and depth information.

[0050] A mapping relationship acquisition module 306, configured to acquire the spatial position relationship between the first imaging device and the second imaging device, and determine the stereo image mapping relationship between the first imaging device and the second imaging device according to the spatial position relationship.

[0051] A coordinate credibility verification module 308, configured to verify the first coordinate information and the second coordinate information according to the stereo image mapping relationship, and determine the credibility of the first coordinate information and the second coordinate information.

[0052] A training data generation module 310, configured to, when the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, use the first coordinate information as the annotation data of the first image data and the second coordinate information as the annotation data of the second image data, so as to form training data for model training according to the training data.

[0053] The implementation manner of the embodiment of the present application is similar to the implementation manner of the above method embodiment. The specific implementation manner can refer to the specific implementation manner of the above method embodiment, which will not be elaborated here.

[0054] The solution of the present application can be applied to a cloud computer, which can also be referred to as cloud computing, a server, etc. The cloud computer can interact with the imaging device to obtain the image data collected by the imaging device. The cloud computer can acquire the images collected by the imaging device, analyze them, and determine the coordinates of the target object in each image. Then, using the stereo image mapping relationship between the images of different imaging devices, the coordinates determined by different imaging devices are verified to determine the credibility of the coordinates analyzed by the image analysis model. When the credibility is relatively high, the corresponding coordinates are used as the annotation of the image data to form training data. It should be noted that in this solution, an image analysis model can be used to screen the collected image data and the analyzed coordinates (remove low-quality analysis results), so as to use more accurate coordinates as annotation data to form high-quality training data.

[0055] Specifically, the cloud computer in this solution can obtain the image data captured by at least two camera devices for photographing a target object. The at least two camera devices include a first camera device and a second camera device. The image data includes the first image data captured by the first camera device and the second image data captured by the second camera device. After obtaining the first image data and the second image data, the first image data and the second image data can be input into an image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data. Both the first coordinate information and the second coordinate information include position information and depth information. Then, the spatial position relationship between the first camera device and the second camera device can be obtained, and the stereo image mapping relationship between the first camera device and the second camera device can be determined based on the spatial position relationship, so as to verify the first coordinate information and the second coordinate information according to the stereo image mapping relationship and determine the credibility of the first coordinate information and the second coordinate information. When the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data for model training based on the training data. The image analysis model in this solution can be trained based on the training data, and other models can also be trained.

[0056] Among them, in the first image data and the second image data captured at the same time, the position of the target object in the three-dimensional space should be fixed. Therefore, the analysis results of the target object in the two images should be consistent. Thus, this solution can use the spatial position relationship between the first camera device and the second camera device to determine the stereo image mapping relationship between the first camera device and the second camera device, so as to perform mapping conversion on the first coordinate information according to the stereo image mapping relationship to determine the mapped coordinate information, and analyze whether the mapped coordinate information is the same as the second coordinate information. If they are the same or the difference is small, it can be determined that the credibility of the first coordinate information and the second coordinate information exceeds the preset credibility threshold. Thus, the first coordinate information is used as the annotation data of the first image data, and the second coordinate information is used as the annotation data of the second image data to form training data. The first camera device and the second camera device in this solution can be camera devices corresponding to a certain road section, so as to collect images of the road and vehicles for analysis to determine training data. This solution does not require users to manually annotate images, which is convenient for generating training data.

[0057] Based on the above embodiments, the present application further provides an electronic device, including: a memory and at least one processor; the memory is used to store computer execution instructions; the at least one processor is used to execute the computer execution instructions stored in the memory, so that the at least one processor executes the method as described in the above embodiments.

[0058] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above data processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, or an optical disc, etc.

[0059] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0061] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0062] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0063] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0064] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0065] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. According to the definition in this article, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0066] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0067] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0068] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A three-dimensional image analysis and recognition method based on cloud computing, characterized in that, The method includes: Obtaining image data captured by at least two imaging devices for photographing a target object, where the at least two imaging devices include a first imaging device and a second imaging device, and the image data includes first image data captured by the first imaging device and second image data captured by the second imaging device; Inputting the first image data and the second image data into an image analysis model to determine first coordinate information of the target object in the first image data and second coordinate information of the target object in the second image data, where both the first coordinate information and the second coordinate information include position information and depth information; Obtaining the spatial position relationship between the first imaging device and the second imaging device, and determining the stereoscopic image mapping relationship between the first imaging device and the second imaging device based on the spatial position relationship; Validating the first coordinate information and the second coordinate information according to the stereoscopic image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information; When the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, using the first coordinate information as the annotation data of the first image data and the second coordinate information as the annotation data of the second image data to form training data for model training based on the training data.

2. The method according to claim 1, characterized in that, The imaging device includes a first optical component for long-distance shooting, a second optical component for close-range shooting, and an image sensor for receiving information from the first optical component and the second optical component; The obtaining of the image data captured by at least two imaging devices for photographing a target object includes: Determining the first image data based on the acquisition data of the image sensor of the first imaging device; Determining the second image data based on the acquisition data of the image sensor of the second imaging device.

3. The method according to claim 2, wherein The determining of the first image data based on the acquisition data of the image sensor of the first imaging device, or the determining of the second image data based on the acquisition data of the image sensor of the second imaging device, includes: Performing data analysis on the acquisition data of the image sensor of the imaging device to determine whether the target object is within the shooting areas corresponding to the first optical component and the second optical component; When it is recognized that the target object is within the shooting areas corresponding to the first optical component and the second optical component, using the corresponding acquisition data as the image data.

4. The method according to claim 3, characterized in that, The validating of the first coordinate information and the second coordinate information according to the stereoscopic image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information includes: Adjusting the first coordinate information according to the stereoscopic image mapping relationship to obtain mapped coordinate information; Comparing the mapped coordinate information with the second coordinate information to determine the credibility of the first coordinate information and the second coordinate information based on the comparison result.

5. The method according to claim 4, characterized in that, The first coordinate information includes a first sub-coordinate and a second sub-coordinate, where the first sub-coordinate corresponds to the first optical component of the first imaging device, the second sub-coordinate corresponds to the second optical component of the first imaging device, the second coordinate information includes a third sub-coordinate and a fourth sub-coordinate, where the third sub-coordinate corresponds to the first optical component of the second imaging device, and the fourth sub-coordinate corresponds to the second optical component of the second imaging device.

6. The method according to claim 5, wherein The process of adjusting the first coordinate information according to the stereoscopic image mapping relationship to obtain the mapped coordinate information can be determined according to the following formula: A = G * X, where A is the mapped coordinate information, G is the mapping matrix representing the stereoscopic image mapping relationship, and X is the first coordinate information; The mapped coordinate information includes the first mapped coordinate corresponding to the first sub-coordinate and the second mapped coordinate corresponding to the second sub-coordinate. Comparing the mapped coordinate information with the second coordinate information to determine the credibility of the first coordinate information and the second coordinate information according to the comparison result includes: Comparing the first mapped coordinate with the third sub-coordinate to determine the first comparison result; Comparing the second mapped coordinate with the fourth sub-coordinate to determine the second comparison result; Determining the credibility of the first coordinate information and the second coordinate information according to the first comparison result and the second comparison result.

7. The method according to claim 6, wherein The method further includes: Obtaining the first fixed position relationship and the first focal length relationship between the first optical component and the second optical component of the first imaging device to determine the first fixed mapping relationship; Determining the third mapped coordinate according to the first sub-coordinate and the first fixed mapping relationship; Comparing the third mapped coordinate with the second sub-coordinate to determine the credibility of the first sub-coordinate and the second sub-coordinate, so as to determine the credibility of the first coordinate information; When the credibility of the first coordinate information is lower than the first credibility threshold, discarding the first coordinate information.

8. The method according to claim 7, wherein The method further includes: Obtaining the second fixed position relationship and the second focal length relationship between the first optical component and the second optical component of the second imaging device to determine the second fixed mapping relationship; Determining the fourth mapped coordinate according to the third sub-coordinate and the second fixed mapping relationship; Comparing the fourth mapped coordinate with the fourth sub-coordinate to determine the credibility of the third sub-coordinate and the fourth sub-coordinate, so as to determine the credibility of the second coordinate information; When the credibility of the second coordinate information is lower than the second credibility threshold, discarding the second coordinate information.

9. A three-dimensional image analysis and recognition system based on a cloud computer, characterized in that, The system includes: An image data acquisition module, configured to acquire image data captured by at least two imaging devices for photographing a target object. The at least two imaging devices include a first imaging device and a second imaging device. The image data includes first image data captured by the first imaging device and second image data captured by the second imaging device; A coordinate information acquisition module, configured to input the first image data and the second image data into an image analysis model to determine the first coordinate information of the target object in the first image data and the second coordinate information of the target object in the second image data. The first coordinate information and the second coordinate information both include position information and depth information; A mapping relationship acquisition module, configured to acquire the spatial position relationship between the first imaging device and the second imaging device, and determine the stereoscopic image mapping relationship between the first imaging device and the second imaging device according to the spatial position relationship; A coordinate credibility verification module, configured to verify the first coordinate information and the second coordinate information according to the stereoscopic image mapping relationship to determine the credibility of the first coordinate information and the second coordinate information; A training data generation module, configured to, when the credibility of the first coordinate information and the second coordinate information exceeds a preset credibility threshold, use the first coordinate information as the annotation data of the first image data and the second coordinate information as the annotation data of the second image data, so as to form training data for model training based on the training data.

10. An electronic device, characterized in that, It includes: a memory and at least one processor; the memory is used to store computer execution instructions; the at least one processor is configured to execute the computer execution instructions stored in the memory, so that the at least one processor executes the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Imaging method, imaging device and imaging system for panoramic three-dimensional image

    CN106990668A

  • Ship scene restoring and positioning method and system based on camera and edge calculation

    CN112907728A