Recognition model training method and apparatus, and mobile intelligent device
By generating images from different angles and uniformly distributing 3D information, the method improves the accuracy of obstacle recognition models by addressing the issue of uneven training data distribution.
Patent Information
- Application Number
- JP2025511429
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-22
- Filing Date
- 2023-03-24
- Publication Date
- 2025-09-02
AI Technical Summary
Existing recognition models trained with unevenly distributed 3D information of obstacles result in inaccurate obstacle recognition due to the uneven distribution of training data, particularly affecting the prediction of azimuth angles.
The method involves generating images from different angles of view by rotating the coordinate system of a first imaging component and determining corresponding three-dimensional information, thereby enriching the data distribution and training a recognition model with more uniformly distributed data.
This approach enhances the accuracy of 3D information prediction by reducing the impact of uneven training data distribution, leading to more precise obstacle recognition.
Smart Images

Figure 2025528892000001_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to Chinese Patent Application No. 202211009619.X, entitled "RECOGNITION MODEL TRAINING METHOD AND APPARATUS, AND MOBILE INTELLIGENT DEVICE," filed with the State Intellectual Property Office of the People's Republic of China on August 22, 2022, the entire contents of which are incorporated herein by reference.
[0002] [Technical field] The present application relates to the field of intelligent driving, and in particular to a perception model training method and apparatus, and a mobile intelligent device. [Background technology]
[0003] Currently, mobile intelligent devices (e.g., intelligent driving vehicles or robotic vacuum cleaners) can acquire environmental information and recognize obstacles by using visual sensors (e.g., cameras). Obstacle recognition based on a monocular camera is a low-cost visual recognition solution. The principle of this solution is that the image captured by the monocular camera is input into a preset recognition model, and three-dimensional (3D) information of each obstacle included in the image, such as 3D coordinates, size, and azimuth angle, can be output by using the recognition model.
[0004] In related solutions, a recognition model is trained by acquiring images and 3D information of obstacles contained in the images. However, since the acquired 3D information of obstacles is unevenly distributed, when obstacle recognition is performed using a recognition model acquired through training based on the unevenly distributed 3D information of obstacles, the acquired 3D information of obstacles is inaccurate. Summary of the Invention
[0005] The present application provides a recognition model training method and apparatus, and a mobile intelligent device, for obtaining more accurate 3D information of obstacles when the recognition model obtained through training is used to perform obstacle recognition.
[0006] To achieve the aforementioned objectives, the present application uses the following technical solutions:
[0007] According to a first aspect, the present application provides a recognition model training method, the method including: acquiring a first image captured by a first photographing component, the first photographing component corresponding to a first angle of view; generating a second image corresponding to a second photographing component based on the first image, the second photographing component being determined based on the first photographing component and the second photographing component corresponding to a second angle of view; determining second three-dimensional information of a first object in the second image based on first three-dimensional information of the first object in the first image; and training a recognition model based on the second image and the second three-dimensional information, the recognition model being used to recognize the object in the image captured by the first photographing component.
[0008] Based on the aforementioned technical solution, an image corresponding to a second imaging component is generated by using an image captured by a first imaging component, and the second imaging component can be determined based on the first imaging component, where the first imaging component corresponds to a first angle of view and the second imaging component corresponds to a second angle of view. In this way, images at different angles of view of the first imaging component can be acquired, thereby enriching the acquired image data. Three-dimensional information of objects in images corresponding to the second imaging component is determined based on three-dimensional information of objects in images captured by the first imaging component, thereby acquiring three-dimensional information of objects in images at different angles of view of the first imaging component, and further acquiring more uniformly distributed three-dimensional information. A recognition model is then trained using the more uniformly distributed three-dimensional information. This reduces the impact of uneven training data distribution on the prediction results of the recognition model. Subsequently, when the recognition model is used to detect objects, the acquired three-dimensional information of the objects is more accurate.
[0009] In one possible design, the second imaging component is a virtual imaging component. Before generating a second image corresponding to the second imaging component based on the first image, the method further includes rotating the coordinate system of the first imaging component by a predetermined angle to acquire the second imaging component. Based on this design, the coordinate system of the first imaging component is rotated by a predetermined angle to acquire the second imaging component. In this manner, a virtual imaging component having a different angle of view from that of the first imaging component is acquired, and the images captured by the first imaging component and three-dimensional information of objects in the images can be converted into images corresponding to the virtual imaging component and three-dimensional information of objects in the images. Therefore, three-dimensional information of objects at different angles of view can be acquired, which results in a more even distribution of acquired training data. This reduces the impact of uneven training data distribution on the trained model.
[0010] In a possible design, the internal parameters of the first imaging component and the second imaging component are the same. Alternatively, the internal parameters of the first imaging component and the second imaging component are different.
[0011] In one possible design, generating a second image corresponding to the second imaging component based on the first image includes determining a first coordinate on a predetermined reference plane of a first pixel in the first image, where the first pixel is an arbitrary pixel in the first image; determining a second coordinate of the first pixel corresponding to the second imaging component based on the first coordinate; determining the second pixel based on the second coordinate; and generating the second image based on the second pixel. Based on this design, the predetermined reference plane is used as a reference. There is a correspondence between the first pixel and a point on the predetermined reference plane, and there is also a correspondence between the second pixel and a point on the predetermined reference plane. Therefore, by using the predetermined reference plane as an intermediary, a pixel corresponding to the second imaging component can be determined based on a pixel in the image captured by the first imaging component, and the image corresponding to the second imaging component is generated based on the second pixel corresponding to the second imaging component.
[0012] In one possible design, the preset reference plane is a spherical surface using the optical center of the first imaging component as the center of the sphere. Based on this design, the preset reference plane is a spherical surface using the optical center of the first imaging component as the center of the sphere. In this way, when the coordinate system of the first imaging component is rotated by a preset angle to acquire the second imaging component, the preset reference plane does not change. The preset reference planes corresponding to the first imaging component and the second imaging component are the same. Therefore, the preset reference plane can be used as a medium for determining pixels in an image corresponding to the second imaging component based on pixels in an image captured by the first imaging component to generate an image corresponding to the second imaging component.
[0013] In one possible design, determining second three-dimensional information of the first object in the second image based on first three-dimensional information of the first object in the first image includes determining the second three-dimensional information based on the first three-dimensional information and a coordinate transformation relationship between the first and second imaging components. Based on this design, the three-dimensional information of the object in the image captured by the first imaging component can be transformed into three-dimensional information of the object in an image corresponding to the second imaging component.
[0014] In one possible design, prior to the step of determining second three-dimensional information of the first object in the second image based on the first three-dimensional information of the first object in the first image, the method further includes the steps of acquiring point cloud data corresponding to the first image and collected by a sensor, the point cloud data including third three-dimensional information of the first object, and determining the first three-dimensional information based on the third three-dimensional information and a coordinate transformation relationship between the first photography component and the sensor. Based on this design, the three-dimensional information of the object collected by the sensor can be converted into three-dimensional information of the object in the image captured by the first photography component.
[0015] In a possible design, the first three-dimensional information, the second three-dimensional information, or the third three-dimensional information includes one or more of a three-dimensional coordinate, a size, and an azimuth angle.
[0016] According to a second aspect, the present application provides a recognition model training apparatus. The recognition model training apparatus includes corresponding modules or units for implementing the aforementioned methods. The modules or units may be implemented by hardware, software, or hardware executing corresponding software. In a possible design, the recognition model training apparatus includes an acquisition unit (also referred to as an acquisition module) and a processing unit (also referred to as a processing module). The acquisition unit is configured to acquire a first image captured by a first photographing component, the first photographing component corresponding to a first angle of view. The processing unit is configured to: generate, based on the first image, a second image corresponding to a second photographing component, the second photographing component being determined based on the first photographing component, the second photographing component corresponding to a second angle of view; determine, based on first three-dimensional information of the first object in the first image, second three-dimensional information of the first object in the second image; and train a recognition model based on the second image and the second three-dimensional information, the recognition model being used to recognize the object in the image captured by the first photographing component.
[0017] In a possible design, the second imaging component is a virtual imaging component. The processing unit is further configured to rotate a coordinate system of the first imaging component by a preset angle to acquire the second imaging component.
[0018] In a possible design, the internal parameters of the first imaging component and the second imaging component are the same. Alternatively, the internal parameters of the first imaging component and the second imaging component are different.
[0019] In a possible design, the processing unit is particularly configured to determine a first coordinate on a predetermined reference plane of a first pixel in a first image, where the first pixel is any pixel in the first image, determine a second coordinate of the first pixel corresponding to a second imaging component based on the first coordinate, determine the second pixel based on the second coordinate, and generate a second image based on the second pixel.
[0020] In a possible design, the preset reference plane is a spherical surface using the optical center of the first imaging component as the sphere center.
[0021] In a possible design, the processing unit is particularly configured to determine the second three-dimensional information based on the first three-dimensional information and a coordinate transformation relationship between the first imaging component and the second imaging component.
[0022] In one possible design, the acquisition unit is further configured to acquire point cloud data corresponding to the first image and collected by the sensor, the point cloud data including third three-dimensional information of the first object, and the processing unit is further configured to determine the first three-dimensional information based on the third three-dimensional information and a coordinate transformation relationship between the first imaging component and the sensor.
[0023] In a possible design, the first three-dimensional information, the second three-dimensional information, or the third three-dimensional information includes one or more of a three-dimensional coordinate, a size, and an azimuth angle.
[0024] According to a third aspect, the present application provides a recognition model training apparatus including a processor, the processor coupled to a memory, the processor configured to execute a computer program stored in the memory to enable the recognition model training apparatus to perform a method according to any one of the first aspect and the design of the first aspect. Optionally, the memory may be coupled to the processor or may be independent of the processor.
[0025] In a possible design, the recognition model training apparatus further includes a communication interface that can be used by the recognition model training apparatus to communicate with another device. For example, the communication interface can be a transceiver, an input / output interface, an interface circuit, an output circuit, an input circuit, a pin, or related circuitry.
[0026] The recognition model training device in the third or fourth aspect may be a computing platform in an intelligent driving system, and the computing platform may be an in-vehicle computing platform or a cloud computing platform.
[0027] According to a fourth aspect, the present application provides an electronic device, the electronic device comprising a recognition model according to either the first aspect or the second aspect and a design of either the first aspect or the second aspect.
[0028] In a possible design, the electronic device further includes a first imaging component according to either the first or second aspect and any one of the designs of the first or second aspect. For example, the first imaging component may be a monocular camera.
[0029] According to a fifth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium including a computer program or instructions, the computer program or instructions, when executed on a recognition model training device, enabling the recognition model training device to perform a method according to any one of the first aspect and the design of the first aspect.
[0030] According to a sixth aspect, the present application provides a computer program product, the computer program product including computer programs or instructions that, when executed on a computer, enable the computer to perform a method according to any one of the first aspect and the design of the first aspect.
[0031] According to a seventh aspect, the present application provides a chip system including at least one processor and at least one interface circuit, the at least one interface circuit configured to perform a transceiver function and transmit instructions to the at least one processor, and when the at least one processor executes the instructions, the at least one processor performs a method according to any one of the first aspect and the design of the first aspect.
[0032] According to an eighth aspect, the present application provides a mobile intelligent device including a first photographing component and a recognition model training apparatus according to any one of the designs of the second aspect or the third aspect and the second aspect or the third aspect, wherein the first photographing component is configured to capture a first image and transmit the first image to the recognition model training apparatus.
[0033] It should be noted that for the technical effects provided by any one of the designs of the second to eighth aspects, please refer to the technical effects provided by the corresponding design of the first aspect, and the details will not be described again in this specification. [Brief explanation of the drawings]
[0034] [Figure 1] FIG. 1 is a schematic diagram of point cloud data according to an embodiment of the present application; [Figure 2] FIG. 1 is a schematic diagram of a monocular camera coordinate system and a lidar coordinate system according to an embodiment of the present application. [Figure 3] 1 is a schematic diagram of visual 3D information according to an embodiment of the present application; [Figure 4]FIG. 2 is a schematic diagram of an azimuth angle of an obstacle according to an embodiment of the present application; [Figure 5] FIG. 1 is a schematic illustration of the distribution of azimuth angles obtained for training a recognition model in a related solution. [Figure 6] 1 is a schematic diagram of the architecture of a system according to an embodiment of the present application; [Figure 7] 1 is a schematic diagram of a structure of a training device according to an embodiment of the present application. [Figure 8] FIG. 1 is a functional block diagram of a vehicle according to an embodiment of the present application. [Figure 9] 1 is a schematic diagram of a camera disposed on a vehicle according to an embodiment of the present application; [Figure 10] 1 is a schematic flowchart of a recognition model training method according to an embodiment of the present application; [Figure 11] FIG. 2 is a schematic diagram of a second imaging component according to an embodiment of the present application. [Figure 12] FIG. 2 is a schematic diagram of a distribution of azimuth angles obtained for training a recognition model according to an embodiment of the present application; [Figure 13] 1 is a schematic flowchart of another recognition model training method according to an embodiment of the present application; [Figure 14] FIG. 2 is a schematic diagram of a preset reference plane according to an embodiment of the present application. [Figure 15] 1 is a schematic diagram of 3D information obtained through prediction by using a recognition model according to an embodiment of the present application; [Figure 16] 1 is a schematic diagram of the structure of a recognition model training apparatus according to an embodiment of the present application; [Figure 17] 1 is a schematic diagram of the structure of a chip system according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION
[0035] In the description of this application, unless otherwise specified, " / " represents an "or" relationship between related objects. For example, A / B may represent A or B. In this application, "and / or" only describes an association relationship for describing related objects and represents that three relationships may exist. For example, A and / or B may represent the following three cases: when only A exists, when both A and B exist, and when only B exists, where A or B may be singular or plural.
[0036] In the description of this application, unless otherwise specified, "a plurality of" means two or more. "At least one of the following items (moieties)" or similar expressions means any combination of these items, including any combination of a single item (moiety) or multiple items (moieties). For example, at least one of a, b, or c can represent a, b, c, a and b, a and c, b and c, and a, b, and c, where a, b, and c may be singular or plural.
[0037] In addition, in order to clearly describe the technical solutions in the embodiments of the present application, terms such as "first" and "second" are used in the embodiments of the present application to distinguish between the same or similar items that basically provide the same function or purpose. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity or execution order, and terms such as "first" and "second" do not indicate clear distinctions.
[0038] In the embodiments of the present application, terms such as "example" or "for example" are used to denote providing an example, illustration, or explanation. Any embodiment or design solution described in the embodiments of the present application as an "example" or "for example" should not be construed as preferred or advantageous over another embodiment or design solution. Rather, the use of terms such as "example" or "for example" is intended to present relevant concepts in a particular way for ease of understanding.
[0039] The features, structures, or characteristics in the present application may be combined in any suitable manner in one or more embodiments. The sequence numbers of the processes described above do not imply an execution order in the embodiments of the present application. The execution order of the processes should be determined based on the functions and internal logic of the processes and should not be construed as any limitation on the implementation process of the embodiments of the present application.
[0040] In some scenarios, some optional features in the embodiments of the present application may be independently implemented without relying on other features to solve corresponding technical problems and achieve corresponding effects. Alternatively, in some scenarios, optional features may be combined with other features based on requirements.
[0041] In this application, unless otherwise specified, the same or similar parts in the embodiments shall refer to each other. In various embodiments of this application, unless otherwise specified or there is no logical contradiction, the terms and / or descriptions in different embodiments shall be consistent and may be mutually referenced, and the technical features in different embodiments may be combined based on their internal logical relationships to form new embodiments. The implementation form of this application is not intended to limit the protection scope of this application.
[0042] In addition, the network architectures and service scenarios described in the embodiments of the present application are intended to more clearly explain the technical solutions in the embodiments of the present application, and do not constitute limitations on the technical solutions provided in the embodiments of the present application. Those skilled in the art may know that with the evolution of network architectures and the emergence of new service scenarios, the technical solutions provided in the embodiments of the present application can also be applied to similar technical problems.
[0043] For ease of understanding, the following first describes relevant technical terms and concepts that may be used in the embodiments of the present application.
[0044] 1. Monocular camera
[0045] A monocular camera is a camera that has only one lens and cannot directly measure depth information.
[0046] 2. Bird's-eye view
[0047] A bird's-eye view is a three-dimensional illustration drawn using high-point perspective, based on the principles of perspective. That is, a bird's-eye view is an image obtained by looking down from a high position.
[0048] 3. Recognition Model
[0049] A recognition model is an information processing system that includes a large number of interconnected processing units (e.g., neurons), and each processing unit in the recognition model includes a corresponding mathematical formula. After data is input to a processing unit, the processing unit executes the mathematical formula included in the processing unit to calculate the input data and generate output data. The input data of each processing unit is the output data of the previous processing unit connected to it, and the output data of each processing unit is the input data of the next processing unit connected to it.
[0050] In the recognition model, after data is input, the recognition model selects a corresponding processing unit for the input data based on the learning and training of the recognition model, performs calculations on the input data by using the processing unit, and determines and outputs the final calculation result. In addition, the recognition model can also continuously learn and evolve in the data calculation process, and continuously optimize the calculation process of the recognition model based on the feedback of the calculation results. The more times the recognition model calculates and trains, the more result feedback it obtains, and the more accurate the calculation result.
[0051] The recorded recognition model in an embodiment of the present application is used to process images captured by the photography component to determine three-dimensional information of objects in the images, including, but not limited to, people and various types of obstacles, such as objects.
[0052] 4. Point cloud data
[0053] When a laser beam is shone on the surface of an object, the reflected laser light carries information such as orientation and distance. When the laser beam scans based on a specific trajectory, information about the reflected laser points is recorded as the laser beam scans. Because the scan is extremely fine, a large number of laser points can be acquired, and each laser point includes 3D coordinates. Therefore, a laser point cloud can be formed, and these laser point clouds may be referred to as point cloud data. In some embodiments, these laser point clouds can be acquired through laser scanning using a lidar. Of course, point cloud data can alternatively be acquired using another type of sensor. This application is not limited in this respect.
[0054] Optionally, the point cloud data may include 3D information of one or more objects, such as 3D coordinates, size, and azimuth. For example, FIG. 1 shows point cloud data according to one embodiment of the present application. Each point in FIG. 1 is a laser point, and position B on the circle is the position of the sensor that acquires the point cloud data. The 3D box with an arrow (e.g., 3D box A) is a 3D model established based on the 3D information of the scanned object. The coordinates of the 3D box are the 3D coordinates of the object. The length, width, and height of the 3D box are the size of the object. The direction of the arrow on the 3D box is the orientation of the object, etc.
[0055] 5. Monocular camera coordinate system
[0056] The coordinate system of the monocular camera may be a three-dimensional coordinate system using the optical center of the monocular camera as the origin. For example, (1) in FIG. 2 is a schematic diagram of the coordinate system of the monocular camera according to an embodiment of the present application. As shown in (1) in FIG. 2, the coordinate system of the monocular camera may be a spatial Cartesian coordinate system using the optical center o of the monocular camera as the origin. The line of sight of the monocular camera is the positive direction of the z-axis, the downward direction is the positive direction of the y-axis, and the rightward direction is the positive direction of the x-axis. Optionally, the y-axis and x-axis may be parallel to the imaging plane of the monocular camera.
[0057] 6. Lidar coordinate system
[0058] The coordinate system of the lidar may be a three-dimensional coordinate system using the laser emission center as the origin. For example, (2) in FIG. 2 is a schematic diagram of the coordinate system of the lidar according to an embodiment of the present application. As shown in (2) in FIG. 2, the coordinate system of the lidar may be a spatial Cartesian coordinate system using the laser emission center o of the lidar as the origin. In the coordinate system of the lidar, the upward direction is the positive direction of the z-axis, the forward direction is the positive direction of the x-axis, and the leftward direction is the positive direction of the y-axis.
[0059] It should be understood that the monocular camera coordinate system shown in (1) of Figure 2 and the lidar coordinate system shown in (2) of Figure 2 are merely examples for the purpose of explanation. The monocular camera coordinate system and the lidar coordinate system may alternatively be set in a different manner or in a different type of coordinate system. This is not a limitation of the present application.
[0060] Currently, a mobile intelligent device (using an autonomous vehicle as an example) can acquire environmental information and recognize obstacles by using a visual sensor (e.g., a camera). Obstacle recognition based on a monocular camera is a low-cost visual recognition solution. The principle of this solution is that an image captured by a monocular camera is input into a preset recognition model, and 3D information in the coordinate system of the monocular camera of each obstacle included in the image can be output by using the recognition model.
[0061] For example, Fig. 3(1) is a schematic diagram of visualizing 3D information output by a recognition model on an image captured by a monocular camera according to an embodiment of the present application. Each 3D box (e.g., 3D box C) shown in Fig. 3(1) is drawn based on 3D information of an obstacle on which the 3D box is located. Fig. 3(2) is a schematic diagram of visualizing 3D information output by a recognition model on a bird's-eye view according to an embodiment of the present application. Similarly, each 3D box (e.g., 3D box D) shown in Fig. 3(2) is drawn based on 3D information of an obstacle on which the 3D box is located.
[0062] Currently, the aforementioned recognition model can be trained by acquiring 3D information of an image and an obstacle in the image. The 3D information of the obstacle in the image can be acquired based on point cloud data. Specifically, the point cloud data corresponding to the image is acquired by using a lidar, and the point cloud data includes 3D information of the obstacle (the obstacle is an obstacle in the image) in the coordinate system of the lidar. Then, extrinsic parameters between the lidar and the monocular camera can be calibrated to convert the 3D information in the coordinate system of the lidar into 3D information in the coordinate system of the monocular camera. Thereafter, the recognition model can be trained by using the acquired image and the 3D information of the obstacle in the converted image in the coordinate system of the monocular camera. It can be understood that the extrinsic parameters refer to the transformation relationship between different coordinate systems. Calibrating the extrinsic parameters between the lidar and the monocular camera is to solve the transformation relationship between the coordinate system of the lidar (the coordinate system shown in (2) of FIG. 2) and the coordinate system of the monocular camera (the coordinate system shown in (1) of FIG. 2).
[0063] In this solution, the acquired 3D information of the obstacle used to train the recognition model is heterogeneous. For example, the 3D information of the obstacle is an azimuth angle, which is usually obtained by summing θ and α. For example, referring to the monocular camera coordinate system shown in (1) of FIG. 2, FIG. 4 is a schematic diagram of θ and α according to one embodiment of the present application. θ is the horizontal included angle between the connecting line between the center point M of the obstacle and the origin o of the monocular camera coordinate system and the z coordinate of the monocular camera coordinate system, where the direction of the arrow N is the orientation of the obstacle. In the monocular camera coordinate system, when the obstacle is rotated around the y axis of the monocular camera coordinate system to the z axis, for example, when the obstacle is moved from position 1 to position 2, α represents the included angle between the orientation of the obstacle (i.e., the direction of the arrow N shown at position 2) and the direction of the x axis, where the origin o of the coordinate system is used as the center, and the connecting line from the origin o to the center point M of the obstacle is used as the radius.
[0064] For example, Figure 5 is a schematic diagram of the distribution of θ and α obtained in this solution. As shown in Figure 5, the horizontal coordinate represents θ, and the vertical coordinate represents α. From Figure 5, it can be seen that θ and α are distributed unevenly. For example, when θ is 0 degrees, α is densely distributed in areas such as -180 degrees, -90 degrees, 0 degrees, 90 degrees, and 180 degrees, while α is sparsely distributed in areas other than the above areas.
[0065] Both θ and α can be obtained by prediction using a recognition model obtained through training. However, in general, θ is predicted roughly accurately, while α is easily affected by uneven training data. Therefore, if a recognition model is trained using unevenly distributed 3D information of an obstacle obtained, the recognition model obtained through training will inaccurately predict α due to the uneven distribution of the training data. For example, since the sparsely distributed α is predicted as the densely distributed α in FIG. 5, the predicted azimuth angle of the obstacle is inaccurate, i.e., the predicted 3D information of the obstacle is inaccurate.
[0066] In response to this, the present application provides a recognition model training method for reducing the influence of uneven distribution of training data on the prediction results of the recognition model, so that when the recognition model obtained through training is used to recognize an obstacle, the obtained 3D information of the obstacle is more accurate.
[0067] For example, Figure 6 is a schematic diagram of the architecture of a system according to one embodiment of the present application. As shown in Figure 6, the system 60 includes a training device 10, an execution device 20, and the like.
[0068] The training device 10 may execute a recognition model to train the recognition model. For example, the training device 10 may be a computing platform in an intelligent driving system. For example, the training device 10 may be various computing devices having computing capabilities, such as a server or a cloud device. The server may be a single server or a server cluster including multiple servers. Alternatively, the training device 10 may be an in-vehicle computing platform. The specific type of the training device 10 is not limited in this application.
[0069] The recognition models trained by the training device 10 are configured on the execution device 20, and the execution device 20 uses the recognition models to recognize various objects, etc. For example, the objects include, but are not limited to, various types of objects, such as pedestrians, vehicles, road signs, animals, buildings, etc. In some embodiments, the recognition models trained by the training device 10 may be configured on multiple execution devices 20, and each execution device 20 may recognize various obstacles, etc. by using the recognition models configured on the execution device 20. In some embodiments, each execution device 20 may be configured with one or more recognition models trained by the training device 10, and each execution device 20 may recognize various obstacles by using the one or more recognition models configured on the execution device 20.
[0070] For example, the execution device 20 may be a device having obstacle recognition requirements, such as an artificial intelligence (AI) device such as a robotic vacuum cleaner or an intelligent vehicle, or may be a device having the capability to control the aforementioned devices having obstacle recognition requirements, such as a desktop computer, a handheld computer, a notebook computer, or an ultra-mobile personal computer (UMPC).
[0071] For example, in an autonomous driving scenario, while an autonomous vehicle is traveling along a predetermined route, the recognition model is used to recognize road signs, driving reference objects, obstacles on the road, etc. in the environment to ensure that the autonomous vehicle travels safely and accurately. Road signs may include graphic or text road signs. Driving reference objects may be buildings or plants. Obstacles on the road may include dynamic objects (e.g., animals, pedestrians, or moving vehicles) or static objects (e.g., stationary vehicles).
[0072] In one possible example, the training device 10 and the execution device 20 are different processors located on different physical devices (eg, a server or servers in a cluster).
[0073] For example, execution device 20 may be a neural network processing unit (NPU), a graphics processing unit (GPU), a central processing unit (CPU), another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0074] The training device 10 may be a GPU, NPU, microprocessor, ASIC, or one or more integrated circuits configured to control program execution of the solutions in the present application.
[0075] It should be understood that Figure 6 is only a simplified exemplary diagram for ease of understanding. In actual application, the above-mentioned system may further include other devices not shown.
[0076] For example, the training device 10 in this embodiment of the present application may be implemented by using different devices. For example, the training device 10 in this embodiment of the present application may be implemented by using the communication device of FIG. 7. FIG. 7 is a schematic diagram of the hardware structure of the training device 10 according to one embodiment of the present application. The training device 10 includes at least one processor 701, a communication line 702, a memory 703, and at least one communication interface 704.
[0077] The processor 701 may be a general-purpose CPU, a microprocessor, an ASIC, or one or more integrated circuits configured to control program execution in the solution of the present application.
[0078] Communication lines 702 may include paths for conveying information between the aforementioned components.
[0079] The communication interface 704 is configured to communicate with another device. In this embodiment of the present application, the communication interface 704 may be a module, a circuit, a bus, an interface, a transceiver, or another device capable of implementing a communication function. Optionally, when the communication interface is a transceiver, the transceiver may be an independently disposed transmitter, and the transmitter may be configured to transmit information to another device. Alternatively, the transceiver may be an independently disposed receiver, and configured to receive information from another device. Alternatively, the transceiver may be a component that integrates the functions of transmitting and receiving information. The specific implementation form of the transceiver is not limited in the embodiments of the present application.
[0080] Memory 703 may be read-only memory (ROM) or another type of static storage device capable of storing static information and instructions, random access memory (RAM) or another type of dynamic storage device capable of storing information and instructions, or may be electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or another compact disc storage device, optical disc storage device (including compact optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disc storage media or another magnetic storage device, or any other medium usable to carry or store expected program code in the form of instructions or data structures and accessible by a computer. Memory 703 may exist independently or be connected to processor 701 via communication line 702. Alternatively, memory 703 may be integrated with processor 701.
[0081] The memory 703 is configured to store computer-executable instructions used to implement the solutions of the present application. The processor 701 is configured to execute the computer-executable instructions stored in the memory 703 to perform the methods provided in the following embodiments of the present application.
[0082] Optionally, the computer-executable instructions in the embodiments of the present application may be referred to as application code, instructions, computer programs, or other names, without being particularly limited in the embodiments of the present application.
[0083] In a specific implementation, in one embodiment, the processor 701 may include one or more CPUs, for example, CPU0 and CPU1 of FIG.
[0084] In a specific implementation, in one embodiment, training device 10 may include multiple processors, such as processor 701 and processor 705 of FIG. 7. Each of the processors may be a single-CPU processor or a multi-CPU processor. A processor herein may be one or more devices, circuits, and / or processing cores configured to process data (e.g., computer program instructions).
[0085] The training device 10 may be a general-purpose device or a specialized device, and the type of training device 10 is not limited in this embodiment of the present application.
[0086] It may be understood that the structure shown in this embodiment of the present application does not constitute a specific limitation on the training device 10. In some other embodiments of the present application, the training device 10 may include more or fewer components than those shown in the figures, some components may be combined, some components may be split, or there may be a different component layout. The components shown in the figures may be implemented by hardware, software, or a combination of software and hardware.
[0087] For example, the execution device 20 is an autonomous vehicle. Figure 8 is a functional block diagram of the vehicle 100. In some embodiments, a trained recognition model is configured in the vehicle 100. During the driving process, the vehicle 100 may recognize objects in the environment by using the recognition model to ensure that the vehicle 100 drives safely and accurately. For example, the objects include, but are not limited to, road signs, driving reference objects, obstacles on the road, etc.
[0088] Vehicle 100 may include various subsystems, such as a traction system 110, a sensor system 120, a control system 130, one or more peripheral devices 140, a power source 150, a computer system 160, and a user interface 170. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple elements. Additionally, all of the subsystems and elements of vehicle 100 may be interconnected in a wired or wireless manner.
[0089] The traction system 110 includes components that provide power to the vehicle 100 for propulsion. In one embodiment, the traction system 110 may include an engine 111, a transmission 112, an energy source 113, and wheels 114. The engine 111 may be an internal combustion engine, an electric motor, an air-compression engine, or a combination of other types of engines, such as a hybrid engine including a gasoline engine and an electric motor, or a hybrid engine including an internal combustion engine and an air-compression engine. The engine 111 converts the energy source 113 into mechanical energy.
[0090] Examples of energy source 113 include gasoline, diesel, another oil-based fuel, propane, another compressed gas-based fuel, ethanol, solar panels, batteries, and other power sources. Energy source 113 may also provide energy to other systems of vehicle 100.
[0091] The transmission 112 may transmit mechanical power from the engine 111 to the wheels 114. The transmission 112 may include a gearbox, a differential, and a drive shaft. In one embodiment, the transmission 112 may further include other components, such as a clutch. The drive shaft may include one or more shafts that may be coupled to one or more wheels 114.
[0092] Sensor system 120 may include several sensors that sense information about the surrounding environment of vehicle 100. For example, sensor system 120 may include a positioning system 121 (the positioning system may be a global positioning system (GPS), a BeiDou system, or another positioning system), an inertial measurement unit (IMU) 122, radar 123, lidar 124, and a camera 125. Sensor data from one or more of these sensors may be used to detect objects and corresponding characteristics of the objects (such as position, shape, orientation, speed, etc.). Such detection and recognition are important functions for the safe operation of vehicle 100.
[0093] Positioning system 121 may be configured to estimate the geographic position of vehicle 100. IMU 122 is configured to sense changes in the position and orientation of vehicle 100 based on inertial acceleration. In one embodiment, IMU 122 may be a combination of an accelerometer and a gyroscope.
[0094] Radar 123 may sense objects in the environment surrounding vehicle 100 by using radio signals. In some embodiments, in addition to sensing objects, radar 123 may be further configured to sense the speed and / or direction of travel of the objects.
[0095] The lidar 124 may use laser light to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the lidar 124 may include one or more laser sources, a laser scanner, one or more detectors, and other system components.
[0096] The camera 125 may be configured to capture multiple images of the environment surrounding the vehicle 100 and multiple images within the vehicle cockpit. The camera 125 may be a still camera or a video camera. In some embodiments of the present application, the camera 125 may be a monocular camera. For example, the monocular camera may include, but is not limited to, a long-focus camera, a medium-focus camera, a short-focus camera, a fisheye camera, etc. For example, FIG. 9 is a schematic diagram of the camera 125 disposed on the vehicle 100 according to one embodiment of the present application. For example, the camera 125 may be disposed in various positions, such as at the head, rear, sides, and top of the vehicle.
[0097] The control system 130 may control the operation of the vehicle 100 and components of the vehicle 100. The control system 130 may include various elements such as a steering system 131, a throttle 132, a braking unit 133, a computer vision system 134, a route control system 135, and an obstacle avoidance system 136.
[0098] The steering system 131 may be operated to adjust the direction of travel of the vehicle 100. For example, in one embodiment, the steering system 131 may be a steering wheel system.
[0099] The throttle 132 is configured to control the operating speed of the engine 111 , which in turn controls the speed of the vehicle 100 .
[0100] The braking unit 133 is configured to control and decelerate the vehicle 100 .
[0101] The computer vision system 134 may be operated to process and analyze images captured by the camera 125 to recognize objects and / or features in the vehicle 100's environment, as well as the physical and facial features of the driver in the vehicle cockpit. The objects and / or features may include traffic signals, road conditions, and obstacles, and the physical and facial features of the driver include the driver's behavior, gaze, facial expression, etc.
[0102] Route control system 135 is configured to determine a driving path for vehicle 100. In some embodiments, route control system 135 may determine a driving path for vehicle 100 by reference to data from sensors, positioning system 121, and one or more predetermined maps.
[0103] The obstacle avoidance system 136 is configured to recognize, evaluate, and avoid or otherwise circumvent potential obstacles in the environment of the vehicle 100.
[0104] Of course, in one example, control system 130 may include additional or alternative components other than those shown and described. Alternatively, some of the above components may be omitted.
[0105] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or a user by using peripheral devices 140. Peripheral devices 140 may include a wireless communication system 141, an onboard computer 142, a microphone 143, and / or a speaker 144. In some embodiments, peripheral devices 140 provide a means for a user of vehicle 100 to interact with user interface 170.
[0106] The wireless communication system 141 may communicate wirelessly with one or more devices directly or via a communication network.
[0107] The power source 150 may provide power to various components of the vehicle 100 .
[0108] Some or all of the functions of vehicle 100 are controlled by computer system 160. Computer system 160 may include at least one processor 161, which executes instructions 1621 stored, for example, in data storage device 162. Computer system 160 may alternatively be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner.
[0109] Processor 161 may be any conventional processor, such as a CPU, an ASIC, or another dedicated device of a hardware-based processor. While FIG. 8 functionally depicts the processor, data storage, and other elements within the same physical enclosure, those skilled in the art should understand that a processor, computer system, or data storage device may actually include multiple processors, computer systems, or data storage devices that may be stored within the same physical enclosure or multiple processors, computer systems, or data storage devices that may be stored in different physical enclosures. For example, a data storage device may be a hard disk drive or another storage medium located in a different physical enclosure. Thus, reference to a processor or computer system should be understood to include reference to a set of processors or computer systems or data storage devices that may operate in parallel, or to a set of processors or computer systems or data storage devices that may not operate in parallel. Unlike using a single processor to perform the steps described herein, some components, such as the steering component and the deceleration component, may each have their own processor, with the processor performing only calculations related to the component's specific function.
[0110] In various aspects described herein, the processor may be located remotely from the vehicle and communicate wirelessly with the vehicle. In other aspects, some processes described herein are executed on a processor located inside the vehicle, while other processes, including performing the steps necessary for a single operation, are executed by a remote processor.
[0111] In some embodiments, data storage 162 may include instructions 1621 (e.g., program logic) that may be executed by processor 161 to perform various functions of vehicle 100, including those described above. Data storage 162 may further include additional instructions, including instructions for transmitting data to, receiving data from, interacting with, and / or controlling one or more of driving system 110, sensor system 120, control system 130, and peripheral devices 140.
[0112] In addition to instructions 1621, data storage device 162 may further store data such as road maps, route information, and vehicle position, direction, speed, and other such vehicle data, as well as other information. Such information may be used by vehicle 100 and computer system 160 when vehicle 100 operates in autonomous, semi-autonomous, and / or manual modes. Optionally, in some embodiments of the present application, data storage device 162 stores trained recognition models.
[0113] For example, in a possible embodiment, the data storage device 162 may be located within the surrounding environment and may acquire images captured by the vehicle based on the camera 125 within the sensor 120, and may acquire 3D information of each obstacle in the image by using the stored recognition model.
[0114] User interface 170 is configured to provide information to or receive information from a user of vehicle 100. Optionally, user interface 170 may interact with one or more input / output devices in set of peripheral devices 140, such as one or more of wireless communication system 141, onboard computer 142, microphone 143, and speaker 144.
[0115] The computer system 160 may control the vehicle 100 based on information obtained from various subsystems (e.g., the driving system 110, the sensor system 120, and the control system 130) and information received from the user interface 170.
[0116] Optionally, one or more of the aforementioned components may be mounted separately from or associated with vehicle 100. For example, data storage device 162 may be partially or completely separate from vehicle 100. The aforementioned components may be coupled to one another to communicate in a wired and / or wireless manner.
[0117] Optionally, the above components are only examples. In actual applications, components in the above modules may be added or removed according to actual requirements. Figure 8 should not be construed as a limitation on this embodiment of the present application.
[0118] The vehicle 100 may be a car, truck, motorcycle, bus, boat, airplane, helicopter, lawn mower, recreational vehicle, playground vehicle, construction equipment, tram, golf cart, train, trolley, etc. This is not particularly limited in the embodiments of the present application.
[0119] In some other embodiments of the present application, the autonomous vehicle may further include a hardware structure and / or a software module, and may implement the functions in the form of a hardware structure, a software module, or both a hardware structure and a software module. Whether the functions in the aforementioned functions are performed by using a hardware structure, a software module, or a combination of a hardware structure and a software module depends on the specific application and design constraints of the technical solution.
[0120] For example, Fig. 10 shows a recognition model training method according to one embodiment of the present application. The method may be performed by the training device 10 shown in Fig. 6 or by a processor of the training device 10, for example, the processor shown in Fig. 7. The method includes the following steps:
[0121] S1001: A first image captured by a first photography component is obtained.
[0122] The first photographing component corresponds to a first angle of view, i.e., the first image is an image corresponding to the first angle of view. For example, the first photographing component may be a monocular camera, including, but not limited to, a long-focus camera, a medium-focus camera, a short-focus camera, a fisheye camera, etc. Optionally, the first photographing camera may alternatively be another device having an image capture function. In this embodiment of the present application, the first photographing component may alternatively be referred to as an image capture device, a sensor, etc. It may be understood that the name of the first photographing component does not constitute a limitation on the function of the first photographing component.
[0123] S1002: A second image corresponding to a second imaging component is generated based on the first image.
[0124] The second shooting component may be determined based on the first shooting component, and the second shooting component corresponds to a second angle of view, i.e., the second image is an image corresponding to the second angle of view. The first angle of view is different from the second angle of view, i.e., the shooting angle corresponding to the first shooting component is different from the shooting angle corresponding to the second shooting component. In other words, the first image is different from the second image.
[0125] In a possible implementation, the first photography component is a virtual photography component. Before step S1002, the coordinate system of the first photography component may be rotated by a preset angle to obtain a second photography component. In other words, the coordinate system of the first photography component is rotated by a preset angle to obtain a new coordinate system, and the photography component corresponding to the new coordinate system is used as the second photography component.
[0126] Optionally, the extrinsic parameters between the second imaging component and the first imaging component may be determined based on one or more of a preset rotation angle, a rotation direction, etc. The extrinsic parameters may be understood to be a transformation relationship between two coordinate systems. Therefore, the extrinsic parameters between the second imaging component and the first imaging component may also be described as a coordinate transformation relationship (or referred to as a coordinate system transformation relationship) between the second imaging component and the first imaging component.
[0127] In some embodiments, the external parameter between the second imaging component and the first imaging component is a rotation matrix related to a preset angle. In the following, the external parameter between the second imaging component and the first imaging component will be described using an example in which the coordinate system of the first imaging component is the coordinate system shown in (1) of Figure 2.
[0128] In a possible example, the coordinate system of the first imaging component is rotated by a preset angle around the x-axis of the coordinate system of the first imaging component to acquire the second imaging component. In this case, the external parameters between the second imaging component and the first imaging component can be expressed as Equation 1.
number
[0129] In Equation 1, X represents an external parameter between the second imaging component and the first imaging component, and ω represents a preset rotation angle, where 0 degrees≦ω≦360 degrees.
[0130] Optionally, there may be errors in the parameters of Equation 1 (e.g., one or more of cos(ω), sin(ω), and −sin(ω)). Thus, in this example, the external parameters between the second imaging component and the first imaging component may alternatively be expressed as Equation 1a:
number
[0131] In Equation 1a, v and u are both preset constants and can be used to eliminate errors in the corresponding parameters. It can be understood that the constants used to eliminate the errors in the parameters can be the same, for example, v and u can be the same, or different, for example, v and u can be different. Of course, the overall error in Equation 1 can alternatively be eliminated by using an error matrix. For example, the error matrix can be a 3x3 matrix, and each element of the matrix is a preset constant. In other words, the expression format of the preset constant used to eliminate the error in Equation 1 is not limited in this application, and the position of the preset constant added to Equation 1 is not limited.
[0132] In another possible example, the coordinate system of the first imaging component is rotated by a preset angle around the y-axis of the coordinate system of the first imaging component to acquire the second imaging component. In this case, the external parameters between the second imaging component and the first imaging component can be expressed as Equation 2.
number
[0133] For an explanation of the parameters in Equation 2, please refer to the explanation of the corresponding parameters in Equation 1.
[0134] In yet another possible example, the coordinate system of the first imaging component is rotated by a preset angle around the z-axis of the first imaging component to acquire the second imaging component. In this case, the external parameters between the second imaging component and the first imaging component can be expressed as Equation 3:
number
[0135] For an explanation of the parameters in Equation 3, please refer to the explanation of the corresponding parameters in Equation 1.
[0136] In the above example, it can be understood that the coordinate system of the first imaging component is rotated around one coordinate axis in the coordinate system of the first imaging component to acquire the second imaging component. Alternatively, the coordinate system of the first imaging component may be rotated around multiple coordinate axes in the coordinate system of the first imaging component to acquire the second imaging component. For example, first, rotate a specific angle around the x-axis, then rotate a specific angle around the y-axis. In another example, first, rotate a specific angle around the z-axis, then rotate a specific angle around the x-axis. In another example, first, rotate a specific angle around the y-axis, then rotate a specific angle around the x-axis, and finally rotate a specific angle around the z-axis. Optionally, rotation around a coordinate axis may be performed one or more times, for example, first, rotate a specific angle around the x-axis, then rotate a specific angle around the y-axis, and then rotate a specific angle around the x-axis. Optionally, the angles of the three rotations may be the same or different.
[0137] When the coordinate system of the first imaging component is rotated around multiple coordinate axes in the coordinate system of the first imaging component to acquire the second imaging component, the external parameters between the second imaging component and the first imaging component can be expressed by multiplying multiple of the above-mentioned Equation 1, Equation 2, and Equation 3. For example, first rotate by a specific angle around the x-axis, and then rotate by a specific angle around the y-axis. The external parameters between the second imaging component and the first imaging component can be expressed by multiplying the above Equation 1 and the above Equation 2, and can be expressed as, for example, Equation 4.
number
[0138] In Equation 4, X represents an external parameter between the second imaging component and the first imaging component. β represents a rotation angle around the x-axis, and 0 degrees≦β≦360 degrees. λ represents a rotation angle around the y-axis, and 0 degrees≦λ≦360 degrees. β and λ may be the same or different. Optionally, a preset angle (i.e., the aforementioned α) may be determined based on the rotation angle around each coordinate axis. For example, in this example, the preset angle may be determined based on β and λ. In some examples, the preset angle may be the sum of β and λ.
[0139] Similarly, when the second imaging component is acquired by one or more rotations around multiple other coordinate axes, the expression of the external parameters between the second imaging component and the first imaging component can be referred to the above-mentioned implementation, and the details will not be described again in this specification.
[0140] Similarly, the parameter errors of Equation 2, Equation 3, and Equation 4 can alternatively be eliminated by adding a preset constant. For specific implementation, please refer to the related implementation of Equation 1a. Examples will not be described again in this specification.
[0141] Optionally, the coordinate system of the first imaging component is rotated around another direction to acquire the second imaging component, instead of rotating around the coordinate axis of the coordinate system of the first imaging component. In other words, in this embodiment of the present application, the coordinate system of the first imaging component may be rotated in any direction by any angle to acquire the second imaging component. This is not limited in the present application.
[0142] For example, referring to the coordinate system of the monocular camera shown in (1) of Figure 2, the coordinate system of the first photography component 1100 may be shown in (1) of Figure 11. For example, the coordinate system shown in (1) of Figure 11 is rotated left by N degrees to obtain a new coordinate system, where N is a positive number. For example, the new coordinate system may be the coordinate system shown in (2) of Figure 11, and the photography component corresponding to the coordinate system shown in (2) of Figure 11 (the photography component 1110 shown in (2) of Figure 11) is the second photography component.
[0143] Optionally, the manner in which the coordinate system of the second imaging component is set may be the same as the manner in which the coordinate system of the first imaging component is set. For example, the optical center of the imaging component is used as the origin of the coordinate system, and the line of sight direction of the imaging component is used as the z-axis direction. Alternatively, the manner in which the coordinate system of the second imaging component is set may be different from the manner in which the coordinate system of the first imaging component is set. For example, the origin of the coordinate system is different, and the coordinate axes are set differently. This is not a limitation of the present application.
[0144] In another possible implementation, the coordinate system of the first imaging component and the coordinate system of the second imaging component may be the same type of coordinate system, for example, both may be spatial Cartesian coordinate systems, or may be different types of coordinate systems.
[0145] In some possible implementations, the intrinsics of the first and second photography components are the same. For example, the intrinsics of the first photography component are directly used as the intrinsics of the second photography component. Preferably, the intrinsics of the first and second photography components are different. For example, different intrinsics are set for the second photography component than for the first photography component. Optionally, the distortion coefficients of the first and second photography components may be the same or different. For a detailed description of the intrinsics and distortion coefficients of photography components, please refer to the description of the intrinsics and distortion coefficients of cameras in the prior art. The details will not be described again in this application.
[0146] S1003: Determine second three-dimensional information of the first object in the second image based on the first three-dimensional information of the first object in the first image.
[0147] Optionally, the first three-dimensional information of the first object in the first image is three-dimensional information of the first object in a coordinate system of the first imaging component, and the second three-dimensional information of the first object in the second image is three-dimensional information of the first object in a coordinate system of the second imaging component.
[0148] Optionally, the second three-dimensional information of the first object may be determined based on the first three-dimensional information of the first object and a coordinate transformation relationship between the first and second photography components. It may be understood that the coordinate transformation relationship between the first and second photography components may be described as an external parameter between the first and second photography components.
[0149] Optionally, before step S1003, the method shown in FIG. 10 may further include the following steps S1003a and S1003b (not shown).
[0150] S1003a: Obtain point cloud data corresponding to a first image and collected by a sensor.
[0151] The point cloud data includes third three-dimensional information of the first object. Optionally, the third three-dimensional information of the first object is three-dimensional information in a coordinate system of a sensor. For example, the sensor may be a sensor configured to collect point cloud data, such as a lidar.
[0152] Optionally, in this embodiment of the present application, the first three-dimensional information, the second three-dimensional information, or the third three-dimensional information includes one or more of three-dimensional coordinates, size, azimuth angle, and the like.
[0153] S1003b: Determine first three-dimensional information based on the third three-dimensional information and the coordinate transformation relationship between the first imaging component and the sensor.
[0154] It may be understood that the coordinate transformation relationship between the first imaging component and the sensor may also be described as an extrinsic parameter between the first imaging component and the sensor.
[0155] S1004: A recognition model is trained based on the second image and the second three-dimensional information.
[0156] The recognition model is used to recognize objects in images captured by the first photography component.
[0157] Optionally, in the training process, the recognition model may further project the second three-dimensional information of the first object onto the pixels of the second image based on the internal parameters, distortion coefficients, etc. of the second photographing component, to train the recognition model. For the description of the imaging principle in this specification, please refer to the description in the prior art. The details will not be described again in this application.
[0158] For example, the second three-dimensional information is an azimuth angle. Referring to the above definition of the azimuth angle, the azimuth angle is usually obtained by summing θ and α. For example, FIG. 12 is a diagram of the distribution of the obtained θ and α of an object according to one embodiment of the present application. As shown in FIG. 12, the horizontal coordinate represents θ, and the vertical coordinate represents α. It can be seen from FIG. 12 that θ and α are distributed unevenly.
[0159] Based on the aforementioned technical solution, an image corresponding to a second imaging component is generated by using an image captured by a first imaging component, and the second imaging component can be determined based on the first imaging component. In this way, images at different angles of view of the first imaging component can be acquired, thereby enriching the acquired image data. Three-dimensional information of objects in images corresponding to the second imaging component is determined based on three-dimensional information of objects in images captured by the first imaging component, thereby acquiring three-dimensional information of objects in images at different angles of view of the first imaging component, and further acquiring more uniformly distributed three-dimensional information. A recognition model is then trained using the more uniformly distributed three-dimensional information. This reduces the impact of uneven training data distribution on the prediction results of the recognition model. Subsequently, when the recognition model is used to detect objects, the acquired three-dimensional information of the objects becomes more accurate.
[0160] In some embodiments, when there are multiple first imaging components, each first imaging component may correspond to one recognition model. For example, for each first imaging component, the recognition model training device may determine a second imaging component corresponding to the first imaging component based on the first imaging component, and train a recognition model corresponding to the first imaging component. Then, the recognition model corresponding to each first imaging component is used to recognize only objects in images captured by the corresponding first imaging component.
[0161] In some other embodiments, when multiple first filming components are present, the multiple first filming components may correspond to one recognition model. For example, for multiple first filming components, the model training device may determine a corresponding second filming component based on the multiple first filming components and train a recognition model corresponding to the multiple first filming components. Subsequently, the recognition model obtained through training may be used to recognize objects in images captured by the multiple first filming components corresponding to the recognition model. For example, the recognition model training device may use the method described in FIG. 10 to generate images corresponding to the second filming component based on images captured by the multiple first filming components, where the second filming component is determined based on the multiple first filming components. Similarly, second three-dimensional information of objects in images corresponding to the second filming component is separately determined based on the first three-dimensional information of objects in images captured by the multiple first filming components. The recognition model is trained based on images corresponding to the second filming component and the second three-dimensional information of objects in images corresponding to the second filming component.
[0162] Based on the solution in this embodiment, both the images captured by each first imaging component and the first three-dimensional information of the object in the image can be used to train the same recognition model. This increases the training data for training the recognition model, improves data utilization, and reduces the cost of collecting training data. In addition, training one recognition model can be applied to multiple first imaging components. In other words, one training model can recognize objects in images captured by multiple first imaging components. In this way, there is no need to train a recognition model for each first imaging component, thereby reducing training costs.
[0163] Optionally, in a possible implementation, step S1002 shown in FIG. 10 may be specifically implemented as steps S1005 to S1008 shown in FIG.
[0164] S1005: A first coordinate on a preset reference plane of a first pixel in a first image is determined.
[0165] Optionally, the preset reference plane may be a spherical surface using the optical center of the first imaging component as the sphere center.
[0166] For example, Fig. 14 is a schematic diagram of a preset reference plane according to an embodiment of the present application. As shown in Fig. 14 (1), the preset reference plane may be a spherical surface K using the optical center of the first imaging component as its center, i.e., a spherical surface K using the origin o of the coordinate system of the first imaging component as its center.
[0167] Optionally, the preset reference plane may alternatively be set in another manner, for example, a spherical surface using a point at a preset distance from the optical center of the first imaging component as the center of the sphere, or a spherical surface using another position of the first imaging component as the center of the sphere. The specific method for setting the preset reference plane is not limited in this application. Optionally, the preset reference plane may be a flat surface or a curved surface. This is also not limited in this application.
[0168] It can be understood that there is a mapping relationship between a point in the coordinate system of the first imaging component and a pixel in the first image. Using point P in the coordinate system of the first imaging component as an example, a pixel in the first image corresponding to point P can be determined as point P1 based on the imaging principle, internal parameters, distortion coefficients, etc. of the first imaging component.
[0169] Point P is normalized to sphere K. As shown in (1) of FIG. 14, a point normalized from point P to sphere K is defined as point P'. Since there is a correspondence between point P and point P1 and between point P and point P', there is also a correspondence between point P' and point P1. It should be understood that any point in the coordinate system of the first imaging component can be normalized to a corresponding point on sphere K. For different points in the coordinate system of the first imaging component, the points normalized to sphere K may be the same. The points normalized to sphere K correspond one-to-one to pixels in the first image. For example, the first pixel is point P1. In this case, the first coordinate of the first pixel on the preset reference plane is the coordinate of point P' in the coordinate system of the first imaging component, i.e., the coordinate of point P' in the coordinate system shown in (1) of FIG. 14.
[0170] Optionally, the first pixel may be any pixel in the first image, and for each pixel in the first image, a first coordinate on a preset reference plane may be determined based on this step. Optionally, the first pixel may alternatively be a pixel obtained by performing a difference process on pixels in the first image. The first pixel is not limited in this application. The difference process is a conventional technique. For details, please refer to existing solutions. The details will not be described in this specification.
[0171] S1006: Based on the first coordinate, a second coordinate of the first pixel corresponding to a second imaging component is determined.
[0172] Referring to the implementation of FIG. 11 , the second imaging component can be acquired by rotating the coordinate system of the first imaging component by a predetermined angle. Therefore, when the second imaging component is acquired by rotating the coordinate system of the first imaging component by a predetermined angle, the sphere K shown in FIG. 14(1) is also rotated by a predetermined angle to acquire the sphere K1 shown in FIG. 14(2). However, in the camera coordinate system of the first imaging component shown in FIG. 14(1), the position of point P does not change even when it is normalized to a point on the sphere K1 shown in FIG. 14(2). However, because the coordinate system of the second imaging component shown in FIG. 14(2) is different from the coordinate system of the first imaging component shown in FIG. 14(1), the coordinate of point P′ shown in FIG. 14(2) in the coordinate system of the second imaging component (the coordinate system shown in FIG. 14(2)) is a second coordinate.
[0173] In some embodiments, a second coordinate of the first pixel corresponding to the second imaging component may be determined based on the first coordinate and an external parameter between the second imaging component and the first imaging component.
[0174] In one possible implementation, for example, the external parameter between the second imaging component and the first imaging component is a rotation matrix related to a preset angle, and the second coordinate system corresponding to the first pixel and the second imaging component can be obtained by left-multiplying the first coordinate by the rotation matrix.
[0175] For example, if the rotation matrix is the matrix of Equation 1, the second coordinate of the first pixel corresponding to the second imaging component can be expressed as Equation 5.
number
[0176] In Equation 5,
number
number
number
[0177] Similarly, the parameter error in Equation 5 may alternatively be removed by adding a preset constant. For specific implementation, please refer to the related implementation of Equation 1a. Examples will not be described again in this specification.
[0178] Similarly, if the rotation matrix is in another format (e.g., the matrix of Equation 2 to Equation 4), please refer to the implementation of Equation 5 for the representation of the second coordinate of the first pixel corresponding to the second imaging component. An example will not be described again in this specification.
[0179] S1007: A second pixel is determined based on the second coordinate.
[0180] Similarly, there is a correspondence between points on the spherical surface K1 and pixels in the second image corresponding to the second imaging component. Based on the imaging principle, points on the spherical surface K1 can be mapped to pixels in the second image. For example, the second coordinates are the coordinates of point P' shown in (2) of FIG. 14 in the coordinate system shown in (2) of FIG. 14. In this case, the second pixel is P2.
[0181] It can be understood that for each first pixel in the first image, the corresponding second pixel can be determined based on the solution in this embodiment of the present application.
[0182] S1008: A second image is generated based on the second pixels.
[0183] For example, a second image can be generated based on the second pixels by calling a remap function in an open source computer vision library (OpenCV).
[0184] For descriptions of other steps in FIG. 13, please refer to the descriptions of the corresponding steps in FIG.
[0185] For example, Fig. 15 (1) is a schematic diagram of an object orientation obtained through prediction by using a recognition model trained in an existing solution, and Fig. 15 (2) is a schematic diagram of an object orientation obtained by using a recognition model obtained through training according to one embodiment of the present application. The vehicles shown in Fig. 15 are objects, and the white rectangular boxes of each vehicle are 3D boxes marked with the vehicle's 3D information predicted based on the recognition model, and the direction of the white arrow is the direction of the vehicle head, i.e., the vehicle's orientation. From Fig. 15, it can be seen that in this embodiment of the present application, the accuracy of the object orientation obtained by using the recognition model obtained through training is higher than the accuracy of the corresponding object orientation predicted by using a model trained in the existing solution.
[0186] The above mainly describes the solutions provided in the embodiments of the present application from the perspective of methods. It can be understood that, to implement the aforementioned functions, the recognition model training apparatus includes a hardware structure and / or software modules for performing corresponding functions. With reference to the units and algorithm steps described in the embodiments disclosed in the present application, the embodiments of the present application can be implemented in the form of hardware or hardware and computer software. Whether the functions are performed by hardware or hardware driven by a computer depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solutions in the embodiments of the present application.
[0187] In an embodiment of the present application, the recognition model training device may be divided into functional modules based on the above-described method embodiment. For example, each functional module may be obtained through division based on each corresponding function, or two or more functions may be integrated into one processing unit. The integrated unit may be implemented in the form of hardware or in the form of a software functional module. It should be noted that in this embodiment of the present application, the division into units is merely an example and represents a logical division of functions. In actual implementation, other division methods may be used.
[0188] 16 is a schematic diagram of the structure of a recognition model training apparatus according to an embodiment of the present application. The recognition model training apparatus 1600 may be configured to implement the method described in the above method embodiment. Optionally, the recognition model training apparatus may be the training device 10 shown in FIG. 6, or may be a module (e.g., a chip) used in a server. The server may be located on a cloud. For example, the recognition model training apparatus 1600 may specifically include an acquiring unit 1601 and a processing unit 1602.
[0189] The acquiring unit 1601 is configured to support the recognition model training apparatus 1600 in performing step S1001 of Fig. 10. Additionally / alternatively, the acquiring unit 1601 is configured to support the recognition model training apparatus 1600 in performing step S1001 of Fig. 13. Additionally / alternatively, the acquiring unit 1601 is further configured to support the recognition model training apparatus 1600 in performing other steps performed by the recognition model training apparatus in this embodiment of the present application.
[0190] Processing unit 1602 is configured to support recognition model training apparatus 1600 in performing steps S1002 to S1004 of Fig. 10. Additionally / alternatively, processing unit 1602 is further configured to support recognition model training apparatus 1600 in performing steps S1005 to S1008 and steps S1003 to S1004 of Fig. 13. Additionally / alternatively, processing unit 1602 is further configured to support recognition model training apparatus 1600 in performing other steps performed by the recognition model training apparatus in this embodiment of the present application.
[0191] Optionally, the recognition model training apparatus 1600 shown in Fig. 16 may further include a communication unit 1603. The communication unit 1603 is configured to support the recognition model training apparatus 1600 in performing steps of communication between the recognition model training apparatus and another device in this embodiment of the present application.
[0192] Optionally, the recognition model training apparatus 1600 shown in Figure 16 may further include a storage unit (not shown in Figure 16) for storing programs or instructions. When the processing unit 1602 executes the programs or instructions, the recognition model training apparatus 1600 shown in Figure 16 can perform the methods in the above-described method embodiments.
[0193] For technical effects of the recognition model training apparatus 1600 shown in Fig. 16, please refer to the technical effects described in the above-mentioned method embodiments. Details will not be described again in this specification. The processing unit 1602 in the recognition model training apparatus 1600 shown in Fig. 16 may be implemented by a processor or a processor-related circuit component, and may be a processor or a processing module. The communication unit 1603 may be implemented by a transceiver or a transceiver-related circuit component, and may be a transceiver or a transceiver module.
[0194] An embodiment of the present application further provides a chip system. As shown in FIG. 17, the chip system includes at least one processor 1701 and at least one interface circuit 1702. The processor 1701 and the interface circuit 1702 may be interconnected using a line. For example, the interface circuit 1702 may be configured to receive a signal from another device. In another example, the interface circuit 1702 may be configured to transmit a signal to another device (e.g., the processor 1701). For example, the interface circuit 1702 may read instructions stored in a memory and transmit the instructions to the processor 1701. When the instructions are executed by the processor 1701, the recognition model training device may be able to perform the steps performed by the recognition model training device in the above-described embodiment. Of course, the chip system may further include another discrete device. This is not particularly limited in the embodiments of the present application.
[0195] Optionally, there may be one or more processors in the chip system. The processor may be implemented using hardware or software. When the processor is implemented using hardware, the processor may be a logic circuit, an integrated circuit, etc. When the processor is implemented using software, the processor may be a general-purpose processor and is implemented by reading software code stored in a memory.
[0196] Optionally, there may be one or more memories in the chip system. The memory may be integrated with the processor or may be located separately from the processor. This is not limited in this application. For example, the memory may be a non-transitory processor such as a read-only memory (ROM). The memory and the processor may be integrated on the same chip or may be located separately on different chips. The type of memory and the way in which the memory and the processor are located are not particularly limited in this application.
[0197] For example, the chip system may be a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a system on a chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or another integrated chip.
[0198] It should be understood that the steps in the foregoing method embodiments may be completed by using instructions in the form of integrated logic circuits of hardware in a processor or software. The steps of the methods disclosed with reference to the embodiments of the present application may be directly executed by a hardware processor, or may be executed through a combination of hardware and software modules in a processor.
[0199] An embodiment of the present application further provides an electronic device, which includes the recognition model in the above-mentioned method embodiment. Optionally, the electronic device may be the above-mentioned execution device.
[0200] In a possible design, the electronic device further includes the first imaging component in the method embodiments described above.
[0201] An embodiment of the present application further provides a computer storage medium, which stores computer instructions that, when executed on a recognition model training device, enable the recognition model training device to perform the method in the above-described method embodiment.
[0202] One embodiment of the present application provides a computer program product, which includes computer programs or instructions, which, when executed on a computer, enable the computer to perform the method in the above-mentioned method embodiments.
[0203] An embodiment of the present application provides a mobile intelligent device, including a first photographing component and the recognition model training apparatus described in the above embodiment, where the first photographing component is configured to capture a first image and send the first image to the recognition model training apparatus.
[0204] In addition, an embodiment of the present application further provides an apparatus. The apparatus may specifically be a chip, a component, or a module. The apparatus may include a processor and a memory connected to each other. The memory is configured to store computer-executable instructions. When the apparatus operates, the processor may execute the computer-executable instructions stored in the memory to enable the apparatus to perform the method in the above-mentioned method embodiments.
[0205] The recognition model training apparatus, mobile intelligent device, computer storage medium, computer program product, or chip provided in the embodiments is configured to execute the corresponding method provided above. Therefore, for the beneficial effects that can be achieved by the recognition model training apparatus, mobile intelligent device, computer storage medium, computer program product, or chip, please refer to the beneficial effects of the corresponding method provided above. The details will not be described again in this specification.
[0206] Based on the above description of the implementation form, those skilled in the art can understand that for convenient and concise description, the above division into functional modules is only used as an example for illustration. In actual application, the above functions can be allocated to different functional modules for implementation according to requirements, that is, the internal structure of the device is divided into different functional modules to implement all or part of the above-described functions.
[0207] In some embodiments provided in this application, it should be understood that the disclosed devices and methods may be implemented in other ways. The embodiments may be combined with or referenced to each other without inconsistency. The described device embodiments are merely examples. For example, the division into modules or units is merely a logical division of function, and other divisions may be used in actual implementation. For example, multiple units or components may be combined or integrated into another device, and alternatively, some features may be ignored or not implemented. In addition, the shown or described mutual couplings or direct couplings or communication connections may be implemented via some interfaces. Indirect couplings or communication connections between devices or units may be implemented electronically, mechanically, or in other forms.
[0208] The units described as separate parts may or may not be physically separate, and the parts shown as units may be one or more physical units, located in one place or distributed in different places. Some or all of the units may be selected according to actual requirements to achieve the objectives of the solutions of the embodiments.
[0209] In addition, the functional units in the embodiments of the present application may be integrated into one processing unit, each of the units may exist physically alone, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0210] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, the integrated unit may be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application may essentially be implemented in the form of a software product, or the portion that contributes to the prior art, or all or part of the technical solutions. The software product may be stored in a storage medium and include instructions for instructing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to perform all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes any medium capable of storing program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0211] The above description is merely a specific implementation form of the present application and does not limit the protection scope of the present application. Any variations or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application shall fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for training a recognition model, comprising: obtaining a first image captured by a first photography component, the first photography component corresponding to a first angle of view; generating a second image corresponding to a second photography component based on the first image, the second photography component being determined based on the first photography component, and the second photography component corresponding to a second angle of view; determining second three-dimensional information of the first object in the second image based on first three-dimensional information of the first object in the first image; training a recognition model based on the second image and the second three-dimensional information, the recognition model being used to recognize objects in images captured by the first photography component; A method comprising:
2. The second photography component is a virtual photography component, and prior to the step of generating a second image corresponding to the second photography component based on the first image, the method further comprises: Rotating the coordinate system of the first imaging component by a preset angle to acquire the second imaging component. The method of claim 1 further comprising:
3. the internal parameters of the first imaging component and the second imaging component are the same; or the internal parameters of the first imaging component and the second imaging component are different; 3. The method according to claim 1 or 2.
4. The step of generating a second image corresponding to a second imaging component based on the first image includes: determining a first coordinate on a predetermined reference plane of a first pixel in the first image, the first pixel being an arbitrary pixel in the first image; determining a second coordinate of the first pixel corresponding to the second imaging component based on the first coordinate; determining a second pixel based on the second coordinate; generating the second image based on the second pixels; 4. The method of claim 1, comprising:
5. The method of claim 4 , wherein the preset reference plane is a spherical surface using the optical center of the first imaging component as the sphere center.
6. The step of determining second three-dimensional information of the first object in the second image based on first three-dimensional information of the first object in the first image comprises: determining the second three-dimensional information based on the first three-dimensional information and a coordinate transformation relationship between the first imaging component and the second imaging component; 6. The method of claim 1, comprising:
7. Prior to the step of determining second three-dimensional information of the first object in the second image based on first three-dimensional information of the first object in the first image, the method further comprises: acquiring point cloud data corresponding to the first image and collected by a sensor, the point cloud data including third three-dimensional information of the first object; determining the first three-dimensional information based on the third three-dimensional information and a coordinate transformation relationship between the first imaging component and the sensor; 7. The method of claim 1, further comprising:
8. A recognition model training device comprising an acquisition unit and a processing unit, The acquisition unit is configured to acquire a first image captured by a first photography component, the first photography component corresponding to a first angle of view; The processing unit generating a second image corresponding to a second photography component based on the first image, the second photography component being determined based on the first photography component, and the second photography component corresponding to a second angle of view; determining second three-dimensional information of the first object in the second image based on first three-dimensional information of the first object in the first image; training a recognition model based on the second image and the second three-dimensional information, the recognition model being used to recognize objects in images captured by the first photography component; and An apparatus configured to:
9. 9. The apparatus of claim 8, wherein the second imaging component is a virtual imaging component, and the processing unit is further configured to rotate a coordinate system of the first imaging component by a preset angle to acquire the second imaging component.
10. the internal parameters of the first imaging component and the second imaging component are the same; or the internal parameters of the first imaging component and the second imaging component are different; 10. The device according to claim 8 or 9.
11. The processing unit determining a first coordinate on a predetermined reference plane of a first pixel in the first image, the first pixel being an arbitrary pixel in the first image; determining a second coordinate of the first pixel corresponding to the second imaging component based on the first coordinate; determining a second pixel based on the second coordinate; generating the second image based on the second pixels; 11. An apparatus according to any one of claims 8 to 10, specifically configured to:
12. The apparatus of claim 11 , wherein the preset reference plane is a spherical surface using the optical center of the first imaging component as its center.
13. the processing unit is particularly configured to determine the second three-dimensional information based on the first three-dimensional information and a coordinate transformation relationship between the first imaging component and the second imaging component.
13. Apparatus according to any one of claims 8 to 12.
14. the acquisition unit is further configured to acquire point cloud data corresponding to the first image and collected by a sensor, the point cloud data including third three-dimensional information of the first object; the processing unit is further configured to determine the first three-dimensional information based on the third three-dimensional information and a coordinate transformation relationship between the first imaging component and the sensor.
14. Apparatus according to any one of claims 8 to 13.
15. 8. A recognition model training apparatus comprising a processor coupled to a memory, the processor configured to execute a computer program stored in the memory to enable the recognition model training apparatus to perform a method according to any one of claims 1 to 7.
16. 16. A mobile intelligent device comprising: a first photography component; and a recognition model training apparatus according to any one of claims 8 to 15, wherein the first photography component is configured to capture a first image and transmit the first image to the recognition model training apparatus.
Citation Information
Patent Citations
Encoding apparatus and encoding method, and decoding apparatus and decoding method
JP2018182755A
Training dataset generation for depth measurement
US20220165027A1