Transforming sensor data to train models for use with different sensor configurations
By transforming sensor data using machine learning and computer graphics techniques, the problem of inaccurate model output under different sensor configurations was solved, enabling accurate identification and application of sensor data under various configurations.
Patent Information
- Application Number
- CN202111532910.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-08
- Filing Date
- 2021-12-15
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In the prior art, due to differences in sensor configuration, a model trained using one vehicle sensor configuration may produce inaccurate outputs under another vehicle sensor configuration, resulting in poor application performance of the model under different sensor configurations.
By using machine learning and computer graphics techniques, sensor data is transformed from one configuration to another, separating the object of interest from the background, and the model is trained under the new configuration to achieve data augmentation and adaptation.
It enables accurate identification and application of sensor data under different configurations, and improves the output performance of the model under different sensor configurations.
Smart Images

Figure CN115035236B_ABST
Abstract
Description
Background Technology
[0001] The information provided in this section is for the purpose of presenting the general background of this disclosure. To the extent described in this section, the work of the currently named inventors and aspects of the description that may not constitute prior art at the time of filing are neither explicitly nor implicitly considered to be prior art of this disclosure.
[0002] This disclosure relates to transforming sensor data from a given sensor configuration to any other reference frame to train models that can be used with different sensor configurations.
[0003] In many applications, data collected by sensors is used to train models (e.g., machine learning-based models). In use, the trained model receives data from sensors and outputs data trained on the data received from the sensors. For example, in automotive applications (e.g., autonomous driving applications), data collected by various sensors (e.g., cameras) mounted on the vehicle is used to train the model. Sensors collect data as the vehicle travels on the road. The collected data is used to train the model. The trained model is deployed in the vehicle. In use, the trained model receives data from sensors and outputs data that the model was trained to produce. Summary of the Invention
[0004] A system includes a processor and a memory storing instructions, which, when executed by the processor, configure the processor to receive first data from a first set of sensors arranged in a first configuration. The instructions configure the processor to transform the first data into second data to train a model to recognize third data captured by a second set of sensors arranged in a second configuration, different from the first configuration. The instructions configure the processor to train the model based on sensing the second data by the second set of sensors to recognize the third data captured by the second set of sensors arranged in the second configuration.
[0005] In another feature, the trained model recognizes the third data captured by the second set of sensors arranged in the second configuration.
[0006] In another feature, at least one of the second group of sensors is different from at least one of the first group of sensors.
[0007] Among other features, the instructions configure the processor to detect one or more objects in the first data and to separate the objects in the first data from the background.
[0008] Among other features, the instructions configure the processor to use a machine learning-based model to transform the perspective of the object from 2D to 3D and to use computer graphics techniques to transform the perspective of the background from 2D to 3D.
[0009] In another feature, the instructions configure the processor to combine the transformed viewpoint of the object and the transformed viewpoint of the background to generate a 3D scene representing the first data.
[0010] In another feature, the instructions configure the processor to train the model based on the 3D scene sensing the first data from the second set of sensors.
[0011] Among other features, the instructions configure the processor to generate a 2D representation of the object from the 3D perspective sensed by the second set of sensors and to generate a 2D representation of the background from the 3D perspective sensed by the second set of sensors.
[0012] In another feature, the instructions configure the processor to combine the 2D representation of the object from the 3D perspective sensed by the second set of sensors with the 2D representation of the background from the 3D perspective.
[0013] In another feature, the instructions configure the processor to train the model based on a combination of the 2D representation of the object from the 3D perspective sensed by the second set of sensors and the 2D representation of the background from the 3D perspective.
[0014] Among other features, a method includes receiving first data from a first set of sensors arranged in a first configuration. The method includes transforming the first data into second data, the second data reflecting the first data as perceived by a second set of sensors arranged in a second configuration. The second configuration differs from the first configuration. The method includes training a model by sensing the second data using the second set of sensors to identify third data captured by the second set of sensors arranged in the second configuration.
[0015] In another feature, the method further includes using a trained model to identify the third data captured by the second set of sensors arranged in the second configuration.
[0016] In another feature, at least one of the second group of sensors is different from at least one of the first group of sensors.
[0017] Among other features, the method further includes: detecting one or more objects in the first data; and separating the objects in the first data from the background.
[0018] Among other features, the method also includes: transforming the perspective of the object from 2D to 3D using a machine learning-based model; and transforming the perspective of the background from 2D to 3D using computer graphics techniques.
[0019] In another feature, the method further includes: combining the transformed viewpoint of the object and the transformed viewpoint of the background to generate a 3D scene representing the first data.
[0020] In another feature, the method further includes training the model based on the 3D scene representing the first data sensed by the second set of sensors.
[0021] Among other features, the method further includes: generating a 2D representation of the 3D perspective of the object sensed by the second set of sensors; and generating a 2D representation of the 3D perspective of the background sensed by the second set of sensors.
[0022] In another feature, the method further includes: combining the 2D representation of the object from the 3D perspective sensed by the second set of sensors with the 2D representation of the background from the 3D perspective.
[0023] In another feature, the method further includes training the model based on a combination of a 2D representation of the object from a 3D perspective sensed by the second set of sensors and a 2D representation of the background from a 3D perspective.
[0024] The present invention also includes the following solutions:
[0025] Option 1. A system comprising:
[0026] Processor; and
[0027] Memory, which stores instructions that, when executed by the processor, configure the processor to:
[0028] Receive first data from a first group of sensors arranged in a first configuration;
[0029] The first data is transformed into second data to train a model to identify third data captured by a second set of sensors arranged in a second configuration, wherein the second configuration differs from the first configuration; and
[0030] The model is trained based on the second set of sensors sensing the second data to identify the third data captured by the second set of sensors arranged in the second configuration.
[0031] Option 2. The system according to Option 1, wherein the trained model identifies the third data captured by the second set of sensors arranged in the second configuration.
[0032] Option 3. The system according to Option 1, wherein at least one of the second group of sensors is different from at least one of the first group of sensors.
[0033] Option 4. The system according to Option 1, wherein the instructions configure the processor to:
[0034] Detect one or more objects in the first data; and
[0035] Separate the object from the background in the first data.
[0036] Option 5. The system according to Option 4, wherein the instructions configure the processor to:
[0037] The object's perspective is transformed from 2D to 3D using a machine learning-based model; and
[0038] Computer graphics technology is used to transform the perspective of the background from 2D to 3D.
[0039] Solution 6. The system according to Solution 5, wherein the instructions configure the processor to combine the transformed viewpoint of the object and the transformed viewpoint of the background to generate a 3D scene representing the first data.
[0040] Option 7. The system according to Option 6, wherein the instructions configure the processor to train the model based on a 3D scene representing the first data sensed by the second set of sensors.
[0041] Option 8. The system according to Option 5, wherein the instructions configure the processor to:
[0042] Generate a 2D representation of the object from the 3D perspective sensed by the second set of sensors; and
[0043] Generate a 2D representation of the 3D perspective of the background sensed by the second set of sensors.
[0044] Option 9. The system according to Option 8, wherein the instructions configure the processor to combine the 2D representation of the object from the 3D perspective sensed by the second set of sensors with the 2D representation of the background from the 3D perspective.
[0045] Option 10. The system according to Option 9, wherein the instructions configure the processor to train the model based on a combination of the 2D representation of the object from the 3D perspective sensed by the second set of sensors and the 2D representation of the background from the 3D perspective.
[0046] Option 11. A method comprising:
[0047] Receive first data from a first group of sensors arranged in a first configuration;
[0048] The first data is transformed into second data to train a model to identify third data captured by a second set of sensors arranged in a second configuration, wherein the second configuration differs from the first configuration; and
[0049] The model is trained based on the second set of sensors sensing the second data to identify the third data captured by the second set of sensors arranged in the second configuration.
[0050] Option 12. The method according to Option 11 further includes: using a trained model to identify the third data captured by the second set of sensors arranged in the second configuration.
[0051] Option 13. The method according to Option 11, wherein at least one of the second group of sensors is different from at least one of the first group of sensors.
[0052] Option 14. The method according to Option 11 further includes:
[0053] Detect one or more objects in the first data; and
[0054] Separate the object from the background in the first data.
[0055] Option 15. The method according to Option 14 further includes:
[0056] The object's perspective is transformed from 2D to 3D using a machine learning-based model; and
[0057] Computer graphics technology is used to transform the perspective of the background from 2D to 3D.
[0058] Option 16. The method according to Option 15 further includes: combining the transformed view of the object and the transformed view of the background to generate a 3D scene representing the first data.
[0059] Option 17. The method according to Option 16 further includes: training the model based on the 3D scene representing the first data sensed by the second set of sensors.
[0060] Option 18. The method according to Option 15 further includes:
[0061] Generate a 2D representation of the object from a 3D perspective sensed by the second set of sensors; and
[0062] Generate a 2D representation of the background from a 3D perspective sensed by the second set of sensors.
[0063] Option 19. The method according to Option 18 further includes: combining the 2D representation of the 3D view of the object sensed by the second set of sensors with the 2D representation of the 3D view of the background.
[0064] Option 20. The method according to Option 19 further includes: training the model based on a combination of a 2D representation of the object from a 3D perspective sensed by the second set of sensors and a 2D representation of the background from a 3D perspective.
[0065] Further applications of this disclosure will become apparent from the detailed description, claims, and drawings. The detailed description and specific examples are intended for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description
[0066] This disclosure will be more fully understood from the detailed description and the accompanying drawings, in which:
[0067] Figure 1 An example of a system according to this disclosure for transforming sensor data to train a model for use with different sensor configurations is shown;
[0068] Figure 2 The present disclosure illustrates a general method for transforming sensor data to train models for use with different sensor configurations.
[0069] Figure 3 This illustrates a method according to the present disclosure for transforming sensor data to train models for use with different sensor configurations. Figure 2 An example of the method; and
[0070] Figure 4 This illustrates a method according to the present disclosure for transforming sensor data to train models for use with different sensor configurations. Figure 2 Another example of the method.
[0071] In the accompanying drawings, reference numerals may be used repeatedly to identify similar and / or identical elements. Detailed Implementation
[0072] Sensors (e.g., cameras) can be configured differently on different types of vehicles (e.g., cars, trucks, SUVs, etc.). Therefore, data sensed by sensors on one type of vehicle (e.g., a car) may differ in many ways from data sensed by sensors on another type of vehicle (e.g., a truck). Consequently, a model trained using data collected from sensors configured for one type of vehicle (e.g., a machine learning-based model) may not produce the correct output when data from sensors configured differently for another type of vehicle is input into the trained model.
[0073] For example, a camera can be mounted on a car in a way that differs from that on a truck or SUV. Therefore, the image of a scene captured by a camera on a car may differ from the image of the same scene captured by a camera on a truck or SUV. For instance, the viewpoint of the image captured by a camera on a car may differ from the viewpoint of the image captured by a camera on a truck or SUV. As a result, a model trained using images captured by cameras configured for one vehicle model (e.g., a machine learning-based model) may not produce accurate outputs when images captured by cameras configured differently for another vehicle model are input into the trained model.
[0074] Therefore, the first data from the first sensor configuration can be used to train the first model; and the first trained model can be used with the first sensor configuration to produce accurate results based on the data sensed by the first sensor configuration. However, the first data cannot be reused to train a second model used with the second sensor configuration. If the first data is used to train the second model, the output of the second trained model, based on the second data received as input from the second sensor configuration, may be inaccurate when used with the second sensor configuration.
[0075] This disclosure provides systems and methods for addressing the aforementioned problems. The system and method transform sensor data collected from one sensor configuration into sensor data viewed from a different reference frame. Specifically, the system and method perform this transformation using viewpoint transformation or machine learning techniques. After transforming the data to the new reference frame, the transformed data can be used to train models (e.g., machine learning-based models) that can be deployed using different sensor configurations.
[0076] Current viewpoint transformation and data augmentation methods are limited to the use of image and computer graphics techniques. In contrast, the system and method of this disclosure complement computer graphics techniques with machine learning to perform reference frame transformations of 3D scenes. This system and method utilize viewpoint transformation and machine learning to transform video, radar, lidar, and other non-static image media.
[0077] More specifically, the system and method perform viewpoint transformation on sensor data by separating the object of interest (OoI) from the background using object detection techniques for each sensor modality. The system and method use machine learning techniques to transform the viewpoint of the OoI and computer graphics techniques to transform the viewpoint of the background. The system and method then recombine the OoI and background with the transformed viewpoint in each sensor modality to perform data augmentation for different sensor configurations.
[0078] Machine learning techniques are used to synthesize missing data from a single object when the viewpoint of a single object (i.e., a single OoI) is transformed. For large background regions containing relatively little salient information, computer graphics techniques can be relatively efficient and sufficiently accurate. Therefore, this system and method transform sensor data aligned across various sensor modalities to a desired sensor configuration. These and other features of the systems and methods disclosed herein will now be described in more detail below.
[0079] Throughout this disclosure, references are made to computer graphics techniques and machine learning techniques used by the systems and methods of this disclosure. For example, computer graphics techniques may include ray tracing. Machine learning techniques may include generative adversarial networks (GANs), neural radiation fields (NeRFs), and generative radiation fields (GRAFs). These techniques are summarized after the description of the systems and methods of this disclosure.
[0080] Figure 1 A system 100 for transforming sensor data into different sensor configurations according to the present disclosure is shown. System 100 includes a first set of sensors 102, a processing module 104, a second set of sensors 106, and a training module 108. Processing module 104 includes an object detection module 110, an object separation module 112, a viewpoint transformation module 114, and a combination module 116.
[0081] The following is for reference. Figures 2-4 This will explain the operation of the various modules in System 100. First, refer to... Figure 2 Briefly describe the operation, and then refer to... Figure 3 and Figure 4 The operation is described in detail. Throughout the following description, the term "controller" refers to one or more modules of processing module 104.
[0082] Figure 2A method 150 for transforming sensor data from one reference frame to a reference frame with a different sensor configuration, according to this disclosure, is illustrated. At 152, a controller (e.g., object detection module 110) receives data from a first sensor (e.g., a first set of sensors 102). At 154, the controller (e.g., elements 112, 114, 116) transforms the data. At 156, a second sensor senses the transformed data. At 158, the controller (e.g., training module 108) uses the transformed data sensed by the second sensor to train a model. At 160, in use, the trained model receives additional data from the second sensor and outputs the correctly trained result. The trained model outputs the result by recognizing additional data as if the model were trained using data directly collected by the second sensor rather than based on data collected by the first sensor.
[0083] Figure 3 A method 200 for transforming sensor data to train a model for use with different sensor configurations, according to this disclosure, is illustrated. At 202, a controller (e.g., object detection module 110) receives first data captured by a first set of sensors (e.g., first set of sensors 102) arranged in a first configuration. At 204, the controller (e.g., object detection module 110) detects an object of interest (OoI) in the first data. At 206, the controller (e.g., object separation module 112) separates the object in the first data from the background.
[0084] At 208, the controller (e.g., view transformation module 114) uses one or more machine learning techniques to transform the view of the object from 2D to 3D. At 210, the controller (e.g., view transformation module 114) uses one or more computer graphics techniques to transform the view of the background from 2D to 3D. At 212, the controller (e.g., combination module 116) combines the transformed 3D views of the object and the background to generate a 3D scene representing the first data.
[0085] At 214, a second set of sensors (e.g., second set of sensors 106) arranged in a second configuration senses the 3D scene to generate a 3D representation of the 3D scene. The arrangement of the second sensors in the second configuration differs from the arrangement of the first set of sensors in the first configuration. At 216, a controller (e.g., training module 108) uses the data sensed by the second set of sensors to train a model (e.g., a machine learning-based model). That is, the controller uses the 3D representation of the 3D scene generated by the second set of sensors to train the model.
[0086] At point 218, in use, the trained model receives additional data from the second set of sensors and outputs the correctly trained result. The trained model outputs the result by recognizing additional data, just as it was trained using additional data directly collected by the second set of sensors, rather than based on data collected by the first set of sensors as described above.
[0087] Figure 4 A method 250 for transforming sensor data from one reference frame to a reference frame with a different sensor configuration, according to this disclosure, is illustrated. Method 250 differs from method 200 in that method 200 combines 3D representations of the object and background from a transformed viewpoint, while method 250 combines 2D representations of the object and background from a transformed viewpoint, as described below. Essentially, method 200 places a 3D OoI within a 3D background and then senses the 3D scene using a second set of sensors, while method 250 senses both the 3D OoI and the 3D background using a second set of sensors and places a 2D representation of the OoI within a 2D representation of the background, as explained below.
[0088] At 252, the controller (e.g., object detection module 110) receives first data captured by a first set of sensors (e.g., first set of sensors 102) arranged in a first configuration. At 254, the controller (e.g., object detection module 110) detects an object of interest in the first data. At 256, the controller (e.g., object separation module 112) separates the object in the first data from the background.
[0089] At 258, the controller (e.g., view transformation module 114) uses one or more machine learning techniques to transform the view of the object from 2D to 3D. At 260, the controller (e.g., view transformation module 114) uses one or more computer graphics techniques to transform the view of the background from 2D to 3D.
[0090] At 262, a second set of sensors (e.g., second set of sensors 106) arranged in a second configuration senses the 3D transformation view of the object and generates a 2D representation of the 3D transformation view of the object. The arrangement of the second sensors in the second configuration differs from the arrangement of the first set of sensors in the first configuration. At 264, the second set of sensors senses the 3D transformation view of the background and generates a 2D representation of the 3D transformation view of the background.
[0091] At 266, the controller (e.g., combination module 116) combines 2D representations of the viewpoints of the 3D transformations of the object and the background. At 268, the controller (e.g., training module 108) uses the combined 2D representations of the viewpoints of the 3D transformations of the object and the background to train a model (e.g., a machine learning-based model).
[0092] At point 270, in use, the trained model receives additional data from the second set of sensors and outputs the correctly trained result. The trained model outputs results by recognizing additional data, just as it was trained using additional data directly collected by the second set of sensors, rather than based on data collected by the first set of sensors as described above.
[0093] The systems and methods described above can be used in many applications. Non-limiting examples of applications include the following. For instance, the systems and methods described above can be used for data augmentation during the training of various machine learning systems.
[0094] In the second use case example, the system and method can be used with V2X (Vehicle-to-All) communication systems and Advanced Driver Assistance Systems (ADAS). V2X is communication between a vehicle and any entity that may affect or be affected by the vehicle. V2X incorporates other more specific types of communication, such as V2I (Vehicle-to-Infrastructure), V2N (Vehicle-to-Network), V2V (Vehicle-to-Vehicle), V2P (Vehicle-to-Pedestrian), V2D (Vehicle-to-Device), and V2G (Vehicle-to-Grid).
[0095] V2X defines a peer-to-peer communication protocol that enhances situational awareness between vehicles. V2X applications range from intersection warnings and alerts to nearby emergency vehicles to blind spot warnings that help prevent lane-change-related accidents in regular and emergency traffic situations. Additionally, road detours for construction, traffic flow, or traffic accidents can be signaled via V2X. Pedestrians can also benefit from V2X safety enhancements via their mobile phones.
[0096] ADAS uses human-machine interfaces to improve a driver's ability to react to hazards on the road. ADAS increases safety and reaction time through warning and automated systems. Some examples of ADAS include forward collision warning, high beam safety systems, lane departure warning, and traffic sign recognition. Current ADAS capabilities are limited by the capabilities of the vehicle's sensors. V2V communication can expand ADAS capabilities by allowing vehicles to communicate directly with each other and share information about relative speed, position, direction of travel, and even control inputs (such as sudden braking, acceleration, or changes in direction). Combining this data with the vehicle's own sensor inputs can create a wider and more detailed picture of the surrounding environment and provide earlier and more accurate warnings or corrective actions to avoid collisions.
[0097] The systems and methods disclosed herein can be used with V2X and ADAS, as described below. For example, a first vehicle manufactured by a first manufacturer can relay a hazard detection indication via V2X. A second vehicle manufactured by a second manufacturer can receive the indication via V2X. Without the systems and methods described above, the second vehicle accepts or rejects the presence of a hazard indicated by the first vehicle. Conversely, if the systems and methods described above are deployed in the second vehicle, the first vehicle can include its sensor data and a short sequence of information about its sensors in the indication. The systems and methods described above in the second vehicle can transform the sensor data of the first vehicle to match its own configuration and use its own model (trained to process sensor data present in the configuration on the first vehicle) to analyze the hazard and draw conclusions independently of the decisions made by the first vehicle. Therefore, the second vehicle can make better decisions about how to handle a hazardous situation, rather than making a binary decision by relying on a hazard indication received from the first vehicle.
[0098] In a third example use case, the infotainment system of a vehicle employing the above system and method can use sensor data collected from the vehicle's sensors and transform the collected data to enhance the field of vision provided to the vehicle occupants. The system and method can also allow the occupants to manipulate the viewpoint of the scene captured by the vehicle's sensors to alter the display of the surrounding environment. For example, on a touchscreen displaying the scene, the occupants can be provided with a menu including various configurations for the vehicle's sensors (i.e., various possible arrangements of the sensors that can be virtually arranged). The occupants can select a configuration, and the system uses the data collected by the vehicle's sensors, transforms the collected data to the selected sensor configuration using the system and method of this disclosure, and displays a new view of the scene on the touchscreen as if the new view were actually captured by the vehicle's sensors arranged in the selected configuration.
[0099] The following is a summary of various computer graphics and machine learning techniques that can be used with the systems and methods described above. For example, in 3D computer graphics, ray tracing is a rendering technique used to generate images by tracing the paths of light rays to pixels in the image plane and simulating the effects of these rays encountering virtual objects. Ray tracing can simulate many optical effects, such as reflection, refraction, scattering, and dispersion phenomena, such as chromatic aberration. Ray tracing can produce a high degree of visual realism, more realistic than typical scanline rendering methods, but it is computationally intensive.
[0100] Path tracing is a form of ray tracing that can produce soft shadows, depth of field, motion blur, caustics, ambient occlusion, and direct lighting. Path tracing is an unbiased rendering method, but it requires tracing a large number of rays to obtain a high-quality reference image without noise artifacts.
[0101] The following are examples of machine learning techniques that can be used to detect and manipulate objects of interest as described in the systems and methods described above in this disclosure. For example, Generative Adversarial Networks (GANs) are a class of machine learning techniques that can be used to synthesize 3D objects. Given a training set, a GAN learns to generate new data with the same statistics as the training set. For example, a GAN trained on a photograph can generate a new photograph that has many realistic characteristics and appears at least superficially realistic.
[0102] As another example, Neural Radiation Field (NeRF) is a fully connected deep network that can be trained to reproduce an input view of a single scene using a rendering loss. The network receives the spatial location and viewing orientation (5D input) and outputs the volumetric density and view-dependent emission radiation at that location. Volumetric rendering is used to discriminately render the new view. To create a 3D scene, NeRF uses many images of the scene taken from different views, and is therefore computationally intensive. Therefore, NeRF is better suited for creating static scenes, such as virtual museum exhibits, compared to dynamically changing environments with many scenes encountered by vehicles while driving.
[0103] In other examples, 3D objects can be represented by a continuous function called Generative Radiation Field (GRAF). GRAF generates a consistent 3D image and trains using only unposed 2D images. GRAF incorporates 3D perception by adding a virtual camera to the model. The generated 3D representation of the object is parameterized by the 3D generator. The virtual camera and a corresponding renderer produce images of the 3D representation. GRAF can render images from different viewpoints by controlling the pose of the virtual camera in the model. GRAF models shape and appearance using two untangled latent codes, allowing for separate modification of them.
[0104] These techniques primarily focus on object manipulation rather than dynamic scenes. However, by combining these techniques (i.e., using computer graphics techniques for the background and machine learning techniques for the OoI), the aforementioned systems and methods can synthesize dynamic scenes, such as those captured by cameras while driving a vehicle. This enables the transformation of sensor data to train models for use with different sensor configurations. Specifically, as described above, the systems and methods separate the OoI from the background and combine these techniques (i.e., using machine learning techniques to transform the perspective of the OoI and using computer graphics techniques to transform the perspective of the background) to transform the sensor data to an additional reference frame.
[0105] The foregoing description is illustrative in nature and is not intended to limit this disclosure, its application, or use. The broad teachings of this disclosure can be implemented in various forms. Therefore, while this disclosure includes specific examples, its true scope should not be so limited, as other modifications will become apparent upon examination of the drawings, specification, and appended claims. It should be understood that one or more steps within the method may be performed in different orders (or simultaneously) without altering the principles of this disclosure. Furthermore, although each of the embodiments described above is described as having certain features, any one or more of those features described with respect to any embodiment of this disclosure may be implemented in features of any of other embodiments and / or combined with features of any of other embodiments, even if such combinations are not explicitly described. In other words, the described embodiments are not mutually exclusive, and substitutions of one or more embodiments for each other remain within the scope of this disclosure.
[0106] Various terms are used to describe spatial and functional relationships between elements (e.g., between modules, circuit elements, semiconductor layers, etc.), including “connection,” “joint,” “link,” “adjacent,” “near,” “on top,” “above,” “below,” and “set.” Unless explicitly described as “direct,” when a relationship between first and second elements is described in the foregoing disclosure, the relationship can be a direct relationship in which no other intermediate elements exist between the first and second elements, or an indirect relationship in which one or more intermediate elements exist between the first and second elements (spatially or functionally). As used herein, at least one of the phrases A, B, and C should be interpreted as indicating a logic using non-exclusive OR (A OR B OR C) and should not be interpreted as indicating “at least one of A, at least one of B, and at least one of C.”
[0107] In the accompanying drawings, the direction of the arrows, as indicated by the arrows, typically represents the flow of information of interest (e.g., data or instructions). For example, when elements A and B exchange various types of information, but the information transmitted from element A to element B is relevant to the illustration, the arrow can point from element A to element B. This unidirectional arrow does not imply that no other information is transmitted from element B to element A. Furthermore, for information sent from element A to element B, element B can send a request for that information or an acknowledgment of receipt of that information to element A.
[0108] In this application, including the following definitions, the term "module" or "controller" may be replaced by the term "circuit". The term "module" may refer to, be part of, or include the following: application-specific integrated circuit (ASIC); digital, analog, or mixed analog / digital discrete circuit; digital, analog, or mixed analog / digital integrated circuit; combinational logic circuit; field-programmable gate array (FPGA); processor circuitry (shared, dedicated, or grouped) that executes code; memory circuitry (shared, dedicated, or grouped) that stores code executed by said processor circuitry; other suitable hardware components that provide said functionality; or combinations of some or all of the above, such as in a system-on-a-chip.
[0109] This module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces connected to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module in this disclosure may be distributed across multiple modules connected via the interface circuits. For example, multiple modules may allow for load balancing. In another example, a server (also referred to as a remote or cloud) module may perform some functions on behalf of a client module.
[0110] The term "code" as used above can include software, firmware, and / or microcode, and can refer to programs, routines, functions, classes, data structures, and / or objects. The term "shared processor circuitry" covers a single processor circuitry that executes some or all of the code from multiple modules. The term "group processor" circuitry includes processor circuitry combined with additional processor circuitry to execute some or all of the code from one or more modules. References to multiprocessor circuitry cover multiprocessor circuitry on discrete dies, multiprocessor circuitry on a single die, multiple cores of a single processor circuit, multiple threads of a single processor circuit, or a combination thereof. The term "shared memory circuitry" covers a single memory circuitry that stores some or all of the code from multiple modules. The term "group memory circuitry" covers memory circuitry that, in combination with additional memory, stores some or all of the code from one or more modules.
[0111] The term memory circuit is a subset of the term computer-readable medium. As used herein, the term computer-readable medium does not cover transient electrical or electromagnetic signals propagating through a medium (such as on a carrier wave); therefore, the term computer-readable medium can be considered tangible and non-transient. Non-limiting examples of non-transient, tangible computer-readable media are non-volatile memory circuits (such as flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (such as static random access memory circuits or dynamic random access memory circuits), magnetic storage media (such as analog or digital magnetic tape or hard disk drives), and optical storage media (such as CDs, DVDs, or Blu-ray discs).
[0112] The apparatus and methods described in this application can be implemented, in part or in whole, by a special-purpose computer created by configuring a general-purpose computer to perform one or more specific functions implemented in a computer program. The aforementioned function blocks, flowchart components, and other elements serve as software specifications that can be routinely converted into computer programs by skilled technicians or programmers.
[0113] A computer program includes processor-executable instructions stored on at least one non-transitory tangible computer-readable medium. A computer program may also include or depend on stored data. A computer program may encompass a basic input / output system (BIOS) that interacts with the hardware of a special-purpose computer, device drivers that interact with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, etc.
[0114] Computer programs may include: (i) descriptive text to be parsed, such as HTML (Hypertext Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation), (ii) assembly code, (iii) object code generated from source code by a compiler, (iv) source code executed by an interpreter, (v) source code compiled and executed by a just-in-time compiler, and so on. As an example only, source code can be written using syntax from languages including C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Java®, Fortran, Perl, Pascal, Curl, OCaml, JavaScript®, HTML5 (Hypertext Markup Language 5th Revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Flash®, Visual Basic®, Lua, MATLAB, SIMULINK, and Python®.
Claims
1. A system comprising: processor; as well as Memory, which stores instructions that, when executed by the processor, configure the processor to: Receive first data from a first group of sensors arranged in a first configuration; The first data is transformed into second data to train a model to identify third data captured by a second set of sensors arranged in a second configuration, wherein the second configuration is different from the first configuration; as well as The model is trained based on the second set of sensors sensing the second data to identify the third data captured by the second set of sensors arranged in the second configuration. The instructions also configure the processor to: Detect one or more objects in the first data; Separate the object from the background in the first data; The object's perspective is transformed from 2D to 3D using a machine learning-based model; and Computer graphics technology is used to transform the perspective of the background from 2D to 3D.
2. The system according to claim 1, wherein, The trained model recognizes the third data captured by the second set of sensors arranged in the second configuration.
3. The system according to claim 1, wherein, At least one of the second group of sensors is different from at least one of the first group of sensors.
4. The system according to claim 1, wherein, The instructions configure the processor to combine the transformed viewpoint of the object and the transformed viewpoint of the background to generate a 3D scene representing the first data.
5. The system according to claim 4, wherein, The instructions configure the processor to train the model based on a 3D scene that represents the first data sensed by the second set of sensors.
6. The system according to claim 1, wherein, The instructions configure the processor to: Generate a 2D representation of the object from the 3D perspective sensed by the second set of sensors; and Generate a 2D representation of the 3D perspective of the background sensed by the second set of sensors.
7. The system according to claim 6, wherein, The instructions configure the processor to combine the 2D representation of the object from the 3D perspective sensed by the second set of sensors with the 2D representation of the background from the 3D perspective.
8. The system according to claim 7, wherein, The instructions configure the processor to train the model based on a combination of the 2D representation of the object from the 3D perspective sensed by the second set of sensors and the 2D representation of the background from the 3D perspective.
9. A method comprising: Receive first data from a first group of sensors arranged in a first configuration; The first data is transformed into second data to train a model to identify third data captured by a second set of sensors arranged in a second configuration, which is different from the first configuration. as well as The model is trained based on the second set of sensors sensing the second data to identify the third data captured by the second set of sensors arranged in the second configuration. The method further includes: Detect one or more objects in the first data; Separate the object from the background in the first data; The object's perspective is transformed from 2D to 3D using a machine learning-based model; and Computer graphics technology is used to transform the perspective of the background from 2D to 3D.
10. The method of claim 9, further comprising: The trained model is used to identify the third data captured by the second set of sensors arranged in the second configuration.
11. The method according to claim 9, wherein, At least one of the second group of sensors is different from at least one of the first group of sensors.
12. The method according to claim 9, further comprising: The transformed viewpoints of the object and the transformed viewpoints of the background are combined to generate a 3D scene representing the first data.
13. The method of claim 12, further comprising: The model is trained based on the 3D scene perceived by the second set of sensors, which represents the first data.
14. The method of claim 9, further comprising: Generate a 2D representation of the object from a 3D perspective sensed by the second set of sensors; as well as Generate a 2D representation of the background from a 3D perspective sensed by the second set of sensors.
15. The method of claim 14, further comprising: The 2D representation of the object from the 3D perspective sensed by the second set of sensors is combined with the 2D representation of the background from the 3D perspective.
16. The method of claim 15, further comprising: The model is trained based on a combination of a 2D representation of the object from a 3D perspective sensed by the second set of sensors and a 2D representation of the background from a 3D perspective.
Citation Information
Patent Citations
Method for searching for 3D model with 2D pictures
CN105930382A
Behavior-guided path planning in autonomous machine applications
CN110618678A
Automatic generation of ground truth data for training or retraining machine learning model
CN112287960A
Object detection system and method incorporating background clutter removal
US20080310677A1
Method and apparatus for estimating the velocity vector of multiple vehicles on non-level and curved roads using a single camera
US7920959B1