Visual perception method and device, storage medium and electronic device
By utilizing bird's-eye view grid index features and a self-attention model in autonomous driving, the problem of multi-camera systems' dependence on intrinsic and extrinsic parameter matrices is solved, achieving efficient and accurate visual perception.
Patent Information
- Application Number
- CN202210618710.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-06-01
AI Technical Summary
In existing technologies, the visual perception methods of multi-camera systems in autonomous driving rely heavily on intrinsic and extrinsic parameter matrices, which leads to inaccurate mapping results during vehicle operation and affects the accuracy of perception tasks.
By extracting features from images acquired by a multi-camera system on a vehicle, and using grid point index features and a self-attention model in the bird's-eye view, grid point features are determined, and the feature map of the bird's-eye view is directly obtained, avoiding dependence on intrinsic and extrinsic parameter matrices.
It improves the processing efficiency and accuracy of perception tasks, reduces the dependence on intrinsic and extrinsic parameter matrices, and ensures the stability of perception results when the vehicle position changes.
Smart Images

Figure CN114882465B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of visual perception technology, and in particular to a visual perception method, apparatus, storage medium, and electronic device. Background Technology
[0002] In the field of autonomous driving, the perception results of multi-camera systems in visual perception systems cannot be directly used for subsequent prediction and control systems. It is necessary to fuse the features from different perspectives in a certain way, such as mapping them to a bird's-eye view (BEV) and uniformly expressing them in the vehicle's coordinate system. Existing technologies usually use the point attention scheme to map the images obtained by the multi-camera system to the bird's-eye view, but this method is highly dependent on the intrinsic and extrinsic parameter matrices. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure is proposed. Embodiments of this disclosure provide a visual perception method, apparatus, storage medium, and electronic device.
[0004] According to one aspect of the present disclosure, a visual perception method is provided, comprising:
[0005] The multi-camera system of the vehicle simultaneously captures multiple images from different perspectives of the surrounding environment of the vehicle, and extracts features from each image to obtain multiple first feature maps.
[0006] Based on the index features of multiple grid points included in the bird's-eye view at the same time corresponding to at least one first feature map, determine the grid point features corresponding to each of the multiple grid points respectively;
[0007] Based on the grid features corresponding to each of the grid points, a second feature map corresponding to the bird's-eye view is determined;
[0008] The second feature map is identified based on the network model corresponding to the preset perception task, and the perception result corresponding to the preset perception task is determined.
[0009] According to another aspect of the embodiments of this disclosure, an in-vehicle visual perception device is provided, comprising:
[0010] The feature extraction module is used to extract features from multiple images of the vehicle's surrounding environment captured by the vehicle's multi-camera system at the same time from different perspectives, and obtain multiple first feature maps.
[0011] The feature correspondence module is used to determine the grid feature corresponding to each of the multiple grid points based on the index features of the multiple grid points included in the bird's-eye view at the same time in the first feature map determined by at least one of the feature extraction modules.
[0012] The feature map determination module is used to determine the second feature map corresponding to the bird's-eye view based on the grid point features determined by the feature correspondence module corresponding to each of the grid points.
[0013] The perception and recognition module is used to recognize the second feature map determined by the feature map determination module based on the network model corresponding to the preset perception task, and to determine the perception result corresponding to the preset perception task.
[0014] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the visual perception method described in any of the above embodiments.
[0015] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising:
[0016] processor;
[0017] Memory used to store the processor's executable instructions;
[0018] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the visual perception method described in any of the above embodiments.
[0019] Based on the above embodiments of this disclosure, a visual perception method, apparatus, storage medium, and electronic device are provided. By determining the index feature corresponding to each grid point in the bird's-eye view, the grid point feature corresponding to each grid point is determined. This eliminates the need to combine intrinsic and extrinsic parameter matrices to determine the second feature corresponding to the bird's-eye view, thus overcoming the problem of excessive reliance on intrinsic and extrinsic parameter matrices in the prior art.
[0020] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0021] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0022] Figure 1This is a flowchart illustrating a visual perception method provided in an exemplary embodiment of this disclosure.
[0023] Figure 2a This is a public announcement Figure 1 The illustrated embodiment is a flowchart of step 102.
[0024] Figure 2b This is a schematic diagram of an exemplary grid-corresponding region in a visual perception method provided by an exemplary embodiment of this disclosure.
[0025] Figure 3 This is a public announcement Figure 2a The illustrated embodiment is a flowchart of step 1022.
[0026] Figure 4 This is a public announcement Figure 3 The illustrated embodiment is a flowchart of step 301.
[0027] Figure 5 This is a public announcement Figure 2a The illustrated embodiment is a flowchart of step 1023.
[0028] Figure 6 This is a schematic diagram of the structure of a visual sensing device provided in an exemplary embodiment of the present disclosure.
[0029] Figure 7 This is a schematic diagram of the structure of a visual sensing device provided in another exemplary embodiment of this disclosure.
[0030] Figure 8 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0031] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0032] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0033] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0034] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0035] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0036] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0037] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0038] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0039] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0040] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0041] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0042] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0043] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0044] Application Overview
[0045] In the process of realizing this disclosure, the inventors discovered that the prior art usually uses a pointing attention scheme to map the images obtained by the vehicle camera onto the bird's-eye view. However, this method has at least the following problems: it is highly dependent on the intrinsic and extrinsic parameter matrices. When the vehicle is moving, the intrinsic and extrinsic parameters of the cameras in the multi-camera system are prone to change, which will cause the mapping results to be incorrect, thus leading to inaccurate perception results of the perception task determined by the bird's-eye view.
[0046] Exemplary methods
[0047] Figure 1 This is a schematic flowchart of a visual perception method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 1 As shown, it includes the following steps:
[0048] Step 101: Extract features from multiple images of the vehicle's surrounding environment captured by the vehicle's multi-camera system at the same time from different perspectives to obtain multiple first feature maps.
[0049] In this embodiment, the vehicle can be any device capable of mounting a multi-camera system, such as a vehicle or an intelligent mobile robot. Since the vehicle's position may change over time (e.g., a vehicle's position changes continuously during travel), this embodiment limits the multiple images to be acquired simultaneously by the multi-camera system. This ensures that the multiple images captured by the multi-camera system represent different perspectives of the vehicle from the same location. The multi-camera system can include multiple cameras, preferably automotive-grade cameras. Each camera corresponds to one image and can be used to capture images containing the vehicle's surrounding environment. For example, when the vehicle is a vehicle, the multi-camera system can consist of multiple surround-view cameras mounted on the vehicle. After acquiring multiple images, a neural network model (e.g., a convolutional neural network) can be used to extract features from each image, resulting in multiple first feature maps, each corresponding to one image.
[0050] Step 102: Based on the index features of multiple grid points included in the bird's-eye view at the same time corresponding to at least one first feature map, determine the grid point features corresponding to each grid point in the multiple grid points.
[0051] In this embodiment, the bird's-eye view is a portion (e.g., a rectangular area) of a plane with a preset z-axis value in a coordinate system centered on the vehicle. Each grid point can correspond to a rectangular area in the world coordinate system. For example, each grid point corresponds to an area of 0.5m*0.5m in the world coordinate system, or an area of 0.6m*1m in the world coordinate system. The specific size of each grid point can be set according to the actual application scenario. The larger the range corresponding to the bird's-eye view, the larger the area corresponding to the grid point can be. The length and width of each grid point can be equal or unequal. In this embodiment, the index features of each grid point in the bird's-eye view can be determined in at least one first feature map in advance through coordinate system transformation, projection transformation, and other operations. When performing perception and recognition based on the current multi-camera system, these index features can be directly obtained to determine the grid point features corresponding to each grid point in the bird's-eye view.
[0052] Step 103: Based on the grid features corresponding to each grid point, determine the second feature map corresponding to the bird's-eye view.
[0053] Optionally, the grid feature of each grid point can be represented as a vector. By concatenating the grid features corresponding to multiple grid points according to the position of each grid point in the bird's-eye view, the second feature map corresponding to the bird's-eye view can be obtained.
[0054] Step 104: Based on the network model corresponding to the preset perception task, identify the second feature map and determine the perception result corresponding to the preset perception task.
[0055] In this embodiment, the preset perception task can be any visual perception task, such as segmentation, detection, or classification. The operation of the perception task in this step is implemented through a network model corresponding to the visual perception task.
[0056] The visual perception method provided in the above embodiments of this disclosure determines the grid feature corresponding to each grid point by determining the index feature corresponding to each grid point in the bird's-eye view. It does not require combining the intrinsic and extrinsic parameter matrices to determine the second feature corresponding to the bird's-eye view, thus overcoming the problem of the prior art's heavy reliance on the intrinsic and extrinsic parameter matrices.
[0057] like Figure 2a As shown above, in the above Figure 1 Based on the illustrated embodiment, step 102 may include the following steps:
[0058] Step 1021: Determine the multiple grid points included in the bird's-eye view.
[0059] Optionally, the size of the bird's-eye view can be determined according to the actual application scenario. Typically, the bird's-eye view covers the same range of the surrounding environment as the multiple images captured by the multi-camera system. Optionally, the bird's-eye view can be divided into multiple grid points according to its size, with each grid point being the same size. For example, each grid point corresponds to an area of 0.5m*0.5m in the world coordinate system.
[0060] Step 1022: Determine the index feature corresponding to each grid point in at least one first feature map, and obtain at least one index feature corresponding to each grid point.
[0061] In this embodiment, each grid point may or may not have a corresponding index region in a first feature map. For example, if a grid point corresponds to the right side of the vehicle in a bird's-eye view, the first feature map corresponding to the image obtained from the left-side view camera may not have the index feature corresponding to that grid point; for example, as... Figure 2b As shown, the grid points in the bird's-eye view on the right correspond to the positions of the vehicle's side doors captured by the front-view camera and the right front-side camera; however, the vehicle's side doors were not captured by the other cameras in the multi-camera system, therefore, the grid points do not have index features in other images.
[0062] Step 1023: Based on at least one index feature corresponding to each grid point, determine the grid point features corresponding to each grid point in the multiple grid points.
[0063] In this embodiment, a neural network model (e.g., a self-attention model) can be used to process at least one index feature corresponding to each grid point. Taking all the index features corresponding to a grid point as input, the output is the grid point feature corresponding to that grid point. By combining the grid point feature with the position of the grid point in the bird's-eye view, the grid point feature of each grid point in the bird's-eye view can be determined, thereby obtaining the second feature map of the bird's-eye view. That is, this embodiment does not require coordinate system transformation involving camera intrinsic and extrinsic parameter matrices to obtain the second feature map corresponding to the bird's-eye view, overcoming the problem of dependence on intrinsic and extrinsic parameter matrices in related technologies, and enabling the bird's-eye view with determined index relationships to obtain the corresponding second feature map at any time.
[0064] like Figure 3 As shown above, in the above Figure 2a Based on the illustrated embodiment, step 1022 may include the following steps:
[0065] Step 301: For each of the at least one first feature maps, based on the intrinsic parameter matrix, extrinsic parameter matrix, and three-dimensional coordinates of the grid points in the multi-camera system corresponding to the first feature map, determine the mapping region of the grid points in the first feature map.
[0066] In this embodiment, after determining the positional relationship between the vehicle and the multi-camera system, the mapping area of the grid point in each first feature map can be determined based on the camera's intrinsic and extrinsic parameter matrices and the three-dimensional coordinates of the grid point. For example, the features in the image coordinate system can be transformed to a coordinate system centered on the vehicle (e.g., the vehicle coordinate system) through coordinate system transformation, so that when performing visual perception tasks, the index features corresponding to each grid point can be directly obtained without having to recombine the camera's intrinsic and extrinsic parameter matrices for coordinate system transformation.
[0067] Step 302: Based on the position of the grid point in at least one mapping region corresponding to at least one first feature map, determine at least one index feature corresponding to the grid point.
[0068] In this embodiment, grid points can be mapped to each of multiple first feature maps. When the location corresponding to the mapped region is within the range of the first feature map, the mapped region is valid, and the feature corresponding to the mapped region is used as the index feature of the grid point in the first feature map. When the location corresponding to the mapped region exceeds the range of the first feature map (e.g., the location coordinates are negative), it indicates that the mapped region is invalid, and the index feature corresponding to the grid point does not exist in the first feature map corresponding to the mapped region. By using the index features determined by all valid mapped regions, at least one index feature corresponding to the grid point is obtained. By pre-establishing the correspondence between grid points and index features, the dependence of the grid point feature corresponding to the grid point on the camera's intrinsic and extrinsic parameter matrices is reduced. Even if the intrinsic and extrinsic parameter matrices change to some extent when the vehicle position changes (e.g., when the vehicle is moving), the method provided in this embodiment can still obtain relatively accurate perception results.
[0069] The process of determining the index features of grid points using intrinsic and extrinsic parameter matrices provided in this embodiment is executed after determining the positional relationship between the vehicle and the multi-camera system, and can be performed before step 101. After determining the index features corresponding to the grid points, the correspondence can be stored. In the actual application of the visual perception method provided in this embodiment, the correspondence between grid points and index features can be directly called, without having to perform the correspondence calculation of grid points and index features every time a perception task is performed. As long as the vehicle and the multi-camera system set on the vehicle do not change, it is not necessary to combine the intrinsic and extrinsic parameter matrices to determine the index features corresponding to the grid points, which improves the processing efficiency of perception tasks and overcomes the problem of heavy reliance on intrinsic and extrinsic parameter matrices in the prior art.
[0070] like Figure 4 As shown above, in the above Figure 3 Based on the illustrated embodiment, step 301 may include the following steps:
[0071] Step 3011: Based on the intrinsic and extrinsic parameter matrices, map the three-dimensional coordinates of the grid points to the image coordinate system corresponding to the first feature map to obtain the image coordinates of the grid points in the image coordinate system.
[0072] Optionally, the three-dimensional coordinates of the grid points can be mapped to the image coordinate system corresponding to the first feature map based on the following formula (1):
[0073] c k =K k ·Rt k ·c 3D Formula (1)
[0074] Where k represents the corresponding vehicle camera number, the multiple cameras included in the multi-camera system can be pre-numbered, and the corresponding number is used to represent the corresponding camera later. For example, the 6 cameras included in the multi-camera system can be numbered clockwise from the front as 1 (front), 2 (front right), 3 (rear right), 4 (rear), 5 (rear left), 6 (front left), etc. k K represents the image coordinates of the grid points mapped to the first feature map. k Let Rt be the intrinsic parameter matrix of the camera numbered k. k c represents the extrinsic parameter matrix of the camera numbered k. 3D Represents the three-dimensional coordinates of the grid points.
[0075] Step 3012: Determine the mapping area using the image coordinates as the center and combining the preset length and preset width.
[0076] In this embodiment, a mapping region is determined based on a preset length and a preset width. For example, when the preset length is Kh and the preset width is Kw, the resulting mapping region is Kh*Kw. The values of Kh and Kw can be the same or different, and their specific values can be set according to the actual application scenario. This embodiment maps grid points to a mapping region based on a certain range, rather than using only image coordinates as mapping coordinates. Therefore, even if inaccurate intrinsic and extrinsic parameters or precision issues cause a shift in the corresponding mapping region, the target can still be covered, making the perception result obtained from the second feature map determined based on this mapping region insensitive to intrinsic and extrinsic parameters.
[0077] Optionally, step 3011 may include:
[0078] Step a1: Based on the intrinsic and extrinsic parameter matrices, the three-dimensional coordinates of the grid points are mapped to the image coordinate system to obtain the precise coordinates of the grid points in the image coordinate system.
[0079] For example, coordinate mapping can be achieved based on the above formula (1) to obtain the precise coordinates of the grid point in the image coordinate system. At this time, the coordinate point may be an integer or a non-integer.
[0080] Step a2: Obtain an integer coordinate as the image coordinate based on the position of the precise coordinates around the precise coordinates.
[0081] In this embodiment, when the precise coordinates are integers, they are directly used as image coordinates, and the mapping area is determined with these image coordinates as the center point. In most cases, the precise coordinates are non-integers. In this case, an integer coordinate can be determined at any position around the precise coordinates to achieve rounding. This allows the grid points to be roughly mapped onto the first feature map of each viewpoint using the camera's intrinsic and extrinsic parameters (allowing for errors), thus obtaining the image coordinates (u, v) under each viewpoint. In this embodiment, the rounding process is equivalent to adding noise to the intrinsic and extrinsic parameter matrix, thus avoiding excessive dependence on the precision of the intrinsic and extrinsic parameter matrix.
[0082] like Figure 5 As shown above, in the above Figure 2a Based on the illustrated embodiment, step 1023 may include the following steps:
[0083] Step 501: Perform an expansion operation on each of the at least one first feature maps to obtain multiple strip features corresponding to each first feature map.
[0084] Optionally, each first feature map can be expanded using a convolution acceleration algorithm (img2col) to obtain expanded features. Expanded features can be understood as including multiple strip features, where each strip feature (one row of expanded features) corresponds to a K*K patch surrounding each feature point on the first feature map, where K is an integer greater than 1, and the specific value can be set according to the actual scenario.
[0085] Step 502: Based on the multiple bar features corresponding to each first feature map, determine the bar index feature corresponding to each index feature in at least one index feature, and obtain at least one bar index feature corresponding to each grid point.
[0086] Since each strip feature obtained by unfolding each first feature map corresponds to a K*K region surrounding a pixel, and combined with the image features corresponding to each grid point (corresponding to a feature point in the first feature map), at least one strip index feature corresponding to each grid point can be determined.
[0087] Step 503: Determine the grid feature corresponding to each grid point based on at least one bar index feature corresponding to each grid point.
[0088] In this embodiment, at least one bar index feature corresponding to each grid point can be input into a self-attention model (e.g., Transformer, a deep learning model with a self-attention mechanism) to obtain the grid point feature (e.g., a 1*1*C′ feature vector). By performing self-attention operation only on at least one bar index feature to obtain the grid point feature, compared to the related technology which requires performing self-attention operation on all feature points in all first feature maps, this embodiment greatly reduces the amount of computation and improves the processing efficiency of the preset perception task.
[0089] In another embodiment, step 1023 may further include: for each grid point in the plurality of grid points, performing feature extraction on at least one index feature corresponding to each grid point based on a self-attention model to obtain grid point features corresponding to each grid point.
[0090] This embodiment is different from the above. Figure 5 The illustrated embodiment omits the step of unfolding the first feature map, and directly inputs the image features of at least one region corresponding to the grid point in at least one first feature map into a self-attention model (e.g., Transformer, a deep learning model with a self-attention mechanism) to obtain the grid point features (e.g., a 1*1*C′ feature vector) corresponding to that grid point. Compared with related technologies, this embodiment greatly reduces the amount of computation by performing self-attention operations on all feature points in all first feature maps. Furthermore, by using strip features as input instead of unfolding the first feature map, the computation process is saved, further improving the processing efficiency of the preset perception task.
[0091] Any of the visual perception methods provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the visual perception methods provided in this disclosure can be executed by a processor, such as by a processor executing any of the visual perception methods mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.
[0092] Exemplary device
[0093] Figure 6 This is a schematic diagram of the structure of a visual sensing device provided in an exemplary embodiment of this disclosure. Figure 6 As shown, the apparatus provided in this embodiment includes:
[0094] The feature extraction module 61 is used to extract features from multiple images of the vehicle's multi-camera system at the same time from different perspectives of the vehicle's surrounding environment, and obtain multiple first feature maps.
[0095] The feature correspondence module 62 is used to determine the grid feature corresponding to each grid point in the multiple grid points based on the index features of the multiple grid points included in the bird's-eye view at the same time and the index features of at least one first feature map determined by the feature extraction module 61.
[0096] The feature map determination module 63 is used to determine the second feature map corresponding to the bird's-eye view based on the grid features determined by the feature correspondence module 62 for each grid point.
[0097] The perception and recognition module 64 is used to recognize the second feature map determined by the feature map determination module 63 based on the network model corresponding to the preset perception task, and to determine the perception result corresponding to the preset perception task.
[0098] The visual perception device provided in the above embodiments of this disclosure determines the grid feature corresponding to each grid point by determining the index feature corresponding to each grid point in the bird's-eye view. It does not require combining the intrinsic and extrinsic parameter matrices to determine the second feature corresponding to the bird's-eye view, thus overcoming the problem of the prior art's heavy reliance on the intrinsic and extrinsic parameter matrices.
[0099] Figure 7 This is a schematic diagram of the structure of a visual sensing device provided in another exemplary embodiment of this disclosure. Figure 7 As shown, the apparatus provided in this embodiment includes:
[0100] Feature correspondence module 62 includes:
[0101] Grid point determination unit 621 is used to determine the plurality of grid points included in the bird's-eye view.
[0102] The index determination unit 622 is used to determine the index features corresponding to each of the multiple grid points in at least one of the first feature maps, thereby obtaining at least one index feature corresponding to each of the grid points.
[0103] The grid feature determination unit 623 is used to determine the grid feature corresponding to each of the plurality of grid points based on at least one index feature corresponding to each of the grid points.
[0104] Optionally, the index determination unit 622 is specifically configured to, for each of the at least one first feature maps, determine the mapping region of the grid point in the first feature map based on the intrinsic parameter matrix, extrinsic parameter matrix of the camera in the multi-camera system corresponding to the first feature map and the three-dimensional coordinates of the grid point; and determine at least one index feature corresponding to the grid point based on the position of the grid point in at least one of the mapping regions corresponding to the at least one first feature map.
[0105] Optionally, when determining the mapping region of the grid point in the first feature map based on the intrinsic parameter matrix, extrinsic parameter matrix of the camera in the multi-camera system corresponding to the first feature map, the index determining unit 622 is used to map the three-dimensional coordinates of the grid point to the image coordinate system corresponding to the first feature map based on the intrinsic parameter matrix and the extrinsic parameter matrix, to obtain the image coordinates of the grid point in the image coordinate system; and determine the mapping region with the image coordinates as the center, combined with a preset length and a preset width.
[0106] Optionally, when the index determining unit 622 maps the three-dimensional coordinates of the grid point to the image coordinate system corresponding to the first feature map based on the intrinsic parameter matrix and the extrinsic parameter matrix to obtain the image coordinates of the grid point in the image coordinate system, it is used to map the three-dimensional coordinates of the grid point to the image coordinate system based on the intrinsic parameter matrix and the extrinsic parameter matrix to obtain the precise coordinates of the grid point in the image coordinate system; and obtain an integer coordinate based on the position of the precise coordinates around the precise coordinates as the image coordinates.
[0107] In some optional embodiments, the grid feature determination unit 623 is specifically configured to perform an expansion operation on each of the first feature maps in at least one first feature map to obtain a plurality of strip features corresponding to each first feature map; based on the plurality of strip features corresponding to each first feature map, determine the strip index feature corresponding to each of the at least one index feature to obtain at least one strip index feature corresponding to each of the grid points; and based on the at least one strip index feature corresponding to each of the grid points, determine the grid feature corresponding to each of the grid points.
[0108] In some alternative embodiments, the grid feature determination unit 623 is specifically used to extract features from at least one index feature corresponding to each of the plurality of grid points based on a self-attention model, so as to obtain the grid feature corresponding to each of the grid points.
[0109] Exemplary electronic devices
[0110] Below, for reference Figure 8 This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device 100 and a second device 200, or a standalone device independent of them, which may communicate with the first and second devices to receive acquired input signals from them.
[0111] Figure 8 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0112] like Figure 8 As shown, the electronic device 80 includes one or more processors 81 and memory 82.
[0113] The processor 81 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 80 to perform desired functions.
[0114] The memory 82 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 81 may execute the program instructions to implement the visual perception methods of the various embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0115] In one example, the electronic device 80 may also include an input device 83 and an output device 84, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0116] For example, when the electronic device is a first device 100 or a second device 200, the input device 83 can be the aforementioned microphone or microphone array for capturing the input signal from the sound source. When the electronic device is a standalone device, the input device 83 can be a communication network connector for receiving the acquired input signals from the first device 100 and the second device 200.
[0117] In addition, the input device 83 may also include, for example, a keyboard, a mouse, etc.
[0118] The output device 84 can output various information to the outside, including determined distance information, direction information, etc. The output device 84 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0119] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device 80 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 80 may include any other suitable components depending on the specific application.
[0120] Exemplary computer program products and computer-readable storage media
[0121] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the visual perception methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0122] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0123] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the visual perception methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0124] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0125] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0126] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0127] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0128] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0129] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0130] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0131] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A visual perception method, comprising: The multi-camera system of the vehicle simultaneously captures multiple images from different perspectives of the surrounding environment of the vehicle, and extracts features from each image to obtain multiple first feature maps. Based on the index features of multiple grid points included in the bird's-eye view at the same time corresponding to at least one first feature map, determine the grid point features corresponding to each of the multiple grid points respectively; Based on the grid features corresponding to each of the grid points, a second feature map corresponding to the bird's-eye view is determined; The second feature map is identified based on the network model corresponding to the preset perception task, and the perception result corresponding to the preset perception task is determined. Based on the index features of multiple grid points included in the bird's-eye view at the same time corresponding to at least one of the first feature maps, determine the grid point features corresponding to each of the multiple grid points, including: Identify the multiple grid points included in the bird's-eye view; Determine the index feature corresponding to each of the multiple grid points in at least one first feature map, and obtain at least one index feature corresponding to each of the grid points; Based on at least one index feature corresponding to each of the grid points, the grid point features corresponding to each of the multiple grid points are determined; the at least one index feature corresponding to each grid point is processed by a neural network model, and all the index features corresponding to the grid point are taken as input to obtain the grid point features corresponding to the grid point.
2. The method according to claim 1, wherein, The step of determining the index feature corresponding to each of the plurality of grid points in at least one of the first feature maps, and obtaining at least one index feature corresponding to each of the grid points, includes: For each of the first feature maps in at least one first feature map, the mapping region of the grid point in the first feature map is determined based on the intrinsic parameter matrix, extrinsic parameter matrix of the camera in the multi-camera system corresponding to the first feature map and the three-dimensional coordinates of the grid point. Based on the position of the grid point in at least one of the mapping regions corresponding to at least one of the first feature maps, at least one index feature corresponding to the grid point is determined.
3. The method according to claim 2, wherein, The step of determining the mapping region of the grid point in the first feature map based on the intrinsic parameter matrix and extrinsic parameter matrix of the camera in the multi-camera system corresponding to the first feature map and the three-dimensional coordinates of the grid point includes: Based on the intrinsic parameter matrix and the extrinsic parameter matrix, the three-dimensional coordinates of the grid points are mapped to the image coordinate system corresponding to the first feature map to obtain the image coordinates of the grid points in the image coordinate system; The mapping area is determined using the image coordinates as the center and in combination with a preset length and a preset width.
4. The method according to claim 3, wherein, The step of mapping the three-dimensional coordinates of the grid points to the image coordinate system corresponding to the first feature map based on the intrinsic parameter matrix and the extrinsic parameter matrix, to obtain the image coordinates of the grid points in the image coordinate system, includes: Based on the intrinsic parameter matrix and the extrinsic parameter matrix, the three-dimensional coordinates of the grid points are mapped to the image coordinate system to obtain the precise coordinates of the grid points in the image coordinate system; Based on the precise coordinates, an integer coordinate is obtained at the position around the precise coordinates as the image coordinates.
5. The method according to any one of claims 1-4, wherein, The step of determining the grid point features corresponding to each of the plurality of grid points based on at least one index feature corresponding to each of the grid points includes: Perform an expansion operation on each of the first feature maps in at least one of the first feature maps to obtain multiple strip features corresponding to each first feature map; Based on multiple strip features corresponding to each of the first feature maps, determine the strip index feature corresponding to each of the index features in at least one of the index features, and obtain at least one strip index feature corresponding to each of the grid points respectively; Based on at least one of the bar index features corresponding to each of the grid points, the grid point features corresponding to each of the grid points are determined.
6. The method according to any one of claims 1-4, wherein, The step of determining the grid point feature corresponding to each of the plurality of grid points based on at least one index feature corresponding to each of the grid points includes: For each of the plurality of grid points, feature extraction is performed on at least one index feature corresponding to each grid point based on a self-attention model to obtain the grid point feature corresponding to each grid point.
7. A visual sensing device, comprising: The feature extraction module is used to extract features from multiple images of the vehicle's surrounding environment captured by the vehicle's multi-camera system at the same time from different perspectives, and obtain multiple first feature maps. The feature correspondence module is used to determine the grid feature corresponding to each of the multiple grid points based on the index feature of the multiple grid points included in the bird's-eye view at the same time corresponding to at least one first feature map determined by the feature extraction module. The feature map determination module is used to determine the second feature map corresponding to the bird's-eye view based on the grid point features determined by the feature correspondence module corresponding to each of the grid points. The perception and recognition module is used to recognize the second feature map determined by the feature map determination module based on the network model corresponding to the preset perception task, and to determine the perception result corresponding to the preset perception task. The feature correspondence module includes: A grid point determination unit is used to determine a plurality of grid points included in the bird's-eye view; An index determination unit is used to determine the index features corresponding to each of the multiple grid points in at least one first feature map, thereby obtaining at least one index feature corresponding to each of the grid points. A grid feature determination unit is used to determine the grid feature corresponding to each of the multiple grid points based on at least one index feature corresponding to each of the grid points; and to process at least one index feature corresponding to each grid point through a neural network model, taking all index features corresponding to the grid point as input and outputting the grid feature corresponding to the grid point.
8. A computer-readable storage medium storing a computer program for performing the visual perception method according to any one of claims 1-6.
9. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the visual perception method according to any one of claims 1-6.