Data transmission method and application apparatus
By receiving image contour information to generate geometric information, the problem of indistinct contours in sparse perception data scenarios is solved, improving perception accuracy, protecting privacy, and reducing data transmission volume.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-10-27
- Publication Date
- 2026-06-04
Smart Images

Figure CN2025130331_04062026_PF_FP_ABST
Abstract
Description
Data transmission method and application device
[0001] This application claims priority to Chinese Patent Application No. 202411718830.8, filed on November 26, 2024, entitled “Data Transmission Method and Application Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communication technology, and in particular to a data transmission method and application device. Background Technology
[0003] With the development of communication technology, communication systems composed of network equipment (such as base stations) and at least one terminal device (such as a user terminal or vehicle terminal) can acquire perception data of the surrounding environment to enable the terminal device to perform perception functions such as detection, localization, recognition, and imaging of target objects. However, in scenarios where perception data is sparse, contour information is not obvious, making it difficult to extract the geometric information of the target object, thus hindering the realization of perception functions. Summary of the Invention
[0004] This application discloses a data transmission method and application device. By using image contour information to assist in the generation of geometric information from perception data, the representation ability and reconstruction accuracy of perception can be improved. It also provides a significant improvement in accuracy in sparse perception scenarios where the contours are not obvious.
[0005] In a first aspect, embodiments of this application disclose a data transmission method, which can be applied to a first device, which can be a terminal device or a network device. The method includes: receiving image contour information; generating geometric information based on the image contour information, as well as perceptual data and / or first information; wherein the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces, and the geometric information is used to indicate M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perceptual space and the boundary points of each of the M three-dimensional surfaces, and the first information is generated based on the perceptual data. Thus, by using image contour information to assist perceptual data in generating geometric information, the representational capability and reconstruction accuracy of perception can be improved, and significant accuracy improvement is also achieved in sparse perception scenarios where contours are not obvious. Furthermore, the feature information of the transmitted data, compared to the transmitted data itself, can protect user privacy and reduce the amount of data transmitted.
[0006] M represents the number of two-dimensional surfaces, and this application does not limit the size of M. Optionally, M can be greater than or equal to 1. When M is greater than 1, boundary points in multiple two-dimensional surfaces can be obtained, and the boundary points in the corresponding three-dimensional surface of each two-dimensional surface can be obtained, which is beneficial for constructing three-dimensional surfaces.
[0007] Optionally, the image contour information includes the (pixel) positions of boundary points in each of the M two-dimensional surfaces. Thus, the boundary points in each of the M two-dimensional surfaces, and the M two-dimensional surfaces themselves, are determined by the positions of the boundary points in each of the M two-dimensional surfaces. This application does not limit the number of boundary points in the two-dimensional surfaces; the number of boundary points in each of the M two-dimensional surfaces can be greater than or equal to 3. That is, if the number of identified boundary points is less than 3, the two-dimensional surfaces corresponding to these boundary points do not need to be obtained. If the number of identified boundary points is greater than or equal to 3, a two-dimensional surface can be determined based on at least 3 of these boundary points. For example, the expression for the two-dimensional surface determined by the 3 boundary points A, B, and C is A + B + C = 0.
[0008] Optionally, the image contour information may also include feature points of each of the M two-dimensional surfaces, excluding boundary points. This facilitates the acquisition of more feature information.
[0009] Optionally, the positions of boundary points in a two-dimensional surface can be represented using an array or a bitmap. Similarly, the positions of boundary points in a three-dimensional surface can be represented using an array or a bitmap.
[0010] Optionally, the image contour information may also include at least one of the following: the number of two-dimensional surfaces M, the number of boundary points in each two-dimensional surface, and the image size of each two-dimensional surface. The image size may include the dimensions of the shape constructed by the boundary points in the two-dimensional surface; if the image is a rectangular image, the image size may include the width and height of the rectangle.
[0011] In conjunction with the first aspect, in some feasible examples, the first information includes a registration result and / or a set of three-dimensional feature points generated based on the perceived data and the three-dimensional feature information set; wherein, the three-dimensional feature information set is generated based on two-dimensional feature pair information, the two-dimensional feature pair information being used to indicate matching pixel pairs between the first image data and the second image data, the three-dimensional feature information set including at least one set of three-dimensional feature information, each set of the three-dimensional feature information including three-dimensional feature points generated from the matching pixel pairs; and one set of three-dimensional feature information in the three-dimensional feature information set including the set of three-dimensional feature points.
[0012] In this embodiment, two-dimensional feature pair information is used to indicate matching pixel pairs between two image data. A matching pixel pair consists of two matching pixels. This application uses two image data as a first image data and a second image data as an example. Optionally, the two-dimensional feature pair information includes the position of the matching pixel pair between the first image data and the second image data. Thus, matching pixel pairs between at least two image data can be determined by the position of the matching pixel pairs between at least two image data.
[0013] Optionally, the positions of the matching pixel pairs can be represented using an array. Alternatively, the positions of the matching pixel pairs can be represented using a bitmap.
[0014] In conjunction with the first aspect, in some feasible examples, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data. The image size of the image data may include the width and height of the image data. Thus, the two-dimensional features of the image data, such as matching pixel pairs between the image data and other image data, can be located based on the image size, which helps improve the accuracy and convenience of data fusion.
[0015] In this embodiment, data fusion can be understood as coordinate alignment, where two-dimensional feature pair information can be obtained based on the first image data. Optionally, when a bitmap represents the matching pixel pairs between the first and second image data, the width and height of the bitmap can be the image size of the first image data. In this way, the second image data can be mapped to the first image data, so that points corresponding to the same spatial location in the two images correspond one-to-one, thereby achieving the purpose of data fusion. The image corresponding to the first image data is used as a reference to determine the pixel points that match the second image data with the first image data.
[0016] In the embodiments of this application, the first image data and the second image data can be images acquired by the same acquisition device, or they can be images acquired by different acquisition devices.
[0017] It should be noted that this application uses first image data and second image data as examples. In reality, multiple image data sets can exist. If the number of image data sets is greater than two, the processing method for the first and second image data sets can be referred to to obtain W blocks of two-dimensional feature pair information. For example, one image data set from multiple image data sets can be matched pairwise with another image data set to obtain W blocks of two-dimensional feature pair information. Here, W is an integer greater than 1, specifically the number of matching two image data sets from multiple image data sets. Each two-dimensional feature pair information set in the W blocks of two-dimensional feature pair information can include N matching pixel pairs, where N is a positive integer.
[0018] Optionally, if image data other than the first image data and the second image data exists, the two-dimensional feature pair information may also include the image size of the image data, and may include the matching pixel pairs obtained by matching the image data with other image data (e.g., the first image data, the second image data, etc.). That is, the two-dimensional feature pair information may include the image size of each image data, and the matching pixel pairs between every two image data.
[0019] In conjunction with the first aspect, in some feasible examples, the registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data; the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data; and the transformation relationship between the perceptual coordinate system and the first camera coordinate system of the perceptual data. Thus, the camera coordinates corresponding to feature points in the image data or perceptual data can be determined based on the registration result. Then, subsequently acquired image data or perceptual data can be fused based on the camera coordinates, which can improve the efficiency of data fusion and improve the accuracy of data processing.
[0020] In this embodiment, the transformation relationship between the first pixel coordinate system and the first camera coordinate system refers to the transformation relationship of the feature point from the first pixel coordinate system to the first camera coordinate system, which can be represented by a first transformation matrix. The transformation relationship between the second pixel coordinate system and the second camera coordinate system refers to the transformation relationship of the feature point from the second pixel coordinate system to the second camera coordinate system, which can be represented by a second transformation matrix. The transformation relationship between the perception coordinate system and the first camera coordinate system refers to the transformation relationship of the feature point from the first camera coordinate system to the perception coordinate system, which can be represented by a third transformation matrix.
[0021] It's understandable that if the coordinate system corresponding to the registration result remains unchanged, data fusion can be performed directly based on that registration result. Otherwise, it's necessary to obtain the registration result again. Obtaining the registration result periodically or triggered by specific events helps improve the accuracy of data fusion.
[0022] Optionally, the registration result can also be used to indicate at least one of the following: the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system. This can improve the efficiency of other data fusion processes.
[0023] In this embodiment, the transformation relationship between the first camera coordinate system and the second camera coordinate system can be represented by a camera transformation matrix, which can be based on the first camera coordinate system. The transformation relationship between the first pixel coordinate system and the perception coordinate system can be obtained by multiplying the first transformation matrix and the third transformation matrix, for example, by multiplying the first transformation matrix and the third transformation matrix. The transformation relationship between the second pixel coordinate system and the perception coordinate system can be obtained by multiplying the second transformation matrix and the third transformation matrix, for example, by multiplying the second transformation matrix and the third transformation matrix. The transformation relationship between the perception coordinate system and the second camera coordinate system refers to the transformation relationship of feature points from the second camera coordinate system to the perception coordinate system.
[0024] It should be noted that the registration results can also be used to indicate the transformation relationship between other coordinate systems, such as the transformation relationship between the first pixel coordinate system and the world coordinate system, the transformation relationship between the second pixel coordinate system and the world coordinate system, etc., without limitation here. In this way, data fusion can be performed based on the transformation relationship between at least two coordinate systems.
[0025] A three-dimensional feature point set can be understood as three-dimensional feature points in the camera coordinate system converted into three-dimensional feature points in the perception coordinate system, or it can be understood as three-dimensional feature points formed by the fusion of two-dimensional feature pairs and perception data. That is, it can be three-dimensional feature points in the camera coordinate system, or it can be three-dimensional feature points in the perception coordinate system, or it can even be three-dimensional feature points in the world coordinate system. There are no restrictions here.
[0026] This application does not limit the method for determining the 3D feature point set. It can calculate the vertical distance between two matching pixels in each matching pixel pair corresponding to each 3D feature point in each set of 3D feature information based on a loss function, and then determine the 3D feature point set based on the vertical distance. For example, the 3D feature points in the set of 3D feature information with the largest minimum vertical distance are the 3D feature point set. Alternatively, based on perceptual data, the 3D feature points in each set of 3D feature information can be fused to obtain the feature points in each set of 3D feature information transformed from the first camera coordinate system to the perceptual coordinate system; then, the chamfer distance between the feature points before and after fusion is determined, and the 3D feature point set is determined based on the chamfer distance. For example, the 3D feature points in the set of 3D feature information with the largest minimum chamfer distance are the 3D feature point set.
[0027] It should be noted that a set of 3D feature points can belong to the optimal or locally optimal set of 3D feature information within a 3D feature information set. That is, the set of 3D feature points contains optimal 3D feature points, and the number or range of these optimal 3D feature points is the largest. 3D feature points in other sets of 3D feature information outside the 3D feature point set may also be optimal, but their number is less than the number of optimal 3D feature points in the 3D feature point set, or their coverage is smaller than the coverage of the optimal 3D feature points in the 3D feature point set. The optimal 3D feature point can be the 3D feature point corresponding to the minimum loss function determined by the vertical distance or chamfer distance.
[0028] In conjunction with the first aspect, in some feasible examples, the 3D feature information set further includes: the number of groups of 3D feature information and / or the number of feature points in each group of 3D feature information. The number of groups of 3D feature information is defaulted to camera parameters, such as the total number of combinations that can be formed from the first pre-selected value group and the second pre-selected value group, or the number of combinations after filtering out some combinations. The filtered combinations can be combinations whose loss function for the corresponding 3D feature points does not meet preset conditions, for example, combinations where the vertical distance or chamfer distance of the corresponding 3D feature points is greater than a preset distance, thus filtering out some combinations with poor 3D features. The number of points in each group of 3D feature information refers to the number of 3D features in that group, which can be understood as the 3D feature points corresponding to the matching pixels in a matching pixel pair. It can be understood that the information of the 3D feature information set can be determined based on the number of groups of 3D feature information and / or the number of points in each group of 3D feature information, which helps improve the accuracy of processing the 3D feature information set.
[0029] In conjunction with the first aspect, in some feasible examples, the three-dimensional feature information set is generated based on the two-dimensional feature pair information, at least one first pre-selected value group, and at least one second pre-selected value group; wherein, the at least one first pre-selected value group is used to indicate pre-selected values of unknown parameters of the first image data, and the at least one second pre-selected value group is used to indicate pre-selected values of unknown parameters of the second image data. Optionally, the unknown parameter of each image data can be the focal length f, for example, the focal length on the U-axis of the pixel coordinate system. Focal length on the V-axis of the pixel coordinate system Furthermore, the unknown parameters of each image data point can also be the physical dimensions of pixels, such as the physical dimension du of a pixel on the U-axis of the pixel coordinate system and the physical dimension dv of a pixel on the V-axis of the pixel coordinate system. In this way, pre-selected values from the image data can be used to assist in determining the three-dimensional feature information set from the two-dimensional feature pair information. For example, the three-dimensional feature information corresponding to different first and second pre-selected value groups in the three-dimensional feature information set can improve the accuracy of obtaining the three-dimensional feature information set.
[0030] Optionally, at least one first preselected value group and at least one second preselected value group can be represented in the form of a list. For example, the list of at least one first preselected value group and at least one second preselected value group may include the initial value, ending value, and step size of the unknown parameter in the first image data and the initial value, ending value, and step size of the unknown parameter in the second image data, etc. Alternatively, the list of at least one first preselected value group and at least one second preselected value group may include a set of preselected values for the unknown parameter on each axis, etc.
[0031] In conjunction with the first aspect, in some feasible examples, the first information includes geometric indication information generated based on the perceived data and the image contour information; wherein the geometric indication information is used to indicate at least one of the following: the number of feature points in each of the three-dimensional surfaces, the index value of each feature point in the perceived data, and each feature point. Optionally, feature points include boundary points and points outside the boundary points, specifically points that can be mapped to a two-dimensional surface. It is understood that sending the index value of each feature point in the perceived data can save resources compared to sending each feature point individually. Compared to sending the index value of each feature point in the perceived data, sending feature points eliminates the need for the second device to determine feature points based on the index value, thereby improving the efficiency of the second device in acquiring feature points.
[0032] Optionally, in some feasible examples, the method may further include receiving the perceived data and / or the first information. That is, the first device receives not only image contour information but also perceived data and / or the first information, such as geometric indication information. Thus, when the first information is geometric indication information, the first device can generate geometric information based on the image contour information and the geometric indication information, thereby improving the efficiency of acquiring geometric information.
[0033] Optionally, the geometric indication information is generated based on image contour information and perceptual data. Alternatively, the geometric indication information can be generated based on image contour information and registration results, or based on image contour information and a set of 3D feature points.
[0034] For example, three-dimensional feature points in the perceptual data are converted into feature points in two-dimensional surfaces, and combined with feature points in the image contour information, so that points on the three-dimensional surfaces fall onto the two-dimensional surfaces, thus obtaining a set of two-dimensional feature points for each two-dimensional surface. The index value of each feature point in each two-dimensional feature point set in the perceptual data is then determined, thereby obtaining the index value of the feature points of each of the M three-dimensional surfaces in the perceptual data. In this way, based on the perceptual data, the feature points of each two-dimensional surface in the image contour information can be registered to obtain geometric indication information.
[0035] It is understandable that, based on image contour information and geometric indicator information, the corresponding three-dimensional surface for each two-dimensional surface in the image contour information can be determined, and the feature points (including boundary points) of each two-dimensional surface can be determined in the three-dimensional surface. Thus, geometric information can be obtained from the boundary points of the three-dimensional surface. For example, based on the feature points of each three-dimensional surface according to the index value, a three-dimensional surface can be fitted, and the boundary points corresponding to the boundary points of the two-dimensional surface in the image contour information can be determined in the three-dimensional surface.
[0036] Optionally, the method may further include: receiving first information. The first information may refer to the foregoing, and may be obtained by the second device that sent the first information, or may be forwarded by the second device to the first device, without limitation.
[0037] In conjunction with the first aspect, in some feasible examples, the method further includes: sending second information, the second information including the geometric information. The second information can be sent to a second device that sends image contour information, or it can be sent to a device other than the second device (such as a data acquisition device or other devices). Thus, the device receiving the second information can perform sensing functions such as target object detection, localization, recognition, and imaging based on the geometric information.
[0038] In conjunction with the first aspect, in some feasible examples, the second information further includes at least one of the following: the number M of the three-dimensional or two-dimensional surfaces, and the number of feature points in each of the three-dimensional or two-dimensional surfaces. Thus, image reconstruction can be achieved based on geometric information and the second information.
[0039] In this embodiment of the application, the second information may not be sent to the acquisition device. Instead, the second information may be sent to the acquisition device without carrying the number M of the three-dimensional or two-dimensional surfaces or the number of feature points in each of the three-dimensional or two-dimensional surfaces, thereby avoiding the acquisition device from repeatedly acquiring this type of information.
[0040] Secondly, embodiments of this application disclose another data transmission method, which can be applied to a second device, which can be a terminal device or a network device. The method includes: acquiring image contour information based on image data; and sending the image contour information. The image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces.
[0041] In conjunction with the second aspect, in some feasible examples, the method further includes: receiving second information; wherein the second information includes the geometric information, the geometric information being used to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perception space and the boundary points of each of the M three-dimensional surfaces.
[0042] In conjunction with the second aspect, in some feasible examples, the second information further includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of feature points in each of the three-dimensional surface or the two-dimensional surface.
[0043] Optionally, the method may further include: sending sensing data and / or first information, wherein the first information is used to indicate feature information generated by the sensing data.
[0044] It should be understood that the second aspect is implemented by the second device. The specific content of the second aspect corresponds to that of the first aspect, and the corresponding features and beneficial effects of the second aspect can be referred to the description of the first aspect. To avoid repetition, detailed descriptions are appropriately omitted here.
[0045] Thirdly, embodiments of this application disclose a communication device, including units, modules, or means for performing the steps of the first aspect, the second aspect, or any of the implementation methods described above. The modules, units, or means can be implemented by software, by hardware, or by a combination of software and hardware.
[0046] Fourthly, embodiments of this application disclose another communication device. This communication device may include one or more processors, which are configured to execute methods described above, either by executing instructions in memory or by using logic circuitry, to perform any of the methods described above or any possible examples.
[0047] In some feasible examples, the communication device may also include interface circuitry, through which the processor communicates with other devices or components.
[0048] In some feasible examples, the communication device also includes the memory.
[0049] In conjunction with the third or fourth aspect, in some feasible examples, the communication device may be a first device or a second device.
[0050] In the embodiments of this application, the first device and the second device can be a terminal device or a network device. The terminal device can be a terminal as a finished product, or a component or module with terminal functions, or a circuit or chip (such as a modem chip, also known as a baseband chip, or a system-on-chip (SoC) chip or system-in-package (SIP) chip containing a modem core), a chip system, or a processor that can be applied to the terminal to perform communication functions. Alternatively, it can be a logical node, logical module, or software that can implement all or part of the terminal functions.
[0051] The network device described in the embodiments of this application may be a network device as a final product, or a component or module with network device functions, or a communication chip (such as a processor, baseband chip, or chip system) that can be applied in a network device.
[0052] Fifthly, embodiments of this application provide a communication system, which includes a first device and a second device. When the first device and the second device are running in the communication system, the first device is used to execute the method described in the first aspect or in feasible examples thereof, and the second device is used to execute the method described in the second aspect or in feasible examples thereof.
[0053] Sixthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed, cause the methods described above or in feasible examples to be implemented.
[0054] In a seventh aspect, embodiments of this application provide a computer program product including instructions that, when executed, cause the methods in any of the above aspects or possible examples to be implemented.
[0055] In conjunction with the sixth or seventh aspect, the instructions can be executed by the computer or the processor in the computer.
[0056] Eighthly, this application provides a chip or chip system including at least one processor for calling and executing instructions stored in a memory, causing a communication device on which the chip or chip system is mounted to perform any of the above-described methods or possible examples.
[0057] Optionally, the chip or chip system may also include memory.
[0058] Ninthly, this application provides another chip, including: an input interface, an output interface, and a processing circuit. The input interface, the output interface, and the processing circuit are connected to the circuit via internal connection paths. The processing circuit is used to execute the method of any of the above aspects or possible examples. Optionally, the chip also includes a memory. The input interface, the output interface, the processor, and the memory are connected via internal connection paths. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method of any of the above aspects or possible examples.
[0059] In a tenth aspect, this application provides another chip system including at least one processor and a communication interface, the communication interface and at least one processor being interconnected via a line, the at least one processor being used to run a computer program or instructions to perform the methods in any of the above aspects or possible examples.
[0060] It should be understood that the implementation and beneficial effects of the above-mentioned aspects can be mutually referenced. Attached Figure Description
[0061] The accompanying drawings used in the embodiments of this application are described below.
[0062] Figures 1A and 1B are schematic diagrams of the architecture of a communication system applied to a data transmission method according to an embodiment of this application;
[0063] Figure 2A is a schematic diagram illustrating the principle of camera imaging according to an embodiment of this application;
[0064] Figure 2B is a schematic diagram of mapping the boundary points of a two-dimensional surface to the boundary points of a three-dimensional surface according to an embodiment of this application;
[0065] Figure 3A is a bitmap of pixels in a two-dimensional surface provided in an embodiment of this application;
[0066] Figure 3B is a bitmap of matching pixel pairs provided in an embodiment of this application;
[0067] Figure 4 is a schematic diagram of W-block two-dimensional feature pair information provided in an embodiment of this application;
[0068] Figure 5 is a schematic diagram of three-dimensional feature information provided in an embodiment of this application;
[0069] Figure 6 is an interactive schematic diagram of a data transmission method provided in an embodiment of this application;
[0070] Figures 7A to 7D are interactive schematic diagrams of another data transmission method provided in the embodiments of this application;
[0071] Figure 8 is a schematic diagram of sensing data provided in an embodiment of this application;
[0072] Figure 9 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;
[0073] Figure 10 is a schematic diagram of another communication device provided in an embodiment of this application;
[0074] Figure 11 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0075] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0076] Please refer to Figure 1A or Figure 1B, which are schematic diagrams of the architecture of a communication system applied to a data transmission method according to an embodiment of this application. The communication system may include, but is not limited to, at least one of the following: Long Term Evolution (LTE) communication system, New Radio (NR) communication system, LTE Advanced (LTE-A) communication system, Device-to-Device (D2D) communication system, Vehicle-to-Everything (V2X) communication system, Machine-to-Machine (M2M) communication system, Internet of Things (IoT) communication system, Narrow Band Internet of Things (NB-IoT) communication system, Integrated Sensing and Communication System, Frequency Division Duplex (FDD) communication system, Time Division Duplex (TDD) communication system, Non-Terrestrial Network (NTN) communication system, Wireless Projection Communication System, Integrated Access and Backhaul (IAB) communication system, Public Land Mobile Network (PLMN) communication system, and Non-Public Network (NPN) communication system. Network (NPN) communication systems, as well as those applied to future communication systems, or non-3rd generation partnership project (3GPP) communication systems, etc.
[0077] As shown in Figure 1A or Figure 1B, the communication system may include a first device 10 and a second device 20. This application does not limit the number of first devices 10 and second devices 20; Figures 1A and 1B illustrate one first device 10 and one second device 20. In practice, the communication system may include one or more first devices 10 and one or more second devices 20.
[0078] Optionally, as shown in Figure 1B, the communication system may further include a data acquisition device, which can be used to acquire image data and / or sensory data, etc. It should be noted that Figure 1B uses a single acquisition device 30 as an example, and the acquisition device 30 and the second device 20 are different devices. In practice, the acquisition device 30 can be either a second device or a first device, such as the second device 20 in Figure 1A, or a communication device not shown in Figures 1A and 1B. When the acquisition device is a second device, the second device can send the image data and / or sensory data acquired by the second device, or characteristic information obtained from data processing, to the first device, which helps to improve data processing efficiency. When the acquisition device is a first device, the first device can process the image data and / or sensory data acquired by the first device, which can improve data processing efficiency.
[0079] This application does not limit the number of image data and sensing data. For example, the number of image data may be two, and the number of sensing data may be one. In the embodiments of this application, all image data and sensing data can be collected by one acquisition device, as shown by the second device 20 in Figure 1A, which can be used as an acquisition device to collect all image data and sensing data; or all image data can be collected by one acquisition device, and all sensing data can be collected by another acquisition device, as shown by the second device 20 in Figure 1B, which can be used as an acquisition device to collect all image data, and the acquisition device 30 in Figure 1B, which is used to collect sensing data; or image data can be collected by at least one acquisition device, and sensing data can be collected by the acquisition device that collects the image data or by other acquisition devices, as shown by the second device 20 in Figure 1B collecting the first image data, the acquisition device 30 in Figure 1B collecting the second image data, and sensing data can be collected by the acquisition device 30 in Figure 1B, the second device 20, or a communication device not shown in Figure 1B, etc. Optionally, the acquisition device for collecting image data can be the same as the acquisition device for collecting sensing data, or the acquisition device for collecting image data can be different from the acquisition device for collecting sensing data.
[0080] Optionally, the data acquisition device can be a terminal device or a data acquisition module within a terminal device. The data acquisition module may include an image acquisition device, such as a camera, for acquiring image data. The data acquisition module may also include a sensing module, such as a module involving radar, Bluetooth, Wi-Fi, 5G, etc., for acquiring sensing data.
[0081] The principle of collecting sensing data can be referenced from the principle of radar. That is, the transmitter sends electromagnetic waves, which are reflected by the object to be sensed and then acquired by the receiver. The acquired reflected signals are further processed into sensing results, which can show information such as the size and outline of the object. The transmitter and receiver that perform the sensing are called sensing nodes.
[0082] Image data can include images, videos, etc., and can be used to identify the two-dimensional image features of objects, such as two-dimensional coordinates and color. Perception data, such as the aforementioned reflection signals and perception results, can include scene information; for example, perception data includes the position, color, and shape of objects inside the vehicle. In this embodiment, image data is an image. Perception data can be point cloud data, also known as laser point cloud (PCD), 3D point cloud, or simply point cloud. It is a collection of massive points representing the spatial distribution and surface characteristics of a target object, obtained by using a laser to acquire the three-dimensional spatial coordinates (usually represented in x, y, z three-dimensional coordinates) of each sampling point on the object's surface in the same spatial reference frame. Compared to images, point clouds, while lacking detailed texture information, contain rich three-dimensional spatial information. In addition to three-dimensional spatial information, point cloud data can also include color information, grayscale values, depth, segmentation results, time delay, etc., which are not limited here.
[0083] This application can be applied to scenarios such as autonomous / assisted driving, V2X, drones, 3D map reconstruction, smart cities, smart homes, factories, healthcare, and maritime sectors. For example, in autonomous driving scenarios, vehicles or drones can generate dynamic maps based on perception data, and / or identify and alert to hazardous events based on perception data.
[0084] This application does not limit the triggering conditions for the acquisition device to collect sensing data. It can be triggered by an event, such as when the sensing coordinate system of the acquisition device changes or the position of the acquisition device changes. Alternatively, it can be triggered by time, for example, by pre-configuring a period, and the acquisition device collects sensing data when the period time arrives.
[0085] As shown in Figure 1A or Figure 1B, the first device 10 can be a network device, and the second device 20 can be a terminal device. In fact, both the first device and the second device can be terminal devices, or both the first device and the second device can be network devices; or the first device can be a terminal device, and the second device can be a network device.
[0086] Terminal devices can connect to network devices wirelessly or via wired connections, enabling uplink (UL) or downlink (DL) communication. Terminal devices can also connect to each other wirelessly or via wired connections, allowing for sidelink (SL) communication.
[0087] The terminal equipment involved in this application is an entity on the user side used to receive or transmit signals, providing voice and / or data to the user. Terminal equipment may also be referred to as a terminal, user equipment (UE), access terminal, UE unit, UE station, mobile device, mobile station, mobile station, mobile terminal, mobile client, mobile unit, remote station, remote terminal, remote unit, wireless unit, wireless communication equipment, user agent, or user device, etc. Among them, the access terminal can be a cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, vehicle-mounted device, wearable device, or terminal in a future communication system, etc. It is sometimes simply referred to as a terminal below. In Figures 1A and 1B and this application document, the terminal equipment is described using a vehicle or UE as an example.
[0088] It should be noted that the terminal device described in the embodiments of this application can be a terminal as a final product, such as the various terminal devices mentioned above, or it can be a component or part with terminal functions, or it can be a communication chip (such as a processor, baseband chip, or chip system, etc.) that can be applied in a terminal. That is to say, components, parts, or chips applied in the above-mentioned devices also belong to terminal devices.
[0089] In Figure 1A or Figure 1B, network devices are exemplified as access network (AN) devices. Access network devices, also known as radio access network (RAN) devices, or simply access networks, are nodes or devices that connect terminal devices to a wireless network. In other words, the access network provides access services to terminal devices, enabling them to access (or connect to) the network. Access networks can support both wired and wireless access.
[0090] Optionally, the access network consists of multiple AN / RAN nodes. AN / RAN nodes can include, but are not limited to: access points (APs), enhanced node Bs (eNBs), home evolved Node Bs (HNBs), baseband units (BBUs), next-generation node Bs (gNBs), transmission reception points (TRPs), transmission points (TPs), or other access nodes, such as wireless relay nodes or wireless backhaul nodes. AN / RAN nodes can be one or more antenna panels, or network nodes constituting gNBs or transmission points, such as BBUs or distributed units (DUs), or devices performing RAN functions in communication systems such as D2D, V2X, M2M, and U2U. AN / RAN nodes can be radio controllers in cloud radio access network (CRAN) scenarios, open RAN (O-RAN or ORAN), or access networks in future communication systems, etc., without any limitations.
[0091] It should be noted that the network device described in the embodiments of this application can be a network device as a final product, such as the various network devices mentioned above, or it can be a component or part with network device functions, or it can be a communication chip (such as a processor, baseband chip, or chip system, etc.) that can be applied in a network device. That is to say, components, parts, or chips applied in the above-mentioned devices also belong to network devices.
[0092] In the communication system shown in Figure 1A or Figure 1B, although the access network and terminal equipment are shown, the application scenario may not be limited to the access network and terminal equipment. For example, it may also include equipment for carrying virtualized network functions, which is obvious to those skilled in the art and will not be described in detail here.
[0093] Furthermore, the number and type of network devices and terminal devices included in the communication system shown in Figure 1A or Figure 1B are merely examples, and the embodiments of this application are not limited thereto. For example, it may also include more or fewer terminal devices communicating with the network devices. As another example, it may also include more or fewer network devices communicating with the terminal devices. For the sake of brevity, they are not described one by one in the accompanying drawings.
[0094] Optionally, the communication system may also include network devices not shown in Figure 1A or Figure 1B, such as core network (CN) devices, data network devices, etc.
[0095] In different communication systems, core network equipment (hereinafter referred to as core network) can correspond to different devices. For example, in a 3G communication system, it can correspond to the Serving GPRS Support Node (SGSN) and / or the Gateway GPRS Support Node (GGSN); in a 4G communication system, it can correspond to the Mobility Management Entity (MME) and / or the Serving Gateway (S-GW); and in a 5G communication system, it can correspond to the aforementioned Policy Control Function (PCF) network elements, Unified Data Management (UDM) network elements, Application Function (AF) network elements, Access and Mobility Management Function (AMF) network elements, Session Management Function (SMF) network elements, Location Management Function (LMF) network elements, and User Plane Function (UPF) network elements, etc.
[0096] Among them, the UPF network element is responsible for managing the transmission of user plane data and quality of service (QoS) control, traffic statistics and other functions. It can perform user data packet forwarding according to the routing rules of the session management network element, such as sending uplink data to the data network or other user plane network elements, and forwarding downlink data to other user plane network elements or (R)AN network elements.
[0097] The AMF (Access Default Mode) network element is responsible for user access management, security authentication, and mobility management. The LMF (Local Mode Default Mode) network element manages and controls location service requests from target terminals and processes location-related information. The SMF (Supply, Service Default Mode) network element manages sessions, allocating and releasing resources for terminal device sessions. The UDM (User Default Mode) network element manages the context of user subscriptions, such as storing terminal device subscription information. The PCF (Policy and Charging Rules Function) network element is responsible for user policy management. Similar to the Policy and Charging Rules Function (PCRF) network element in LTE, it is primarily responsible for policy authorization, quality of service (QoS), and generating charging rules, and distributing these rules to the UPF (User Default Mode) network element via the SMF network element to complete the installation of the corresponding policies and rules. The AF (Application Default Mode) network element can be a third-party application control platform or the operator's own equipment. The AF network element is responsible for application management and can provide services to multiple application servers.
[0098] In this embodiment, the data network device is hereinafter referred to as the data network. The data network is used to provide business services to users. Generally, the client is a terminal, and the server is the data network. The data network provided by the data network may include a private network, such as a local area network (LAN). The data network may also include an external network not managed by an operator, such as the Internet. Alternatively, the data network may include a proprietary network jointly deployed by operators, such as a network providing Internet Protocol Multimedia Subsystem (IMS) services.
[0099] In some embodiments, the first device, the second device, the acquisition device, the network device, and the terminal device may also be referred to as communication devices, which may be a general-purpose device or a special-purpose device. The embodiments of this application do not specifically limit this.
[0100] In this embodiment, the terminal device or network device includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also referred to as main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Furthermore, this embodiment does not specifically limit the specific structure of the execution entity of the method provided in this embodiment, as long as it can communicate according to the method provided in this embodiment by running a program that records the code of the method provided in this embodiment. For example, the execution entity of the method provided in this embodiment can be a terminal device or a network device, or a functional module in the terminal device or network device that can call and execute a program.
[0101] Furthermore, various aspects or features of this application can be implemented as methods, apparatus, or articles of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" as used herein encompasses a computer program accessible from any computer-readable device, carrier, or medium. For example, computer-readable media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical discs (e.g., compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memory (EPROMs), cards, sticks, or key drives, etc.). The various storage media described herein may represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0102] To facilitate understanding of the embodiments of this application, definitions of technical terms that may appear in the embodiments of this application are given below. The terminology used in the implementation section of this application is only used to explain specific embodiments of this application and is not intended to limit this application.
[0103] 1. The process of image formation is essentially a transformation of several coordinate systems.
[0104] Please refer to Figure 2A, which is a schematic diagram illustrating the principle of camera imaging according to an embodiment of this application. As shown in Figure 2A, the camera coordinate system is based on the optical center O. c With the origin as the axis, the direction of the optical axis as the z-axis, and the x and y directions parallel to the image as the x and y axes, respectively, the x, y, and z axes can be referred to as X, Y, Z, and X, respectively. c Y c Z c The unit is length. The pixel coordinate system takes the vertex (u, v) of the image as the origin, and the u and v directions are parallel to the x and y directions, respectively. Its x-axis and y-axis can be called the U-axis and V-axis, respectively, and the unit is pixels.
[0105] In this embodiment of the application, the transformation relationship between the pixel coordinate system and the camera coordinate system can be obtained based on the following formula (1), and the schematic diagram can be shown in Figure 2A.
[0106] Where u is the coordinate of the pixel (or pixel point or feature point) on the U-axis of the pixel coordinate system, and v is the coordinate of the pixel on the V-axis of the pixel coordinate system. i It is a point (i.e., the aforementioned pixel) in the camera coordinate system X c The coordinates on the x-axis, y i It is the Y-axis of the point in the camera coordinate system. c The coordinates on the (y) axis, z i It is the Z-axis of the point in the camera coordinate system. c The coordinates on the (z) axis. K can be called the transformation matrix from the pixel coordinate system to the camera coordinate system. This transformation matrix is used to indicate the transformation relationship from 2D pixel coordinates to 3D camera coordinates. K can be represented by the following equation (2):
[0107] Where w is the width of the image and h is the height of the image. u It is the X-axis of the point in the camera coordinate system. c The focal length on the (x) axis, f v It is the point in the Y-axis of the camera coordinate system c Focal length on the (y) axis. du is the physical size of a pixel on the U-axis (u direction) of the pixel coordinate system, and dv is the physical size of a pixel on the V-axis (v direction) of the pixel coordinate system.
[0108] In this embodiment of the application, the transformation relationship between the camera coordinate system and the perception coordinate system can be obtained based on the following formula (3) so that the points in the camera coordinate system are displayed in the form of a three-dimensional point cloud.
[0109] Where, x iy i and z i As mentioned above, this will not be repeated here. si It is the x-coordinate of a point on the perceptual coordinate system, y si It is the y-coordinate of the point on the perceptual coordinate system, z-coordinate. si M is the coordinate of the point on the z-axis of the sensing coordinate system. M can be called the transformation matrix from the camera coordinate system to the sensing coordinate system. This transformation matrix is used to indicate the transformation relationship from 3D camera coordinates to 3D sensing coordinates. M can be represented by the following equation (4):
[0110] Where R is a 3x3 rotation matrix, and T is a 3x1 translation matrix, or translation vector.
[0111] 2. Image contour information refers to the contour or boundary information of an object in an image. In the embodiments of this application, image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces. Here, M represents the number of two-dimensional surfaces, and this application does not limit the size of M. Optionally, M can be greater than or equal to 1. When M is greater than 1, boundary points in multiple two-dimensional surfaces can be obtained, and the boundary points in the corresponding three-dimensional surface of each two-dimensional surface can be obtained, which is beneficial for constructing three-dimensional surfaces.
[0112] This application does not limit the number of boundary points in a two-dimensional surface; the number of boundary points for each of the M two-dimensional surfaces can be greater than or equal to 3. That is, if the number of identified boundary points is less than 3, the two-dimensional surfaces corresponding to these boundary points do not need to be obtained. If the number of identified boundary points is greater than or equal to 3, a two-dimensional surface can be determined based on at least 3 of these boundary points. For example, the expression for the two-dimensional surface determined by the 3 boundary points A, B, and C is A + B + C = 0.
[0113] This application uses a two-dimensional surface as an example. Referring to Figure 2B, the four boundary points A, B, C, and D in a two-dimensional surface of the two-dimensional image can be determined according to the camera coordinate system. This two-dimensional image includes M two-dimensional surfaces.
[0114] Optionally, the image contour information includes the (pixel) positions of boundary points in each of the M two-dimensional surfaces. Thus, the boundary points in each of the M two-dimensional surfaces, and the M two-dimensional surfaces themselves, are determined by the positions of the boundary points in each of the M two-dimensional surfaces.
[0115] The positions of boundary points in a two-dimensional surface can be represented using an array. For example, the positions of boundary points in M two-dimensional surfaces are {{A1, B1, ...}, {A2, B2, ...}, ..., {A...}. M B M...}}. Here, the data set within each curly brace of the array representing the positions of boundary points in M two-dimensional surfaces represents the position of each boundary point in a given two-dimensional surface. {A1, B1, ...} represents the position of each boundary point in the first two-dimensional surface, and so on, {A...}... M B M , ...} represents the position of each boundary point in the Mth two-dimensional surface.
[0116] Alternatively, the positions of boundary points in a two-dimensional surface can be represented using a bitmap.
[0117] Optionally, the image contour information may also include feature points of each of the M two-dimensional surfaces, excluding boundary points.
[0118] For example, please refer to Figure 3A, which is a bitmap of pixels in a two-dimensional surface provided in an embodiment of this application. This bitmap is used to indicate the positions of boundary points in a two-dimensional surface, and also to indicate the positions of pixels in the two-dimensional surface other than the boundary points. As shown in Figure 3A, the rectangular image corresponding to the two-dimensional surface can be processed into multiple grids according to its width (w) and height (h), with each grid representing the position of a pixel in the rectangular image. If the value in the grid is 1, it indicates that a pixel exists at that position. If the value in the grid is 0, it indicates that a pixel does not exist at that position (or was not detected). For example, A... 11 This indicates that a pixel exists at this location, A 14 and A 41 These indicate that no pixels were captured at that location. Thus, the image contour of a two-dimensional surface can be determined based on the boundary pixels of the rectangular image, as shown in Figure 3A, A. 11 A 21 A 32 A 23 A 12 A 13 The image outline formed by connecting the corresponding pixels.
[0119] It should be noted that the pixels in the two-dimensional surface shown in Figure 3A include not only boundary points but also pixels outside the boundary points. These pixels can be referred to as feature points of the two-dimensional surface. This facilitates the acquisition of more feature information. In some other feasible examples, the bitmap of the two-dimensional surface in the image contour information can only show the boundary points of the two-dimensional surface. For example, boundary points are represented by 1 in the bitmap, and other pixels are represented by 0, thereby determining the location of the boundary points of the two-dimensional surface.
[0120] Optionally, the image contour information may also include at least one of the following: the number of two-dimensional surfaces M, the number of boundary points in each two-dimensional surface, and the image size of each two-dimensional surface. Further, the image contour information may also include the number of feature points in each two-dimensional surface.
[0121] The image size can include the dimensions of the shape constructed from the boundary points in the two-dimensional plane. For example, as shown in Figure 2B, if the shape constructed from the boundary points in the two-dimensional plane is a rectangular image, then the image size can include the width (w) and height (h) of the rectangle. The number of boundary points in each two-dimensional plane can be indicated by an array, such as {N1, N2, ..., N...}. M}, N1 represents the number of boundary points in the first two-dimensional surface, and so on, N M Let N be the number of boundary points in the Mth two-dimensional surface. This application does not limit the number of boundary points in each two-dimensional surface; for example, N1, N2, ..., N... M Any two of them can be equal or unequal.
[0122] 3. Two-dimensional feature pair information refers to the two-dimensional feature information of matching pixel pairs in at least two image data (at least two images in this application). In the embodiments of this application, two-dimensional feature pair information is used to indicate matching pixel pairs between two image data. A matching pixel pair consists of two matching pixels. This application does not limit the matching method of pixels between image data (or images). The following example uses two image data as the first image data and the second image data. The first image data and the second image data can be input into a local feature transformer (LoFTR) matching model to obtain the matching pixel pairs between the first image data and the second image data.
[0123] The matching pixel pair between the first image data and the second image data includes a first pixel in the first image data and a second pixel in the second image data that matches the first pixel. That is, if the first pixel has no matching pixel in the second image data, the matching pixel pair between the first and second image data does not include the matching pixel pair of the first pixel. If the first pixel has a matching pixel in the second image data, and that matching pixel is the second pixel, then the matching pixel pair between the first and second image data includes the matching pixel pair of the first pixel, and this matching pixel pair includes both the first and second pixel.
[0124] It should be noted that the above example uses the matching pixel pair of the first pixel in the first image data. When the first pixel and the second pixel match, the matching pixel pair of the first pixel can also be referred to as the matching pixel pair of the second pixel.
[0125] Optionally, the two-dimensional feature pair information includes the positions of matching pixel pairs between the first image data and the second image data. Thus, matching pixel pairs between at least two image data can be determined by the positions of matching pixel pairs between at least two image data.
[0126] Optionally, the positions of the matching pixel pairs can be represented using an array.
[0127] When the number of matching pixel pairs is greater than 1, for example, if the number of matching pixel pairs is N, then N matching pixel pairs can be represented as an array, such as {(p 11 p 21 ), ..., (p 1N p 2N )}. Each set of parentheses represents a matching pixel pair, and the two values within the parentheses are the positions of the two pixels in that matching pixel pair within the pixel coordinate system of the image data. (p 11 p 21 (p) represents the first matching pixel pair, and so on, (p) 1N p 2N ) represents the Nth matching pixel pair.
[0128] With (p 11 p 21 (p) provides an example. 11 p 21 p in ) 11 This can be the position of the first pixel in the first image data within the pixel coordinate system of the first image data, for example, p 11 This includes the position of the first pixel on the u-axis of the pixel coordinate system of the first image data and the position of the first pixel on the v-axis of the pixel coordinate system of the first image data. 21 This can be the position of the second pixel in the second image data that matches the first pixel in the second image data within the pixel coordinate system of the second image data, for example, p 21 This includes the position of the second pixel on the u-axis of the pixel coordinate system of the second image data and the position of the second pixel on the v-axis of the pixel coordinate system of the second image data.
[0129] Alternatively, the positions of matching pixel pairs can be represented using a bitmap.
[0130] For example, please refer to Figure 3B, which is a bitmap of matched pixel pairs provided in an embodiment of this application. This bitmap is used to indicate the position of the matched pixel pair between the first image data and the second image data in the pixel coordinate system of the first or second image data. As shown in Figure 3B, the rectangular image can be processed into multiple grids according to the width (w) and height (h) of the rectangular image corresponding to the first or second image data. Each grid represents the position of a pixel in the rectangular image. If the value in the grid is 1, it indicates that the pixel at that position is a matched pixel between the first and second image data. If the value in the grid is 0, it indicates that the pixel at that position is a mismatched pixel between the first and second image data. (e.g., P) 11P represents the position of a pixel that matches between the first image data and the second image data. 14 and P 41 These represent the positions of a pixel that does not match between the first image data and the second image data.
[0131] In some feasible examples, two-dimensional feature pair information can also be used to indicate the image size of the first image data and / or the image size of the second image data. The image size of the image data can include its width and height. As shown in Figure 3B, the two-dimensional feature pair information to which the matching pixel pair in the bitmap representation belongs includes the width and height of the image size. Thus, the two-dimensional features of the image data, such as matching pixel pairs between this image data and other image data, can be located based on the image size, which helps improve the accuracy and convenience of data fusion.
[0132] In this embodiment, data fusion can be understood as coordinate alignment. The pixel coordinate system of the first image data can be called the first pixel coordinate system, and the pixel coordinate system of the second image data can be called the second pixel coordinate system. The camera coordinate system of the first image data can be called the first camera coordinate system, and the camera coordinate system of the second image data can be called the second camera coordinate system. The transformation matrix between the first pixel coordinate system and the first camera coordinate system can be called the first transformation matrix, such as K1, which is used to indicate the transformation relationship of feature points from the first pixel coordinate system to the first camera coordinate system. The transformation matrix between the second pixel coordinate system and the second camera coordinate system can be called the second transformation matrix, such as K2, which is used to indicate the transformation relationship of feature points from the second pixel coordinate system to the second camera coordinate system. The first pixel coordinate system can be understood as the coordinate system of pixels in the first image data, and the second pixel coordinate system can be understood as the coordinate system of pixels in the second image data. The first camera coordinate system can be understood as the camera coordinate system when the acquisition device acquires the first image data, and the second camera coordinate system can be understood as the camera coordinate system when the acquisition device acquires the second image data.
[0133] In this embodiment, two-dimensional feature pair information can be obtained based on the first image data. Optionally, when a bitmap represents the matching pixel pairs between the first and second image data, the width and height of the bitmap can be the image size of the first image data. In this way, the second image data can be mapped to the first image data, so that points corresponding to the same spatial location in the two images correspond one-to-one, thereby achieving data fusion. The image corresponding to the first image data is used as a reference to determine the matching pixels between the second and first image data.
[0134] It is understandable that the first camera coordinate system and the second camera coordinate system can be different, and the first pixel coordinate system and the second pixel coordinate system can also be different. Since the first image data and the second image data are different, an image with overlapping areas between the first image data and the second image data is needed to obtain the matching pixels between them.
[0135] In the embodiments of this application, the first image data and the second image data can be images acquired by the same acquisition device, or they can be images acquired by different acquisition devices.
[0136] It should be noted that this application uses first image data and second image data as examples. In reality, multiple image data sets can exist. If the number of image data sets is greater than two, the processing method for the first and second image data sets can be referred to to obtain W blocks of two-dimensional feature pair information. For example, one image data set from multiple image data sets can be matched pairwise with another image data set to obtain W blocks of two-dimensional feature pair information. Here, W is an integer greater than 1, specifically the number of matching two image data sets from multiple image data sets. Each two-dimensional feature pair information set in the W blocks of two-dimensional feature pair information can include N matching pixel pairs, where N is a positive integer.
[0137] Taking an example with three image data sets, the image data includes a first image data set, a second image data set, and a third image data set. The third image data set can be matched with the first image data set, and it can also be matched with the second image data set. As shown in Figure 4, the first two-dimensional feature pair information obtained by matching the first image data set and the second image data set is as follows: {(p 11 p 21 ), ..., (p 1N p 2N The second two-dimensional feature pair information obtained by matching the first image data and the third image data is as follows: {(p 11 p 21 ), ..., (p 1N p 2N The second and third image data can be matched to obtain a third block of two-dimensional feature pairs, such as {(p 21 p 31 ), ..., (p 2N p 3N )}.
[0138] Optionally, if the number of matching pixel pairs in one of the W blocks of two-dimensional feature pairs is less than N, it can be indicated by default non-matching pixel pairs (e.g., (-1, -1)). If the number of matching pixel pairs in one of the W blocks of two-dimensional feature pairs is greater than N, N matching pixel pairs can be selected from that block of two-dimensional feature pairs for indication. Alternatively, the value of N can be increased so that each block of two-dimensional feature pairs in the W blocks includes the increased number of N matching pixel pairs.
[0139] Optionally, if the overlap area between the first image data and the second image data is less than a first threshold, the number of image data can be greater than 2. If the overlap area between the first image data and the second image data is greater than the first threshold, the number of image data can be equal to 2.
[0140] This application does not limit the first threshold, for example, 80%. It can be understood that the greater the overlap between two image data sets, the more matching pixel pairs there are, and the more features can be fused. Conversely, the smaller the overlap between two image data sets, the fewer matching pixel pairs there are, and the fewer features can be fused. That is, when the overlap between the first and second image data sets is greater than the first threshold, most of the image features can be obtained based on the first and second image data sets. When the overlap between the first and second image data sets is less than the first threshold, the image data can include other image data or perceptual data besides the first and second image data sets to obtain more fusion features, which is beneficial for subsequent feature fusion.
[0141] Optionally, if image data other than the first image data and the second image data exists, the two-dimensional feature pair information may also include the image size of the image data, and may include the matching pixel pairs obtained by matching the image data with other image data (e.g., the first image data, the second image data, etc.). That is, the two-dimensional feature pair information may include the image size of each image data, and the matching pixel pairs between every two image data.
[0142] 4. Three-dimensional feature information, used to indicate the three-dimensional features of feature points.
[0143] In this embodiment, the three-dimensional feature information can be generated based on two-dimensional feature pair information, used to indicate the three-dimensional features corresponding to at least two two-dimensional image data. The three-dimensional feature information includes three-dimensional feature points corresponding to matching pixel pairs, which can be used to represent the depth information of the matching pixels in the matching pixel pair, and can be understood as the coordinates of the matching pixels on the z-axis. This application does not limit the method for generating the three-dimensional feature information; optionally, the transformation matrix M between the first camera coordinate system and the second camera coordinate system can be obtained. cam Based on M cam Obtain the three-dimensional feature information corresponding to the two-dimensional feature pair information.
[0144] Among them, M cam Used to indicate the transformation relationship between the first camera coordinate system and the second camera coordinate system. This application is for obtaining M cam The method is not limited, M cam This can be obtained using computer vision techniques, such as Pycolmap and OpenCV. Optionally, the transformation matrix M between the first and second camera coordinate systems... cam The first camera coordinate system can be used as a reference. That is, under the first camera coordinate system, the point cloud features corresponding to each pixel in the matching pixel pair are determined, thereby obtaining the three-dimensional feature information.
[0145] Based on M cam A schematic diagram illustrating the acquisition of 3D feature information corresponding to 2D feature pairs can be found in Figure 5. In Figure 5, cam1 represents the first camera coordinate system, and cam2 represents the second camera coordinate system. 1i v 1i (u) represents a pixel in the first image data. 2i v 2i ) indicates that in the second image data, (u) 1i v 1i The matched pixels. d represents (u 1i v 1i ) and (u 2i v 2i The vertical distance between (x) and (x). 1i y 1i , z 1i ) represents (u 1i v 1i The coordinates are in the first camera coordinate system. As shown in Figure 5, using the first camera coordinate system as a reference, the point cloud features corresponding to the matching pixel pairs between the first image data and the second image data can be obtained, thereby obtaining the three-dimensional feature information between the first image data and the second image data.
[0146] Optionally, the first transformation matrix may have at least one first preselected value group, and the second transformation matrix may have at least one second preselected value group. The at least one first preselected value group is used to indicate preselected values of unknown parameters of the first image data, and the at least one second preselected value group is used to indicate preselected values of unknown parameters of the second image data.
[0147] In this embodiment, the unknown parameter for each image data can be the focal length f, for example, the focal length on the U-axis of the pixel coordinate system. Focal length on the V-axis of the pixel coordinate system And so on. The unknown parameters for each image data point can also be the physical size of the pixels, such as the physical size *du* of the pixel on the U-axis of the pixel coordinate system, or the physical size *dv* of the pixel on the V-axis of the pixel coordinate system. Thus, given the physical size of the pixels, according to the aforementioned formula... The physical dimensions of pixels can determine an unknown focal length; or, given a known focal length, the physical dimensions of unknown pixels can be determined using this formula and the focal length. This application uses focal length and pixel physical dimensions as examples of unknown parameters of image data; in reality, unknown parameters can also be other information.
[0148] Optionally, at least one first preselected value group and at least one second preselected value group can be represented in the form of a list.
[0149] For example, the list of at least one first pre-selected value group and at least one second pre-selected value group may include the initial value, ending value, and step size of the unknown parameter in the first image data, and the initial value, ending value, and step size of the unknown parameter in the second image data, etc. Taking the focal length as an example, the list of at least one first pre-selected value group and at least one second pre-selected value group is as follows: {f u1_0 f v1_0 f u2_0 f v2_0 f u1_n f v1_n f u2_n f v2_n s u1 s v1 s u2 s v2}wait.
[0150] Among them, f u1_0 This can represent the initial value of the focal length on the U-axis of the first pixel coordinate system, f. v1_0 This can represent the initial value of the focal length on the V-axis of the first pixel coordinate system. u1_n f can represent the ending value of the focal length on the U-axis of the first pixel coordinate system. v1_n This can represent the ending value of the focal length on the V-axis of the first pixel coordinate system. s u1The step size s can represent the focal length on the U-axis of the first pixel coordinate system. v1 This can represent the step size of the V-axis focal length in the first pixel coordinate system. u2_0 This can represent the initial value of the focal length on the U-axis of the second pixel coordinate system, f. v2_0 This can represent the initial value of the focal length on the V-axis of the second pixel coordinate system. u2_n f can represent the ending value of the focal length on the U-axis of the second pixel coordinate system. v2_n This can represent the ending value of the focal length on the V-axis of the second pixel coordinate system. s u2 The step size s can represent the focal length on the U-axis of the second pixel coordinate system. v2 It can represent the step size of the focal length on the V-axis of the second pixel coordinate system.
[0151] It is understandable that, when the initial and final values of the unknown parameters are unequal, multiple unknown parameters of the first image data are determined based on the initial, final, and step values of the unknown parameters of the first image data. Similarly, multiple unknown parameters of the second image data can be determined based on the initial, final, and step values of the unknown parameters of the second image data. When the initial and final values of the unknown parameters of the first image data are equal, the number of groups (or individuals) in the first pre-selected value group or the number of candidate values for the unknown parameters of the first image data is 1. Alternatively, when the step size of the unknown parameters of the first image data is 0, and only the initial or final values of the unknown parameters of the first image data are included, the number of groups (or individuals) in the first pre-selected value group or the number of candidate values for the unknown parameters of the first image data is 1. Similarly, the number of groups (or individuals) in the second pre-selected value group or the number of candidate values for the unknown parameters of the second image data can be determined.
[0152] For example, a list of at least one first set of preselected values and at least one second set of preselected values, or may include a set of preselected values for the unknown parameter on each axis, such as etc. Among them, This represents at least one pre-selected value of the focal length on the U-axis of the first pixel coordinate system. This represents at least one pre-selected value of the focal length on the V-axis of the first pixel coordinate system. u2_0 , ...} represents at least one pre-selected value of the focal length on the U-axis of the second pixel coordinate system, {f v2_0 , ...} represents at least one pre-selected value of the focal length on the V-axis of the second pixel coordinate system.
[0153] It is understandable that, given that the first transformation matrix has at least one first pre-selected value group and the second transformation matrix has at least one second pre-selected value group, it is possible to base it on M. camFor each first pre-selected value group and each second pre-selected value group, the point cloud features corresponding to the matching pixel pairs in the two-dimensional feature pair information are obtained to obtain a three-dimensional feature information set. This three-dimensional feature information set includes at least one set of three-dimensional feature information. Each set of three-dimensional feature information is two-dimensional feature pair information (or M determined by the two-dimensional feature pair information). cam The three-dimensional features generated are different from those generated by a first set of pre-selected values and a second set of pre-selected values. In other words, the first set of pre-selected values and / or the second set of pre-selected values are different between any two sets of three-dimensional feature information.
[0154] If the number of groups in the first preselected value group is 1 and the number of groups in the second preselected value group is 1, then the number of groups in the three-dimensional feature information is also 1, meaning the three-dimensional feature information set includes one set of three-dimensional feature information. Otherwise, if the number of groups in the first preselected value group is greater than 1, or the number of groups in the second preselected value group is greater than 1, then the number of groups in the three-dimensional feature information is also greater than 1, meaning the three-dimensional feature information set can include at least two sets of three-dimensional feature information.
[0155] In some feasible examples, the 3D feature information set may also include the number of groups of 3D feature information and / or the number of points in each group of 3D feature information.
[0156] The number of 3D feature information groups is assumed to be the camera parameters, such as the total number of combinations that can be formed from the first and second pre-selected value groups, or the number of combinations after filtering out some combinations. The filtered combinations can be those whose loss function for the corresponding 3D feature points does not meet preset conditions, such as combinations where the vertical distance or chamfer distance between the corresponding 3D feature points is greater than a preset distance, thus filtering out some combinations with poor 3D features. The number of points in each group of 3D feature information refers to the number of 3D features in that group, which can be understood as the 3D feature points corresponding to the matching pixels in a matching pixel pair. It can be understood that the number of 3D feature information groups and / or the number of points in each group can determine the information of the 3D feature information set, which helps improve the accuracy of processing the 3D feature information set.
[0157] 5. A three-dimensional feature point set can be understood as three-dimensional feature points in the camera coordinate system converted into three-dimensional feature points in the perception coordinate system, or it can be understood as three-dimensional feature points formed by the fusion of two-dimensional feature pairs and perception data. That is, it can be three-dimensional feature points in the camera coordinate system, or it can be three-dimensional feature points in the perception coordinate system, or it can even be three-dimensional feature points in the world coordinate system. There are no restrictions here.
[0158] In some feasible examples, the 3D feature point set includes 3D feature points from a set of 3D feature information in the 3D feature information set. That is, the 3D feature point set is the set of 3D feature points corresponding to that set of 3D feature information.
[0159] This application does not limit the method for determining the three-dimensional feature point set. The vertical distance between two matching pixels in each matching pixel pair corresponding to each three-dimensional feature point in each set of three-dimensional feature information can be calculated based on a loss function, as shown by d in Figure 5. The three-dimensional feature point set is then determined based on the vertical distance. For example, the three-dimensional feature points in the set of three-dimensional feature information with the largest minimum vertical distance constitute the three-dimensional feature point set. The loss function can be Mahalanobis distance loss, Euclidean distance loss, Chebyshev discrepancy loss, Hamming distance loss, etc., and is not limited here.
[0160] Alternatively, based on the perceived data, the 3D feature points in each group of 3D feature information in the 3D feature information set can be fused to obtain the feature points in each group of 3D feature information transformed from the first camera coordinate system to the perceived coordinate system; then, the chamfer distance of the feature points before and after fusion can be determined, and the 3D feature point set can be determined based on the chamfer distance. For example, the 3D feature points in the group of 3D feature information with the largest minimum chamfer distance are the 3D feature point set.
[0161] Here, the sensing coordinate system can be understood as the coordinate system of the acquisition device when acquiring sensing data. Based on the sensing data, the method for fusing 3D feature points in each group of 3D feature information in the 3D feature information set can be exemplified using a group of 3D feature information, such as the i-th group of 3D feature information. For example, the sensing data and the i-th group of 3D feature information are matched to obtain the transformation relationship between the first camera coordinate system and the sensing coordinate system, and / or the relative scaling ratio s of the 3D feature points in the sensing data and the i-th group of 3D feature information; based on this transformation relationship and / or s, the 3D feature points of the i-th group of 3D feature information transformed from the first camera coordinate system to the sensing coordinate system are determined.
[0162] In this embodiment, the transformation relationship between the first camera coordinate system and the perception coordinate system can be represented by a transformation matrix between the perception coordinate system of the perception data and the first camera coordinate system. This transformation matrix can be called the third transformation matrix, such as M mentioned above, and is used to indicate the transformation relationship of feature points from the first camera coordinate system to the perception coordinate system. M and s can be obtained based on the maximum likelihood function, which can be the correlation function of Pycpd, such as L(M,s|X,Y). Here, X is the perception data, and Y is the i-th group of three-dimensional feature information.
[0163] The method for determining the three-dimensional feature points of the i-th group of three-dimensional feature information from the first camera coordinate system to the perception coordinate system based on the transformation relationship and s can be based on the following formula (5): X′=sMY (5)
[0164] Where X′ represents the three-dimensional feature point of the i-th group of three-dimensional feature information transformed from the first camera coordinate system to the perception coordinate system.
[0165] The chamfer distance can be understood as the loss function of feature points before and after fusion. The chamfer distance can be referred to as formula (6) shown below.
[0166] Thus, the chamfer distance between three-dimensional feature points can be calculated according to formula (6), and a set of three-dimensional feature information selected from the three-dimensional feature information set can also be used as the three-dimensional feature point set based on the chamfer distance. It should be noted that the three-dimensional feature point set can belong to the optimal or locally optimal three-dimensional feature information in the three-dimensional feature information set. That is, the set of three-dimensional feature information to which the three-dimensional feature point set belongs has the optimal three-dimensional feature points, and the number of optimal three-dimensional feature points is the largest or the range is the widest. The three-dimensional feature points in other sets of three-dimensional feature information outside the three-dimensional feature point set may also be the optimal three-dimensional feature points, but their number is less than the number of optimal three-dimensional feature points in the three-dimensional feature point set, or their coverage is less than the coverage of the optimal three-dimensional feature points in the three-dimensional feature point set. Among them, the optimal three-dimensional feature point can be the three-dimensional feature point corresponding to the minimum loss function determined by the aforementioned vertical distance or chamfer distance.
[0167] 6. Registration results, used to indicate the transformation relationship between different coordinate systems, may include the transformation matrix between the pixel coordinate system and the camera coordinate system, the transformation matrix between the camera coordinate system and the perception coordinate system, etc. The transformation relationship between different coordinate systems can be the transformation relationship obtained after data fusion and registration (or alignment or optimization). Therefore, the transformation matrix between different coordinate systems can be understood as the transformation matrix after registration (or alignment or optimization).
[0168] This application does not limit the method for obtaining the registration result; it can be obtained based on a 3D feature information set and perceptual data. Specifically, it can be obtained by optimizing the transformation matrix between various coordinate systems based on a 3D feature point set generated from the 3D feature information set and perceptual data. In the embodiments of this application, the first pixel coordinate system and the first camera coordinate system of the first image data, and the second pixel coordinate system of the second image data are used as examples. Thus, based on the 3D feature information set and the 3D feature point set generated from the perception data, the first transformation matrix between the first pixel coordinate system and the first camera coordinate system, the second transformation matrix between the second pixel coordinate system and the second camera coordinate system, and the third transformation matrix between the first camera coordinate system and the perception coordinate system can be registered (or aligned or optimized) to obtain the registered (or aligned or optimized) first transformation matrix, second transformation matrix, and third transformation matrix. This allows us to obtain the registered (or aligned or optimized) transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the registered (or aligned or optimized) transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, and the registered (or aligned or optimized) transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data.
[0169] Optionally, the registration result can be the optimal first transformation matrix, second transformation matrix, and third transformation matrix. Alternatively, the registration result can indicate the first and second transformation matrices through indices in the optimal combinations corresponding to at least one first preselected value group and at least one second preselected value group. It is understood that indicating the first and second transformation matrices through indices can save signaling.
[0170] Optionally, under resource constraints, a portion of the 3D feature information can be sent first. If the 3D feature points in this portion meet the preset performance requirements, the remaining 3D feature information does not need to be sent. Otherwise, the remaining 3D feature information, or a portion of the remaining 3D feature information, can continue to be sent. In this way, sending 3D feature information of matching pixel pairs in an additive manner helps improve performance.
[0171] 7. Configuration information, used to indicate data registration parameters to the object of data transmission. Optionally, configuration information is sent to the device matching two image data to indicate at least one of the following: the number of matching pixel pairs, at least one first pre-selected value group, and at least one second pre-selected value group. This helps to improve the accuracy of data processing.
[0172] Furthermore, when the number of image data points is greater than two, the number of image data points can be indicated through configuration information. The image data includes at least the first image data and the second image data. This improves the accuracy of data processing.
[0173] Optionally, the configuration information may also include information about the three-dimensional feature information set, such as the number of groups of three-dimensional feature information and / or the number of points in each group of three-dimensional feature information.
[0174] 8. Geometric information refers to the feature information of the three-dimensional surface corresponding to the image contour information in the perceptual space. In the embodiments of this application, geometric information is used to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perceptual space and the boundary points of each of the M three-dimensional surfaces. The perceptual space can be understood as a three-dimensional coordinate system. As shown in Figure 2B, the two-dimensional surface can be projected onto the three-dimensional surface of the camera's three-dimensional space, and the boundary points in the two-dimensional surface each have corresponding (mapped) boundary points in the three-dimensional surface. For example, boundary point A corresponds to boundary point A', boundary point B corresponds to boundary point B', boundary point C corresponds to boundary point C', and boundary point D corresponds to boundary point D'. The positions of the boundary points in the three-dimensional surface can be represented by an array or by a bitmap, as described in the description of the representation of the boundary points of the two-dimensional surface, and will not be repeated here.
[0175] This application does not limit the method for obtaining geometric information, and may include the following four methods, among which:
[0176] Method 1: Generate geometric information based on image contour information and perceptual data.
[0177] Method 2: Generate geometric information based on image contour information and 3D feature point set.
[0178] Method 3: Generate geometric information based on image contour information and registration results.
[0179] Method 4: Generate geometric information based on image contour information and geometric indication information.
[0180] As mentioned earlier, the registration results and 3D feature point sets can be generated based on perceptual data and 3D feature information sets. The following example demonstrates the generation of geometric information from perceptual data and image contour information. 3D feature points in the perceptual data can be converted into feature points in a 2D surface. Combined with feature points from the image contour information, points on the 3D surface are mapped onto the 2D surface, resulting in a 2D feature point set for each 2D surface. Based on the corresponding 3D feature point set for each 2D feature point set, a 3D surface is fitted. The boundary points of each 2D surface are projected into the camera's 3D space, and the intersection of the projection into the camera's 3D space and the 3D surface is taken as the boundary points of the 3D surface. Thus, 3D surfaces and their boundary points are obtained, yielding geometric information. Furthermore, perceptual points can be sampled from the obtained bounded 3D surfaces to obtain a dense point cloud, completing the densification of perceptual points. The registration results can be used to generate 3D feature points. The 3D feature points or 3D feature point sets generated by the registration results can be converted into feature points in a 2D surface, thus performing the above steps to obtain geometric information.
[0181] In this embodiment, geometric indication information is used to indicate geometric information. Optionally, the geometric indication information may include at least one of the following: the number of feature points in each of the M three-dimensional surfaces, the index value of each feature point in the perceptual data, and each feature point...
[0182] Feature points can include boundary points, specifically points that can be mapped to a two-dimensional surface. The number of feature points in each of the M three-dimensional surfaces can be represented by a list, such as {n1, ..., n}. M} Where n1 represents the number of feature points in the first 3D surface, and so on, n M Let n be the number of feature points in the Mth 3D surface. This application does not limit the number of feature points in each 3D surface; for example, n1, ..., n M Any two of them can be equal or unequal. The number of feature points in each three-dimensional surface can be equal to the number of feature points in the corresponding two-dimensional surface.
[0183] The index value of each feature point in the perceptual data can be represented by a list, such as {{i 1_1 i 1_2 , ...}, ..., {i M_1 B M_2 , ...}}. Where, {i 1_1 i 1_2 , ...} represents the set of index values of feature points in the first 3D surface in the perceptual data, and so on, {i M_1 B M_2 , ...} represents the set of index values of feature points in the M-th 3D surface in the perceptual data.
[0184] Optionally, geometric indication information can be generated based on image contour information and perception data.
[0185] For example, three-dimensional feature points in the perceptual data are converted into feature points in two-dimensional surfaces, and combined with feature points in the image contour information, so that points on the three-dimensional surfaces fall onto the two-dimensional surfaces, thus obtaining a set of two-dimensional feature points for each two-dimensional surface. The index value of each feature point in each two-dimensional feature point set in the perceptual data is then determined, thereby obtaining the index value of the feature points of each of the M three-dimensional surfaces in the perceptual data. In this way, based on the perceptual data, the feature points of each two-dimensional surface in the image contour information can be registered to obtain geometric indication information.
[0186] It is understandable that, based on image contour information and geometric indicator information, the corresponding three-dimensional surface for each two-dimensional surface in the image contour information can be determined, and the feature points (including boundary points) of each two-dimensional surface can be determined in the three-dimensional surface. Thus, geometric information can be obtained from the boundary points of the three-dimensional surface. For example, based on the feature points of each three-dimensional surface according to the index value, a three-dimensional surface can be fitted, and the boundary points corresponding to the boundary points of the two-dimensional surface in the image contour information can be determined in the three-dimensional surface.
[0187] This application proposes a data transmission method that uses image contour information to assist in the generation of geometric information from perception data. This method can improve the representational ability and reconstruction accuracy of perception, and also provides a significant improvement in accuracy in sparse perception scenarios where contours are not obvious.
[0188] Please refer to Figure 6, which is an interactive schematic diagram of a data transmission method provided in an embodiment of this application. The first and second devices in Figure 6 can be described with reference to Figures 1A and 1B. The method includes, but is not limited to, the following steps S601 and S602, wherein:
[0189] S601, the first device receives image contour information from the second device, the image contour information being used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces.
[0190] Correspondingly, the second device sends image contour information to the first device.
[0191] In some feasible examples, the method may further include: a second device acquiring image contour information based on image data. The image data may be the first image data or the second image data described previously, or it may be image data other than the first and second image data; no limitation is made here. The method for acquiring image contour information can be referred to the description in FIG2B, and will not be repeated here.
[0192] S602. The first device generates geometric information based on image contour information, as well as sensing data and / or first information. The geometric information is used to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the sensing space and the boundary points of each of the M three-dimensional surfaces. The first information is used to indicate the feature information generated by the sensing data.
[0193] The perceptual data and geometric information can be referred to the aforementioned definitions and will not be repeated here. As can be seen from the method described above, after the second device sends the image contour information to the first device, the first device can generate geometric information based on the image contour information and the perceptual data to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perceptual space and the boundary points of each of the M three-dimensional surfaces.
[0194] In a first feasible example, the first information includes a registration result and / or a set of three-dimensional feature points generated based on perceptual data and a set of three-dimensional feature information. The set of three-dimensional feature information is generated based on two-dimensional feature pair information, which indicates matching pixel pairs between first image data and second image data. The set of three-dimensional feature information includes at least one set of three-dimensional feature information, each set including three-dimensional feature points generated from the matching pixel pairs. One of the at least one sets of three-dimensional feature information includes the set of three-dimensional feature points.
[0195] The first feasible example can be referred to the descriptions of Method 1, Method 2, and Method 3 mentioned above, and will not be repeated here. It can be understood that after the second device sends the image contour information to the first device, the first device can generate geometric information based on the image contour information and the registration result and / or the three-dimensional feature point set generated by the perception data, so as to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perception space and the boundary points of each of the M three-dimensional surfaces.
[0196] In some feasible examples, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data. Thus, the two-dimensional features of the image data, such as matching pixel pairs between the image data and other image data, can be located based on the image size, which helps improve the accuracy and convenience of data fusion.
[0197] In some feasible examples, the three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information. This helps to improve the accuracy of processing the three-dimensional feature information set.
[0198] In some feasible examples, the three-dimensional feature information set is generated based on the two-dimensional feature pair information, the at least one first pre-selected value group, and the at least one second pre-selected value group. The at least one first pre-selected value group indicates pre-selected values for unknown parameters of the first image data, and the at least one second pre-selected value group indicates pre-selected values for unknown parameters of the second image data. Thus, by generating the three-dimensional feature information set based not only on the two-dimensional feature pair information but also on the pre-selected values of unknown parameters of the first and second image data, the accuracy of obtaining the three-dimensional feature information set can be improved.
[0199] In some feasible examples, the registration result is used to indicate the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, and the transformation relationship between the perceptual coordinate system and the first camera coordinate system of the perceptual data. Thus, the camera coordinates corresponding to feature points in the image data or perceptual data can be determined based on the registration result. For example, the camera coordinates corresponding to pixels in the image data can be determined based on the transformation relationship between the first pixel coordinate system and the first camera coordinate system, and the image data acquired by the acquisition device corresponding to the first image data. Similarly, the camera coordinates corresponding to pixels in the image data can be determined based on the transformation relationship between the second pixel coordinate system and the second camera coordinate system, and the image data acquired by the acquisition device corresponding to the second image data. The camera coordinates corresponding to feature points in the perceptual data can also be determined based on the transformation relationship between the perceptual coordinate system and the first camera coordinate system, and the perceptual data acquired by the acquisition device corresponding to the perceptual data. Then, the subsequently acquired image data or perceptual data can be fused based on the camera coordinates, which can improve the efficiency of data fusion and improve the accuracy of data processing.
[0200] In this embodiment, the transformation relationship between the first pixel coordinate system and the first camera coordinate system refers to the transformation relationship of the feature point from the first pixel coordinate system to the first camera coordinate system, which can be represented by a first transformation matrix. The transformation relationship between the second pixel coordinate system and the second camera coordinate system refers to the transformation relationship of the feature point from the second pixel coordinate system to the second camera coordinate system, which can be represented by a second transformation matrix. The transformation relationship between the perception coordinate system and the first camera coordinate system refers to the transformation relationship of the feature point from the first camera coordinate system to the perception coordinate system, which can be represented by a third transformation matrix.
[0201] Optionally, the registration result can also be used to indicate at least one of the following: the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system. This can improve the efficiency of other data fusion processes.
[0202] In this embodiment, the transformation relationship between the camera coordinate system and the second camera coordinate system can be represented by a camera transformation matrix, which can be based on the first camera coordinate system. The transformation relationship between the first pixel coordinate system and the perception coordinate system can be obtained by multiplying the first and third transformation matrices, such as by multiplying the first and third transformation matrices. The transformation relationship between the second pixel coordinate system and the perception coordinate system can be obtained by multiplying the second and third transformation matrices, such as by multiplying the second and third transformation matrices. The transformation relationship between the perception coordinate system and the second camera coordinate system refers to the transformation relationship of feature points from the second camera coordinate system to the perception coordinate system.
[0203] It should be noted that the registration results can also be used to indicate the transformation relationship between other coordinate systems, such as the transformation relationship between the first pixel coordinate system and the world coordinate system, the transformation relationship between the second pixel coordinate system and the world coordinate system, etc., without limitation here. In this way, data fusion can be performed based on the transformation relationship between at least two coordinate systems.
[0204] In some feasible examples, the registration result is obtained based on periodic or triggered information. That is, the acquisition device can be triggered to acquire image data and sensing data when the periodic time arrives, so that the first device and the second device can obtain the registration result according to the method provided in this application. Alternatively, it can be triggered by an event, such as when the sensing coordinate system of the acquisition device changes or the position of the acquisition device changes, so that the first device and the second device can obtain the registration result according to the method provided in this application.
[0205] It's understandable that if the coordinate system corresponding to the registration result remains unchanged, data fusion can be performed directly based on that registration result. Otherwise, it's necessary to obtain the registration result again. Obtaining the registration result periodically or triggered by specific events helps improve the accuracy of data fusion.
[0206] This application does not limit the method for acquiring (or generating) two-dimensional feature pair information, three-dimensional feature information set, registration result, and three-dimensional feature point set. The apparatus for acquiring two-dimensional feature pair information, three-dimensional feature information set, registration result, and three-dimensional feature point set can be referred to as a processing apparatus. The processing apparatus can be a terminal device or a network device. The following descriptions address different application scenarios and different diagrams, wherein:
[0207] In the first application scenario, the data acquisition device, the first device, and the second device are the same processing device.
[0208] In this application scenario, the processing device acquires at least two image data; obtains W blocks of two-dimensional feature pair information based on the at least two image data, each block of two-dimensional feature pair information is used to indicate the matching pixel pair between the two image data; and generates a three-dimensional feature information set based on the W blocks of two-dimensional feature pair information.
[0209] Optionally, the processing device can also acquire sensing data and obtain registration results and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the sensing data.
[0210] Furthermore, the processing device can also fuse the subsequently acquired sensory data based on the registration results and / or the three-dimensional feature point set.
[0211] It is understandable that in the first application scenario, a processing device is used to acquire image data and sensory data, and to process different sensory data to obtain two-dimensional feature pair information, three-dimensional feature information set, registration result and / or three-dimensional feature point set of these sensory data.
[0212] In the second application scenario, the data acquisition device is a device other than the first and second devices, and the processing device is either the first or the second device. As shown in Figure 1B, the data acquisition device 30 is a device other than the first device 10 and the second device 20, and the processing device can be either the first device 10 or the second device 20.
[0213] In the third application scenario, the data acquisition device is at least one of the first device and the second device, and the processing device is a device other than the data acquisition device. As shown in Figure 1A, when the data acquisition device is the first device 10, the processing device is the second device 20; or when the data acquisition device is the second device 20, the processing device is the first device 10.
[0214] In the second and third application scenarios, the acquisition device can be one or more terminal devices or acquisition modules within terminal devices. As shown in Figure 7A, the processing device (first device or second device) receives at least two image data from at least one acquisition device; the processing device acquires W blocks of two-dimensional feature pair information based on the at least two image data, each block of two-dimensional feature pair information being used to indicate matching pixel pairs between the two image data; the processing device generates a three-dimensional feature information set based on the W blocks of two-dimensional feature pair information.
[0215] Optionally, the processing device may also receive sensing data from the acquisition device; the processing device may acquire registration results and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the sensing data; and the processing device may send the registration results and / or the set of three-dimensional feature points to the acquisition device.
[0216] It is understandable that in the second and third application scenarios, the acquisition device sends image data and perception data to the processing device, and the processing device separately acquires two-dimensional feature pair information, three-dimensional feature information set, registration result and / or three-dimensional feature point set.
[0217] The fourth application scenario, regardless of whether the data acquisition device is the first device or the second device, and the processing device is the first device or the second device, can include the following six methods, among which:
[0218] Method 1: The second device acquires two-dimensional feature pair information, while the first device acquires a three-dimensional feature information set, registration results, and / or a three-dimensional feature point set;
[0219] Method 2: The second device acquires two-dimensional feature pair information, the first device acquires a three-dimensional feature information set, and the second device acquires the registration result and / or a three-dimensional feature point set;
[0220] Method 3: The first device acquires two-dimensional feature pair information, and the second device acquires a three-dimensional feature information set, registration results, and / or a three-dimensional feature point set;
[0221] Method 4: The first device acquires two-dimensional feature pair information, the second device acquires a three-dimensional feature information set, and the first device acquires the registration result and / or a three-dimensional feature point set;
[0222] Method 5: The first device acquires two-dimensional feature pair information and a three-dimensional feature information set, and the second device acquires the registration result and / or a three-dimensional feature point set;
[0223] Method 6: The second device acquires two-dimensional feature pair information and a three-dimensional feature information set, while the first device acquires the registration result and / or a three-dimensional feature point set.
[0224] It is understood that in the fourth application scenario, the first device and the second device respectively acquire at least one of the following: two-dimensional feature pair information, three-dimensional feature information set, registration result, and three-dimensional feature point set. In the above six methods, image data and / or perceived data can be acquired by the first device, the second device, or other devices. Specifically, method one can refer to Figure 7B, method two to Figure 7C, and method six to Figure 7D. In method three, the first device can be the second device in Figure 7B, and the second device in method three can be the first device in Figure 7B. In method four, the first device can be the second device in Figure 7C, and the second device in method four can be the first device in Figure 7C. In method five, the first device can be the second device in Figure 7D, and the second device in method five can be the first device in Figure 7D.
[0225] As shown in Figures 7B to 7D, image data and sensory data can be acquired by a data acquisition device, or by a first device or a second device. When the number of image data points is equal to 2, W can be 1. When the number of image data points is greater than 2, M can be greater than 1. Different devices for acquiring the first image data, the second image data, and the sensory data can include the following three cases, wherein:
[0226] Scenario 1: UE#1 collects the first image data, UE#2 collects the second image data, and UE#3 collects the perceived data.
[0227] Scenario 2: UE#1 collects the first image data and the second image data, while UE#3 collects the perceived data.
[0228] Scenario 3: UE#1 collects the first image data and perception data, and UE#2 collects the second image data.
[0229] The above three examples are applicable to the six methods described above, and at least one of the first and second devices in the six methods can be one of UE#1, UE#2, and UE#3. That is, the first device can be one of UE#1, UE#2, and UE#3, and the second device can be any device among UE#1, UE#2, and UE#3 other than the first device, or any device other than UE#1, UE#2, and UE#3. Alternatively, the second device can be one of UE#1, UE#2, and UE#3, and the first device can be any device among UE#1, UE#2, and UE#3 other than the second device, or any device other than UE#1, UE#2, and UE#3.
[0230] In the third or fourth application scenario, and in some feasible examples, the method may further include: synchronizing configuration information between the first and second devices. The configuration information includes at least one of the following: the number of matched pixel pairs, at least one first pre-selected value group, and at least one second pre-selected value group. This facilitates the determination of the three-dimensional feature information set, such as the three-dimensional feature information corresponding to different first and second pre-selected value groups within the three-dimensional feature information set, thereby improving the accuracy of feature information acquisition.
[0231] In some feasible examples, the configuration information also includes the number of image data, which includes at least the first image data and the second image data. If image data other than the first and second image data is not included, the configuration information may not include the number of image data, thus allowing the acquisition of two-dimensional feature pair information between the two image data after receiving the first and second image data. If the configuration information includes the number of image data, the number of pairwise combinations of the image data can be determined based on this number, thereby determining the number of two-dimensional feature pairs and improving the accuracy of feature information acquisition.
[0232] If other image data exists, it can include W blocks of two-dimensional feature pair information, used to indicate matching pixel pairs between two images in the other image data, the first image data, and the second image data. Optionally, the two-dimensional feature pair information can also be used to indicate the image size of the other image data. In this way, the two-dimensional features of the image can be located based on the image size of the image data, which helps to improve the accuracy and convenience of data fusion.
[0233] In a second feasible example, the first information includes geometric indication information generated based on the perceived data and the image contour information. The geometric indication information indicates the number of feature points in each of the three-dimensional surfaces and / or the index value of each feature point in the perceived data, and the number of each feature point. Referring to the description of Method Four above, it will not be repeated here.
[0234] In this example, the geometric indication information may be information generated by the second device based on image contour information and perceptual data. Alternatively, the geometric indication information may be information generated by the second device based on image contour information and registration results, or information generated based on image contour information and a set of three-dimensional feature points.
[0235] Optionally, the method may further include: the first device receiving sensing data and / or first information from the second device. Correspondingly, the second device sends the sensing data and / or first information to the first device. That is, the second device not only sends image contour information to the first device, but may also send sensing data and / or first information, such as geometric indication information. Thus, when the first information is geometric indication information, the first device can generate geometric information based on the image contour information and the geometric indication information, thereby improving the efficiency of acquiring geometric information.
[0236] It should be noted that the above examples of the two types of first information are just that—examples. In reality, the first information can also be other types of information generated based on perceptual data, and this is not limited here. Furthermore, the first information mentioned above can be implemented individually or in combination.
[0237] Referring to the perceptual data shown in Figure 8, in a simulation implementation, the loss function for calculating the chamfer distance of boundary points based on sparse perceptual data X and the original perceptual data O is 0.00153. The perceptual data F reconstructed from image data and sparse (e.g., sampling rate of 0.1) perceptual data, compared with the original perceptual data O, has a loss function of 0.00038 for calculating the chamfer distance of boundary points. Thus, perceptual data reconstructed from image data can improve the accuracy of feature information extraction from perceptual data.
[0238] In another simulation, perceptual data obtained by fusing registered 3D feature points with perceptual data O at a sampling rate of 0.1, and comparing it with the original perceptual data, the loss function for calculating the chamfer distance of boundary points is 0.00102. Perceptual data F' reconstructed from image data, perceptual data X at a sampling rate of 0.1, and geometric information generated from registered 3D feature points, and compared with the original perceptual data O to generate geometric information, has a loss function for calculating the chamfer distance of boundary points of 0.0006. Thus, perceptual data reconstructed from image data can improve the accuracy of feature information extraction from perceptual data.
[0239] In some feasible examples, the method may further include: the first device sending second information to the second device. The second information includes geometric information. Similarly, the first device may also send the second information to a device other than the second device (such as a data acquisition device or other devices). Thus, the device receiving the second information can perform sensing functions such as detection, localization, recognition, and imaging of the target object based on the geometric information.
[0240] In some feasible examples, the second information also includes at least one of the following: the number M of the three-dimensional or two-dimensional surfaces, and the number of feature points in each of the three-dimensional or two-dimensional surfaces. Thus, image reconstruction can be achieved based on the geometric information and the second information.
[0241] In this embodiment of the application, the second information may not be sent to the acquisition device. Instead, the second information may be sent to the acquisition device without carrying the number M of the three-dimensional or two-dimensional surfaces or the number of feature points in each of the three-dimensional or two-dimensional surfaces, thereby avoiding the acquisition device from repeatedly acquiring this type of information.
[0242] It is understandable that, in the method shown in Figure 6, using image contour information to assist in generating geometric information from perceptual data can improve the representational ability and reconstruction accuracy of perception, and also significantly improve accuracy in sparse perception scenarios where contours are not obvious. Furthermore, the feature information transmitted can protect user privacy and reduce data transmission volume compared to the transmitted data itself.
[0243] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.
[0244] Please refer to Figure 9, which is a schematic diagram of a communication device provided in an embodiment of this application. The communication device may include a transceiver unit 901 and a processing unit 902. The transceiver unit 901 may be a unit with signal input (receiving) or output (transmitting) functions, used for transmitting signals to other devices or other units within a device. The processing unit 902 may be a unit with processing functions, such as one or more processors, used for executing instructions (or code or programs). The communication device may be a first device or a second device. The first device and the second device may be a terminal device or a network device.
[0245] When the communication device is the first device, the communication device includes:
[0246] The transceiver unit 901 is used to receive image contour information; wherein, the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces;
[0247] The processing unit 902 is used to generate geometric information based on the image contour information, as well as the perception data and / or the first information; wherein the geometric information is used to indicate the boundary points of the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perception space and the boundary points of each of the M three-dimensional surfaces, and the first information is used to indicate the feature information generated based on the perception data.
[0248] In some feasible examples, the first information includes a registration result and / or a set of three-dimensional feature points generated based on the perceived data and the three-dimensional feature information set; wherein the three-dimensional feature information set is generated based on two-dimensional feature pair information, the two-dimensional feature pair information being used to indicate matching pixel pairs between the first image data and the second image data, the three-dimensional feature information set including at least one set of three-dimensional feature information, each set of the three-dimensional feature information including three-dimensional feature points generated from the matching pixel pairs; one set of the three-dimensional feature information in the at least one set of three-dimensional feature information includes the set of three-dimensional feature points.
[0249] In some feasible examples, the registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, the transformation relationship between the perceptual coordinate system and the first camera coordinate system of the perceptual data, the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perceptual coordinate system, the transformation relationship between the second pixel coordinate system and the perceptual coordinate system, and the transformation relationship between the perceptual coordinate system and the second camera coordinate system.
[0250] In some feasible examples, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.
[0251] In some feasible examples, the three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.
[0252] In some feasible examples, the three-dimensional feature information set is generated based on the two-dimensional feature pair information, at least one first preselected value group, and at least one second preselected value group; wherein the at least one first preselected value group is used to indicate preselected values of unknown parameters of the first image data, and the at least one second preselected value group is used to indicate preselected values of unknown parameters of the second image data.
[0253] In some feasible examples, the first information includes geometric indication information generated based on the perceptual data and the image contour information; wherein the geometric indication information is used to indicate at least one of the following: the number of feature points in each of the three-dimensional surfaces and / or the index value of each of the feature points in the perceptual data, and each of the feature points.
[0254] Optionally, the transceiver unit 901 is further configured to receive the sensed data and / or the first information.
[0255] In some feasible examples, the transceiver unit 901 is also configured to receive second information; wherein the second information includes the geometric information.
[0256] In some feasible examples, the second information further includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of feature points in each of the three-dimensional surface or the two-dimensional surface.
[0257] When the communication device is a second device, the communication device includes:
[0258] The processing unit 902 is used to acquire image contour information based on image data; wherein, the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces;
[0259] The transceiver unit 901 is used to transmit the image contour information.
[0260] Optionally, the transceiver unit 901 is further configured to transmit sensing data and / or first information, wherein the first information is used to indicate feature information generated by the sensing data.
[0261] In some feasible examples, the transceiver unit 901 is further configured to receive second information; wherein the second information includes geometric information, the geometric information being used to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perception space and the boundary points of each of the M three-dimensional surfaces.
[0262] In some feasible examples, the second information further includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of feature points in each of the three-dimensional surface or the two-dimensional surface.
[0263] The implementation of the above-mentioned transceiver unit 901 and processing unit 902 can be referred to the relevant description of the method embodiment shown in FIG6, which will not be repeated here.
[0264] Please refer to Figure 10, which is a schematic diagram of another communication device provided in an embodiment of this application. As shown in Figure 10, the communication device may include a processor 111 and a storage medium 112. The processor 111 may also be called a processing unit, which can implement certain control functions. The storage medium 112 may also be called a storage unit or a memory. Instructions 114 are stored on the storage medium 112. The instructions 114 can be executed on the processor 111, causing the communication device to perform any of the methods described in Figure 6 of this application embodiment.
[0265] Optionally, the processor 111 may include instructions 113 that can be executed on the processor 111 to cause the communication device to perform any of the methods described in FIG6 in the embodiments of this application.
[0266] The communication device can be a first device or a second device. The first device and the second device can be terminal devices or network devices, used to implement the method described in the method embodiments. However, the scope of the device described in this application is not limited thereto; the communication device can be a standalone device or part of a larger device. For example, the communication device can be:
[0267] (1) An independent integrated circuit IC, or chip, or chip system or subsystem;
[0268] (2) A collection of one or more ICs, wherein the collection of ICs may optionally include a storage component for storing data and / or instructions;
[0269] (3) ASIC, such as modems;
[0270] (4) Modules that can be embedded in other devices.
[0271] Please refer to Figure 11, which is a schematic diagram of a terminal device provided in an embodiment of this application. For ease of explanation, Figure 11 only shows the main components of the terminal device. As shown in Figure 11, the terminal device includes a processor, a memory, a control circuit, an antenna, and input / output devices. The processor is mainly used to process communication protocols and communication data, control the entire terminal device, execute software programs, and process the data of the software programs. The memory is mainly used to store software programs and data. The radio frequency circuit is mainly used for the conversion between baseband signals and radio frequency signals and the processing of radio frequency signals. The antenna is mainly used for transmitting and receiving radio frequency signals in the form of electromagnetic waves. Input / output devices, such as touch screens, displays, and keyboards, are mainly used to receive user input data and output data to the user.
[0272] When the terminal device is powered on, the processor can read the software program from the storage unit, parse and execute the instructions of the software program, and process the data of the software program. When data needs to be transmitted wirelessly, the processor performs baseband processing on the data to be transmitted and outputs the baseband signal to the radio frequency (RF) circuit. The RF circuit processes the baseband signal to obtain the RF signal and transmits the RF signal outward in the form of electromagnetic waves through the antenna. When data is sent to the terminal device, the RF circuit receives the RF signal through the antenna. This RF signal is further converted into a baseband signal and output to the processor. The processor converts the baseband signal back into data and processes the data.
[0273] For ease of explanation, Figure 11 shows only one memory and processor. In actual terminal devices, multiple processors and memories may exist. Memory may also be referred to as storage medium or storage device, etc., and the embodiments of this application do not limit this.
[0274] In one embodiment, the antenna is used to perform the operations performed by the transceiver unit 901 in the above embodiment. The processor is used to perform the operations performed by the processing unit 902 in the above embodiment.
[0275] This application also provides a computer-readable storage medium storing instructions or computer programs that, when executed, can implement the relevant steps in the data transmission method provided in the above-described method embodiments.
[0276] This application also provides a computer program product, which includes instructions or a computer program that, when executed, causes one or more steps in any of the above-described data transmission methods to be performed. If the constituent modules of the aforementioned devices are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0277] The instructions or computer program can be executed by a computer or processor, without limitation.
[0278] This application provides a chip or chip system including at least one processor for calling and running instructions stored in a memory, causing a communication device with the chip installed to perform any of the above methods or to execute the steps of the processing unit 902.
[0279] This application embodiment also provides another chip, including a processor and a memory, wherein the processor is used to call and run instructions stored in the memory, causing a communication device with the chip installed to perform any of the above methods, or to perform the steps of the processing unit 902.
[0280] This application embodiment also provides another chip, including: an input interface, an output interface, and a processing circuit. The input interface, the output interface, and the processing circuit are connected via internal connection paths. The processing circuit is used to execute any of the methods described above. Optionally, the chip also includes a memory. The input interface, the output interface, the processor, and the memory are connected via internal connection paths. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute any of the methods described above, or to execute the steps of processing unit 902.
[0281] This application also provides another chip system, including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a line. The at least one processor is used to run a computer program or instructions to perform any of the methods described above, or to execute the steps of processing unit 902. This chip system may be composed of chips, or may include chips and other discrete devices.
[0282] This application also provides a communication system, which includes a first device and a second device, as detailed in Figure 6. The first device and the second device in this application can be a terminal device or a network device.
[0283] It should be understood that the memory mentioned in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be a hard disk drive (HDD), a solid-state drive (SSD), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be RAM, which is used as an external cache. Memory is any other medium capable of carrying or storing desired program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in the embodiments of this application can also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.
[0284] It should also be understood that the processor mentioned in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor, or any conventional processor, etc.
[0285] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) is integrated into the processor.
[0286] It should be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0287] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments provided herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0288] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0289] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0290] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0291] The steps in the methods of this application can be adjusted, combined, or deleted according to actual needs. Each step in each embodiment can be partially performed (for example, the terminal device may not perform the steps performed by the terminal device in the above embodiments). The execution order of different steps can be changed. The embodiments described herein can be combined with other embodiments, different embodiments can be combined with each other, and different steps of different embodiments herein can be combined.
[0292] The modules / units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0293] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments.
[0294] In this application, it may refer to a communication protocol or specification, such as the 3GPP communication protocol.
[0295] In this application, unless otherwise specified, "at least one" means "one or more".
[0296] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0297] In the embodiments of this application, "including" can refer to a relationship of inclusion or an equality relationship. For example, A includes B, which could mean that A includes B and may also include other content, or that A and B are the same content.
[0298] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0299] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
Claims
1. A data transmission method, characterized in that, include: Receive image contour information; wherein the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces; Based on the image contour information, as well as the perceptual data and / or the first information, geometric information is generated; wherein, the geometric information is used to indicate the boundary points of the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perceptual space and the boundary points of each of the M three-dimensional surfaces, and the first information is used to indicate the feature information generated by the perceptual data.
2. The method according to claim 1, characterized in that, The first information includes geometric indication information generated based on the perceived data and the image contour information; The geometric indication information is used to indicate at least one of the following: the number of feature points in each of the three-dimensional surfaces, the index value of each feature point in the perceptual data, and each feature point.
3. The method according to claim 1, characterized in that, The first information includes registration results and / or a set of three-dimensional feature points generated based on the perceived data and the three-dimensional feature information set; The three-dimensional feature information set is generated based on two-dimensional feature pair information. The two-dimensional feature pair information is used to indicate matching pixel pairs between the first image data and the second image data. The three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated from the matching pixel pairs. The set of three-dimensional feature information in the three-dimensional feature information set includes the set of three-dimensional feature points.
4. The method according to claim 3, characterized in that, The registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, the transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data, the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.
5. The method according to claim 3, characterized in that, The two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.
6. The method according to claim 3, characterized in that, The three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.
7. The method according to any one of claims 3 to 6, characterized in that, The three-dimensional feature information set is generated based on the two-dimensional feature pair information, at least one first pre-selected value group, and at least one second pre-selected value group; Wherein, the at least one first preselected value group is used to indicate the preselected value of the unknown parameter of the first image data, and the at least one second preselected value group is used to indicate the preselected value of the unknown parameter of the second image data.
8. The method according to any one of claims 1 to 7, characterized in that, Also includes: Send a second message; wherein the second message includes the geometric information.
9. The method according to claim 8, characterized in that, The second information also includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of boundary points in each of the three-dimensional surface or the two-dimensional surface.
10. A data transmission method, characterized in that, include: Image contour information is obtained based on image data; wherein, the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces; Send the image contour information.
11. The method according to claim 10, characterized in that, Also includes: Receive second information; wherein the second information includes geometric information, the geometric information being used to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perception space and the boundary points of each of the M three-dimensional surfaces.
12. The method according to claim 11, characterized in that, The second information also includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of boundary points in each of the three-dimensional surface or the two-dimensional surface.
13. A data transmission device, characterized in that, include: A transceiver unit is used to receive image contour information; wherein the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces; The processing unit is configured to generate geometric information based on the image contour information, as well as the perceptual data and / or first information; wherein the geometric information is used to indicate the boundary points of the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perceptual space and the boundary points of each of the M three-dimensional surfaces, and the first information is used to indicate the feature information generated by the perceptual data.
14. The apparatus according to claim 13, characterized in that, The first information includes geometric indication information generated based on the perceived data and the image contour information; The geometric indication information is used to indicate at least one of the following: the number of feature points in each of the three-dimensional surfaces and / or the index value of each feature point in the perceptual data, and each feature point.
15. The apparatus according to claim 13, characterized in that, The first information includes registration results and / or a set of three-dimensional feature points generated based on the perceived data and the three-dimensional feature information set; The three-dimensional feature information set is generated based on two-dimensional feature pair information. The two-dimensional feature pair information is used to indicate matching pixel pairs between the first image data and the second image data. The three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated from the matching pixel pairs. One set of three-dimensional feature information in the at least one set of three-dimensional feature information includes the set of three-dimensional feature points.
16. The apparatus according to claim 15, characterized in that, The registration result is used to indicate the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, the transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data, the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.
17. The apparatus according to claim 15, characterized in that, The two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.
18. The apparatus according to claim 15, characterized in that, The three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.
19. The apparatus according to any one of claims 15 to 18, characterized in that, The three-dimensional feature information set is generated based on the two-dimensional feature pair information, at least one first pre-selected value group, and at least one second pre-selected value group; Wherein, the at least one first preselected value group is used to indicate the preselected value of the unknown parameter in the first image data, and the at least one second preselected value group is used to indicate the preselected value of the unknown parameter in the second image data.
20. The apparatus according to any one of claims 13 to 19, characterized in that, The transceiver unit is also used to send second information; wherein the second information includes the geometric information.
21. The apparatus according to claim 20, characterized in that, The second information also includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of boundary points in each of the three-dimensional surface or the two-dimensional surface.
22. A data transmission device, characterized in that, include: A processing unit is used to acquire image contour information based on image data; wherein the image contour information is used to indicate M two-dimensional surfaces and the boundary points of each of the M two-dimensional surfaces; A transceiver unit is used to send the image contour information.
23. The apparatus according to claim 22, characterized in that, The transceiver unit is further configured to receive second information; wherein the second information includes geometric information, the geometric information being used to indicate the M three-dimensional surfaces corresponding to each of the M two-dimensional surfaces in the perception space and the boundary points of each of the M three-dimensional surfaces.
24. The apparatus according to claim 23, characterized in that, The second information also includes at least one of the following: the number M of the three-dimensional surface or the two-dimensional surface, and the number of boundary points in each of the three-dimensional surface or the two-dimensional surface.
25. A communication device, characterized in that, The communication device includes a processor and a storage medium storing instructions that, when executed by the processor, cause the method according to any one of claims 1 to 12 to be performed.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, cause the method of any one of claims 1 to 12 to be implemented.
27. A computer program product, characterized in that, The computer program product includes instructions that, when executed, cause the method of any one of claims 1 to 12 to be implemented.
28. A chip or chip system, characterized in that, Includes a processor for retrieving and executing instructions stored in a memory, causing a communication device with a chip mounted to perform the method as described in any one of claims 1 to 12.
29. A communication system, characterized in that, It includes a first device and a second device, the first device being used to perform the method according to any one of claims 1 to 10, and the second device being used to perform the method according to claim 11 or 12.