Data transmission method and application apparatus

By generating a three-dimensional feature information set and performing data fusion, the problem of low efficiency in data transmission and fusion of different modalities is solved, and the accuracy of data processing and image reconstruction is improved.

WO2026113777A1PCT designated stage Publication Date: 2026-06-04HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-10-27
Publication Date
2026-06-04

Smart Images

  • Figure CN2025130302_04062026_PF_FP_ABST
    Figure CN2025130302_04062026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a data transmission method and an application apparatus, which are applied to the technical field of sensing. The method comprises: first, receiving two-dimensional feature pair information obtained by means of performing pixel matching on at least two pieces of image data; and then, on the basis of the two-dimensional feature pair information, generating a three-dimensional feature information set, wherein the three-dimensional feature information set comprises at least one set of three-dimensional feature information, and each set of three-dimensional feature information comprises three-dimensional feature points generated using matched pixel pairs. Furthermore, the three-dimensional feature information set can also be fused with sensing data, so as to obtain fused three-dimensional feature points. In this way, data and / or feature information of the data can be transmitted, thereby facilitating data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Data transmission method and application device

[0001] This application claims priority to Chinese Patent Application No. 202411708993.8, filed on November 26, 2024, entitled “Data Transmission Method and Application Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of communication technology, and in particular to a data transmission method and application device. Background Technology

[0003] With the development of sensing technology, communication devices such as network equipment and terminal devices can acquire different types of sensing data, such as image data and sensing data acquired by communication devices based on the surrounding environment. Data fusion can be performed based on the feature information of different modalities of data to obtain more feature information, which is beneficial for further data processing. Therefore, effectively transmitting data or the feature information of data is a technical problem to be solved by those skilled in the art. Summary of the Invention

[0004] This application discloses a data transmission method and application device that can transmit data and / or data feature information, which is helpful for data fusion.

[0005] Firstly, this application discloses a first data transmission method, which can be applied to a first device, which can be a terminal device or a network device. The method includes: receiving two-dimensional feature pair information from a second device, the two-dimensional feature pair information indicating matching pixel pairs between first image data and second image data; generating a three-dimensional feature information set based on the two-dimensional feature pair information, the three-dimensional feature information set including at least one set of three-dimensional feature information, each set of three-dimensional feature information including three-dimensional feature points generated from the matching pixel pairs. Thus, it is possible to transmit three-dimensional feature information obtained by fusing at least two image data, facilitating further data fusion.

[0006] In conjunction with the first aspect, in some embodiments, the method may further include: receiving sensing data from the second device; and obtaining a registration result and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the sensing data. The three-dimensional feature point set can be understood as three-dimensional feature points converted from the camera coordinate system to the sensing coordinate system, or as three-dimensional feature points resulting from the fusion of two-dimensional feature pair information and sensing data. That is, it can be three-dimensional feature points in the camera coordinate system, or in the sensing coordinate system, or even in the world coordinate system; no limitation is made here. Thus, the three-dimensional feature information obtained by fusing at least two image data sets can be further fused with the sensing data, which helps to improve the representational ability of the fused data and the accuracy of image reconstruction.

[0007] In conjunction with the first aspect, in some embodiments, the method may further include fusing the three-dimensional feature point set and the perceptual data. This can enhance the representational capability of the fused perceptual data, thereby improving the accuracy of image reconstruction.

[0008] Optionally, the method may further include fusing the first image data and / or the second image data based on a three-dimensional feature point set and / or registration results. This can enhance the representational capability of the fused image data and improve the accuracy of image reconstruction.

[0009] In conjunction with the first aspect, in some embodiments, the method may further include: sending the registration result and / or the three-dimensional feature point set to the second device. Optionally, the first device may also send the registration result and / or the three-dimensional feature point set to other devices (such as acquisition devices). This eliminates the need to acquire data from different modalities for each data fusion, thus improving the efficiency of subsequent data fusion.

[0010] In conjunction with the first aspect, in some embodiments, the registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, and the transformation relationship between the perceptual coordinate system and the first camera coordinate system of the perceptual data. Thus, the camera coordinates corresponding to feature points in the image data or perceptual data can be determined based on the registration result.

[0011] In this embodiment, the transformation relationship between the first pixel coordinate system and the first camera coordinate system refers to the transformation relationship of the feature point from the first pixel coordinate system to the first camera coordinate system, which can be represented by a first transformation matrix. The transformation relationship between the second pixel coordinate system and the second camera coordinate system refers to the transformation relationship of the feature point from the second pixel coordinate system to the second camera coordinate system, which can be represented by a second transformation matrix. The transformation relationship between the perception coordinate system and the first camera coordinate system refers to the transformation relationship of the feature point from the first camera coordinate system to the perception coordinate system, which can be represented by a third transformation matrix.

[0012] Optionally, the registration result is also used to indicate at least one of the following: the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.

[0013] In this embodiment, the transformation relationship between the first camera coordinate system and the second camera coordinate system can be represented by a camera transformation matrix, which can be based on the first camera coordinate system. The transformation relationship between the first pixel coordinate system and the perception coordinate system can be obtained by multiplying the first transformation matrix and the third transformation matrix, for example, by multiplying the first transformation matrix and the third transformation matrix. The transformation relationship between the second pixel coordinate system and the perception coordinate system can be obtained by multiplying the second transformation matrix and the third transformation matrix, for example, by multiplying the second transformation matrix and the third transformation matrix. The transformation relationship between the perception coordinate system and the second camera coordinate system refers to the transformation relationship of feature points from the second camera coordinate system to the perception coordinate system.

[0014] It should be noted that the registration results can also be used to indicate the transformation relationship between other coordinate systems, such as the transformation relationship between the first pixel coordinate system and the world coordinate system, the transformation relationship between the second pixel coordinate system and the world coordinate system, etc., without limitation here. In this way, data fusion can be performed based on the transformation relationship between at least two coordinate systems.

[0015] In conjunction with the first aspect, in some embodiments, the three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in the three-dimensional feature information set. Thus, the registration results can be registered (or aligned or optimized) based on the three-dimensional feature points in this set of three-dimensional feature information, which helps to improve the representational capability of the fused data and the accuracy of image reconstruction.

[0016] It should be noted that a set of 3D feature points can belong to the optimal or locally optimal set of 3D feature information within a 3D feature information set. That is, the set of 3D feature points contains optimal 3D feature points, and the number or range of these optimal 3D feature points is the largest. 3D feature points in other sets of 3D feature information outside the 3D feature point set may also be optimal, but their number is less than the number of optimal 3D feature points in the 3D feature point set, or their coverage is smaller than the coverage of the optimal 3D feature points in the 3D feature point set. The optimal 3D feature point can be the 3D feature point corresponding to the minimum loss function determined by the vertical distance or chamfer distance.

[0017] In conjunction with the first aspect, in some embodiments, the method may further include: synchronizing configuration information with the second device; wherein the configuration information includes at least one of the following: the number of matching pixel pairs, at least one first pre-selected value group, and at least one second pre-selected value group, wherein the at least one first pre-selected value group is used to indicate pre-selected values ​​of unknown parameters of the first image data, and the at least one second pre-selected value group is used to indicate pre-selected values ​​of unknown parameters of the second image data. Optionally, the unknown parameter of each image data may be the focal length f, for example, the focal length on the U-axis of the pixel coordinate system. Focal length on the V-axis of the pixel coordinate system Furthermore, the unknown parameters for each image data point can also be the physical dimensions of the pixels, such as the physical dimension du of a pixel on the U-axis of the pixel coordinate system, or the physical dimension dv of a pixel on the V-axis of the pixel coordinate system. In this way, feature information of the data can be obtained based on the synchronized configuration information, thereby improving the accuracy of feature information acquisition.

[0018] In conjunction with the first aspect, in some embodiments, the method may further include: generating the three-dimensional feature information set based on the two-dimensional feature pair information, the at least one first pre-selected value group, and the at least one second pre-selected value group. Thus, the three-dimensional feature information set can be determined using pre-selected values ​​from image data to assist the two-dimensional feature pair information. For example, different first and second pre-selected value groups in the three-dimensional feature information set correspond to different three-dimensional feature information, which can improve the accuracy of obtaining the three-dimensional feature information set.

[0019] In conjunction with the first aspect, in some embodiments, the configuration information further includes the number of image data, which includes at least the first image data and the second image data. Thus, the number of pairwise combinations of image data can be determined based on the number of image data, thereby determining the number of two-dimensional feature pairs and improving the accuracy of feature information acquisition.

[0020] In conjunction with the first aspect, in some embodiments, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data. Thus, the two-dimensional features of the image can be located based on the image size of the image data, which helps to improve the accuracy and convenience of data fusion.

[0021] In conjunction with the first aspect, in some embodiments, the three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information. This helps to improve the accuracy of processing the three-dimensional feature information set.

[0022] In conjunction with the first aspect, in some embodiments, the registration result is obtained based on periodic or triggered information. That is, the acquisition device can be triggered to acquire image data and sensory data when the periodic time arrives, so that the first device and the second device can obtain the registration result according to the method provided in this application. Alternatively, it can be triggered by an event, such as when the sensory coordinate system of the acquisition device changes, or when the position of the acquisition device changes, so that the first device and the second device can obtain the registration result according to the method provided in this application.

[0023] Optionally, if the coordinate system corresponding to the registration result remains unchanged, data fusion can be performed directly based on the registration result. If the coordinate system corresponding to the registration result changes, a new registration result can be obtained, and data fusion can be performed based on this newly obtained registration result. Obtaining the registration result periodically or triggered by specific events helps improve the accuracy of data fusion.

[0024] Secondly, embodiments of this application disclose a second data transmission method, which can be applied to a second device, which can be a terminal device or a network device. The method includes: acquiring two-dimensional feature pair information, the two-dimensional feature pair information being used to indicate matching pixel pairs between first image data and second image data; and sending the two-dimensional feature pair information to a first device.

[0025] In conjunction with the second aspect, in some embodiments, the method further includes: sending sensing data to the first device.

[0026] In conjunction with the second aspect, in some embodiments, the method further includes: receiving the registration result of the first device and / or a set of three-dimensional feature points.

[0027] In conjunction with the second aspect, in some embodiments, the registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, and the transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data.

[0028] Optionally, the registration result is also used to indicate at least one of the following: the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.

[0029] In conjunction with the second aspect, in some embodiments, the three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in a three-dimensional feature information set; wherein, the three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated by the matching pixel pair.

[0030] In conjunction with the second aspect, in some embodiments, the method further includes: synchronizing configuration information with the first device; wherein the configuration information includes at least one of the following: the number of matching pixel pairs, at least one first preselected value group and at least one second preselected value group, the at least one first preselected value group being used to indicate a preselected value of an unknown parameter in a first pixel coordinate system of the first image data, and the at least one second preselected value group being used to indicate a preselected value of the unknown parameter in a second pixel coordinate system of the second image data.

[0031] In conjunction with the second aspect, in some embodiments, the configuration information further includes the number of image data, which includes at least the first image data and the second image data.

[0032] In conjunction with the second aspect, in some embodiments, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.

[0033] In conjunction with the second aspect, in some embodiments, the three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.

[0034] In conjunction with the second aspect, in some embodiments, the registration result is obtained based on periodic or triggered information.

[0035] Optionally, the second aspect is implemented by a second device. The specific content of the second aspect corresponds to that of the first aspect, and the corresponding features and beneficial effects of the second aspect can be referred to the description of the first aspect. To avoid repetition, detailed descriptions are appropriately omitted here.

[0036] Thirdly, embodiments of this application disclose a communication device, including units, modules, or means for performing the steps of the first aspect, the second aspect, or any of the implementation methods described above. The modules, units, or means can be implemented by software, by hardware, or by a combination of software and hardware.

[0037] Fourthly, embodiments of this application disclose another communication device. This communication device may include one or more processors, which are configured to execute methods described above, either by executing instructions in memory or by using logic circuitry, to perform any of the methods described above or any possible examples.

[0038] In some embodiments, the communication device may further include an interface circuit, wherein the processor is configured to communicate with other devices or components through the interface circuit.

[0039] In some embodiments, the communication device further includes the memory.

[0040] In conjunction with the third or fourth aspect, in some embodiments, the communication device may be a first device or a second device.

[0041] In the embodiments of this application, the first device and the second device can be a terminal device or a network device. The terminal device can be a terminal as a final product, a component or module with terminal functions, or a communication chip (such as a processor, baseband chip, or chip system) that can be applied in a terminal. The network device can be a network device as a final product, a component or module with network device functions, or a communication chip (such as a processor, baseband chip, or chip system) that can be applied in a network device.

[0042] Fifthly, embodiments of this application provide a communication system, which includes a first device and a second device. When the first device and the second device are running in the communication system, the first device is used to execute the method described in the first aspect or its embodiments, and the second device is used to execute the method described in the second aspect or its embodiments.

[0043] Sixthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed, cause the methods described above or in the embodiments thereof to be implemented.

[0044] In a seventh aspect, embodiments of this application provide a computer program product including instructions that, when executed, cause the methods in any of the above aspects or possible examples to be implemented.

[0045] In conjunction with the sixth or seventh aspect, the instructions can be executed by the computer or the processor in the computer.

[0046] Eighthly, this application provides a chip or chip system including at least one processor for calling and executing instructions stored in a memory, causing a communication device on which the chip or chip system is mounted to perform any of the above-described methods or possible examples.

[0047] Optionally, the chip or chip system may also include memory.

[0048] Ninthly, this application provides another chip, including: an input interface, an output interface, and a processing circuit. The input interface, the output interface, and the processing circuit are connected via internal connection paths. The processing circuit is used to execute the method of any of the above aspects or possible examples. Optionally, the chip also includes a memory. The input interface, the output interface, the processor, and the memory are connected via internal connection paths. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method of any of the above aspects or possible examples.

[0049] In a tenth aspect, this application provides another chip system including at least one processor and a communication interface, the communication interface and at least one processor being interconnected via a line, the at least one processor being used to run a computer program or instructions to perform the methods in any of the above aspects or possible examples.

[0050] Optionally, the implementation and beneficial effects of the above-mentioned aspects can be referenced from each other. Attached Figure Description

[0051] The accompanying drawings used in the embodiments of this application are described below.

[0052] Figures 1A and 1B are schematic diagrams of the architecture of a communication system applied to a data transmission method according to an embodiment of this application;

[0053] Figure 2 is a schematic diagram illustrating the principle of camera imaging provided in an embodiment of this application;

[0054] Figure 3 is a bitmap of matching pixel pairs provided in an embodiment of this application;

[0055] Figure 4 is a schematic diagram of W-block two-dimensional feature pair information provided in an embodiment of this application;

[0056] Figure 5 is a schematic diagram of three-dimensional feature information provided in an embodiment of this application;

[0057] Figure 6 is a flowchart illustrating a data transmission method provided in an embodiment of this application;

[0058] Figure 7A is a schematic diagram of the perception data fusion process before and after according to an embodiment of this application;

[0059] Figure 7B is a schematic diagram of the fusion of first image data and second image data provided in an embodiment of this application;

[0060] Figures 8A to 8D are interactive schematic diagrams of another data transmission method provided in the embodiments of this application;

[0061] Figure 9 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0062] Figure 10 is a schematic diagram of another communication device provided in an embodiment of this application;

[0063] Figure 11 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0064] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0065] In this application, unless otherwise specified, "at least one" means "one or more".

[0066] Please refer to Figure 1A or Figure 1B, which are schematic diagrams of the architecture of a communication system applied to a data transmission method according to an embodiment of this application. The communication system may include, but is not limited to, at least one of the following: Long Term Evolution (LTE) communication system, New Radio (NR) communication system, LTE Advanced (LTE-A) communication system, Device-to-Device (D2D) communication system, Vehicle-to-Everything (V2X) communication system, Machine-to-Machine (M2M) communication system, Internet of Things (IoT) communication system, Narrow Band Internet of Things (NB-IoT) communication system, Integrated Sensing and Communication System, Frequency Division Duplex (FDD) communication system, Time Division Duplex (TDD) communication system, Non-Terrestrial Network (NTN) communication system, Wireless Projection Communication System, Integrated Access and Backhaul (IAB) communication system, Public Land Mobile Network (PLMN) communication system, and Non-Public Network (NPN) communication system. Network (NPN) communication systems, as well as those applied to future communication systems, or non-3rd generation partnership project (3GPP) communication systems, etc.

[0067] As shown in Figure 1A or Figure 1B, the communication system may include a first device 10 and a second device 20. This application does not limit the number of first devices 10 and second devices 20; Figures 1A and 1B illustrate one first device 10 and one second device 20. In practice, the communication system may include one or more first devices 10 and one or more second devices 20.

[0068] As shown in Figure 1B, optionally, the communication system may also include a data acquisition device, which can be used to acquire image data and / or sensory data, etc. It should be noted that Figure 1B uses a single acquisition device 30 as an example, and the acquisition device 30 and the second device 20 are different devices. In practice, the acquisition device 30 can be either a second device or a first device, such as the second device 20 in Figure 1A, or a communication device not shown in Figures 1A and 1B. When the acquisition device is a second device, the second device can send the image data and / or sensory data acquired by the second device, or characteristic information obtained from data processing, to the first device, which helps improve data processing efficiency. When the acquisition device is a first device, the first device can process the image data and / or sensory data acquired by the first device, which can improve data processing efficiency.

[0069] This application does not limit the number of image data and sensing data. For example, the number of image data may be two, and the number of sensing data may be one. In the embodiments of this application, all image data and sensing data can be collected by one acquisition device, as shown by the second device 20 in FIG1A, which can be used as an acquisition device to collect all image data and sensing data; or all image data can be collected by one acquisition device, and all sensing data can be collected by another acquisition device, as shown by the second device 20 in FIG1B, which can be used as an acquisition device to collect all image data, and the acquisition device 30 in FIG1B, which is used to collect sensing data; or image data can be collected by at least one acquisition device, and sensing data can be collected by the acquisition device that collects the image data or by other acquisition devices, as shown by the second device 20 in FIG1B collecting the first image data, the acquisition device 30 in FIG1B collecting the second image data, and sensing data can be collected by the acquisition device 30 in FIG1B, the second device 20, or a communication device not shown in FIG1B, etc. Optionally, the acquisition device for collecting image data can be the same as the acquisition device for collecting sensing data, or the acquisition device for collecting image data can be different from the acquisition device for collecting sensing data.

[0070] Optionally, the data acquisition device can be a terminal device or a data acquisition module within a terminal device. The data acquisition module may include an image acquisition device, such as a camera, for acquiring image data. The data acquisition module may also include a sensing module, such as a module used to acquire sensing data in systems involving radar, Bluetooth, Wi-Fi, 5G, etc.

[0071] The principle of collecting sensing data can be referenced from the principle of radar. That is, the transmitter sends electromagnetic waves, which are reflected by the object to be sensed and then acquired by the receiver. The acquired reflected signals are further processed into sensing results, which can show information such as the size and outline of the object. The transmitter and receiver that perform the sensing are called sensing nodes.

[0072] Image data can include images, videos, etc., and can be used to identify the two-dimensional image features of objects, such as two-dimensional coordinates and color. Perception data, such as the aforementioned reflection signals and perception results, can include scene information; for example, perception data includes the position, color, and shape of objects inside the vehicle. In this embodiment, image data is an image. Perception data can be point cloud data, also known as laser point cloud, 3D point cloud, or simply point cloud. It is a collection of massive points representing the spatial distribution and surface characteristics of a target object, obtained by using a laser to acquire the three-dimensional spatial coordinates (usually represented in x, y, z three-dimensional coordinates) of each sampling point on the object's surface in the same spatial reference frame. Compared to images, point clouds, while lacking detailed texture information, contain rich three-dimensional spatial information. In addition to three-dimensional spatial information, point cloud data can also include some or all of color information, grayscale values, depth, segmentation results, and time delay, etc., which are not limited here.

[0073] This application can be applied to scenarios such as autonomous / assisted driving, V2X, drones, 3D map reconstruction, smart cities, smart homes, factories, healthcare, and maritime sectors. For example, in autonomous driving scenarios, vehicles or drones can generate dynamic maps based on perception data, and / or identify and alert to hazardous events based on perception data.

[0074] This application does not limit the triggering conditions for the acquisition device to collect sensing data. It can be triggered by an event, such as when the sensing coordinate system of the acquisition device changes or the position of the acquisition device changes. Alternatively, it can be triggered by time, for example, by pre-configuring a period, and the acquisition device collects sensing data when the period time arrives.

[0075] As shown in Figure 1A or Figure 1B, the first device 10 can be a network device, and the second device 20 can be a terminal device. In fact, both the first device and the second device can be terminal devices, or both the first device and the second device can be network devices; or the first device can be a terminal device, and the second device can be a network device.

[0076] Terminal devices and network devices, network devices and network devices, and terminal devices and terminal devices can communicate using licensed spectrum, unlicensed spectrum, or both simultaneously. This application does not limit the spectrum resources used by terminal devices and network devices.

[0077] Terminal devices can connect to network devices wirelessly or via wired connections, enabling uplink (UL) or downlink (DL) communication. Terminal devices can also connect to each other wirelessly or via wired connections, allowing for sidelink (SL) communication.

[0078] The terminal equipment involved in this application is an entity on the user side used to receive or transmit signals, providing voice and / or data to the user. Terminal equipment can be a terminal, user equipment (UE), access terminal, UE unit, UE station, mobile device, mobile station, mobile station, mobile terminal, mobile client, mobile unit, remote station, remote terminal, remote unit, wireless unit, wireless communication equipment, user agent, or user device, etc. Among them, the access terminal can be a cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication capabilities, computing device or other processing device connected to a wireless modem, vehicle-mounted device, wearable device, or terminal in a future communication system, etc. Hereinafter, it is sometimes simply referred to as a terminal. In Figures 1A and 1B and this application document, the terminal equipment is described using a vehicle or UE as an example.

[0079] It should be noted that the terminal device described in the embodiments of this application can be a terminal as a final product, such as the various terminal devices mentioned above, or it can be a component or part with terminal functions, or it can be a communication chip (such as a processor, baseband chip, or chip system, etc.) that can be applied in a terminal. That is to say, components, parts, or chips applied in the above-mentioned devices also belong to terminal devices.

[0080] In Figure 1A or Figure 1B, network devices are exemplified as access network (AN) nodes. An access network node can be a radio access network (RAN) device, which connects terminal devices to the wireless network. In other words, the access network provides access services to terminal devices, enabling them to access (or connect to) the network. The access network can support both wired and wireless access.

[0081] Optionally, the access network includes multiple AN / RAN nodes. AN / RAN nodes may include, but are not limited to: access points (APs), enhanced node Bs (eNBs), home evolved Node Bs (HNBs), baseband units (BBUs), next-generation node Bs (gNBs), transmission reception points (TRPs), transmission points (TPs), or other access nodes, such as wireless relay nodes or wireless backhaul nodes. AN / RAN nodes may be one or more antenna panels, or network nodes constituting gNBs or transmission points, such as BBUs or distributed units (DUs), or devices performing RAN functions in communication systems such as D2D, V2X, M2M, and U2U. AN / RAN nodes can be radio controllers in cloud radio access network (CRAN) scenarios, open RAN (O-RAN or ORAN), or access networks in future communication systems, etc., without any limitations.

[0082] The network device described in the embodiments of this application can be a network device as a final product, such as the various network devices mentioned above, or it can be a component or part with network device functions, or it can be a communication chip (such as a processor, baseband chip, or chip system, etc.) that can be applied in a network device.

[0083] In the communication system shown in Figure 1A or Figure 1B, although access network nodes (such as the first device 10) and terminal devices (such as the second device 20) are shown, the application scenario may not be limited to including access networks and terminal devices. For example, it may also include devices for carrying virtualized network functions, which are obvious to those skilled in the art and will not be described in detail here.

[0084] Furthermore, the number and type of network devices and terminal devices included in the communication system shown in Figure 1A or Figure 1B are merely examples, and the embodiments of this application are not limited thereto. For example, it may also include more or fewer terminal devices communicating with the network devices. As another example, it may also include more or fewer network devices communicating with the terminal devices. For the sake of brevity, they are not described one by one in the accompanying drawings.

[0085] Optionally, the communication system may also include network devices not shown in Figure 1A or Figure 1B, such as core network (CN) devices, data network devices, etc.

[0086] In different communication systems, core network equipment (hereinafter referred to as core network) can correspond to different devices. For example, in a 3G communication system, it can correspond to the Serving GPRS Support Node (SGSN) and / or the Gateway GPRS Support Node (GGSN); in a 4G communication system, it can correspond to the Mobility Management Entity (MME) and / or the Serving Gateway (S-GW); and in a 5G communication system, it can correspond to the aforementioned Policy Control Function (PCF) network elements, Unified Data Management (UDM) network elements, Application Function (AF) network elements, Access and Mobility Management Function (AMF) network elements, Session Management Function (SMF) network elements, Location Management Function (LMF) network elements, and User Plane Function (UPF) network elements, etc.

[0087] Among them, the UPF network element is responsible for managing the transmission of user plane data and quality of service (QoS) control, traffic statistics and other functions. It can perform user data packet forwarding according to the routing rules of the session management network element, such as sending uplink data to the data network or other user plane network elements, and forwarding downlink data to other user plane network elements or (R)AN network elements.

[0088] The AMF (Access Default Mode) network element is responsible for user access management, security authentication, and mobility management. The LMF (Local Mode Default Mode) network element manages and controls location service requests from target terminals and processes location-related information. The SMF (Supply, Service Default Mode) network element manages sessions, allocating and releasing resources for terminal device sessions. The UDM (User Default Mode) network element manages the context of user subscriptions, such as storing terminal device subscription information. The PCF (Policy and Charging Rules Function) network element is responsible for user policy management. Similar to the Policy and Charging Rules Function (PCRF) network element in LTE, it is primarily responsible for policy authorization, quality of service (QoS), and generating charging rules, and distributing these rules to the UPF (User Default Mode) network element via the SMF network element to complete the installation of the corresponding policies and rules. The AF (Application Default Mode) network element can be a third-party application control platform or the operator's own equipment. The AF network element is responsible for application management and can provide services to multiple application servers.

[0089] In this embodiment, the data network device is hereinafter referred to as the data network. The data network is used to provide business services to users. Generally, the client is a terminal, and the server is the data network. The data network provided by the data network may include a private network, such as a local area network (LAN). The data network may also include an external network not managed by an operator, such as the Internet. Alternatively, the data network may include a proprietary network jointly deployed by operators, such as a network providing Internet Protocol Multimedia Subsystem (IMS) services.

[0090] In some embodiments, the first device 10, the second device 20, the acquisition device, the network device, and the terminal device may also be referred to as communication devices, which may be a general-purpose device or a special-purpose device. This application does not specifically limit this.

[0091] In this embodiment, the terminal device or network device includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also referred to as main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Furthermore, this embodiment does not specifically limit the specific structure of the execution entity of the method provided in this embodiment, as long as it can communicate according to the method provided in this embodiment by running a program that records the code of the method provided in this embodiment. For example, the execution entity of the method provided in this embodiment can be a terminal device or a network device, or a functional module in the terminal device or network device that can call and execute a program.

[0092] Furthermore, various aspects or features of this application can be implemented as methods, apparatus, or articles of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" as used herein encompasses a computer program accessible from any computer-readable device, carrier, or medium. For example, computer-readable media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical discs (e.g., compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memory (EPROMs), cards, sticks, or key drives, etc.). The various storage media described herein may represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.

[0093] To facilitate understanding of the embodiments of this application, definitions of technical terms that may appear in the embodiments of this application are given below. The terminology used in the implementation section of this application is only used to explain specific embodiments of this application and is not intended to limit this application.

[0094] (1) The process of object imaging is essentially a transformation of several coordinate systems.

[0095] Please refer to Figure 2, which is a schematic diagram illustrating the principle of camera imaging according to an embodiment of this application. As shown in Figure 2, the camera coordinate system is based on the optical center O. c With the origin as the axis, the direction of the optical axis as the z-axis, and the x and y directions parallel to the image as the x and y axes, respectively, the x, y, and z axes can be referred to as X, Y, Z, and X, respectively. c Y c Z c The unit is length. The pixel coordinate system takes the vertex (u, v) of the image as the origin, and the u and v directions are parallel to the x and y directions, respectively. Its x-axis and y-axis can be called the U-axis and V-axis, respectively, and the unit is pixels.

[0096] In the embodiments of this application, the transformation relationship between the pixel coordinate system and the camera coordinate system can be obtained or satisfied based on the following formula (1), and the schematic diagram can be shown in Figure 2.

[0097] Where u is the coordinate of the pixel (or pixel point or feature point) on the U-axis of the pixel coordinate system, and v is the coordinate of the pixel on the V-axis of the pixel coordinate system. i It is a point (i.e., the aforementioned pixel) in the X coordinate system. c The coordinates on the x-axis, y i It is the Y-axis of the point in the camera coordinate system. c The coordinates on the (y) axis, z i It is the Z-axis of the point in the camera coordinate system. c The coordinates on the (z) axis. K is the transformation matrix from the pixel coordinate system to the camera coordinate system, which indicates the transformation relationship from 2D pixel coordinates to 3D camera coordinates. K satisfies formula (2):

[0098] Where w is the width of the image and h is the height of the image. u It is the X-axis of the point in the camera coordinate system. c The focal length on the (x) axis, f v It is the point in the Y-axis of the camera coordinate system c Focal length on the (y) axis. du is the physical size of a pixel on the U-axis (u direction) of the pixel coordinate system, and dv is the physical size of a pixel on the V-axis (v direction) of the pixel coordinate system.

[0099] In the embodiments of this application, the transformation relationship between the camera coordinate system and the perception coordinate system can be obtained or satisfied based on the following formula (3) so that the points in the camera coordinate system are displayed in the form of a three-dimensional point cloud.

[0100] Where, x i yi and z i As mentioned above, this will not be repeated here. si It is the x-coordinate of a point on the perceptual coordinate system, y si It is the y-coordinate of the point on the perceptual coordinate system, z-coordinate. si This is the coordinate of the point on the z-axis of the sensing coordinate system. M is the transformation matrix from the camera coordinate system to the sensing coordinate system, which indicates the transformation relationship from 3D camera coordinates to 3D sensing coordinates. M satisfies formula (4):

[0101] Where R is a 3x3 rotation matrix, and T is a 3x1 translation matrix, or translation vector.

[0102] (2) Two-dimensional feature pair information, used to indicate two-dimensional feature information of matching pixel pairs in at least two image data (e.g., at least two images).

[0103] In this embodiment, two-dimensional feature pair information is used to indicate matching pixel pairs between the first image data and the second image data. A matching pixel pair consists of matching pixels. This application does not limit the matching method for pixels between image data (or images). For example, the first image data and the second image data can be input into a local feature transformer (LoFTR) matching model to obtain matching pixel pairs between the first image data and the second image data.

[0104] The matching pixel pair between the first image data and the second image data includes a first pixel in the first image data and a second pixel in the second image data that matches that first pixel. That is, if the first pixel has no matching pixel in the second image data, the matching pixel pair between the first and second image data does not include the matching pixel pair of the first pixel. If the first pixel has a matching pixel in the second image data, and that matching pixel is the second pixel, then the matching pixel pair between the first and second image data includes the matching pixel pair of the first pixel, and this matching pixel pair includes both the first and second pixel. The above example uses the matching pixel pair of the first pixel in the first image data. In the case where the first pixel and the second pixel match, the matching pixel pair of the first pixel can also be referred to as the matching pixel pair of the second pixel.

[0105] Alternatively, the position of the matching pixel pair can be represented using values ​​from an array.

[0106] When the number of matching pixel pairs is greater than 1, for example, if the number of matching pixel pairs is N, then N matching pixel pairs can be represented as an array, such as {(p11 p 21 ), ..., (p 1N p 2N )}, or (p 1i p 2i ), i∈{1,…,N}. Here, each bracket represents a matching pixel pair, and the two values ​​within the brackets represent the positions of the two pixels in the matching pixel pair within their respective pixel coordinates in the image data. (p 11 p 21 (p) represents the first matching pixel pair, and so on, (p) 1N p 2N (p) represents the Nth matching pixel pair. 11 p 21 (p) provides an example. 11 p 21 p in ) 11 This can be the position of the first pixel in the first image data within the pixel coordinate system of the first image data, for example, p 11 This includes the position of the first pixel on the u-axis of the pixel coordinate system of the first image data and the position of the first pixel on the v-axis of the pixel coordinate system of the first image data. 21 This can be the position of the second pixel in the second image data that matches the first pixel in the second image data within the pixel coordinate system of the second image data, for example, p 21 This includes the position of the second pixel on the u-axis of the pixel coordinate system of the second image data and the position of the second pixel on the v-axis of the pixel coordinate system of the second image data.

[0107] Alternatively, the position of the matching pixel pair can be represented using values ​​from a bitmap.

[0108] For example, please refer to Figure 3, which is a bitmap of matching pixel pairs provided in an embodiment of this application. This bitmap is used to indicate the position of the matching pixel pair between first image data and second image data in the pixel coordinate system of the first or second image data. As shown in Figure 3, the image can be processed into multiple grids according to its width (w) and height (h), each grid representing the position of a pixel in the image. If the value in the grid is 1, it indicates that the pixel at that position is a matching pixel between the first and second image data. If the value in the grid is 0, it indicates that the pixel at that position is a non-matching pixel between the first and second image data. For example, P... 11 P represents the position of a pixel that matches between the first image data and the second image data. 14 and P 41 These represent the positions of a single pixel that does not match between the first and second image data.

[0109] In some embodiments, the two-dimensional feature pair information can also be used to indicate the image size of the first image data and / or the image size of the second image data. The image size of the image data can include the width and height of the image data. As shown in Figure 3, the two-dimensional feature pair information to which the matching pixel pair in the bitmap representation belongs includes the width and height of the image size. Thus, the two-dimensional features of the image data, such as matching pixel pairs between the image data and other image data, can be located based on the image size of the image data, which helps improve the accuracy and convenience of data fusion.

[0110] In this embodiment, data fusion can be understood as coordinate alignment. The pixel coordinate system of the first image data can be called the first pixel coordinate system, and the pixel coordinate system of the second image data can be called the second pixel coordinate system. The camera coordinate system of the first image data can be called the first camera coordinate system, and the camera coordinate system of the second image data can be called the second camera coordinate system. The transformation matrix between the first pixel coordinate system and the first camera coordinate system can be called the first transformation matrix, such as K1, which is used to indicate the transformation relationship of feature points from the first pixel coordinate system to the first camera coordinate system. The transformation matrix between the second pixel coordinate system and the second camera coordinate system can be called the second transformation matrix, such as K2, which is used to indicate the transformation relationship of feature points from the second pixel coordinate system to the second camera coordinate system. The first pixel coordinate system can be understood as the coordinate system of pixels in the first image data, and the second pixel coordinate system can be understood as the coordinate system of pixels in the second image data. The first camera coordinate system can be understood as the camera coordinate system when the acquisition device acquires the first image data, and the second camera coordinate system can be understood as the camera coordinate system when the acquisition device acquires the second image data.

[0111] In this embodiment, two-dimensional feature pair information can be obtained based on the first image data. Optionally, when a bitmap represents the matching pixel pairs between the first and second image data, the width and height of the bitmap can be the image size of the first image data. In this way, the second image data can be mapped to the first image data, so that points corresponding to the same spatial location in the two images correspond one-to-one, thereby achieving data fusion. The image corresponding to the first image data is used as a reference to determine the pixel points that match the second image data with the first image data.

[0112] Optionally, the first camera coordinate system and the second camera coordinate system can be different, and the first pixel coordinate system and the second pixel coordinate system can also be different. The first image data and the second image data are different image data. When the first image data and the second image data are different, an image with overlapping areas between the first image data and the second image data is needed to obtain matching pixels between them.

[0113] In the embodiments of this application, the first image data and the second image data can be images acquired by the same acquisition device, or they can be images acquired by different acquisition devices.

[0114] This application uses first image data and second image data as examples. In some embodiments, multiple image data may also exist. If the number of image data is greater than 2, the processing method for the first and second image data can be referred to, such as matching one image data with another image data in pairs to obtain W blocks of two-dimensional feature pair information. Here, W is an integer greater than 1, specifically the number of matching two image data in the multiple image data. Each two-dimensional feature pair information in the W blocks of two-dimensional feature pair information may include N matching pixel pairs, where N is a positive integer.

[0115] Taking an example with three image data sets, the image data includes a first image data set, a second image data set, and a third image data set. The third image data set can be matched with the first image data set, and it can also be matched with the second image data set. As shown in Figure 4, the first two-dimensional feature pair information obtained by matching the first image data set and the second image data set is as follows: {(p 11 p 21 ), ..., (p 1N p 2N The second two-dimensional feature pair information obtained by matching the first image data and the third image data is as follows: {(p 11 p 31 ), ..., (p 1N p 3N The second and third image data can be matched to obtain a third block of two-dimensional feature pairs, such as {(p 21 p 31 ), ..., (p 2N p 3N )}.

[0116] Optionally, if the number of matching pixel pairs in one of the W blocks of two-dimensional feature pairs is less than N, it can be indicated by default non-matching pixel pairs (e.g., (-1, -1)). If the number of matching pixel pairs in one of the W blocks of two-dimensional feature pairs is greater than N, N matching pixel pairs can be selected from that block of two-dimensional feature pairs for indication. Alternatively, the value of N can be increased so that each block of two-dimensional feature pairs in the W blocks includes the increased number of N matching pixel pairs.

[0117] Optionally, if the overlap area between the first image data and the second image data is less than a first threshold, the number of image data can be greater than 2. If the overlap area between the first image data and the second image data is greater than the first threshold, the number of image data can be equal to 2.

[0118] This application does not limit the first threshold, for example, 80%. The more overlapping areas between two image data sets, the more matching pixel pairs, and the more features can be fused. The fewer overlapping areas between two image data sets, the fewer matching pixel pairs, and the fewer features can be fused. That is, when the overlapping area between the first and second image data sets is greater than the first threshold, most of the image features can be obtained based on the first and second image data sets. When the overlapping area between the first and second image data sets is less than the first threshold, the image data can include other image data or perceptual data in addition to the first and second image data sets to obtain more fusion features, which is beneficial for subsequent feature fusion.

[0119] Optionally, if image data other than the first image data and the second image data exists, the two-dimensional feature pair information may further include the image size of the image data, and may include matching pixel pairs obtained by matching the image data with other image data (e.g., the first image data, the second image data, etc.). That is, the two-dimensional feature pair information may include the image size of each image data, and the matching pixel pairs between every two image data.

[0120] (3) Three-dimensional feature information, used to indicate the three-dimensional features of feature points.

[0121] In this embodiment, the three-dimensional feature information can be generated based on two-dimensional feature pair information, used to indicate the three-dimensional features corresponding to at least two two-dimensional image data. The three-dimensional feature information includes three-dimensional feature points corresponding to matching pixel pairs, which can be used to represent the depth information of the matching pixels in the matching pixel pair, and can be understood as the coordinates of the matching pixels on the z-axis. This application does not limit the method for generating the three-dimensional feature information; optionally, the camera transformation matrix M between the first camera coordinate system and the second camera coordinate system can be obtained. cam Based on M cam Obtain the three-dimensional feature information corresponding to the two-dimensional feature pair information.

[0122] Among them, M cam Used to indicate the transformation relationship between the first camera coordinate system and the second camera coordinate system. This application is for obtaining M cam The method is not limited, M camThis can be obtained using computer vision techniques, such as Pycolmap and OpenCV. Optionally, the transformation matrix M between the first and second camera coordinate systems... cam The first camera coordinate system can be used as a reference. That is, under the first camera coordinate system, the point cloud features corresponding to each pixel in the matching pixel pair are determined, thereby obtaining the three-dimensional feature information.

[0123] Based on M cam A schematic diagram illustrating the acquisition of 3D feature information corresponding to 2D feature pairs can be found in Figure 5. In Figure 5, cam1 represents the first camera coordinate system, including the X1, Y1, and Z1 axes. cam2 represents the second camera coordinate system. The first pixel coordinate system of the first image data includes the U1 and V1 axes, and the width of the first image data is w1, and the height is h1. The second pixel coordinate system of the second image data includes the U2 and V2 axes, and the width of the second image data is w2, and the height is h2. (u 1i v 1i (u) represents a pixel in the first image data. 2i v 2i ) indicates that in the second image data, (u) 1i v 1i The matched pixels. d represents (u 1i v 1i ) and (u 2i v 2i The vertical distance between (x) and (x). 1i y 1i , z 1i ) represents (u 1i v 1i The coordinates are in the first camera coordinate system. As shown in Figure 5, using the first camera coordinate system as a reference, the point cloud features corresponding to the matching pixel pairs between the first image data and the second image data can be obtained, thereby obtaining the three-dimensional feature information between the first image data and the second image data.

[0124] Optionally, the first transformation matrix may have at least one first preselected value group, and the second transformation matrix may have at least one second preselected value group. The at least one first preselected value group is used to indicate preselected values ​​of unknown parameters of the first image data, and the at least one second preselected value group is used to indicate preselected values ​​of unknown parameters of the second image data.

[0125] In this embodiment, the unknown parameter for each image data can be the focal length f, for example, the focal length on the U-axis of the pixel coordinate system. Focal length on the V-axis of the pixel coordinate system And so on. The unknown parameters for each image data point can also be the physical size of the pixels, such as the physical size *du* of the pixel on the U-axis of the pixel coordinate system, or the physical size *dv* of the pixel on the V-axis of the pixel coordinate system. Thus, given the physical size of the pixels, according to the aforementioned formula... The physical dimensions of pixels can determine an unknown focal length; or, given a known focal length, the physical dimensions of unknown pixels can be determined using this formula and the focal length. This application uses focal length and pixel physical dimensions as examples of unknown parameters of image data. In some embodiments, the unknown parameters may also be other information.

[0126] Optionally, at least one first preselected value group and at least one second preselected value group can be represented in the form of a list.

[0127] For example, the list of at least one first pre-selected value group and at least one second pre-selected value group may include the initial value, ending value, and step size of the unknown parameter of the first image data, and the initial value, ending value, and step size of the unknown parameter of the second image data, etc. Taking the focal length as an example, the list of at least one first pre-selected value group and at least one second pre-selected value group is as follows: {f u1_0 f v1_0 f u2_0 f v2_0 f u1_n f v1_n f u2_n f v2_n s u1 s v1 s u2 s v2}wait.

[0128] Among them, f u1_0 This can represent the initial value of the focal length on the U-axis of the first pixel coordinate system, f. v1_0 This can represent the initial value of the focal length on the V-axis of the first pixel coordinate system. u1_n f can represent the ending value of the focal length on the U-axis of the first pixel coordinate system. v1_n This can represent the ending value of the focal length on the V-axis of the first pixel coordinate system. u1 The step size s can represent the focal length on the U-axis of the first pixel coordinate system. v1 This can represent the step size of the V-axis focal length in the first pixel coordinate system. u2_0 This can represent the initial value of the focal length on the U-axis of the second pixel coordinate system, f. v2_0 This can represent the initial value of the focal length on the V-axis of the second pixel coordinate system. u2_n f can represent the ending value of the focal length on the U-axis of the second pixel coordinate system. v2_n This can represent the ending value of the focal length on the V-axis of the second pixel coordinate system. s u2The step size s can represent the focal length on the U-axis of the second pixel coordinate system. v2 It can represent the step size of the focal length on the V-axis of the second pixel coordinate system.

[0129] When the initial and final values ​​of the unknown parameters are unequal, multiple unknown parameters of the first image data are determined based on the initial, final, and step values ​​of the unknown parameters of the first image data. Similarly, multiple unknown parameters of the second image data are determined based on the initial, final, and step values ​​of the unknown parameters of the second image data. When the initial and final values ​​of the unknown parameters of the first image data are equal, the number of groups (or individuals) in the first pre-selected value group or the number of candidate values ​​for the unknown parameters of the first image data is 1. Alternatively, when the step size of the unknown parameters of the first image data is 0, and only the initial or final values ​​of the unknown parameters of the first image data are included, the number of groups (or individuals) in the first pre-selected value group or the number of candidate values ​​for the unknown parameters of the first image data is 1. Similarly, the number of groups (or individuals) in the second pre-selected value group or the number of candidate values ​​for the unknown parameters of the second image data can be determined.

[0130] For example, a list of at least one first preselected value group and at least one second preselected value group, or a set of each preselected value group, such as etc. Among them, This represents at least one pre-selected value of the focal length on the U-axis of the first pixel coordinate system. This represents at least one pre-selected value of the focal length on the V-axis of the first pixel coordinate system. u2_0 , ...} represents at least one pre-selected value of the focal length on the U-axis of the second pixel coordinate system, {f v2_0 , ...} represents at least one pre-selected value of the focal length on the V-axis of the second pixel coordinate system.

[0131] If the first transformation matrix has at least one first preselected value group and the second transformation matrix has at least one second preselected value group, then it can be based on M. cam For each first pre-selected value group and each second pre-selected value group, the point cloud features corresponding to the matching pixel pairs in the two-dimensional feature pair information are obtained to obtain a three-dimensional feature information set. This three-dimensional feature information set includes at least one set of three-dimensional feature information. Each set of three-dimensional feature information is two-dimensional feature pair information (or M determined by the two-dimensional feature pair information). cam This is a 3D feature generated from a first set of pre-selected values ​​and a second set of pre-selected values. In other words, the first set of pre-selected values ​​and / or the second set of pre-selected values ​​are different between any two sets of 3D feature information.

[0132] If the number of groups in the first preselected value group is 1 and the number of groups in the second preselected value group is 1, then the number of groups in the three-dimensional feature information is also 1, meaning the three-dimensional feature information set includes one set of three-dimensional feature information. Otherwise, if the number of groups in the first preselected value group is greater than 1, or the number of groups in the second preselected value group is greater than 1, then the number of groups in the three-dimensional feature information is also greater than 1, meaning the three-dimensional feature information set can include at least two sets of three-dimensional feature information.

[0133] In some embodiments, the three-dimensional feature information set may further include the number of groups of three-dimensional feature information and / or the number of points in each group of three-dimensional feature information.

[0134] The number of 3D feature information groups is assumed to be the camera parameters, such as the total number of combinations that can be formed from the first and second pre-selected value groups, or the number of combinations after filtering out some combinations. The filtered combinations can be those whose loss function for the corresponding 3D feature points does not meet preset conditions, such as combinations where the vertical distance or chamfer distance between the corresponding 3D feature points is greater than a preset distance, thus filtering out some combinations with poor 3D features. The number of points in each group of 3D feature information refers to the number of 3D features in that group, which can be understood as the 3D feature points corresponding to the matching pixels in a matching pixel pair. Based on the number of 3D feature information groups and / or the number of points in each group, the information of the 3D feature information set can be determined, which helps improve the accuracy of processing the 3D feature information set.

[0135] (4) The three-dimensional feature point set can be understood as the conversion of three-dimensional feature points in the camera coordinate system to three-dimensional feature points in the perception coordinate system, or it can be understood as the three-dimensional feature points resulting from the fusion of two-dimensional feature pair information and perception data. That is, it can be three-dimensional feature points in the camera coordinate system, or it can be three-dimensional feature points in the perception coordinate system, or even three-dimensional feature points in the world coordinate system, without limitation here. In some embodiments, the three-dimensional feature point set includes three-dimensional feature points in a set of three-dimensional feature information in the three-dimensional feature information set. That is, the three-dimensional feature point set is the set corresponding to the three-dimensional feature points in the set of three-dimensional feature information.

[0136] This application does not limit the method for determining the three-dimensional feature point set. The vertical distance between two matching pixels in a matching pixel pair corresponding to a three-dimensional feature point can be calculated based on a loss function, as shown by the vertical distance d in Figure 5. Thus, the three-dimensional feature point set can be determined based on the vertical distance. For example, the three-dimensional feature points in the set of three-dimensional characteristic information with the largest minimum vertical distance constitute the three-dimensional feature point set. The loss function can be Mahalanobis distance loss, Euclidean distance loss, Chebyshev discrepancy loss, Hamming distance loss, etc., and is not limited here.

[0137] Optionally, based on the sensing data, the 3D feature points in each group of 3D feature information in the 3D feature information set can be fused to obtain the feature points of each group of 3D feature information transformed from the first camera coordinate system to the sensing coordinate system; then, the chamfer distance of the feature points before and after fusion can be determined, and the 3D feature point set can be determined based on the chamfer distance. For example, the 3D feature points in the group of 3D feature information with the largest minimum chamfer distance are the 3D feature point set.

[0138] Here, the sensing coordinate system can be understood as the coordinate system of the acquisition device when acquiring sensing data. Based on the sensing data, the method for fusing 3D feature points in each group of 3D feature information in the 3D feature information set can be exemplified using a group of 3D feature information, such as the i-th group of 3D feature information. For example, the sensing data and the i-th group of 3D feature information are matched to obtain the transformation relationship between the first camera coordinate system and the sensing coordinate system, and / or the relative scaling ratio s of the 3D feature points in the sensing data and the i-th group of 3D feature information; based on this transformation relationship and / or s, the 3D feature points of the i-th group of 3D feature information transformed from the first camera coordinate system to the sensing coordinate system are determined.

[0139] In this embodiment, the transformation relationship between the first camera coordinate system and the perception coordinate system can be represented by a transformation matrix between the perception coordinate system of the perception data and the first camera coordinate system. This transformation matrix can be called the third transformation matrix, such as M mentioned above, and is used to indicate the transformation relationship of feature points from the first camera coordinate system to the perception coordinate system. M and / or s can be obtained based on the maximum likelihood function, which can be the correlation function of Pycpd, such as L(M,s|X,Y). Where X is the perception data and Y is the i-th group of three-dimensional feature information.

[0140] The method for determining the three-dimensional feature points of the i-th group of three-dimensional feature information from the first camera coordinate system to the perception coordinate system based on the transformation relationship and s can be based on the formula (5) shown below or satisfy formula (5), X′=sMY (5)

[0141] Where X′ represents the three-dimensional feature point of the i-th group of three-dimensional feature information transformed from the first camera coordinate system to the perception coordinate system.

[0142] The chamfer distance can be understood as the loss function of the feature points before and after fusion. The chamfer distance can be referred to as formula (6) shown below or satisfy formula (6).

[0143] Thus, the chamfer distance between the three-dimensional feature points can be calculated according to formula (6), and a set of three-dimensional feature information can be selected from the three-dimensional feature information set as the three-dimensional feature point set based on the chamfer distance. It should be noted that the three-dimensional feature point set can belong to the optimal or locally optimal three-dimensional feature information in the three-dimensional feature information set. That is, the set of three-dimensional feature information to which the three-dimensional feature point set belongs has the optimal three-dimensional feature points, and the number of optimal three-dimensional feature points is the largest or the range is the widest. The three-dimensional feature points in other sets of three-dimensional feature information outside the three-dimensional feature point set may also be the optimal three-dimensional feature points, but their number is less than the number of optimal three-dimensional feature points in the three-dimensional feature point set, or their coverage is less than the coverage of the optimal three-dimensional feature points in the three-dimensional feature point set. Among them, the optimal three-dimensional feature point can be the three-dimensional feature point corresponding to the minimum loss function determined by the aforementioned vertical distance or chamfer distance.

[0144] (5) The registration result is used to indicate the transformation relationship between different coordinate systems, which may include the transformation matrix between the pixel coordinate system and the camera coordinate system, the transformation matrix between the camera coordinate system and the perception coordinate system, etc. Among them, the transformation relationship between different coordinate systems can be the transformation relationship obtained by registration (or alignment or optimization) after data fusion, and the transformation matrix between different coordinate systems can be understood as the transformation matrix after registration (or alignment or optimization).

[0145] This application does not limit the method for obtaining the registration result; it can be obtained based on a 3D feature information set and perceptual data. Specifically, it can be obtained by optimizing the transformation matrix between various coordinate systems based on the 3D feature point set obtained from the 3D feature information set and perceptual data. In the embodiments of this application, the first pixel coordinate system and the first camera coordinate system of the first image data, and the second pixel coordinate system of the second image data are used as examples. Thus, based on the 3D feature information set and the 3D feature point set obtained from the perception data, the first transformation matrix between the first pixel coordinate system and the first camera coordinate system, the second transformation matrix between the second pixel coordinate system and the second camera coordinate system, and the third transformation matrix between the first camera coordinate system and the perception coordinate system can be registered (or aligned or optimized) to obtain the registered (or aligned or optimized) first transformation matrix, second transformation matrix, and third transformation matrix. This allows us to obtain the registered (or aligned or optimized) transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the registered (or aligned or optimized) transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, and the registered (or aligned or optimized) transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data.

[0146] Optionally, the registration result can be the optimal first transformation matrix, second transformation matrix, and third transformation matrix. Alternatively, the registration result can indicate the first and second transformation matrices through indices in the optimal combinations corresponding to at least one first preselected value group and at least one second preselected value group. Indicating the first and second transformation matrices through indices can save signaling.

[0147] Optionally, under resource constraints, a portion of the 3D feature information can be sent first. If the 3D feature points in this portion meet the preset performance requirements, the remaining 3D feature information does not need to be sent. Otherwise, the remaining 3D feature information, or a portion of the remaining 3D feature information, can continue to be sent. In this way, sending 3D feature information of matching pixel pairs in an additive manner helps improve performance.

[0148] (6) Configuration information is used to indicate data registration parameters to the object of data transmission. Optionally, configuration information is sent to the device matching two image data to indicate at least one of the following: the number of matching pixel pairs, at least one first pre-selected value group, and at least one second pre-selected value group. This helps to improve the accuracy of data processing.

[0149] Furthermore, when the number of image data points is greater than two, the number of image data points can be indicated through configuration information. The image data includes at least the first image data and the second image data. This improves the accuracy of data processing.

[0150] Optionally, the configuration information may also include information about the three-dimensional feature information set, such as the number of groups of three-dimensional feature information and / or the number of points in each group of three-dimensional feature information.

[0151] This application proposes a data transmission method capable of transmitting data or feature information obtained from data fusion, which facilitates data fusion. Furthermore, transmitting the feature information can protect user privacy and reduce data transmission volume compared to transmitting the data itself.

[0152] Please refer to Figure 6, which is a flowchart illustrating a data transmission method according to an embodiment of this application. The first and second devices in Figure 6 can be described with reference to Figures 1A and 1B. The method includes, but is not limited to, the following steps S601 and S602, wherein:

[0153] S601, the first device receives two-dimensional feature pair information, which is used to indicate matching pixel pairs between the first image data and the second image data.

[0154] In some embodiments, the second device sends two-dimensional feature pair information to the first device, causing the first device to receive the two-dimensional feature pair information from the second device. When the second device is not a data acquisition device, the two-dimensional feature pair information sent by the second device to the first device can be determined by the second device or by the data acquisition device. In other embodiments, the data acquisition device can send the two-dimensional feature pair information to the first device. Alternatively, step S601 can be replaced by the first device acquiring the two-dimensional feature pair information. This application does not limit the device used to acquire the two-dimensional feature pair information.

[0155] The two-dimensional feature pair information can be referred to the foregoing and is not limited thereto. In some embodiments, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data. In this way, the two-dimensional features of the image data, such as the matching pixel pairs between the image data and other image data, can be located according to the image size of the image data, which is beneficial to improving the accuracy and convenience of data fusion.

[0156] S602, the first device generates a three-dimensional feature information set based on two-dimensional feature pair information. The three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated by matching pixel pairs.

[0157] The three-dimensional feature information set can be referred to the foregoing and is not limited here. In some embodiments, the three-dimensional feature information set further includes: the number of groups of three-dimensional feature information and / or the number of feature points in each group of three-dimensional feature information. This helps to improve the accuracy of processing the three-dimensional feature information set.

[0158] The method shown in Figure 6 can transmit three-dimensional feature information obtained by fusing at least two image data, which is beneficial for further data fusion.

[0159] In some embodiments, the method may further include: synchronizing configuration information between the first device and the second device. The configuration information includes at least one of the following: the number of matched pixel pairs, at least one first pre-selected value group, and at least one second pre-selected value group. The at least one first pre-selected value group indicates pre-selected values ​​for unknown parameters of the first image data, and the at least one second pre-selected value group indicates pre-selected values ​​for unknown parameters of the second image data. Thus, feature information of the data can be obtained based on the synchronized configuration information, improving the accuracy of feature extraction.

[0160] In other embodiments, the first device can also synchronize configuration information with other devices. These other devices can be devices other than the second device, such as acquisition devices for acquiring other image data, acquisition devices for acquiring sensory data, devices for obtaining two-dimensional feature pairs, devices for generating three-dimensional feature information sets, devices for generating registration results and / or three-dimensional feature points, etc. This facilitates improved accuracy in feature extraction.

[0161] In some embodiments, the method may further include: the first device generating the three-dimensional feature information set based on the two-dimensional feature pair information, the at least one first pre-selected value group, and the at least one second pre-selected value group. Thus, the three-dimensional feature information set can be determined by using pre-selected values ​​from image data to assist the two-dimensional feature pair information. For example, the three-dimensional feature information corresponding to different first and second pre-selected value groups in the three-dimensional feature information set can improve the accuracy of obtaining the three-dimensional feature information set.

[0162] After the first and second devices synchronize their configuration information, the second device can generate two-dimensional feature pair information based on the number of matched pixel pairs, and the first device can generate a three-dimensional feature information set based on the two-dimensional feature pair information, at least one first pre-selected value group, and at least one second pre-selected value group, which can improve the accuracy of obtaining feature information.

[0163] In some embodiments, the configuration information further includes the number of image data, which includes at least the first image data and the second image data. If other image data besides the first and second image data are not included, the configuration information may not include the number of image data, thus allowing the acquisition of two-dimensional feature pair information between the two image data after receiving the first and second image data. If the configuration information includes the number of image data, the number of pairwise combinations of the image data can be determined based on this number, thereby determining the number of two-dimensional feature pairs and improving the accuracy of feature information acquisition.

[0164] If other image data exists, it can include W blocks of two-dimensional feature pair information, used to indicate matching pixel pairs between two images in the other image data, the first image data, and the second image data. Optionally, the two-dimensional feature pair information can also be used to indicate the image size of the other image data. In this way, the two-dimensional features of the image can be located based on the image size of the image data, which helps to improve the accuracy and convenience of data fusion.

[0165] In some embodiments, the method may further include: a first device receiving sensing data from a second device; the first device acquiring a registration result and / or a set of three-dimensional feature points based on a three-dimensional feature information set and the sensing data. Correspondingly, the second device sends the sensing data to the first device. The sensing data may not originate from the second device, but rather from a data acquisition device that acquires the sensing data, or be forwarded by a network device, etc., and is not limited thereto. Thus, the three-dimensional feature information obtained by fusing at least two image data sets can be further fused with the sensing data, which is beneficial for improving the representational capability of the fused data and improving the accuracy of image reconstruction.

[0166] In some embodiments, the method may further include: a first device fusing the three-dimensional feature point set and the perceptual data. Similarly, a second device may also fuse the three-dimensional feature point set and the perceptual data. This can improve the representational capability of the fused perceptual data, thereby improving the accuracy of image reconstruction.

[0167] For example, referring to Figure 7A, whether it is sparse or dense perceptual data, the fusion adds multiple feature points compared to the original data, thereby improving the representational ability of the fused data and helping to improve the accuracy of image reconstruction.

[0168] Optionally, the method may further include: a first device or a second device fusing the first image data and / or the second image data based on a three-dimensional feature point set and / or registration results. This can improve the representational capability of the fused image data and help improve the accuracy of image reconstruction.

[0169] For example, referring to Figure 7B, after the first image data and the second image data are fused, multiple feature points are added compared to before fusion, which can improve the representational ability of the fused data and help improve the accuracy of image reconstruction.

[0170] In some embodiments, the method may further include: a first device sending a registration result and / or a set of three-dimensional feature points to a second device. Correspondingly, the second device receives the registration result and / or the set of three-dimensional feature points from the first device. Optionally, the first device may also send the registration result and / or the set of three-dimensional feature points to other devices (such as an acquisition device). In this way, it is unnecessary to acquire data from different modalities for each data fusion, which is beneficial for improving the efficiency of subsequent data fusion.

[0171] In some embodiments, the three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in the three-dimensional feature information set. Thus, the registration results can be registered (or aligned or optimized) based on the three-dimensional feature points in this set of three-dimensional feature information, which helps to improve the representational capability of the fused data and the accuracy of image reconstruction.

[0172] In some embodiments, the registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, and the transformation relationship between the perceptual coordinate system and the first camera coordinate system of the perceptual data. Thus, the camera coordinates corresponding to feature points in the image data or perceptual data can be determined based on the registration result. For example, the camera coordinates corresponding to pixels in the image data can be determined based on the transformation relationship between the first pixel coordinate system and the first camera coordinate system, and the image data acquired by the acquisition device corresponding to the first image data. Similarly, the camera coordinates corresponding to pixels in the image data can be determined based on the transformation relationship between the second pixel coordinate system and the second camera coordinate system, and the image data acquired by the acquisition device corresponding to the second image data. The camera coordinates corresponding to feature points in the perceptual data can also be determined based on the transformation relationship between the perceptual coordinate system and the first camera coordinate system, and the perceptual data acquired by the acquisition device corresponding to the perceptual data.

[0173] Optionally, the transformation relationship between the first pixel coordinate system and the first camera coordinate system is represented by a first transformation matrix, the transformation relationship between the second pixel coordinate system and the second camera coordinate system is represented by a second transformation matrix, and the transformation relationship between the perception coordinate system and the first camera coordinate system is represented by a third transformation matrix.

[0174] For example, simulation tests can be performed with the first transformation matrix being K1 and the third transformation matrix being M1. Registration results can be obtained without changing the coordinate system. In this registration result, the first transformation matrix can be K2, and the third transformation matrix can be M2, where:

[0175] Thus, based on K1 and K2, the transformation matrix K between the pixel coordinate system and the camera coordinate system can be determined. u and f v The normalized mean square error (NMSE) is 9.7 * 10⁻⁶. -6 Based on M1 and M2, the NMSE of the rotation matrix R of the transformation matrix M from the camera coordinate system to the sensor coordinate system can be determined to be 4.3 * 10^- ... -3The NMSE of the translation vector T is 0.2. By implementing the method provided in this application for data fusion and obtaining the registration results of image data, the accuracy of obtaining the transformation matrix from the pixel coordinate system to the camera coordinate system of the image data is improved.

[0176] Optionally, the registration result is also used to indicate at least one of the following: the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.

[0177] In this embodiment, the transformation relationship between the first camera coordinate system and the second camera coordinate system can be represented by a camera transformation matrix, which can be based on the first camera coordinate system. The transformation relationship between the first pixel coordinate system and the perception coordinate system can be obtained by multiplying the first transformation matrix and the third transformation matrix, for example, by multiplying the first transformation matrix and the third transformation matrix. The transformation relationship between the second pixel coordinate system and the perception coordinate system can be obtained by multiplying the second transformation matrix and the third transformation matrix, for example, by multiplying the second transformation matrix and the third transformation matrix. The transformation relationship between the perception coordinate system and the second camera coordinate system refers to the transformation relationship of feature points from the second camera coordinate system to the perception coordinate system.

[0178] It should be noted that the registration results can also be used to indicate the transformation relationship between other coordinate systems, such as the transformation relationship between the first pixel coordinate system and the world coordinate system, the transformation relationship between the second pixel coordinate system and the world coordinate system, etc., without limitation here. In this way, data fusion can be performed based on the transformation relationship between at least two coordinate systems.

[0179] In some embodiments, the registration result is obtained based on periodic or triggered information. That is, the acquisition device can be triggered to acquire image data and sensing data when the periodic time arrives, so that the first device and the second device can obtain the registration result according to the method provided in this application. Alternatively, it can be triggered by an event, such as when the sensing coordinate system of the acquisition device changes or the position of the acquisition device changes, so that the first device and the second device can obtain the registration result according to the method provided in this application.

[0180] Optionally, if the coordinate system corresponding to the registration result remains unchanged, data fusion can be performed directly based on the registration result. If the coordinate system corresponding to the registration result changes, a new registration result can be obtained, and data fusion can be performed based on the newly obtained registration result. Obtaining the registration result periodically or triggered by specific events helps improve the accuracy of data fusion.

[0181] Figure 6 illustrates the generation of a 3D feature information set, registration result, and 3D feature point set using the first device. Optionally, the 2D feature pair information, 3D feature information set, registration result, and 3D feature point set can be generated (or acquired) by one or more processing devices. The processing device can be at least one of the first device, second device, acquisition device, etc. The processing device can be a terminal device or a network device.

[0182] The solutions described in this application are described below for different application scenarios and different diagrams, wherein:

[0183] First application scenario: The data acquisition device, the first device, and the second device are the same processing device.

[0184] In this application scenario, the processing device acquires at least two image data; obtains W blocks of two-dimensional feature pair information based on the at least two image data, each block of two-dimensional feature pair information is used to indicate the matching pixel pair between the two image data; and obtains a three-dimensional feature information set based on the W blocks of two-dimensional feature pair information.

[0185] Optionally, the processing device can also collect sensing data and obtain registration results and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the sensing data.

[0186] Furthermore, the processing device can also fuse the subsequently acquired sensory data based on the registration results and / or the three-dimensional feature point set.

[0187] In the first application scenario, a processing device is used to acquire image data and sensor data, and to process different sensor data to obtain two-dimensional feature pair information, three-dimensional feature information set, registration result and / or three-dimensional feature point set of these sensor data.

[0188] Second application scenario: The data acquisition device is a device other than the first device and the second device, and the processing device is either the first device or the second device. As shown in Figure 1B, the data acquisition device 30 is a device other than the first device 10 and the second device 20, and the processing device can be either the first device 10 or the second device 20.

[0189] The third application scenario: The data acquisition device is at least one of the first device and the second device, and the processing device is a device other than the data acquisition device. As shown in Figure 1A, when the data acquisition device is the first device 10, the processing device is the second device 20; or when the data acquisition device is the second device 20, the processing device is the first device 10.

[0190] In the second and third application scenarios, the acquisition device can be one or more terminal devices or acquisition modules within terminal devices. As shown in Figure 8A, the processing device (first device or second device) receives at least two image data from at least one acquisition device; the processing device acquires W blocks of two-dimensional feature pair information based on the at least two image data, each block of two-dimensional feature pair information being used to indicate matching pixel pairs between the two image data; the processing device acquires a three-dimensional feature information set based on the W blocks of two-dimensional feature pair information.

[0191] Optionally, the processing device may also receive sensing data from the acquisition device; the processing device may acquire registration results and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the sensing data; and the processing device may send the registration results and / or the set of three-dimensional feature points to the acquisition device.

[0192] In the second and third application scenarios, the acquisition device sends image data and perception data to the processing device, and the processing device separately acquires two-dimensional feature pair information, three-dimensional feature information set, registration result and / or three-dimensional feature point set.

[0193] Fourth application scenario: Regardless of whether the data acquisition device is the first device or the second device, the processing device is either the first device or the second device. The fourth application scenario can include the following six methods, among which:

[0194] Method 1: The second device acquires two-dimensional feature pair information, and the first device acquires a three-dimensional feature information set, registration result, and / or three-dimensional feature point set, as described in Figure 6.

[0195] Method 2: The second device acquires two-dimensional feature pair information, the first device acquires a three-dimensional feature information set, and the second device acquires the registration result and / or a three-dimensional feature point set;

[0196] Method 3: The first device acquires two-dimensional feature pair information, and the second device acquires a three-dimensional feature information set, registration results, and / or a three-dimensional feature point set;

[0197] Method 4: The first device acquires two-dimensional feature pair information, the second device acquires a three-dimensional feature information set, and the first device acquires the registration result and / or a three-dimensional feature point set;

[0198] Method 5: The first device acquires two-dimensional feature pair information and a three-dimensional feature information set, and the second device acquires the registration result and / or a three-dimensional feature point set;

[0199] Method Six: The second device acquires two-dimensional feature pair information and a three-dimensional feature information set, while the first device acquires the registration result and / or a three-dimensional feature point set.

[0200] In the fourth application scenario, the first device and the second device respectively acquire at least one of the following: two-dimensional feature pair information, three-dimensional feature information set, registration result, and three-dimensional feature point set. In the above six methods, image data and / or perceived data can be acquired by the first device, the second device, or other devices. Specifically, Method 1 can be referenced to Figure 8B, Method 2 to Figure 8C, and Method 6 to Figure 8D. The first device in Method 3 can be the second device in Figure 8B, and the second device in Method 3 can be the first device in Figure 8B. The first device in Method 4 can be the second device in Figure 8C, and the second device in Method 4 can be the first device in Figure 8C. The first device in Method 5 can be the second device in Figure 8D, and the second device in Method 5 can be the first device in Figure 8D.

[0201] As shown in Figures 8B to 8D, image data and sensory data can be acquired by a data acquisition device, or by a first device or a second device. When the number of image data is equal to 2, M can be 1. When the number of image data is greater than 2, M can be greater than 1. Different devices for acquiring the first image data, the second image data, and the sensory data can include the following three cases, wherein:

[0202] Scenario 1: UE#1 collects the first image data, UE#2 collects the second image data, and UE#3 collects the perceived data.

[0203] Scenario 2: UE#1 collects the first image data and the second image data, while UE#3 collects the perceived data.

[0204] Scenario 3: UE#1 collects the first image data and perception data, and UE#2 collects the second image data.

[0205] The above three examples are applicable to the six methods described above, and at least one of the first and second devices in the six methods can be one of UE#1, UE#2, and UE#3. That is, the first device can be one of UE#1, UE#2, and UE#3, and the second device can be any device among UE#1, UE#2, and UE#3 other than the first device, or any device other than UE#1, UE#2, and UE#3. Alternatively, the second device can be one of UE#1, UE#2, and UE#3, and the first device can be any device among UE#1, UE#2, and UE#3 other than the second device, or any device other than UE#1, UE#2, and UE#3.

[0206] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.

[0207] Please refer to Figure 9, which is a schematic diagram of a communication device provided in an embodiment of this application. The communication device may include a transceiver unit 901 and a processing unit 902. The transceiver unit 901 may be a unit with signal input (receiving) or output (transmitting) functions, used for transmitting signals to other devices or other units within a device. The processing unit 902 may be a unit with processing functions, such as one or more processors, used for executing instructions (or code or programs). The communication device may be a first device or a second device. The first device and the second device may be a terminal device or a network device.

[0208] In the first embodiment, the communication device can be a first device, wherein:

[0209] The transceiver unit 901 is used to receive two-dimensional feature pair information from the second device, the two-dimensional feature pair information being used to indicate matching pixel pairs between the first image data and the second image data;

[0210] The processing unit 902 is used to generate a three-dimensional feature information set based on the two-dimensional feature pair information. The three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated by the matching pixel pair.

[0211] In some embodiments, the transceiver unit 901 is further configured to receive sensing data from the second device;

[0212] The processing unit 902 is also used to obtain registration results and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the perception data.

[0213] In some embodiments, the processing unit 902 is further configured to fuse the three-dimensional feature point set and the perception data.

[0214] In some embodiments, the transceiver unit 901 is further configured to send the registration result and / or the three-dimensional feature point set to the second device.

[0215] In some embodiments, the registration result is used to indicate at least one of the following transformation relationships: between the first pixel coordinate system and the first camera coordinate system of the first image data, between the second pixel coordinate system and the second camera coordinate system of the second image data, between the perceptual coordinate system and the first camera coordinate system of the perceptual data, between the first camera coordinate system and the second camera coordinate system, between the first pixel coordinate system and the perceptual coordinate system, between the second pixel coordinate system and the perceptual coordinate system, and between the perceptual coordinate system and the second camera coordinate system.

[0216] In some embodiments, the three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in the three-dimensional feature information set.

[0217] In some embodiments, the transceiver unit 901 is further configured to synchronize configuration information with the second device; wherein the configuration information includes at least one of the following: the number of matching pixel pairs, at least one first preselected value group, and at least one second preselected value group, wherein the at least one first preselected value group is used to indicate the preselected value of the unknown parameter of the first image data, and the at least one second preselected value group is used to indicate the preselected value of the unknown parameter of the second image data.

[0218] In some embodiments, the processing unit 902 is further configured to generate the three-dimensional feature information set based on the two-dimensional feature pair information, the at least one first preselected value group, and the at least one second preselected value group.

[0219] In some embodiments, the configuration information further includes the number of image data, wherein the image data includes at least the first image data and the second image data.

[0220] In some embodiments, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.

[0221] In some embodiments, the three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.

[0222] In some embodiments, the registration result is obtained based on periodic or triggered information.

[0223] In the first embodiment, the communication device may be a second device, wherein:

[0224] Processing unit 902 is used to acquire two-dimensional feature pair information, which is used to indicate matching pixel pairs between first image data and second image data;

[0225] The transceiver unit 901 is used to send the two-dimensional feature pair information to the first device.

[0226] In some embodiments, the transceiver unit 901 is further configured to send sensing data to the first device.

[0227] In some embodiments, the transceiver unit 901 is further configured to receive the registration result and / or three-dimensional feature point set of the first device.

[0228] In some embodiments, the registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, the transformation relationship between the perceptual coordinate system and the first camera coordinate system of the perceptual data, the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perceptual coordinate system, the transformation relationship between the second pixel coordinate system and the perceptual coordinate system, and the transformation relationship between the perceptual coordinate system and the second camera coordinate system.

[0229] In some embodiments, the three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in a three-dimensional feature information set; wherein, the three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated by the matching pixel pair.

[0230] In some embodiments, the transceiver unit 901 is further configured to synchronize configuration information with the first device; wherein the configuration information includes at least one of the following: the number of matching pixel pairs, at least one first preselected value group and at least one second preselected value group, wherein the at least one first preselected value group is used to indicate the preselected value of the unknown parameter of the first image data, and the at least one second preselected value group is used to indicate the preselected value of the unknown parameter of the second image data.

[0231] In some embodiments, the configuration information further includes the number of image data, wherein the image data includes at least the first image data and the second image data.

[0232] In some embodiments, the two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.

[0233] In some embodiments, the three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.

[0234] In some embodiments, the registration result is obtained based on periodic or triggered information.

[0235] The implementation of the above-mentioned transceiver unit 901 and processing unit 902 can be referred to the relevant description of the method embodiment shown in FIG6, which will not be repeated here.

[0236] Please refer to Figure 10, which is a schematic diagram of another communication device provided in an embodiment of this application. As shown in Figure 10, the communication device may include a processor 111 and a storage medium 112. The processor 111 may also be called a processing unit, which can implement certain control functions. The storage medium 112 may also be called a storage unit or a memory. Instructions 114 are stored on the storage medium 112. The instructions 114 can be executed on the processor 111, causing the communication device to perform any of the methods described in Figure 6 of this application embodiment.

[0237] Optionally, the processor 111 may include instructions 113 that can be executed on the processor 111 to cause the communication device to perform any of the methods described in FIG6 in the embodiments of this application.

[0238] The communication device can be a first device or a second device. The first device and the second device can be terminal devices or network devices, used to implement the methods described in the method embodiments. However, the scope of the devices described in this application is not limited thereto; the communication device can be a standalone device or part of a larger device. For example, the communication device can be:

[0239] (1) An independent integrated circuit IC, or chip, or chip system or subsystem;

[0240] (2) A collection of one or more ICs, wherein the collection of ICs may optionally include a storage component for storing data and / or instructions;

[0241] (3) ASIC, such as modems;

[0242] (4) Modules that can be embedded in other devices.

[0243] Please refer to Figure 11, which is a schematic diagram of a terminal device provided in an embodiment of this application. For ease of explanation, Figure 11 only shows the main components of the terminal device. As shown in Figure 11, the terminal device includes a processor, a memory, a control circuit, an antenna, and input / output devices. The processor is mainly used to process communication protocols and communication data, control the entire terminal device, execute software programs, and process the data of the software programs. The memory is mainly used to store software programs and data. The radio frequency circuit is mainly used for the conversion between baseband signals and radio frequency signals and the processing of radio frequency signals. The antenna is mainly used for transmitting and receiving radio frequency signals in the form of electromagnetic waves. Input / output devices, such as touch screens, displays, and keyboards, are mainly used to receive user input data and output data to the user.

[0244] When the terminal device is powered on, the processor can read the software program from the storage unit, parse and execute the instructions of the software program, and process the data of the software program. When data needs to be transmitted wirelessly, the processor performs baseband processing on the data to be transmitted and outputs the baseband signal to the radio frequency (RF) circuit. The RF circuit processes the baseband signal to obtain the RF signal and transmits the RF signal outward in the form of electromagnetic waves through the antenna. When data is sent to the terminal device, the RF circuit receives the RF signal through the antenna. This RF signal is further converted into a baseband signal and output to the processor. The processor converts the baseband signal back into data and processes the data.

[0245] For ease of explanation, Figure 11 shows only one memory and processor. In actual terminal devices, multiple processors and memories may exist. Memory may also be referred to as storage medium or storage device, etc., and the embodiments of this application do not limit this.

[0246] In one embodiment, the antenna is used to perform the operations performed by the transceiver unit 901 in the above embodiment. The processor is used to perform the operations performed by the processing unit 902 in the above embodiment.

[0247] This application also provides a computer-readable storage medium storing instructions or computer programs that, when executed, can implement the relevant steps in the data transmission method provided in the above-described method embodiments.

[0248] This application also provides a computer program product, which includes instructions or a computer program that, when executed, causes one or more steps in any of the above-described data transmission methods to be performed. If the constituent modules of the aforementioned devices are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0249] The instructions or computer program can be executed by a computer or processor, without limitation.

[0250] This application provides a chip or chip system including at least one processor for calling and running instructions stored in a memory, causing a communication device with the chip installed to perform any of the above methods or to execute the steps of the processing unit 902.

[0251] This application embodiment also provides another chip, including a processor and a memory, wherein the processor is used to call and run instructions stored in the memory, causing a communication device with the chip installed to perform any of the above methods, or to perform the steps of the processing unit 902.

[0252] This application embodiment also provides another chip, including: an input interface, an output interface, and a processing circuit. The input interface, the output interface, and the processing circuit are connected via internal connection paths. The processing circuit is used to execute any of the methods described above. Optionally, the chip also includes a memory. The input interface, the output interface, the processor, and the memory are connected via internal connection paths. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute any of the methods described above, or to execute the steps of processing unit 902.

[0253] This application also provides another chip system, including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a line. The at least one processor is used to run a computer program or instructions to perform any of the methods described above, or to execute the steps of processing unit 902. This chip system may be composed of chips, or may include chips and other discrete devices.

[0254] This application also provides a communication system, which includes a first device and a second device, as detailed in Figure 6. The first device and the second device in this application can be a terminal device or a network device.

[0255] Optionally, the memory mentioned in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be a hard disk drive (HDD), a solid-state drive (SSD), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be RAM, which is used as an external cache. Memory is any other medium capable of carrying or storing desired program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in the embodiments of this application can also be a circuit or any other device capable of implementing a storage function for storing program instructions and / or data.

[0256] Optionally, the processor mentioned in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or any conventional processor.

[0257] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) is integrated into the processor.

[0258] It should be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0259] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments provided herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0260] In the several embodiments provided in this application, the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0261] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0262] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0263] The steps in the methods of this application can be adjusted, combined, or deleted according to actual needs. Each step in each embodiment can be partially performed (for example, the terminal device may not perform the steps performed by the terminal device in the above embodiments). The execution order of different steps can be changed. The embodiments described herein can be combined with other embodiments, different embodiments can be combined with each other, and different steps of different embodiments herein can be combined.

[0264] The modules / units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0265] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments.

[0266] In this application, it may refer to a communication protocol or specification, such as the 3GPP communication protocol.

[0267] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0268] In the embodiments of this application, "including" can refer to a relationship of inclusion or an equality relationship. For example, A includes B, which could mean that A includes other content besides B, or that A and B are the same content.

[0269] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0270] In the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

Claims

1. A data transmission method, characterized in that, include: The first device receives two-dimensional feature pair information from the second device, the two-dimensional feature pair information being used to indicate matching pixel pairs between the first image data and the second image data; The first device generates a three-dimensional feature information set based on the two-dimensional feature pair information. The three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated by the matching pixel pairs.

2. The method according to claim 1, characterized in that, Also includes: The first device receives sensing data from the second device; The first device acquires registration results and / or a set of three-dimensional feature points based on the three-dimensional feature information set and the sensing data.

3. The method according to claim 2, characterized in that, Also includes: The first device fuses the three-dimensional feature point set and the perceived data.

4. The method according to claim 2 or 3, characterized in that, Also includes: The first device sends the registration result and / or the three-dimensional feature point set to the second device.

5. The method according to any one of claims 2 to 4, characterized in that, The registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, the transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data, the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.

6. The method according to any one of claims 2 to 5, characterized in that, The three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in the three-dimensional feature information set.

7. The method according to any one of claims 1 to 6, characterized in that, Also includes: The first device and the second device are configured with synchronous information; wherein the configuration information includes at least one of the following: the number of matching pixel pairs, at least one first preselected value group, and at least one second preselected value group, wherein the at least one first preselected value group is used to indicate the preselected value of the unknown parameter of the first image data, and the at least one second preselected value group is used to indicate the preselected value of the unknown parameter of the second image data.

8. The method according to claim 7, characterized in that, Also includes: The first device generates the three-dimensional feature information set based on the two-dimensional feature pair information, the at least one first pre-selected value group, and the at least one second pre-selected value group.

9. The method according to claim 7 or 8, characterized in that, The configuration information also includes the number of image data, which includes at least the first image data and the second image data.

10. The method according to any one of claims 1 to 9, characterized in that, The two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.

11. The method according to any one of claims 1 to 10, characterized in that, The three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.

12. The method according to any one of claims 2 to 11, characterized in that, The registration results are obtained based on periodic or triggered information.

13. A data transmission method, characterized in that, include: The second device acquires two-dimensional feature pair information, which is used to indicate matching pixel pairs between the first image data and the second image data. The second device sends the two-dimensional feature pair information to the first device.

14. The method according to claim 13, characterized in that, Also includes: The second device sends sensing data to the first device.

15. The method according to claim 14, characterized in that, Also includes: The second device receives the registration result and / or three-dimensional feature point set from the first device.

16. The method according to claim 15, characterized in that, The registration result is used to indicate at least one of the following: the transformation relationship between the first pixel coordinate system and the first camera coordinate system of the first image data, the transformation relationship between the second pixel coordinate system and the second camera coordinate system of the second image data, the transformation relationship between the perception coordinate system and the first camera coordinate system of the perception data, the transformation relationship between the first camera coordinate system and the second camera coordinate system, the transformation relationship between the first pixel coordinate system and the perception coordinate system, the transformation relationship between the second pixel coordinate system and the perception coordinate system, and the transformation relationship between the perception coordinate system and the second camera coordinate system.

17. The method according to claim 15 or 16, characterized in that, The three-dimensional feature point set includes three-dimensional feature points from a set of three-dimensional feature information in the three-dimensional feature information set; The three-dimensional feature information set includes at least one set of three-dimensional feature information, and each set of three-dimensional feature information includes three-dimensional feature points generated by the matching pixel pairs.

18. The method according to any one of claims 13 to 17, characterized in that, Also includes: The second device and the first device are synchronized with configuration information; wherein the configuration information includes at least one of the following: the number of matching pixel pairs, at least one first preselected value group and at least one second preselected value group, wherein the at least one first preselected value group is used to indicate the preselected value of the unknown parameter of the first image data, and the at least one second preselected value group is used to indicate the preselected value of the unknown parameter of the second image data.

19. The method according to claim 18, characterized in that, The configuration information also includes the number of image data, which includes at least the first image data and the second image data.

20. The method according to any one of claims 13 to 19, characterized in that, The two-dimensional feature pair information is also used to indicate the image size of the first image data and / or the image size of the second image data.

21. The method according to any one of claims 17 to 20, characterized in that, The three-dimensional feature information set further includes: the number of groups of the three-dimensional feature information and / or the number of feature points in each group of the three-dimensional feature information.

22. The method according to any one of claims 15 to 21, characterized in that, The registration results are obtained based on periodic or triggered information.

23. A communication device, characterized in that, Includes units for performing the method as described in any one of claims 1 to 22.

24. A communication device, characterized in that, The communication device includes a processor and a storage medium storing instructions that, when executed by the processor, cause the method according to any one of claims 1 to 22 to be performed.

25. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, cause the method of any one of claims 1 to 22 to be implemented.

26. A computer program product, characterized in that, The computer program product includes instructions that, when executed, cause the method of any one of claims 1 to 22 to be implemented.

27. A chip or chip system, characterized in that, Includes a processor for retrieving and executing instructions stored in a memory, causing a communication device with a chip mounted to perform the method as described in any one of claims 1 to 22.

28. A communication system, characterized in that, It includes a first device and a second device, the first device being used to perform the method according to any one of claims 1 to 12, and the second device being used to perform the method according to any one of claims 13 to 22.