Camera calibration

The semantic hierarchical multi-modal matching mechanism addresses the inefficiencies and inaccuracies of traditional camera calibration methods by using semantic hierarchical information matching to establish accurate correspondences between real-world objects and image pixels, achieving efficient and robust camera calibration and localization.

WO2025111810A1PCT designated stage expired Publication Date: 2025-06-05ALCATEL LUCENT SHANGHAI BELL CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/134787
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing camera calibration methods for monocular cameras in factory environments are inaccurate and inefficient, requiring significant manual intervention and often introducing measurement errors. Additionally, traditional methods rely on special markers and can be prone to mismatch and noise, especially when dealing with large areas like a 20m ×20m field.

Method used

A semantic hierarchical multi-modal matching mechanism is proposed for monocular camera calibration and localization. This mechanism uses two-layer semantic hierarchical information matching from objects to points, allowing for accurate and robust correspondence between real-world objects and image pixels. It reduces communication burden by transmitting multi-modal semantic elements rather than large data sets and can select suitable static or moving objects automatically without the need for special markers.

Benefits of technology

The proposed mechanism achieves accurate and efficient camera calibration and localization, improving matching accuracy and robustness to environmental noise. It reduces communication payload significantly and can perform calibration automatically, making it suitable for various scenarios, including factory environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023134787_05062025_PF_FP_ABST
    Figure CN2023134787_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to camera calibration. In an aspect, a device obtains first data including first semantic information of an object and first coordinate information of the object in a physical space, and obtains second data including second semantic information of the object and second coordinate information of the object in an image of the physical space captured by a camera. With the matching between the first and second semantic information, the device determines calibration data of the camera based on the first and second coordinate information. As such, the accurate matching of corresponding points from the real world and image pixels is implemented for camera calibration and localization, and the communication burden is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

CAMERA CALIBRATIONFIELD

[0001] Various example embodiments relate to the field of communication and in particular, to devices, methods, apparatuses, and a computer readable storage medium for camera calibration or localization.BACKGROUND

[0002] In certain scenarios, e.g. a factory environment, accurate object localization is an essential function and plays a very important role in environmental monitoring and safety. Camera localization is a relatively good choice due to camera’s accessibility. For example, a camera localization technique may match 2D point features in the captured image (s) to 3D map features stored in the 3D map as 3D points.SUMMARY

[0003] In general, example embodiments of the present disclosure provide a solution related to camera calibration and localization, especially for monocular camera calibration and localization.

[0004] In a first aspect, there is provided a first device. The first device includes at least one processor and at least one memory storing indications that, when executed by the at least one processor, cause the first device at least to: obtain first data including first semantic information of a first object and first coordinate information of the first object in a physical space; obtain second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera; determine, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; and based on determining that the first object and the second object are the same object, determine calibration data of the camera based on the first coordinate information and the second coordinate information.

[0005] In a second aspect, there is provided a second device. The second device includes at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the second device at least to: obtain sensed data of a first  object in a physical space by sensing the first object; extract, from the sensed data, first semantic information of the first object; extract, from the sensed data, first coordinate information of the first object in the physical space; and transmit, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.

[0006] In a third aspect, there is provided a third device. The third device includes at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the third device at least to: obtain an image of a physical space captured by a camera; extract, from the image, second semantic information of a second object in the physical space; extract, from the image, second coordinate information of the second object in the image; and transmit, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.

[0007] In a fourth aspect, there is provided a method. The method includes obtaining, at a first device, first data including first semantic information of a first object and first coordinate information of the first object in a physical space; obtaining second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera; determining, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; and based on determining that the first object and the second object are the same object, determining calibration data of the camera based on the first coordinate information and the second coordinate information.

[0008] In a fifth aspect, there is provided a method. The method includes obtaining, at a second device, sensed data of a first object in a physical space by sensing the first object; extracting, from the sensed data, first semantic information of the first object; extracting, from the sensed data, first coordinate information of the first object in the physical space; and transmitting, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.

[0009] In a sixth aspect, there is provided a method. The method includes obtaining, at a third device, an image of a physical space captured by a camera; extracting, from the image,  second semantic information of a second object in the physical space; extracting, from the image, second coordinate information of the second object in the image; and transmitting, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.

[0010] In a seventh aspect, there is provided an apparatus. The apparatus includes means for obtaining, at a first device, first data including first semantic information of a first object and first coordinate information of the first object in a physical space; means for obtaining second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera; means for determining, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; and means for based on determining that the first object and the second object are the same object, determining calibration data of the camera based on the first coordinate information and the second coordinate information.

[0011] In an eighth aspect, there is provided an apparatus. The apparatus includes means for obtaining, at a second device, sensed data of a first object in a physical space by sensing the first object; means for extracting, from the sensed data, first semantic information of the first object; means for extracting, from the sensed data, first coordinate information of the first object in the physical space; and means for transmitting, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.

[0012] In a ninth aspect, there is provided an apparatus. The apparatus includes means for obtaining, at a third device, an image of a physical space captured by a camera; means for extracting, from the image, second semantic information of a second object in the physical space; means for extracting, from the image, second coordinate information of the second object in the image; and means for transmitting, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.

[0013] In a tenth aspect, there is provided a non-transitory computer readable medium including program instructions for causing an apparatus to perform at least the method  according to any of the above fourth, fifth, and sixth aspects.

[0014] In an eleventh aspect, there is provided a first device. The first device includes obtaining circuitry configured to obtain first data including first semantic information of a first object and first coordinate information of the first object in a physical space; obtaining circuitry configured to obtain second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera; determining circuitry configured to determine, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; and determining circuitry configured to, based on determining that the first object and the second object are the same object, determine calibration data of the camera based on the first coordinate information and the second coordinate information.

[0015] In a twelfth aspect, there is provided a second device. The second device includes obtaining circuitry configured to obtain sensed data of a first object in a physical space by sensing the first object; extracting circuitry configured to extract, from the sensed data, first semantic information of the first object; extracting circuitry configured to extract, from the sensed data, first coordinate information of the first object in the physical space; and transmitting circuitry configured to transmit, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.

[0016] In a thirteenth aspect, there is provided a third device. The third device includes obtaining circuitry configured to obtain an image of a physical space captured by a camera; extracting circuitry configured to extract, from the image, second semantic information of a second object in the physical space; extracting circuitry configured to extract, from the image, second coordinate information of the second object in the image; and transmitting circuitry configured to transmit, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.

[0017] In a fourteenth aspect, there is provided a computer program including instructions, which, when executed by an apparatus, cause the apparatus at least to perform at least the method according to any one of the above fourth, fifth and sixth aspects.

[0018] It is to be understood that the summary section is not intended to identify key or  essential features of embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of example embodiments of the present disclosure will become easily comprehensible through the following description.BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Some example embodiments will now be described with reference to the accompanying drawings, in which:

[0020] FIG. 1 illustrates an example environment in which embodiments of the present disclosure may be implemented;

[0021] FIG. 2 shows a signaling chart illustrating a process for camera calibration and localization according to some embodiments of the present disclosure;

[0022] FIGS. 3A and 3B show schematic diagrams illustrating examples of 3D semantic data for camera calibration and localization according to some embodiments of the present disclosure;

[0023] FIGS. 4A and 4B show schematic diagrams illustrating examples of 2D semantic data for camera calibration and localization according to some embodiments of the present disclosure;

[0024] FIG. 5 shows a flowchart illustrating a process for camera calibration and localization according to some embodiments of the present disclosure;

[0025] FIG. 6 illustrates a flowchart of a method implemented at a first device according to some embodiments of the present disclosure;

[0026] FIG. 7 illustrates a flowchart of a method implemented at a second device according to some embodiments of the present disclosure;

[0027] FIG. 8 illustrates a flowchart of a method implemented at a third device according to some embodiments of the present disclosure;

[0028] FIG. 9 illustrates a simplified block diagram of a device that is suitable for implementing embodiments of the present disclosure; and

[0029] FIG. 10 illustrates a block diagram of an example computer readable medium in accordance with some embodiments of the present disclosure.

[0030] Throughout the drawings, the same or similar reference numerals represent the same  or similar element.DETAILED DESCRIPTION

[0031] Principles of the present disclosure will now be described with reference to some example embodiments. It is to be understood that these embodiments are described for the purpose of illustration and help those skilled in the art to understand and implement example embodiments of the present disclosure, without suggesting any limitation as to the scope of the disclosure. The disclosure described herein can be implemented in various manners other than the ones described below.

[0032] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skills in the art to which this disclosure belongs.

[0033] References in the present disclosure to “one embodiment, ” “an embodiment, ” “an example embodiment, ” and the like indicate that the embodiment described may include a particular feature, structure, or characteristic, but it is not necessary that every embodiment includes the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0034] It shall be understood that although the terms “first” and “second” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0035] The terminology used herein is for the purpose of describing particular embodiments and is not intended to be limiting of example embodiments. As used herein, the singular forms “a” , “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” , “comprising” , “has” , “having” , “includes” and / or “including” , when used herein, specify the presence of stated features, elements, and / or components etc., but do not preclude the  presence or addition of one or more other features, elements, components and / or combinations thereof. As used herein, “at least one of the following: <a list of two or more elements>” and “at least one of <a list of two or more elements>” and similar wording, where the list of two or more elements are joined by “and” or “or” , mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.

[0036] As used in this application, the term “circuitry” may refer to one or more or all of the following:

[0037] (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and

[0038] (b) combinations of hardware circuits and software, such as (as applicable) :

[0039] (i) a combination of analog and / or digital hardware circuit (s) with software / firmware and

[0040] (ii) any portions of hardware processor (s) with software (including digital signal processor (s) ) , software, and memory (ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and

[0041] (c) hardware circuit (s) and or processor (s) , such as a microprocessor (s) or a portion of a microprocessor (s) , that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.

[0042] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0043] As used herein, the term “communication network” refers to a network following any suitable communication standards, such as Long Term Evolution (LTE) , LTE-Advanced (LTE-A) , Wideband Code Division Multiple Access (WCDMA) , High-Speed Packet Access (HSPA) , Narrow Band Internet of Things (NB-IoT) and so on. Furthermore, the communications between a terminal device and a network device in the communication  network may be performed according to any suitable generation communication protocols, including, but not limited to, the first generation (1G) , the second generation (2G) , 2.5G, 2.75G, the third generation (3G) , the fourth generation (4G) , 4.5G, the future fifth generation (5G) communication protocols, and / or any other protocols either currently known or to be developed in the future. Embodiments of the present disclosure may be applied in various communication systems. Given the rapid development in communications, there will of course also be future type communication technologies and systems with which example embodiments of the present disclosure may be embodied. It should not be seen as limiting the scope of the present disclosure to only the aforementioned system.

[0044] As used herein, the term “network device” or "network element" refers to a node in a communication network via which a terminal device accesses the network and receives services therefrom. The communication network may include a core network (CN) . The network device or element in CN (also referred to as core network element) may refer to location management function (LMF) . The communication network may include a radio access network (RAN) . The network device in RAN may refer to a base station (BS) or an access point (AP) , for example, a node B (NodeB or NB) , an evolved NodeB (eNodeB or eNB) , a new radio (NR) next generation NodeB (also referred to as a gNB) , a Remote Radio Unit (RRU) , a radio header (RH) , a remote radio head (RRH) , a relay, a low power node such as a femto, a pico, and so forth, depending on the applied terminology and technology.

[0045] The term “terminal device” refers to any end device that may be capable of wireless communication. By way of example rather than limitation, a terminal device may also be referred to as a user equipment (UE) , a Subscriber Station (SS) , a Portable Subscriber Station, a Mobile Station (MS) , or an Access Terminal (AT) . The terminal device may include, but not limited to, a mobile phone, a cellular phone, a smart phone, voice over IP (VoIP) phones, wireless local loop phones, a tablet, a wearable terminal device, a personal digital assistant (PDA) , portable computers, desktop computer, image capture terminal devices such as digital cameras, gaming terminal devices, music storage and playback appliances, vehicle-mounted wireless terminal devices, wireless endpoints, mobile stations, laptop-embedded equipment (LEE) , laptop-mounted equipment (LME) , USB dongles, smart devices, wireless customer-premises equipment (CPE) , an Internet of Things (loT) device, a watch or other wearable, a head-mounted display (HMD) , a vehicle, a drone, a medical device and applications (e.g., remote surgery) , an industrial device and applications (e.g., a robot and / or other wireless devices operating in an industrial and / or an automated processing chain contexts) , a  consumer electronics device, a device operating on commercial and / or industrial wireless networks, and the like. In the following description, the terms “terminal device” , “terminal” , “user equipment” and “UE” may be used interchangeably.

[0046] Various cameras can be used to implement the camera localization. Compared with a stereo or depth camera, a monocular camera cannot measure an object’s distance, but thanks to the simplicity of its structure and universality, it is still widely used in the localization field. Monocular camera calibration is a fundamental prerequisite to guarantee the localization performance. It can use a calibration model to present the relationship between the 3D coordinates in the real world of an observed object and the 2D coordinates of its projection in the image. Thus, the calibration model for object localization in real world from camera is used. In general, some known points in the real world and their projections in the image are used to compute such camera calibration model parameters.

[0047] For example, in a factory scenario, workers draw a certain size of grid on the ground. Some physical points from the ground and corresponding pixels in an image from the camera are selected. By a pre-training or least-squares solution, camera calibration model parameters can be computed. With the computed calibration model, a worker for example can be localized on the field in real-world coordinates from the middle point of the bottom of their bounding box in the image.

[0048] Radar sensors have good accuracy for measuring the ranges of targets in the environment. Camera calibration assisted with a Radar / Lidar has also been widely used in localization works. That is, the Radar / Lidar provides the 3D coordinates of the corresponding points (CPs) of image pixels to compute the calibration model. Thus, a series of CPs can be set, and their 3D coordinates and 2D image projections are known in the radar coordinate system and in the camera coordinate system, respectively. The camera calibration model can be estimated from these 3D-2D CPs.

[0049] This solution needs some special geometrical objects to obtain 3D-2D CPs or geometric constraints between them. For example, the corners of a wall are extracted as the 3D-2D CPs in a scene to calculate the calibration model parameters. The corners of a wall are ubiquitous, so this solution is widely available. One strategy is utilized to estimate the 3D-2D line correspondences and compute the calibration model. A camera calibration approach is presented to automatize the parameter estimation by utilizing segmentation information from images and point clouds.

[0050] The good accuracy of targets’ ranges measurement from the radar may be utilized to mark the 3D coordinates of objects in the real world. But with this solution, the data of the 3D point cloud from the Radar / Lidar is huge, it imposes a burden on communication. Additionally, such solution may need the accurate matching of the CPs from the Radar / Lidar 3D and image 2D coordinate systems. The points on corners, lines or areas can be used for the matching, but these features don’ t have sufficient constraints, and they are prone to mismatch and to be disturbed by noise.

[0051] In addition, a traditional manual calibration method may require significant manual intervention and often introduces measurement errors. It is also time-consuming. A 20m ×20m field takes about half or one day to do calibration. Thus, the traditional manual calibration method is inaccurate and very inefficient.

[0052] In view of the above analysis and discussions, some example embodiments of the present disclosure provide a semantic hierarchical multi-modal matching mechanism for monocular camera calibration and localization. With two-layer semantic hierarchical information matching from objects to points, the corresponding points relationship between the real world and the image pixels is more accurate and robust to environment disturbance. Multi-modal semantic elements (SEs) relevant to calibration task is transmitted on the channel, rather than transmitting various data of greater size, which can reduce the burden of the communication greatly. The proposed mechanism does not need any special marker, it can select a suitable static or moving object from a structure environment quickly and automatically.

[0053] Accordingly, some embodiments of the present disclosure provide a solution to implement accurate and effective camera calibration and localization. In this solution, a device can obtain first data including first semantic information of an object and first coordinate information of the object in a physical space, and can obtain second data including second semantic information of the object and second coordinate information of the object in an image of the physical space captured by a camera. With the matching between the first and second semantic information, the device can determine calibration data of the camera based on the first and second coordinate information. As such, a semantic hierarchical multi-modal matching mechanism for camera calibration and localization is provided, thereby implementing the accurate matching of CPs from the real world and image pixels and reducing the communication burden. Principles and embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0054] FIG. 1 illustrates a schematic diagram of an example environment 100 in which some embodiments of the present disclosure can be implemented. As shown in FIG. 1, the environment 100 may include a plurality of devices such as a device 110, a device 120 and a device 130. As used herein, in various example embodiments and for ease of discussion, the device 110 can also be referred to as a first device 110, the device 120 can also be referred to as a second device 120, the device 130 can also be referred to as a third device 130. The device 110 may communicate with the device 120 and 130. It is to be understood that the number of devices in FIG. 1 is given for the purpose of illustration without suggesting any limitations to the present disclosure. The environment 100 may include any suitable number of devices adapted for implementing example embodiments of the present disclosure.

[0055] In some embodiments, the device 110 may be any type of communication device in the communication network, for example, the mobile terminal or the LMF as described above, or another kind of computing device (e.g. a server, a PC, etc. ) . Additionally or alternatively, the device 110 may include, for example, a memory or a computer readable storage medium configured to store information, data, or the like for enabling the device 110 to carry out various functions in accordance with an example embodiment of the present disclosure. The configuration of the device 120 or 130 may be same or similar with that of the device 110, so the same or similar part will not be described repeatedly herein.

[0056] In some embodiments, the functions of devices 110, 120 and 130 may be combined and implemented as a device, or the devices 110, 120 and 130 may be integrated into an apparatus. The integrated device or apparatus may be embodied by or otherwise associated with any communication device or computing device.

[0057] More details of some embodiments of the present disclosure will be described with reference to FIG. 2. FIG. 2 shows a signaling chart illustrating a process for camera calibration and localization according to some embodiments of the present disclosure. For the purpose of discussion, the process 200 will be described with reference to FIG. 1. The process 200 may involve the devices 110, 120 and 130 in FIG. 1. It is to be appreciated that any graphic elements, numerical values, and descriptive text in these figures are for the purpose of illustration without suggesting any limitations.

[0058] In the process 200, the device 110 obtains (201) first data transmitted (202) from the device 120, for example, the device having capability of joint communication and sensing (JCAS) . While the device 110 receives the first data from the device 120 in the example of  FIG. 2, it is merely an example manner in which the device 110 obtains the first data. In some other embodiments, the device 110 can obtain the first data from any other one or more appropriate devices, apparatuses, entities, or parties. Alternatively, in some further embodiments, the device 110 can obtain the first data by locally generating the first data. The first data includes first semantic information of a first object and first coordinate information of the first object in a physical space, and can be used for determining calibration data of a camera which captures an image of the physical space. Although some example embodiments are described by taking a camera as an example of an image capturing device, it is understood that various example embodiments of the present disclosure can also be applied to any other devices which can capture an image. The first semantic information is extracted (203) from sensed data of the first object in a physical space at the device 120 side, the sensed data being obtained (204) by sensing the first object. The first coordinate information is extracted (205) from the sensed data at the device 120 side.

[0059] In some embodiments, the first data may be received from the device 120 via a first interface. The first interface may include an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field may include a content for an object label of the first object. The at least one characteristic field may include a content for at least one characteristic of the first object. The plurality of coordinate fields may include three-dimensional (3D) coordinates of a plurality of points of the first object.

[0060] In addition, the device 110 obtains (206) second data transmitted (207) from the device 130, for example, the device having capability of vision artificial intelligent (AI) processing. While the device 110 receives the second data from the device 130 in the example of FIG. 2, it is merely an example manner in which the device 110 obtains the second data. In some other embodiments, the device 110 can obtain the second data from any other one or more appropriate devices, apparatuses, entities, or parties. Alternatively, in some further embodiments, the device 110 can obtain the second data by locally generating the first data. The second data includes second semantic information of a second object and second coordinate information of the second object in the image of the physical space captured by the camera, and can be used for determining the calibration data of the camera. At the device 120 side, the image of the physical space captured by the camera is obtained (208) , the second semantic information of the second object in the physical space is extracted (209) from the image, and the second coordinate information of the second object in the image is extracted (210) from the image.

[0061] In some embodiments, the second data may be received from the device 130 via a second interface. The second interface may include an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field may include a content for an object label of the second object. The at least one characteristic field may include a content for at least one characteristic of the second object. The plurality of coordinate fields may include two-dimensional (2D) coordinates of a plurality of points of the second object.

[0062] Then, based on the first semantic information and the second semantic information, the device 110 determine (211) whether the first object and the second object are a same object. Based on determining that the first object and the second object are the same object, the device 110 determine (212) calibration data of the camera based on the first coordinate information and the second coordinate information.

[0063] In some embodiments, based on determining that the first semantic information matches the second semantic information, the device 110 may determine that the first object and the second object are the same object. Alternatively, based on determining that the first semantic information does not match the second semantic information, the device 110 may determine that the first object and the second object are different objects.

[0064] More specifically, as an example, the device 110 may determine whether a type of the first semantic information and a type of the second semantic information are a same type. Based on determining that the type of the first semantic information and the type of the second semantic information are the same type, the device 110 may determine whether a corresponding content for the type of the first semantic information is same with a corresponding content for the type of the second semantic information. Based on determining that the corresponding content for the type of the first semantic information is same with the corresponding content for the type of the second semantic information, the device 110 may determine that the first object and the second object are the same object.

[0065] In some embodiments, the type of the first semantic information may include an object label of the first object. Further, the type of the first semantic information may include at least one characteristic of the first object. The at least one characteristic of the first object may be predetermined for the first object, and may be stored in the device 120. For example, the at least one characteristic of the first object may include at least one of a size, a shape, a moving direction, or a moving speed of the first object. The corresponding  content for the type of the first semantic information may include at least one of a value, a symbol, or a text.

[0066] In some embodiments, the type of the second semantic information may include an object label of the second object. Further, the type of the second semantic information may include at least one characteristic of the second object. The at least one characteristic of the second object is predetermined for the second object, and may be stored in the device 130. For example, the at least one characteristic of the second object may include at least one of a size, a shape, a moving direction, or a moving speed of the second object. The corresponding content for the type of the second semantic information may include at least one of a value, a symbol, or a text.

[0067] For the determination that the first object and the second object are the same object, more details will be further described here with reference to FIGS. 3B and 4B. For example, the device 110 may learn that the object label of the first object is a symbol of car, and that the object label of the second object also is a symbol of car, thereby determining that the first object and the second object are the same object. In other words, if there is no characteristic filed and object labels of the first and second objects are the same, the first and second objects are the same object. Preferably, as another example, the device 110 may learn that the object label of the first object is a symbol of car and a moving speed of the first object is a value of 2, and that the object label of the second object also is a symbol of car and a moving speed of the second object also is a value of 2, thereby determining that the first object and the second object are the same object. In yet another example, the device 110 may learn that the first one of fields included in the first semantic information of the first object is a symbol of car, and that the first one of fields included in the second semantic information of the second object also is a symbol of car, thereby determining that the first object and the second object are the same object.

[0068] With the determination that the first object and the second object are the same object, based on determining that no previous calibration data of the camera is stored in the device 110, the device 110 may calculate the calibration data based on the first coordinate information and the second coordinate information. For example, the device 110 may determine a set of 3D coordinates of a set of points of the same object from the first coordinate information, and determine a set of 2D coordinates of the set of points of the same object corresponding to the set of 3D coordinates from the second coordinate information. It is to be understood that the set of points of the same object may be predetermined for the same  object. For example, the set of points of the first object is predetermined for the first object and stored in the device 120, and the set of points of the second object is predetermined for the second object and stored in the device 130. Next, based on the set of 3D coordinates and the set of 2D coordinates, the device 110 may calculate at least one of a rotation matrix or a translation vector of the camera.

[0069] Alternatively, with the determination that the first object and the second object are the same object, based on determining that previous calibration data of the camera is stored in the device 110, the device 110 may update or optimize the previous calibration data based on the first coordinate information and the second coordinate information to obtain the calibration data. For example, the device 110 may determine the set of 3D coordinates and the set of 2D coordinates, and determine another set of 3D coordinates by projecting the second set of 2D coordinates to the physical space based on the previous calibration data. Next, the device 110 may adjust the previous calibration data to minimize a difference between the set of 3D coordinates and another set of 3D coordinates.

[0070] Now some more embodiments of the present disclosure herein will be further described in detail with reference to FIGS. 3A, 3B, 4A, 4B, and 5 below for the purpose of clearer understanding. FIG. 3A shows a multi-modal 3D SEs interface definition. FIG. 3B shows an example of multi-modal 3D SEs interface. FIG. 4A shows a multi-modal 2D SEs interface definition. FIG. 4B shows an example of multi-modal 2D SEs interface. The semantic hierarchical multi-modal matching and calibration model generating workflow is illustrated in FIG. 5.

[0071] As described above, in order to get the accurate matching of CPs from the real world and image pixels and reduce the communication burden, a semantic hierarchical multi-modal matching mechanism for monocular camera calibration and localization is proposed. This mechanism may use the sensing capability of joint communication and sensing (JCAS) technology and the artificial intelligent (AI) processing capability of vision.

[0072] As shown in FIG. 5, the JCAS may sense objects in a scene via a JCAS sensing radar. In particular, the JCAS may sense certain typical characteristics of the objects (e.g. a size, a shape, a position, a moving direction and a speed, etc. ) . Then, the JCAS may extract (501) semantic data of the objects from the sensed data. For example, the JCAS may distinguish and extract SEs for different objects in the 3D real world. The SEs may include an object label, multiple typical characteristic values and coordinates of the indicated  specific points for different object. The specific characteristics and points may be determined for the specific object. Also, the number of the characteristics and specific points may be determined for the specific object. For example, a hydrant or a car may be sensed. The characteristic of the hydrant may be its shape, and the transmitted characteristic data may be a shape value. The specific points are chosen from significant points of the object, like the points on the head, bottom or corners of the hydrant. The transmitted data corresponding to these points may include their 3D positions. The characteristics of the car may include its shape, moving direction and moving speed. The specific points may be on its car light, wheel or head.

[0073] Next, the JCAS may transmit the extracted semantic data to a camera calibration model generator (CCMG) via a multi-modal 3D SEs interface. The multi-modal 3D SEs interface is defined for transmitting the semantic data from the JCAS to the CCMG. In this case, the 3D SEs interface may include the object label, the typical characteristics and the coordinates of the indicated specific points of the sensed objects. As shown in FIG. 3A, the defined typical characteristics of the object may include an object label 310 and characteristics 1-M 321-32m, and the coordinates of the defined specific points of the object may include positions of points 1-N 331-33n. In an example shown in FIG. 3B, the object label 310 may be a hydrant, the typical characteristic may be the shape of the hydrant 321, and the coordinates of the indicated specific points of the hydrant may include (x1, y1, z1) 331, (x2, y2, z2) 332. . . (xk, yk, zk) 33k. In another example shown in FIG. 3B, the object label 310 may be a car, the typical characteristic may be its shape 321, moving direction 322 and moving speed 323, and the coordinates of the indicated specific points of the car may include (x1, y1, z1) 331, (x2, y2, z2) 332. . . (xn, yn, zn) 33n. Here, the k and n can be any integer greater than or equal to 6.

[0074] The vision AI processing entity (VAIPE) may process images from a camera. Similarly, the VAIPE may detect the objects from the images, and extract (502) semantic data from the detected objects. For example, the VAIPE may recognize and extract SEs for different objects in 2D image as semantic data. Then, the VAIPE may transmit the extracted SEs to the CCMG via a multi-modal 2D SEs interface. The multi-modal 2D SEs interface is defined for transmitting the semantic data from the VAIPE to the CCMG. In this case, the 2D SEs interface may include an object label, multiple typical characteristics and coordinates of the indicated specific points of the detected objects. The typical characteristics and specific points are also determined for the objects. As shown in FIG.  4A, the defined typical characteristics of the object may include an object label 410 and characteristics 1-M 412-42m, and the coordinates of the defined specific pixels of the object may include positions of points 1-N 431-43n. In an example shown in FIG. 4B, the object label 410 may be a hydrant, the typical characteristic may be the shape of the hydrant 421, and the coordinates of the indicated specific pixels of the hydrant may include (u1, v1) 431, (u2, v2) 432 . . . (uk, vk) 43k. In another example shown in FIG. 4B, the object label 410 may be a car, the typical characteristic may be its shape 421, moving direction 422 and moving speed 423, and the coordinates of the indicated specific points of the car may include (u1, v1) 431, (u2, v2) 432 . . . (un, vn) 43n. Here, the k and n can be any integer greater than or equal to 6.

[0075] In the VAIPE, the selection of the typical characteristics and the specific points of the object should be consistent with that of the multi-modal 3D SEs interface definition, because the JCAS and the camera which needs to be calibrated should have an overlapped field of vision (FOV) , i.e., the camera which needs to be calibrated have the overlapped FOV with the JCAS and the matched image pixels and point cloud are in the overlapped FOV. The transmitted data of the characteristic may include a corresponding characteristic value. For example, the transmitted data may include the 2D image pixel positions corresponding to the indicated points.

[0076] It is to be understood that the term “multi-modal” herein refers to the original data being the point cloud or image herein, i.e., the SEs interface including data with different types, and the extracted SEs are semantic data (text, value, etc. ) . In addition, the characteristics and lists of the indicated specific points for different objects may be stored in a knowledge base (KB) of the JCAS and a KB of the VAIPE. The KBs in JCAS and VAIPE should be consistent so that the CCMG can compare them correctly.

[0077] The CCMG is designed for this semantic hierarchical multi-modal matching and calibration model computation or optimization. In particular, one hierarchical matching scheme can be implemented in the CCMG. For example, the SEs received (503) by the CCMG from the JCAS and VAIPE should be synchronized. The objects which have sufficient points for calibration model are chosen for matching. The object in 3D coordinates in the real world and its 2D image projection are matched based on the SEs information (for example, the object label and characteristic (s) ) . Thus, according to the received SEs, the objects in the 3D real world and 2D image can be matched (504) easily.

[0078] Further, based on the matching result of objects, the specific points of each object defined in the KB will be used for CPs matching. The matched object has the same list of the specific points in the 3D point cloud and 2D image, so the 3D-2D CPs from the received SEs are matched (505) as a 3D-2D CP pair. The matched 3D-2D CPs can be used to determine the calibration data for the camera. For example, the position coordinates of 3D-2D CPs may be used for calibration model computation or optimization, depending on whether a calibration model does exist (506) .

[0079] If the calibration model does not exist, the CCMG can compute (507) a camera calibration model for localization using the position coordinates of 3D-2D CPs received from the SEs interface and output it. The camera calibration model computation process may refer to a pinhole camera model.

[0080] wherein [u, v, 1] T denotes the camera image coordinates, and [x, y, z, 1] T is the point cloud coordinates in the real world. The camera calibration model parameters which are needed to be obtained include a rotation matrix and a translation vector as shown in equation (1) . The intrinsic is assumed to be known. Thus, with certain suitable 3D-2D CP pairs, the camera calibration model can be computed.

[0081] If the calibration model already exists, the model can be further optimized. For example, based on the position coordinates received from the SEs interface, the computed calibration model can be optimized further for better accuracy. In particular, image pixels from the matched 3D-2D CPs are projected back to 3D coordinates of the real world based on the stored calibration model. That is, the 3D position in real world of the image pixel can be predicted or computed (508) using the calibration model. The error between the predicted and real positions can be trained for calibration model optimization. For example, the error between the back-projected point set Vpre from the image and the real point set Vreal from the JCAS can be used to optimize (509) the calibration model through a training loss function as shown in equation (2) .

[0082] The training loss function is:

[0083] wherein n is the number of points in a set. By adjusting R and t to minimize the difference  between the back-projected points and the real points, the optimal camera calibration model can be determined.

[0084] In an example for verification, this mechanism has been adopted for the factory scenario localization. The test and verification results are fully meeting the actual demand for factory production. This mechanism is also planned to be further promoted to other factory environment. Except for factory scenarios, this mechanism also can meet other localization requirement because of the promising JCAS technology, ubiquitous monocular camera devices and good generality and flexibility of this mechanism.

[0085] An example of the test and verification result of camera calibration model computation using the proposed semantic hierarchical multi-modal matching mechanism is shown in Table 1 below.

[0086] Table 1. Test and verification result of camera calibration model computation

[0087] The Table 1 gives the test and verification result of camera calibration model computation for our factory scenario. In an example, the resolution ratio of the test image is 3840×2160, that is about 8 million pixels in one image. The actual test and verification scenario is a 16m×30m rectangular field.

[0088] As shown in the Table 1, the first two columns are the selected 3D-2D CP pairs for testing from the matched object in the 3D point cloud and 2D image. It is to be understood that the number of the selected 3D-2D CP pairs is given for the purpose of illustration without suggesting any limitations to the present disclosure. The 3D point cloud is obtained from the sensed objects via the JCAS, and the 2D image is obtained from the camera. Test points are in the overlapped FOV of the JCAS and camera. The test points are used to compute the camera calibration model.

[0089] The data of columns 3, 4, and 5 is used for verification. The column 3 is one 2D image pixel from the camera. The column 4 is the computed physical position based on the image pixel of column 3 and the calibration model. The column 5 gives the actual position  of such pixel. In the case that the position of the object on the ground for the localization application of the factory scenario is verified, the z-axis coordinate values of the verified points are ignored. The verified points are in the FOV of the camera regardless of whether they are within the range of the FOV of the JCAS.

[0090] As shown in the Table 1, the corresponding verified point of columns 3, 4, and 5 presents the calibration model verification result. The calculated position using the obtained camera calibration module for the 2D image pixel (2247, 577) is (8.89071146m, 5.2063376m) , the real position of the verified point in the physical space is (8.4m, 5.2m) , and the error for the verified point is about 0.5m. By verifying a large number of points, it can also be concluded that the average error of the verified points is about 0.5m.

[0091] Table 2. Test and verification result of camera calibration model optimization

[0092] The Table 2 is the test and verification result of the camera calibration model optimization. The resolution ratio of the image is also 3840×2160. The actual test and verification scenario is a 16m×30m rectangular field of the factory scenario.

[0093] As shown in the Table 2, the column 1 is the 2D image pixels of the test points. The error between the back-projection of such 2D image pixels and their corresponding 3D point cloud in the column 2 are used for calibration model optimization. Based on the optimized calibration model, the corresponding verified point of column 3, 4, and 5 shows the verification result. The calculated position using the optimized camera calibration model for the 2D image pixel (2305, 885) is (8.15172176m, 5.24838043m) . The real position of the verified point in the physical space is (8.4m, 5.2m) , the error for the verified point is about 0.2m. By verifying a large number of points, it can also be concluded that the average error of the verified points is about 0.2m.

[0094] With the solutions as described above, some advantages can be obtained. For example, for camera calibration model computation and optimization, the proposed semantic hierarchical multi-modal matching mechanism performs matching operations for the objects and 3D-2D CPs, which improving the matching accuracy. In addition, this mechanism is  robust to dynamic environment noise. The proposed calibration and localization mechanism is suitable for different scenario. There is no need for any special marker. It can achieve the matching for a suitable static or moving object and points from the structure environment quickly and perform the calibration automatically. The multi-modal SEs are extracted and transmitted on the channel, which reduces the communication payload to about 1%when compared to the traditional scheme. The camera is calibrated using the overlapped FOV of the camera and JCAS. Once the camera is calibrated, other FOV of the camera also can be localized.

[0095] FIG. 6 shows a flowchart of an example method 600 implemented at a first device in accordance with some embodiments of the present disclosure. For the purpose of discussion, the method 600 will be described from the perspective of the device 110 with reference to FIG. 1.

[0096] At block 610, the device 110 obtains first data including first semantic information of a first object and first coordinate information of the first object in a physical space. At block 620, the device 110 obtains second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera. At block 630, based on the first semantic information and the second semantic information, the device 110 determines whether the first object and the second object are a same object. At block 640, based on determining that the first object and the second object are the same object, the device 110 determines calibration data of the camera based on the first coordinate information and the second coordinate information.

[0097] In some embodiments, based on determining that the first semantic information matches the second semantic information, the device 110 may determine that the first object and the second object are the same object. Alternatively, based on determining that the first semantic information does not match the second semantic information, the device 110 may determine that the first object and the second object are different objects.

[0098] In some embodiments, the device 110 may determine whether a type of the first semantic information and a type of the second semantic information are a same type. Based on determining that the type of the first semantic information and the type of the second semantic information are the same type, the device 110 may determine whether a corresponding content for the type of the first semantic information is same with a corresponding content for the type of the second semantic information. Based on  determining that the corresponding content for the type of the first semantic information is same with the corresponding content for the type of the second semantic information, the device 110 may determine that the first object and the second object are the same object.

[0099] In some embodiments, the type of the first semantic information may comprise an object label of the first object. In addition, the type of the first semantic information may further comprise at least one characteristic of the first object.

[0100] In some embodiments, the at least one characteristic of the first object includes at least one of a size, a shape, a moving direction, or a moving speed of the first object.

[0101] In some embodiments, the corresponding content for the type of the first semantic information includes at least one of a value, a symbol, or a text.

[0102] In some embodiments, the type of the second semantic information may comprise an object label of the second object. In addition, the type of the second semantic information may further comprise at least one characteristic of the second object.

[0103] In some embodiments, the at least one characteristic of the second object includes at least one of a size, a shape, a moving direction, or a moving speed of the second object.

[0104] In some embodiments, the corresponding content for the type of the second semantic information includes at least one of a value, a symbol, or a text.

[0105] In some embodiments, the device 110 may receive the first data from the device 120 having capability of joint communication and sensing (JCAS) .

[0106] In some embodiments, the device 110 may receive the second data from the device 130 having capability of vision artificial intelligent (AI) processing.

[0107] In some embodiments, the first data may be received from the device 120 via a first interface. The first interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the first object, the at least one characteristic field includes a content for at least one characteristic of the first object, and the plurality of coordinate fields include three-dimensional (3D) coordinates of a plurality of points of the first object.

[0108] In some embodiments, the second data may be received from the device 130 via a second interface. The second interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the second object, the at least one characteristic field includes  a content for at least one characteristic of the second object, and the plurality of coordinate fields include two-dimensional (2D) coordinates of a plurality of points of the second object.

[0109] In some embodiments, the at least one characteristic of the first object is predetermined for the first object. Alternatively, the at least one characteristic of the second object is predetermined for the second object.

[0110] In some embodiments, based on determining that no previous calibration data of the camera is stored in the device 110, the device 110 may calculate the calibration data based on the first coordinate information and the second coordinate information. Alternatively, based on determining that previous calibration data of the camera is stored in the device 110, the device 110 may update the previous calibration data based on the first coordinate information and the second coordinate information to obtain the calibration data.

[0111] In some embodiments, the device 110 may determine, from the first coordinate information, a set of 3D coordinates of a set of points of the same object. The device 110 may determine, from the second coordinate information, a set of 2D coordinates of the set of points of the same object corresponding to the set of 3D coordinates. The device 110 may calculate, based on the set of 3D coordinates and the set of 2D coordinates, at least one of a rotation matrix or a translation vector of the camera.

[0112] In some embodiments, the device 110 may determine, from the first coordinate information, a first set of 3D coordinates of a set of points of the same object. The device 110 may determine, from the second coordinate information, a second set of 2D coordinates of the set of points of the same object corresponding to the first set of 3D coordinates. The device 110 may determine a third set of 3D coordinates by projecting the second set of 2D coordinates to the physical space based on the previous calibration data. The device 110 may adjust the previous calibration data to minimize a difference between the first set of 3D coordinates and the third set of 3D coordinates.

[0113] In some embodiments, the set of points of the same object is predetermined for the same object.

[0114] FIG. 7 shows a flowchart of an example method 700 implemented at a second device in accordance with some embodiments of the present disclosure. For the purpose of discussion, the method 700 will be described from the perspective of the device 120 with reference to FIG. 1.

[0115] At block 710, the device 120 obtains sensed data of a first object in a physical space  by sensing the first object. At block 720, the device 120 extracts, from the sensed data, first semantic information of the first object. At block 730, the device 120 extracts, from the sensed data, first coordinate information of the first object in the physical space. At block 740, the device 120 transmits, to the device 110, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.

[0116] In some embodiments, a type of the first semantic information may comprise an object label of the first object. In addition, the type of the first semantic information may further comprise at least one characteristic of the first object.

[0117] In some embodiments, the at least one characteristic of the first object includes at least one of a size, a shape, a moving direction, or a moving speed of the first object.

[0118] In some embodiments, a corresponding content for the type of the first semantic information includes at least one of a value, a symbol, or a text.

[0119] In some embodiments, the first data may be transmitted to the device 110 via a first interface. The first interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the first object, the at least one characteristic field includes a content for at least one characteristic of the first object, and the plurality of coordinate fields include three-dimensional (3D) coordinates of a plurality of points of the first object.

[0120] In some embodiments, the at least one characteristic of the first object is predetermined for the first object and stored in the device 120.

[0121] In some embodiments, the first coordinate information comprises a set of 3D coordinates of a set of points of the first object.

[0122] In some embodiments, the set of points of the first object is predetermined for the first object and stored in the device 120.

[0123] In some embodiments, the device 120 has capability of joint communication and sensing (JCAS) .

[0124] FIG. 8 shows a flowchart of an example method 800 implemented at a third device in accordance with some embodiments of the present disclosure. For the purpose of discussion, the method 800 will be described from the perspective of the device 130 with reference to FIG. 1.

[0125] At block 810, the device 130 obtains an image of a physical space captured by a camera. At block 820, the device 130 extracts, from the image, second semantic information of a second object in the physical space. At block 830, the device 130 extracts, from the image, second coordinate information of the second object in the image. At block 840, the device 130 transmits, to the device 110, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.

[0126] In some embodiments, a type of the second semantic information may comprise an object label of the second object. In addition, the type of the second semantic information may further comprise at least one characteristic of the second object.

[0127] In some embodiments, the at least one characteristic of the second object includes at least one of a size, a shape, a moving direction, or a moving speed of the second object.

[0128] In some embodiments, a corresponding content for the type of the second semantic information includes at least one of a value, a symbol, or a text.

[0129] In some embodiments, the second data may be transmitted to the device 110 via a second interface. The second interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the second object, the at least one characteristic field includes a content for at least one characteristic of the second object, and the plurality of coordinate fields include two-dimensional (2D) coordinates of a plurality of points of the second object.

[0130] In some embodiments, the at least one characteristic of the second object is predetermined for the second object and stored in the device 130.

[0131] In some embodiments, the second coordinate information comprises a set of 2D coordinates of a set of points of the second object.

[0132] In some embodiments, the set of points of the second object is predetermined for the second object and stored in the device 130.

[0133] In some embodiments, the device 130 has capability of vision artificial intelligent (AI) processing.

[0134] In some embodiments, an apparatus capable of performing the method 600 (for example, the device 110) may include means for performing the respective steps of the method 600. The means may be implemented in any suitable form. For example, the  means may be implemented in a circuitry or software module.

[0135] In some embodiments, the apparatus includes: means for obtaining first data including first semantic information of a first object and first coordinate information of the first object in a physical space; means for obtaining second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera; means for determining, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; and means for based on determining that the first object and the second object are the same object, determining calibration data of the camera based on the first coordinate information and the second coordinate information.

[0136] In some embodiments, the means for determining whether the first object and the second object are the same object may include means for based on determining that the first semantic information matches the second semantic information, determining that the first object and the second object are the same object; or means for based on determining that the first semantic information does not match the second semantic information, determining that the first object and the second object are different objects.

[0137] In some embodiments, the means for determining that the first object and the second object are the same object based on determining that the first semantic information matches the second semantic information may include means for determining whether a type of the first semantic information and a type of the second semantic information are a same type; means for based on determining that the type of the first semantic information and the type of the second semantic information are the same type, determining whether a corresponding content for the type of the first semantic information is same with a corresponding content for the type of the second semantic information; and means for based on determining that the corresponding content for the type of the first semantic information is same with the corresponding content for the type of the second semantic information, determining that the first object and the second object are the same object.

[0138] In some embodiments, the type of the first semantic information may comprise an object label of the first object. In addition, the type of the first semantic information may further comprise at least one characteristic of the first object.

[0139] In some embodiments, the at least one characteristic of the first object includes at least one of a size, a shape, a moving direction, or a moving speed of the first object.

[0140] In some embodiments, the corresponding content for the type of the first semantic information includes at least one of a value, a symbol, or a text.

[0141] In some embodiments, the type of the second semantic information may comprise an object label of the second object. In addition, the type of the second semantic information may further comprise at least one characteristic of the second object.

[0142] In some embodiments, the at least one characteristic of the second object includes at least one of a size, a shape, a moving direction, or a moving speed of the second object.

[0143] In some embodiments, the corresponding content for the type of the second semantic information includes at least one of a value, a symbol, or a text.

[0144] In some embodiments, the means for obtaining the first data may include means for receiving the first data from the device 120 having capability of joint communication and sensing (JCAS) .

[0145] In some embodiments, the means for obtaining the second data may include means for receiving the second data from the device 130 having capability of vision artificial intelligent (AI) processing.

[0146] In some embodiments, the first data may be received from the device 120 via a first interface. The first interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the first object, the at least one characteristic field includes a content for at least one characteristic of the first object, and the plurality of coordinate fields include three-dimensional (3D) coordinates of a plurality of points of the first object.

[0147] In some embodiments, the second data may be received from the device 130 via a second interface. The second interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the second object, the at least one characteristic field includes a content for at least one characteristic of the second object, and the plurality of coordinate fields include two-dimensional (2D) coordinates of a plurality of points of the second object.

[0148] In some embodiments, the at least one characteristic of the first object is predetermined for the first object. Alternatively, the at least one characteristic of the second object is predetermined for the second object.

[0149] In some embodiments, the means for determining the calibration data of the camera  may include means for based on determining that no previous calibration data of the camera is stored in the device 110, calculating the calibration data based on the first coordinate information and the second coordinate information; or means for based on determining that previous calibration data of the camera is stored in the device 110, updating the previous calibration data based on the first coordinate information and the second coordinate information to obtain the calibration data.

[0150] In some embodiments, the means for calculating the calibration data may include means for determining, from the first coordinate information, a set of 3D coordinates of a set of points of the same object; means for determining, from the second coordinate information, a set of 2D coordinates of the set of points of the same object corresponding to the set of 3D coordinates; means for calculating, based on the set of 3D coordinates and the set of 2D coordinates, at least one of a rotation matrix or a translation vector of the camera.

[0151] In some embodiments, the means for updating the previous calibration data may include means for determining, from the first coordinate information, a first set of 3D coordinates of a set of points of the same object; means for determining, from the second coordinate information, a second set of 2D coordinates of the set of points of the same object corresponding to the first set of 3D coordinates; means for determining a third set of 3D coordinates by projecting the second set of 2D coordinates to the physical space based on the previous calibration data; and means for adjusting the previous calibration data to minimize a difference between the first set of 3D coordinates and the third set of 3D coordinates.

[0152] In some embodiments, the set of points of the same object is predetermined for the same object.

[0153] In some embodiments, the apparatus further includes means for performing other steps in some embodiments of the method 600. In some embodiments, the means includes at least one processor and at least one memory including computer program code, the at least one memory and computer program code configured to, with the at least one processor, cause the performance of the apparatus.

[0154] In some embodiments, an apparatus capable of performing the method 700 (for example, the device 120) may include means for performing the respective steps of the method 700. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or software module.

[0155] In some embodiments, the apparatus includes means for obtaining sensed data of a  first object in a physical space by sensing the first object; means for extracting, from the sensed data, first semantic information of the first object; means for extracting, from the sensed data, first coordinate information of the first object in the physical space; and means for transmitting, to the device 110, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.

[0156] In some embodiments, a type of the first semantic information may comprise an object label of the first object. In addition, the type of the first semantic information may further comprise at least one characteristic of the first object.

[0157] In some embodiments, the at least one characteristic of the first object includes at least one of a size, a shape, a moving direction, or a moving speed of the first object.

[0158] In some embodiments, a corresponding content for the type of the first semantic information includes at least one of a value, a symbol, or a text.

[0159] In some embodiments, the first data may be transmitted to the device 110 via a first interface. The first interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the first object, the at least one characteristic field includes a content for at least one characteristic of the first object, and the plurality of coordinate fields include three-dimensional (3D) coordinates of a plurality of points of the first object.

[0160] In some embodiments, the at least one characteristic of the first object is predetermined for the first object and stored in the device 120.

[0161] In some embodiments, the first coordinate information comprises a set of 3D coordinates of a set of points of the first object.

[0162] In some embodiments, the set of points of the first object is predetermined for the first object and stored in the device 120.

[0163] In some embodiments, the device 120 has capability of joint communication and sensing (JCAS) .

[0164] In some embodiments, the apparatus further includes means for performing other steps in some embodiments of the method 700. In some embodiments, the means includes at least one processor and at least one memory including computer program code, the at least one memory and computer program code configured to, with the at least one processor, cause  the performance of the apparatus.

[0165] In some embodiments, an apparatus capable of performing the method 800 (for example, the device 130) may include means for performing the respective steps of the method 800. The means may be implemented in any suitable form. For example, the means may be implemented in a circuitry or software module.

[0166] In some embodiments, the apparatus includes means for obtaining an image of a physical space captured by a camera; means for extracting, from the image, second semantic information of a second object in the physical space; means for extracting, from the image, second coordinate information of the second object in the image; and means for transmitting, to the device 110, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.

[0167] In some embodiments, a type of the second semantic information may comprise an object label of the second object. In addition, the type of the second semantic information may further comprise at least one characteristic of the second object.

[0168] In some embodiments, the at least one characteristic of the second object includes at least one of a size, a shape, a moving direction, or a moving speed of the second object.

[0169] In some embodiments, a corresponding content for the type of the second semantic information includes at least one of a value, a symbol, or a text.

[0170] In some embodiments, the second data may be transmitted to the device 110 via a second interface. The second interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields. The object label field includes a content for an object label of the second object, the at least one characteristic field includes a content for at least one characteristic of the second object, and the plurality of coordinate fields include two-dimensional (2D) coordinates of a plurality of points of the second object.

[0171] In some embodiments, the at least one characteristic of the second object is predetermined for the second object and stored in the device 130.

[0172] In some embodiments, the second coordinate information comprises a set of 2D coordinates of a set of points of the second object.

[0173] In some embodiments, the set of points of the second object is predetermined for the second object and stored in the device 130.

[0174] In some embodiments, the device 130 has capability of vision artificial intelligent (AI) processing.

[0175] In some embodiments, the apparatus further includes means for performing other steps in some embodiments of the method 800. In some embodiments, the means includes at least one processor and at least one memory including computer program code, the at least one memory and computer program code configured to, with the at least one processor, cause the performance of the apparatus.

[0176] FIG. 9 is a simplified block diagram of a device 900 that is suitable for implementing embodiments of the present disclosure. The device 900 may be provided to implement the communication device, for example the device110 or the device 120 as shown in FIG. 1. As shown, the device 900 includes one or more processors 910, one or more memories 920 coupled to the processor 910, and one or more communication modules 940 coupled to the processor 910.

[0177] The communication module 940 is for bidirectional communications. The communication modules 940 have at least one antenna to facilitate communication. The communication interface may represent any interface that is necessary for communication with other network elements.

[0178] The processor 910 may be of any type suitable to the local technical network and may include one or more of the following: general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on multicore processor architecture, as non-limiting examples. The device 900 may have multiple processors, such as an application specific integrated circuit chip that is slaved in time to a clock which synchronizes the main processor.

[0179] The memory 920 may include one or more non-volatile memories and one or more volatile memories. Examples of the non-volatile memories include, but are not limited to, a Read Only Memory (ROM) 924, an electrically programmable read only memory (EPROM) , a flash memory, a hard disk, a compact disc (CD) , a digital video disk (DVD) , and other magnetic storage and / or optical storage. Examples of the volatile memories include, but are not limited to, a random access memory (RAM) 922 and other volatile memories that will not last in the power-down duration.

[0180] A computer program 930 includes computer executable instructions that are executed by the associated processor 910. The program 930 may be stored in the ROM  1020. The processor 910 may perform any suitable actions and processing by loading the program 930 into the RAM 922.

[0181] The embodiments of the present disclosure may be implemented by means of the program 930 so that the device 900 may perform any process of the disclosure as discussed with reference to FIGS. 2 to 5. The embodiments of the present disclosure may also be implemented by hardware or by a combination of software and hardware.

[0182] In some embodiments, the program 930 may be tangibly contained in a computer readable medium which may be included in the device 900 (such as in the memory 920) or other storage devices that are accessible by the device 900. The device 900 may load the program 930 from the computer readable medium to the RAM 922 for execution. The computer readable medium may include any types of tangible non-volatile storage, such as ROM, EPROM, a flash memory, a hard disk, CD, DVD, and the like. FIG. 10 shows an example of the computer readable medium 1000 in form of CD or DVD. The computer readable medium has the program 930 stored thereon.

[0183] Generally, various embodiments of the present disclosure may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device. While various aspects of embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representations, it is to be understood that the block, apparatus, system, technique or method described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0184] Some example embodiments of the present disclosure also provide at least one computer program product tangibly stored on a non-transitory computer readable storage medium. The computer program product includes computer-executable instructions, such as those included in program modules, being executed in a device on a target real or virtual processor, to carry out the methods as described above with reference to FIGS. 2-8. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, or the like that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or split  between program modules as desired in various embodiments. Machine-executable instructions for program modules may be executed within a local or distributed device. In a distributed device, program modules may be located in both local and remote storage media.

[0185] Program code for carrying out methods of example embodiments of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0186] In the context of the present disclosure, the computer program codes or related data may be carried by any suitable carrier to enable the device, apparatus or processor to perform various processes and operations as described above. Examples of the carrier include a signal, computer readable medium, and the like.

[0187] The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , an optical fiber, a portable compact disc read-only memory (CD-ROM) , an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. The term “non-transitory, ” as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM) .

[0188] Further, while operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous.  Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the present disclosure, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0189] Although example embodiments of the present disclosure have been described in languages specific to structural features and / or methodological acts, it is to be understood that the present disclosure defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1.A first device comprising:at least one processor; andat least one memory storing indications that, when executed by the at least one processor, cause the first device at least to:obtain first data including first semantic information of a first object and first coordinate information of the first object in a physical space;obtain second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera;determine, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; andbased on determining that the first object and the second object are the same object, determine calibration data of the camera based on the first coordinate information and the second coordinate information.2.The first device of claim 1, wherein the first device is caused to determine whether the first object and the second object are the same object by:based on determining that the first semantic information matches the second semantic information, determining that the first object and the second object are the same object; orbased on determining that the first semantic information does not match the second semantic information, determining that the first object and the second object are different objects.3.The first device of claim 1, wherein the first device is caused to determine that the first object and the second object are the same object based on determining that the first semantic information matches the second semantic information by:determining whether a type of the first semantic information and a type of the second semantic information are a same type;based on determining that the type of the first semantic information and the type of the second semantic information are the same type, determining whether a corresponding content for the type of the first semantic information is same with a corresponding content for the type of the second semantic information; andbased on determining that the corresponding content for the type of the first semantic information is same with the corresponding content for the type of the second semantic information, determining that the first object and the second object are the same object.4.The first device of claim 3, wherein the type of the first semantic information comprises the following:an object label of the first object; andat least one characteristic of the first object.5.The first device of claim 3, wherein the at least one characteristic of the first object includes at least one of a size, a shape, a moving direction, or a moving speed of the first object.6.The first device of any of claims 3-5, wherein the corresponding content for the type of the first semantic information includes at least one of a value, a symbol, or a text.7.The first device of any of claims 3-6, wherein the type of the second semantic information comprises the following:an object label of the second object; andat least one characteristic of the second object.8.The first device of claim 7, wherein the at least one characteristic of the second object includes at least one of a size, a shape, a moving direction, or a moving speed of the second object.9.The first device of any of claims 3-8, wherein the corresponding content for the type of the second semantic information includes at least one of a value, a symbol, or a text.10.The first device of any of claims 1-9, wherein the first device is caused to obtain the first data by:receiving the first data from a second device having capability of joint communication and sensing (JCAS) .11.The first device of any of claims 1-10, wherein the first device is caused to  obtain the second data by:receiving the second data from a third device having capability of vision artificial intelligent (AI) processing.12.The first device of claim 10 or 11, wherein the first data is received from the second device via a first interface,wherein the first interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields, andwherein the object label field includes a content for an object label of the first object, the at least one characteristic field includes a content for at least one characteristic of the first object, and the plurality of coordinate fields include three-dimensional (3D) coordinates of a plurality of points of the first object.13.The first device of any of claims 10-12, wherein the second data is received from the third device via a second interface,wherein the second interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields, andwherein the object label field includes a content for an object label of the second object, the at least one characteristic field includes a content for at least one characteristic of the second object, and the plurality of coordinate fields include two-dimensional (2D) coordinates of a plurality of points of the second object.14.The first device of any of claims 4-13, wherein at least one of the following:the at least one characteristic of the first object is predetermined for the first object; orthe at least one characteristic of the second object is predetermined for the second object.15.The first device of any of claims 1-14, wherein the first device is caused to determine the calibration data of the camera by:based on determining that no previous calibration data of the camera is stored in the first device, calculating the calibration data based on the first coordinate information and the second coordinate information; orbased on determining that previous calibration data of the camera is stored in the first device, updating the previous calibration data based on the first coordinate information and the second coordinate information to obtain the calibration data.16.The first device of claim 15, wherein the first device is caused to calculate the calibration data by:determining, from the first coordinate information, a set of 3D coordinates of a set of points of the same object;determining, from the second coordinate information, a set of 2D coordinates of the set of points of the same object corresponding to the set of 3D coordinates; andcalculating, based on the set of 3D coordinates and the set of 2D coordinates, at least one of a rotation matrix or a translation vector of the camera.17.The first device of claim 15, wherein the first device is caused to update the previous calibration data by:determining, from the first coordinate information, a first set of 3D coordinates of a set of points of the same object;determining, from the second coordinate information, a second set of 2D coordinates of the set of points of the same object corresponding to the first set of 3D coordinates;determining a third set of 3D coordinates by projecting the second set of 2D coordinates to the physical space based on the previous calibration data; andadjusting the previous calibration data to minimize a difference between the first set of 3D coordinates and the third set of 3D coordinates.18.The first device of claim 16 or 17, wherein the set of points of the same object is predetermined for the same object.19.A second device comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the second device at least to:obtain sensed data of a first object in a physical space by sensing the first object;extract, from the sensed data, first semantic information of the first object;extract, from the sensed data, first coordinate information of the first object in the physical space; andtransmit, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.20.The second device of claim 19, wherein a type of the first semantic information comprises the following:an object label of the first object; andat least one characteristic of the first object.21.The second device of claim 20, wherein the at least one characteristic of the first object includes at least one of a size, a shape, a moving direction, or a moving speed of the first object.22.The second device of claim 20 or 21, wherein a corresponding content for the type of the first semantic information includes at least one of a value, a symbol, or a text.23.The second device of any of claims 19 -22, wherein the first data is transmitted to the first device via a first interface,wherein the first interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields, andwherein the object label field includes a content for an object label of the first object, the at least one characteristic field includes a content for at least one characteristic of the first object, and the plurality of coordinate fields include three-dimensional (3D) coordinates of a plurality of points of the first object.24.The second device of any of claims 20-23, wherein the at least one characteristic of the first object is predetermined for the first object and stored in the second device.25.The second device of any of claims 19-24, wherein the first coordinate information comprises a set of 3D coordinates of a set of points of the first object.26.The second device of claim 25, wherein the set of points of the first object is predetermined for the first object and stored in the second device.27.The second device of any of claims 19-26, wherein the second device has capability of joint communication and sensing (JCAS) .28.A third device comprising:at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, cause the third device at least to:obtain an image of a physical space captured by a camera;extract, from the image, second semantic information of a second object in the physical space;extract, from the image, second coordinate information of the second object in the image; andtransmit, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.29.The third device of claim 28, wherein a type of the second semantic information comprises the following:an object label of the second object; andat least one characteristic of the second object.30.The third device of claim 29, wherein the at least one characteristic of the second object includes at least one of a size, a shape, a moving direction, or a moving speed of the second object.31.The third device of claim 29 or 30, wherein a corresponding content for the type of the second semantic information includes at least one of a value, a symbol, or a text.32.The third device of any of claims 28-31, wherein the second data is transmitted to the first device via a second interface,wherein the second interface includes an object label field, at least one characteristic field, and a plurality of coordinate fields, andwherein the object label field includes a content for an object label of the second object, the at least one characteristic field includes a content for at least one characteristic of the second object, and the plurality of coordinate fields include two-dimensional (2D) coordinates of a plurality of points of the second object.33.The third device of any of claims 29-32, wherein the at least one characteristic of the second object is predetermined for the second object and stored in the third device.34.The third device of any of claims 28-33, wherein the second coordinate information comprises a set of 2D coordinates of a set of points of the second object.35.The third device of claim 34, wherein the set of points of the second object is predetermined for the second object and stored in the third device.36.The third device of any of claims 28-35, wherein the third device has capability of vision artificial intelligent (AI) processing.37.A method comprising:obtaining, at a first device, first data including first semantic information of a first object and first coordinate information of the first object in a physical space;obtaining second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera;determining, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; andbased on determining that the first object and the second object are the same object, determining calibration data of the camera based on the first coordinate information and the second coordinate information.38.A method comprising:obtaining, at a second device, sensed data of a first object in a physical space by sensing the first object;extracting, from the sensed data, first semantic information of the first object;extracting, from the sensed data, first coordinate information of the first object in the physical space; andtransmitting, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.39.A method comprising:obtaining, at a third device, an image of a physical space captured by a camera;extracting, from the image, second semantic information of a second object in the physical space;extracting, from the image, second coordinate information of the second object in the image; andtransmitting, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.40.An apparatus comprising:means for obtaining, at a first device, first data including first semantic information of a first object and first coordinate information of the first object in a physical space;means for obtaining second data including second semantic information of a second object and second coordinate information of the second object in an image of the physical space captured by a camera;means for determining, based on the first semantic information and the second semantic information, whether the first object and the second object are a same object; andmeans for based on determining that the first object and the second object are the same object, determining calibration data of the camera based on the first coordinate information and the second coordinate information.41.An apparatus comprising:means for obtaining, at a second device, sensed data of a first object in a physical space by sensing the first object;means for extracting, from the sensed data, first semantic information of the first object;means for extracting, from the sensed data, first coordinate information of the first object in the physical space; andmeans for transmitting, to a first device, first data including the first semantic information and the first coordinate information, wherein the first data is used for determining calibration data of a camera which captures an image of the physical space.42.An apparatus comprising:means for obtaining, at a third device, an image of a physical space captured by a camera;means for extracting, from the image, second semantic information of a second object in the physical space;means for extracting, from the image, second coordinate information of the second object in the image; andmeans for transmitting, to a first device, second data including the second semantic information and the second coordinate information, wherein the second data is used for determining calibration data of the camera.43.A non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the method of any of claims 37-39.

Citation Information

Patent Citations

  • Vehicle visual positioning method and system based on two-dimensional semantic map

    CN116295457A

  • Exercise evaluation instrument

    CN116490126A

  • Calibration based on semantic objects

    US20220185331A1

  • Method and apparatus for calibrating extrinsic parameter of a camera

    US20230206500A1