Computer-implemented method for updating a representation of a spatial scene

The method improves the transformation of point clouds into 3D mesh models by distinguishing static and dynamic elements, reducing computational load and enhancing accuracy and stability in real-time applications.

DE102024210705A1Pending Publication Date: 2026-05-07SIEMENS HEALTHINEERS AG
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
SIEMENS HEALTHINEERS AG
Filing Date
2024-11-07
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing methods for transforming point clouds into 3D mesh models in real-time applications, such as collision avoidance systems, face significant computational effort and limited measurement accuracy, complicating the detection of movements and changes in the scene.

Method used

A method that distinguishes between static and dynamic elements in a spatial scene by comparing positional information of recording points with a tolerance threshold, excluding points below the threshold from updates, and defining a spatial environment to detect occlusion, thereby reducing computational load and improving accuracy.

Benefits of technology

Enhances the accuracy and stability of 3D models by reducing redundancy and inefficiency, allowing better detection of genuine changes and reducing computational effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Computer-implemented method for updating a representation of a spatial scene, comprising the following steps: S1) Receiving pixels of the representation of the spatial scene, S2) Recording a plurality of recording points of the scene with a camera, whereby each recording point is assigned position information, S3) Assigning recording points to corresponding pixels in the representation based on the respective position information, S4) Determining any deviation by comparing the position information of a given recording point with the position information of the corresponding image point, S5) Excluding those recording points from the update whose deviation is below a tolerance threshold, characterized by S6) Combining a plurality of pixels of the representation into a partial representation, S7) if the deviation for a subset of the partial representation pixels is above the tolerance threshold and for another subset of the partial representation pixels is below the tolerance threshold, define a spatial environment for a partial representation pixel of the subset with a deviation above the tolerance threshold, S8) Determine whether at least one recording point lies within a volume defined by the specified spatial environment and the camera's image sensor, and S9) Exclude the partial display pixel from the update if the deviation of at least one recording point located in the volume exceeds a motion threshold, S10) Update only the pixels of the display that are not excluded from the update, exclusively based on recording points that are not excluded from the update, S11) Providing the updated display.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a computer-implemented method for updating a representation of a spatial scene, as well as a corresponding system, computer program, and computer program product.

[0002] Computer-implemented methods for updating a representation of a spatial scene are known from the prior art. Known solutions for capturing and updating a spatial scene include recording point clouds with a suitable camera, such as a depth camera, a time-of-flight camera, or a LiDAR (Light Imaging Detection and Ranging) camera. Such cameras allow the acquisition of 3D information only for those elements of a spatial scene that are visible to the camera. Since a complete 3D scene cannot be captured in this way, they are also referred to as 2.5D cameras. When analyzing point clouds of a spatial scene acquired with 2.5D or LiDAR cameras, surfaces and edges can be detected and represented in the form of a 3D mesh model.A recorded scene can be continuously updated to ensure that moving objects and changes are detected, for example, as part of a collision avoidance system. Collision avoidance systems are used, for example, in robot-assisted X-ray angiography systems.

[0003] It is known from the prior art to segment points in a point cloud captured by a camera. This segmentation leads to the subdivision of the points into sub-point clouds. The points of a representation are thus divided into sub-representations. Segmentation can, for example, lead to the subdivision into sub-point clouds or sub-representations that correspond to elements of the spatial scene, such as objects or people. Furthermore, it is known to model such elements or the associated sub-point clouds as 3D models, for example, as 3D mesh models. Through such segmentation and modeling, the points of the segmented sub-point cloud, sub-representation, or 3D model (in particular, a 3D mesh model) of the respective element of the spatial scene are assigned.

[0004] A particular problem in real-time applications, such as collision avoidance, is the considerable computational effort required to transform point clouds into 3D mesh models. This is compounded by the limited measurement accuracy of the image sensors used. Both of these factors complicate the detection of movements and changes in the scene, which should ideally occur as quickly and without latency as possible. Both of these factors are integrated into motion planning to prevent collisions.

[0005] The invention is based on the finding that static elements of a scene cause significant redundancy when transforming point clouds into 3D mesh models: For static elements of the scene that do not move (e.g. walls, tables, devices), the post-processing effort for the transformation into a mesh model is not required.

[0006] The technical problem underlying the invention is to improve and accelerate methods for transforming a point cloud into a 3D mesh model for updating a spatial scene, to increase their accuracy and to reduce redundancy and inefficiency of such methods by improving the distinction between static and dynamic elements of the scene.

[0007] The invention solves this problem by means of a method having the features of the independent claim as well as by means of a system, a computer program and a computer program product having the features of the further independent claims.

[0008] According to the invention, a method comprising the process steps of claim 1 is proposed.

[0009] The computer-implemented method according to the invention for updating a representation of a spatial scene comprises the following steps: S1) Receiving pixels of the representation of the spatial scene, S2) Recording a plurality of recording points of the scene with a camera, whereby each recording point is assigned position information, S3) Assigning recording points to corresponding pixels in the representation based on the respective position information, S4) Determining any deviation by comparing the position information of a given recording point with the position information of the corresponding image point, S5) Excluding those recording points from the update whose deviation is below a tolerance threshold, characterized by S6) Combining a plurality of pixels of the representation into a partial representation, S7) if the deviation for a subset of the partial representation pixels is above the tolerance threshold and for another subset of the partial representation pixels is below the tolerance threshold, define a spatial environment for a partial representation pixel of the subset with a deviation above the tolerance threshold, S8) Determine whether at least one recording point lies within a volume defined by the specified spatial environment and the camera's image sensor, and S9) Exclude the partial display pixel from the update if the deviation of at least one recording point located in the volume exceeds a motion threshold, S10) Update only the pixels of the display that are not excluded from the update, exclusively based on recording points that are not excluded from the update, S11) Providing the updated display.

[0010] The term "scene" is to be understood broadly. In particular, it can refer to an actual spatial scene that can be captured by one or more cameras. A spatial scene can include, for example, people, objects, the floor and walls of a room, etc. The spatial scene can be, for example, a building or a room within a building. Specifically, the spatial scene can be a radiology or interventional room containing diagnostic or therapeutic equipment or an imaging device, as well as patients or medical personnel.

[0011] The term "representation" is to be understood broadly. In particular, it can refer to a computer-readable representation of image data. This can be, for example, a 2D representation in which each pixel corresponds to a single image point. It can also be a 3D representation in which each voxel corresponds to a single image point. A 3D representation is related to a 2D representation of the same object in that depth information is assigned to the image points of the 2D representation. The three dimensions of a 3D representation can, for example, correspond to those of a Cartesian coordinate system. The third dimension can also be depth information that is determined by a 2D camera and assigned to the image points. Such cameras are also called 2.5D cameras because they can capture a complete 2D image, but the third dimension is inherently incomplete.Information about the third dimension of a pixel, the depth information, can only be determined for the pixels of the 2D image. However, pixels that lie on top of each other from the camera's perspective obscure or shadow each other. The camera can therefore only capture the obscuring pixel, not the obscured pixel, which is why a complete 3D image is not created.

[0012] The term "camera" is used here in a broad sense. In particular, it can refer to any camera capable of capturing 2D image points and assigning them additional depth or spatial information. This could include, for example, a standard 2.5D camera, a time-of-flight camera, or a LiDAR camera. A camera comprises an image sensor. It may also include lenses or mirrors to adjust the focal length, aperture, and depth of field, which affect the optical properties and the angle of view.

[0013] Camera points are assigned to pixels in the display being updated if they appear at the same or a similar angle to a pixel in the display from the camera's perspective. A camera point can therefore appear either instead of a pixel in the display or in the immediate vicinity of a pixel in the display's field of view.

[0014] In addition to a viewing angle, each point in the image is assigned depth information. Using this depth information allows us to determine whether a point in the image is located at the same spatial position or in close proximity to a pixel in the image. If a point in the image appears at the same viewing angle as a pixel, it obscures the pixel. In this case, if the depth information of the point in the image matches that of the pixel, it can be assumed that the point is unchanged or a static element of the scene. Otherwise, the element may have moved towards or away from the camera, or it may have been obscured by another element.

[0015] Comparing positional information is a broad term and can refer specifically to comparing the position of image points or recording points. The position is determined by the viewing angle from which the camera captures a recording point and the depth information the camera detects for that point. The positional information is therefore initially stored in the camera's reference coordinate system. If spatial deviations between recording points are to be determined, either the deviations or the positional information must be transformed into a suitable coordinate system, such as a Cartesian coordinate system.

[0016] The tolerance threshold should be chosen to account for the accuracy of the positional information assigned to the recording points by the camera. Apparent deviations in the depth or spatial position of a recording point or pixel, which may be caused by measurement inaccuracies of the camera, should be suppressed by the tolerance threshold. This avoids unnecessary computational effort and improves the informativeness and quality of the image update.

[0017] The method according to the invention enables better detection of unchanged, static elements. Apparent changes, which can be caused, for example, by partial shading of such elements by other, moving elements in the scene, should not lead to unintended changes in the static elements. Therefore, point cloud acquisition points that correspond to such apparent changes are not used to update the static elements by re-transformation of the 3D mesh model. Changes to pixels of a static element that represent only an apparent change to the static element are also not used for re-transformation of the static element. This selective update process significantly reduces the computational load and improves the efficiency of the system.

[0018] The method uses one or more cameras to capture points and determine their positions in space. It compares these points with an existing 3D model of the environment and adjusts the 3D model if it detects that no points are found in the vicinity of the vertices. The method also considers the possibility of shadowing or occlusion: Consider a model vertex generated from a previous point cloud by a camera whose spatial environment no longer contains points in the current point cloud of that camera. The method then checks whether there are any current points near the vertex's projection to the camera. If the distance between the vertex and such a current point is large, the method assumes occlusion by a new object rather than movement of the segmented object.In this case, the vertex is considered still valid, and the system does not perform a new segmentation. This saves computation time and avoids unnecessary changes to the model.

[0019] This method enhances a digital 3D model of a spatial scene by comparing point clouds captured by a camera with the scene itself. It differs from the prior art in that it considers not only the deviation between the camera's points and the image points, but also their spatial relationships. The method detects whether several closely located image points form a common surface or structure of a scene element. If only some, but not all, of these image points show a significant deviation from the camera's points, which can be explained by occlusion of the line of sight, the method assumes that the element's surface or structure is still valid and does not require updating. If the image points have been previously segmented, the method can also check whether the camera's points deviate from the previously segmented image components.If the segmentation was previously modeled as a 3D mesh model with vertices, the method can also check whether the sampling points deviate from the previously created vertices. In this case, the method checks each previously created vertex and expects to find sampling points in its vicinity (the size of which is determined by the resolution of the camera and the distance to it) in the current iteration.

[0020] To determine whether there is vignetting or occlusion, the method checks if there are any image points located within a specific area around the image points affected by deviations. This area extends from the camera to the image points in question. It has the shape of a cone if it forms a circle or sphere around the deviating image points. If such image points exist, and if they also show a significant deviation from the image points above a motion threshold, then the method assumes that vignetting is present. In the case of vignetting, the deviating image points obscure the image points belonging to a static element of the scene. If vignetting is present, the method does not update the surface or structure of the static element despite the deviating image points.

[0021] If such sampling points show no significant deviation, then it is possible that the object has moved. In this case, it makes sense to discard the object model and re-segment this area of ​​sampling points. A reasonable distance threshold for this assumption could be, for example, 50 cm. Only if the distance of a new sampling point to the vertex is sufficiently large, for example, greater than the threshold of 50 cm, is it assumed that the sampling point is not due to a movement of the segmented object, but rather to shadowing or occlusion by a new object.

[0022] A technical advantage of this method is that it increases the accuracy and stability of the 3D model. The method prevents small errors or disturbances in the camera images from leading to large updates in the 3D model. It is better able to distinguish between genuine changes in the real scene and erroneous conclusions due to shadowing or measurement inaccuracies.

[0023] The updated representation is provided by the process to be evaluated, for example, in a collision avoidance system, or to be displayed on a screen, for example.

[0024] A further development of the procedure provides that a partial representation is retained unchanged if no partial representation pixel is allowed to be updated.

[0025] A further development of the procedure stipulates that a partial representation is retained unchanged if only a limited proportion of the partial representation's pixels are permitted for updating, in particular a proportion of no more than half of the partial representation's pixels. This provides a plausible and easily determinable and verifiable criterion for deciding whether a partial representation should be retained or discarded.

[0026] This can also mean, in particular, that a partial representation is retained if not all partial representation pixels are confirmed by current acquisition points. If none of the current acquisition points in the spatial area of ​​the partial representation are to be used for updating, the assignment of the existing partial representation pixels to the partial representation can be retained, and all partial representation pixels do not need to be subjected to a current transformation into the 3D model of the scene.

[0027] A further development of the procedure provides that the environment of the partial representation pixel is defined as circular or spherical.

[0028] If the area surrounding the pixel is defined as circular or spherical, a conical volume results, assuming an isotropic viewing angle across the camera's aperture. A spherical area appears circular from the camera's perspective, while a circular area only appears circular if its axis is aligned with the camera. The important thing is to define an area where a conical volume can be assumed, at least approximately, because a cone has a geometry that is easy to describe. This simple geometric description simplifies the calculation and thus reduces computational effort. As a result, the calculation requires fewer resources and can be performed faster.

[0029] The cone can be designed depending on the camera's field of view; that is, with a wider field of view, the cone can taper more rapidly, and with a narrower field of view, it can taper less rapidly. A camera's field of view depends on the camera's focal length (or the focal length of the camera lens) and the size of the camera's image sensor. The shorter the focal length, the wider the field of view. The larger the image sensor, the wider the field of view. The wider the field of view, the larger the area of ​​a scene captured by the camera. In other words, the field of view indicates the angle at which a bundle of rays traveling from points in a scene to the camera diverges or converges.If a constant or isotropic image angle is assumed across the camera's image aperture, a point or a circular or spherical spatial perimeter is imaged by a cone-shaped beam of light in the camera.

[0030] A further development of the procedure provides that the environment is defined in such a way that it includes all partial representation pixels whose deviation exceeds the tolerance threshold.

[0031] The environment could, for example, be defined by connecting the outermost questionable partial representation pixels with deviations exceeding the tolerance threshold by lines, resulting in an environment in the shape of a polygon or polyhedron. This would allow for a particularly simple and efficient definition of the environment. Alternatively, the environment could be defined by fitting a circle, sphere, ellipse, or ellipsoid to encompass the questionable partial representation pixels. The questionable partial representation pixels should be tightly enclosed, and a margin can be specified as an upper limit for the distance between outer partial representation pixels. Defining the environment as a circle or sphere would be particularly advantageous, as this would result in a conical volume with the advantages explained above.

[0032] A further development of the procedure envisages that the comparison of position information is carried out in such a way that a spatial distance is determined as the deviation.

[0033] Especially when the image display update is primarily based on spatial changes, the distance between the capture points and the resulting pixels is a crucial criterion. Spatial changes of points are mainly reflected in changes in position, which can be measured as distance. It is therefore advisable to limit oneself to distance as the decisive criterion and ignore any other available information. This avoids additional computational effort and prevents any inaccuracies in the potentially available additional information from affecting the result.

[0034] A further development of the procedure provides that the tolerance threshold is 5 cm to 50 cm, preferably 15 cm.

[0035] A further development of the procedure provides that the movement threshold is 35 cm to 65 cm, preferably 50 cm.

[0036] A further development of the method proposes that the motion threshold be set as half the distance of the partial display pixel to the camera. This provides a plausible assumption for the motion threshold, and the motion threshold can be set easily and quickly for each partial display pixel.

[0037] A further development of the procedure envisages that the representation of the spatial scene has the format of a 3D mesh model.

[0038] A 3D mesh model is a commonly used digital surface model that can be used to digitally represent an object or person in 3D. Typically, a 3D mesh model consists of vertices, edges, and faces. The vertices are used as coordinates. The edges connect adjacent vertices. The faces are bounded by the edges and have the shape of polygons. Alternatively, a 3D mesh model can also be created using other methods, for example, to make the model's surfaces smoother or more flexible. However, such alternative models are often more difficult to handle and require more computation. Using a 3D mesh model offers the advantage that these models are widely available and that extensive knowledge and existing applications can be utilized.

[0039] A further development of the procedure provides that, in order to obtain a partial representation, the representation is segmented, whereby the segmentation determines contiguous regions of the representation, whereby pixels belonging to a contiguous region are grouped together to form a partial representation.

[0040] Segmentation is a well-known technique in digital image processing. It involves identifying conceptually related regions within an image, which can represent, for example, individual objects or people. The grouping of neighboring pixels or voxels belonging to such a region is called segmentation. Semantic segmentation, in which an image is divided into segments belonging to specific classes, can be particularly important in spatial scenes. Here, each pixel or data point is assigned a class, such as person or table. Using segmentation to obtain a partial representation has the advantage that segmentation techniques are well-known and widely understood, allowing access to extensive knowledge and existing applications.

[0041] Further features and advantages will become apparent from the dependent patent claims and from the following description of exemplary embodiments with reference to figures.

[0042] The figures show: Fig. 1 Spatial scene Fig. 2. Exemplary representation of the shadow casting problem Fig. 3. Methods according to the invention Fig. 4 pixels when shadows are cast Fig. 5 Circular environment of a shadowed pixel and associated volume

[0043] In Fig. Figure 1 is a schematic example of a spatial scene. A 2.5D camera SC is positioned to capture a spatial scene. The spatial scene includes a robot-assisted X-ray machine 37, a patient table 32, and a display 41. A patient is located on the patient table 32, and people P are present in the spatial scene. The 2.5D camera SC is connected to a control unit PU via a data connection SIG. The data connection SIG uses conventional technology and can operate both wirelessly and via a cable.

[0044] The control unit PU is connected to the X-ray unit 37 and the display 41 via a data connection S. It can be configured to control the X-ray unit 37 and to display X-ray images on the display 41. The control unit PU is also connected to a computer unit 42 via a data connection 26. The computer unit 42 is used for evaluating and analyzing image information captured by the camera SC. It is configured to transform point clouds into a 3D mesh model and to update this model based on continuously acquired additional point clouds.

[0045] Fig. Figure 2 shows an exemplary illustration of a shadow casting problem. The depicted spatial scene includes a table 33 as a static element. As a result of a previous segmentation, the table 33 was recognized as an independent element of the scene and stored as a partial representation of the 3D model of the scene. Through the segmentation, the partial representation pixels 34 were assigned to the partial representation.

[0046] A person P is located between table 33 and camera SC. Person P obscures part of the image of table 33 captured by camera SC. Thus, person P casts a shadow on part of table 33 from the perspective of camera SC, resulting in a shadow problem. The extent of the shadow is outlined by lines 38. The recording points 35 within lines 38 are caused by person P. They are recorded instead of partial representation pixels 34 of table 33. Therefore, the recording by camera SC no longer fully represents the partial representation of table 33.

[0047] In Fig. Figure 3 shows a schematic representation of the method according to the invention.

[0048] In step S1), pixels representing a spatial scene are received.

[0049] In step S2, a multiple new viewpoints from the scene are captured with a camera. The camera assigns depth information or position information to each viewpoint.

[0050] In step S3, for each pixel (which may be assigned to a previous segmentation, such as a table), it is checked whether there are any current recording points in its vicinity. The size of this vicinity can be chosen depending on factors such as the camera's resolution and its distance from the pixel.

[0051] In step S4), any deviation is determined by comparing the position information of a respective recording point with the position information of the respective assigned image point.

[0052] In step S5, acquisition points where the deviation is below a tolerance threshold are excluded from further updates. By excluding such acquisition points from the update, within a tolerance that can, for example, account for measurement inaccuracies, the transformation of new acquisition points is prevented, and the existing image points are retained.

[0053] In step S6), a plurality of pixels of the representation are combined into a partial representation.

[0054] In step S7, it is determined whether the calculated deviation for a subset of the previously used partial representation pixels exceeds the tolerance threshold and for another subset of the partial representation pixels falls below the tolerance threshold. If this is the case, a spatial environment is determined for a previously used partial representation pixel in the subset with a deviation exceeding the tolerance threshold.

[0055] The process checks whether the partial representation, which could be, for example, the surface of an object, is still uniformly covered with sampling points corresponding to the previous pixels. In other words, it checks whether most of the vertices of the partial representation represented by the earlier pixels now have current sampling points in their vicinity. If so, the current sampling points near the partial representation are dropped, and the partial representation is retained as "still current." Only then are the remaining current sampling points examined and, for example, re-segmented. In this way, a partial representation of a static table, i.e., a table mesh model, can persist for a long time, and computation time is saved by retaining the earlier pixels.

[0056] However, this can lead to a problem where the coverage of the previous pixels by the currently captured pixels is no longer uniform. For example, the coverage might show a more or less large gap or hole. In this case, there are many previous pixels in whose vicinity no current pixels can be found. Two possible causes for this are: Firstly, the object represented by the partial representation might no longer be present in the scene; secondly, a new object might have moved in front of the previously existing object.

[0057] If the object is no longer present in the scene, its partial representation and the resulting mesh model should be removed. If a new object has moved in front of the previously existing object, then at least parts of the partial representation would now be obscured. However, the previous object might still be present in the scene. This is particularly likely if the previous object remained static and unchanged in the scene for a relatively long time and / or if new viewpoints are not spatially close to the object or its partial representation. In this case, the partial representation should not be removed initially.

[0058] One problem is how to distinguish between the two cases "object disappeared" and "obscured object." In this context, the further question arises as to how long an object must have remained unchanged to be considered static, and / or how far current image points must be from the previous image points of the previous object to be considered "not spatially close." The following two procedural steps examine whether one or more current image points obscure previous image points and are not spatially close to them.

[0059] In step S8, it is determined whether at least one current recording point lies within a volume defined by the specified spatial environment and the camera's image sensor. The extent of the spatial environment can decrease relative to the camera's viewing angle, or it can change relative to the size ratio of the specified spatial environment to the camera's image sensor.

[0060] The process checks whether, for a pixel whose surroundings are free of new sampling points, any sampling points can be found within the vicinity of the projection line to the camera. If not, the pixel is considered obsolete. Otherwise, in the following step (S9), it is checked whether these sampling points associated with the pixel have a minimum distance to the pixel. If so, shadowing is assumed, and the pixel is accepted as still valid despite the surroundings being devoid of sampling points.

[0061] If not, the possibility that the partial representation has changed at this point and a re-segmentation is appropriate must be considered.

[0062] With a typical 2.5D camera, its resolution or measurement accuracy can be used to determine the spatial environment. Measurement accuracy tolerances can be on the order of approximately 1.5 mm at a distance of 50 cm and 5 cm at a distance of 5 m. Therefore, the greater the distance of a recording point from the camera, the greater the resolution-related blurring. The point blurs in space according to a probability density. This blurring is accounted for by the spatial environment and can be approximately described as a sphere or circle. Its radius should then be chosen depending on the camera's measurement accuracy. Additional optical parameters can also be considered, such as field of view, angle of view, fisheye imaging geometries, and distance to the camera.

[0063] A volume is defined between the spatial environment of the image point and the camera lens. If a sphere or circle is chosen as the spatial environment, the volume is a cone. If new image points lie within this volume, it can be assumed that the image points of the previous partial representation have been obscured by new image points, or that the previously existing object has been obscured by a new object. The previous image point, on the basis of which the volume is defined, could therefore be considered still, but obscured. On the other hand, it is of course possible that the previously existing object has simply moved closer to the camera, meaning that the current image points still belong to the partial representation.

[0064] An occlusion or temporary occlusion of the previous object should only be assumed if the distance between the occluding camera points and the corresponding occluded image points is large, especially in relation to the size of the occluded object. In such a case, the assumption that the occluding camera points can actually be explained by a deformation of the supposedly occluded object appears implausible.

[0065] If recordings from other cameras are available, these may further support the assumption that the object was concealed.

[0066] Furthermore, occlusion should be assumed if the extent of the occlusion increases, at least initially. An initially increasing extent of occlusion would correspond to a newly captured object gradually moving in front of the previously existing object over time.

[0067] Furthermore, occlusion is assumed if the current recording points within the defined volume are significantly closer to the camera than to the previously existing object. In this case, a new object would typically maintain a certain minimum distance from an existing object to avoid colliding with it during its movement.

[0068] In step S9, the partial display pixel is excluded from the update if the deviation from at least one acquisition point located within the volume, particularly a conical volume, exceeds a motion threshold. This prevents partial display pixels that are merely shaded or obscured from being replaced by acquisition points resulting from the shaded or obscured position.

[0069] In step S10), the update is performed exclusively on existing pixels of the display that are not excluded from the update, and exclusively on the basis of recording points that are not excluded from the update.

[0070] If too many of the existing pixels in a partial representation need to be updated, the previously created partial representation is removed from the model. The new acquisition points are then retained and used for segmenting new objects.

[0071] Conversely, if a large proportion of the existing pixels in the current partial representation are excluded from the update and are therefore treated as still current, all new acquisition points in the vicinity of the partial representation are removed instead. This simplifies the point cloud formed from the acquisition points. The image within the area of ​​the partial representation is then retained unchanged.

[0072] In step S11, the updated display of the control unit PU is provided. The control unit PU can, for example, display the updated display on display 41 or use it in a collision avoidance calculation.

[0073] Based on Fig. Section 4 explains the altered and unchanged partial representation pixels of a partial representation when shadows are cast. As before, using the example of... Fig. As explained in Figure 2, the recording points 35 of person P obscure previously recorded partial representation image points 36 of table 33. According to the method of the invention, it is determined in this situation that some partial representation image points of the static element table 33 have remained unchanged, while other partial representation image points 36 of table 33 have been altered or are no longer present. The aim of the method is now to determine whether the partial representation of table 33 should be discarded as part of an update. In this case, all recording points that are in the spatial area of ​​the partial representation of table 33 would have to be analyzed and segmented in order to update the 3D model of the scene.Conversely, if it can be determined that the partial representation of table 33 should remain, all recording points in the spatial area of ​​the partial representation of table 33 can be assumed to be unchanged and do not need to be re-analyzed, segmented and transformed into the 3D model.

[0074] Based on Fig.Figure 5 explains, using a circular environment of a shadowed static partial representation pixel and an environment defined in relation to this, how the inventive method determines whether the partial representation of the table 33 should be assumed to be unchanged. The aim of the analysis is to determine whether a part of the table 33 is obscured by a change in the scene. In the illustration, this change is represented by person P. As an example, the previous partial representation pixel 40 of the table 33 is considered, which is either not present or significantly altered in the current image from camera SC. A circular or spherical spatial environment 43 is defined around the previous partial representation pixel 40. Then, the volume between the environment 43 and the camera or the camera's image sensor 39 is considered.If there are recording points within this volume, the deviation of the spatial position of these recording points from the previous partial representation pixel 40 is determined. If the deviation of the spatial position of such recording points exceeds a movement threshold, it can be assumed that an element has moved between table 33 and the camera. In the representation, person P has moved between table 33 and the camera. In this case, such recording points are not used to update the partial representation of table 33. Instead, the current partial representation pixels belonging to table 33 are assumed to be unchanged and are not subjected to any analysis or transformation into the 3D model.

[0075] The preceding description is intended to include persons of male, female or other gender identities, regardless of the grammatical gender of a particular term.

Claims

[1] Computer-implemented method for updating a representation of a spatial scene, comprising the steps: S1) Receiving pixels of the representation of the spatial scene, S2) Recording a plurality of recording points of the scene with a camera, whereby each recording point is assigned position information, S3) Assigning recording points to corresponding pixels in the representation based on the respective position information, S4) Determining any deviation by comparing the position information of a given recording point with the position information of the corresponding image point, S5) Excluding those recording points from the update whose deviation is below a tolerance threshold, characterized by S6) Combining a plurality of pixels of the representation into a partial representation, S7) if the deviation for a subset of the partial representation pixels is above the tolerance threshold and for another subset of the partial representation pixels is below the tolerance threshold, define a spatial environment for a partial representation pixel of the subset with a deviation above the tolerance threshold, S8) Determine whether at least one recording point lies within a volume defined by the specified spatial environment and the camera's image sensor, and S9) Exclude the partial display pixel from the update if the deviation of at least one recording point located in the volume exceeds a motion threshold, S10) Update only the pixels of the display that are not excluded from the update, exclusively based on recording points that are not excluded from the update, S11) Providing the updated display. [2] Method according to claim 1, wherein a partial representation is retained unchanged if no partial representation pixel is allowed to be updated. [3] Method according to claim 1, wherein a partial representation is retained unchanged if at most a limited proportion of the partial representation pixels is allowed to be updated, in particular a proportion of at most half of the partial representation pixels. [4] Method according to one of the preceding claims, wherein the environment of the partial representation pixel is defined as circular or spherical. [5] Method according to any of the preceding claims, wherein the environment is defined such that it includes all partial representation pixels whose deviation is above the tolerance threshold. [6] Method according to one of the preceding claims, wherein the comparison of position information is carried out in such a way that a spatial distance is determined as the deviation. [7] Method according to any of the preceding claims, wherein the tolerance threshold is 5 cm to 50 cm, preferably 15 cm. [8] Method according to one of the preceding claims, wherein the movement threshold is 35 cm to 65 cm, preferably 50 cm. [9] Method according to any one of the preceding claims 1 to 7, wherein the motion threshold is set as half the distance of the partial display pixel to the camera. [10] Method according to any of the preceding claims, wherein the representation of the spatial scene has the format of a 3D mesh model. [11] Method according to one of the preceding claims, wherein to obtain a partial representation the representation is segmented, wherein by the segmentation connected regions of the representation are determined, wherein pixels belonging to a connected region are grouped together to form a partial representation. [12] Provisioning unit for providing an updated representation of a spatial scene, comprising a computing unit for performing the steps of a procedure according to one of the preceding procedure claims. [13] Computer program which causes a computing device to carry out the steps of a procedure according to one of the preceding procedure claims when it is executed on the computing device. [14] Electronically readable data carrier on which a computer program is stored according to the preceding computer program claim.

Citation Information

Patent Citations

  • Live updates in a networked remote collaboration session

    WO2022178238A1