Method for generating enrichment data for enriching a view of an environment, and corresponding rendering device, server equipment, system and computer program
Patent Information
- Application Number
- EP2024703325
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-02
- Filing Date
- 2024-02-01
- Publication Date
- 2025-12-10
AI Technical Summary
In cluttered environments, especially industrial settings, it is difficult for operators equipped with rendering devices to distinguish the presence and location of objects or other individuals, leading to increased accident risks due to reduced visibility of potentially hazardous machinery or coworkers.
A method for generating enrichment data to enhance the view of a rendering device, which includes obtaining location information for objects of interest and generating data representing their positions and obscured parts, using visual, audio, or haptic indications to improve visibility, even if they are hidden by obstacles.
This approach enhances the user's awareness of objects and individuals within the environment, reducing the risk of accidents by providing a more comprehensive and accurate representation of the surroundings, even if they are not directly visible.
Smart Images

Figure EP2024052481_08082024_PF_FP
Abstract
Description
[0001] Description
[0002] Title of the invention: Method for generating data for enriching a view of an environment, corresponding rendering device, server equipment, system and computer program
[0003] Technical field of the invention
[0004] This application concerns the field of virtual and / or augmented reality.
[0005] In particular, it concerns the generation of data for enriching a rendering of a real environment in which there is at least one user equipped with a rendering device mounted on their head, such as virtual and / or augmented reality glasses or headsets.
[0006] It applies in particular, but not exclusively, to an operator's working environment cluttered by the presence of objects and in which the visibility of other operators is reduced.
[0007] Prior art
[0008] In a cluttered environment and / or one with reduced visibility, it is difficult for an individual to distinguish the presence of certain objects or other individuals. In an industrial environment, this lack of visibility of potentially dangerous machine tools or other personnel present in a building, leads to an increased risk of accident for an operator. For example, when two operators are required to work together on the maintenance of the same machine tool, they need to know what the other is doing and where they are at all times.
[0009] Today, we know of communication solutions that allow technicians equipped with mobile terminals, such as cell phones or walkie-talkies, to communicate with each other to keep each other informed of each other's actions. However, these solutions are insufficient to prevent certain accidents.
[0010] This request improves the situation.
[0011] Statement of the invention
[0012] To this end, the present application proposes a method for generating enrichment data for at least one view of at least one part of an environment from at least one first point of view, said view being intended to be rendered at least partially by at least one rendering device, said method comprising:
[0013] - obtaining location information of at least one object of interest (in said environment,
[0014] - a generation of enrichment data (of said view, at least from the location information obtained, said enrichment data comprising at least data representative of a position of said object of interest relative to said first point of view and data representative of an occulted part of said object of interest.
[0015] The present application thus proposes a completely new and inventive approach allowing an enriched visualization of a part of a real environment in which there is at least one object of interest. This approach enriches the view of the part of the environment by adding to it in particular information on the position of the object of interest and information on the hidden part of this object. In particular, the present application proposes to improve the visibility of a real environment for a user equipped with a rendering device mounted on his head and through which he views this environment, which is not possible, for example, with solutions based solely on voice.
[0016] These include, for example, visual, audio, or haptic indications that complement the rendering of the view when it is displayed on the rendering device. The view thus enriched can, for example, help a user of the rendering device to monitor the movement of one or more objects of interest in the environment and to prevent risky situations in terms of accidents.
[0017] In this way, the present application can contribute to helping a user of a rendering device to become aware, on the part of the enriched environment rendered by the rendering device, of the presence and position of objects of interest present in the real environment, even if they are not visible, or partially, in reality from his own point of view (they can be hidden by a physical object such as a partition, furniture, a machine, a pallet of stacked boxes, etc.).
[0018] It is understood that the term environment here refers to a real place, closed or open, for example a factory, a warehouse, a residential building, a shopping center, etc., in which all types of objects of interest in the broad sense may be located. The term view of the environment rendered by the rendering device is used for any representation of this real environment displayed on the screen of the user's rendering device, whether virtual, mixed or augmented. The term object of interest here refers to any type of object likely to be of interest to a user of the rendering device, i.e. a construction machine, a robot, an animal, or an individual. In this context, a human being or an animal is considered as an "object" of detection or location by a suitable technique using sensors.In at least some embodiments, the object of interest is likely to move, for example to move in the environment and / or to perform movements, for example rotation and / or twisting on itself, that is to say without necessarily changing position. It may be at least partially visible from the first point of view, or even totally obscured by physical obstacles (furniture, machines, partitions, etc.) and / or by lack of light.
[0019] In at least some embodiments, the location information is obtained several times, for example periodically, for example once per second, and the part of the environment rendered by the rendering device is updated when new location information is available, so as to present at a time to a user of the rendering device an enriched view consistent with the reality of the ground.
[0020] According to the present application, the rendering device may correspond to any type of user terminal equipped with suitable communication and rendering means.
[0021] For example, in at least some embodiments, a virtual view of the environment comprising enrichment data generated by the invention is superimposed as an additional layer on that of the view of the part of the real environment displayed by the rendering device, which may itself be real, virtual or mixed. According to a first example, the rendering device is a fixed terminal, of the workstation or surveillance type from which a user, or supervisor, can monitor at least part of the real environment and / or the activity taking place there. It may be included in the environment or remote. In this case, the view of the part of the real environment rendered by the device is obtained from images captured by one or more cameras arranged in different positions of the real environment. The point of view is for example that of one of the cameras in question.
[0022] According to at least one embodiment, the at least one rendering device is a device for supervising said environment, and / or a plurality of mobile rendering devices. The rendering device may be, for example, a device for overall supervision of said environment, for example, terminal equipment such as a fixed or mobile computer, which is not necessarily located in the environment and may be remote. The user is, for example, a supervisor responsible for monitoring the activity within the environment and in particular ensuring the safety of the individuals therein. In this case, the point of view may be the point of view of a camera whose video data was used to produce the view produced by the rendering device or, more generally, the origin of a reference frame of the environment. Additionally or alternatively, the enriched view(s) may be rendered by a plurality of mobile rendering devices.
[0023] According to at least one embodiment, at least one rendering device among the plurality of mobile rendering devices is a mobile terminal equipment of said environment and comprises at least one camera configured to capture video data from the first viewpoint, the rendered view being generated at least partially from video data of said camera.
[0024] According to at least one embodiment, the rendering device is a mobile device such as a head-mounted device (HMD), such as virtual and / or augmented reality glasses or headsets. In this case, the user of this device is located inside the real environment and can move around in it. For example, this is a maintenance operator responsible for working on machine tools present in the environment. In this case, the rendering device is present in the environment and the point of view is that of the user's rendering device.
[0025] In at least one embodiment, the method for generating enrichment data further comprises generating the enrichment data of a virtual graphical representation of said environment, intended to be rendered by said device for supervising said environment, thus enabling, for example, a user, for example a supervisor, to carry out overall supervision of said environment.
[0026] According to at least one embodiment, the method comprises:
[0027] - a detection of said at least one object of interest at least in said view,
[0028] - obtaining information describing said at least one detected object of interest, and
[0029] - use of the description information to generate at least some of the view enrichment data.
[0030] According to one example, the object of interest is a physical object at least partially visible in the view. For example, it is a part of a machine tool, some faces of which are not visible because they are masked by other constituent parts of the machine tool. Once the object has been detected and recognized, description information, such as information on shape, size, structure, material, etc., is for example obtained and used to enrich the view. The enrichment data may contribute, in at least some embodiments, to enhancing the physical object in the view displayed to the user. For example, its contours appear highlighted, a noise produced by the object of interest is amplified, textual description information (name, reference, etc.) or graphics are embedded superimposed on the real view, indicating where exactly the part in front of him is located and what part it is.
[0031] This description information is stored in a memory, for example.
[0032] In at least some embodiments, video data of the portion of the environment captured by cameras of the environment from distinct viewpoints may be used to detect and recognize the object of interest.
[0033] In another example, the object of interest is an individual present in the environment. It is assumed that at least one part of its body is detectable on the view or on other video sequences captured by other cameras. When this body part has been detected and recognized as such, information describing the human body (anatomical data, positions of key points such as knees, ankles, head, hands, etc.) can be used to enrich the view displayed by the rendering device at the level of the individual's position.
[0034] In yet another example, the object of interest is a part of the body of the user of the rendering device, for example, his arms.
[0035] In yet another example, the object of interest is an electromechanical device, or a part of such a device, for example an articulated handling arm.
[0036] According to yet another aspect of the invention, the method comprises: obtaining a graphical representation of the occluded part of the object of interest in the view of the part of the environment rendered by the rendering device at least as a function of the position of the object of interest relative to the viewpoint and the description information; the enrichment data of the occluded part comprise the graphical representation obtained.
[0037] For example, the description information includes a manipulable 3D graphical representation model of the object of interest, which can be displayed with the view (such as on the margin of the view).
[0038] For example, the 3D graphic representation of the object of interest has been previously constructed and stored in memory. The user can view it from different angles or even obtain additional information about this object and, for example, view hidden parts of this part. If it is a part of a machine, he could also access information relating to interactions of this part with other parts of the machine, for example included in a description manual accessible via his rendering device.
[0039] An advantage, in at least some embodiments, is that rendering this graphical representation can help the user of the rendering device to visualize the occluded portion of the object of interest and therefore have a more complete perception of this object of interest.
[0040] When the object of interest is a physical object, view enrichment can, for example, help a maintenance operator understand how a part is arranged with one or more other parts of a machine tool.
[0041] In the case where the object of interest is an individual for example, the determined occluded part comprises a part of the individual's body and the graphical representation of this occluded part may comprise, at least in certain embodiments, an avatar of the part of the body of this individual. An advantage, at least in certain embodiments, may be to help the user of the rendering device to become aware of the posture and / or a movement of this individual and therefore to perceive possible risks of collision or even accident. For example, this avatar is obtained from positions of the occluded part and an inverse kinematics technique which defines the members as non-deformable solids of given length connected to each other by joints and subjected to given maximum stresses.One advantage, at least in some embodiments, may be to help produce a relatively lifelike avatar in the sense that the individual's body will adopt realistic postures.
[0042] When the object of interest is a part of the body of the user of the HMD rendering device, for example, their arms, the view may be enhanced, at least in some embodiments, from the arms of an avatar to the positions of the arms in the view. In this way, the user can see, for example, where their arms are located even if they are partly obscured by parts of the machine tool they are repairing.
[0043] Of course, the object of interest can also be a robot or an articulated handling arm, part of which is occluded. A graphical representation of the occluded part of this robot can be, as previously described, inserted into the generated enrichment data to increase the perception of the part of the environment seen by the user.
[0044] According to at least one embodiment, the location information of said at least one object of interest is obtained from at least one position sensor placed near and / or on said at least one object.
[0045] Such an embodiment may assist in locating the object of interest, whether or not it is visible in the view, for example when it is outside the field of view of the camera(s) that captured the video data from which the view was generated.
[0046] In at least some embodiments, such a position sensor may also help identify the object of interest. For example, it is a radio-frequency identification (RFID) device configured to emit a radio signal in response to a radio signal received from at least one tag reading device placed in the environment, the latter being configured to read at least one piece of identification information contained in the tag and determine a relative position of the object of interest with respect to the reading device placed in the environment.
[0047] For example, when the object of interest is a physical object, typically a part of a machine tool, a robot, a trolley, a piece of furniture, etc., it can generally be equipped with means of identification, in particular for inventory purposes. When the object of interest is another rendering device, for example the HMD device of a user who is in the environment, it can also be equipped with at least one position sensor, as well as other sensors, for example cameras, an inertial system, etc.
[0048] The location of an object of interest can also be obtained by other means (in particular in the case where the object of interest is not equipped with a position sensor), for example by processing images from video sequences captured by cameras placed in the environment and / or carried by the rendering device. This is the case, for example, of an individual who is not wearing an HMD rendering device or of physical objects which are not part of the inventory (unlike those listed above) but which could constitute dangerous obstacles. According to at least one embodiment, the method further comprises:
[0049] - obtaining a map of at least part of the environment, comprising at least location information for a plurality of physical objects present in the environment; and in that the generation of the enrichment data comprises at least the insertion of data representative of positions of said physical objects relative to the first point of view. Thus, in at least certain embodiments, this map of the environment provides information on the positions of physical objects, for example immobile ones, present in the environment which may constitute potential obstacles to the visibility of other users. They correspond, for example, to partitions, doors, furniture, machines, tools, boxes, trolleys, and more generally any type of physical object. They do not necessarily correspond to a point of interest for a user.In at least some embodiments, this mapping may also indicate at least one danger zone.
[0050] Alternatively, other geolocation solutions are possible to locate (for example precisely) objects of interest and people who may be equipped with location devices, based on a satellite positioning solution, for example of the GPS (Global Positioning System) or Glonass (Global Navigation Satellite System) type outdoors, or a geolocation solution using ultra-wideband radio technology or UWB (Ultra Wide Band) or a direction detection technique such as BDF (Bluetooth Direction Finding) for indoors, or not be equipped with such devices. In this case, geolocation is implemented using sensors present in the environment, such as RGB-D cameras, for example, and whose position is known in a reference frame of the environment.
[0051] Regarding the layout of the environment, in terms of partitions, doors, aisles, a map can be constructed from 3D video data captured by one or more users as they walk around the premises, systematically and exhaustively (initial construction) or as they intervene (on-going construction). Matching video images acquired by different users at different times can help detect areas that are found in the different images and then label them using a set of labels depending on the type of environment considered (for example, the set of labels can be adapted to a site and be different for a machine maintenance site, for a production line, for a goods storage warehouse, etc.).
[0052] According to at least one example, the enrichment data includes, in addition to the positions of the physical objects, identification information of these physical objects obtained from the mapping or image matching previously discussed.
[0053] Knowledge of this mapping can for example be used to help obtain (for example determine) the occluded part of an object of interest in the view displayed by the rendering device.
[0054] According to at least one embodiment, the detection of said at least one object of interest implements an artificial intelligence module configured to use at least one previously learned model and associating said 3D video data with position and classification information of said object of interest.
[0055] An advantage of using an artificial intelligence module may be, at least in certain embodiments, to take advantage of its processing power and integration of a quantity of data which may be very large and its capacity, once trained, to help produce reliable, precise and reproducible results.
[0056] In some embodiments, such an artificial intelligence module may make it possible to obtain position information of one or more key points of the object in the view and at least one class or type of the object. This may be particularly interesting, at least in some embodiments, in particular for objects of interest that are not equipped with position sensors (such as RFID sensors). For example, when the object of interest is a physical object, the artificial intelligence module may produce as output the class of the part and its position in the environment. From the class obtained, it is possible, for example, to retrieve the corresponding description information from memory.
[0057] When the object of interest is an individual, the artificial intelligence module can produce, at least in some embodiments, the position of key anatomical points and recognize that it is a human being.
[0058] According to at least one embodiment, the method comprises training (for example prior) of said model from a training database comprising a plurality of pairs of data, a said pair associating 3D video data with at least one piece of position information of an object of interest and one piece of object class information.
[0059] For physical objects, a learning base can be built by manually labeling video sequences of items from a list of objects of interest that the user is likely to encounter when present in the environment. For example, this is a catalog of parts of an industrial machine.
[0060] For humans, the training base can be a database, for example, a public database, presenting labeled images or videos of individuals in different positions. The labels indicate, for example, the different parts of the body (limbs, head, trunk).
[0061] In some embodiments, the learning is updated "continuously" using data collected over time and user evaluations of the quality and accuracy of the generated enrichment data (and in particular location and identification information of objects / users / individuals).
[0062] The present application also relates to a data device for enriching at least one view of at least one part of an environment from at least a first point of view, said view being intended to be rendered at least partially by at least one rendering device, said device being configured to implement:
[0063] - obtaining location information of at least one object of interest in said environment,
[0064] - a generation of enrichment data for said view, at least from the location information obtained, said enrichment data comprising at least data representative of a position of said object of interest relative to said first point of view and data representative of an obscured part of said object of interest.
[0065] In certain embodiments, the device according to the present application can generate enrichment data for several rendering devices, including the rendering devices (for example of the HMD type) of several users present in the environment and / or for a rendering device of a supervising user, called a supervision device, which is not necessarily present in the environment.
[0066] Thus, a supervisor can, for example, visualize an enriched global view of the environment on which the positions of the users present are indicated. In this way, the supervisor can, for example, easily locate them (for example, locate them all at a glance). In the event of an accident, such embodiments can help the supervisor to understand the situation more quickly and therefore manage the emergency. The position of an injured user can, for example, be signaled very quickly (or even immediately) to other users located nearby. Knowing and visualizing their position, they can, for example, intervene without delay. The rendering devices can, in certain embodiments, be equipped with at least one vital signs sensor (cardiac, fall, etc.) to automate accident detection.
[0067] In certain embodiments of the invention, the enrichment data generation device is configured to generate enrichment data for a virtual graphical representation of said environment, intended to be rendered by said supervision device (ET, DISP) of said environment.
[0068] According to at least one embodiment of the invention, said device is configured to implement the steps of the method for generating enrichment data of at least one view of at least one part of an environment as described previously.
[0069] According to at least one embodiment of the invention, said device is integrated into server equipment connected via a communication network to at least one rendering device adapted to produce at least a partial rendering of a view of at least one environment from at least a first point of view.
[0070] In certain embodiments of the invention, said server equipment comprises a module configured to generate a virtual graphical representation of the environment intended to be rendered by said device for supervising said environment. For example, this virtual graphical representation can be enriched with enrichment data generated by said device for generating enrichment data.
[0071] The server equipment and the device have at least the same advantages as those conferred by the aforementioned generation method.
[0072] It can be placed in the environment or outside this environment, for example in a remote communication network for example organized according to a cloud architecture (from the English, "Cloud Computing"). In particular, the communication network can be adapted to allow communications between the server equipment and the rendering device(s) having a reasonable latency (for example without latency), that is to say in quasi real time. In certain embodiments, the invention makes it possible to generate enrichment data of several renderings of the real environment, for example those of the users of HMD devices who evolve (move, work, etc.) in the real environment and that of the supervisor.
[0073] According to at least one embodiment of the present application, the aforementioned server equipment can itself be integrated into a system for supervising an environment comprising at least one rendering device adapted to produce at least a partial rendering of a view of the environment from at least a first point of view and at least one sensor configured to collect at least one location information of at least one object of interest present in the environment.
[0074] In some embodiments, said system may comprise several rendering devices:
[0075] - a fixed rendering device intended to be used by a supervisor in charge of the activity and safety of the environment, to visualize one or more views of parts of the environment according to one or more points of view corresponding to video cameras placed in the environment;
[0076] - one or more mobile rendering devices of user(s) present in the environment through which each person can, for example, view the part of the environment visible from their own point of view.
[0077] According to at least one embodiment of the present application, said supervision system comprises at least one mobile rendering device provided with an electronic tag and at least one location device configured to implement the reception of a signal comprising identification information from said electronic tag, the determination of a location of said electronic tag based on the identification information received, and the transmission of said location to said enrichment data generation device.
[0078] The location device provides the enrichment data generation device with the location(s) of the mobile rendering devices equipped with electronic tags, which may contribute to a reduction in the computational load of said enrichment data generation device.
[0079] In at least one embodiment, said location device is configured to emit, prior to receiving the signal comprising identification information from said electronic tag, a request signal to said electronic tag, such an embodiment allows the use of passive RFID (Radio Frequency Identification) type electronic tags.
[0080] Alternatively or cumulatively, electronic labels can emit (for example regularly) their respective identification data.
[0081] The invention also relates to a computer program product comprising program code instructions for implementing the method of generating enrichment data as described above, in any of its embodiments, when executed by a processor.
[0082] A program may use any programming language, and may be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0083] The invention finally relates to a recording medium readable by a computer on which is recorded the computer program comprising program code instructions for executing the steps of the method according to the invention as described above in any one of its embodiments.
[0084] Such a recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage medium, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording medium, for example a mobile medium (memory card) or a hard disk or an SSD.
[0085] On the other hand, such a recording medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means, so that the computer program contained therein is remotely executable. The program according to the invention may in particular be downloaded over a network, for example the Internet.
[0086] Alternatively, the recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to perform or to be used in performing the above method.
[0087] According to an exemplary embodiment, the present technique is implemented by means of software and / or hardware components. In this regard, the term "module" may correspond in this document to a software component, a hardware component or a set of hardware and software components.
[0088] A software component corresponds to one or more computer programs, one or more sub-programs of a program, or more generally to any element of a program or software capable of implementing a function or a set of functions, as described below for the module concerned. Such a software component is executed by a data processor of a physical entity (terminal, server, gateway, set-top-box, router, etc.) and is likely to access the hardware resources of this physical entity (memories, recording media, communication buses, electronic input / output cards, user interfaces, etc.). Subsequently, resources are understood to mean all sets of hardware and / or software elements supporting a function or a service, whether individual or combined.
[0089] Similarly, a hardware component is any element of a hardware assembly capable of implementing a function or set of functions, as described below for the module concerned. It may be a programmable hardware component or one with an integrated processor for running software, for example an integrated circuit, a smart card, a memory card, an electronic card for running firmware, etc.
[0090] Each component of the system described above can of course implement its own software or hardware modules.
[0091] The different embodiments mentioned above can be combined with each other for the implementation of the present technique.
[0092] Brief description of the drawings
[0093] Other characteristics and advantages of the present invention will emerge from the description given below, with reference to the appended drawings which illustrate an exemplary embodiment thereof without any limiting character. In the figures:
[0094] • [Fig. 1] schematically illustrates an example of architecture of a system for supervising an environment, comprising at least one device for generating data for enriching at least one view of the environment, according to at least one embodiment of the invention;
[0095] • [Fig. 2] describes in the form of a flowchart the steps of a method for generating data for enriching a view of an environment implemented by the device according to at least one embodiment of the invention;
[0096] • [Fig. 3] schematically illustrates a graphic representation of an individual's body of which at least one part has been detected in the view, according to at least one embodiment of the invention;
[0097] • [Fig. 4A] and [Fig. 4B] schematically illustrate artificial intelligence modules each implementing a previously learned model to detect objects of interest, according to at least one embodiment of the invention;
[0098] • [Fig. 5] schematically illustrates a view enriched with images of a scene in which the parts of an individual's body, hidden by a wall, are represented by the corresponding parts of an avatar of this individual, according to at least one embodiment of the invention;
[0099] • [Fig. 6] schematically illustrates the data flows between the equipment and devices of the system according to at least one embodiment of the invention;
[0100] • [Fig. 7] illustrates an example of hardware structure of a device for generating data for enriching a view of an environment according to at least one embodiment of the invention.
[0101] Description of the invention
[0102] The present application relates to a device for generating data for enriching a part of a real environment rendered by an augmented reality (AR) and / or virtual reality (VR) and / or mixed reality (MR) rendering device. This is, for example, a rendering device of the HMD (Head Mounted Display) type intended to be mounted on a user's head, and through which the user can view a real, augmented environment. This may be, for example, augmented reality glasses or a virtual reality headset. Haptic effects (for example, vibrations) and / or audio may also be rendered.
[0103] Of course, the invention is not limited to HMD devices, but applies to any terminal equipment, for example a smartphone, a computer, a tablet, etc., provided that it is suitable for producing a virtual, augmented or mixed reality rendering.
[0104] As a reminder, augmented reality (AR) is a technology that allows virtual elements (for example, in 3D and / or in real time) to be integrated into a real environment. The principle is to combine the virtual and the real and give the illusion of perfect integration to the user. According to this principle, augmented reality glasses are equipped with transparent lenses that allow the user to see with their own eyes a view of the part of the environment located in their field of vision (real view). Virtual elements are superimposed on this real view and enrich the user's experience. Virtual reality (VR) is a technology that allows a virtual scene to be digitally simulated by computer, for example at least by sight. It allows, for example, the user of a virtual reality headset to experience immersion in an artificial world.A 3D video of the part of the real environment located in the user's field of vision is captured in real time by at least one camera placed on the headset and displayed on the headset's screen instead of what the user would see with their own eyes. Mixed reality MR is a technology that captures the environment around a user, transcribes it in 3D and allows 3D objects to be inserted and manipulated, as in the case of VR. According to these principles, additional virtual elements are integrated into the 3D video to enrich the view of the environment.
[0105] The principle of the invention consists in particular in obtaining location information of at least one object of interest likely to move in the real environment and in using it to generate enrichment data, for example audio, haptic and / or visual, of a portion of the environment rendered by the rendering device. Such enrichment data may in particular comprise, in the case of a rendering device of the HMD type, visual indicators of relative positions of the object(s) of interest with respect to the rendering device (or its user) in the real environment, whether they are visible from the point of view of this user or masked (i.e. hidden) by the presence of objects.
[0106] Here, the term “object of interest” refers to an object likely to move in the real environment of the rendering device, or of its user, such as another electronic or mechanical device (a robot or construction machine for example) or a living being (an animal or an individual). In the latter case, this individual may or may not be equipped with a rendering device, different or similar to the rendering device of the present application. This other rendering device may, for example, be an HMD-type rendering device worn by this individual. If so, this second rendering device makes it possible to display a view of part of the environment from its own point of view on a screen of its HMD device placed in front of its eyes.
[0107] The present application applies to both an indoor and outdoor environment. It may be particularly interesting for a work environment, in which several operators intervene on physical objects, for example machine tools, which are potentially dangerous. It can indeed help, at least in certain embodiments, to signal at any time to the user of the rendering device the position of objects of interest located near him, such as individuals, whether or not users of rendering devices or robots and can thus help to prevent (or limit) the occurrence of work accidents.
[0108] In relation to Figure 1, a simplified example of the architecture of a supervision system S of an ENV environment according to an embodiment of the invention is presented by way of illustration. It is assumed here that it is an indoor environment, for example a factory building or a warehouse or any type of premises. Inside this ENV environment, a first user OP1 equipped with a rendering device HMD1 is considered. This rendering device is a head-mounted HMD1 rendering device HMD1 in the illustrated example. The user is viewing a view of the ENV environment on the screen of the rendering device HMD1 placed in front of his eyes. For example, it was obtained from one or more video cameras fixed on the HMD device and corresponds to his field of vision.
[0109] At least one object of interest is likely to be present and mobile in the ENV environment. In the illustrated example, this is a plurality of individuals OP2 to OP4. Optionally, at least one of these individuals may also be equipped with a rendering device identical to or different from the rendering device of the user OP1. It is assumed in the following that the individuals are also equipped with rendering devices HMD2 to HMD4, for example HMD-type devices mounted on their heads.
[0110] It is assumed that the environment ENV also comprises a plurality of physical objects, arranged in different locations, which can reduce the visibility of the user OP1 depending on his position, his size and / or his orientation in particular. These are, for example, generally immobile physical objects, participating in an arrangement of the environment, for example a partition OBJ1, a door, a large piece of furniture such as the machine tool OBJ2, a cabinet, a production line (not shown), generally fixed. They can also be physical objects, for example smaller in size, likely to be moved, such as for example a trolley OBJ3, a box OBJ4, a chair or a table (not shown), etc. They can also be objects of interest (not shown) that the user has in line of sight on his screen and that he wishes to manipulate, for example to carry out a maintenance operation.
[0111] According to the embodiment detailed in relation to FIG. 1, the system S comprises the rendering device HMD1 to HMD4 of the user OP1 and the individuals OP2 to OP4.
[0112] The system S may also comprise at least one sensor, for example a plurality of sensors, SNS_ENV1 to SNS_ENV4, placed in the environment ENV and configured to collect measurement data comprising at least location information of the HDM device 1 (and / or its user OP1), the HDM devices 2 to HDM4 (and / or individuals OP2, OP3, OP4 using these devices) and more generally of objects of interest present in the environment. To do this, the at least one sensor comprises for example at least one RFID (Radio Frequency Identification) location and identification device configured to communicate via a radio link with at least one electronic tag TAG placed on at least one object of interest, when it is in its field of action (for example approximately from 1 cm to 50 m), retrieve the identification information contained in this tag and locate it precisely by triangulation.For example, the wireless communication technology used belongs to a group including radio technologies such as UWB, BLE (from the English, “Bluetooth Low Energy”, Wi-Fi, 5G etc.
[0113] Of course, depending on the embodiments, other types of sensors suitable for providing location information may be used.
[0114] Thus, the system may comprise at least one camera, here a plurality of video cameras placed according to different points of view and configured to capture 3D video sequences of the ENV environment. In this regard, the 3D video sequences provided by these cameras may be used to detect and locate objects of interest (physical objects or living beings) present in the ENV environment but not equipped with electronic tags TAG.
[0115] The system S may also include temperature sensors, smoke detectors, or more generally any other type of sensor capable of collecting measurement data relating to a state of the environment or the elements which constitute it at a given moment.
[0116] According to the illustrated example, the system S comprises a device 100 for generating data for enriching a view of a part of the environment ENV, for example displayed on a screen of the rendering device (for example HMD1) of at least one user (for example OP1). Such a device is configured to obtain information relating to its own location (or that of its user), from data collected by sensors coupled to the rendering device HMD1 worn by the user and location information of at least one object of interest likely to move in said environment. This location information may be collected by sensors placed in the environment or on and / or in the rendering device of the user and / or on and / or in the objects of interest.The device 100 for generating enrichment data is configured, at least from the location information obtained, to determine a relative position of the object of interest from the point of view of the device 100 (or in other words its reference frame) and then generate enrichment data for said view, at least from the determined relative position. Said enrichment data comprises at least visual, audio, haptic, etc. indications of the position of said at least one object of interest relative to the user OP1 and / or to his rendering device. For example, the device 100 is configured to transmit said enrichment data to the rendering device of the view concerned, for example to the rendering device HMD1 of the user OP1 for display.
[0117] The device 100 thus implements the method for generating enrichment data for at least one view of a portion of the environment rendered by a rendering device according to the invention which will be detailed below in relation to FIG. 2. According to the example of FIG. 1, each of the operators OP1 to OP4 is a user of a rendering device HM1 to HMD4 and the device 100 is configured to generate enrichment data for the views viewed by each of them. Obviously, in other implementations it can generate enrichment data rendered on a single rendering device of the system S. In certain embodiments, the device 100 can also generate enrichment data for a global rendering of the environment ENV for a supervisor user SV in charge of the security of the environment ENV and equipped with a rendering device, for example the terminal equipment ET.This SV supervisor and its terminal equipment ET can be present in the ENV environment or remote and connected to the other elements of the supervision system S by a communication network (not shown).
[0118] In the example of Figure 1, the system S comprises a server equipment S which integrates the device 100. Alternatively, the device 100 may be independent of the server equipment S, but connected to it by a wired or wireless link. Alternatively, the server equipment ES may be placed outside the environment ENV, in a remote communication network, for example organized according to a cloud architecture, provided that the communications between the server equipment and the rendering devices can be done with limited latency (for example without latency), i.e. in near real time.
[0119] According to yet another option, the device 100 can be integrated into a rendering device, for example the rendering device HMD1 of the user OP1 or the rendering device ET of the supervisor SV.
[0120] According to at least one exemplary embodiment of FIG. 1, the server equipment ES comprises a transmitting / receiving E / R module configured for example to receive the measurement data collected by at least one sensor (for example the plurality of sensors SNS_ENV1 to SNS_ENV4) placed in the ENV environment and / or at least one sensor of at least one rendering device (for example sensors of the rendering devices of the users SNS_OP1 to SNS_OP4).
[0121] The server equipment ES may also comprise a transmission / reception module (merged with or distinct from the module below) configured to transmit and receive data to or from other elements of the system S, for example receiving the measurement data from the sensors and / or transmitting the enrichment data provided by the device 100 to the rendering device HMD1, comprising at least one rendering indication of at least one relative position (with respect to OP1) of at least one object of interest present in the environment, such as for example at least one of the other users OP2-OP4. The server equipment ES may also comprise a memory M configured to store at least the measurement data received from the sensors and the rendering enrichment data.
[0122] In some embodiments, the server equipment ES may further comprise an artificial intelligence module MIA1 configured to detect and / or recognize objects of interest present in the environment and visible at least partially in 3D video sequences acquired by cameras placed in the environment and / or on and / or in the rendering devices HMD1 to HMD4 of the users OP1 to OP4. It is noted that the module MIA1 may also be configured to detect and / or recognize obstacles present on the path of one of these users (cash register, forklift, etc.). According to at least one embodiment, the module MIA1 is configured to analyze the images of the video sequences to determine in which context the object(s) of interest that it has detected are used.For example, if the MIA1 module detects that the user has an identifiable object of interest in front of them and that their hands are close to it, it may decide that the user wishes to intervene or is in the process of intervening on this object and produce this information as output so that enrichment data is generated to increase the view of this object of interest for the user. Alternatively, it may also be configured to produce as output information that another user is already intervening on this object of interest, so that specific enrichment data is generated to signal this to the user.
[0123] The server equipment ES may also comprise an artificial intelligence module MIA2, possibly distinct from the previous module MIA1, configured to detect and / or recognize parts (for example limbs) of the body of individuals at least partially visible in the video sequences (or other parts of non-human objects of interest in certain embodiments).
[0124] Alternatively, the transmission / reception module, the memory, the artificial intelligence module(s) may be integrated into the device 100 itself.
[0125] According to the exemplary embodiment of Figure 1, the server equipment ES comprises a module JUM configured to generate a virtual graphical representation of the environment ENV, called digital twin, which a user, for example the supervisor SV, can view on a display device DISP of his terminal equipment ET and in which he can navigate as he wishes to carry out monitoring. As previously mentioned, the device 100 can, in certain embodiments, be configured to generate enrichment data for one or more views of this digital twin and transmit them to the rendering device ET. In relation to Figure 2, a method for generating data for enriching the rendering of at least one rendering device according to at least one embodiment of the invention is now described.In the detailed example, this involves enriching a view of at least part of the environment ENV displayed on a screen of a rendering device HMD1 of at least one user OP1. Such a method is for example implemented by the device 100, integrated for example in the server equipment ES of FIG. 1. In the illustrated example, the method may comprise obtaining 21 location information LOC of the rendering device HMD1 and / or of the user OP1, obtaining 22 location information LOC_OI of at least one object of interest present in the environment ENV and likely to move there, such as at least one of the other users OP2 to OP4 or a robot. This location information may be obtained, for example, from data received from the SNS_HMD1 sensors of the rendering devices HMD1 to HMD4 worn by the user OP1 and the other users OP2 to OP4 and / or by other sensors SNS_ENV1 to SNS_ENV4 placed in the environment.For example, this location information can be obtained from RFID location and identification devices placed in the environment, as previously described in relation to FIG. 1. In certain embodiments, it can be obtained multiple times, for example periodically, for example once per second. It can be associated with time information representative of the current time instant of the measurement. It can also be obtained non-periodically, for example as a function of the movements of the user and / or objects of interest (when passing near a sensor such as an RFID reader for example).
[0126] In the illustrated example, the method may then comprise a determination 23 of the positions of the user OP1 (and / or of his rendering device) and of the objects of interest Ol such as the other users OP2 to OP4 in the reference frame Ref of the environment ENV, from the location information obtained. They are for example stored in memory, for example in memory M of the server equipment ES. This determination 23 may for example be triggered by obtaining new location information at 21 and 22. Relative positions of the objects of interest, with respect to the user OP1 (in a reference frame of the user centered on his rendering device HMD1) may also be determined in at least certain embodiments.
[0127] At 29, enrichment data can be generated at least from these relative positions. They include, for example, coordinates and description information of a rendering indicator, such as a visual indicator to be inserted into the view displayed on the screen of the rendering device HMD1. This is, for example, in the case of a visual indicator, description information (color, size, position, orientation, etc.) of an arrow pointing to the position of an object of interest Ol, such as another user present in the field of vision of the user OP1. It can also be visual indicators representative of danger zones, the presence of colleagues, critical information (breakdown, danger, etc.). Alternatively, it is also possible to introduce lateral vibrations to indicate to the user in which direction the danger zone, the colleague, etc., is located, characteristic sound signals or simply a voice message from a supervisor.
[0128] The user's field of view FOV can, for example, be determined from the orientation of his head and general knowledge about the human visual system. The orientation of the head can be obtained, for example, from inertial data collected by at least one sensor of the inertial measurement module IMU (from the English, "Inertial Measurement Unit") type placed on the rendering device HMD1 and configured to measure an acceleration, an angular velocity and / or an orientation using accelerometers, gyroscopes and / or magnetometers.
[0129] In some embodiments, the presence of another user or an object of interest moving in the ENV environment may be signaled even if it is outside the field of view FOV of the user OP1, for example behind their back.
[0130] The method may finally comprise the transmission 30 of the generated enrichment data at least to the rendering device HMDI of the user OP1 for the purpose of rendering them. For example, if the device HMD1 is of the augmented reality glasses type, the generated enrichment data (or more precisely the elements that they describe) will be superimposed on the real view of the environment that the user OP1 perceives through the glasses lenses. According to another example, if the device HMD1 is of the virtual reality headset type, they will be integrated by / during the updating of the virtual view generated by the headset. This transmission may be optional in certain embodiments, for example when the generation device is also the rendering device.
[0131] According to the exemplary embodiment of Figure 2, the method comprises obtaining 20 a 3D MAP of the environment ENV. This is for example a map that locates and identifies physical objects present in the environment ENV, such as partitions, doors, windows, rooms, paths, furniture, etc. These physical objects are generally fixed and participate in the layout of the premises. They are likely to constitute obstacles to the visibility of other users OP2 to OP4. This MAP may take into account data from a plan of the building, if applicable.
[0132] Optionally, this MAP mapping can be obtained using a localization and matching technique, for example SLAM (Simultaneous Location and Mapping) for example a technique described in the article by Ruan et al., entitled "GP-SLAM+: real-time 3D lidar SLAM based on improved regionalized Gaussian process map reconstruction", published in August 2020 by arXiv, according to which one or more operators, for example users OP1 to OP4, survey the ENV environment and capture 3D video sequences of this environment. The video sequences obtained are used to reconstruct a 3D view or map of the building, in particular by detecting common regions between the different sequences.This operation can be carried out, in an initialization phase, prior to the implementation of the method according to the invention, on the basis of a systematic and exhaustive survey of the operators, so as to obtain an initial version of the MAP mapping of the environment. It can in certain embodiments be completed as it goes along by taking into account the 3D video sequences acquired by the users OP1 to OP4 during their interventions in the ENV environment.
[0133] It is thus understood that such a 3D MAP mapping comprises location and identification information for physical objects OP that comprise the environment ENV (for example the physical objects that constitute it). In certain embodiments, this location and identification information can be taken into account when generating the enrichment data ENH (at 29). More precisely, the visual indications of the other users OP2 to OP4 can specify whether, from the point of view of the user OP1, they are placed in front of or behind one or more of the physical objects OP located and identified by the mapping, therefore whether they are visible or at least partially hidden.
[0134] According to the exemplary embodiment of Figure 2, the method also comprises obtaining 24, from at least one CAM camera of said rendering device HMD1 of the operator OP1, 3D video data 3DV representative of the part of the environment perceived by the user OP1 from his position and according to his point of view. Here, we consider in particular the video sequences acquired over a time period between a previous time instant and the current time instant.
[0135] From the 3DV video data obtained, the method can implement the detection 26 of at least one object of interest present in the view, at least using the 3D video data obtained. This is for example a physical object, such as a particular part of a machine tool which the user OP1 has approached and towards which he at least directs his gaze, or even his hands and on which he must intervene. It can also be a fixed physical object, which has not for example been mapped, or even a living being such as an individual who is not equipped with an HMD rendering device (and therefore a TAG label) or an animal. According to another example, the user could blink to designate an object of interest placed in front of him for which he wishes an enrichment of the view currently being viewed.In some embodiments, a relative position of said at least one object of interest with respect to the user OP1 is determined, from the obtained 3D video data.
[0136] In this regard, according to at least one embodiment, the detection can be configured so as to characterize the objects of interest sought. For example, it is possible to define, using a radius centered on the rendering device, an area in which the detected objects of interest are taken into consideration and presented to the user of the rendering device. Too close, part of the information would be missing; too far, it would be overloaded.
[0137] In some embodiments, the step of detecting at least one object of interest can be carried out by an artificial intelligence module MIA1 implementing a previously learned model. Such a module is trained to produce position and classification information, or even identification of the object of interest. It will be detailed further in relation to FIG. 4 A.
[0138] According to the exemplary embodiment of Figure 2, the method comprises at 28 obtaining description information DESC of the detected object of interest Ol. This is for example a technical sheet or a manipulable 3D virtual representation of the object of interest. For example, this description information DESC is stored in the memory M of the server equipment ES, for example within a virtual catalog of objects of interest present in the environment ENV.
[0139] This description information DESC can be taken into account at 29 to generate additional enrichment data relating to the detected object of interest. For example, this additional enrichment data can define a visual outline indicator of the object of interest to be superimposed on the displayed view and / or a command to insert the 3D virtual graphical representation of this manipulable object of interest in a corner of the screen of the HMD1 rendering device.
[0140] This additional enrichment data is aggregated at 29 with the other enrichment data (comprising the relative position of at least one object(s) of interest (for example at least one of the other users) with respect to the user), before being transmitted at 30 to the rendering device HMD1 of this user for display.
[0141] In some embodiments, the detection information of the objects of interest is also stored in memory. It can thus be used to enrich the views of the environment of other users and / or that of the digital twin, and / or for the purpose of building a history.
[0142] According to the exemplary embodiment of FIG. 2, the 3DV video data captured at 24 are further used to detect and / or recognize at 25 one or more parts of the body of the user OP1 and / or one or more parts of an individual who is not equipped with an HMD rendering device or even of another user OP2-OP4 of an HMD rendering device. They can also in certain embodiments be used to detect and / or recognize at 25 one or more parts of an object of interest, such as a device having a body and / or rigid portions, possibly articulated and / or interconnected, of known shape and / or possible movements, for example, such as an electromechanical device (such as a robot) referenced in a knowledge base).
[0143] With regard to the user OP1, this concerns, for example, his hands, arms, feet and legs. Indeed, when the user OP1 approaches a machine tool to manipulate a particular part of this machine, his hands may be partially visible on the video data acquired by the camera placed on his rendering device HMD1. If he tilts his head, his feet or legs may at least partially appear on the video sequence. It is noted that this step may, in certain embodiments, take into account video data from other cameras, for example fixed cameras of the ENV environment and oriented so that the user is in their field of vision and that parts of his body not present on the video data captured by his rendering device HMD1 are visible.
[0144] As for individuals (such as other users OP2 to OP4), these are, for example, those who are at least partially visible in the video data captured by the camera of the rendering device HMD1 and possibly other cameras placed in the ENV environment over the considered time period.
[0145] Of course, the detection of objects of interest in the environment and near the rendering device can be carried out from other input data than one or more video sequences. For example, data acquired by other types of sensors than a video camera, such as a laser (LIDAR), a sonar, a microphone, a pressure sensor, a motion detector, can be used instead of or in addition to the video sequence(s).
[0146] In certain embodiments, this detection step can be carried out by an artificial intelligence module MIA2 implementing a previously learned MOD2 model. It will be detailed further in relation to Figure 4B.
[0147] Once detected and identified, the part of an object of interest (for example the part of the body of the user OP1 and / or of at least one other individual, for example of the other users OP2-OP4, or the part of an object of interest other than human), is positioned in a reference frame of the user OP1 and, in certain embodiments in a reference frame of the environment ENV. The location and possibly identification information obtained can be stored in memory in association with the current time instant (associated with the captured video sequence) and the identification of the user concerned.
[0148] According to the exemplary embodiment of Figure 2, the method comprises at 27 obtaining, at least from the detection information obtained, a virtual 3D graphic representation of the identified part(s) of the body of the user OP1 and / or parts of other possible objects of interest (such as parts of the body of other individuals). In other words, in certain embodiments, it is a matter of generating an avatar of a point of interest such as the user OP1 (and / or the other possible individual(s)) which can be inserted and animated in one or more views of the environment ENV intended to be viewed by the user himself and / or other users of other rendering devices in quasi-real time.
[0149] This 3D graphic representation or avatar may include, for example, at least some of the following information:
[0150] - Point of interest identifier,
[0151] - Model of the point of interest;
[0152] - Location information (Latitude, longitude and elevation),
[0153] - Absolute position, calculated relative to the environment's Ref reference frame,
[0154] - A list of parts to display containing both their position (X, Y, Z) and a rotation to apply (pitch, yaw, roll),
[0155] - Timestamp associated with the last location and identification information obtained.
[0156] In some embodiments, each part of a point of interest (head, trunk, limb for an individual) can be defined independently of the others. For example, an arm is composed of an elbow, a forearm, a wrist, a palm of the hand, fingers, themselves composed of phalanges. It is therefore necessary to precisely position and animate each of these elements.
[0157] For example, we define different parts of the user's body as follows:
[0158] {
[0159] Base of the spine: "SpineBase":{"location":{"x":0.71 ,"y":8.82,"z":19.7},"rotation":{"x":-0.012,"y":0.999,"z":0}},
[0160] Mid-spine: "SpineMid":{"location":{"x":0.64,"y":11.77,"z":19.69}, "rotation":!" / ':^.012, "y":1,"z":0}},
[0161] Neck: "Neck":{"location":{"x":0.69,"y":14.45,"z":19.52},"rotation":{"x":0.024,"y":0.997,"z":-0.052}}, Head: "Head":{"location":{"x":0.37,"y":16.28, "z":19.8},"rotation":{"x":0,"y":0,"z":0}},
[0162] Left shoulder: "ShoulderLeft":{"location":{"x":-1 .1 ,"y":13,"z":19.16},"rotation":{"x":0.836,"y":- 0.538, "z":0.057}},
[0163] Left elbow: "ElbowLeft":{"location":{"x":-0.96,"y":13.56,"z":16.15},"rotation":{"x":-0.636,"y":- 0.003, "z":-0.027}},}
[0164] In some embodiments, as described in connection with Figure 3, the relative positions of the elements that make up a part of an individual's body (trunk, limbs, head) can be determined using a technique called inverse kinematics, for example described in the document by Ramadoss et al., entitled “Whole-Body Human Kinematics Estimation using Dynamical Inverse
[0165] Kinematics and Contact-Aided Lie Group Kalman Filter”, published by arXiv in May 2022.
[0166] In this technique, limbs are rigid solids of a given length, connected to each other by joints. The joints support a certain maximum torsion, which varies depending on the direction or speed. The movements are assumed to be sufficiently regular to appear natural. Such a technique can help produce a relatively realistic avatar in the sense that the individual's body will only adopt realistic postures.
[0167] It is understood that the 3DV video data obtained from the HMD1 rendering device of the user OP1 and / or other cameras placed in the ENV environment do not necessarily provide an image of the body, in its entirety, of the user OP1 or of another user OP2-OP4 present in his field of vision. Nevertheless, the inverse kinematics technique makes it possible to extrapolate the invisible parts on the video data.
[0168] According to the exemplary embodiment of Figure 2, the relative positions of the user OP1 and the other users OP2-OP4 determined at 22 are used at 23 to determine which avatars are to be inserted into the enrichment data ENH intended for the user OP1. In certain embodiments, the knowledge of the layout of the environment provided by the mapping MAP obtained at 20, optionally the location and / or identification information of the objects of interest OOI and / or the location information and / or the identification and / or location information of at least one part of the body of one or more other users present in the part of the environment rendered by the rendering device, obtained at 25, are used at 29 to determine the parts of the objects of interest hidden in the view rendered by the rendering device of the user OP1.Knowledge of this occluded part of the object of interest can then be used to generate specific enrichment data, intended to complete the rendering of this object of interest. For example, in the case of an individual for whom it has been determined that at least one of his limbs is occluded in the view, the corresponding limb of the avatar generated for this individual can be selected to be inserted into the enrichment data intended for the rendering device HDM1 of the user OP1. In this way, the enrichment data makes it possible to represent the occluded body part of the other user in the view displayed on the rendering device HMD1 of the user OP1.
[0169] In certain embodiments, as described in relation to FIG. 4A, the 3D video sequence(s) captured by the camera(s) of the rendering device HMD1 of the user OP1 are provided as input to an artificial intelligence module MIA1 implementing a data analysis model MOD1 obtained by machine learning and intended to locate and identify one or more objects of interest Ol from input video data.
[0170] For example, this MIA1 module is a neural network or any other artificial intelligence module capable of performing the same functions, based for example on one of the following techniques
[0171] - decision tree, for example Random Forests type,
[0172] - boosted decision trees (or “Gradient boosted decision trees”, in English), for example of the Catboost, XGBoost or LightGBM type
[0173] - support vector machine (or “Support Vector Machine” in English);
[0174] - Multi-layer Perceptron type neural network;
[0175] - etc.
[0176] In certain embodiments, this learning or training may have been implemented during a prior step (not shown in FIG. 2), which made it possible to construct the analysis model MOD1 of the module MAI1 from a set of training data. The analysis model MOD1 thus constructed is used at 25 by the module MAI to identify and locate one or more objects of interest OOI present in the field of vision of the user OP1, and in particular those at least partially obscured.
[0177] It is noted that the training of the MOD1 model and the analysis of video data using the MOD1 model to locate and identify objects of interest in the input data, for example 3D video data, are presented here in two stages, or in two distinct phases, for simplicity. It is understood, however, that the training can be carried out several times (in particular in parallel with or after the analysis) and that the analysis can be continuous.
[0178] Thus, in certain embodiments, the learning may comprise an “upstream” learning phase (initial and prior to the analysis phase) for learning to analyze 3D video data captured by one or more cameras from the point of view of the rendering device and defining parameters (or weights) to then enable, from any other 3D video data captured by other cameras of other rendering devices, presented as input to the module MD1, the identification and localization of the object(s) of interest present. The upstream learning phase is therefore based on a learning set or base comprising a plurality of labeled video data (i.e. labeled). In certain embodiments, a label may comprise the positions of key points of the object of interest in the video data and an identifier of the object of interest.For example, the training base may include labeled video data of each of the parts in a machine tool parts catalog, from several distinct viewpoints.
[0179] The learning phase may also include on-the-fly learning from the data collected by at least one device for rendering a view of at least part of the ENV environment, in order to refine the configuration of the artificial intelligence module MIA resulting from the upstream learning (ground reality). The two learnings may be carried out on the same equipment (for example by the device 100, by the server equipment ES, and / or by different equipment (for example the upstream learning and / or the on-the-fly learning may be carried out on another equipment, for example remote, dedicated to this task, I and having suitable computing and memory capacities.
[0180] According to at least some embodiments, the artificial intelligence module MIA1 receives as input in the form of an input vector V1 IN, the 3D video data transmitted by the rendering device HMD1 to the device 100. For example, such a vector V1 IN comprises a temporal sequence of this video data of given duration, for example equal to 1s. In some embodiments, the vector V1 IN may comprise additional information relating to a pose of the user's head for example, so as to take into account the direction of his gaze.
[0181] According to at least some embodiments, the prediction module MIA1 may be configured to produce as output a vector V1 OUT comprising location and identification information of one or more objects of interest, associated with one or more time instants. Of course, it may happen that no object of interest is detected, for example when the user OP1 is wandering in the environment ENV and is not working on a given machine tool.
[0182] According to the exemplary embodiment of FIG. 2, the output vector V1 OUT is used at 27 to obtain description information DESC, comprising for example a 3D representation model of the identified object of interest, for example from a local or remote memory, for example structured as a database indexed by identifiers of the objects of interest. The description data obtained are then used to generate additional enrichment data for the user's view at 29, aggregate them with the other enrichment data previously generated, in particular position information, and transmit them at 30 to the user OP1. In certain embodiments, the location and identification information of the object of interest determined by the module MIA1 are stored in memory in association with the current time instant.
[0183] In relation to Figure 4B, the 3D video sequence(s) captured by the camera(s) of the rendering device HMD1 of the user OP1 as well as video sequences captured by other cameras placed in the environment ENV, are provided as input to an artificial intelligence module MIA2 implementing a data analysis model MOD2 obtained by machine learning and intended to locate and identify parts of an object of interest (such as parts of the body of one or more other users). The considerations on the implemented neural network and the learning made in relation to Figure 4A apply.
[0184] In this case, the learning base may comprise, in certain embodiments, a plurality of video sequences of objects of interest labeled with location information and identification of key points of these objects (for example for individuals their head, trunk, limbs). In certain embodiments, the base covers a plurality of possible postures or positions of the objects of interest and the learning video sequences were captured in the ENV environment.
[0185] In relation to Figure 5, two images 11, 12 of a view of a scene captured by a camera are shown. In image 11, the upper body of an individual is visible through a window F and the lower part of his body is hidden by an MR wall. In image 12, it is completely hidden by the MR wall. The view of the scene has been enhanced by embedding in images 11 and 12, in place of the parts of the body hidden by the MR wall, the corresponding parts of a virtual graphic representation or avatar of this individual. It is understood that the insertion of the avatar of the individual can help a user of a rendering device that would display this scene to him, to know at any time the position and posture of this individual and therefore help him to follow his movements and actions more easily.
[0186] In relation to Figure 6, the data flows within a system S are now described according to another exemplary embodiment. In the foregoing and in relation to Figures 2 to 5, an example of implementation of the invention has been described for a particular rendering device HMD1. In relation to Figure 1, a system S has been described comprising a plurality of rendering devices HMD1-HMD4 carried by a plurality of users OP1-OP4 present in an environment ENV. According to the exemplary embodiment of Figure 6, the server equipment ES of the system S is configured to simultaneously and in real time enrich the views displayed on the rendering devices HMD1 to HMD4 of each of the users OP1 to OP4. In other words, it is configured to implement the method of generating enrichment data for each of the views displayed on the rendering devices HMD1 to HMD4 of the users OP1 to OP4.In this example, it comprises a device 100 according to the invention.
[0187] In some embodiments, it is further configured to enrich the digital twin view of the ENV environment displayed on the DISP rendering device intended for overall supervision of this ENV environment, for example by a supervisor responsible for the security of the environment and its users.
[0188] In relation to Figure 6, RFID location and identification EMT devices placed in the ENV environment emit several times, for example periodically, for example once per second, a “ping” type request signal, to which electronic tags of the TAG type placed on the rendering devices HMD1 to HMD4 of the users OP1 to OP4 and present in the field of action of the EMT devices respond by transmitting the identification information that they contain. The EMT devices determine a location of each of the TAG tags by triangulation and transmit it to the device 100. Such EMT location devices can emit signals to the electronic tags of the TAG type (for example tags that it is desired to locate in the ENV environment) and / or receive signals from such electronic tags.In some embodiments, the electronic tags may be passive (with reference to their use), for example in the case of RFID type tags. Other technologies may be used, for example Ultra Wide Band, Bluetooth Low Energy or Wifi. In the case of RFID type tags, said EMT location device may be configured to transmit, to the electronic tags TAG placed on the rendering devices HMD1 to HMD4, a request signal, receive a signal comprising identification information from said electronic tags, determine the locations of said electronic tags, based on the received identification information, and transmit said locations to said enrichment data generation device.
[0189] In some embodiments, the electronic tags may be active (with reference to their use), for example in the case of RFID type tags. Thus, the electronic tags TAG placed on the rendering devices HMD1 to HMD4 may emit (for example regularly) their respective identification data. This emission may be done at their initiative (and not in response to a request). Thus, the location device EMT may determine the locations of the tags without performing a prior step of emitting the request signal.
[0190] The device 100 can in turn determine at 22 position information of at least some (for example each) of the devices HMD1 to HMD4 in a reference frame Ref of the environment, based on the locations transmitted by said location device EMT. It can also obtain measurement data MES collected by at least one sensor. These or this at least one SNS sensor can be of various types and include for example one or more video cameras placed for example at different points of the environment ENV, one or more temperature sensors, one or more smoke detectors, one or more motion detectors, for example laser-based, one or more volumetric sensors of the LIDAR type for measuring a distance, etc.
[0191] In some embodiments, the device 100 is also configured to collect data transmitted by programmable logic controllers (PLCs), which are programmed components used for controlling or regulating machines or installations. Such commands can be implemented in very different sectors, for example to control production installations, injection presses. Such data can help the device 100 to know the status of the machines and more generally of the equipment present in the environment. It can also, in some embodiments, control them remotely.
[0192] In certain embodiments, the device 100 receives at 24 data from the sensor(s), such as sensors SNS_HMD1 to SNS_HMD4 embedded on the rendering devices HMD1 to HMD4, comprising for example inertial data (acceleration, speed, and / or orientation) and / or captured 3D video data captured by at least one camera placed in the upper and front part of the corresponding rendering device and configured to film the part of the environment located in the field of vision of the user OP1 to OP4.
[0193] It is also assumed that the device 100 has previously obtained, for example from a memory M, a mapping MAP of the environment ENV, which includes position and identification information of physical elements or objects arranged in the environment ENV (partitions, machines, mobilization, doors, windows, aisles, etc.).
[0194] The information in this mapping can, in certain embodiments, make it possible to determine, using the position of another user present in the user's field of vision, whether he is actually visible or whether he is obscured by an element of the environment (such as a partition for example). This MAP mapping can also, in certain embodiments, be used to create a virtual 3D representation of the ENV environment, called a digital twin JUM. This digital twin can, in certain embodiments, be rendered (for example, displayed) on the rendering device ET placed or not in the ENV environment. In certain embodiments, a user of this device ET, for example an SV operator responsible for the supervision and security of the ENV environment, can navigate in this digital twin, as during a virtual tour.It can also display specific information about certain machines, such as their status.
[0195] In some embodiments, the device 100 may be configured to predict a trajectory of at least one object of interest moving in the environment. This prediction may, for example, take into account sensor data, such as, for example, speed and acceleration data, relating to this or these object(s) of interest. For example, from the position information obtained for each of the rendering devices HMD1 to HMD4, other sensor data, such as, for example, speed and acceleration data of at least some (for example, each) of the rendering devices HMD1 to HMD4, the device 100 may be configured to predict a trajectory of an object of interest. According to another example, it is configured to take into account previous movements of this object of interest, its current trajectory and the view of the part of the environment to predict where it will go next.
[0196] According to another embodiment, the data collected by all the sensors for all the moving objects of interest are used to predict the trajectories of each of these objects of interest and avoid collisions, by generating alert signals and enriching the views of the different users with these signals.
[0197] The device 100 can thus help to prevent possible collisions, the device 100 can be configured to transmit, in the event of an established risk of collision, an alert message to at least one of the users concerned. For example, this message is a voice message sent to their mobile terminals. It can also be a visual alert indicator or transmitted as data enriching the view displayed on the screens of their respective rendering devices.
[0198] In some embodiments, the device 100 may be configured to locate and / or identify objects of interest viewed by at least some of the users OP1 to OP4 of the rendering devices HMD1 to HMD4. This location / identification may take into account, for example, video data received from the cameras of the rendering devices HMD1 to HMD4 of these users for a given time period (between two successive time instants of obtaining the location information, for example of duration of the order of a few seconds (for example 1 second, 2 seconds, 5 seconds, etc.). In some embodiments, it may use (in addition to the video data, for example) orientation data obtained from the inertial sensors (gyroscope) of the rendering devices HMD1 to HMD4 to determine a pose of the head of at least one user and the orientation of his gaze, which may help to delimit a more restricted area of interest in his field of vision.In certain embodiments, the device 100 may use the artificial intelligence module MIA1 integrated for example into the server equipment ES which implements a previously learned model MOD1 to carry out these identification and localization operations from the captured video data, as previously described in relation to FIG. 4A. The identification information of the objects of interest detected for at least one user is associated with position information of these objects of interest in the reference frame Ref of the environment and stored in memory. In certain embodiments, this information is associated with a virtual 3D graphical representation manipulated by the corresponding object of interest available in memory, for example in the catalog.
[0199] The video data received from the cameras of the rendering devices HMD1 to HMD4, and in certain embodiments the video data from other cameras (SNS_ENV1 to SNS_ENV4) placed in the environment can also be used to locate and identify at least certain parts of the body of the user of the rendering device and / or other individuals visible in the video. For this operation, the device 100 can call upon the artificial intelligence module MIA2 (of the server equipment ES for example), configured to recognize parts of the body of an individual which appear in one or more video sequences which are provided to it as input, as described in relation to FIG. 4B.The identified and located body parts can in certain embodiments be associated with a user by exploiting the position information previously obtained at 21 (in particular the TAG label received from the position sensor of its rendering device) and stored in memory M.
[0200] In some embodiments, a virtual graphical representation or avatar of at least one object of interest (individual or electromechanical device for example) can be generated from the identification information and positions of the parts of the object of interest detected (therefore visible) on the video data. This virtual graphical representation can for example, in some embodiments, be based at least partially on an anatomical model independently defining each part of the object of interest as a solid, for example according to an inverse kinematics technique, as previously described in relation to FIG. 3. It is understood that this technique can help to extrapolate the positions of parts (such as limbs of a user) not visible on the video sequences.
[0201] For at least one user OP1 to OP4, his position information and that of at least one object of interest (such as one of the other users of rendering devices), the mapping information MAP and the identification and location information of parts of the user's own body and / or parts of other objects of interest (such as individuals visible in the view displayed on the screen of his rendering device HMD1 to HMD4) and other video collected in the environment, can be aggregated, in certain embodiments, to determine (for example using depth information) whether other parts of the user's own body or parts of at least one object of interest (such as one or more other individuals) are occluded.If necessary, the corresponding parts of their avatars can be obtained (e.g. from memory M) and inserted as ENH enrichment data of the view displayed on the user's rendering device.
[0202] The enrichment data thus generated for a user of a rendering device can be aggregated and then transmitted to the rendering device, when the device 100 is not integrated into the rendering device.
[0203] In some embodiments, specific enrichment data may be generated for the digital twin view JUM of the ENV environment displayed on the monitoring device DISP.
[0204] Finally, in relation to Figure 7, an example of hardware structure of a device 100 for generating enrichment data for at least one view of at least one part of an environment (ENV) from at least one first viewpoint is presented. Said view is intended to be rendered at least partially by at least one rendering device. Said device comprises at least one module for obtaining location information for at least one object of interest in said environment and a module for generating enrichment data for said view, at least from the location information obtained. Said enrichment data comprises at least data representative of a position of said object of interest relative to said first viewpoint and data representative of an obscured part of said object of interest.
[0205] In some embodiments, the device 100 may comprise a module for detecting said at least one object of interest at least in said view, a module for obtaining description information of said at least one detected object of interest, and the module for generating the enrichment data is configured to use the description information to generate at least some of the enrichment data of the view.
[0206] In certain embodiments, the device may comprise a module for obtaining a graphical representation of the occluded part of the object of interest in the view of the part of the environment rendered by the rendering device at least as a function of the position of the object of interest relative to the point of view and the description information and the generation module is configured to generate enrichment data of the occluded part comprising the graphical representation obtained.
[0207] In some embodiments, the device 100 may comprise a module for obtaining a virtual graphical representation of the identified object of interest, the module for generating the enrichment data being configured to insert said virtual graphical representation of the identified object of interest into the view of the environment viewed by the user. In some embodiments, the module for detecting said at least one object of interest is configured to implement an artificial intelligence module configured to use at least one previously learned model and associating said 3D video data with at least one position information and optionally classification and / or identification information of said object of interest.
[0208] The term "module" can correspond to a software component as well as to a hardware component or a set of hardware and software components, a software component itself corresponding to one or more computer programs or sub-programs or more generally to any element of a program capable of implementing a function or a set of functions.
[0209] More generally, the device 100 may comprise a random access memory 103 (for example a RAM memory), a processing unit 102 equipped for example with a processor, and controlled by a computer program Pg1, representative of the previously mentioned modules, stored in a read-only memory 101 (for example a ROM memory or a hard disk). Upon initialization, the code instructions of the computer program are for example loaded into the random access memory 103 before being executed by the processor of the processing unit 102.The RAM 103 may also contain, for example, the determined position information, the location and identification information of the objects of interest, the location and identification information of the body parts of the user of the device and / or parts of objects of interest (such as those of individuals, for example other users), the virtual graphical representations of the objects of interest, the virtual graphical representations of the body parts of the users, the mapping of the environment, etc.
[0210] Figure 7 illustrates only one particular way, among several possible ways, of producing the device 100 so that it performs the steps of the method for generating data for enriching a view of an environment of at least one user of a rendering device as detailed above, in relation to Figure 2, in its different embodiments. Indeed, these steps can be carried out indifferently on a reprogrammable computing machine (a PC computer, a DSP processor or a microcontroller) executing a program comprising a sequence of instructions, or on a dedicated computing machine (for example a set of logic gates such as an FPGA or an ASIC, or any other hardware module) included in the device 100.
[0211] In the case where the device 100 is produced with a reprogrammable computing machine, the corresponding program (i.e. the sequence of instructions) may be stored in a removable storage medium (such as for example an SD card, a USB key, a CD-ROM or a DVD-ROM) or not, this storage medium being partially or totally readable by a computer or a processor.
[0212] The various embodiments have been described above in relation to a device 100 integrated into a rendering device and / or server equipment placed or not inside the environment. The invention is however not limited to this example, the server equipment being able to be remote, for example in a cloud network (from the English, “Cloud Computing”), provided that reduced latency conditions are guaranteed.
[0213] The invention just presented provides certain advantages in at least some embodiments. It can help to 'enrich' the view of an environment displayed to a user of a rendering device. According to at least one embodiment, it can provide the user with a view of the portion of the environment rendered through obstacles, which can help the user to better understand the environment and can help to limit the risks of collision and more generally of accidents.
[0214] The exemplary embodiments detailed above relate to an industrial environment. Of course, the invention is not limited to this particular case and applies more generally to any environment in which physical obstacles reduce the perception of a user of a rendering device displaying a view of a part of this environment.
Claims
Claims 1. Method for generating enrichment data of at least one view of at least one part of an environment (ENV) from at least one first point of view, said view being intended to be rendered at least partially by at least one rendering device, said method comprising: - obtaining (22) location information of at least one object of interest (Ol) in said environment (ENV), - a generation (29) of enrichment data (ENH) of said view, at least from the location information obtained, said enrichment data comprising at least data representative of a position of said object of interest relative to said first point of view and data representative of an occulted part of said object of interest.
2. Method according to the preceding claim, characterized in that the at least one rendering device is a supervision device (ET, DISP) of said environment (ENV), and / or a plurality of mobile rendering devices (HMD1-HMD4).
3. Method according to the preceding claim, characterized in that at least one rendering device among the plurality of mobile rendering devices (HMD1-HMD4) is a mobile terminal equipment of said environment (ENV) and comprises at least one camera configured to capture video data from the first point of view, the rendered view being generated at least partially from video data of said camera.
4. Method according to one of claims 2 or 3, characterized in that it further comprises the generation of enrichment data for a virtual graphic representation of said environment (ENV), intended to be rendered by said supervision device (ET, DISP) of said environment (ENV).
5. Method according to any one of the preceding claims, characterized in that it comprises: - a detection of said at least one object of interest (Ol) at least in said view, - obtaining description information (DESC) of said at least one detected object of interest, and - use of the description information to generate at least some of the view enrichment data.
6. Method according to the preceding claim, characterized in that it comprises: obtaining a graphic representation of the occulted part of the object of interest (Ol) in the view of the part of the environment rendered by the rendering device at least as a function of the position of the object of interest (Ol) relative to the point of view and the description information and in that the enrichment data of the occulted part comprise the graphic representation obtained.
7. Method according to any one of the preceding claims, characterized in that the location information of said at least one object of interest (Ol) is obtained from at least one position sensor placed near and / or on said at least one object.
8. Method according to any one of the preceding claims, characterized in that it further comprises: - obtaining (20) a map (MAP) of at least a part of the environment (ENV), comprising at least location information of a plurality of physical objects present in the environment; and in that the generation (28) of the enrichment data comprises at least the insertion of data representative of positions of said physical objects relative to said first point of view.
9. Method according to claim 5, characterized in that the detection of said at least one object of interest (O1) implements an artificial intelligence module (MIA1, MIA2) configured to use at least one model (MODI, MOD2) previously learned and associating said 3D video data with position and classification information of said object of interest.
10. Device (100) for generating enrichment data of at least one view of at least one part of an environment (ENV) from at least a first point of view, said view being intended to be rendered at least partially by at least one rendering device, said device being configured to implement: - obtaining location information of at least one object of interest (Ol) in said environment (ENV), - a generation of enrichment data (ENH) of said view, at least from the location information obtained, said enrichment data comprising at least data representative of a position of said object of interest relative to said first point of view and data representative of an occulted part of said object of interest.
11. Device (100) for generating enrichment data according to the preceding claim, characterized in that the at least one rendering device is a supervision device (ET, DISP) of said environment (ENV), and / or a plurality of mobile rendering devices (HMD1-HMD4).
12. Device (100) for generating enrichment data according to the preceding claim, characterized in that it is further configured to generate enrichment data of a virtual graphic representation of said environment (ENV), intended to be rendered by said supervision device (ET, DISP) of said environment (ENV).
13. Server equipment (ES) connected via a communication network to at least one rendering device (ET, DISP, HMD1-HMD4) adapted to produce at least a partial rendering of a view of at least one environment from at least a first point of view, characterized in that it comprises a device (100) for generating enrichment data according to any one of claims 10 to 12.
14. Server equipment (ES) according to the preceding claim, characterized in that it comprises a module (JUM) configured to generate a virtual graphic representation of the environment (ENV) intended to be rendered by said supervision device (ET, DISP) of said environment (ENV).
15. System (S) for supervising an environment (ENV), characterized in that it comprises at least one rendering device (ET, DISP, HMD1-HMD4) adapted to produce at least a partial rendering of a view of the environment from at least a first point of view, at least one sensor configured to collect at least one location information of at least one object of interest present in the environment and server equipment (ES) according to one of claims 13 or 14.
16. Supervision system (S) according to the preceding claim, characterized in that said at least one rendering device is a mobile rendering device provided with an electronic tag (TAG), and in that said supervision system (S) further comprises at least one location device (EMT) configured to implement: a reception of a signal comprising identification information from said electronic tag (TAG), a determination of a location of said electronic tag (TAG) as a function of the identification information received, and a transmission of said location to said device (100) for generating enrichment data.
17. Computer program comprising instructions which, when these instructions are executed by a processor, cause the latter to implement the steps of the method for generating enrichment data, according to any one of claims 1 to 9.
18. A computer-readable recording medium on which a computer program according to claim 17 is recorded.