Methods for obtaining multimedia elements and enriched rendering, electronic devices, systems, computer program products and corresponding media

The method of creating a multimedia element from an audiovisual stream with digital twin parameters addresses communication challenges in asynchronous settings by enriching object rendering, enhancing understanding and clarity in professional interactions.

FR3160839A1Pending Publication Date: 2025-10-03ORANGE SA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
FR2024003113
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing asynchronous communications face challenges in understanding and conveying complex information due to the lack of face-to-face interaction, particularly in professional settings, where videoconferencing does not fully resolve communication difficulties.

Method used

A method involving the creation of a multimedia element from an audiovisual stream that includes extracting parameters from a digital twin of an object of interest, allowing for enriched rendering of this object in a physical environment, enhancing communication by providing context and movement information.

Benefits of technology

Facilitates clearer communication by enabling users to describe and understand objects of interest in a physical environment, improving asynchronous communication by incorporating real-time and historical digital twin data for enhanced understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000035_0000
    Figure 00000035_0000
  • Figure 00000035_0001
    Figure 00000035_0001
  • Figure 00000036_0000
    Figure 00000036_0000
Patent Text Reader

Abstract

Methods for obtaining a multimedia element and enriched rendering, electronic devices, system, computer program products and corresponding supports The invention relates on the one hand to a method for obtaining a multimedia element, comprising: Obtaining an audiovisual stream representing a physical environment; Creating a multimedia element giving access to a value of a first parameter of a digital twin relating to a first object of interest of the audiovisual stream, the value being extracted from the digital twin during a capture of the audiovisual stream.And on the other hand, a method for enriched rendering of a first object of interest of a physical environment, comprising: - obtaining a multimedia element, said multimedia element giving access to a first value of a first parameter of a digital twin relating to said first object of interest; - obtaining a first audiovisual representation of a first portion of said digital twin taking into account the first value obtained; rendering of the first audiovisual representation in the physical environment. The invention also relates to the corresponding electronic devices, computer program products and media. Figure for the abstract: Fig. 5.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Methods for obtaining multimedia elements and enriched rendering, electronic devices, system, computer program products and corresponding supports 1. Technical field

[0001] The present application relates to the field of telecommunications and more precisely to the field of information sharing (for example via communications (such as asynchronous communications) between users via communication terminals. It relates in particular to a method for obtaining a multimedia element and a method for rendering such a multimedia element, implemented respectively by one or more electronic devices, as well as the corresponding electronic devices, computer program products and recording media. 2. State of the art

[0002] The use of electronic means to share information between individuals is becoming increasingly widespread. In particular, various communications are distinguished, such as synchronous communications (such as telephone calls or video conferences) and asynchronous communications (by electronic mail, for example).

[0003] Asynchronous communications offer many advantages. They allow you to send information to a recipient without having to arrange a meeting with them in advance or even worry about their availability. They therefore help to overcome the constraints of presence or schedules. They have become particularly useful, for example, in the professional world with the globalization of exchanges or the rise of teleworking.

[0004] Whether digital communications are synchronous or asynchronous, they can give rise to difficulties in understanding between the parties. In practice, it is often easier to "get a message across" (for example, express a feeling and / or give explanations) face to face. As a result, users often resort to videoconferencing (or video messaging) solutions. However, the use of video techniques does not resolve all communication difficulties between parties, particularly in the case of asynchronous communications and / or complex situations.

[0005] The object of the present application is to propose improvements to at least some of the drawbacks of the state of the art. 3. Statement of the invention

[0006] The present application aims to improve the situation using a method for obtaining a multimedia element, said method comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - A creation of a multimedia element giving access to at least one value of at least one first parameter of a digital twin relating to at least one first object of interest of said audiovisual stream, said value being extracted from said digital twin during a capture of said audiovisual stream.

[0007] According to the embodiments, said physical environment is a real environment or a virtual representation of a real environment for example.

[0008] In at least some embodiments, said multimedia element further comprises at least one audiovisual content obtained from said captured audiovisual stream.

[0009] In at least some embodiments, said audiovisual content is obtained by extraction from said captured audiovisual stream.

[0010] In at least some embodiments, said audiovisual content is obtained by transforming at least a portion of said captured audiovisual stream.

[0011] In at least some embodiments, said transformation comprises a replacement of at least one object of interest of a scene of said captured stream by a virtual representation of said object of interest.

[0012] In at least some embodiments, said multimedia element comprises at least two values ​​of said first parameter, said at least two values ​​corresponding to at least one fluctuation of said first parameter during said capture.

[0013] In at least some embodiments, said first object is selected, from among the objects of interest detected in said audiovisual stream, taking into account a current position and / or a current orientation of said first object in at least one scene of said audiovisual stream, relative to at least one position and / or an orientation of at least one second object of interest of said audiovisual stream.

[0014] In at least some embodiments, said second object relates to an individual present in said audiovisual stream. Said second object may for example correspond to a representation of the individual himself (real character) or to a virtual representation of this individual (an avatar for example).

[0015] In at least some embodiments, said at least one position of said second comprises a plurality of positions, representative of a movement of said second object in said audiovisual stream.

[0016] In at least some embodiments, said first object, respectively said second object, is selected from among the objects of interest detected in said stream audiovisual, taking into account a visual and / or auditory similarity between said first object, respectively said second object, and at least one reference object.

[0017] In at least some embodiments, said first object, respectively said second object, is selected, from among the objects of interest detected in said audiovisual stream, taking into account a visual and / or audio designation of said first object, respectively said second object, in at least one scene of said audiovisual stream.

[0018] In at least some embodiments, said first object, respectively said second object, is selected, from among the objects of interest detected in said audiovisual stream, taking into account a manual designation of said first object, respectively said second object, via a user interface.

[0019] In at least some embodiments, said manual designation is performed during and / or after said capture.

[0020] In at least some embodiments, said created multimedia element comprises values ​​relating to at least two parameters of said digital twin and said method comprises, after said creation, filtering said values ​​of said parameters.

[0021] In at least some embodiments, said filtering takes into account said objects of interest to which said parameters relate.

[0022] For example, thanks to filtering, it is possible to keep only parameter values ​​relating to certain object(s) of interest or to delete parameter values ​​relating to certain object(s) of interest.

[0023] According to another aspect, the present application relates to a method for enriched rendering of at least one first object of interest of a physical environment, said method comprising: - obtaining a multimedia element, said multimedia element providing access to at least a first value of at least a first parameter of a digital twin relating to at least said first object of interest; - obtaining at least a first audiovisual representation of at least a first portion of said digital twin taking into account said first value obtained; - a rendering of said first audiovisual representation in said physical environment.

[0024] In at least some embodiments, said first value is a value, at a time prior to said obtaining of said multimedia element, of said parameter.

[0025] In at least some embodiments, said at least a first portion of said digital twin comprises said first object of interest

[0026] In at least some embodiments, said multimedia element provides access to at least one audiovisual content relating to said first object of interest and the method comprises a rendering, jointly with said rendering of said first audiovisual representation, of at least a portion of said audiovisual content.

[0027] In at least some embodiments, the enriched rendering method comprises rendering, jointly with said rendering of said first audiovisual representation, a second audiovisual representation of at least a second portion of said digital twin taking into account at least one current value of said parameter of said digital twin.

[0028] In at least some embodiments, said at least one second portion of said digital twin comprises said first object of interest.

[0029] In at least some embodiments, said enriched rendering method comprises highlighting said at least one first object of interest in said first, respectively second, audiovisual representation.

[0030] In at least some embodiments, said rendering of said first audiovisual representation is implemented conditionally taking into account a presence of a speaker in said physical environment.

[0031] In at least some embodiments, said rendering of said first audiovisual representation is implemented conditionally taking into account a profile of said present speaker.

[0032] In at least certain embodiments, said rendering of said first audiovisual representation is implemented conditionally taking into account a geographical proximity between said present speaker and a piece of equipment in said physical environment.

[0033] In at least some embodiments, said equipment belongs to a group comprising: - a rendering device implementing said enriched rendering method; - a device for rendering said first and / or second audiovisual representation; - said at least one first object of interest.

[0034] The characteristics, presented in isolation in the present application in connection with certain embodiments of at least one of the methods of obtaining and / or rendering of the present application, can be combined with each other according to other embodiments of this method.

[0035] According to another aspect, the present application also relates to an electronic device adapted to implement at least one of the methods of obtaining and / or enriched rendering of the present application in any of its embodiments.

[0036] For example, the present application thus relates to an electronic device comprising at least one processor configured to obtain a multimedia element, said obtaining comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - A creation of a multimedia element giving access to at least one value of at least one first parameter of a digital twin relating to at least one first object of interest of said audiovisual stream, said value being extracted from said digital twin during a capture of said audiovisual stream.

[0037] For example, the present application thus relates to an electronic device comprising at least one processor configured for an enriched rendering of at least one first object of interest of a physical environment, an implementation of said enriched rendering comprising: - obtaining a multimedia element, said multimedia element providing access to at least a first value of at least a first parameter of a digital twin relating to at least said first object of interest; - obtaining at least a first audiovisual representation of at least a first portion of said digital twin taking into account said first value obtained; - a rendering of said first audiovisual representation in said physical environment.

[0038] According to another aspect, the present application also relates to a telecommunications system comprising at least one first electronic device adapted to implement the method of obtaining the present application in any one of its embodiments and at least one second electronic device adapted to implement the method of rendering the present application in any one of its embodiments.

[0039] Thus, the present application relates for example to a telecommunications system comprising at least one first electronic device comprising at least one processor configured to obtain a multimedia element, said obtaining comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - A creation of a multimedia element giving access to at least a first value of at least a first parameter of a digital twin relating to at least a first object of interest of said audiovisual stream, said first value being extracted from said digital twin during a capture of said audiovisual stream;

[0040] And at least one second electronic device comprising at least one processor configured for an enriched rendering of said at least one first object of interest, an implementation of said enriched rendering comprising: - obtaining said multimedia element; - obtaining at least a first audiovisual representation of at least a first portion of said digital twin taking into account said first value; - a rendering of said first audiovisual representation in said physical environment.

[0041] The present application also relates to a computer program comprising instructions for implementing the various embodiments of at least one of the above enriched obtaining and / or rendering methods, when the program is executed by a processor and a recording medium readable by an electronic device and on which the computer program and the corresponding information medium are recorded.

[0042] For example, the present application thus relates to a computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, a method for obtaining a multimedia element, said method comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - A creation of a multimedia element giving access to at least one value of at least one first parameter of a digital twin relating to at least one first object of interest of said audiovisual stream, said value being extracted from said digital twin during a capture of said audiovisual stream.

[0043] For example, the present application thus relates to a computer program comprising instructions for implementing, when the program is executed by a processor of an electronic device, a method for enriched rendering of at least one first object of interest of a physical environment, said method comprising: - obtaining a multimedia element, said multimedia element giving access to at least a first value of at least a first parameter of a digital twin relating to at least said first object of interest; - obtaining at least a first audiovisual representation of at least a first portion of said digital twin taking into account said first value obtained; - a rendering of said first audiovisual representation in said physical environment.

[0044] For example, the present application also relates to an information medium readable by a processor of an electronic device and on which is recorded a computer program comprising instructions for implementing, when the program is executed by the processor, a method for obtaining a multimedia element, said method comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - A creation of a multimedia element giving access to at least one value of at least one first parameter of a digital twin relating to at least one first object of interest of said audiovisual stream, said value being extracted from said digital twin during a capture of said audiovisual stream.

[0045] For example, the present application also relates to a recording medium readable by a processor of an electronic device and on which is recorded a computer program comprising instructions for implementing, when the program is executed by the processor, a method for enriched rendering of at least a first object of interest of a physical environment, said method comprising: - obtaining a multimedia element, said multimedia element providing access to at least a first value of at least a first parameter of a digital twin relating to at least said first object of interest; - obtaining at least a first audiovisual representation of at least a first portion of said digital twin taking into account said first value obtained; - a rendering of said first audiovisual representation in said physical environment.

[0046] The above-mentioned programs may use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0047] The information (or recording) media mentioned in the present application may be any entity or device capable of storing the program. For example, a medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means.

[0048] Such a storage means may for example be a hard disk, flash memory, etc.

[0049] On the other hand, an information carrier may be a transmissible carrier such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or other means. A program according to the invention can in particular be downloaded from an Internet-type network.

[0050] Alternatively, an information (or recording) medium may be an integrated circuit in which a program is incorporated, the circuit being adapted to execute or to be used in the execution of any one of the embodiments of at least one of the methods which are the subject of the present patent application.

[0051] Generally speaking, by obtaining an element, is meant in the present application for example a reception of this element from a communication network, an acquisition of this element (via for example user interface elements or sensors), a creation of this element by various processing means such as by copying, encoding, decoding, transformation etc. and / or an access of this element from a local or remote storage medium accessible to at least one device implementing, at least partially, this obtaining. 4. Brief description of the drawings

[0052] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:

[0053] [Fig.l] presents a simplified view of a system, cited as an example, in which at least certain embodiments of the enriched obtaining and rendering methods of the present application can be implemented,

[0054] [Fig.2] presents a simplified view of a device suitable for implementing at least certain embodiments of the method for obtaining and / or rendering enriched in the present application,

[0055] [Fig.3A] shows an overview of the method of obtaining the present application, in certain of its embodiments.

[0056] [Fig.3B] shows an overview of the method of obtaining the present application, in certain of its embodiments.

[0057] [Fig.4] presents an overview of the enriched rendering method of the present application, in certain of its embodiments.

[0058] [Fig.5] illustrates an example of obtaining a multimedia element and enriched rendering of a first object of interest from this multimedia element according to at least certain embodiments of the methods of the present application. 5. Description of the methods of implementation

[0059] The present application aims to provide a new solution, capable of facilitating communications, in particular asynchronous communications, between users of a telecommunications system. It can, for example, help a user to describe, optionally in addition to a message comprising audiovisual content, at least one object of interest in a physical, real environment (or alternatively a virtual representation of this environment) which may be diverse depending on the embodiments. It may for example be a factory, a garden, or a public place. This physical environment will also be called a “capture environment” in this application.

[0060] More specifically, the present application proposes, according to a first aspect, to create, at the initiative of a “producer” user, a multimedia element obtained in particular from an audiovisual stream (for example a video sequence) representing a first portion of the physical environment, this first portion comprising an object of interest (for the “producer” user in particular) or being located near such an object. The multimedia element comprises (or gives access to) information relating to the object of interest, obtained during the capture of the audiovisual stream, and coming from a digital twin. This information may for example relate to the current state of the representation of the object of interest in the digital twin during the capture of the audiovisual stream, and / or to its identification, and / or to its positioning. The information may comprise at least one audiovisual representation of this object of interest etc.

[0061] Optionally, the multimedia element may comprise (or provide access to) audiovisual content obtained from at least a portion of the captured audiovisual stream.

[0062] The digital twin may be, for example, a digital twin specific to the object of interest or a digital twin relating to at least the first portion of the physical environment (for example, a digital twin of the complete physical environment (i.e. as a whole) including optionally avatars of people present in the environment). Alternatively, the multimedia element may comprise, not the information relating to the object of interest in the digital twin but a link giving access to this information. The multimedia element therefore comprises information from a digital twin relating to at least one object of interest of the audiovisual stream and may optionally comprise audiovisual content originating at least partially from this stream and linked to this object of interest.

[0063] According to a second aspect, at least some data related to the created multimedia element can then be rendered (rendered) at a later time to a “consumer” user. The rendering can vary according to the embodiments. For example, in certain embodiments, it may involve a rendering of at least some of the information of the digital twin (such as parameter values) included in the multimedia element (or alternatively to which the multimedia element provides access) and a joint rendering of at least one audiovisual representation (i.e. audio and / or visual) of at least one object of interest to which at least one information of the digital twin obtained from the multimedia element relates. In some embodiments, the rendering may further comprise rendering (together with the digital twin information and the audiovisual representation) audiovisual content that the multimedia element comprises (or alternatively to which it provides access). As explained above, the information that the multimedia element comprises (or alternatively to which it provides access) was obtained during the capture of an audiovisual stream comprising an object of interest to which at least some of this information relates. The rendered content may in particular have been obtained from the captured audiovisual stream.

[0064] In some embodiments (e.g., when the information in the media element is time-stamped), the rendering of this information may, for example, be time-synchronized with the rendered audiovisual representation and / or with the optionally rendered content. In some embodiments, the rendering of data related to the media element may include a joint rendering of digital twin information included in the media element and current-time digital twin information equivalent to the rendered information included in the media element (such as the joint rendering of digital twin parameter values ​​included in the media element (and relating to the time of capture) and the current values ​​of these parameters).

[0065] For example, the rendering may include information from the digital twin representative of the position of the object of interest in the physical environment at the time of capturing the audiovisual stream. According to another example, the rendering may include a rendering of information from the digital twin relating to the state of the object of interest at the time of capturing and information from the equivalent digital twin (same parameter) relating to the current state (at the time of rendering) of the object of interest.

[0066] The rendering may also optionally include other information included in the multimedia element (for example at least one timestamp relating to the capture of the audiovisual stream, or information relating to the state of the digital twin during the capture of the audiovisual stream).

[0067] The content obtained from the captured stream may for example represent a first object of interest and a second object of interest of the physical environment, the second object of interest being for example an individual (real or an avatar), such as the producing user or a third party, present in the physical environment and designating (via visual or vocal indications for example) or interacting with the first object of interest.

[0068] Thus, thanks to the method of the present application, an individual can for example be filmed (in an audiovisual stream) while he points with his finger or his gaze at certain elements of the physical environment, which will then be rendered (for example during of a subsequent rendering of at least a portion of the film), in their position when the film was captured. Rendering can also be done in the physical environment, so as to render the elements designated "in situ" and optionally (in the case of mixed reality rendering) by highlighting their current position. Such an embodiment can help to highlight a movement of an object of interest between its capture and its rendering.

[0069] Such a method can thus help a producer user to generate explanations more simply and at the same time facilitate their understanding by a consumer user.

[0070] According to the embodiments, the multimedia element may be intended to be communicated (i.e. transmitted) to a third party (“recipient”) via a communication (for example asynchronous), or to be simply made available, for an enriched rendering of an object of interest of this multimedia element, to a potential third party “consumer”. For example, the enriched rendering may be implemented at the request of a third party (authenticated for example) or without even an initiative of the third party (for example it may be a continuous rendering (loop broadcast) or a rendering carried out following a detection of a passage near a device for enriched rendering of the object of interest.

[0071] By way of example, the present application may in particular be applied when objects are in motion during the capture of an audiovisual stream (the message of a “producer” user). Indeed, it may be important to save these movements for the proper understanding of the message. The present application proposes a solution making it possible, for example, to create (for subsequent enriched rendering) a “message” in situ taking into account possible changes, between the time of capture of the stream and the rendering of the “message”, in the physical environment where the stream containing this message was captured.

[0072] The present application will now be described in more detail in connection with [Fig. 1].

[0073] [Fig.l] represents a telecommunications system 100 in which certain embodiments of the invention can be implemented. The system 100 comprises one or more electronic devices, at least some of which can communicate with each other via one or more communication networks 180, 182, possibly interconnected (via for example an interconnection device 170 (also called a gateway or "gateway" according to the English terminology), such as a local area network or LAN (Local Area Network) and / or a wide area network, or WAN (Wide Area Network). Examples of networks may include a corporate or home LAN network and / or a WAN network of the internet type, or cellular, GSM - Global System for Mobile Communications, UMTS - Universal Mobile Telecommunications System, Wifi - Wireless, etc.

[0074] As illustrated in [Fig.l], the system 100 may also comprise one or more electronic devices, such as a terminal 110, 112, (such as a laptop, a smartphone, and / or a tablet), an augmented, mixed or virtual reality headset, a connected object 140, and / or a server 130, for example an application server, a storage device 150. The system may also comprise management and / or network interconnection elements (not shown). These electronic devices may be associated with at least one user 160, 162 (for example via a user account accessible by login), some of the electronic devices 110, 112 being able to be associated with the same user 160 or shared by several users 160, 162.

[0075] In the present application, a connected object (or "smart object" according to English terminology) is an electronic device having components (hardware or software) giving it the ability to perform a function (action, provision of a service, acquisition of information about its physical environment; etc.), means (such as a microcontroller) for controlling these components to perform this function, and means of communication within a communication network making it possible in particular to obtain and / or provide information related to the performance of the function. For example, a connected object can acquire information relating to its current physical environment and transmit it over the network, or receive a command from the communication network to perform an action.

[0076] Some of the devices of the system 100 can process (for example create and / or store and / or modify and / or manipulate and / or delete) information relating to a physical environment 120 where one or more objects 140, 142 are present. Some 140 of these objects can be connected objects belonging to the system 100 (therefore “devices” of the system 100), other objects 142 can be “unconnected” objects not comprising means of communication. It is noted that the notion of object here includes not only inert structures but also living structures, such as trees, animals or individuals.

[0077] The information relating to the physical environment may be representative of a current (current) state of the physical environment and / or of a state, an activity or a positioning (location, orientation, mobility, etc.) of at least some of the objects (connected or not) that it contains or which are located near this physical environment.

[0078] For example, certain information may relate to objects 146 not present in the physical environment but having an impact on this physical environment.

[0079] In the present application, the term "digital twin" of a real object will be used to designate a modeling of this real object by a set of information relating to this real object. This information may in particular be representative of its current state, and / or its links or interactions with other objects. Similarly, the term "digital twin" of a physical environment comprising at least one real object will be used to designate a modeling of this physical environment by a set of information relating to the physical environment in its entirety, and / or to at least some of the real objects that it contains and / or to their interactions. Thus, a digital twin is a digital representation, at a current time, of the object or the physical environment with which it is associated.In the same way that the physical environment may include, or contain, one or more objects, the digital twin of the physical environment may include digital twins of objects in the physical environment to which it corresponds, or have access, at least partially, to the digital twins of objects in the physical environment.

[0080] The digital twin can therefore evolve as the object or physical environment it represents evolves.

[0081] It is noted that the information relating to a real object can be acquired via the real object itself (if it has means of communication with at least one other device of the system 100) or via another object or device (such as for example a camera external to this object but having this object in its field of vision and therefore able to transmit information relating to a state or a movement of this object).

[0082] A digital twin may, according to another example, be supplied periodically or continuously with data collected via sensors placed on and / or in and / or near the real object or the physical environment, or with data resulting from an inspection by an operator, at at least a certain instant, of this object.

[0083] Examples of sensors include microphones, motion sensors, pressure sensors, eye-tracking devices, cameras (such as multi-sensor cameras capable of producing a 3D representation (e.g., wireframe) of a real object in the physical environment and of detecting and locating real objects in the physical environment).

[0084] In certain embodiments, a digital twin can be obtained and / or enriched by mapping (via spatial recognition for example)

[0085] Thus, unlike conventional digital modeling, a digital twin is configured to provide, at any time, information on the current operating state of the object or the physical environment to which it is connected.

[0086] [Fig.2] illustrates a simplified structure of an electronic device 200 of the system 100, for example the device 110, 112, 130 or 140 of [Fig.l], adapted to implement the principles of the present application. Depending on the embodiments, it may be a server, and / or a terminal.

[0087] The device 200 comprises in particular at least one memory M 210. The device 200 may in particular comprise a buffer memory, a volatile memory, for example of the RAM type (for “Random Access Memory” according to English terminology), and / or a non-volatile memory (for example of the ROM type (for “Read Only Memory” according to English terminology). The device 200 may also comprise a processing unit UT 220, equipped for example with at least one processor P 222, and controlled by a computer program PG 212 stored in memory M 210. At initialization, the code instructions of the computer program PG are for example loaded into a RAM memory before being executed by the processor P.Said at least one processor P 222 of the processing unit UT 220 can in particular implement, individually or collectively, any one of the embodiments of at least one of the methods for obtaining and / or enriching rendering of the present application (described in particular in relation to figures 3A, 3B, 4 and 5), according to the instructions of the computer program PG.

[0088] The device may also comprise, or be coupled to, at least one input / output module LO 230, such as a communication module, allowing for example the device 200 to communicate with other devices of the system 100, via wired or wireless communication interfaces, and / or such as a module for interfacing with a user of the device (also called more simply in this application “user interface”).

[0089] By user interface of the device, we mean for example an interface integrated into the device 200, or a part of a third-party device coupled to this device by wired or wireless communication means. For example, it may be a secondary screen of the device, a camera (allowing acquisition of gesture commands from an operator for example) or a set of speakers connected by wireless technology to the device.

[0090] A user interface may in particular be a user interface, called an “output” user interface, adapted to a rendering (or to the control of a rendering) of an output element of a computer application used by the device 200, for example an application running at least partially on the device 200 or an “online” application running at least partially remotely, for example on the server 140 of the system 100. Examples of output user interfaces of the device include one or more screens, in particular at least one graphical screen (touch screen for example), one or more speakers, a connected headset (in particular an augmented, mixed or virtual reality headset).

[0091] By rendering, we mean here a restitution (or “output” according to English terminology) on at least one user interface, in any form, for example comprising textual, audio and / or video components, or a combination of such components.

[0092] Furthermore, a user interface may be a so-called “input” user interface, adapted to acquiring a command from a user of the device 200. This may in particular be an action to be performed in connection with an item rendered by the device 200, and / or a command to be transmitted to a computer application used by the device 200, for example an application running at least partially on the device 200 or an “online” application running at least partially remotely, for example on the server 140 of the system 100. Examples of input user interfaces of the device 200 include a sensor, an audio and / or video acquisition means (for example a microphone, a camera (webcam), an eye-tracking device), a means for acquiring a command (key(s), for example a keyboard, button, mouse, actuator of a touch screen), etc.

[0093] Said at least one microprocessor of the device 200 may in particular be adapted to implement at least one of the methods of the present application.

[0094] Thus, said at least one microprocessor of the device 200 can in particular be adapted for obtaining a multimedia element, said obtaining comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - A creation of a multimedia element giving access to at least one value of at least one first parameter of a digital twin relating to at least one first object of interest of said audiovisual stream, said value being extracted from said digital twin during a capture of said audiovisual stream.

[0095] Thus, said at least one microprocessor of the device 200 can in particular be adapted for an enriched rendering of at least one first object of interest of a physical environment, an implementation of said enriched rendering comprising: - obtaining a multimedia element, said multimedia element giving access to at least one first value of at least one first parameter of a digital twin relating to at least said first object of interest; - obtaining at least a first audiovisual representation of at least a first portion of said digital twin taking into account said first value obtained; - a rendering of said first audiovisual representation in said physical environment.

[0096] It is noted that in certain embodiments, the device may be adapted to both the implementation of the obtaining method and the enriched rendering method of the present application. In other embodiments, on the contrary, the obtaining of the multimedia element and the enriched rendering may be implemented by different devices.

[0097] Some of the above input-output modules are optional and may therefore be absent from the device 200 in certain embodiments. In particular, if the present application is sometimes detailed in connection with a device communicating with at least one second device of the system 100, at least one of the obtaining and / or rendering methods may also be implemented locally by a device (for example when the digital twin of the physical environment is local to the device and the latter locally implements both the obtaining method and the enriched rendering method of the present application (at different times for example)).

[0098] On the contrary, in some of its embodiments, the at least one of the obtaining and / or rendering methods can be implemented in a distributed manner between at least two devices 110, 112, 130, 140, 150, 170 of the system 100.

[0099] The term "module" or the term "component" or "element" of the device is understood here to mean a hardware element, in particular wired, or a software element, or a combination of at least one hardware element and at least one software element. The methods according to the invention can therefore be implemented in various ways, in particular in wired form and / or in software form.

[0100] Certain embodiments of the obtaining method 300 of the present application are now presented in more detail. Certain embodiments are presented in connection with Figures 3A and 5, other embodiments are presented in connection with Figures 3B and 5. The same numbering is used in Figures 3A and 3B to designate identical or similar elements. The method 300 of Figures 3A and / or 3B can for example be implemented by the electronic device 200 illustrated in [Fig.2].

[0101] Certain embodiments of the rendering method of the present application are then detailed in connection with Figures 4 and 5. The rendering method, in the embodiments detailed in connection with Figures 4 and 5, can in particular be used to process elements resulting from the method for obtaining Figures 3A and 5 or Figures 3B and 5.

[0102] [Fig.5] illustrates, in a simplified manner, an example of implementation of the obtaining and rendering methods at different times.

[0103] As shown in [Fig.3A], the method 300 may comprise (step 320) temporally synchronized joint recordings 321, 322 of audiovisual content (recording 321) obtained from an audiovisual stream capturing a scene (audio and / or video) of a physical environment, and information from at least one digital twin (recording 322), in its current state (i.e. during the capture of the stream). This information relates in particular to at least one physical object present or in the vicinity of the captured scene.

[0104] In the example of [Fig. 5], a recording 321 representing in particular a first interlocutor 160 is triggered at time t1. In the illustrated example, this recording 321 corresponds to an acquisition of a video sequence (or alternatively of a plurality of images) capturing the scene 510-t1. This scene takes place in a portion of a physical environment (such as the physical environment 120 of [Fig. 1]). The portion of the environment comprises several objects (140, 142, 144, 160) whose respective positions (140-t1, 142-t1, 144-t1, 160-t1) at time t1 are illustrated in [Fig. 5]. Thus the interlocutor 160 is himself an object of interest in this example.

[0105] In the example of [Fig.5], the recording representing in particular the interlocutor 160 and the recording of the parameters of the digital twin of the environment both stop at a time t2.

[0106] At least some information relating to the capture of the scene (such as a location (position, orientation, field width, etc.) of the capture) can be recorded (saved) in association with the video sequence.

[0107] Furthermore, at least some information from the digital twin of the physical environment is recorded (stored) in parallel with the recording of the video sequence. This information may, for example, relate to objects present in the captured scene (such as objects 140, 142, 144, 160).

[0108] As illustrated in [Fig.3A], in certain embodiments, the same command 310 can trigger both the recording 321 of the audiovisual content and the recording 322 of parameters of the digital twin of the environment). Similarly, the same command 330 can trigger both the stopping of the recording of the audiovisual content and that of the parameters of the digital twin.

[0109] In other embodiments, separate commands may control each of the recordings.

[0110] A command to start or end (stop) recording may for example be received via an input interface of the device 200 (for example via an actuation of a physical element of the device such as a physical button) or via an interface communication interface of the device 200 (for example, a command coming from the telephone or a headset of the “producer” user and received via a wireless communication interface of the device 200).

[0111] For the sake of simplicity, it has been considered in [Fig. 5] that between the times t1 and t2 the objects 140, 142, 144 and 160 had not moved. Of course, in certain embodiments, at least some of the objects of interest (or other elements of the captured scene) (in particular the speaker 160 in our example) may be moving during the recording. In certain embodiments, the origin, direction and / or angle of the shot may also fluctuate during the recording (scanning of a scene or zooming in on a speaker for example).

[0112] According to the embodiments, the recording 332 of the parameters of the digital twin can be carried out at regular intervals during the recording of the audiovisual content, and / or during particular events (start and / or end of recording, change of state of a parameter (fluctuation of its value), such as a status and / or a position, occurrence of an unexpected event (intrusion alert, and / or fire, and / or breakdown of equipment in the physical environment, etc.). Recording several times of a parameter of an object of interest can help to highlight a change of state and / or a movement and / or a modification, and / or an alteration of an object of interest during its capture.Recording the value of a parameter of the digital twin only when it fluctuates can help to limit the memory occupation of the recording of the digital twin, which can be important in certain embodiments, for example in certain embodiments where the digital twin is very complex (so in the updates require for example significant processing times), and / or where the multimedia element must be generated very quickly (following the occurrence of a critical event for example), and / or where the obtaining method is executed on a device having very limited memory and / or processing capacities).

[0113] In certain embodiments where the multimedia element comprises one (or more) video components, the recording 321 of the audiovisual content may limit the recording to an area of ​​the scene located near a reference point (for example an object of interest such as a speaker present in the scene). For example, the recorded area may be a circular or rectangular area centered on the speaker whose radius, respectively the width / height, is a constant corresponding to a “maximum” distance (used as a threshold) from the speaker. The recording of the audiovisual content may also comprise a recording of at least one metadata related to the captured audiovisual stream (such as a start and / or end timestamp of recording, the links with certain elements of the digital twin, the 3D coordinates or the states of the elements in the space at the time of recording, the characteristics of the device that captured the audiovisual content, a type of encoding of the audiovisual content, etc.).

[0114] In the example of [Fig.3A], the method may comprise a recording 332 of all the parameters of the digital twin then a filtering 360 of the parameters of the recorded digital twin, to retain only some of these parameters, taking into account the recorded audiovisual content (and possibly filtered 371 also). These filterings of recorded parameter(s) as well as of the recorded audiovisual content may be optional in certain embodiments. They are detailed more precisely below.

[0115] In certain embodiments, the method may comprise a detection 340 of at least one object of interest in the recorded audiovisual content. This detection 340 may implement an analysis 341 of the audiovisual content, via for example image processing and / or audio processing techniques. For example, the analysis 341 may comprise image processing (motion analysis for example) to detect 342 a gestural, verbal or visual designation (by a look for example), by an interlocutor (present in at least one captured scene), of at least one region of the acquired scene. A more detailed image processing (shape detection for example) may further be carried out on the designated region to detect one or more objects therein (potentially constituting one or more objects of interest).Cumulatively or alternatively, the analysis 341 may comprise an audio analysis to detect 342 in a speech of the interlocutor one or more keywords semantically corresponding to objects detected (by image processing) in the acquired scene. The interlocutor himself may potentially constitute an object of interest.

[0116] This analysis may for example result in a determination (or detection) of at least one object(s) of interest in the recorded audiovisual stream. In certain embodiments, it may also involve detecting in the acquired scene (via shape recognition for example) at least one particular object assimilated (for example by a prior configuration) as being of interest to the speaker and / or to a potential consumer of the video sequence. This may for example be a critical object (from the point of view of its operation and / or due to its dangerousness), and / or an object occupying a significant part of the captured scene, a moving or noisy object in the captured scene and / or a dangerous zone not to be crossed and / or not to be exceeded for persons who do not have the authorization, etc.It can also be portions of an object of interest more particularly likely to designate and / or interact between another object of interest in the scene (as for example in the case where one of the objects of interest is an interlocutor present in the scene, his hands, his legs, his eyes, etc.), that it is. therefore to follow (track) with precision (for example to detect other objects of interest as explained above). For example, those portions of the object of interest more likely to designate and / or interact with another object of interest in the scene can be followed with more precision than other portions of the object of interest.

[0117] Once at least one object of interest has been detected, the method may comprise a matching 350 of the object of interest detected in the stream with at least one piece of information (parameter) of the recorded digital twin (recording 322).

[0118] The object of interest may be a connected object 140 having its own digital twin, or a connected object 140 to which at least one parameter of the digital twin of at least one portion of the environment relates. The parameters corresponding to the object in the digital twin may include its identifier, and / or descriptors of this object such as its position (location and orientation), its dimensions, its shape, its color, its state, its movement (speed, acceleration, direction, rotation, etc.). The object of interest may also be a non-connected object 142, in which case an estimate of certain parameters describing it (for example its position relative to connected objects of the digital twin) may be extracted from the acquired scene.

[0119] The recorded parameters of the digital twin may correspond not only to objects of interest in the captured scene but also to objects (such as objects 146 located near the captured scene) represented in the digital twin and having an impact on the captured scene, such as objects at the origin of a source of light or noise in the captured scene.

[0120] As illustrated in [Fig.3A], the method 300 may comprise a 360 filtering of at least one recorded parameter of the digital twin. The 360 ​​filtering may comprise a deletion 361 of at least one recorded parameter of the digital twin. For example, it may involve deleting at least one recorded parameter not corresponding to any object of interest in the content, or corresponding to an object previously declared (by parameterization for example) as of no interest (such as a broom or shavings in an industrial environment for example).

[0121] It may also involve selecting the objects of interest to be retained based on their position (or their orientation) relative to a reference position (or a reference orientation). For example, in the case of a video sequence, the method may comprise filtering the recorded parameters of the digital twin to retain only those relating to objects of interest located near a reference object present in the video (whose position is considered to be a reference position). This reference object may, for example, be a speaker delivering an audio and / or gestural message in the audiovisual content. For example, all digital twin parameters relating to objects of interest located at a distance greater than a first distance from this object (for example a constant distance fixed by parameterization) can be deleted.

[0122] In the case of a video sequence, in embodiments where the position and / or orientation of an interlocutor is used as a reference, and where the latter moves, all the parameters of the digital twin relating to objects located at an instant in the sequence near the interlocutor can for example be preserved.

[0123] As illustrated in [Fig.3A], the method may comprise a modification 370 of the audiovisual content. This modification may be optional in certain embodiments. The modification may for example comprise a filtering 371 of the content to remove certain portions (temporal or spatial) and / or at least one component (audio, visual, etc.) of the audiovisual content so as to retain, for example, only image portions representing an object of interest, and / or only a portion of an audio sequence of the speaker relating to an object of interest, and / or only an audio component and / or only a visual component.

[0124] The modification 370 may also comprise in certain embodiments obtaining modified audiovisual content corresponding to an extraction 372 of a portion (to be kept in this modified content) of the captured content (for example an extraction of a silhouette by “cutting out” the silhouette in the captured scene).

[0125] The modification may comprise (in addition or alternatively) a transformation 373 of a portion of the content. For example, after cutting out a silhouette of a character in the audiovisual content, it may be replaced by an avatar of this character, or a hologram of this character. The avatar may in particular be correlated to the captured character, to its movements and its expressions.

[0126] Related data may also be saved in association with the (possibly modified) content, for example in the form of metadata. Examples of such data may include an identifier of an object (or portion of an object) present in the (possibly modified) audiovisual content, information characterizing the capture, positioning information of an object of the modified content, etc. It is noted that this may be absolute positioning information (for example a GPS position) or relative positioning information (for example a distance and / or an orientation relative to an object of the digital twin). In particular, for a character, the positioning information may include information relating to the position and / or movement of certain parts of the character's body (such as at least one of its limbs (hands, arms, legs, feet, etc.), its head, its mouth, its eyes, its facial expression, etc.

[0127] In some embodiments, the captured audiovisual content may be 3D content.

[0128] Filtering of the content may be similar or consequential (correlated) to filtering of the recorded parameters of the digital twin and vice versa, where filtering performed in the content may result in filtering to be performed on the corresponding parameters of the digital twin and vice versa.

[0129] In some embodiments, the filtering 360 of the digital twin may comprise a rendering 362 of at least some elements of the digital twin. Similarly, the modification 370 of the audiovisual content may comprise a rendering 374 of at least a portion and / or a component of the audiovisual content.

[0130] This rendering 362, 374 of at least certain elements of the digital twin and / or of at least one portion and / or one component of the audiovisual content can be carried out before and / or after an automatic deletion of elements of the digital twin and / or of at least one modification of a portion and / or of a component of the audiovisual content (as explained above (when it exists)). Such embodiments can make it possible to obtain a validation 363, 375 by an operator (for example by the interlocutor present on the recorded content) of the elements to be kept of the digital twin and / or of the portions and / or components to be kept of the audiovisual content.In particular, the method may comprise receiving (via a user interface or a communication interface) a command to delete certain elements from the digital twin and / or to add certain deleted elements, and / or a command to delete certain portions and / or components of the audiovisual content and / or to add certain deleted portions or components.

[0131] These two renderings may be optional in certain embodiments. These two renderings may be performed jointly or separately. The deletion, respectively the addition, of an object of interest may in particular cause the deletion, respectively the addition, of the parameters of the digital twin associated with the object of interest.

[0132] The method 300 may also comprise a creation 380 of a multimedia element comprising the recorded elements of the digital twin (possibly filtered). The multimedia element may also comprise the audiovisual content (possibly modified).

[0133] The audiovisual content and the recorded elements of the digital twin may in particular be associated 381 in the multimedia element taking into account their respective acquisition times so as to enable them to be synchronized temporally. For example, in embodiments where the elements of the digital twin and the audiovisual content are time-stamped, the association may take into account their respective timestamps. More precisely, the association may take into account a proximity (for example an equality) between their respective timestamps. The association 381 may also take into account the objects of interest to which they relate.

[0134] As illustrated in [Fig.3A], the method may also comprise a storage 382 of the created multimedia element (for example for a subsequent rendering on the device implementing the obtaining method, or for a rendering from a device coupled thereto) and / or a transmission 383 of this multimedia element to at least one other device (for example another device of the system 100 implementing the rendering method described in connection with FIGS. 4 and 5).

[0135] It is noted that in certain embodiments, the transmission 383 of the multimedia element may comprise a transmission of at least one designation of at least one recipient user of the multimedia element. The multimedia element (or a link to such an element) may for example be transmitted in an electronic message (via various electronic messaging tools depending on the embodiments) to at least one recipient user.

[0136] This transmission of a designation of at least one recipient may be optional in certain embodiments.

[0137] In certain embodiments (not illustrated in [Fig.3A]) the filtering 360 of the recorded elements of the digital twin (respectively the modification of the content 370) may occur after the creation of the multimedia element. In such a case, the deletion of a component of the multimedia element may result in the deletion of a component associated with it in the multimedia element.

[0138] Certain embodiments have been presented above in connection with Figures 3A and 5. In the embodiments illustrated in [Fig.3A], the analysis of the content and the detection of objects of interest via this content, and therefore the filtering of the twin accordingly, are carried out after the end of the recording of the content (i.e. after the time t2 according to [Fig.5]). Such embodiments can make it possible to limit the processing during the content recording phase and therefore to limit the load peaks of the processor of the device 200.

[0139] Other embodiments are now detailed in connection with Figures 3B and 5.

[0140] [Fig.3B] thus illustrates a variant of the embodiments illustrated in [Fig.3A]. In this variant, in particular, the joint recording 390 differs from the joint recording 320 of [Fig.3A]. Thus, the audiovisual content is analyzed as it is recorded 391 (i.e. between times t1 and t2 according to [Fig.5]), to detect 392 objects of interest, match 393 information from the digital twin relating to these objects of interest and record 394 this information as it is recorded. The detection 392 and the matching 393 may be similar to those of the detection 341 and the matching 350 described in connection with [Fig.3A].

[0141] Embodiments such as illustrated in [Fig.3B] can help to limit the quantity of information from the recorded digital twin and therefore make it possible to limit the memory size necessary for implementing the obtaining method.

[0142] As explained in connection with [Fig.3A], the recording 394 of the parameters of the digital twin can be carried out at regular intervals during the recording of the audiovisual content, and / or during particular events (start and / or end of recording, and / or fluctuation of a parameter, and / or occurrence of an unexpected event for example). Such modes can help to limit the size of the multimedia element created.

[0143] The recording start and end commands 310, 330 may be similar to those described in [Fig.3A].

[0144] In the embodiment illustrated in [Fig.3B], a modification 370 of the acquired video content and / or a filtering 360 of the information from the digital twin can also be carried out once the recording is finished. Depending on the embodiments, this filtering can be identical to or different from the filtering of the embodiments of [Fig.3A]. In particular, since the information representative of the recorded digital twin relates only to the objects of interest of the recorded video content, it may be optional to further filter the recorded digital twin information. Optionally, rendering and validation by an operator can be implemented in a similar manner to what has been described in relation to [Fig.3A].

[0145] The creation of the multimedia element 380 may also be similar to that described in connection with [Fig.3A].

[0146] The recording of audiovisual content captured in the real world has been detailed above. Alternatively, it may be a capture of a scene from a virtual environment, the interlocutor in this case not being a human interlocutor but an avatar present in the virtual environment from which the scene originates.

[0147] [Fig. 4] illustrates certain embodiments of the rendering method 400 of the present application. The method 400 may for example be implemented by an electronic device such as the device 200 illustrated in [Fig. 2]. In the remainder of the description, for simplicity, we will refer to the device implementing the rendering method 400 as a “rendering device”, even if the device implementing (at least partially) the rendering method 400 (and controlling the rendering) and the device actually performing the rendering may be different. Indeed, as described previously in connection with [Fig. 2], the device implementing the rendering method may be coupled to another device performing the rendering (such as a secondary screen). In the remainder of the application, we will refer to the “rendering environment” to designate the physical environment of the device performing the rendering (and also for simplicity, as explained above, for designate the physical environment of the device in which the rendering process is implemented).

[0148] As illustrated in [Fig.4], the method 400 may comprise obtaining 410 at least one multimedia element. For example, it may involve receiving this multimedia element via a communication interface of the rendering device or reading access to a storage area, local or remote, of this multimedia element, such as a database (for example the element 150 of [Fig.l]), such as a database common to several devices of the system 100). The multimedia element (or a link to such an element) may for example be included in an electronic message received by a user of the rendering device.

[0149] As illustrated, the method may comprise an extraction 420 of data from (and / or from) the obtained multimedia element.

[0150] For example, the extraction 420 may comprise an extraction 422 (of and / or from the multimedia element) of information from a digital twin, at least some of this information relating to at least one object of interest of a physical environment. The information relating to an object of interest may in particular comprise an identifier of the object of interest. According to the embodiments, the physical environment to which the multimedia element refers may be the rendering environment or another environment.

[0151] For example, in some embodiments, the rendering method to be implemented to render (render) a multimedia element representing and / or giving access to information (at a time preceding the current time) of a digital twin of at least one object of interest located in the rendering environment, and optionally giving access to audiovisual content acquired (at this previous time) at the same location.

[0152] In certain embodiments, the rendering method can be implemented to restore (render) a multimedia element received from a third party and representing information (at a time preceding the current time) of a digital twin of at least one object of interest located in a physical environment remote from the rendering environment. Such embodiments can for example be applied for a transmission of a report (including this multimedia element) from an operator to a remote supervisor, or for a transmission of instructions (including this multimedia element) from a trainer to an operator (having to apply these instructions in a rendering environment close to the physical environment of capture of the instructions (case of twin factories for example).

[0153] Optionally, the extraction 420 may comprise an extraction 421 of (and / or from) the multimedia element of an audiovisual content and / or an extraction 423 of other complementary information, for example in connection with the content, such as metadata associated with the content and describing, for example, the context in which the content was shot, and / or at least one identifier and / or positioning information, in the content, of the object of interest in the physical environment.

[0154] Some of this data (including the audiovisual content itself and at least some of its metadata) may be optional in certain embodiments. Embodiments where the multimedia element does not contain audiovisual content may make it possible to limit the size of the multimedia element and consequently to limit the associated processing time, which may contribute to obtaining a smoother rendering of information from the digital twin than when, in conjunction with their rendering, a rendering of audiovisual content (for example video) is performed.

[0155] The additional information may, for example, make it possible to associate information from the digital twin extracted from the multimedia element with an object of interest present in the content. They may also make it possible, according to another example where the rendering environment corresponds to the physical environment to which the multimedia element relates, to determine a point of view (positioning, and / or orientation, and / or field width, etc.) identical (or almost identical) to the point of view of capturing a scene of this content (and / or from which it originates).The at least one positioning information may comprise, for example, in the case where the audiovisual content has a visual component (such as a video sequence), a designation of at least one image of the content, as well as the coordinates of at least one pixel representing the object of interest in this image (for example an offset in English terminology) relative to a reference pixel (for example the top left corner) of this image).

[0156] In some embodiments, the rendering associated with the multimedia element (and which will be described below) can be performed at the command of a user (for example a recipient of a message including the multimedia element (or a link to such an element)).

[0157] In other embodiments, as illustrated in [Fig.4], the rendering can be performed automatically. In such embodiments, it can be performed systematically upon obtaining the multimedia element, or conditionally, when certain criteria are met. Certain criteria may relate to a possible prior validation by an operator. Furthermore, certain criteria may relate to the current context of the rendering device. Thus, as illustrated in [Fig.4], in certain embodiments, the method may comprise a verification 430 of adequacy of the current context of the rendering device for rendering the multimedia element. More precisely, the verification may comprise an acquisition 431 of contextual data from the rendering device. This may involve, for example, acquiring data relating to the positioning of the rendering device (or a device coupled thereto and implementing the actual rendering), and / or to detect a possible presence in the vicinity of the rendering device (or a device coupled thereto and implementing the actual rendering), and / or to identify and / or physically recognize this person, and / or to acquire a current timestamp. The method may also comprise a verification 432 of rendering criteria based on at least some of the acquired contextual data. For example, the rendering may only be carried out when at least one of the verified criteria (or all, depending on the embodiments) are met.

[0158] According to a first example, applicable in particular when the rendering environment corresponds to the capture environment, one of the criteria to be verified may be a point of view of the rendering device (rendering) identical to the point of view of a capture device at the origin of at least part of the multimedia element.

[0159] According to a second example, the rendering may only be performed upon detection of a presence in the vicinity of the rendering device (or of a device coupled thereto and implementing the actual rendering). According to a third example, applicable in particular when the rendering environment corresponds to the capture environment, the rendering may only be performed upon detection of a presence in the vicinity of the current position of at least one object of interest to which at least some of the data extracted from the multimedia element relate and / or in the vicinity of a position of such an object, as extracted from the multimedia element (i.e. simply the current position of the object of interest during the capture of the audiovisual stream via which it was detected).A criterion may for example be the presence of the recipient of the message giving access to the multimedia element, or of any individual, at a distance less than a first distance (such as a constant value configured beforehand and used as a threshold value) from the rendering device, and / or of at least one object of interest to which data extracted from the multimedia element and / or a reference point of the physical (or virtual) environment are related, such as a portion of the physical environment corresponding to the shooting of the content.

[0160] Thus, in embodiments where the rendering environment corresponds to the capture environment, the rendering can be proposed via a visual or vocal message to an individual approaching a location corresponding to a current positioning of at least one of the objects of interest of the content, and / or to a positioning of at least one of the objects of interest of the content during the capture of the content, and / or a location corresponding to the position of the shot during the initial recording of the content, (or be carried out automatically when the individual approaches).

[0161] According to another example, applicable in particular when the rendering environment corresponds to the capture environment, a rendering criterion may be a presence in the current (current) physical environment of at least one object of interest of the multimedia element. Such an embodiment can help to limit the rendering of “obsolete” messages, thus causing unnecessary consumption of CPU resources on the one hand and unnecessary disruption for a consumer user on the other hand. In other embodiments, on the contrary, such a criterion may not apply, so as for example to help a consumer user to notice the disappearance of an object from the physical environment.

[0162] In some embodiments, where at least one criterion relates to an authenticated individual, the verification 432 of rendering criteria may comprise an authentication of the individual whose presence has been detected or requesting a rendering of the multimedia element.

[0163] It is noted that in certain embodiments the rendering criteria applicable to the rendering of information from the digital twin and to the rendering of the audiovisual content may differ from each other.

[0164] As highlighted above, the verification 430 of adequacy of the current context of the rendering device to the rendering of the multimedia element may be optional, at least in certain embodiments.

[0165] As illustrated in figures 4 and 5, the method 400 may comprise a rendering 440 of at least certain data of the multimedia element (for example after verification of rendering criteria as set out above).

[0166] The rendering may in particular comprise a joint rendering of at least one portion of the audiovisual content extracted from the multimedia element (rendering 441) and of at least one of the pieces of information of the digital twin extracted from the multimedia element (rendering 442);

[0167] For example, one of the objects of interest rendered may be an interlocutor (such as the user producing the audiovisual content).

[0168] The rendering of at least one object of interest in the content may also be conditional. For example, it may take into account rendering criteria similar to certain criteria to be verified cited as examples above (see verification 432 of rendering criteria). For example, only objects of interest whose positioning information in the multimedia element corresponds to a position close to the current position of a user “consuming” the multimedia element may be rendered.

[0169] In the example of [Fig.5], the rendering 440 of the multimedia element is carried out between the times t3 and t4 in the presence of the “consumer” user 162-t3. The rendering 440 of the multimedia element comprises a rendering 441 of an audiovisual content (a video sequence) 160-t1 extracted from the multimedia element and representing the first speaker 160 between the times t1 and t2. For example, it may be a clipping of the speaker in the captured video sequence, or an avatar representing the speaker and animated to take into account the movements made by the speaker during the capture of the video sequence. According to Figures 4 and 5, a rendering 442 of at least some information from the digital twin is performed jointly (for example synchronized) with the rendering of the speaker 160. The rendering 442 may for example include an audiovisual representation (virtual or real) of at least one object of interest 140, 160 (elements 140-tl, and 160-tl respectively). Depending on the embodiments, this visual representation may correspond to an extraction of a portion of the captured audiovisual content or even to a virtual representation generated from the information from the digital twin during the capture. For simplicity, we will speak hereinafter in both cases of “captured” visual representation (or rendering of an object “as captured”)).In the case where the rendering device is adapted to a rendering of the “augmented reality” or virtual reality type, the rendering of the object(s) of interest as captured (140-tl, 160-tl) can for example be carried out in “superposition” with their current rendering / appearance 140-t3, 160-t3. Thus, the consumer user 162 can for example consume a scene in which the speaker 160-tl appears designating the object of interest 140-tl, the same scene also comprising the element of interest 140-t3 in its current position. The rendering of an object of interest (and / or its representation (graphic or sound) can for example give access to the parameter of the digital twin extracted from the multimedia element and concerning this object of interest.

[0170] In embodiments where the rendering comprises a rendering of an at least partial visual representation of the digital twin (in its previous and / or current state), the consumer user can for example thus see in situ the recommendations of an interlocutor present in the audiovisual content (for example 3D) rendered (avatar or more realistic reconstruction of his face and / or his body). The consumer user can possibly move in the physical environment (actually or virtually in the case of a virtual representation of this physical environment) near one of the objects of interest rendered for example) to clearly understand and see the important elements mentioned by the interlocutor.Thus, the present application offers new possibilities to a “consumer” user compared to prior art solutions such as video recordings where the position, orientation and shooting angle of the video recording are, when they are reproduced, only those of the capture.

[0171] According to one example, moving objects of interest related to the content may appear in 3D as the content is consumed. In some embodiments, rendering an object of interest may include highlighting (via an audible and / or visual indicator, such as a highlight, a particular color, text, etc.) a fluctuation of at least some information (position, state, etc.) relating to the object of interest. For example, the method may understand a comparison between the current state of these objects in the real environment and their state when captured and, in the event of fluctuation, the objects concerned can be rendered with a highlighting of this fluctuation, so as to alert a “consumer” user of this fluctuation.

[0172] The indication may be rendered on the object of interest as captured and / or (in embodiments where the current scene is rendered in mixed reality) on the object of interest in its current "state".

[0173] Such embodiments, which make it possible to highlight objects of interest present in the initial scene which have been deleted, or on the contrary which are still present at the time of rendering in the current scene, can help a consumer user to judge whether the rendered multimedia element is of interest or not in the rendering environment.

[0174] Furthermore, such embodiments may help a user to find an object that has been moved since capture and is possibly outside the portion of the environment currently viewed by the consuming user.

[0175] In the case of a periodic recording of information from the digital twin in the multimedia element, the method may comprise an interpolation of the values ​​of the digital twin parameters during rendering (for example an interpolation of the positions of the objects of interest (and their orientation) between two recordings).

[0176] It is noted that if [Fig.5] illustrates in identical time intervals [t1; t2] and [t3; t4], for simplicity, in certain embodiments these intervals may be different. In this case, the method may comprise an acceleration or on the contrary a slowing down of the rendering of the multimedia element (rendering of the information of the digital twin and optionally of the content) to correspond to a desired restitution duration. In addition, the user consuming the content may have the possibility of acting on the rendering of the multimedia element (pausing, and / or rewinding, and / or slowing down the restitution speed (for a restitution in particular with a speed lower than a “nominal” speed corresponding to the speed of capture of the content) and / or on the contrary accelerating the restitution speed (for a restitution with a speed higher than the nominal speed in particular).In such embodiments, the rendering of the digital twin information may be adapted accordingly, such that the digital twin information remains temporally synchronized with the rendered content.

[0177] The enriched obtaining and rendering methods can find applications in many fields, professional or personal, to deal for example with situations where a “producer” user must explain a situation to a “consumer” user in relation to the physical environment which surrounds him. can thus cite areas such as industry, maintenance, digital business, archiving and sharing of information, and / or geolocation

[0178] According to a first example, the methods of the present application can offer an alternative to sending an audio message between operators of a production line (“When you arrive tomorrow and you are in front of the cutter, the one at the back of the factory, on the right, be careful: the activation lever of the cutter on the left is very difficult to lower. It is dangerous. Instead, use the button that you will find below the emergency zone. Not the push button but the small one right next to it”). Thus, thanks to the invention, it is possible to propose a rendering of a multimedia message that more easily highlights an object of interest (here the “small push button”) and thus avoid errors of judgment that could have serious consequences for a “consumer” operator.

[0179] According to a second example, the methods of the present application may help describe an object of interest in a museum or historic building.

[0180] According to a third example, the methods of the present application can help in the conditional rendering (after the success of a “consumer” user in solving a puzzle for example) of a step message during a geolocation game.

[0181] In a variant (not illustrated), the multimedia element can be enriched during / after an implementation of the rendering method by applying (again) the method for obtaining a multimedia element. For example, to the information of a first digital twin and to the optional audiovisual content already contained in the multimedia element, and representative of at least a first object of interest during a first capture period, other information of a second digital twin and optionally a second audiovisual content, representative of at least a second object of interest during a second capture period according to the obtaining method already described, can be added. Such a variant can for example allow a consumer of a multimedia element to respond to a user who is a producer of this multimedia element by transmitting to him an enriched version of this multimedia element.

Claims

Claims

1. Method for obtaining a multimedia element, said method comprising: - Obtaining an audiovisual stream representing at least a portion of a physical environment; - Creating a multimedia element giving access to at least one value of at least one first parameter of a digital twin relating to at least one first object of interest of said audiovisual stream, said value being extracted from said digital twin during a capture of said audiovisual stream^

2. Obtaining method according to claim 1 where said multimedia element further comprises at least one audiovisual content obtained from said captured audiovisual stream.

3. Obtaining method according to claim 2 where said audiovisual content is obtained by extraction from said captured audiovisual stream.

4. Obtaining method according to claim 2 or 3 where said audiovisual content is obtained by transforming at least a portion of said captured audiovisual stream.

5. Obtaining method according to one of claims 1 to 4 where said first object is selected, from among the objects of interest detected in said audiovisual stream, taking into account a current position and / or a current orientation of said first object in at least one scene of said audiovisual stream relative to at least one position and / or an orientation of at least one second object of interest of said audiovisual stream.

6. Obtaining method according to claim 5 where said first object, respectively said second object, is selected, from among the objects of interest detected in said audiovisual stream, taking into account a visual and / or auditory similarity between said first object, respectively said second object, and at least one reference object.

7. Obtaining method according to claim 5 or 6 where said first object, respectively said second object, is selected, from among the objects of interest detected in said audiovisual stream, taking into account a visual and / or audio designation of said first object, respectively of said second object, in at least one scene of said audiovisual stream.

8. Obtaining method according to one of claims 5 to 7 where said first object, respectively said second object, is selected, from among the objects of interest detected in said audiovisual stream, taking into account a manual designation of said first object, respectively said second object, via a user interface.

9. Obtaining method according to one of claims 1 to 8 where said created multimedia element comprises values ​​relating to at least two parameters of said digital twin and where said method comprises, after said creation, a filtering of said values ​​of said parameters.

10. Method for enriched rendering of at least one first object of interest of a physical environment, said method comprising: - obtaining a multimedia element, said multimedia element giving access to at least one first value of at least one first parameter of a digital twin relating to at least said first object of interest; - obtaining at least one first audiovisual representation of at least one first portion of said digital twin taking into account said first value obtained; - rendering said first audiovisual representation in said physical environment.

11. An enriched rendering method according to claim 10 wherein said multimedia element provides access to at least one audiovisual content relating to said first object of interest and where the method comprises a rendering, jointly with said rendering of said first audiovisual representation, of at least a portion of said audiovisual content.

12. An enriched rendering method according to claim 10 or 11 wherein the method comprises rendering, jointly with said rendering of said first audiovisual representation, a second audiovisual representation of at least a second portion of said digital twin taking into account at least one current value of said parameter of said digital twin.

13. The enriched rendering method of claim 12 wherein said method comprises highlighting said at least one first object of interest in said first, respectively second, audiovisual representation.

14. An enhanced rendering method according to claim 13 wherein said rendering of said first audiovisual representation is implemented conditionally taking into account a presence of a speaker in said physical environment.

15. An enriched rendering method according to claim 14 wherein said rendering of said first audiovisual representation is implemented conditionally taking into account a profile of said present speaker.

16. An enhanced rendering method according to claim 14 or 15 wherein said rendering of said first audiovisual representation is implemented conditionally taking into account a geographical proximity between said present speaker and equipment in said physical environment.

17. An enriched rendering method according to claim 16 wherein said equipment belongs to a group comprising: - a rendering device implementing said rendering method; - a device for rendering said first and / or second audiovisual representation; - said at least one first object of interest.

Citation Information

Patent Citations

  • Systems and methods for spatial conversion and synchronization between geolocal augmented reality and virtual reality modalities associated with real-world physical locations

    US11656835B1

  • Systems and methods for attaching synchronized information between physical and virtual environments

    US20200160607A1

  • System and method for immersive training using augmented reality using digital twins and smart glasses

    US20240071003A1