Method for processing an audiovisual stream, corresponding electronic device and computer program product.
The method addresses the challenges of high computing power in 3D image reproduction by detecting and spatializing objects within a volumetric context, dynamically adapting graphic parameters, and enhancing user interaction on portable devices, achieving a more immersive and efficient 3D visualization.
Patent Information
- Application Number
- FR2024003492
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-04
- Publication Date
- 2025-10-10
AI Technical Summary
Existing systems for capturing and reproducing three-dimensional images face challenges with high computing power and storage requirements, limiting their use on portable devices, and lack intuitive and interactive methods for volumetric representation.
A method for processing audiovisual streams that includes detecting objects of interest within a volumetric capture context, spatializing these objects considering a volumetric rendering context, and dynamically adapting graphic parameters to enhance the rendering process, allowing for more immersive and interactive 3D visualizations on portable devices.
The method provides a more faithful and immersive representation of objects, improves detection precision, reduces computational load, and enhances user interaction by adapting to the capture and rendering contexts, resulting in a more realistic and efficient 3D visualization experience.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for processing an audiovisual stream, corresponding electronic device and computer program product.
[0001] 1. Field of the invention
[0002] The present application relates to the field of image and audiovisual stream processing, with potential application in various contexts such as the capture, manipulation and / or broadcasting of 3D images and / or videos.
[0003] It relates to a method for processing an audiovisual stream.
[0004] It also relates to the corresponding electronic device, computer program product and system.
[0005] 2. State of the art
[0006] Stereoscopy, developed in the 19th century, allows the perception of depth in a scene by presenting two slightly different images to each eye. Since then, methods for capturing and volumetrically reproducing images and videos have evolved, finding applications in fields as diverse as medicine and the security of official documents. Holography, which creates three-dimensional images by recording light interference, differs from stereoscopy in its ability to reproduce an image with correct perspective and parallax, without requiring special equipment for viewing. However, contemporary holographic systems face challenges for nomadic and interactive use. They require high computing power and storage capacity, which is restrictive for portable devices.The incorporation of small holographic displays remains a technical challenge, limiting user interaction. These systems can also face problems of insufficient brightness and high power consumption.
[0007] Alternative systems to traditional holography exist, such as those based on "Pepper's ghost." The principle is based on a 19th-century theatrical technique invented by John Henry Pepper, which allows a three-dimensional optical illusion to be created by the reflection of an object on a transparent surface arranged in such a way that the reflected image appears to exist in space. Applied to screens, this technique can generate volumetric representations without resorting to heavy and expensive equipment. However, the methods implementing this technique depend on the format of the audiovisual streams to be processed. These are therefore mainly pre-recorded content.
[0008] To date, there is no system that allows capturing an audiovisual stream and process for a volumetric representation, in a simple, intuitive and controlled manner, with a portable terminal such as a smartphone or a tablet.
[0009] The present application thus aims to propose improvements to at least some of the drawbacks of the state of the art.
[0010] 3. Statement of the invention
[0011] The present application aims to improve the situation using a method for processing an audiovisual stream, said method comprising: - obtaining at least one image from a first audiovisual stream, - a detection of at least one object of interest in said image taking into account of a volumetric capture context, - a spatialization of said at least one detected object of interest taking into account a volumetric rendering context.
[0012] In the remainder of this application, an object of interest will be designated as a portion of a two-dimensional (or 2D) or three-dimensional (or 3D) image which can be the subject of an analysis, recognition or interaction by the method of the invention.
[0013] In the present application, a capture designates a process of collecting audiovisual data, in particular three-dimensional data, such as images, videos, sounds and / or spatial information, from a real environment, using devices such as microphones, cameras with charge-coupled optical-digital sensors (CCD or "Charge Couple Device" in English), telemetric sensors such as light detection and ranging sensors (LiDAR or "Light Detection And Ranging" in English) and / or (ToF or "Time of Flight" in English). LiDAR emits laser pulses and measures the time it takes for each pulse to return after hitting an object, which makes it possible to calculate the distance and generate an accurate depth map of the scene.Unlike LiDAR which measures time for each point individually, ToF cameras measure time for an entire scene simultaneously, resulting in a depth map.
[0014] The volumetric capture context refers to all the technical, spatial and / or operational parameters and conditions that can impact one or more captures. This context includes the geometry of the capture space, the position and orientation, in this space, of the objects of which an audiovisual representation is to be captured, as well as the position, angle and technical characteristics of the capture equipment used. The volumetric capture context can also include environmental conditions such as brightness at the time of capture. Ambient light, whether natural or artificial, can influence the way in which images are captured. It can thus affect visibility, contrast, shadows and the color of objects in the scene.
[0015] Rendering is understood to mean in the present application a process by which a computer generates a rendering (or "output" according to the English terminology) on at least one user interface, in any form, for example comprising textual, audio and / or video components, or a combination of such components. This process may involve a conversion of three-dimensional data into 2D images, but it may also concern a rendering of 2D data, such as textures or vectors. Rendering may also include a simulation of light, shadows, reflection and refraction to help give objects a realistic and / or artistic appearance. It may be carried out on the fly and / or pre-calculated.
[0016] The volumetric rendering context refers, in the present application, to all the parameters, techniques and display conditions related to the rendering of three-dimensional images from spatial data. This includes, for example, rendering equipment allowing a display perceived or produced in 3 dimensions, but also the conditions of orientation, stability and position of the equipment which can affect the perception of the image. It can also be the physical environment in which the rendering(s) can be carried out, such as the ambient brightness and / or clarity.
[0017] The method can thus help to obtain, in at least certain embodiments, a more faithful and immersive representation of the objects of interest than certain traditional 3D visualization methods, by taking into account the volumetric contexts of capture and rendering. The implementation of such a method can for example help, in at least certain embodiments, to improve the rendering of the depth of the objects, and therefore their realism, in 3D visualization applications.
[0018] Taking into account the volumetric context of capture when detecting the object of interest makes it possible to obtain information on a captured scene, which can result, for example, in an improvement in the quality and relevance of the audiovisual processing. In addition, this consideration can allow an improvement in the precision in the detection, by distinguishing an object of interest from the elements surrounding the object.
[0019] Furthermore, knowledge of the position and orientation of the object of interest in space can help adapt the rendering consistently with the user's perspective, which can help improve the user's visual experience and sense of immersion. This adaptation can be useful for object tracking in the event of rapid or unexpected movements. Detection, taking into account the volume, can help estimate sizes and determine a spatial relationship between objects.
[0020] In some embodiments, the method may allow for customizing the content or information displayed based on the position and orientation of the detected object, thus providing a more targeted and relevant user experience. Finally, by precisely identifying the object of interest, the method can help focus processing resources where they are needed, thus helping to reduce the computational load and improving energy efficiency.
[0021] Considering a volumetric rendering context during the spatialization step offers the advantage of being able to help continuously adjust the projected image to correct any distortion or mismatch that might arise during the rendering process. This dynamic approach provides ongoing information about the rendering, which can be used to make ongoing corrections to the image, helping to bring the final result closer to the original intent.
[0022] In at least one embodiment, said spatialization comprises a creation of at least one first spatialized image of said object of interest, said creation comprising: - a selection of a visual object from among candidate visual objects taking into account the characteristics of said detected object of interest, - a substitution of said object of interest in said image of said first stream or in said first spatialized image by said selected visual object.
[0023] Selecting a visual object based on the characteristics of the detected object of interest can allow for personalization, by adapting the spatialized image to the specificities of the object. Substituting the object of interest with a selected visual object can also improve the visual quality of the representation, by using high-resolution 3D models and / or textures that can be more detailed than the original image captured with a lower resolution.
[0024] In some embodiments, the method comprises rendering the at least one first spatialized image. Rendering spatialized images can help improve the quality of the volumetric visualization, providing a more realistic and interactive experience.
[0025] In certain embodiments, the method comprises a plurality of renderings on the same time window of said first spatialized image. This embodiment can for example be used for rendering on a type of display device requiring several copies of the spatialized image to be rendered so that it can be perceived as three-dimensional.
[0026] In at least one embodiment, said method comprises, during said rendering, a dynamic adaptation of at least one graphic parameter of said first audiovisual stream and / or of at least one portion of said at least one first spatialized image. The dynamic adaptation of the graphic parameters makes it possible to ensure an image quality adapted to variations in the conditions of use and to the specificities of the user's equipment. It offers the possibility of adjusting the visual quality of the stream audiovisual and spatialized images based on viewing conditions, such as ambient lighting and / or screen characteristics, thus improving the overall viewing experience. It can also improve the readability of displayed elements, particularly when viewing text, graphics and / or fine details.
[0027] In some embodiments, said at least one graphics parameter belongs to a group comprising: - a graphic parameter corresponding to a resolution of said first spatialized image; - a graphic parameter corresponding to an image refresh rate of said first audiovisual stream; - a contrast adjustment parameter; - a brightness adjustment parameter; - a white balance adjustment parameter; - a sharpness adjustment parameter; - a combination of at least two of the above parameters.
[0028] Adjusting the resolution and refresh rate can help adjust the smoothness of volumetric images, potentially providing a higher quality visual experience without compromising performance. By adjusting these graphics parameters, the method can, for example, allow bandwidth usage to be controlled, which can be advantageous in environments with limited or unstable internet connections. Adjusting graphics parameters, such as brightness, contrast, or saturation, can, for example, help meet specific visual preferences or content requirements.
[0029] In certain embodiments, the rendering of said at least one first spatialized image takes into account said volumetric capture context and / or said at least one detected object of interest.
[0030] By taking into account the volumetric environment at the time of capture, the method can make it possible to produce a rendering that is closer to reality by helping to improve spatial coherence between the real environment and the virtual environment and can, for example, help to adapt the spatialized image to the user's specific environment; for example, by adjusting the scale or perspective of the objects so that they fit naturally into their physical space. By taking into account the object of interest detected during rendering, the method can help to obtain more realistic interactions between the participants and the virtual or spatialized objects, for example, as if these interactions occurred in the real world.
[0031] In certain embodiments, when said at least one detected object of interest is a living being, said selection of said visual object takes into account a similarity, in terms of bodily expression, between said at least one detected object of interest and said visual object.
[0032] Selecting visual objects based on body expression similarity can help make interactions more natural and intuitive, thereby improving user communication and engagement. The method can help more accurately capture and convey the nuances of nonverbal communication, which can help provide a more authentic human interaction. The matching between the detected body expressions and the selected visual object can contribute, in some embodiments, to a more natural and intuitive user experience by reducing the dissonance between real movements and their virtual representation.
[0033] In certain embodiments, said dynamic adaptation takes into account said visual object. Visual coherence between the detected object of interest and its spatialized representation can thus be promoted, thus helping for example to improve the integration of the object in the audiovisual stream containing at least one spatialized image.
[0034] In some embodiments, the method comprises transmitting a second audiovisual stream containing said at least one first spatialized image. Transmitting an audiovisual stream enriched with volumetric content can help to share immersive experiences potentially with other users, extending the possibilities for communication and collaboration. In addition, the method can, in some embodiments, have the advantage of allowing part of the processing of the audiovisual stream to be transferred to a device having the necessary processing capacity by performing and / or completing the detection of at least one object of interest and / or the spatialization of the at least one object of interest.
[0035] In some embodiments, the method comprises:
[0036] - reception of a third audiovisual stream;
[0037] - a rendering of a second spatialized image obtained from said third stream at audiovisual received.
[0038] Receiving an audiovisual stream containing at least one spatialized image can allow a user to benefit from a volumetric viewing experience, regardless of the terminal he is using. The fact that the images are already spatialized can make it possible, in certain embodiments, to reduce the computational load on the user's receiving terminal, which can be particularly advantageous for terminals with limited processing capabilities.
[0039] In some embodiments, the method is implemented by a first electronic device during a communication session with at least one second electronic device.
[0040] The use of electronic devices to implement the method ensures that the technology can be integrated and used in a wide range of devices, depending on the embodiments, thus promoting greater interoperability between different systems and platforms. A device can thus process and adapt content according to its technical characteristics and the needs of its user, thus contributing, for example, to offering a personalized experience during the communication session.
[0041] The characteristics presented in isolation in the present application in connection with certain embodiments of the method of the present application can be combined with each other according to other embodiments of the present method.
[0042] According to another aspect, the present application also relates to an electronic device comprising at least one processor configured to implement the method of the present application in any of its embodiments.
[0043] Thus, in certain embodiments, the present application relates to an electronic device comprising at least one processor adapted to: - obtaining at least one image from a first audiovisual stream, - a detection of at least one object of interest in said image taking into account of a volumetric capture context, - a spatialization of said at least one detected object of interest taking into account a volumetric rendering context.
[0044] According to another aspect, the present application also relates to a system comprising at least one processor configured to implement the method of the present application in any of its embodiments.
[0045] Thus, in certain embodiments, the present application relates to a data processing system for processing an audiovisual stream comprising:
[0046] a first electronic device comprising at least one processor adapted to: - obtaining at least one image from a first audiovisual stream, - a detection of at least one object of interest in said image taking into account of a volumetric capture context, - a spatialization of said at least one detected object of interest taking into account a volumetric rendering context. - a transmission of a second audiovisual stream containing said at least one first spatialized image,
[0047] a second electronic device comprising at least one processor adapted to: - a reception of said second audiovisual stream containing at least said first spatialized image; - a rendering of said first spatialized image.
[0048] The present application also relates to a computer program comprising instructions for implementing the various embodiments of the above method, when the computer program is executed by a processor and a medium recording readable by an electronic device and on which the computer program is recorded.
[0049] For example, the present application thus relates to a computer program comprising instructions for implementing, when the computer program is executed by a processor of an electronic device, a method for processing an audiovisual stream, said method comprising: - obtaining at least one image from a first audiovisual stream, - a detection of at least one object of interest in said image taking into account a volumetric capture context, - a spatialization of said at least one detected object of interest taking into account a volumetric rendering context.
[0050] The above-mentioned program may use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0051] The recording (or information) media mentioned in the present application may be any entity or device capable of storing the program. For example, a medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or even a magnetic recording means.
[0052] Such a storage means may for example be a hard disk, a flash memory, etc.
[0053] On the other hand, an information medium may be a transmissible medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. A program according to the invention may in particular be downloaded from a network such as the Internet.
[0054] Alternatively, an information carrier may be an integrated circuit in which a program is incorporated; in the present application, the circuit is adapted to execute or to be used in the execution of any of the embodiments of the method which is the subject of the present patent application.
[0055] 4. Brief description of the drawings
[0056] Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:
[0057] [Fig.l] presents a simplified view of a system adapted to implement at least certain embodiments of the method of the present application,
[0058] [Fig.2] presents a simplified view of a device suitable for implementing at least certain embodiments of the method for processing an audiovisual stream of this request,
[0059] [Fig. 3] presents an overview of the method for processing an audiovisual stream of the present application, in certain of its embodiments.
[0060] [Fig.4] presents a timing diagram of the method for processing an audiovisual stream of the present application, in certain of its embodiments.
[0061] [Fig.5] shows an example of a communication between two participants using two devices for processing an audiovisual stream.
[0062] [Fig.6] presents an example of a plurality of renderings on the same time window of a spatialized image.
[0063] [Fig.7] presents an example of a device for processing an audiovisual stream allowing dynamic adaptation of graphic parameters of an audiovisual stream and / or of at least one portion of at least one spatialized image.
[0064] 5. Description of the embodiments
[0065] The present application aims to propose a volumetric visualization of an audiovisual stream by at least one electronic device using image processing techniques and taking advantage of the physical environment in which said device is located.
[0066] As used herein, the term 'audiovisual stream' refers to data comprising images and, optionally, associated audio data. Although in some embodiments the method may process both components, some embodiments may focus exclusively on the visual portion.
[0067] The present application relates in particular to a method for processing an audiovisual stream making it possible to detect and spatialize objects of interest in the stream while taking into account the capture and / or rendering context.
[0068] In particular, unlike certain prior art solutions proposing bulky, energy-consuming devices and / or devices limited to the viewing of pre-recorded content, the present application proposes a method for processing an audiovisual stream that can be adapted to connected portable terminals, this adaptation being able to be dynamic.
[0069] The method comprises obtaining at least one image within the stream. It also identifies at least one object of interest in the image by taking into account the volumetric context of the capture necessary to obtain the image. Finally, the method comprises a spatialization of the detected object of interest by taking into account, for example, the volumetric context of rendering of the spatialized image.
[0070] In particular, in certain embodiments where the detected object of interest is a living being, the present application may take into account a bodily expression (for example a head movement or a grimace which may signify a communication problem) grasping) of the user(s) in order to adapt the processing of the audiovisual stream accordingly, the consideration being able, for example, to take into account a similarity between a bodily expression and pre-established and / or evolving bodily expression models (for example via an artificial intelligence learning system).
[0071] We now describe, by way of example, a telecommunications system in which the audiovisual stream processing method of the present application can be implemented.
[0072] [Fig.l] illustrates such a telecommunication system 100. This system comprises one or more audiovisual stream processing devices (110, 111, 115), which may be digital terminals such as smartphones, tablets or portable or fixed computers (110, 111), or servers 115. These processing devices may integrate or be associated with CCD and / or LiDAR and / or ToF opto-digital capture equipment (120, 121, 122) which have the capacity to capture visual and / or metric information of a scene external (101, 102) to the equipment. In certain embodiments, at least some of the opto-digital capture equipment (120, 121) may be integrated into the stream processing devices (110, 111).In some embodiments, at least some of the opto-digital capture equipment 122 may be external to the stream processing devices and transmit the captured images to at least one stream processing device 110 via a network interface, wired (USB, RJ45, optical (single-mode or multi-mode fiber) or wireless (4G, 5Gn, WiFi, Bluetooth).
[0073] Further, in some embodiments, the telecommunications system 100 may be configured to at least partially centralize processing of the audiovisual stream within a device 115, such as a server.
[0074] In certain embodiments, the system may also include connected terminals (150, 160) which may interact with the audiovisual stream processing devices 115. A transmitting connected terminal 150 may for example send an audiovisual stream to an audiovisual stream processing device 115, which after processing this audiovisual stream, may transmit an audiovisual stream containing at least one spatialized image to another processing device 111 and / or a receiving connected terminal 160.
[0075] [Fig.2] illustrates a simplified structure 200 of an electronic device 200, adapted to implement the principles of the present application. Depending on the embodiments, it may be a server and / or a terminal. The device 200 may for example correspond to an audiovisual stream processing device (110, 111, 115) of the system 100 in [Fig.l].
[0076] The device 200 notably comprises at least one memory M 210. The device 200 may notably comprise a buffer memory, a volatile memory, for example example of RAM type (for "Random Access Memory" according to English terminology), and / or a non-volatile memory for example of ROM type (for "Read Only Memory" according to English terminology). The memory can for example be used to temporarily or permanently store images, and / or video, and / or sounds and / or audiovisual streams captured before, during and after the processing of an audiovisual stream according to the present application. This includes for example original 2D images captured, spatialized images, as well as, optionally, the associated audio data. The memory can also contain data relating to visual objects such as 2D and / or 3D models as well as parameters and configuration data necessary for the processing of audiovisual streams, such as processing based on object detection algorithms, spatialization, and volumetric display graphic parameters.
[0077] The device 200 may also comprise a processing unit UT 220, equipped for example with at least one processor P 222, and controlled by a computer program PG 212 stored in memory M 210. At initialization, the code instructions of the computer program PG are for example loaded into a RAM memory before being executed by the processor P. Said at least one processor P 222 of the processing unit UT 220 may in particular implement, individually or collectively, any one of the embodiments of the method of the present application (described in particular in relation to [Fig. 3]), according to the instructions of the computer program PG.
[0078] In certain embodiments, the device 200 may also comprise means for receiving an audiovisual stream from at least one network interface of the device 200. The network interface may communicate with a network in a wired manner (such as Ethernet or USB) or wirelessly (such as WiFi, Bluetooth, or 4G / 5G).
[0079] Similarly, in certain embodiments, the device 200 has the capacity to be able to transmit at least one audiovisual stream originating from the network 140.
[0080] By user interface (or “human-machine interface”) of the device, we mean for example an interface integrated into the device 200, or a part of a third-party device coupled to this device by wired or wireless communication means.
[0081] A user interface may in particular be a user interface, called an “output” user interface, adapted to a rendering (or to the control of a rendering) of an output element of a computer application used by the device 200, for example an application running at least partially on the device 200 or an “online” application running at least partially remotely, for example an application accessible via the device 200. Examples of output user interfaces of the device include one or more screens, in particular at least one graphical screen (touch screen for example), one or more speakers, a connected headset, one or more in light indicator(s) such as light-emitting diodes (or LEDs for "Light Electronic Display" according to English terminology). In addition, the device 200 may be equipped with (or coupled to) at least one piece of rendering equipment (130, 131), such as a display screen adapted to render an audiovisual stream containing at least one spatialized image. The rendering equipment may, for example, be a piece of "Pepper's phantom" type equipment.
[0082] By rendering, as detailed above, we mean here a restitution (or “output” according to English terminology) on at least one user interface, in any form, for example comprising textual, audio and / or video components, or a combination of such components.
[0083] Furthermore, a user interface may be a so-called “input” user interface, adapted to acquiring a command from a user of the device 200. This may in particular be an action to be performed and / or a command to be transmitted to a computer application used by the device 200, for example an application running at least partially on the device 200. An “input” user interface may also be adapted to acquiring at least one configuration parameter linked to a display device.
[0084] Examples of input user interface of the device 200 include a sensor, an audio and / or video acquisition means (microphone, camera (webcam) for example), a keyboard, a mouse.
[0085] In certain embodiments, the device may comprise (or be coupled to) means for capturing an audiovisual or visual stream, such as at least one opto-digital capture device suitable for capturing an image or a video, such as an integrated camera 120 or connected 122 to the processing device 200, for example, a front camera of a smartphone or a computer webcam. In embodiments where the device comprises (or is coupled to) means for capturing independently a visual stream and an audio stream, the device may further comprise means for synchronizing these visual and audio streams.
[0086] In some embodiments, the device 200 may also comprise (separately or in addition to the above means) specialized capture interfaces, such as LiDAR and / or ToF sensors for the acquisition of volumetric data.
[0087] Said at least one microprocessor of the device 200 may in particular be adapted to implement a method for processing an audiovisual stream, said method comprising: - obtaining at least one image from a first audiovisual stream, - a detection of at least one object of interest in said image taking into account of a volumetric capture context, - a spatialization of said at least one detected object of interest taking into account a volumetric rendering context.
[0088] Some of the above input-output modules are optional and may therefore be absent from the device 200 in certain embodiments. In particular, if the present application is sometimes detailed in connection with a device communicating with at least one second device of the system 100, the method may also be implemented locally by a device, for example using application output elements not requiring exchanges between devices (case of a so-called “stand-alone” device for example).
[0089] On the contrary, in some of its embodiments, the method can be implemented in a distributed manner between at least two devices (101, 102) of the system 100.
[0090] We now describe in connection with [Fig. 3], in a simplified manner, the method 300 for processing an audiovisual stream of the present application, in certain of its embodiments. The method 300 can be implemented for example by the device 200 described above.
[0091] As illustrated in [Fig.3], the method 300 comprises obtaining 310 at least one audiovisual stream. This obtaining may depend on the embodiments. Thus, in certain embodiments, the obtaining may be carried out for example via a capture by a capture device of the device 200. Obtaining an audiovisual stream may be carried out, in certain embodiments, by connecting for example to a streaming service such as a videoconferencing bridge, where the data is transmitted continuously from a remote server. The device 200 may communicate for example, in certain embodiments, with other devices, such as smartphones, tablets or surveillance cameras, to obtain their audiovisual stream. The method may make it possible, in certain embodiments, to obtain an audiovisual stream by importing pre-recorded video and audio files stored locally or accessible via the network from a storage location..
[0092] In some embodiments, obtaining comprises storing the audiovisual stream in the device 200. This may involve, for example, the use of a buffer memory designed to temporarily store data in transit, which may allow continuous and smooth processing of images, videos and / or sounds. By dynamically adjusting the size of the buffer according to the data rate and the processing capabilities of the device, the method can effectively reduce the risk of slowdown or interruption in the processing of the audiovisual stream.
[0093] Following the obtaining of at least one image from the audiovisual stream, the method 300 may comprise a detection 320 of at least one object of interest. The method may for example comprise a recognition and a determination of at least one location of at least one particular element (object of interest) in an image or a obtained video sequence. This detection can be carried out in various ways depending on the embodiments. It can implement, for example, various speech recognition and / or computer vision and artificial intelligence techniques. Convolutional neural networks (CNNs) or other deep learning architectures can, for example, be trained to recognize and localize types of objects of interest (audio and / or visual). The objects of interest can be predefined during the training phase, where a neural network is exposed to many labeled (tagged) images or audio samples representing these objects. The neural network can also learn to recognize distinctive characteristics of the objects of interest and to distinguish these objects of interest from other audio and / or visual elements in new unlabeled images or audio samples.Thus, the predefinition of objects of interest may be integrated, in certain embodiments, into the deep learning model, which may allow them to be detected (i.e., recognized and localized (accurately)) in audiovisual streams. Such techniques may help to obtain a classification of objects of interest even under complex conditions due, for example, in the case of image processing, to the presence of obstacles masking the object or backgrounds that may cause the objects to be confused with their environment. Image analysis algorithms may, for example, be used to analyze the audiovisual stream and detect patterns, shapes, or colors that correspond to predefined objects of interest. Pattern recognition techniques may also be applied to help identify types of objects based on their geometry, texture, or outline.For example, such techniques can be used to detect human faces or everyday objects.
[0094] In the case where the audiovisual stream comprises a video sequence, once an object of interest has been detected in a portion (image for example) of the stream, the method 300 may comprise tracking of the object, carried out in real time, through the different images of the sequence which follow the first image where the object was detected. The method may thus make it possible to maintain a focus on an object of interest detected in an audiovisual stream.
[0095] This detection can help to analyze the content of the audiovisual stream and react accordingly, whether to modify the image, interact with the user or make decisions based on the detected objects.
[0096] As illustrated in [Fig.3], the method 300 may comprise a spatialization 330 of at least one detected object. By spatialization of an image or portion of an image (such as an object of interest), we mean a transformation of at least one image, respectively the portion of the image considered, into a three-dimensional representation (or spatialized image) which can thus be perceived by a user as a 3-dimensional image. Among the techniques for transforming 2D images into 3D representation, stereoscopy can for example be used by obtaining two images captured from slightly different angles as illustrated in the left part of [Fig.6]. From a smartphone, a stereoscopic capture, that is to say a simultaneous capture of two audiovisual streams representing the same scene (here a user), is carried out by means of two cameras connected to the smartphone: a front camera 120 surmounted by a periscopic system allowing the orientation of the capture towards the user and another camera 122 connected for example to the USB port of the terminal 110 and oriented towards the user.Depth sensors such as LiDAR and / or ToF cameras can also be used to provide direct measurements that, when combined with a 2D image, allow for the construction of a 3D image-based model. Convolutional neural networks can also be trained to estimate depth from a single image, while motion-based feature tracking and reconstruction techniques analyze the displacement of interest points to model the 3D scene. For example, Neural Radiance Fields (NeRFs) can also be used to accurately model light interactions in a three-dimensional scene. By training a neural network with a set of 2D images taken from different angles, NeRFs learn to synthesize new views of the scene, providing depth and volume perception.
[0097] As illustrated in [Fig.4], in one embodiment, the spatialization 330, performed by a device 200 DI of a detected object of interest 320 in an audiovisual stream obtained 310 from an upstream user U1, may comprise a selection 332 of a visual object taking into account the characteristics of the object of interest considered. For example, if the system detects a face as an object of interest, the selected visual object may be a 3D model of a face. It may for example be an avatar, at least some characteristics of which correspond to visual characteristics of the detected object of interest, such as the shape of the face, the position of the eyes, the nose and the mouth, so as to create a volumetric representation corresponding at least partially to the detected face.
[0098] The method may comprise a substitution 334 where, following the selection of a 3D model, the object of interest is replaced in the image by an object corresponding to this model. This substitution may be based for example on depth information previously obtained 310 by a LiDAR sensor or by the application of 3D reconstruction techniques. This step may also require for example the estimation of the geometry of the object from different angles or the exploitation of depth data. from the mentioned sensors.
[0099] When spatializing the object of interest, the method may also include processing the audio component of the audiovisual stream to create an immersive experience consistent with the spatialized image. This processing may include, in embodiments where the resulting audiovisual stream has an audio component, adjusting the captured sound so that it corresponds to the spatial position of the object in the virtual or real environment, thereby creating consistency between the audio and the visual. Depending on the position and / or orientation of the detected object of interest, the sound may be modified to reflect its position in space. For example, if the object moves to the left of the screen, the sound will also be moved to the left in the audio mix. For example, the method may improve the perception of the context of a scene comprising several objects and / or people.
[0100] In some embodiments, the method may comprise rendering 340 one or more spatialized images.
[0101] The rendering can be carried out on the device 200, for example the device DI in [Fig.4] which can be a terminal such as a smartphone or a tablet. The user U1 of this device can thus view in three dimensions a scene containing one or more objects of interest resulting from the capture through a display device such as the devices 130 or 131 of [Fig.l].
[0102] The rendering 340 may be optional in certain embodiments, such as for example when the method comprises (or precedes) a transmission of at least one spatialized image to another device, or when the rendering is carried out deferred (for example during a subsequent request from a user).
[0103] It is noted that, in certain embodiments, the method 300 may comprise processing of the selected visual object, carried out for example upstream and / or during the spatialization 330 to produce a volumetric rendering. Volumetric rendering refers to the creation of a three-dimensional representation which simulates the depth and volume of an object, thus helping to provide a more realistic perception by the user. In certain embodiments, the method may include adding lighting effects, shadows and / or textures to improve the realism of the spatial representation for example.
[0104] In at least one embodiment, the rendering 340 may take into account the volumetric context of audiovisual stream capture. For example, the rendering may take into account spatial and depth information obtained during shooting with volumetric capture equipment such as a LiDAR and / or ToF. During capture, the volumetric capture equipment may, for example, measure the distance between the volumetric capture equipment and objects in the scene. This depth data may then be synchronized with the captured 2D images to create a 3D model of the scene and can allow, during rendering, the application of light, shadow and texture effects to increase the realism of the 3D scene. It can also involve taking into account the position, intensity and / or color of light sources in a scene to create lighting effects. At least some of the spatial information can be representative of a position of the object of interest as well as, optionally, an evolution of this position over time. In such embodiments, static objects can for example require fewer computing resources for rendering than moving objects, thus allowing optimization of the performance of the device 200 according to the state of the object of interest.
[0105] Thus, in some embodiments, the rendering 340 may use not only the visual data, but also three-dimensional data to create an image that reflects the volume and structure of the object of interest. For example, the method may be applied to a capture of a scene for an augmented reality application, where the volumetric context may then be used to place virtual objects consistently with the physical environment.
[0106] In some embodiments, once the image spatialization has been performed, as illustrated in [Fig.4], the method 300 may comprise a plurality of renderings 340 of the spatialized image over the same time window. For example, the method may comprise creating several copies of the same spatialized image and jointly displaying the spatialized image and its copies.
[0107] In the remainder of the application, we will designate by spatialized audiovisual stream an audiovisual stream containing at least one spatialized image.
[0108] As illustrated in [Fig.4], the method 300 may comprise a dynamic adaptation 342 of graphic and / or audio parameters during the rendering of the spatialized audiovisual stream. The dynamic adaptation comprises, for example, an ongoing modification of the graphic parameters in response to variations in the viewing conditions or to changes in the content of the stream. It may apply, depending on the embodiments, to the stream as a whole, including the sequences of images and / or sound, to a specific portion of an image and / or an audio track, or to a combination of the two.
[0109] Adapting audio parameters may, for example, include adjusting volume, equalization, or sound spatialization to ensure consistency with the spatialized image and improve user immersion. For example, in a noisy environment, the method may increase the volume of dialogue while reducing background noise, or adjust the direction from which the sound comes to match the position of objects on the screen.
[0110] Examples of graphical parameters of an image include, but are not limited to to, parameters relating to: - the resolution of an image, which can be defined by the number of pixels in width and height. The higher the resolution, the more detail the image contains - the refresh rate, which expresses the number of times an image is updated on the screen per second, measured for example in hertz (Hz). A higher refresh rate can help make movement smoother and reduce eye strain. - contrast, which can be defined as the ratio between the lightest and darkest parts of an image. High contrast can help make elements in the image easier to distinguish. - brightness, which refers to the intensity of light emitted or reflected by the image. Adjusting brightness can help make the image more visible depending on the ambient lighting. - white balance, which adjusts color reproduction so that objects that are white in reality appear white in the image. The color temperature of the image and its overall rendering can thus be adapted according to the white balance. - sharpness, which improves the clarity of details in an image. An appropriate sharpness setting makes it easier to distinguish textures and contours.
[0111] In some embodiments, by adjusting these graphics parameters separately or in combination, the method 300 can help control the use of bandwidth required to transmit an image or video, which can be advantageous in environments with limited or unstable Internet connections. For example, reducing the resolution of an image decreases the number of pixels to be transmitted, which helps reduce the data size and therefore the bandwidth required for the transmission or transmission of the image. Similarly, adjusting the refresh rate of an audiovisual stream impacts the number of images transmitted per second. For example, decreasing the refresh rate can help reduce the amount of data to be sent and adapt it to the capacity of its network connection.This adjustment may also help maintain the smoothness of the audiovisual stream and ensure synchronization between the audio and video components, depending on the performance of the device 200.
[0112] In certain embodiments, the method may comprise dynamic adaptation of the brightness and / or contrast of the rendering of the spatialized image in relation to the ambient brightness of the environment in which the device is located.
[0113] These may be, for example, embodiments where obtaining a flow au- diovisual, according to the method, is implemented in a location of low ambient light. An increase in the brightness of the rendering and an adjustment of its contrast can for example be carried out according to the method to improve the visibility of the spatialized image for a user. Conversely, in certain embodiments, the method may comprise a reduction in the brightness of the rendering, in particular situations where this does not significantly alter the quality of the perceived image.
[0114] In some embodiments, the adaptation may target a specific portion of the spatialized image. For example, if an area of the image requires special attention, such as a face, the graphics parameters of that area may be adjusted independently of the rest of the image to improve clarity and / or lighting.
[0115] The fact that the adaptation is carried out continuously can help to adjust the graphic parameters of the image as it goes along, from the first rendering. In certain embodiments, the method can comprise an opto-digital capture of the rendering, so as, for example, to make it possible to detect the differences between the projected image and the way in which it is perceived on the rendering device. For example, the method can comprise, in certain embodiments, a modification of the source image obtained 310 taking into account this feedback information. Such embodiments can help to compensate, for example, perspective effects, optical distortions or variations linked to said graphic parameters which could affect the quality of the visualization.For example, the rendering capture device may correspond to the capture device that was used to obtain the image in the step 310 of obtaining the audiovisual stream.
[0116] According to one example, in embodiments where the user's image is volumetrically projected using the method, if the rendering capture device captures a projected image and detects that its colors are washed out due to ambient lighting, a feedback loop may be implemented. The method may then automatically adjust the white balance and / or brightness to compensate for the effects of lighting and improve color fidelity.
[0117] Similarly, in another example, if the projected image appears blurry or details are lost due to inadequate resolution, the method may, for example, dynamically increase sharpness and adjust resolution so that details are clearer and more precise.
[0118] Dynamic adaptation 342 of graphics parameters may for example be useful in environments where viewing conditions change frequently, where brightness and contrast may require constant adjustments to maintain image readability in the face of changes in natural or artificial light. Thus, dynamic adaptation 342 may help maintain image quality. on the fly, adjusting graphics settings based on continuous feedback obtained by monitoring the rendered image.
[0119] In some embodiments, the method may also include tracking (continuous or intermittent (periodic for example)) of the physical context in which the user device is located. The tracking may include, for example, capturing current values of ambient data such as brightness, contrast, and color temperature. This data may be used, in some embodiments, to dynamically adapt the graphical parameters of the spatialized image, in order to help improve the visibility of the image according to the immediate environment of the user. Optionally, this adaptation may also take into account the specific characteristics of the rendering equipment used, such as its resolution, its refresh rate, and its contrast capabilities, to help improve the integration of the spatialized image in the rendering.
[0120] In some embodiments, the dynamic adaptation 342 may also take into account the selected visual object 332 during the spatialization 330. In such embodiments, the method 300 includes an adjustment of graphics parameters according to the visual object, which may help to adapt the rendering performance, for example by allocating resources of the device 200 more efficiently to maintain high image quality without overloading the device integrating the method. This approach may allow, in some embodiments, a customization of the rendering, by adapting not only the appearance of the visual object but also its rendering according to the preferences and needs of the user. For example, in some embodiments, during a videoconference where speakers speak remotely, the method may adapt the sharpness, brightness and contrast of the avatar (i.e.the visual object) or the spatialized image of at least one speaker. In such embodiments can help ensure that their representation is clear and consistent with their actual bodily expression, thereby potentially improving the quality of communication. It can also involve, for example, dynamic adjustment of the graphic parameters to highlight avatars or spatialized objects that could be actively used or manipulated by users based on their bodily expressions.
[0121] In certain embodiments, the method 300 may comprise a transmission 350 of the spatialized audiovisual stream. For example, in connection with FIGS. 1 and 4, if the device 200 DI is a connected terminal, the stream containing at least one spatialized image of the user 101 U1 may be transmitted from his smartphone via a network 140. It may also be a network server 115 which, after having processed an audiovisual stream using the method, transmits a spatialized stream via said network 140.
[0122] The spatialized audiovisual stream thus emitted can be received by a receiving device. The receiving device may be a D2 device adapted to rendering the received stream (such as a device implementing the method of the present application) as illustrated in [Fig.4] or a server.
[0123] In certain embodiments, the method may comprise a rendering 340 of the spatialized audiovisual stream on the device 200 (if said device 200 is a connected terminal 160 such as a smartphone for example).
[0124] It is noted that in certain embodiments, the method can be implemented at least partially by a device 200 corresponding to a server, such as the server 115 of [Fig.l]. Thus, such a server can be configured to implement the method 300 so as to centralize the processing of at least one audiovisual stream received according to the method of the present application. For example, the method can comprise a 360 reception of images and / or videos from terminals (110, 150). This centralization can make it possible to optimize the processing resources by concentrating the complex and computationally intensive operations on dedicated equipment 115, capable of managing large quantities of data.In such embodiments, certain devices (110, 111) can for example implement certain steps of the method, such as local capture or final rendering of the spatialized audiovisual stream, while benefiting from the processing capabilities of the server 115 (which will implement certain steps of the method).
[0125] According to another example, the method implemented on the server may comprise a 360 reception of at least one audiovisual stream already spatialized relative to a first object of interest (and originating from a device implementing the processing method of the present application) but of which a complement of spatialization relative to a second object of interest may be carried out on the server 115.
[0126] An implementation of the method in a non-distributed manner can help simplify the updating of a computer application implementing the method.
[0127] In some embodiments, the method may be implemented during a videoconference. For example, the method may be implemented at least partially on a server integrating videoconference bridge functionalities, to process audiovisual streams of the participants separately or in combination.
[0128] Thus, the method can be implemented by a first device 200 during a communication session with at least one second device 200. For example, as illustrated in [Fig.5], in the case of a telecommunication between two users 101 and 102 each equipped with a device 200 for processing audiovisual streams 110, respectively 111, a volumetric representation of the users can be displayed jointly on the rendering equipment 130 and 131.
[0129] The method of the present application may also assist, in at least some embodiments, in better assessing participant nonverbal communication in detecting a user's body expressions and selecting or creating a 3D avatar that faithfully reproduces these expressions. For example, during a videoconference meeting, the method can obtain an image of the participant in order to detect their face or their entire body as an object of interest by taking into account the volumetric context of capture (for example, the participant's position in the room and the distance from the camera). The method can detect a participant's face and select a 3D avatar that matches their characteristics (such as hairstyle, face shape, or even expression). The personalized avatar is then substituted for the real image of the participant, allowing for a personalized representation of the participant in the videoconference environment.This may also be the case for a face during a video conference where, via the spatialization step, the method can create a three-dimensional model of the face which can then be oriented in any desired direction. This would allow, for example, to present the face of the interlocutor facing the camera, even if the person is physically turned to the side, thus improving engagement and eye contact during a video conference. This ability to adjust the orientation of a captured face is particularly useful in situations where maintaining eye contact is important for communication, such as in professional meetings. The method of the present application can thus help, in at least some embodiments, to a more natural and effective communication experience, by simulating a face-to-face presence of the participants.
[0130] The method of the present application may help, in at least some embodiments, to improve a user's interpretation of a scene by improving the sharpness of an object of interest such as a face that would initially be blurred via capture, by reconstructing it in three dimensions in a partial or total manner to increase its precision. The rendering may also, for example, comprise an application of visual effects to the spatialized image such as increasing the contrast, saturation of colors and / or the application of a luminous outline around the objects of interest to make them more visible and distinguish them from the rest of the scene. This may also be the case, for example, during a remote presentation, where a whiteboard may be detected as an object of interest by the method.For example, automatic brightness and contrast adjustment can be performed so that the board annotations are clearly visible to all participants. The process can also help to adapt the sharpness and / or resolution of graphics, for example, for better interpretation by participants. It can also be an online presentation where the process can detect a product held by the presenter and substitute it with a better quality interactive 3D model. Participants can then view the product from different angles or with different configurations, improving . thus the presentation experience. In a meeting focused on data analysis, the method can detect graphs or charts presented by participants as objects of interest and substitute them with interactive volumetric visualizations, allowing participants to manipulate and examine the data more intuitively. If a participant shows a physical object such as a product prototype, the method can render a spatialized version of that object that takes into account its actual position and size, allowing other participants to see the object as if it were present on their own desk.
[0131] The method of the present application can help, in at least certain embodiments, to facilitate the visualization of the rendering according to the specificities of a rendering equipment as in the case of equipment requiring multiple and simultaneous renderings. Two examples are illustrated in [Fig.6]. On the left of [Fig.6], 4 identical spatialized images are rendered simultaneously on the display screen of the terminal and by reflection using the principle of Pepper's ghost, on the four faces of a transparent truncated inverted pyramid (130, 131) thus producing a single image perceived as three-dimensional. On the right of [Fig.6], the same spatialized image is rendered simultaneously on 4 tablets arranged around a transparent truncated inverted pyramid reproducing the image perceived as three-dimensional.
[0132] The method of the present application may also help, in at least certain embodiments, to lower the energy consumption necessary for processing a multimedia stream by reducing the brightness of the rendering. This may for example be a rendering equipment of the “Pepper phantom” type applied to a connected terminal such as a smartphone and / or a tablet. Such an example of rendering is illustrated in [Fig.5] where a user U1 can see his own face and perceive it in 3 dimensions thanks to a mirror placed on a transparent blade positioned at approximately 45° from the screen of the terminal 110, the mirror allowing the capture of the user by the front camera 120 of the terminal. A dark room associated with the rendering equipment may consist of a darkened space or compartment around the transparent blade.The purpose of this darkroom would be to minimize the amount of stray light reaching the transparent slide, by absorbing ambient light and reducing unwanted reflections. Thus, the image projected onto the transparent slide would benefit from higher contrast. The process can reduce the brightness of the rendering by adapting it to the viewing conditions thus created.
[0133] The method of the present application can help, in at least certain embodiments, to improve the quality of the rendering of the spatialized audiovisual stream by allowing its capture by adapting the graphic parameters of the spatialized image on the fly upon obtaining the rendering, thus constituting a feedback loop. In the example illustrated in [Fig.7], the device 200 is a smartphone or tablet type terminal placed horizontally. The left part of [Fig.7] represents the terminal in profile along its length and the right part represents the terminal in profile along its width. A transparent screen is placed at 45° to the screen of the device to produce a “Pepper ghost” type display. As also shown in [Fig.5], a mirror 510 is placed on the screen so that the front camera 120 can capture the image of the user 101. A semi-transparent glass 710 placed perpendicular to the mirror 510 and to the transparent screen 130 can make it possible to capture the final rendering.Furthermore, spatializing an object of interest while taking into account a volumetric rendering context represents an advantage in processing audiovisual streams, particularly when projecting an image onto non-planar surfaces, such as, for example, a transparent cone used in Pepper's Ghost setup. During rendering, this approach can allow the image to be adapted and deformed so that it perfectly matches the shape of the projection medium, thus ensuring that the three-dimensional illusion is preserved and the projected image is not distorted or inappropriately altered. When projecting an image onto a conical surface, for example, a simple planar projection would result in significant distortion, as the image must expand to cover a larger area as it moves away from the apex of the cone.By taking into account the volumetric rendering context, the method can pre-calculate the adjustments necessary to make the final image appear correct from the observer's perspective. This involves distorting the original image in anticipation of how it will be stretched onto the conical surface, so that, once projected, it appears as a faithful and undistorted representation of the object of interest.
Claims
Claims
1. Method for processing an audiovisual stream, said method comprising: - obtaining at least one image from a first audiovisual stream, - detecting at least one object of interest in said image taking into account a volumetric capture context, - spatializing said at least one detected object of interest taking into account a volumetric rendering context.
2. Method according to claim 1 where said spatialization comprises a creation of at least one first spatialized image of said object of interest, said creation comprising: - a selection of a visual object from among candidate visual objects taking into account the characteristics of said detected object of interest, - a substitution of said object of interest in said image of said first stream or in said first spatialized image by said selected visual object.
3. The method of claim 2 wherein the method comprises rendering said at least one first spatialized image.
4. Method according to claim 2 where the method comprises a plurality of renderings on the same time window of said first spatialized image.
5. Method according to one of claims 3 to 4 where said method comprises, during said rendering, a dynamic adaptation of at least one graphic parameter of said first audiovisual stream and / or of at least one portion of said at least one first spatialized image.
6. Method according to claim 5 where said at least one graphic parameter belongs to a group comprising: - a graphic parameter corresponding to a resolution of said first spatialized image; - a graphic parameter corresponding to an image refresh rate of said first audiovisual stream; - a contrast adjustment parameter; - a brightness adjustment parameter; - a white balance adjustment setting; - a sharpness adjustment setting; - a combination of at least two of the above settings.
7. Method according to one of claims 3 to 6 where the rendering of said at least one first spatialized image takes into account said volumetric capture context and / or said at least one detected object of interest.
8. Method according to one of claims 2 to 7 wherein when said at least one detected object of interest is a living being, said selection of said visual object takes into account a similarity, in terms of bodily expression, between said at least one detected object of interest and said visual object.
9. Method according to one of claims 5 to 8 wherein said dynamic adaptation takes into account said visual object.
10. Method according to one of claims 1 to 9 where the method comprises a transmission of a second audiovisual stream containing said at least one first spatialized image.
11. Method according to one of claims 1 to 9 where the method comprises: - a reception of a third audiovisual stream; - a rendering of a second spatialized image obtained from said third audiovisual stream received.
12. Method according to one of claims 1 to 11 wherein the method is implemented by a first electronic device during a communication session with at least one second electronic device.
13. Electronic device comprising at least one processor adapted to: - obtaining at least one image from a first audiovisual stream, - detecting at least one object of interest in said image taking into account a volumetric capture context, - spatializing said at least one detected object of interest taking into account a volumetric rendering context.
14. Data processing system for processing an audiovisual stream comprising: a first electronic device comprising at least one processor adapted to: - obtaining at least one image from a first audiovisual stream, - a detection of at least one object of interest in said image taking into account a volumetric capture context, - a spatialization of said at least one detected object of interest taking into account a volumetric rendering context. - a transmission of a second audiovisual stream containing said at least one first spatialized image, a second electronic device comprising at least one processor adapted to: - a reception of said second audiovisual stream containing at least said first spatialized image; - a rendering of said first spatialized image.
Citation Information
Patent Citations
Optimizations for dynamic object instance detection, segmentation, and structure mapping
EP3493106A1
Motion-assisted image segmentation and object detection
US20200193609A1
System and method to enhance distant people representation
US20230306698A1