Spatial audio distributed processing method and device

The method enables a secondary device to independently manage spatial audio by creating an acoustic scene modeling engine based on received information, addressing performance and delay issues in immersive environments.

WO2025263775A1PCT designated stage Publication Date: 2025-12-26SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004893
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-02
Filing Date
2025-04-10
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently generating and managing spatial audio in immersive environments without relying on resources of the primary device or external servers, leading to delays and performance issues.

Method used

A method and device that allows a secondary electronic device to independently create, update, and manage an acoustic scene modeling engine of a first space by receiving spatial audio information from a first electronic device, enabling it to generate and reproduce audio of interest based on this engine, thus improving usability and responsiveness to environmental changes.

Benefits of technology

Enhances the usability and performance of immersive spatial audio services by allowing the secondary device to operate independently, reducing delays and improving responsiveness to changes in sound sources and network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004893_26122025_PF_FP_ABST
    Figure KR2025004893_26122025_PF_FP_ABST
Patent Text Reader

Abstract

Various embodiments of the present disclosure relate to a spatial audio distributed processing method and device. To this end, an electronic device may: virtually participate in a first space in which the electronic device is not located, on the basis of an external input; request spatial audio information of the first space from a first electronic device located in the first space; receive the spatial audio information of the first space from the first electronic device; and generate a first-space acoustic scene modeling engine on the basis of the spatial audio information of the first space.
Need to check novelty before this filing date? Find Prior Art

Description

Spatial audio distributed processing method and device

[0001] The present disclosure relates to a method and device for distributed processing of spatial audio. Embodiments of the present disclosure relate to a method and device in which an electronic device (hereinafter, referred to as a second electronic device) located in a second space can create, update, and manage an acoustic scene modeling engine of a first space in which the electronic device virtually participates, and can independently generate and play audio of interest corresponding to a direction of interest or a sound source of interest of the second electronic device in the first space based on the acoustic scene modeling engine of the first space created by the electronic device.

[0002] Advances in communication and media technologies have made it possible to experience the immersive or realistic sensations provided by Virtual Reality (VR) and Augmented Reality (AR) services using various electronic devices (e.g., head-mounted displays (HMDs). Furthermore, active development and standardization are underway to apply room acoustics technology to the metaverse virtual environment. The MPEG-I Immersive Audio standard is standardizing immersive realistic audio rendering technology and metadata encoding required for rendering to provide spatial audio services that enable listeners to freely move and interact in a three-dimensional acoustic space.

[0003] Through immersive spatial sound services that provide optimal spatial sound based on the user's location, users can immerse themselves in a three-dimensional sound space and experience a sense of realism as if they were actually there.

[0004] The present disclosure provides a method and device for distributed processing of spatial audio. Specifically, various embodiments of the present disclosure relate to a method and device in which an electronic device located in a second space can create, update, and manage an acoustic scene modeling engine of a first space in which the electronic device virtually participates, and can independently generate and reproduce audio of interest corresponding to a direction of interest or a sound source of interest of the electronic device in the first space based on the acoustic scene modeling engine of the first space created by the electronic device.

[0005] In order to solve the above-described technical problem, according to one embodiment of the present disclosure, an electronic device may include a communication circuit; at least one processor including a processing circuit; and a memory including one or more storage media storing instructions. When the instructions are individually or collectively executed by the at least one processor, they may cause the electronic device to virtually participate in a first space in which the electronic device is not located based on an external input. When the instructions are individually or collectively executed by the at least one processor, they may cause the electronic device to request spatial audio information of the first space from a first electronic device located in the first space. When the instructions are individually or collectively executed by the at least one processor, they may cause the electronic device to receive spatial audio information of the first space from the first electronic device. When the instructions are individually or collectively executed by the at least one processor, they may cause the electronic device to generate a first spatial acoustic scene modeling engine based on the spatial audio information of the first space.

[0006] In addition, according to one embodiment of the present disclosure, a method of operating an electronic device may include: an operation of virtually participating in a first space in which the electronic device is not located based on an external input; an operation of requesting spatial audio information of the first space from a first electronic device located in the first space; an operation of receiving spatial audio information of the first space from the first electronic device; and an operation of generating a first spatial audio scene modeling engine based on the spatial audio information of the first space.

[0007] In addition, according to one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing the method may be included.

[0008] According to various embodiments of the present disclosure, a second electronic device located in a second space can independently generate, update, and manage an acoustic scene modeling engine of the first space by receiving all spatial audio information of the first space from an electronic device (hereinafter, “first electronic device”) located in the first space. In other words, the second electronic device located in the second space can independently generate, update, and manage an acoustic scene modeling of the first space without relying on the first electronic device located in the first space.

[0009] According to various embodiments of the present disclosure, a second electronic device can independently generate, manage, and reproduce audio of interest of a first space based on an acoustic scene modeling engine of a first space generated by the second electronic device, without relying on resources (e.g., power, processor, memory, communication, battery, etc.) of the first electronic device or an external server. Accordingly, delay and performance issues can be improved compared to conventional passive operations of receiving audio of interest corresponding to a direction / source of interest of the second electronic device from a first electronic device by relying on an acoustic scene modeling engine of the first space located in the first electronic device, and the usability of an immersive spatial audio service can be enhanced.

[0010] According to various embodiments of the present disclosure, the second electronic device can generate, manage, and reproduce audio of interest of the first space based on the acoustic scene modeling engine of the first space that it has generated, so that it can more quickly and efficiently respond to various variable situations such as changes in sound sources of the first space (e.g., changes in sound sources such as sound generation / removal / movement, changes in sound sources such as sounds), changes in the position, direction, movement speed, movement distance, etc. of the second electronic device in the first space, changes in the direction / sound source of interest of the second electronic device, and changes in network performance between the first electronic device and the second electronic device, thereby improving the usability of an immersive spatial sound service.

[0011] The effects that can be obtained from the exemplary embodiments of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly derived and understood by those skilled in the art to which the exemplary embodiments of the present disclosure pertain from the following description. In other words, unintended effects resulting from implementing the exemplary embodiments of the present disclosure can also be derived by those skilled in the art from the exemplary embodiments of the present disclosure.

[0012] FIG. 1 schematically illustrates a spatial audio distributed processing environment according to one embodiment of the present disclosure.

[0013] FIG. 2 illustrates 6DoF user interaction for determining direction of interest or sound source of interest of an electronic device according to one embodiment of the present disclosure.

[0014] FIG. 3 is a block diagram of an electronic device located in a second space according to one embodiment of the present disclosure.

[0015] FIG. 4 is a schematic flowchart of an operation method of an electronic device located in a second space according to one embodiment of the present disclosure.

[0016] FIG. 5 illustrates a UI screen in which an electronic device located in a second space displays all sound sources in a first space, according to one embodiment of the present disclosure.

[0017] FIG. 6 schematically illustrates a spatial audio distributed processing environment in which a third electronic device (600) is added to the first space (110) illustrated in FIG. 1 according to one embodiment of the present disclosure.

[0018] FIG. 7 is a schematic flowchart of an operation method of an electronic device located in a second space when a third electronic device (600) is added to a first space (110) according to one embodiment of the present disclosure.

[0019] FIG. 8 is a block diagram of an electronic device within a network environment according to various embodiments.

[0020] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0021] FIG. 1 schematically illustrates a spatial audio distributed processing environment according to one embodiment of the present disclosure.

[0022] In the spatial audio distributed processing environment illustrated in FIG. 1, a first electronic device (130) located in a first space (110) can virtually invite a second electronic device (140) (a second user) located in a second space (120), which is an external space, to the first space (110) based on a first user input of the first electronic device (130), etc. At least one sound source (e.g., a first sound source, a second sound source, a third sound source) may exist in the first space (110) either physically or virtually. According to one embodiment, the first electronic device (130) and the second electronic device (140) may include various wearable devices such as a head-mounted display (HMD) or a wristwatch, but are not limited thereto, and may include various types of electronic devices including at least one microphone (e.g., a smart phone, a smart tablet, a PDA, a kiosk, an electronic picture frame, a navigation device, a smart TV, etc.). The first electronic device (130) and the second electronic device (140) can provide an augmented reality (AR) service, a virtual reality (VR) service, or a metaverse virtual environment service.

[0023] According to one embodiment, the second electronic device (140) may determine whether the second electronic device (140) operates in a mode in which the second electronic device (140) models an acoustic scene of the first space (110) (hereinafter, referred to as the second electronic device acoustic scene modeling mode). According to one embodiment, the second electronic device acoustic scene modeling mode may be automatically determined by default, determined based on a user input of the second electronic device (140), or determined based on recognition of a predetermined object or predetermined sound source of the first space (110), but is not limited thereto. For example, when the second electronic device (140) participates in a first space (110) in which a real-time user experience is important, such as a game space, the second electronic device (140) may determine to operate in the second electronic device acoustic scene modeling mode when virtually participating in the first space (110) by an invitation from the first electronic device (130). The second electronic device (140) can determine whether the second electronic device is operating in the acoustic scene modeling mode.

[0024] According to one embodiment, when the second electronic device is in the sound scene modeling mode, the second electronic device (140) may request spatial audio information of the first space (110) from the first electronic device (130) and receive spatial audio information of the first space (110) from the first electronic device (130). The requesting and receiving of the spatial audio information may be performed in real time. The spatial audio information of the first space may include sound source data input to at least one microphone (131, 132, 133) of a first electronic device (130) located in the first space (110), a position, direction, echo, reverberation, sensitivity, performance, movement direction, movement speed of at least one microphone (131, 132, 133) of the first electronic device (130), a change amount or change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the movement direction and the movement speed, a relative distance between at least one microphone (131, 132, 133) of the first electronic device (130), sensing data acquired from a camera or at least one sensor of the first electronic device (130), and at least one of a change amount or change direction of the sensing data. The at least one sensor may include, but is not limited to, a proximity sensor, a gyro sensor, a gesture sensor, a pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an ambient light sensor.

[0025] According to one embodiment, when a plurality of electronic devices are included in the first space (110), the second electronic device (140) may request spatial audio information of the first space (110) from each of the plurality of electronic devices. The second electronic device (140) may receive spatial audio information of the first space (110) from each of the plurality of electronic devices. The request and reception of the spatial audio information may be performed in real time. The spatial audio information of the first space may include sound source data input to at least one microphone of each electronic device located in the first space (110), a position, a direction, an echo, a reverberation, a sensitivity, a performance, a moving direction, a moving speed of at least one microphone of each electronic device, a change amount or a change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the moving direction and the moving speed, a relative distance between at least one microphone of each electronic device, sensing data acquired from a camera or at least one sensor of each electronic device, and at least one of a change amount or a change direction of the sensing data. The at least one sensor may include, but is not limited to, a proximity sensor, a gyro sensor, a gesture sensor, a pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an ambient light sensor.

[0026] According to one embodiment, the second electronic device (140) can generate an acoustic scene modeling engine of the first space (110) based on spatial audio information of the first space (110). The second electronic device (140) can increase the accuracy of acoustic scene modeling of the first space (110) by obtaining spatial audio information of the first space (110) from a plurality of electronic devices included in the first space (110). After the acoustic scene modeling engine of the first space (110) is generated in the second electronic device (140), the second electronic device (140) can update the acoustic scene modeling engine of the first space (110) based on spatial audio information of the first space (110) received in real time. The requesting and receiving of the spatial audio information and the updating of the acoustic scene modeling engine can be continuously performed in real time.

[0027] According to one embodiment, the second electronic device (140) can generate and play audio of interest corresponding to the direction of interest or sound source of interest (e.g., a third sound source) of interest of the second user based on an acoustic scene modeling engine of the first space (110) that the second electronic device (140) generates or updates itself as if the second electronic device (140) is actually in the first space (110). A method of determining the direction of interest or sound source of interest of the second user will be described below with reference to FIG. 2. The second electronic device (140) can generate audio of interest having a high SNR (Single-Noise Ratio) of the direction of interest or sound source of interest of the second user by tracking, separating, and performing post-processing including beamforming, echo, reverberation, and noise processing of the direction of interest or sound source of interest of the second user.

[0028] In addition, according to one embodiment, the second electronic device (140) may be placed at a virtual location in the first space (110). The virtual location of the second electronic device (140) may be set to the same location as the first electronic device (130) or to a predetermined location different from the location of the first electronic device (130) based on an external input. The second electronic device (140) may generate audio of interest of the second user corresponding to the virtual location in real time by post-processing the relative location change vector value between its own virtual location in the first space (110) and the location of the first electronic device (130) to generate audio of interest. That is, the second electronic device (140) may generate audio of interest of the second user corresponding to the virtual location by performing post-processing including an operation of mathematically multiplying the relative location change vector value between the virtual location of the second electronic device (140) and the location of the first electronic device (130) to the audio of interest generated based on the location of the first electronic device (130).

[0029] For example, a first space (110) may be a company conference room, and a second space (120) may be a home. A first user of a first electronic device (130) may be holding a meeting in the company conference room, and multiple sound sources (e.g., a first sound source, a second sound source, and a third sound source) may be present in the company conference room. A second electronic device (140) at home may virtually participate in the meeting in the company conference room upon an invitation from the first electronic device (130) based on an input from the first user. The second electronic device (140) may receive spatial audio information corresponding to the company conference room space from the first electronic device (130). The second electronic device (140) may independently generate an acoustic scene modeling engine corresponding to the company conference room based on the received spatial audio information. When a new sound source is added to the company conference room, the second electronic device (140) may receive spatial audio information including information corresponding to the added sound source from the first electronic device (130). The second electronic device (140) can update the acoustic scene modeling engine corresponding to the company conference room based on the received spatial audio information. In addition, when a new electronic device is added to the company conference room, the second electronic device (140) can receive spatial audio information corresponding to the company conference room space from the new electronic device. The second electronic device (140) can update the acoustic scene modeling engine corresponding to the company conference room based on the received spatial audio information. The second electronic device (140) can continuously acquire spatial audio information corresponding to the company conference room space and update the acoustic scene modeling engine in real time. Based on the acoustic scene modeling engine for the company conference room, the second electronic device (140) can automatically generate and play audio of interest corresponding to the direction of interest or the sound source of interest of the second user among multiple sound sources present in the company conference room space.

[0030] According to one embodiment, the second electronic device (140) can determine whether the first electronic device (130) operates in a mode for modeling an acoustic scene of the first space (110) (hereinafter, referred to as the first electronic device acoustic scene modeling mode). According to one embodiment, the first electronic device acoustic scene modeling mode may be automatically determined by default, determined based on a user input of the second electronic device (140), or determined based on recognition of a predetermined object or predetermined sound source of the first space (110), but is not limited thereto. According to one embodiment, the second electronic device (140) can determine whether the first electronic device operates in the acoustic scene modeling mode.

[0031] According to one embodiment, when the first electronic device is in an acoustic scene modeling mode, the first electronic device (130) can obtain spatial audio information of the first space (110). The spatial audio information of the first space may include sound source data input to at least one microphone (131, 132, 133) of the first electronic device (130), position, direction, echo, reverberation, sensitivity, performance, movement direction, movement speed of at least one microphone (131, 132, 133) of the first electronic device (130), change amount or change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to change in the movement direction and the movement speed, a relative distance between at least one microphone (131, 132, 133) of the first electronic device (130), sensing data acquired from a camera or at least one sensor of the first electronic device (130), and at least one of change amount or change direction of the sensing data.

[0032] According to one embodiment, when a plurality of electronic devices are included in a first space (110), the first electronic device (130) may request spatial audio information of the first space (110) from each of the plurality of electronic devices. The first electronic device (130) may receive spatial audio information of the first space (110) from each of the plurality of electronic devices. The request and reception of the spatial audio information may be performed in real time. The spatial audio information of the first space may include sound source data input to at least one microphone of each electronic device located in the first space (110), a position, a direction, an echo, a reverberation, a sensitivity, a performance, a moving direction, a moving speed of at least one microphone of each electronic device, a change amount or a change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the moving direction and the moving speed, a relative distance between at least one microphone of each electronic device, sensing data acquired from a camera or at least one sensor of each electronic device, and at least one of a change amount or a change direction of the sensing data. The at least one sensor may include, but is not limited to, a proximity sensor, a gyro sensor, a gesture sensor, a pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an ambient light sensor.

[0033] According to one embodiment, the first electronic device (130) can generate and update an acoustic scene modeling engine of the first space (110) based on spatial audio information of the first space (110). The first electronic device (130) can continuously perform acquisition of the spatial audio information and updating of the acoustic scene modeling engine in real time. The first electronic device (130) can increase the accuracy of acoustic scene modeling of the first space (110) by acquiring spatial audio information of the first space (110) from a plurality of electronic devices included in the first space (110).

[0034] According to one embodiment, the first electronic device (130) can generate and play the audio of interest of the first user corresponding to the direction of interest or the sound source of interest (e.g., the first sound source) of the first user based on the acoustic scene modeling engine of the first space (110) that it has generated or updated. A method of determining the direction of interest or the sound source of interest of the first user will be described below with reference to FIG. 2. The first electronic device (130) can generate the audio of interest with a high SNR (Single-Noise Ratio) of the direction of interest or the sound source of interest by tracking, separating, and performing post-processing including beamforming, echo, reverberation, and noise processing of the direction of interest or the sound source of interest. In addition, the first electronic device (130) can generate the audio of interest of the first user to be shared with the second electronic device (140) of the virtual location of the first space (110). The first electronic device (130) can generate audio of interest of the first user to be shared with the second electronic device (140) of the virtual location by post-processing and reflecting the relative position change vector value with respect to the virtual location of the second electronic device (140). The virtual location of the second electronic device (140) can be set to the same location as the first electronic device (130) or to a predetermined location different from the location of the first electronic device (130) based on an external input. In addition, the first electronic device (130) can generate audio of interest by customizing and optimizing spatial acoustic parameters for each spatial acoustic scene or each sound source object. The first electronic device (130) can generate audio of interest with customizing and optimizing acoustic parameters and sound effects by differently applying additional effects corresponding to the receiving device, the user, and the service situation, and can share the audio of interest with the second electronic device (140). The additional effects can include, but are not limited to, reverberation parameters, sound delay effects, and calculation of diffraction paths of sound source-listener pairs for static / dynamic sound sources.

[0035] According to one embodiment, the first electronic device (130) can transmit the first user's interest audio corresponding to the virtual location to the second electronic device (140). The second electronic device (140) can play the received first user's interest audio.

[0036] According to one embodiment, in the first electronic device sound scene modeling mode, when the second electronic device (140) wants to play audio of interest corresponding to the direction of interest or sound source of interest of a second user of the second electronic device (140), the second electronic device (140) may request the first electronic device (130) to play audio of interest corresponding to the direction of interest or sound source of interest of the second user by providing information indicating the direction of interest or sound source of interest of the second user to the first electronic device (130). A method of determining the direction of interest or sound source of interest of the second user will be described below with reference to FIG. 2.

[0037] According to one embodiment, the first electronic device (130) may generate audio of interest of the second user corresponding to the direction of interest or sound source of interest of the second user based on the acoustic scene modeling engine of the first space (110) that it has generated or updated. At this time, the first electronic device (130) may post-process and reflect a relative position change vector value with respect to the virtual position of the second electronic device (140), thereby generating audio of interest of the second user to be shared with the second electronic device (140) of the virtual position. The first electronic device (130) may transmit the audio of interest of the second user corresponding to the virtual position to the second electronic device (140). The second electronic device (140) may play the received audio of interest of the second user.

[0038] FIG. 2 illustrates 6DoF user interaction for determining direction of interest or sound source of interest of an electronic device according to one embodiment of the present disclosure.

[0039] According to one embodiment, the first electronic device (130) or the second electronic device (140) can determine the direction of interest of the user or the sound source of interest based on at least one of a head-related transfer function (HRTF), a user input, and a preset condition, but it will be understood by those skilled in the art that the present invention is not limited thereto. The preset condition can include, but is not limited to, a sound source with the highest energy, a sound source in a predetermined direction, and a sound source in the direction of a predetermined object.

[0040] According to one embodiment, a user of the first electronic device (130) or a user of the second electronic device (140) can freely move his or her head in a three-dimensional acoustic space. User interaction, such as movement of the user's head (200), can be expressed as a 6DoF (Degree of Freedom) movement. Referring to FIG. 2, the 6DoF movement can include a rotational movement in the x-axis (roll), y-axis (pitch), and z-axis (yaw) directions in a three-dimensional space, and a translational movement parallel to each of the three axes.

[0041] According to one embodiment, the first electronic device (130) or the second electronic device (140) can sense the 6DoF movement based on the user's head (200) and measure the Head Related Transfer Function (HRTF) based on the sensed user's head movement, thereby determining the user's direction of interest or the sound source of interest. The HRTF can be defined as a transfer function until an acoustic signal radiated from the location of a sound source is transmitted to both ears of the user. The HRTF can define a phenomenon in which more complex path characteristics, such as diffraction on the head surface and reflection by the pinna, change according to the sound transmission direction, in addition to simple path differences due to the inter-aural level difference (ILD) and the inter-aural time difference (ITD) between the two ears. Therefore, the first electronic device (130) or the second electronic device (140) can determine the user's direction of interest or the sound source of interest through the HRTF.

[0042] FIG. 3 is a block diagram of an electronic device located in a second space according to one embodiment of the present disclosure.

[0043] According to one embodiment, the electronic device (300) located in the second space (120) may be an electronic device corresponding to the second electronic device (140) illustrated in FIG. 1. In the operation of the electronic device described in FIG. 3, portions overlapping with those described in FIGS. 1 and 2 may be omitted. The electronic device (300) may include additional components in addition to the illustrated components, or may omit at least one of the illustrated components.

[0044] According to one embodiment, the electronic device (300) may include a communication circuit (not shown), at least one processor (not shown) including a processing circuit; and a memory (not shown) including one or more storage media for storing commands. The at least one processor may execute at least one command constituting a scene module (310), a sound source management module (330), a sound source determination module (340), a sound source generation module (350), and a sound source reproduction module (360).

[0045] According to one embodiment, the electronic device (300) may be virtually invited to the first space (110) based on a first user input of a first electronic device (130) located in the first space (110). It will be understood by those skilled in the art that there may be various communication methods for inviting to the first space (110) using the above communication circuit.

[0046] According to one embodiment, the electronic device (300) can determine whether the electronic device (300) operates in a mode (second electronic device acoustic scene modeling mode) in which the electronic device (300) models an acoustic scene of the first space (110). According to one embodiment, the second electronic device acoustic scene modeling mode may be automatically determined by default, determined based on a user input of the electronic device (300), or determined based on recognition of a predetermined object or a predetermined sound source of the first space (110), but is not limited thereto. The electronic device (300) can determine whether it operates in the second electronic device acoustic scene modeling mode. Hereinafter, a case in which the electronic device (300) operates in the second electronic device acoustic scene modeling mode will be described.

[0047] According to one embodiment, the scene module (310), when executed by the at least one processor, may request spatial audio information of the first space (110) from the first electronic device (130) and obtain spatial audio information of the first space (110) from the first electronic device (130). The spatial audio information of the first space may include sound source data input to at least one microphone (131, 132, 133) of a first electronic device (130) located in the first space (110), a position, direction, echo, reverberation, sensitivity, performance, movement direction, movement speed of at least one microphone (131, 132, 133) of the first electronic device (130), a change amount or change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the movement direction and the movement speed, a relative distance between at least one microphone (131, 132, 133) of the first electronic device (130), sensing data acquired from a camera or at least one sensor of the first electronic device (130), and at least one of a change amount or change direction of the sensing data. The scene module (310) can obtain a timestamp corresponding to the spatial audio information of the first space from the first electronic device (130). The scene module (310) can generate a reference timestamp based on the time at which the spatial audio information of the first space (110) is obtained. The scene module (310) can generate a reference timestamp by applying a network delay and a processing delay to the timestamp obtained from the first electronic device (130). However, it will be understood by those skilled in the art that the applied delay is not limited thereto and may vary. The spatial audio information can be generated using a custom field of the MPEG-I Immersive Audio standard, but is not limited thereto.

[0048] According to one embodiment, the scene module (310), when executed by the at least one processor, may request spatial audio information of the first space (110) from a server in real time and obtain spatial audio information of the first space (110) from the server in real time. The server may be located in the cloud or in the first space (110). The server may obtain spatial audio information of the first space (110) from at least one electronic device of the first space (110).

[0049] According to one embodiment, the scene module (310) may generate and update the sound scene modeling engine (320) of the first space (110) based on the spatial audio information of the first space (110). The sound scene modeling engine (320) of the first space (110) may include information modeling which sound source is producing a sound at which location in the first space (110) and at what volume. The sound scene modeling engine (320) may model the sound scene of the first space (110) based on the reference timestamp generated based on the time at which the spatial audio information of the first space (110) is acquired. The output value of the sound scene modeling engine (320) may include at least one of the position, sound volume, movement direction, movement change amount, change acceleration, change amount of sound volume and direction, and change acceleration of each sound source in the first space (110). The output value of the sound scene modeling engine (320) may be output in real time.

[0050] According to one embodiment, the scene module (310) can update the acoustic scene modeling engine (320) based on spatial audio information of the first space (110). The scene module (310) can obtain spatial audio information of the first space (110) in real time and update the acoustic scene modeling engine (320) based on the spatial audio information in real time.

[0051] In one embodiment, the scene module (310) stores the acoustic scene modeling engine (320) of the first space (110) and can be reused in the future.

[0052] According to one embodiment, the sound source management module (330) may, when executed by the at least one processor, separately manage all sound sources of the first space (110) based on the sound scene modeling engine (320) of the first space (110). The sound source management module (330) may determine the location, size, movement direction, movement speed, etc. of each sound source in relation to the timestamp.

[0053] According to one embodiment, the sound source of interest determination module (340) can determine the direction of interest or sound source of interest of a user of the electronic device (300) by at least one of a head-related transfer function (HRTF), a user input, and a preset condition when executed by the at least one processor, but it will be understood by those skilled in the art that the preset condition may include, but is not limited to, a sound source with the highest energy, a sound source in a predetermined direction, and a sound source in the direction of a predetermined object.

[0054] According to one embodiment, the sound source generation module (350), when executed by the at least one processor, can generate audio of interest corresponding to the direction of interest or sound source of interest of the user of the electronic device (300), determined by the sound source of interest determination module (340) among all sound sources of the first space (110) managed by the sound source management module (330). According to one embodiment, since the acoustic scene modeling engine (320) of the first space (110) exists in the electronic device (300), audio of interest corresponding to the direction of interest or sound source of interest of the first user of the first electronic device (300) can be generated independently of the direction of interest or sound source of interest of the first user of the electronic device (300). The sound source generation module (350) can generate audio of interest having a high SNR (Single-Noise Ratio) of the direction of interest of the user or the sound source of interest by tracking, separating, and performing post-processing including beamforming, echo, reverberation, and noise processing of the direction of interest of the user or the sound source of interest.

[0055] According to one embodiment, the sound source generation module (350) can generate audio of interest of a user of the electronic device (300) corresponding to the virtual location by post-processing a relative position change vector value between its own virtual location and the location of the first electronic device (130) in the first space (110) to generate audio of interest. The virtual location of the electronic device (300) can be set to the same location as the first electronic device (130) or to a predetermined location different from the location of the first electronic device (130) based on an external input.

[0056] According to one embodiment, at least one of the scene module (310), sound source management module (330) and sound source generation module (350) described above may be executed by at least one other electronic device (e.g., server) that is communicatively connected to the electronic device (300).

[0057] According to one embodiment, the audio playback module (360) can play audio of interest generated by the audio generation module (350).

[0058] FIG. 4 is a schematic flowchart of an operation method of an electronic device located in a second space according to one embodiment of the present disclosure.

[0059] According to one embodiment, the electronic device located in the second space (120) may be an electronic device corresponding to the electronic device (300) illustrated in FIG. 3. In the operation of the electronic device described in FIG. 4, portions overlapping with those described in FIGS. 1 to 3 may be omitted. Some of the operations illustrated in FIG. 4 may be omitted, and operations not illustrated in FIG. 4 may be added.

[0060] Referring to FIG. 4, in operation 410 according to one embodiment, the electronic device (300) can virtually participate in a first space (110) where the electronic device (300) is not located based on an external input.

[0061] In operation 420 according to one embodiment, the electronic device (300) may determine whether the electronic device (300) is in an operation mode in which the electronic device (300) generates an acoustic scene modeling engine of the first space (110). If the electronic device (300) is in an operation mode in which the electronic device (300) generates an acoustic scene modeling engine of the first space (110) (second electronic device acoustic scene modeling mode), the electronic device may move to operation 430, and if the first electronic device (130) is in an operation mode in which the first electronic device (130) generates an acoustic scene modeling engine of the first space (110) (first electronic device acoustic scene modeling mode), the electronic device may move to operation 480. The operation mode may be determined based on a user input of the electronic device (300), or may be determined based on recognition of a predetermined object or a predetermined sound source, but is not limited thereto.

[0062] In operation 430 according to one embodiment, the electronic device (300) may request spatial audio information of a first space from a first electronic device (130) located in the first space (110). The spatial audio information of the first space (110) may include sound source data input to at least one microphone (131, 132, 133) of the first electronic device (130), a position, a direction, an echo, a reverberation, a sensitivity, a performance, a moving direction, a moving speed of at least one microphone (131, 132, 133) of the first electronic device (130), a change amount or a change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the moving direction and the moving speed, a relative distance between at least one microphone (131, 132, 133) of the first electronic device (130), sensing data acquired from a camera or at least one sensor of the first electronic device (130), and at least one of a change amount or a change direction of the sensing data.

[0063] In operation 440 according to one embodiment, the electronic device (300) may receive spatial audio information of the first space (110) from the first electronic device (130).

[0064] In operation 450 according to one embodiment, the electronic device (300) may generate a first spatial audio scene modeling engine based on spatial audio information of the first space (110). The operation of generating the first spatial audio scene modeling engine may be initiated interchangeably with the operation of the electronic device (300) copying the first spatial audio scene modeling engine or the operation of moving the first spatial audio scene modeling engine to the electronic device (300).

[0065] In operation 460 according to one embodiment, the electronic device (300) may generate audio of interest in the first space (110) corresponding to the direction of interest of the user of the electronic device (300) or the sound source of interest based on the first spatial acoustic scene modeling engine. The direction of interest of the user of the electronic device (300) or the sound source of interest may be independent from the direction of interest of the user of the first electronic device (130) or the sound source of interest. That is, at a specific point in time, the direction of interest of the user of the electronic device (300) or the sound source of interest may be different from the direction of interest of the user of the first electronic device (130) or the sound source of interest. The sound source of interest or the direction of interest may be determined by at least one of a head-related transfer function (HRTF), a user input of the electronic device (300), and a preset condition, but is not limited thereto. The preset condition may include, but is not limited to, a sound source with the highest energy, a sound source in a predetermined direction, and a sound source in the direction of a predetermined object. According to one embodiment, the electronic device (300) can generate audio of interest of a user of the electronic device (300) corresponding to the virtual location by post-processing a vector value of a relative position change between its own virtual location and the location of the first electronic device (130) in the first space (110) to generate audio of interest.

[0066] In operation 470 according to one embodiment, the electronic device (300) can play the audio of interest of the first space generated above.

[0067] In operation 480 according to one embodiment, the electronic device (300) may operate in a first electronic device acoustic scene modeling mode. The first electronic device (130) may generate and update an acoustic scene modeling engine of the first space (110) based on spatial audio information of the first space (110) that it has acquired. When the electronic device (300) wants to play audio of interest corresponding to a direction of interest or a sound source of interest of a user (or a second user) of the electronic device (300), the electronic device (300) may request audio of interest corresponding to the direction of interest or the sound source of interest of the user from the first electronic device (130) by providing information corresponding to the direction of interest or the sound source of interest of the user to the first electronic device (130). The first electronic device (130) may generate audio of interest of the user corresponding to the direction of interest or the sound source of interest based on the acoustic scene modeling engine of the first space (110) that it has generated or updated. At this time, the first electronic device (130) can generate the user's interest audio to be shared with the electronic device (300) of the virtual location by post-processing and reflecting the relative position change vector value with respect to the virtual location of the electronic device (300). The first electronic device (130) can transmit the user's interest audio corresponding to the virtual location to the electronic device (300). The electronic device (300) can play the received user's interest audio.

[0068] According to one embodiment, the electronic device (300) may obtain at least one sound source of the first space (110) from the first electronic device (130) as an optional operation based on a user input, even in an operating mode in which the electronic device (300) generates an acoustic scene modeling engine of the first space (110). The first electronic device (130) may generate at least one sound source of the first space (110) based on the acoustic scene modeling engine of the first space (110) that it has generated or updated, and transmit the at least one sound source of the first space (110) to the electronic device (300) in response to a request from the electronic device (300). Even if the operating mode in which the electronic device (300) generates the sound scene modeling engine of the first space (110) operates unstably, the electronic device (300) can stably reproduce the direction of interest or the sound source of interest of the user of the electronic device (300) based on at least one sound source of the first space (110) received from the first electronic device (130).

[0069] FIG. 5 illustrates a UI screen in which an electronic device located in a second space displays all sound sources in a first space, according to one embodiment of the present disclosure.

[0070] As described above, the electronic device (300) can generate and update the sound scene modeling engine (320) of the first space (110) based on the spatial audio information of the first space (110) obtained from the first electronic device (130) located in the first space. The sound scene modeling engine (320) of the first space (110) can include modeled information such as which sound source is producing a sound at which location in the first space (110) and at what volume. The electronic device (300) can separately manage all sound sources of the first space (110) based on the sound scene modeling engine (320) of the first space (110). Accordingly, the electronic device (300) can determine at least one state of each sound source, such as the location, size, movement direction, and movement speed, at a specific time, and can display at least one state of each sound source (e.g., the location of the sound source) by configuring it into a UI screen (500) as in FIG. 5. The electronic device (300) can also display its own virtual location (510) in the first space (110) by configuring it as a UI screen (500). The UI screen (500) of FIG. 5 is merely an example, and it will be understood by those skilled in the art that various UI screen configurations that can express at least one state of sound sources existing in the first space (110) and the virtual location of the electronic device (300) are possible.

[0071] Additionally, according to one embodiment, the electronic device (300) may display on a UI screen whether the electronic device (300) (or the second electronic device (140)) is in a mode of modeling an acoustic scene of the first space (110) at a specific point in time, or whether the first electronic device (130) is in a mode of modeling an acoustic scene of the first space (110).

[0072] When the first electronic device (130) operates in a mode for modeling the sound scene of the first space (110), and the electronic device (300) plays the spatial audio received from the first electronic device (130), the electronic device (300) cannot independently configure and display all sound sources of the first space (110) as a UI screen without depending on the first electronic device (1300). That is, when the first electronic device operates in the sound scene modeling mode, the electronic device (300) (or the second electronic device (140)) can request the first electronic device (130) for the location and movement information of each sound source of the first space (110) and obtain the information in response to the request, thereby configuring the UI screen. In addition, when the first electronic device operates in the sound scene modeling mode, the timestamp corresponding to each sound source is a timestamp generated based on the time at which the first electronic device (130) models the sound scene of the first space (110), so that the electronic device (300) can be configured and displayed as a UI screen with a delay equal to the time it takes for the first electronic device (130) to obtain the location and movement information of each sound source from the first electronic device (130).

[0073] On the other hand, when the electronic device (300) (or the second electronic device (140)) operates in a mode for modeling the sound scene of the first space (110) (i.e., the second electronic device sound scene modeling mode), the electronic device (300) can determine by itself which sound source is producing a sound at which location in the first space (110) and at what volume based on the sound scene modeling engine (320) of the first space (110) generated by the electronic device (300), and thus the electronic device (300) can independently configure and display all sound sources of the first space (110) as a UI screen in relation to the first electronic device (1300). In addition, when operating in the second electronic device sound scene modeling mode, the timestamp corresponding to each sound source is a timestamp generated based on the time at which the electronic device (300) acquired spatial audio information of the first space (110) (e.g., a timestamp to which network delay and processing delay are applied), so the electronic device (300) can configure and display a UI screen including all sound sources of the first space (110) without a separate delay.

[0074] FIG. 6 schematically illustrates a spatial audio distributed processing environment in which a third electronic device (600) is added to the first space (110) illustrated in FIG. 1 according to one embodiment of the present disclosure. FIG. 7 schematically illustrates a flowchart of an operation method of an electronic device (140) (or a second electronic device) located in a second space when a third electronic device (600) is added to the first space (610) according to one embodiment of the present disclosure.

[0075] According to one embodiment, the electronic device (140) located in the second space (120) may be an electronic device corresponding to the electronic device (300) illustrated in FIG. 3. In the operation of the electronic device described in FIG. 7, portions overlapping with those described in FIGS. 1 to 3 may be omitted. Some of the operations illustrated in FIG. 7 may be omitted, and operations not illustrated in FIG. 7 may be added.

[0076] As described above with reference to FIG. 1, while the second electronic device (140) is operating in a mode of modeling an acoustic scene of the first space (110), a new third electronic device (600) may be added to the first space (110). The third electronic device (600) may be actually located in the first space (110), or may be located in a space different from the first space (e.g., the third space) but may be virtually added to a predetermined location in the first space. The second electronic device (140) may display by configuring a UI screen that a new third electronic device (600) has been added to the first space (110). The second electronic device (140) may display by configuring a UI screen the actual or virtual location of the third electronic device (600) in the first space (110).

[0077] In operation 710 according to one embodiment, the second electronic device (140) may request spatial audio information of the first space (110) that the third electronic device (600) can obtain from the third electronic device (600). The spatial audio information of the first space (110) that the third electronic device (600) can obtain may include sound source data input through at least one microphone of the third electronic device (600), a position, a direction, an echo, a reverberation, a sensitivity, a performance, a moving direction, a moving speed of at least one microphone of the third electronic device (600), a change amount or a change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the moving direction and the moving speed, a relative distance between at least one microphone of the third electronic device (600), sensing data obtained from a camera or at least one sensor of the third electronic device (600), and at least one of a change amount or a change direction of the sensing data. The at least one sensor may include, but is not limited to, a proximity sensor, a gyro sensor, a gesture sensor, a pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an ambient light sensor.

[0078] According to one embodiment, when a third electronic device (600) outputs a predetermined sound source through its speaker, the sound source (e.g., a fourth sound source) output from the speaker may be added to the first space (110). In this case, the second electronic device (140) may request spatial audio information of the first space that additionally reflects the fourth sound source from all electronic devices in the first space.

[0079] In operation 720 according to one embodiment, the second electronic device (140) can receive spatial audio information of a first space (110) that can be acquired by the third electronic device (600) from the third electronic device (600).

[0080] In one embodiment, when the second electronic device (140) cannot directly communicate with the third electronic device (600), the second electronic device (140) may request and receive spatial audio information of the first space (110) that the third electronic device (600) can obtain from another electronic device (e.g., the first electronic device (130)) or a server in the first space (100). The server may be located in the first space (100) or in the cloud.

[0081] According to one embodiment, when a third electronic device (600) is added to the first space (110), the third electronic device (600) can transmit spatial audio information of the first space (110) that the third electronic device (600) can obtain to all electronic devices participating in the first space (100) (e.g., the first electronic device (130), the second electronic device (140)). If the second electronic device (140) cannot directly communicate with the third electronic device (600), the spatial audio information of the first space (110) that the third electronic device (600) can obtain can be transmitted to the second electronic device (140) through another electronic device (e.g., the first electronic device (130)) of the first space (100) or a server. Accordingly, the second electronic device (140) can receive spatial audio information of the first space (110) that can be acquired by the third electronic device (600) from the third electronic device (600) through direct communication or indirect communication.

[0082] In operation 730 according to one embodiment, the second electronic device (140) can update the first spatial audio scene modeling engine based on the spatial audio information of the first space (110) received from the third electronic device (600). The second electronic device (140) generated the first spatial audio scene modeling engine based on the spatial audio information of the first space (110) received from the first electronic device (130) in operation 450 described above with reference to FIG. 4. In operation 730 according to one embodiment, the second electronic device (140) can update the first spatial audio scene modeling engine generated in operation 450 based on the spatial audio information of the first space (110) received from the third electronic device (600).

[0083] In operation 740 according to one embodiment, the second electronic device (140) may generate audio of interest of the first space (110) corresponding to the direction of interest or sound source of interest of the user of the second electronic device (140) based on the updated first spatial acoustic scene modeling engine.

[0084] In operation 750 according to one embodiment, the second electronic device (140) can play the audio of interest of the first space (110) generated above.

[0085] According to one embodiment, by additionally reflecting spatial audio information of the first space (110) that can be acquired by the third electronic device (600) into the first spatial sound scene modeling engine, a first spatial sound scene in which the third electronic device (600) is added to the actual or virtual location of the first space (110) can be modeled, and the accuracy of the first spatial sound scene modeling engine can be improved.

[0086] FIG. 8 is a block diagram of an electronic device within a network environment according to various embodiments.

[0087] Referring to FIG. 8, in a network environment (800), an electronic device (801) may communicate with an electronic device (802) via a first network (898) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (804) or a server (808) via a second network (899) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (801) may communicate with the electronic device (804) via the server (808). According to one embodiment, the electronic device (801) may include a processor (820), a memory (830), an input module (850), an audio output module (855), a display module (860), an audio module (870), a sensor module (876), an interface (877), a connection terminal (878), a haptic module (879), a camera module (880), a power management module (888), a battery (889), a communication module (890), a subscriber identification module (896), or an antenna module (897). In some embodiments, the electronic device (801) may omit at least one of these components (e.g., the connection terminal (878)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (876), the camera module (880), or the antenna module (897)) may be integrated into one component (e.g., the display module (860)).

[0088] The processor (820) may, for example, execute software (e.g., a program (840)) to control at least one other component (e.g., a hardware or software component) of the electronic device (801) connected to the processor (820) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (820) may store commands or data received from other components (e.g., a sensor module (876) or a communication module (890)) in a volatile memory (832), process the commands or data stored in the volatile memory (832), and store result data in a non-volatile memory (834). According to one embodiment, the processor (820) may include a main processor (821) (e.g., a central processing unit or an application processor) or an auxiliary processor (823) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (821). For example, when the electronic device (801) includes the main processor (821) and the auxiliary processor (823), the auxiliary processor (823) may be configured to use less power than the main processor (821) or to be specialized for a given function. The auxiliary processor (823) may be implemented separately from the main processor (821) or as a part thereof.

[0089] The auxiliary processor (823) may control at least a portion of functions or states associated with at least one component (e.g., a display module (860), a sensor module (876), or a communication module (890)) of the electronic device (801), for example, on behalf of the main processor (821) while the main processor (821) is in an inactive (e.g., sleep) state, or together with the main processor (821) while the main processor (821) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (823) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (880) or a communication module (890)). In one embodiment, the auxiliary processor (823) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (801) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (808)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0090] The memory (830) can store various data used by at least one component (e.g., the processor (820) or the sensor module (876)) of the electronic device (801). The data can include, for example, software (e.g., the program (840)) and input data or output data for commands related thereto. The memory (830) can include volatile memory (832) or non-volatile memory (834).

[0091] The program (840) may be stored as software in the memory (830) and may include, for example, an operating system (842), middleware (844), or an application (846).

[0092] The input module (850) can receive commands or data to be used in a component of the electronic device (801) (e.g., a processor (820)) from an external source (e.g., a user) of the electronic device (801). The input module (850) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0093] The audio output module (855) can output audio signals to the outside of the electronic device (801). The audio output module (855) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0094] The display module (860) can visually provide information to an external party (e.g., a user) of the electronic device (801). The display module (860) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (860) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0095] The audio module (870) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (870) can acquire sound through the input module (850), output sound through the sound output module (855), or an external electronic device (e.g., electronic device (802)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (801).

[0096] The sensor module (876) can detect the operating status (e.g., power or temperature) of the electronic device (801) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (876) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0097] The interface (877) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (801) with an external electronic device (e.g., the electronic device (802)). In one embodiment, the interface (877) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0098] The connection terminal (878) may include a connector through which the electronic device (801) may be physically connected to an external electronic device (e.g., the electronic device (802)). In one embodiment, the connection terminal (878) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0099] The haptic module (879) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (879) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0100] The camera module (880) can capture still images and videos. According to one embodiment, the camera module (880) may include one or more lenses, image sensors, image signal processors, or flashes.

[0101] The power management module (888) can manage the power supplied to the electronic device (801). According to one embodiment, the power management module (888) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0102] A battery (889) may power at least one component of the electronic device (801). In one embodiment, the battery (889) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0103] The communication module (890) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (801) and an external electronic device (e.g., electronic device (802), electronic device (804), or server (808)), and the performance of communication through the established communication channel. The communication module (890) may operate independently from the processor (820) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (890) may include a wireless communication module (892) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (894) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (804) via a first network (898) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (899) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (892) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (896) to verify or authenticate the electronic device (801) within a communication network such as the first network (898) or the second network (899).

[0104] The wireless communication module (892) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (892) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (892) may support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (892) may support various requirements specified in the electronic device (801), an external electronic device (e.g., the electronic device (804)), or a network system (e.g., the second network (899)). According to one embodiment, the wireless communication module (892) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0105] The antenna module (897) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (897) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (897) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (898) or the second network (899), may be selected from the plurality of antennas by, for example, the communication module (890). A signal or power may be transmitted or received between the communication module (890) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (897).

[0106] According to various embodiments, the antenna module (897) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0107] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0108] According to one embodiment, commands or data may be transmitted or received between the electronic device (801) and an external electronic device (804) via a server (808) connected to a second network (899). Each of the external electronic devices (802 or 804) may be the same or a different type of device as the electronic device (801). According to one embodiment, all or part of the operations executed in the electronic device (801) may be executed in one or more of the external electronic devices (802, 804, or 808). For example, when the electronic device (801) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (801) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (801). The electronic device (801) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (801) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (804) may include an Internet of Things (IoT) device. The server (808) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (804) or the server (808) may be included in the second network (899).The electronic device (801) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0109] According to one embodiment of the present disclosure, an electronic device may include a communication circuit; at least one processor including a processing circuit; and a memory including one or more storage media storing instructions. When the instructions are individually or collectively executed by the at least one processor, the instructions may cause the electronic device to virtually participate in a first space in which the electronic device is not located based on an external input. When the instructions are individually or collectively executed by the at least one processor, the instructions may cause the electronic device to request spatial audio information of the first space from a first electronic device located in the first space. When the instructions are individually or collectively executed by the at least one processor, the instructions may cause the electronic device to receive spatial audio information of the first space from the first electronic device. When the instructions are individually or collectively executed by the at least one processor, the instructions may cause the electronic device to generate a first spatial acoustic scene modeling engine based on the spatial audio information of the first space.

[0110] According to one embodiment, when the instructions are individually or collectively executed by the at least one processor, the electronic device may be caused to: generate, based on the first spatial acoustic scene modeling engine, audio of interest of the first space corresponding to a direction of interest or a sound source of interest of a user of the electronic device; and reproduce the generated audio of interest of the first space.

[0111] According to one embodiment, the direction of interest or sound source of interest of the user of the electronic device may be different from the direction of interest or sound source of interest of the user of the first electronic device.

[0112] According to one embodiment, the spatial audio information of the first space may include sound source data input to at least one microphone of the first electronic device; a position, direction, echo, reverberation, performance, movement direction, and movement speed of at least one microphone of the first electronic device; a change amount or a change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the movement direction and the movement speed; a relative distance between at least one microphone of the first electronic device; sensing data acquired from a camera or at least one sensor of the first electronic device; and at least one of a change amount or a change direction of the sensing data.

[0113] According to one embodiment, the sound source of interest or the direction of interest is determined by at least one of a head-related transfer function (HRTF), a user input of the electronic device, and a preset condition; and the preset condition may include a sound source with the highest energy, a sound source in a predetermined direction, and a sound source in the direction of a predetermined object.

[0114] According to one embodiment, when the instructions are individually or collectively executed by the at least one processor, the electronic device may be caused to: determine whether the electronic device is in an operating mode for generating the first spatial audio scene modeling engine; and, based on the determination, if the electronic device is in the operating mode, request spatial audio information of the first space from the first electronic device.

[0115] According to one embodiment, the operation mode may be an operation mode determined based on a user input of the electronic device, or an operation mode determined based on recognition of a predetermined object or a predetermined sound source.

[0116] According to one embodiment, when the instructions are individually or collectively executed by the at least one processor, the electronic device may cause: to request spatial audio information of the first space from a third electronic device added to the first space; to receive spatial audio information of the first space from the third electronic device; and to update the first spatial audio scene modeling engine based on the spatial audio information of the first space received from the third electronic device.

[0117] According to one embodiment, the third electronic device may be added or virtually added to the first space.

[0118] According to one embodiment, when the electronic device is in an operation mode that generates the first spatial audio scene modeling engine, the timestamp corresponding to the spatial audio information of the first space may be a timestamp generated based on the time at which the electronic device received the spatial audio information of the first space.

[0119] In addition, according to one embodiment of the present disclosure, a method of operating an electronic device may include: an operation of virtually participating in a first space in which the electronic device is not located based on an external input; an operation of requesting spatial audio information of the first space from a first electronic device located in the first space; an operation of receiving spatial audio information of the first space from the first electronic device; and an operation of generating a first spatial audio scene modeling engine based on the spatial audio information of the first space.

[0120] According to one embodiment, the operating method of the electronic device may further include an operation of generating audio of interest of the first space corresponding to a direction of interest or a sound source of interest of a user of the electronic device based on the first spatial acoustic scene modeling engine; and an operation of playing the generated audio of interest of the first space.

[0121] According to one embodiment, the direction of interest or sound source of interest of the user of the electronic device may be different from the direction of interest or sound source of interest of the user of the first electronic device.

[0122] According to one embodiment, the spatial audio information of the first space may include sound source data input to at least one microphone of the first electronic device; a position, direction, echo, reverberation, performance, movement direction, and movement speed of at least one microphone of the first electronic device; a change amount or a change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the movement direction and the movement speed; a relative distance between at least one microphone of the first electronic device; sensing data acquired from a camera or at least one sensor of the first electronic device; and at least one of a change amount or a change direction of the sensing data.

[0123] According to one embodiment, the sound source of interest or the direction of interest is determined by at least one of a head-related transfer function (HRTF), a user input of the electronic device, and a preset condition; and the preset condition may include a sound source with the highest energy, a sound source in a predetermined direction, and a sound source in the direction of a predetermined object.

[0124] According to one embodiment, the method of operating an electronic device may further include an operation of determining whether the electronic device is in an operating mode for generating the first spatial audio scene modeling engine. An operation of requesting spatial audio information of the first space from a first electronic device located in the first space may be performed based on the determination if the electronic device is in the operating mode.

[0125] According to one embodiment, the operation mode may be an operation mode determined based on a user input of the electronic device, or an operation mode determined based on recognition of a predetermined object or a predetermined sound source.

[0126] According to one embodiment, the operating method of the electronic device may further include: requesting spatial audio information of the first space from a third electronic device added to the first space; receiving spatial audio information of the first space from the third electronic device; and updating the first spatial audio scene modeling engine based on the spatial audio information of the first space received from the third electronic device.

[0127] According to one embodiment, the third electronic device may be added or virtually added to the first space.

[0128] According to one embodiment, when the electronic device is in an operation mode that generates the first spatial audio scene modeling engine, the timestamp corresponding to the spatial audio information of the first space may be a timestamp generated based on the time at which the electronic device received the spatial audio information of the first space.

[0129] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, display devices, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0130] The various embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. For example, a component expressed in the singular should be understood to include a concept including plural components unless the context clearly indicates only the singular. It should be understood that the term "and / or" used in this document encompasses any and all possible combinations of one or more of the listed items. The terms "comprise," "have," "consist of," and the like used in this disclosure are intended to specify only the presence of a feature, component, part, or combination thereof described in this disclosure, and the use of such terms does not exclude the presence or addition of one or more other features, components, parts, or combinations thereof. In this document, phrases such as "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can each include any one of the items listed together in that phrase, or all possible combinations thereof. Terms such as "first", "second", or "first" or "second" may be used merely to distinguish the corresponding element from other corresponding elements and do not limit the corresponding elements in any other respect (e.g., importance or order).

[0131] The terms "part" or "module" used in various embodiments of this document may include units implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. The "part" or "module" may be an integrally formed component or a minimum unit or part of the component that performs one or more functions. For example, according to one embodiment, the "part" or "module" may be implemented in the form of an application-specific integrated circuit (ASIC).

[0132] The term "if" as used in various embodiments of this document may be interpreted to mean "when", "when", "in response to determining", or "in response to detecting", depending on the context. Similarly, "if it is determined that" or "if ~ is detected" may be interpreted to mean "upon determining", "in response to determining", or "upon detecting", or "in response to detecting", depending on the context.

[0133] The program executed by the electronic device (300) described in this document may be implemented as hardware components, software components, and / or a combination of hardware components and software components. The program may be executed by any system capable of executing computer-readable instructions.

[0134] Software may include a computer program, code, instructions, or a combination of one or more of these, which can configure a processing device to perform a desired operation or command the processing device, either independently or collectively. Software may be implemented as a computer program including instructions stored on a computer-readable storage medium. Examples of the computer-readable storage medium include magnetic storage media (e.g., read-only memory (ROM), random-access memory (RAM), floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROMs, digital versatile discs (DVDs)). The computer-readable storage medium may be distributed across network-connected computer systems so that the computer-readable code can be stored and executed in a distributed manner. Computer programs can be distributed online (e.g., by download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0135] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices, communication circuit; At least one processor comprising a processing circuit; and Comprising a memory comprising one or more storage media for storing instructions; When said instructions are individually or collectively executed by said at least one processor, said electronic device causes: Virtually participate in a first space where the electronic device is not located based on external input; Request spatial audio information of the first space to a first electronic device located in the first space; Receive spatial audio information of the first space from the first electronic device; An electronic device that causes a first spatial audio scene modeling engine to be generated based on spatial audio information of the first space.

2. In paragraph 1, When said instructions are individually or collectively executed by said at least one processor, said electronic device causes: Based on the first spatial sound scene modeling engine, generating audio of interest in the first space corresponding to a direction of interest or a sound source of interest of a user of the electronic device; An electronic device that causes the audio of interest in the first space generated above to be played.

3. In paragraph 2, An electronic device wherein the direction of interest or sound source of interest of the user of the electronic device is different from the direction of interest or sound source of interest of the user of the first electronic device.

4. In any one of paragraphs 1 to 3, The spatial audio information of the first space above is An electronic device comprising: sound source data input by at least one microphone of the first electronic device; position, direction, echo, reverberation, sensitivity, performance, movement direction, and movement speed of at least one microphone of the first electronic device; a change amount or change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to a change in the movement direction and the movement speed; a relative distance between at least one microphone of the first electronic device; sensing data acquired from a camera or at least one sensor of the first electronic device; and at least one of a change amount or change direction of the sensing data.

5. In any one of paragraphs 2 to 4, The sound source of interest or the direction of interest is determined by at least one of a head-related transfer function (HRTF), a user input of the electronic device, and a preset condition; An electronic device, wherein the above preset conditions include a sound source having the highest energy, a sound source in a predetermined direction, and a sound source in a predetermined object direction.

6. In any one of paragraphs 1 to 5, When said instructions are individually or collectively executed by said at least one processor, said electronic device causes: Determining whether the electronic device is in an operating mode for generating the first spatial acoustic scene modeling engine; An electronic device that, based on the above judgment, causes the first electronic device to request spatial audio information of the first space when in the above operation mode.

7. In paragraph 6, An electronic device, wherein the above operation mode is determined based on a user input of the electronic device, or is an operation mode determined based on recognition of a predetermined object or a predetermined sound source.

8. In paragraph 1, When said instructions are individually or collectively executed by said at least one processor, said electronic device causes: Request spatial audio information of the first space to a third electronic device added to the first space; Receive spatial audio information of the first space from the third electronic device; An electronic device that causes the first spatial audio scene modeling engine to be updated based on spatial audio information of the first space received from the third electronic device.

9. In paragraph 8, The third electronic device is an electronic device that is actually located in the first space and added or virtually added.

10. In paragraph 1, An electronic device, wherein when the electronic device is in an operation mode that generates the first spatial audio scene modeling engine, a timestamp corresponding to the spatial audio information of the first space is a timestamp generated based on the time at which the electronic device receives the spatial audio information of the first space.

11. In the method of operating an electronic device, An action of virtually participating in a first space where the electronic device is not located based on external input; An action of requesting spatial audio information of the first space to a first electronic device located in the first space; An operation of receiving spatial audio information of the first space from the first electronic device; and A method comprising the operation of generating a first spatial audio scene modeling engine based on spatial audio information of the first space.

12. In paragraph 11, An operation of generating audio of interest in the first space corresponding to a direction of interest or a sound source of interest of a user of the electronic device based on the first spatial sound scene modeling engine; and A method further comprising an action of playing the audio of interest of the first space generated above.

13. In paragraph 12, A method wherein the direction of interest or sound source of interest of the user of the electronic device is different from the direction of interest or sound source of interest of the user of the first electronic device.

14. In any one of paragraphs 11 to 13, The spatial audio information of the first space above is A method comprising: sound source data input by at least one microphone of the first electronic device; position, direction, echo, reverberation, sensitivity, performance, movement direction, and movement speed of at least one microphone of the first electronic device; change amount or change direction of at least one of the position, the direction, the echo, the reverberation, and the performance due to change in the movement direction and the movement speed; a relative distance between at least one microphone of the first electronic device; sensing data acquired from a camera or at least one sensor of the first electronic device; and at least one of change amount or change direction of the sensing data.

15. In any one of paragraphs 12 to 14, The sound source of interest or the direction of interest is determined by at least one of a head-related transfer function (HRTF), a user input of the electronic device, and a preset condition; A method wherein the above preset conditions include a sound source having the highest energy, a sound source in a predetermined direction, and a sound source in a predetermined object direction.

Citation Information

Patent Citations

  • Rendering method and device for low complexity, low bit rate 6dof hoa

    JP2023060836A

  • Spatial Audio Representation and Rendering

    JP2023527022A

  • Method, device and computer program product for calculating the optimal insulin dose

    KR1020230127950A

  • Pharmaceutical Formulation

    KR102481517B1

  • A Novel Fructose C4 epimerases and Preparation Method for producing Tagatose using the same

    KR102558526B1