Electronic apparatus and controlling method thereof
By using a camera to identify objects and adjust sound characteristics, the electronic device generates stereoscopic sound, overcoming space limitations in portable devices to produce high-quality stereo or surround sound.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2021-11-23
- Publication Date
- 2026-07-21
AI Technical Summary
Portable electronic devices face challenges in generating stereoscopic sound due to limited space for high-performance microphones, limiting their ability to receive sounds with characteristics sufficient for outputting stereo or surround sound.
An electronic device uses a camera to capture images, identify objects and their locations, classify sound based on audio sources, and adjust sound characteristics to generate multiple channels of sound, enabling stereo or surround sound generation using a general microphone.
The solution effectively generates stereoscopic sound by identifying and classifying sound sources, allowing portable devices to produce high-quality stereo or surround sound without the need for multiple microphones, enhancing audio output capabilities.
Smart Images

Figure 112021135200817-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to an electronic device and a control method, and more specifically, to an electronic device and a control method for generating a plurality of channels of sound from an input sound. Background Technology
[0002] With the advancement of electronic technology, electronic devices capable of performing various functions are becoming widespread. For example, in the past, electronic devices output 4-poly or 16-poly sound, but recently, they can output stereo sound and surround sound. In addition, in the past, electronic devices output low-resolution images such as VGA or XGA resolution, but recently, they can output high-resolution images such as Full-HD and Ultra-HD.
[0003] Furthermore, with the advancement of communication technology, electronic devices can transmit and receive large amounts of data. Consequently, it has become commonplace for users to upload or download video data containing high-performance images and sound using electronic devices.
[0004] However, to output high-performance sound, electronic devices must be equipped with multiple microphones or surround microphones capable of receiving various sounds. Nevertheless, due to limitations in the size of portable electronic devices and component placement space, it is difficult to install high-performance microphones; even if multiple microphones are installed, there are limitations in receiving sounds with characteristics sufficient for outputting stereo sound.
[0005] Therefore, there is a need for technology that generates stereoscopic sound containing multiple channels using a microphone typically mounted on portable electronic devices. The problem to be solved
[0006] The present disclosure is intended to solve the aforementioned problems, and the purpose of the present disclosure is to provide an electronic device and a control method for generating stereoscopic sound based on sound input via a general microphone. means of solving the problem
[0007] According to one embodiment of the present disclosure, an electronic device includes a camera for capturing an image, a microphone for receiving one channel of sound, and a processor for generating a plurality of channels of sound based on the input sound. The processor identifies an object and the location of the object from the captured image, classifies the input sound based on an audio source, assigns it to the corresponding identified object, copies the classified sound to generate two channels of sound, adjusts the characteristics of the generated two channels of sound based on the audio source assigned to the identified object and the location of the identified object, and mixes the two channels of sound with the adjusted characteristics according to the audio source to generate two channels of stereo sound.
[0008] Alternatively, the electronic device includes a camera for capturing an image, a microphone for receiving sound input, and a processor for generating sound of multiple channels based on the input sound, wherein the processor identifies an object and the location of the object from the captured image, classifies the input sound based on an audio source and assigns it to the corresponding identified object, extracts a bass sound based on the input sound and estimates a rear sound, clusters the assigned sound based on the location of the identified object, adjusts the features of the extracted bass sound, the estimated rear sound, and the clustered sound, and generates surround sound by assigning the sound with adjusted features according to the audio source to each channel.
[0009] According to one embodiment of the present disclosure, a control method for an electronic device comprises the steps of capturing an image, receiving sound input, and generating a plurality of channels of sound based on the input sound, wherein the step of generating a plurality of channels of sound involves identifying an object and the location of the object from the captured image, classifying the input sound based on an audio source, assigning it to the corresponding identified object, copying the classified sound to generate a 2-channel sound, adjusting the characteristics of the generated 2-channel sound based on the audio source assigned to the identified object and the location of the identified object, and mixing the 2-channel sound with the adjusted characteristics according to the audio source to generate a 2-channel stereo sound.
[0010] Alternatively, a control method for an electronic device includes the steps of capturing an image, receiving sound input, and generating sound of multiple channels based on the input sound, wherein the step of generating sound of multiple channels includes identifying an object and the location of the object from the captured image, classifying the input sound based on an audio source and assigning it to the corresponding identified object, extracting a bass sound based on the input sound and estimating a rear sound, clustering the assigned sound based on the location of the identified object, adjusting the features of the extracted bass sound, the estimated rear sound, and the clustered sound, and assigning the sound with adjusted features according to the audio source to each channel to generate surround sound. Brief explanation of the drawing
[0011] FIG. 1 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure. FIG. 2 is a block diagram illustrating the specific configuration of an electronic device according to one embodiment of the present disclosure. FIGS. 3a to 3c are drawings illustrating the process of matching an object and a sound according to one embodiment of the present disclosure. FIGS. 4a to 4c are drawings illustrating the process of manually matching an object and a sound according to one embodiment of the present disclosure. FIGS. 5 and 6 are drawings illustrating a process for generating stereo sound according to one embodiment of the present disclosure. FIGS. 7a and 7b are drawings illustrating a process of clustering sound according to one embodiment of the present disclosure. FIG. 8 is a drawing illustrating the process of matching sound to a rear object according to one embodiment of the present disclosure. FIG. 9 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure. FIG. 10 is a flowchart illustrating the process of generating stereo sound according to one embodiment of the present disclosure. FIG. 11 is a flowchart illustrating the process of generating surround sound according to one embodiment of the present disclosure. Specific details for implementing the invention
[0012] Hereinafter, various embodiments are described in more detail with reference to the attached drawings. The embodiments described in this specification may be modified in various ways. Specific embodiments may be depicted in the drawings and described in detail in the detailed description. However, specific embodiments disclosed in the attached drawings are intended only to facilitate understanding of various embodiments. Accordingly, the technical concept is not limited by specific embodiments disclosed in the attached drawings, and it should be understood that it includes all equivalents or substitutions that fall within the concept and scope of the disclosure.
[0013] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but these components are not limited by the aforementioned terms. The aforementioned terms are used solely for the purpose of distinguishing one component from another.
[0014] In this specification, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. When a component is described as being “connected” or “connected” to another component, it should be understood that it may be directly connected to or connected to that other component, or that there may be other components in between. On the other hand, when a component is described as being “directly connected” or “directly connected” to another component, it should be understood that there are no other components in between.
[0015] Meanwhile, a "module" or "part" for a component as used in this specification performs at least one function or operation. Furthermore, a "module" or "part" may perform a function or operation by hardware, software, or a combination of hardware and software. Additionally, a plurality of "modules" or a plurality of "parts," excluding a "module" or "part" that must be performed on specific hardware or on at least one processor, may be integrated into at least one module. A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0016] In describing the present disclosure, the order of each step should be understood as non-limiting, unless the preceding step must logically and temporally be performed prior to the subsequent step. That is, except for such exceptional cases, the essence of the disclosure is not affected even if the process described in the subsequent step is performed prior to the process described in the preceding step, and the scope of the rights should be defined regardless of the order of the steps. Furthermore, the designation "A or B" in this specification is defined to mean not only selectively referring to either A or B, but also including both A and B. Additionally, the term "included" in this specification has a meaning that includes additional components beyond the elements listed as included.
[0017] This specification describes only the essential components necessary for explaining the present disclosure and does not mention components unrelated to the essence of the present disclosure. Furthermore, the mentioned components should not be interpreted in an exclusive sense to include only those components, but should be interpreted in a non-exclusive sense to include other components as well.
[0018] Furthermore, in describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is abbreviated or omitted. Meanwhile, each embodiment may be implemented or operated independently, but each embodiment may also be implemented or operated in combination.
[0019] FIG. 1 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present disclosure.
[0020] Referring to FIG. 1, the electronic device (100) includes a camera (110), a microphone (120), and a processor (130).
[0021] The camera (110) can generate an image by capturing the surrounding environment of the electronic device (100). For example, the image may include objects. Additionally, the image may include still images and videos, etc. As an example, one camera (110) may be placed on the rear of the electronic device (100), and multiple cameras of different types performing different functions may be placed. Alternatively, one or more cameras (110) may be placed on the front of the electronic device (100). For example, the camera (110) may include a CCD sensor and a CMOS sensor. Additionally, the camera (110) may include an RGB camera, a depth camera, a wide-angle camera, a telephoto camera, etc.
[0022] The microphone (120) receives external sound. For example, one microphone (120) may be placed in the electronic device (100), or multiple microphones (120) may be placed. For example, the microphone (120) may include a standard microphone, a surround microphone, a directional microphone, etc.
[0023] The processor (130) controls each component of the electronic device (100). For example, the processor (130) controls the camera (110) to capture an image and controls the microphone (120) to receive sound input. Additionally, the processor (130) generates sound of multiple channels based on the input sound. For example, the processor (130) can receive mono sound input and generate stereo sound. Alternatively, the processor (130) can receive mono sound or stereo sound input and generate surround sound. That is, sound of multiple channels refers to three-dimensional sound, and three-dimensional sound may include stereo sound, surround sound, etc.
[0024] The processor (130) identifies an object and the location of the object from a captured image. Then, the processor (130) classifies the input sound based on the audio source and assigns it to the corresponding object. For example, the captured image may be a video. And, the object may include a person, a car, etc. In the present disclosure, the object may be a target that generates sound. As an example embodiment, when an electronic device (100) captures a singer in a video, the electronic device (100) may capture an image of the singer and receive the vocal sound sung by the singer. The processor (130) may identify the singer, who is an object, from the image and identify the location of the singer within the image. Then, the processor (130) may separate the input sound into individual sounds. The processor (130) may classify and identify the audio source corresponding to the separated sound based on frequency characteristics. The processor (130) may identify the object and classify the sound based on an artificial intelligence model. An audio source may refer to a type of sound. For example, when an electronic device (100) receives a singer's vocal sound along with car noise sounds and people's conversation sounds, the processor (130) can separate the input sounds into individual sounds. The processor (130) can classify the sounds into car noise sounds, conversation sounds, and vocal sounds based on the audio source.
[0025] The processor (130) can assign the classified sound to a corresponding object. For example, the processor (130) can identify a singer and identify a vocal sound. Then, the processor (130) can assign the vocal sound to the singer.
[0026] Meanwhile, the electronic device (100) requires two channels of sound to generate stereo sound using the input mono sound. The processor (130) can generate two channels of sound by copying the sound to generate stereo sound based on the input mono sound. In stereo sound, the two channels may refer to a left channel sound and a right channel sound. Also, for the user to perceive stereo sound, the two channels of sound must be output with differences in intensity, time, etc. Accordingly, the processor (130) can adjust the characteristics of the two channels of sound based on the location of the audio source and the identified object. For example, the processor (130) can adjust the sound panning, time delay, phase delay, intensity adjustment, amplitude adjustment, spectral change, etc., of the two channels of sound to a preset location. The processor (130) can generate two channels of stereo sound by mixing the two channels of sound with characteristics adjusted according to the audio source.
[0027] Additionally, the electronic device (100) can generate surround sound using the input sound. As described above, the processor (130) identifies objects and the locations of the objects from the captured image. Then, the processor (130) can classify the input sound based on the audio source and assign it to the corresponding objects. The processor (130) can extract the bass sound from the input sound and estimate the rear sound. Additionally, the processor (130) can cluster the assigned sounds based on the locations of the identified objects. Clustering may mean dividing the image into specific regions and classifying sounds occurring in the same region into a single group based on the location of the objects. For example, if the processor (130) divides the image into left, center, and right regions, it can cluster the sounds into left region sounds, center region sounds, and right region sounds based on the location of the objects.
[0028] The processor (130) can generate surround sound by adjusting the features of the extracted bass sound, the estimated rear sound, and the clustered sound, and assigning the feature-adjusted sound to each channel. Each channel generating surround sound may refer to a 3.1 channel, a 5.1 channel, etc. The input sound for generating surround sound may be a sound containing multiple channels. If the input sound is a mono sound, the processor (130) may include a process of generating a left sound and a right sound by copying the rear sound or the clustered sound.
[0029] FIG. 2 is a block diagram illustrating the specific configuration of an electronic device according to one embodiment of the present disclosure.
[0030] Referring to FIG. 2, the electronic device (100) may include a camera (110), a microphone (120), a processor (130), an input interface (140), a communication interface (150), a sensor (160), a display (170), a speaker (180), and a memory (190). Since the camera (110) and the microphone (120) are the same as those described in FIG. 1, a detailed description is omitted.
[0031] The input interface (140) can receive control commands from a user. For example, the input interface (140) may include a keypad, a touchpad, a touchscreen, etc. Alternatively, the input interface (140) may include an input / output port to receive data. For example, the input interface (140) may receive video including sound and images. If the input interface (140) includes an input / output port, the input / output port may include ports such as HDMI (High-Definition Multimedia Interface), DP (DisplayPort), RGB, DVI (Digital Visual Interface), USB (Universal Serial Bus), Thunderbolt, LAN, AUX, etc. The input interface (140) may also be referred to as an input unit, an input module, etc. If the input interface (140) performs an input / output function, it may also be referred to as an input / output unit, an input / output module, etc.
[0032] The communication interface (150) can communicate with an external device. For example, the communication interface (150) can communicate with an external device using at least one of the communication methods of Wi-Fi, Wi-Fi Direct, Bluetooth, Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), and LTE (Long Term Evolution). The communication interface (150) described above may be referred to as a communication unit, a communication module, a transceiver, etc.
[0033] The sensor (160) can detect objects around the electronic device (100). The processor (130) can recognize a control command based on the detected signal and perform a control operation corresponding to the recognized control command. Additionally, the sensor (160) can detect surrounding environment information of the electronic device (100). The processor (130) can perform a corresponding control operation based on the surrounding environment information detected by the sensor (160). For example, the sensor (160) may include an accelerometer, a gravity sensor, a gyroscope, a geomagnetic sensor, a direction sensor, a motion recognition sensor, a proximity sensor, a voltmeter, an ammeter, a barometer, a hygrometer, a thermometer, an illuminance sensor, a heat detection sensor, a touch sensor, an infrared sensor, an ultrasonic sensor, etc.
[0034] The display (170) can output data processed by the processor (130) as an image. The display (170) can display a captured image and can display a marker indicating a separated sound in the form of text or an image. For example, the display (170) can be implemented as an LCD (Liquid Crystal Display), OLED (Organic Light Emitting Diode), flexible display, touch screen, etc. If the display (170) is implemented as a touch screen, the electronic device (100) can receive control commands through the touch screen.
[0035] The speaker (180) outputs a voice signal after voice processing has been performed. For example, multiple speakers (180) may be placed in the electronic device (100), and the processor (130) may output stereoscopic sound by assigning sound to each channel based on the location of the placed speakers (180). Additionally, the speaker (180) may output information regarding user input commands, information related to the status of the electronic device (100), or information related to operation, etc., as voice or notification sounds.
[0036] The memory (190) can store data, algorithms, etc. that perform the functions of the electronic device (100), and can store programs, instructions, etc. that run on the electronic device (100). For example, the memory (190) can store an image processing artificial intelligence algorithm and a sound processing artificial intelligence algorithm. The processor (130) can identify objects from a captured image using the image processing artificial intelligence algorithm. In addition, the processor (130) can process input sound and generate stereoscopic sound using the sound processing artificial intelligence algorithm. The algorithms stored in the memory (190) can be loaded into the processor (130) under the control of the processor (130) to perform an object identification process or a sound processing process. For example, the memory (190) can be implemented in the form of ROM, RAM, HDD, SSD, memory card, etc.
[0037] So far, the configuration of the electronic device (100) has been described. Below, the process of matching the object included in the image with the sound is described.
[0038] FIGS. 3a to 3c are drawings illustrating the process of matching an object and a sound according to one embodiment of the present disclosure.
[0039] Referring to Fig. 3a, an image of a concert scene is shown.
[0040] The electronic device (100) can film a concert scene as a video. As an example, the image may include a cello player (11), a guitar player (12), and a singer (13). The cello player (11), the guitar player (12), and the singer (13) may refer to objects included in the image. The electronic device (100) can identify objects from the filmed image. For example, the electronic device (100) may include an image processing artificial intelligence algorithm. The electronic device (100) can identify objects from the filmed image using an image processing artificial intelligence algorithm.
[0041] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0042] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0043] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of operations of the previous layer and the multiple weights. The multiple weights possessed by the multiple neural network layers may be optimized by the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), You Only Look Once (YOLO), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0044] As described above, the electronic device (100) can identify objects of a cello player (11), a guitar player (12), and a singer (13) from a captured image using an image processing artificial intelligence algorithm.
[0045] FIG. 3b illustrates a process of separating input sounds. An electronic device (100) may receive ambient mixed sounds (30) while filming a concert. The received mixed sounds (30) may include cello sounds (31), guitar sounds (32), and vocal sounds (33). For example, the electronic device (100) may separate each sound from the received sounds using a sound processing artificial intelligence algorithm. As an example, the sound processing artificial intelligence algorithm may include an Independent Component Analysis (ICA) model. An ICA model can decompose a multivariate signal into independent additional sub-components using a cocktail effect. For example, an ICA model can decompose a mixed sound signal of length T seconds composed of three sources into three sound signals of length T seconds.
[0046] The electronic device (100) can separate the mixed sound signal, classify the audio source corresponding to each sound based on frequency characteristics, and identify it using sound source information. The electronic device (100) can classify and identify the audio source using a sound processing artificial intelligence algorithm. As illustrated in FIG. 3b, the electronic device (100) can first separate the input mixed sound (30) into sound a, sound b, and sound c. Then, the electronic device (100) can classify and identify the separated sound a as a cello sound (31), sound b as a guitar sound (32), and sound c as a vocal sound. In one embodiment, the electronic device (100) can store sound source information and identify the separated sound based on the stored sound source information. Alternatively, the electronic device (100) can transmit the separated sound to an external device containing sound source information. The external device may identify the audio source and transmit the identified sound source information to the electronic device (100).
[0047] Figure 3c shows a diagram in which an object and a sound are matched.
[0048] As described above, the electronic device (100) can identify objects. The electronic device (100) can separate mixed sounds and identify each separated sound. The electronic device (100) can assign (or match) each identified sound to a corresponding object. For example, a cello sound (31) can be assigned to a cello player object (11), a guitar sound (32) can be assigned to a guitar player object (12), and a vocal sound (33) can be assigned to a singer object (13). In one embodiment, the electronic device (100) can display the identified object and the corresponding sound. The electronic device (100) can display a mark indicating the sound corresponding to the object along with the identified object.
[0049] Meanwhile, the electronic device (100) may not be able to identify the separated sound.
[0050] FIGS. 4a to 4c are drawings illustrating the process of manually matching an object and a sound according to one embodiment of the present disclosure.
[0051] Referring to FIG. 4a, a diagram is shown in which an object is identified and a sound corresponding to the object is assigned. For example, an electronic device (100) can identify a cello player object (21) from a captured image and separate four sounds from an input mixed sound. However, the electronic device (100) may not be able to identify one of the separated sounds (41). If the electronic device (100) cannot identify the separated sound, it may display a pre-set indicator on the classified sound. In one embodiment, the pre-set indicator may be text such as "unknown sound". That is, if the electronic device (100) cannot identify an object corresponding to the classified sound, it may display the sound (41) with the pre-set indicator displayed along with the identified object.
[0052] Referring to FIG. 4b, the process of manually matching an object and a sound by a user is illustrated. The electronic device (100) may receive a command from the user (1) to move the marker of the sound (41) that has a preset indicator displayed. For example, the command received from the user (1) may be a drag-and-drop command, but is not limited thereto. The electronic device (100) may move the marker of the sound (41) that has a preset indicator displayed to a cello player object (21) to which no sound is assigned, according to the user's command. The electronic device (100) may assign the sound (41) that has a preset indicator displayed to the cello player object (21) according to the user's command.
[0053] As illustrated in FIG. 4c, through the process described above, the electronic device (100) can match an unidentified sound (41) to one object (21) and match the identified object and the classified sound in a 1:1 ratio.
[0054] FIGS. 5 and 6 are drawings illustrating a process for generating stereo sound according to one embodiment of the present disclosure.
[0055] Referring to FIG. 5, the process of copying an input sound (51) into multiple channels is illustrated. As described above, to generate stereo sound, sound from multiple channels (e.g., left channel, right channel) is required. However, if the electronic device (100) includes a single microphone, the input sound is mono sound. Therefore, the electronic device (100) can generate two channels of sound, a left sound (51a) and a right sound (51b), by copying the mono sound to generate stereo sound. The electronic device (100) can classify each individual sound separated by the method described above according to the audio source and assign it to an identified object. Additionally, the electronic device (100) can generate two channels of sound by copying the classified mono sound.
[0056] Referring to FIG. 6, an example of identifying the location of an object is illustrated. To generate stereo sound, the sound characteristics of each channel must be adjusted according to the distance of the identified object. For example, the volume of the sound corresponding to an object located close to the user must be large, and the volume of the sound corresponding to an object located far from the user must be small. Additionally, the sound of the sound corresponding to an object located to the left of the user must be formed in the left area of the user, and the sound of the sound corresponding to an object located to the right of the user must be formed in the right area of the user.
[0057] For example, the location of an object can be obtained using a triangulation, a LiDAR sensor, or a ToF sensor. If the electronic device (100) includes a sensor, the location of the object can be obtained based on a signal detected by the sensor. Alternatively, as shown in FIG. 6, the electronic device (100) can identify the location of the object using a triangulation. For example, the electronic device (100) can obtain distances D1, D2, and D3 from each object (22, 23, 24). D1, D2, and D3 may be absolute distances or relative distances. Additionally, the electronic device (100) can obtain separation distances X1, X2, and X3 from the left speaker to each object (22, 23, 24). X1, X2, and X3 may be relative distances. The location of the object (22, 23, 24) from the left speaker can be obtained using the Pythagorean theorem. The positions of the objects (22, 23, 24) from the right speaker can also be obtained in a similar manner.
[0058] The electronic device (100) can adjust the characteristics of the two channels of sound assigned to the object based on the location of the acquired object (22, 23, 24). For example, the electronic device (100) can adjust the characteristics of the sound in a manner such as sound panning, time delay, phase delay, intensity adjustment, amplitude adjustment, and spectral change to a preset position. In one embodiment, the sound corresponding to the guitar player's object (24) can form a sound image in the left area by delaying the right channel, delaying the phase, or weakening the intensity or amplitude. Alternatively, the electronic device (100) can form a sound image in the left area by adjusting the left channel in the opposite way to the method of adjusting the characteristics of the right channel described above.
[0059] So far, the process of generating stereo sound using the input mono sound of the electronic device (100) has been explained. Below, the process of generating surround sound is explained.
[0060] FIGS. 7a and 7b are drawings illustrating a process of clustering sound according to one embodiment of the present disclosure.
[0061] Referring to FIG. 7a, the electronic device (100) can receive mixed sounds from the surrounding area of the shooting location while shooting a video including a waterfall (61) and a bird (62). For example, the mixed sounds may include a waterfall sound (71), a bird sound (72), ambient noise sound (73), etc.
[0062] The electronic device (100) can identify objects and the locations of objects from images captured in the same manner as described above. The electronic device (100) can classify input sounds based on audio sources and assign them to the corresponding identified objects. As an example, as shown in FIG. 7a, the electronic device (100) can identify a waterfall object (61) and a bird object (61), assign a waterfall sound (71) corresponding to the identified waterfall object (61), and assign a bird sound (72) corresponding to the identified bird object (62). Furthermore, the electronic device (100) can divide the image into pre-set regions to create a surround channel, and cluster sounds assigned to objects included in the same region among each divided region into the same group.
[0063] Referring to FIG. 7b, an example of clustering sounds by region is illustrated. For example, if an electronic device (100) generates three channels of surround sound, the sound can be divided into a left channel, a center channel, and a right channel. Thus, the electronic device (100) can divide the image into a left region (3), a center region (5), and a right region (7), and identify which of the regions the classified sound belongs to. For example, the electronic device (100) can identify that the waterfall sound (71) belongs to the center region (5), and the bird sound (72) and noise sound (73) belong to the right region (7). Thus, the electronic device (100) can cluster the waterfall sound (71) into a group in the center region (5), and cluster the bird sound (72) and noise sound (73) into a group in the right region (7). If the electronic device (100) generates 5 channels of surround sound, the image can be divided into more detailed areas and the classified sound can be included in each area.
[0064] Meanwhile, if the electronic device (100) generates 3.1 channel or 5.1 channel surround sound, bass sound can be extracted. For example, the electronic device (100) can extract bass sound by low-pass filtering the input mixed sound. The surround sound may include sound generated from rear objects in addition to sound generated from front objects captured by the camera.
[0065] FIG. 8 is a drawing illustrating the process of matching sound to a rear object according to one embodiment of the present disclosure.
[0066] Referring to FIG. 8, a guitar player object, a singer object, and a drum player object may be located in front of the electronic device (100), and a car object (81) and a conversationalist object (82) may be located in the rear. If the electronic device (100) includes a surround camera, the objects located in the rear may also be photographed, so the objects located in the rear can be identified and classified sounds can be matched. Alternatively, if the electronic device (100) includes a camera placed in the front and a camera placed in the rear, the objects located in the rear can also be photographed using the camera placed in the rear, so the objects located in the rear can be identified and classified sounds can be matched.
[0067] However, if the electronic device (100) includes only a camera positioned at the front, it cannot capture a car object (81) and a conversational object (82) located at the rear. However, the mixed sound input to the electronic device (100) may include a car sound (91) and a conversational sound (92). Therefore, the electronic device (100) can estimate a sound other than the sound assigned to the object identified in the image as a rear sound.
[0068] Alternatively, the electronic device (100) may manually estimate the rear sound by the user. For example, the electronic device (100) may separate sound a and sound b from the input mixed sound. Then, the electronic device (100) may identify the separated sounds based on frequency characteristics and sound source information. However, since the electronic device (100) does not find an object corresponding to the identified sound, it may display a pre-set indicator on the identified sound. As an example, the electronic device (100) may display an indicator such as "unknown car sound" on the car sound (91) and "unknown conversation sound" on the conversation sound (92). The electronic device (100) may receive a command from the user (1) to move the indicator of the car sound (91) on which the indicator is displayed. The electronic device (100) may move the indicator of the car sound (91) on which the pre-set indicator is displayed to a pre-set area of the screen according to the user's command. The electronic device (100) can estimate that the sound corresponds to an object located at the rear when the mark of the car sound (91) moves to a preset area. In one embodiment, the electronic device (100) can estimate that the sound corresponds to a left rear object when the mark of the car sound (91) moves to a preset left area, and can estimate that the sound corresponds to a right rear object when the mark of the conversation sound (92) moves to a preset right area.
[0069] The electronic device (100) can generate surround sound by adjusting the features of the extracted bass, estimated rear sound, and clustered sound and assigning them to each channel. The sound feature adjustment process for generating surround sound may be the same as the sound feature adjustment process for generating stereo sound.
[0070] So far, various embodiments of electronic devices generating stereoscopic sound have been described. Below, a method for controlling the electronic device is described.
[0071] FIG. 9 is a flowchart illustrating a method for controlling an electronic device according to an embodiment of the present disclosure, FIG. 10 is a flowchart illustrating a process for generating stereo sound according to an embodiment of the present disclosure, and FIG. 11 is a flowchart illustrating a process for generating surround sound according to an embodiment of the present disclosure. The following description will be made with reference to FIGS. 9 to 11 together.
[0072] The electronic device captures an image and receives sound input (S910), and generates sound of multiple channels based on the received sound (S910). For example, the electronic device can receive mono sound of one channel and generate stereo sound.
[0073] The electronic device can identify objects and the locations of objects from a captured image (S1010). The electronic device can identify objects and the locations of objects based on an image processing artificial intelligence model.
[0074] The electronic device can classify the input sound based on the audio source and assign it to the corresponding identified object (S1020). For example, the electronic device may receive a mixed sound in which various sounds are mixed. The electronic device can separate the input sound into individual sounds. The electronic device can identify the audio source corresponding to each separated sound based on frequency characteristics. The electronic device can classify each sound based on the identified audio source.
[0075] Meanwhile, if the electronic device fails to identify an object corresponding to a classified sound, it may display a mark of the classified sound along with the identified object and a preset indicator. The electronic device may match the classified sound with the preset indicator displayed to the identified object according to the input user's command.
[0076] The electronic device can generate a 2-channel sound by copying the classified sound (S1030). For example, the 2-channel sound may be a left channel sound and a right channel sound. The electronic device can adjust the characteristics of the generated 2-channel sound based on the audio source assigned to the identified object and the location of the identified object (S1040). For example, the electronic device can adjust the characteristics of the 2-channel sound by applying methods such as sound panning to a preset position, time delay, phase delay, intensity adjustment, amplitude adjustment, and spectral change. The electronic device can adjust the characteristics of the 2-channel sound based on a sound processing artificial intelligence model.
[0077] The electronic device can generate 2-channel stereo sound by mixing 2-channel sound with features adjusted according to an audio source (S1050). The generated 2-channel stereo sound can be stored in memory and output to a speaker. Alternatively, the electronic device can transmit the generated 2-channel stereo sound along with an image to an external device.
[0078] Alternatively, the electronic device may receive 1-channel mono sound or stereo sound and generate surround sound.
[0079] The electronic device can identify objects and their locations from captured images (S1110). Additionally, the electronic device can classify input sounds based on audio sources and assign them to corresponding identified objects (S1120). Since the process of identifying objects and classifying sounds to assign them to corresponding objects is identical to the process described above, a detailed explanation is omitted.
[0080] The electronic device can extract a bass sound and estimate a rear sound based on the input sound (S1130). For example, the electronic device can extract a bass sound by low-pass filtering the input sound. Additionally, the electronic device can estimate a sound other than the sound assigned to an object identified in the image as a rear sound. Alternatively, the electronic device can display an indicator on a separate sound that does not match the object. When the sound with the indicator moves to a preset area on the screen according to a user command, the electronic device can estimate the separate sound that does not match the object as a rear sound.
[0081] The electronic device can cluster assigned sounds based on the location of identified objects (S1140). For example, the electronic device can divide an image into multiple regions based on the number of channels of the surround sound to be generated. Then, the electronic device can cluster sounds assigned to objects included in the same region among each divided region into the same group.
[0082] The electronic device can adjust the characteristics of the extracted bass sound, the estimated rear sound, and the clustered sound (S1150). For example, the electronic device can adjust the characteristics of the sound of each channel by applying methods such as sound panning to a preset position, time delay, phase delay, intensity adjustment, amplitude adjustment, and spectral change. The electronic device can adjust the characteristics of the sound of each channel based on a sound processing artificial intelligence model.
[0083] The electronic device can generate surround sound by assigning sounds with features adjusted according to the audio source to each channel (S1160). The electronic device can store, output, or transmit the generated surround sound to an external device.
[0084] A control method for an electronic device according to the various embodiments described above may be provided as a computer program product. The computer program product may include the S / W program itself or a non-transitory computer readable medium on which the S / W program is stored.
[0085] A non-transient readable medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, the various applications or programs described above may be stored and provided on non-transient readable media such as CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.
[0086] Furthermore, although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure. Explanation of the symbols
[0087] 100: Electronic device 110: Camera 120: Microphone 130: Processor
Claims
Claim 1 An electronic device comprising: a camera for capturing an image; a display; an input interface; a microphone for receiving sound of one channel; and a processor for generating sound of multiple channels based on the input sound; wherein the processor identifies an object and the location of the object from the captured image, classifies the input sound based on an audio source and assigns it to the corresponding identified object, copies the classified sound to generate sound of two channels, adjusts the characteristics of the generated sound of two channels based on the audio source assigned to the identified object and the location of the identified object, mixes the sound of two channels with adjusted characteristics according to the audio source to generate stereo sound of two channels, and, if the processor fails to identify an object corresponding to the classified sound, controls the display to display an indicator pre-set on the classified sound and the identified object, and matches the classified sound with the pre-set indicator displayed to the identified object according to a user command input through the input interface. Claim 2 An electronic device according to claim 1, wherein the processor adjusts the characteristics of the sound of the two channels based on at least one of sound panning to a preset position, time delay, phase delay, intensity adjustment, amplitude adjustment, and spectral change. Claim 3 An electronic device according to claim 1, wherein the processor separates the input sound into individual sounds, identifies an audio source corresponding to each of the separated sounds based on frequency characteristics, and classifies each of the sounds based on the identified audio source. Claim 4 delete Claim 5 delete Claim 6 An electronic device according to claim 1, wherein the processor identifies the object and the location of the object based on an image processing artificial intelligence model, and adjusts the characteristics of the generated 2-channel sound based on a sound processing artificial intelligence model. Claim 7 An electronic device comprising: a camera for capturing an image; a display; an input interface; a microphone for receiving sound input; and a processor for generating sound of multiple channels based on the input sound, wherein the processor identifies an object and the location of the object from the captured image, classifies the input sound based on an audio source and assigns it to the corresponding identified object, extracts a bass sound and estimates a rear sound based on the input sound, clusters the assigned sound based on the location of the identified object, adjusts the features of the extracted bass sound, the estimated rear sound, and the clustered sound, and generates surround sound by assigning the sound with adjusted features according to the audio source to each channel, wherein if the processor fails to identify an object corresponding to the classified sound, the processor controls the display to display a pre-set indicator on the classified sound and the identified object, and matches the classified sound with the pre-set indicator displayed to the identified object according to a user command input through the input interface. Claim 8 In claim 7, the processor is an electronic device that estimates a sound other than the sound assigned to an object identified in the image as the rear sound. Claim 9 delete Claim 10 delete Claim 11 A method for controlling an electronic device comprising: a step of capturing an image and receiving sound input; and a step of generating a plurality of channels of sound based on the input sound; wherein the step of generating a plurality of channels of sound includes identifying an object and the location of the object from the captured image, classifying the input sound based on an audio source and assigning it to the corresponding identified object, copying the classified sound to generate a 2-channel sound, adjusting the characteristics of the generated 2-channel sound based on the audio source assigned to the identified object and the location of the identified object, and mixing the 2-channel sound with adjusted characteristics according to the audio source to generate a 2-channel stereo sound; and wherein, if the object corresponding to the classified sound cannot be identified, a pre-set indicator on the classified sound and the identified object are displayed, and the classified sound with the pre-set indicator displayed is matched to the identified object according to a command input by a user. Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 A method for controlling an electronic device, comprising: a step of capturing an image and receiving sound input; and a step of generating sound of multiple channels based on the input sound; wherein the step of generating sound of multiple channels includes identifying an object and the location of the object from the captured image, classifying the input sound based on an audio source and assigning it to the corresponding identified object, extracting a bass sound and estimating a rear sound based on the input sound, clustering the assigned sound based on the location of the identified object, adjusting the features of the extracted bass sound, the estimated rear sound, and the clustered sound, and generating surround sound by assigning the sound with adjusted features according to the audio source to each channel; and wherein, if the object corresponding to the classified sound cannot be identified, the method includes displaying a pre-set indicator on the classified sound and the identified object, and matching the classified sound with the pre-set indicator displayed to the identified object according to a command from an input user. Claim 18 delete Claim 19 delete Claim 20 delete