Video processing apparatus and method
Patent Information
- Application Number
- CN202180066099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-19
- Filing Date
- 2021-09-28
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2041-09-28
AI Technical Summary
[0004]因为普通的音频信号获取设备(例如,麦克风)只能获取二维音频信号,所以可以从二维音频信号中获得单独的声源,并且考虑到声源的运动,通过混合和监控来生成三维音频信号,但是这是非常困难且耗时的任务
Smart Images

Figure CN116210233B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video processing, and more specifically, to the field of generating three-dimensional audio signals. More specifically, this disclosure relates to the field of generating three-dimensional audio signals, including multiple channels, from two-dimensional audio signals based on artificial intelligence (AI). Background Technology
[0002] Audio signals are typically two-dimensional audio signals, such as 2-channel, 5.1-channel, 7.1-channel, and 9.1-channel audio signals.
[0003] However, because two-dimensional audio signals have uncertain audio information or no audio information (audio information of the height component) in the height direction, it is necessary to generate three-dimensional information (n-channel audio signal or multi-channel audio signal, where n is an integer greater than 2) to provide a spatial stereo effect of sound.
[0004] Because ordinary audio signal acquisition devices (e.g., microphones) can only acquire two-dimensional audio signals, it is possible to obtain individual sound sources from two-dimensional audio signals and generate three-dimensional audio signals by mixing and monitoring, taking into account the movement of the sound sources. However, this is a very difficult and time-consuming task.
[0005] Therefore, there is a need for a method to generate a three-dimensional audio signal by using a video signal corresponding to a two-dimensional audio signal together with the two-dimensional audio signal. Summary of the Invention
[0006] Technical issues
[0007] This disclosure provides a way to more easily generate three-dimensional audio signals by using two-dimensional audio signals and video information corresponding to the two-dimensional audio signals.
[0008] Technical solution
[0009] According to embodiments of the present disclosure, a video processing apparatus includes: a memory storing one or more instructions; and at least one processor configured to execute one or more instructions stored in the memory, wherein the at least one processor is configured to: analyze a video signal comprising multiple images based on a first deep neural network (DNN), generate multiple feature information for each time and frequency, and extract a first height component and a first planar component corresponding to motion of an object in the video from the video signal based on a second DNN, extract a second planar component corresponding to motion of a sound source in the audio from a first audio signal that does not have a height component using a third DNN, generate a second height component from the first height component, the first planar component and the second planar component, output a second audio signal including the second height component based on the feature information, and synchronize the second audio signal with the video signal and output the signal.
[0010] At least one processor may be configured to synchronize a video signal with a first audio signal when generating multiple feature information for each time and frequency, generate M one-dimensional image feature mapping information (where M is an integer greater than or equal to 1) based on the motion of an object in the video from the video signal by using a first DNN for generating multiple feature information for each time and frequency, and generate multiple feature information for each time and frequency including the M image feature mapping information for time and frequency by performing frequency-related tiling on the one-dimensional image feature mapping information.
[0011] When extracting the first height component and the first planar component based on the second DNN, and extracting the second planar component based on the third DNN, at least one processor may be configured to: synchronize the video signal with the first audio signal; extract N+M feature mapping information (where N and M are integers greater than or equal to 1) corresponding to the horizontal motion in the video relative to time in the video by using a (2-1)th DNN for extracting feature mapping information corresponding to the horizontal motion in the video corresponding to the first height component; and extract N+M feature mapping information (where N and M are integers greater than or equal to 1) corresponding to the vertical motion in the video relative to time in the video by using a (2-2)th DNN for extracting feature mapping information corresponding to the vertical motion in the video corresponding to the first planar component. ; and by using a third DNN for extracting feature mapping information corresponding to horizontal motion in the audio corresponding to the second planar component, extracting N+M feature mapping information corresponding to horizontal motion in the audio from the first audio signal, and when generating the second height component from the first height component, the first planar component, and the second planar component, at least one processor is configured to: generate N+M correction mapping information relative to time corresponding to the second height component based on the feature mapping information corresponding to horizontal motion in the video corresponding to vertical motion in the video and the feature mapping information corresponding to horizontal motion in the audio; and generate N+M correction mapping information relative to time and frequency corresponding to the second height component by performing frequency-related tiling on the N+M correction mapping information relative to time.
[0012] When a second audio signal including a second height component is output based on feature information, at least one processor can be configured to: generate time and frequency information about the two channels by performing a frequency conversion operation on the first audio signal; generate N audio feature mapping information about time and frequency (where N is an integer greater than or equal to 1) from the time and frequency information about the two channels by using a (4-1)DNN for generating audio features in the first audio signal; generate N+M audio / image integrated feature mapping information based on M image feature mapping information about time and frequency included in multiple feature information at each time and frequency, and N audio feature mapping information about time and frequency (where N is an integer greater than or equal to 1); and generate frequency domain feature mapping information by using a (4-1)DNN for generating frequency domain feature mapping information. The (4-2)th DNN of the second audio signal generates a frequency domain second audio signal for the n-channel from N+M audio / image integrated feature mapping information (where n is an integer greater than 2); by using the (4-3)th DNN for generating audio correction mapping information, audio correction mapping information for the n-channel is generated from N+M correction mapping information corresponding to the N+M audio / image integrated feature mapping information and the second height component in terms of time and frequency; a corrected frequency domain second audio signal for the n-channel is generated by performing correction on the frequency domain second audio signal for the n-channel based on the audio correction mapping information for the n-channel; and a second audio signal for the n-channel is output by performing inverse frequency conversion on the corrected frequency domain second audio signal for the n-channel.
[0013] When generating N+M correction mapping information with respect to time, at least one processor can be configured to: generate a fourth value of the N+M correction mapping information with respect to time based on a ratio set by considering the relationship between a first value of feature mapping information corresponding to horizontal motion in the video, a second value of feature mapping information corresponding to horizontal motion in the audio, and a third value of feature mapping information corresponding to vertical motion in the video; and generate N+M correction mapping information with respect to time including the fourth value.
[0014] The second audio signal is output based on the fourth DNN used to output the second audio signal. The first DNN used to generate multiple feature information for each time and frequency, the second DNN used to extract the first height component and the first planar component, the third DNN used to extract the second planar component, and the fourth DNN used to output the second audio signal are trained based on the comparison results of the first trained two-dimensional audio signal and the first frequency domain trained three-dimensional audio signal reconstructed based on the first trained two-dimensional audio signal and the first frequency domain trained three-dimensional audio signal obtained by frequency conversion of the first trained three-dimensional audio signal.
[0015] The second audio signal is output based on the fourth DNN used to output the second audio signal, and the N+M correction mapping information about time and frequency can be corrected based on user input parameter information. Based on the comparison results of the first training two-dimensional audio signal, the first training image signal and the frequency domain training reconstructed three-dimensional audio signal reconstructed based on user input parameter information and the first frequency domain training three-dimensional audio signal obtained by frequency conversion of the first training three-dimensional audio signal, the first DNN used to generate multiple feature information for each time and frequency, the second DNN used to extract the first height component and the first planar component, the third DNN used to extract the second planar component and the fourth DNN used to output the second audio signal can be trained.
[0016] The first training two-dimensional audio signal and the first training image signal are obtained from a portable terminal that is the same device as the video processing device or a different device connected to the video processing device, and the first training three-dimensional audio signal can be obtained from a surround sound microphone included in or installed on the portable terminal.
[0017] The parameter information of the first to fourth DNNs, obtained as the training results of the first, second, third, and fourth DNNs, can be stored in the video processing device or received from a terminal connected to the video processing device.
[0018] A video processing method of a video processing apparatus according to an embodiment of the present disclosure includes: analyzing a video signal comprising multiple images based on a first deep neural network (DNN) to generate multiple feature information for each time and frequency; extracting a first height component and a first planar component corresponding to the motion of an object in the video signal from the video signal based on a second DNN; extracting a second planar component corresponding to the motion of a sound source in the audio signal from a first audio signal that does not have a height component based on a third DNN; generating a second height component from the first height component, the first planar component, and the second planar component; outputting a second audio signal including the second height component based on the feature information; and synchronizing the second audio signal with the video signal and outputting the signal.
[0019] Generating multiple feature information for each time and frequency includes: synchronizing the video signal with a first audio signal; using a first DNN for generating multiple feature information for each time and frequency; generating M one-dimensional image feature mapping information (where M is an integer greater than or equal to 1) based on the motion of objects in the video from the video signal; and generating multiple feature information for each time and frequency, including M image feature mapping information with respect to time and frequency, by performing frequency-related tiling on the one-dimensional image feature mapping information.
[0020] Extracting the first height component and the first planar component based on the second DNN and extracting the second planar component based on the third DNN may include: synchronizing the video signal with the first audio signal; extracting N+M feature mapping information (where N and M are integers greater than or equal to 1) corresponding to the horizontal motion in the video relative to time in the video by using the (2-1) DNN for extracting feature mapping information corresponding to the horizontal motion in the video corresponding to the first height component; extracting N+M feature mapping information (where N and M are integers greater than or equal to 1) corresponding to the vertical motion in the video relative to time in the video by using the (2-2) DNN for extracting feature mapping information corresponding to the vertical motion in the video corresponding to the first planar component; and extracting N+M feature mapping information (where N and M are integers greater than or equal to 1) corresponding to the horizontal motion in the video relative to time in the video by using the (2-2) DNN for extracting feature mapping information corresponding to the horizontal motion in the video. The third DNN, which provides feature mapping information corresponding to horizontal motion in the audio of the second plane component, extracts N+M feature mapping information (where N and M are integers greater than or equal to 1) corresponding to horizontal motion in the audio from the first audio signal. Generating the second height component from the first height component, the first plane component, and the second plane component includes: generating N+M time-dependent correction mapping information corresponding to the second height component based on feature mapping information corresponding to horizontal motion in the video, feature mapping information corresponding to vertical motion, and feature mapping information corresponding to horizontal motion in the audio; and generating N+M time- and frequency-dependent correction mapping information corresponding to the second height component by performing frequency-dependent tiling on the time-dependent N+M correction mapping information.
[0021] The output of a second audio signal including a second height component based on feature information may include: obtaining time and frequency information about the two channels by performing a frequency conversion operation on the first audio signal; generating N audio feature mapping information (where N is an integer greater than or equal to 1) relative to time and frequency from the time and frequency information about the two channels by using a (4-1)DNN for generating audio features in the first audio signal; generating N+M audio / image integrated feature mapping information based on M image feature mapping information about time and frequency and N audio feature mapping information about time and frequency (where N is an integer greater than or equal to 1) from multiple feature information included in each time and frequency; and using a frequency conversion operation to generate audio features in the first audio signal. The (4-2)th DNN of the second audio signal in the domain obtains the frequency domain second audio signal for the n-channel (where n is an integer greater than 2) from N+M audio / image integrated feature mapping information; by using the (4-3)th DNN for generating audio correction mapping information, audio correction mapping information for the n-channel corresponding to the second height component is generated from the N+M audio / image integrated feature mapping information; by performing correction on the frequency domain second audio signal for the n-channel based on the N+M audio / image correction mapping information for the n-channel, a corrected frequency domain second audio signal for the n-channel is generated; and by performing inverse frequency conversion on the corrected frequency domain second audio signal for the n-channel, a second audio signal for the n-channel is output.
[0022] The second audio signal is output based on the fourth DNN used to output the second audio signal. The first DNN used to generate multiple feature information for each time and frequency, the second DNN used to extract the first height component and the first planar component, the third DNN used to extract the second planar component, and the fourth DNN used to output the second audio signal are trained based on the comparison results of the first trained two-dimensional audio signal and the first frequency domain trained three-dimensional audio signal reconstructed based on the first trained two-dimensional audio signal and the first frequency domain trained three-dimensional audio signal obtained by frequency conversion of the first trained three-dimensional audio signal.
[0023] A computer-readable recording medium according to an embodiment of the present disclosure has a program thereon recorded for performing the method. Attached Figure Description
[0024] To provide a more complete understanding of the accompanying figures cited in this article, a brief description of each figure is provided.
[0025] Figure 1 This is a block diagram illustrating the configuration of a video processing apparatus according to an embodiment.
[0026] Figure 2 This is a diagram illustrating the detailed operation of the video feature information generation unit 110 according to an embodiment.
[0027] Figure 3 This is a diagram used to describe a first deep neural network (DNN) 300 according to an embodiment.
[0028] Figure 4 This is a diagram used to describe a detailed description of the correction information generation unit 120 according to an embodiment.
[0029] Figures 5a to 5b This is a graph used to describe the theoretical background, from which the domain matching parameter α is derived. inf [Mathematical Formula 1].
[0030] Figure 5c This is a diagram describing an algorithm for estimating the height component of a sound source within an audio signal, which is necessary for generating a three-dimensional audio signal by analyzing the motion of objects within a video signal and the motion of sound sources within a two-dimensional audio signal.
[0031] Figure 6a This is a diagram used to describe the (2-1)th DNN 600.
[0032] Figure 6b This is a diagram used to describe the (2-2)th DNN 650.
[0033] Figure 7 This is a diagram used to describe the third DNN 700.
[0034] Figure 8 This is a diagram illustrating the detailed operation of the three-dimensional audio output unit 130 according to an embodiment.
[0035] Figure 9 This is a diagram used to describe the (4-1) DNN 900 according to the embodiment.
[0036] Figure 10 This is a diagram used to describe the (4-2) DNN 1000 according to the embodiment.
[0037] Figure 11 This is a diagram used to describe the (4-3) DNN 1100 according to the embodiment.
[0038] Figure 12 This is a diagram used to describe the methods for training the first DNN, second DNN, third DNN, and fourth DNN.
[0039] Figure 13 This is a diagram used to describe the methods for training the first DNN, second DNN, third DNN, and fourth DNN by taking into account user parameter signals.
[0040] Figure 14It is a flowchart describing the process of training the first DNN, the second DNN, the third DNN, and the fourth DNN by the training device 1400.
[0041] Figure 15 It is a flowchart describing the process of training the first DNN, the second DNN, the third DNN, and the fourth DNN by taking user parameters into account by the training device 1500.
[0042] Figure 16 This is a diagram used to describe the process by which a user collects data for training using a user terminal 1610.
[0043] Figure 17 This is a flowchart describing a video processing method according to an embodiment. Detailed Implementation
[0044] According to an embodiment, a three-dimensional audio signal can be generated by using a two-dimensional audio signal and its corresponding video signal.
[0045] However, the effects that can be achieved by the apparatus and method for processing video according to the embodiments are not limited to those mentioned above, and other effects not mentioned can be clearly understood by those skilled in the art from the following description.
[0046] Because this disclosure allows for various variations and numerous embodiments, certain embodiments will be shown in the accompanying drawings and described in detail in the written description. However, this is not intended to limit this disclosure to specific implementations, but rather to be understood to include all variations, equivalents, and alternatives that fall within the concepts and technical scope of this disclosure.
[0047] When describing embodiments, detailed descriptions of relevant known technologies are omitted when they obscure key points. Furthermore, numbers used in describing embodiments (e.g., first, second, etc.) are merely identification symbols used to distinguish one element from another.
[0048] Furthermore, in this disclosure, when an element is referred to as “connected” or “attached” to another element, the element may be directly connected or attached to the other element, but it should be understood that these elements may be connected or attached to each other with another intermediate element in between, unless otherwise stated.
[0049] Furthermore, in this disclosure, when an element is expressed as "device / or (unit)", "module", etc., it can indicate that two elements are combined into one element, or that for each subdivided function, one element can be divided into two or more elements. In addition, each of the elements described below, in addition to its primary function, can perform some or all of the functions of other elements, and some of the primary functions of each element can be specifically performed by other elements.
[0050] Furthermore, in this disclosure, "deep neural network (DNN)" is a representative example of an artificial neural network model that simulates a cranial neural network, and is not limited to artificial neural network models that use a specific algorithm.
[0051] Furthermore, in this disclosure, "parameters" are values used in the computation of each layer included in a neural network, and may include, for example, weights (and biases) used when applying input values to a particular arithmetic expression. Parameters can be represented in matrix form. Parameters are values set as a result of training and can be updated as needed with additional training data.
[0052] Furthermore, in this disclosure, "first DNN" refers to a DNN used to analyze a video signal comprising multiple images and generate multiple feature information for each time and frequency; "second DNN" refers to a DNN used to extract a first height component and a first planar component corresponding to the motion of an object in the video signal from the video signal; and "third DNN" can refer to a DNN used to extract a second planar component corresponding to the motion of a sound source within the video from a first audio signal that does not have a height component. "Second DNN" and "third DNN" can also refer to DNNs used to generate correction information between audio features in the two-dimensional audio signal and image features in the video signal from the video signal and a two-dimensional audio signal corresponding to the video signal. In this case, the correction information between the audio features in the audio signal and the image features in the video signal is information corresponding to the second height component included in the three-dimensional audio signal described below, and can be information used to match height components that are inconsistent between the domains of the video / audio signal. "Fourth DNN" can refer to a DNN used to output a second audio signal including a second height component from a first audio signal that does not have a height component based on multiple feature information for each time and frequency. In this case, the second height component can be generated from the first height component, the first planar component, and the second planar component. Meanwhile, the "second DNN" may include a "2-1" DNN for generating feature information corresponding to motion in the horizontal direction of the video signal, and a "2-2" DNN for generating feature information corresponding to motion in the vertical direction of the video signal.
[0053] The "third DNN" can be used to generate feature information corresponding to the horizontal motion of a two-dimensional audio signal.
[0054] The “fourth DNN” may include the “4-1” DNN for generating audio feature information from a two-dimensional audio signal, the “4-2” DNN for generating a three-dimensional audio signal from audio / video integrated feature information that integrates audio and image feature information, and the “4-3” DNN for generating frequency correction information based on the aforementioned audio / video integrated feature information and correction information.
[0055] The embodiments based on the technical concept of this disclosure are described in detail below.
[0056] Figure 1 This is a block diagram illustrating the configuration of a video processing apparatus according to an embodiment.
[0057] As mentioned above, in order to provide a spatial stereo effect for sound, a method for easily generating three-dimensional audio signals with a large number of audio signal channels is necessary.
[0058] like Figure 1 As shown, the video processing apparatus 100 can generate a three-dimensional audio signal 103 by using a two-dimensional audio signal 102 and a video signal 101 corresponding to the two-dimensional audio signal 102 as inputs. Here, the two-dimensional audio signal 102 refers to an audio signal in which the audio information in the height direction (audio information of the height component) is uncertain or not included, and the audio information in the left-right and front-back directions (audio information of the planar component) is certain, such as a 2-channel, 5.1-channel, 7.1-channel, and 9.1-channel audio signal. For example, the two-dimensional audio signal 102 may be stereo audio including a left (L) channel and a right (R) channel.
[0059] In this case, the two-dimensional audio signal 102 can be output through an audio signal output device located at the same height, so that the user can feel the spatial stereo effect of the sound in the left-right and front-back directions.
[0060] Meanwhile, the three-dimensional audio signal 103 represents an audio signal that includes audio information in the height direction as well as audio information in the left-right and front-back directions. For example, the three-dimensional audio signal 103 can be a 4-channel surround sound audio signal including W channel, X channel, Y channel and Z channel, but is not limited to this. Here, the W channel signal can indicate the sum of the intensities of the omnidirectional sound sources, the X channel signal can indicate the intensity difference between the front and back sound sources, the Y channel signal can indicate the intensity difference between the left and right sound sources, and the Z channel signal can indicate the intensity difference between the top and bottom sound sources.
[0061] In other words, when the channels are configured to effectively include audio signals in the height direction (the height component of the audio signal), a three-dimensional audio signal can typically include a multi-channel surround sound audio signal with more channels than a two-channel audio signal. In this case, the three-dimensional audio signal can be output through audio signal output devices located at different heights, allowing the user to experience a spatial stereo effect of sound in the vertical direction (height direction) as well as the left, right, and front-back directions.
[0062] In embodiments of this disclosure, a three-dimensional audio signal 103 can be generated from a two-dimensional audio signal 102 by the following steps: obtaining image feature information (feature information for each time and frequency) from a video signal 101 corresponding to the two-dimensional audio signal, and generating features (corresponding to a second height component) of vertical (height direction) motion of a sound source (corresponding to an object in the video) that is not obviously included in the two-dimensional audio signal, based on the motion features (corresponding to a first height component and a first planar component) of an object in the video (corresponding to a sound source in the audio) included in the image feature information.
[0063] Furthermore, there may be slight differences between the audio and video domains. In other words, in video, the motion information of an object in the left-right (X-axis) and vertical (Z-axis) directions is relatively clear, but the motion information in the front-back (Y-axis) direction is uncertain. This is because, due to the nature of video, it is difficult to include information related to the front-back direction in the motion information of an object in video.
[0064] Therefore, errors may occur when generating a three-dimensional audio signal from a two-dimensional audio signal using motion information of objects in a video. Furthermore, when the two-dimensional audio signal is a 2-channel stereo signal, the motion information of the sound source (corresponding to the object) in the two-dimensional audio signal is relatively clear in the left-right (X-axis) and front-back (Y-axis) directions, but the motion information in the vertical direction (Z-axis) is uncertain.
[0065] Therefore, when correction is performed by considering the difference between the motion information of objects in the video in the left-right (X-axis) direction (horizontal direction) and the motion information of the sound source in the two-dimensional audio signal in the left-right (X-axis) direction (horizontal direction) (difference between the audio domain and the image domain), a three-dimensional audio signal can be effectively generated and output from the two-dimensional audio signal using the video signal. Meanwhile, the image feature information generation unit 110, the correction information generation unit 120, and the three-dimensional audio output unit 130 in the video processing apparatus 100 can be implemented based on artificial intelligence (AI), and the AI used in the image feature information generation unit 110, the correction information generation unit 120, and the three-dimensional audio output unit 130 can be implemented as a DNN.
[0066] refer to Figure 1 The video processing apparatus 100 according to an embodiment may include an image feature information generation unit 110, a correction information generation unit 120, a three-dimensional audio output unit 130, and a synchronization unit 140. This disclosure is not limited thereto, and as such... Figure 1 As shown, the video processing apparatus 100 according to the embodiment may further include a frequency conversion unit 125. Optionally, the frequency conversion unit 125 may be included in the three-dimensional audio output unit 130.
[0067] exist Figure 1 In this embodiment, the image feature information generation unit 110, the correction information generation unit 120, the three-dimensional audio output unit 130, and the synchronization unit 140 are shown as separate elements. However, the image feature information generation unit 110, the correction information generation unit 120, the three-dimensional audio output unit 130, and the synchronization unit 140 can be implemented using a single processor. In this case, they can be implemented using a dedicated processor, or using a combination of a general-purpose processor (such as an application processor (AP), a central processing unit (CPU), and a graphics processing unit (GPU)) and software. Furthermore, the dedicated processor may include memory for implementing embodiments of this disclosure, or may include a memory processing unit for using external memory.
[0068] The image feature information generation unit 110, the correction information generation unit 120, the three-dimensional audio output unit 130, and the synchronization unit 140 can also be configured using multiple processors. In this case, they can be implemented using a combination of dedicated processors, or using a combination of multiple general-purpose processors (such as AP, CPU, or GPU) and software.
[0069] The image feature information generation unit 110 can obtain image feature information from the video signal 101 corresponding to the two-dimensional audio signal 102. The image feature information is information about components (for each time / frequency) related to corresponding features in which motion exists, such as objects in the image, and can be feature information for each time and frequency. The corresponding object can correspond to a sound source in the two-dimensional audio signal 102; therefore, the image feature information can be visual feature pattern mapping information corresponding to the sound source used to generate three-dimensional audio.
[0070] The image feature information generation unit 110 can be implemented based on AI. The image feature information generation unit 110 can analyze a video signal including multiple images and generate multiple feature information for each time and frequency based on a first DNN. See below for reference. Figure 3 An example describing the first DNN.
[0071] The image feature information generation unit 110 can synchronize the video signal with the two-dimensional audio signal using a first DNN, and obtain M one-dimensional image feature mapping information (M is an integer greater than or equal to 1) from the video signal 101 based on the (position or) motion of objects in the video. In other words, the M samples can indicate feature patterns corresponding to the (position or) motion of objects in the video. In other words, the one-dimensional image feature mapping information can be generated from at least one frame (or frame window). Simultaneously, by repeatedly obtaining the one-dimensional image feature mapping information, two-dimensional image feature information (feature information at each time step) with multiple frame windows can be obtained.
[0072] The image feature information generation unit 110 can perform tiling on frequencies and fill all frequency windows with the same value, thereby obtaining three-dimensional image feature mapping information (feature information for each time and frequency) with image features, frame windows, and frequency window components. In other words, M image feature mapping information for time and frequency can be obtained. Here, the frequency window represents the type of frequency index, which indicates the frequency (range) corresponding to the value of each sample. Furthermore, the frequency window represents the type of frame index, which indicates the frame (range) corresponding to the value of each sample.
[0073] The following is for reference. Figure 2 The detailed operation of the image feature information generation unit 110 is described below, and references are made below. Figure 3 An example describing the first DNN.
[0074] The correction information generation unit 120 can generate correction information between audio features in the audio signal 102 and image features in the video signal 102 from the video signal 101 and the two-dimensional audio signal 102. The audio features in the two-dimensional audio signal 102 can represent feature components corresponding to the motion of a sound source (corresponding to an object) in the audio. The correction information generation unit 120 can be implemented based on AI. The correction information generation unit 120 can extract a first height component and a first planar component corresponding to the motion of an object (corresponding to a sound source) in the video signal 101 based on a second DNN, and extract a second planar component corresponding to the motion of a sound source in the audio from the two-dimensional audio signal 102, which does not have a height component, based on a third DNN. The correction information generation unit 120 can generate correction information corresponding to the second height component from the first height component, the first planar component, and the second planar component.
[0075] In other words, the correction information generation unit 120 can generate correction information from the video signal and the corresponding two-dimensional audio signal by using a second DNN and a third DNN. See below for reference. Figures 6a to 7 Examples describing the second and third DNNs.
[0076] The correction information generation unit 120 can synchronize the video signal 101 with the two-dimensional audio signal 102 and obtain feature information corresponding to the horizontal motion in the video (corresponding to the first planar component) and feature information corresponding to the vertical motion in the image (corresponding to the first height component).
[0077] The correction information generation unit can obtain feature information (corresponding to the second plane component) that corresponds to the horizontal motion in the audio signal from the two-dimensional audio signal.
[0078] Specifically, the correction information generation unit 120 can obtain N+M feature mapping information (N and M are integers greater than or equal to 1) corresponding to the horizontal motion in the video relative to time in the video signal 101 using the (2-1)th DNN. In other words, the two-dimensional mapping information includes multiple frame window components and N+M feature components corresponding to the motion.
[0079] Simultaneously, the correction information generation unit 120 can obtain N+M feature mapping information (N and M are integers greater than or equal to 1) corresponding to the vertical motion in the video relative to time from the video signal 101 using the (2-2)th DNN. In other words, it can obtain two-dimensional mapping information including multiple frame window components and N+M feature components corresponding to the motion.
[0080] Meanwhile, see below for reference Figure 6a and Figure 6b Examples describing the (2-1)th DNN and the (2-2)th DNN.
[0081] The correction information generation unit 120 can use a third DNN to obtain feature mapping information corresponding to horizontal motion in the video from the two-dimensional audio signal 102. In other words, it can obtain two-dimensional mapping information including multiple frame window components and N+M feature components corresponding to the motion. Meanwhile, refer to the following... Figure 7 An example describing a third DNN.
[0082] The correction information generation unit 120 can generate time-specific correction information based on feature information corresponding to horizontal motion in the video, feature information corresponding to vertical motion in the video, and feature information corresponding to horizontal motion in the audio.
[0083] Specifically, the correction information generation unit 120 can obtain N+M correction mapping information relative to time based on N+M feature mapping information corresponding to horizontal and vertical motion in the image relative to time and feature mapping information corresponding to horizontal motion in the audio. In this case, a fourth value of the N+M correction mapping information relative to time can be obtained based on a ratio set by considering the relationship between a first value of the feature mapping information corresponding to horizontal motion in the image, a second value of the feature mapping information corresponding to horizontal motion in the audio, and a third value of the feature mapping information corresponding to vertical motion in the image, and N+M correction mapping information relative to time including the fourth value can be generated.
[0084] The correction information generation unit 120 can perform frequency-dependent tiling on the correction information relative to time, and obtain correction information relative to time and frequency. For example, the correction information generation unit 120 can correct mapping information including multiple frame window components, multiple frequency window components, and N+M correction parameter components. In other words, the correction information generation unit 120 fills the correction parameter components with the same value relative to all frequency windows, so that the three-dimensional correction mapping information has correction parameters (or domain matching parameters), frame windows, and frequency window components.
[0085] The following is for reference. Figure 4 The detailed operation of the correction information generation unit 120 is described.
[0086] The frequency conversion unit 125 can convert the two-dimensional audio signal 102 into a frequency domain two-dimensional audio signal according to various conversion methods, such as short-time Fourier transform (STFT). The two-dimensional audio signal 102 includes samples divided according to the channel and time, and the frequency domain signal includes samples divided according to the channel, time, and frequency windows.
[0087] The three-dimensional audio output unit 130 can generate and output a three-dimensional audio signal based on a frequency-domain two-dimensional audio signal, image feature information (multiple feature information for each time and frequency), and correction information. The three-dimensional audio output unit 130 can be implemented based on AI. The three-dimensional audio output unit 130 can generate and output a three-dimensional audio signal based on a frequency-domain two-dimensional audio signal, image feature information (multiple feature information for each time and frequency), and correction information. See below for reference. Figures 9 to 11 An example describing the fourth DNN.
[0088] The three-dimensional audio output unit 130 can perform frequency conversion operations on two-dimensional signals and obtain time and frequency information for two channels. However, this disclosure is not limited thereto, and as described above, when the frequency conversion unit 125 exists separately from the three-dimensional audio output unit 130, frequency domain two-dimensional audio signal information can be obtained from the frequency conversion unit 125 without performing frequency conversion operations.
[0089] Frequency domain two-dimensional audio signal information can include time (frame window) and frequency information (frequency window) for two channels. In other words, frequency domain two-dimensional audio signal information can include sample information divided by frequency windows and time.
[0090] The three-dimensional audio output unit 130 can generate audio feature information about time and frequency from the time and frequency of the two channels. Specifically, the three-dimensional audio output unit 130 can use the (4-1)th DNN to generate N audio feature mapping information about time and frequency from the time and frequency information of the two channels. See below for reference. Figure 9 An example describing the (4-1)th DNN.
[0091] The three-dimensional audio output unit 130 can generate integrated audio / image feature information based on audio feature information about time and frequency (audio feature information for each time and frequency) and image feature information about time and frequency (image feature information for each time and frequency). Specifically, the three-dimensional audio output unit 130 can generate N+M integrated audio / image feature mapping information based on M image feature mapping information about time and frequency and N audio feature mapping information about time and frequency.
[0092] The three-dimensional audio output unit 130 can generate an n-channel frequency domain three-dimensional audio signal (n is an integer greater than 2) from audio / image integrated feature mapping information. Specifically, the three-dimensional audio output unit 130 can use the (4-2)th DNN to generate a frequency domain three-dimensional audio signal for n channels from N+M audio / image integrated feature mapping information. See below for reference. Figure 10 Describe an example of the (4-2)th DNN.
[0093] The three-dimensional audio output unit 130 can obtain n-channel audio correction information based on audio / image integrated feature information and time and frequency correction information. Specifically, the three-dimensional audio output unit 130 can obtain n-channel audio correction mapping information (for frequency correction information) from N+M audio / image integrated feature mapping information for time and frequency and N+M correction mapping information for time and frequency.
[0094] The three-dimensional audio output unit 130 can perform correction on the frequency-domain three-dimensional audio signal of the n-channel based on the audio correction mapping information of the n-channel, and obtain the corrected frequency-domain three-dimensional audio signal of the n-channel. In this case, a three-dimensional audio signal including a second height component can be output. Specifically, the second height component is a height component generated by correcting the height component included in the frequency-domain three-dimensional audio signal of the n-channel based on the correction information, and therefore can be a component that well reflects the motion of the sound source in the audio. The three-dimensional audio output unit 130 can perform inverse frequency conversion on the frequency-domain three-dimensional audio signal for n-channel correction, and generate and output a three-dimensional audio signal for the n-channel.
[0095] The following is for reference. Figure 8 This describes the detailed modules and operation of the three-dimensional audio output unit 130.
[0096] Simultaneously, a first DNN, a second DNN, a third DNN, and a fourth DNN can be trained based on the comparison results between the first frequency domain-trained reconstructed three-dimensional audio signal reconstructed from the first training two-dimensional audio signal and the first training image signal, and the first frequency domain-trained three-dimensional audio signal obtained by frequency conversion of the first training three-dimensional audio signal. (See below for reference.) Figure 12 Describe the training of the first DNN, second DNN, third DNN and fourth DNN.
[0097] Simultaneously, time and frequency correction information can be modified based on user (input) parameter information. In this case, the first DNN, second DNN, third DNN, and fourth DNN can be trained based on the comparison results between the frequency domain training reconstructed three-dimensional audio signal reconstructed based on the first training two-dimensional audio signal, the first training image signal, and user parameter information, and the first frequency domain training three-dimensional audio signal obtained by frequency conversion of the first training three-dimensional audio signal. See below for reference. Figure 13 The description further considers the training of the first, second, third, and fourth DNNs based on user input parameters.
[0098] Meanwhile, the first training two-dimensional audio signal and the first training image signal can be the same device as the video processing device (or the training device described below), or can be obtained from a portable terminal that is a different terminal connected to the video processing device (or the training device described below). The first training three-dimensional audio signal can be obtained from a surround sound microphone included or installed on the portable terminal. See below for reference. Figure 16 Describe how a portable terminal acquires training signals.
[0099] Meanwhile, the parameter information of the first to third DNNs obtained as the training results of the first DNN, the second DNN, the third DNN and the fourth DNN can be stored in the video processing device, or can be received from a terminal connected to the video processing device (or the training device described below).
[0100] Synchronization unit 140 can synchronize video signal 101 with three-dimensional audio signal 103, and output synchronized three-dimensional audio and video signals. (See below for reference.) Figures 3 to 11 The description includes detailed modules of the image feature information generation unit 110, correction information generation unit 120 and three-dimensional audio output unit 130 included in the video processing apparatus 100, detailed operation of the modules, and first DNN to fourth DNN included in the image feature information generation unit 110, correction information generation unit 120 and three-dimensional audio output unit 130.
[0101] Figure 2 This is a diagram illustrating the detailed operation of the image feature information generation unit 110 according to an embodiment.
[0102] refer to Figure 2 The image feature information generation unit 110 may include a synchronization unit 210, a first DNN 220, and a tiling unit 230.
[0103] First, the synchronization unit 210 can synchronize the video signal V(t, h, w, 3) with the two-dimensional audio signal. In other words, because the sampling frequency of the two-dimensional audio signal (e.g., 48 kHz) and the sampling frequency of the video signal (e.g., 48 kHz) (e.g., 60 Hz) are different from each other, and in particular, the sampling frequency of the audio signal is significantly greater than the sampling frequency of the image signal, a synchronization operation can be performed to match the samples of the two-dimensional audio signal with the samples (frames) of its corresponding video signal.
[0104] The first DNN 220 can be used to obtain image feature information V from the synchronization signal V(t, h, w, 3). inf A DNN of (1, 1, M') is used. In this case, the image feature information can be M one-dimensional image feature information. The tiling unit 230 can use the first DNN 220 to accumulate M' one-dimensional image feature information for each frame window in order to obtain M' two-dimensional image feature information V with respect to multiple frame windows (τ) (i.e., time). inf (1, τ, and M').
[0105] The tiling unit 230 can process M' two-dimensional image feature information V about multiple frame windows. infTile the frequency components on (1, τ, and M') to obtain three-dimensional image feature information V with respect to multiple frame windows (τ) (i.e., time) and multiple frequency windows (f) (i.e., frequency). inf (f, τ, and M'). In other words, through the two-dimensional image feature information V inf By filling all frequency components with the same image feature values (1, τ, M'), three-dimensional image feature information V can be obtained. inf (1, τ, M').
[0106] Figure 3 This is a diagram used to describe the first DNN 300 according to an embodiment.
[0107] The first DNN 300 may include at least one convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer processes the input data with filters of a predetermined size to obtain feature data. The parameters of the filters in the convolutional layer can be optimized through the training process described below. The pooling layer is used to extract and output feature values from only some of the features of all samples in order to reduce the size of the input data, and may include max pooling layers and average pooling layers, etc. The fully connected layer is a layer where neurons in one layer connect to all neurons in the next layer, and is used for classifying features.
[0108] The reduction layer, as an example of a pooling layer, can primarily represent a pooling layer used to reduce the data size of the input image before it is fed into the convolutional layer.
[0109] refer to Figure 3 The video signal 301 is input into the first DNN 300. The video signal 301 includes samples divided into input channels, time, height, and width. In other words, the video signal 301 can be four-dimensional data of samples. Each sample in the video signal 301 can be a pixel value. The input channels of the video signal 301 are RGB channels, and can be 3, but are not limited to this.
[0110] Figure 3 The size of video signal 301 is shown as (t, h, w, 3), which means that the duration of video signal 301 is t, the number of input channels is 3, the height of the image is h, and the width of the image is w. The duration t represents the number of frames, and each frame corresponds to a certain time period (e.g., 5 ms). The size of video signal 301 as (t, h, w, 3) is merely an example, and according to embodiments, the size of video signal 301, the size of the signal input to each layer, and the size of the signal output from each layer can be modified differently. For example, h and w can be 224, but are not limited to this.
[0111] As a result of the processing in reduction layer 301, video signal 301 can be reduced, and a first intermediate signal can be obtained. In other words, through reduction, the number of samples divided by the height (h) and width (w) of video signal 301 is reduced, and the height and width of video signal 301 are reduced. For example, the height and width of video signal 301 can be 112, but are not limited to this.
[0112] The first convolutional layer 320 consists of c filters of size a×b and processes the reduced image signal (first intermediate signal) 302. For example, as a result of the processing of the first convolutional layer 320, a second intermediate signal 303 of size (112, 112, c) can be obtained. In this case, the first convolutional layer 320 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained. The first layer and the second layer can be identical to each other. However, this disclosure is not limited thereto, and the second layer is a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer is Parametric Rectified Linear Unit (PReLU), and the parameters of the activation function can be trained together.
[0113] By using the first pooling layer 330, pooling can be performed on the second intermediate signal 303. For example, as a result of the processing of the pooling layer 330, a third intermediate signal (14, 14, c) can be obtained.
[0114] The second convolutional layer 340 processes the input signal using f filters of size d×e. As a result of the processing of the second convolutional layer 340, a fourth intermediate signal 305 of size (14, 14, f) is obtained.
[0115] Meanwhile, the third convolutional layer 350 can be a 1×1 convolutional layer. The third convolutional layer 350 can be used to adjust the number of channels. As a result of the processing of the third convolutional layer 350, a fifth intermediate signal 306 with a size of (14, 14, g) can be obtained.
[0116] The first fully connected layer 360 can classify the input feature signal and output a one-dimensional feature signal. As the processing result of the first fully connected layer 360, an image feature signal 307 of size (1, 1, M') can be obtained.
[0117] According to an embodiment of this disclosure, the first DNN 300 obtains an image feature signal corresponding to the motion of an image object (corresponding to a sound source) from the video signal 301. In other words, although Figure 3The first DNN 300 shown includes three convolutional layers, one reduction layer, one pooling layer, and one fully connected layer. However, this is merely an example, and when an image feature signal 307 comprising M image features can be obtained from the video signal 301, various modifications can be made to the number of convolutional layers, the number of reduction layers, the number of pooling layers, and the number of fully connected layers included in the first DNN 300. Similarly, the number and size of filters used in each convolutional layer can be modified in various ways, and the connection order and method between each layer can also be modified in various ways.
[0118] Figure 4 This is a diagram illustrating the detailed operation of the correction information generation unit 120 according to an embodiment.
[0119] refer to Figure 4 The correction information generation unit 120 may include a synchronization unit 410, a second and a third DNN 420, a correction mapping information generation unit 430, and a tiling unit 440.
[0120] refer to Figure 4 The synchronization unit 410 can synchronize the video signal V(t, h, w, 3) with the two-dimensional audio signal. In other words, it can perform a synchronization operation to match samples of the two-dimensional audio signal with samples (frames) of its corresponding image signal.
[0121] The (2-1) DNN 421 can be a DNN used to generate feature mapping information m_v_H(1, τ, N+M') corresponding to the motion of the image in the horizontal direction from the video signal V(t, H, w, 3). In this case, the feature mapping information corresponding to the motion of the image in the horizontal direction can be N+M' image feature information relative to two-dimensional time (frame window) (N and M' are integers greater than or equal to 1).
[0122] The (2-2)DNN 422 can be a DNN used to generate feature mapping information m_v_V(1, τ, N+M') (corresponding to the first height component) corresponding to the motion of the image in the vertical direction from the synchronized video signal V(t, h, w, 3). In this case, the feature mapping information corresponding to the motion of the image in the vertical direction can be N+M' image feature information relative to two-dimensional time (frame window) (N and M' are integers greater than or equal to 1).
[0123] The third DNN 423 can be used to extract data from a two-dimensional audio signal A. In_2D(t, 2) Generate a DNN that generates feature mapping information m_a_H (1, τ, N+M') corresponding to the motion of the audio in the horizontal direction (corresponding to the second plane component). In this case, the feature mapping information corresponding to the motion of the audio in the horizontal direction can be N+M' image feature information relative to two-dimensional time (frame window) (N and M' are integers greater than or equal to 1).
[0124] The correction mapping information generation unit 430 can obtain correction mapping information α from the feature mapping information m_v_H (1, τ, N+M') corresponding to the motion of the image in the horizontal direction, the feature mapping information m_v_V (1, τ, N+M') corresponding to the motion of the image in the vertical direction, and the feature mapping information m_a_H (1, τ, N+M') corresponding to the motion of the audio in the horizontal direction. inf (1, τ, N+M'). Specifically, the correction mapping information generation unit 430 can obtain the correction mapping information α according to the [Mathematical Formula 1] shown below. inf (1, τ, N+M').
[0125] [Mathematical Formula 1]
[0126]
[0127] [Mathematical Formula 1] is based on the theoretical background described below. (See below for reference.) Figures 5a to 5b The description derives from it the parameter α used to obtain the domain matching. inf The theoretical background of [Mathematical Formula 1].
[0128] refer to Figure 5a and Figure 5b Even when the number of motion information mv1 and mv2 of objects in an image is the same as in Case 1 510 and Case 2 520, there may still be cases where the motion degree S of the sound source (corresponding to the object in the image) in Case 1 510 does not correspond to the motion degree S of the sound source (corresponding to the object) in the image in Case 2 520. This is because for virtually every image scene included in the image sensor and camera imaging system, distortion occurs due to differences in the degree of depth perspective deformation, and the information of the sound source object in the image and the motion information of the sound source in the audio are essentially not related to each other.
[0129] Therefore, correction parameters (or domain matching parameters) can be obtained to resolve inconsistencies in motion information, rather than using feature information corresponding to the motion of objects in the image to generate 3D audio.
[0130] In other words, while motion information of objects in an image can be used in the left-right (X-axis) and vertical (Z-axis) directions, motion information in the front-back (Y-axis) direction is uncertain. Therefore, when using the corresponding motion information to generate 3D audio, the error can be significant.
[0131] Meanwhile, in the motion information of the sound source in audio, motion information in the left-right direction (X-axis direction) and the front-back direction (Y-axis direction) can be used, but motion information in the vertical direction (Z-axis direction) may be uncertain.
[0132] To address this inconsistency in motion information, correction parameters can be obtained based on the motion information along the X-axis, which shares a common deterministic characteristic.
[0133] In this case, by comparing the relatively precise X-axis direction information from the motion information of the sound source in the audio with the X-axis direction information from the motion information of the object in the image, the object motion information in the Z-axis direction of the image can be corrected based on the sound source motion information in the Z-axis direction of the audio domain (domain matching). For example, the X-axis and Y-axis direction information mv1_x and mv1_z included in the motion information of the object in the image of Case 1 510 are (10, 2), and the X-axis direction information Smv1_x included in the motion information of the sound source in the audio of Case 1 510 is 5. Based on the proportional expression, the Z-axis direction information Smv1_y of the sound source in the audio can be obtained as 1. In the motion information of the object in the image of Case 2520, the information mv1_x and mv1_z in the X and Z axes are (10, 2), and the motion information of the sound source in the audio of Case 2520, including the information Smv1_x in the X-axis direction, is 8. Based on the scaling expression, the information Smv1_y in the Z-axis direction of the sound source in the audio can be obtained as 1.6. In other words, based on the scaling expression Smv1_x : mv1_x = Smv1_z : mv1_z, it can be Smv1_z = Smv1_x * mv1_z / mv1_x. In this case, the value of Smv1_z can be used as a correction parameter.
[0134] Based on the above method for deriving correction parameters, the above [Mathematical Formula 1] can be derived. The tiling unit 440 can obtain correction mapping information α by tiling the frequency components of the two-dimensional N+M' correction mapping information received from the correction mapping information generation unit 430. inf (f, t, N+M'). In other words, through the two-dimensional correction mapping information α inf By filling all frequency components with the same image feature values (f, t, N+M'), the three-dimensional correction mapping information α can be obtained.inf (1, t, N+M').
[0135] Figure 5c This is a diagram describing an algorithm for estimating the height component of a sound source in an audio signal, which is necessary for analyzing the motion of objects in video signals and the motion of sound sources in two-dimensional audio to generate three-dimensional audio signals.
[0136] refer to Figure 5c The video processing apparatus 100 can analyze a video signal and extract feature information related to a first height component and a first planar component, which are related to the motion of objects in the video. Simultaneously, the video processing apparatus 100 can analyze a two-dimensional audio signal and extract feature information related to a second planar component, which is related to the motion of a sound source in the two-dimensional audio signal. The video processing apparatus 100 can estimate second height component feature information related to the motion of the sound source based on the feature information of the first height component, the first planar component, and the second planar component. The video processing apparatus 100 can output a three-dimensional audio signal including the second height component from the two-dimensional audio signal based on the feature information related to the second height component. In this case, the feature information related to the second height component can correspond to the above-mentioned reference. Figure 4 The description of the correction mapping information.
[0137] Figure 6a This is a diagram used to describe the (2-1)th DNN 600.
[0138] The (2-1) DNN 600 may include at least one convolutional layer, a pooling layer, and a fully connected layer. A reduction layer is an example of a pooling layer and can primarily represent a pooling layer used to reduce the data size of the input image before it is fed into the convolutional layer.
[0139] refer to Figure 6a The video signal 601 is input into the (2-1)th DNN 600. The video signal 601 includes samples divided into input channels, time, height, and width. In other words, the video signal 601 can be four-dimensional data of the samples.
[0140] The magnitude of the video signal 601 (t, h, w, 3) is merely an example, and according to embodiments, the magnitude of the video signal 601, the magnitude of the signal input to each layer, and the magnitude of the signal output from each layer can be modified differently. For example, h and w can be 224, but are not limited thereto.
[0141] The first intermediate signal 602 is obtained by reducing the video signal 601 using a reduction layer 610. In other words, by reducing, the number of samples divided by the height (h) and width (w) of the video signal 601 is reduced, and the height and width of the video signal 601 are reduced. For example, the height and width of the first intermediate signal 602 can be 112, but are not limited to this.
[0142] The first convolutional layer 615 processes the reduced image signal using c filters of size a×b. In this case, to obtain the feature components corresponding to motion in the horizontal direction, a filter of size 3×1 in the horizontal direction can be used. For example, as a result of the processing of the first convolutional layer 615, a second intermediate signal 603 of size (112, 112, c) can be obtained. In this case, the first convolutional layer 615 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained together. The first layer and the second layer can be the same layer, but are not limited thereto, and the second layer can be a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer is PReLU, and the parameters of the activation function can be trained together.
[0143] The third intermediate signal 604 can be obtained by performing pooling on the second intermediate signal 603 using the first pooling layer 620. For example, as a result of the processing of the pooling layer 620, the third intermediate signal (14, 14, c) can be obtained, but it is not limited thereto.
[0144] The second convolutional layer 625 can process the input signal using f filters of size d×e, thereby obtaining the fourth intermediate signal 605. As a result of the processing of the second convolutional layer 625, a fourth intermediate signal 605 of size (14, 14, f) can be obtained, but it is not limited to this.
[0145] Meanwhile, the third convolutional layer 630 can be a 1×1 convolutional layer. The third convolutional layer 630 can be used to adjust the number of channels. As a result of the processing of the third convolutional layer 630, a fifth intermediate signal 606 of size (14, 14, g) can be obtained, but it is not limited to this.
[0146] The first fully connected layer 635 can output a one-dimensional feature signal by classifying the input feature signal. As a result of the processing of the first fully connected layer 635, a feature component signal 607 corresponding to a horizontal movement of magnitude (1, 1, N+M') can be obtained.
[0147] According to embodiment (2-1), the DNN 600 obtains image feature signals 607 from the video signal 601 corresponding to the horizontal motion of the image object (corresponding to the sound source). In other words, although... Figure 6aThe (2-1) DNN 600 shown includes three convolutional layers, one reduction layer, one pooling layer, and one fully connected layer. However, this is merely an example, and when a feature signal 607 comprising N+M' image features in the horizontal direction can be obtained from the video signal 601, various modifications can be made to the number of convolutional layers, the number of reduction layers, the number of pooling layers, and the number of fully connected layers included in the (2-1) DNN 600. Similarly, the number and size of filters used in each convolutional layer can be modified in various ways, and the connection order and method between each layer can also be modified in various ways.
[0148] Figure 6b This is a diagram used to describe the (2-2)th DNN 650.
[0149] The (2-2)DNN 650 may include at least one convolutional layer, a pooling layer, and a fully connected layer. A reduction layer is an example of a pooling layer and can primarily represent a pooling layer used to reduce the data size of the input image before it is fed into the convolutional layer.
[0150] refer to Figure 6b The video signal 651 is input into the (2-2)th DNN 650. The video signal 651 includes samples divided into input channels, time, height, and width. In other words, the image signal 651 can be four-dimensional data of the samples.
[0151] The magnitude of the video signal 651 (t, h, 2, 3) is merely an example, and according to embodiments, the magnitude of the video signal 651, the magnitude of the signal input to each layer, and the magnitude of the signal output from each layer can be modified differently. For example, h and w can be 224. This disclosure is not limited thereto.
[0152] The first intermediate signal 652 is obtained by reducing the video signal 651 using a reduction layer 660. In other words, by reducing the signal, the number of samples divided by the height (h) and width (w) of the image signal 651 is reduced, and the height and width of the video signal 651 are reduced. For example, the height and width of the first intermediate signal 652 can be 112, but are not limited thereto.
[0153] The first convolutional layer 665 processes the reduced image signal using c filters of size a×b. In this case, to obtain the feature components corresponding to motion in the vertical direction, a filter of size 1×3 in the vertical direction can be used. For example, as a result of the processing of the first convolutional layer 665, a second intermediate signal 653 of size (112, 112, c) can be obtained. In this case, the first convolutional layer 665 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained. The first layer and the second layer can be the same layer. However, this disclosure is not limited thereto, and the second layer can be a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer is PReLU, and the parameters of the activation function can be trained together.
[0154] The second intermediate signal 653 can be pooled using the first pooling layer 670. For example, as a result of the processing of the pooling layer 670, a third intermediate signal (14, 14, c) can be obtained, but it is not limited thereto.
[0155] The second convolutional layer 675 processes the input signal using f filters of size d×e, thereby obtaining a fourth intermediate signal 655. As a result of the processing of the second convolutional layer 675, a fourth intermediate signal of size (14, 14, f) can be obtained, but this disclosure is not limited thereto.
[0156] Meanwhile, the third convolutional layer 680 can be a 1×1 convolutional layer. The third convolutional layer 680 can be used to adjust the number of channels. As a result of the processing of the third convolutional layer 680, a fifth intermediate signal 656 of size (14, 14, g) can be obtained.
[0157] The first fully connected layer 685 can output a one-dimensional feature signal by classifying the input feature signal. As a result of the processing of the first fully connected layer 685, a feature component signal 657 corresponding to a horizontal movement of magnitude (1, 1, N+M') can be obtained.
[0158] According to the second-second DNN 650 of this embodiment, an image feature signal 657 corresponding to the vertical motion of the image object (sound source) is obtained from the video signal 651. In other words, although... Figure 6bThe example shown is of a (2-2)DNN 650 comprising three convolutional layers, one reduction layer, one pooling layer, and one fully connected layer. However, this is merely an example, and when an image feature signal 657 comprising N+M' image features in the horizontal direction can be obtained from the video signal 651, various modifications can be made to the number of convolutional layers, the number of reduction layers, the number of pooling layers, and the number of fully connected layers included in the first DNN 600. Similarly, the number of filter sizes used in each convolutional layer can be modified in various ways, and the connection order and method between each layer can also be modified in various ways.
[0159] Figure 7 This is a diagram used to describe the third DNN 700.
[0160] The third DNN 700 may include at least one convolutional layer, a pooling layer, and a fully connected layer. A reduction layer is an example of a pooling layer and can primarily represent a pooling layer used to reduce the data size of the input image before it enters the convolutional layer.
[0161] refer to Figure 7 The two-dimensional audio signal 701 is input to the third DNN 700. The two-dimensional audio signal 701 includes samples divided into input channels and time. In other words, the two-dimensional audio signal 701 can be two-dimensional data of samples. Each sample in the two-dimensional audio signal 701 can be an amplitude. The input channels of the two-dimensional audio signal 701 can be two channels, but are not limited to this.
[0162] Figure 7 The size of the two-dimensional audio signal 701 is shown as (t, 2), but this indicates that the duration of the two-dimensional audio signal 701 is t, and the number of input channels is 2. The size of the two-dimensional audio signal 701 as (t, 2) is merely an example, and according to embodiments, the size of the two-dimensional audio signal 701, the size of the signal input to each layer, and the size of the signal output from each layer can be modified differently.
[0163] The first convolutional layer 710 processes the two-dimensional audio signal 701 using b filters (one-dimensional filters) of size a×1. For example, as a result of the processing by the first convolutional layer 710, a first intermediate signal 702 of size (512, 1, b) can be obtained. In this case, the first convolutional layer 710 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained together. The first layer and the second layer may be the same layer. However, this disclosure is not limited thereto, and the second layer may be a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer is PReLU, and the parameters of the activation function can be trained together.
[0164] The first intermediate signal 702 can be pooled using the first pooling layer 720. For example, as a result of the pooling layer 720, a second intermediate signal 703 of size (28, 1, b) can be obtained.
[0165] The second convolutional layer 730 processes the input signal using d filters of size c×1. As a result of the processing of the second convolutional layer 730, a third intermediate signal 704 of size (28, 1, d) can be obtained.
[0166] Meanwhile, the third convolutional layer 740 can be a 1×1 convolutional layer. The third convolutional layer 740 can be used to adjust the number of channels. As a result of the processing of the third convolutional layer 740, a fourth intermediate signal 705 with a size of (28, 1, g) can be obtained.
[0167] The first fully connected layer 750 can output a one-dimensional feature signal by classifying the input feature signal. As a result of the processing of the first fully connected layer 750, a feature component signal 706 corresponding to the horizontal movement in the (1, 1, N+M') direction can be obtained.
[0168] According to an embodiment of this disclosure, a third DNN 700 obtains an audio feature signal 706 from a two-dimensional audio signal 701 that corresponds to the horizontal motion of a two-dimensional audio source (corresponding to an object in a video). In other words, although... Figure 7 The diagram shows a third DNN 700 comprising three convolutional layers, one pooling layer, and one fully connected layer. However, this is merely an example, and the number of convolutional layers, pooling layers, and fully connected layers included in the third DNN 700 can be modified in various ways when an audio feature signal 706 comprising N+M' audio features in the horizontal direction can be obtained from a two-dimensional audio signal 701. Similarly, the number and size of filters used in each convolutional layer can be modified in various ways, as can the connection order and method between each layer.
[0169] Figure 8 This is a diagram illustrating the detailed operation of the three-dimensional audio output unit 130 according to an embodiment.
[0170] refer to Figure 8 The three-dimensional audio output unit 130 may include a frequency conversion unit 810, a (4-1) DNN 821, an audio / image feature integration unit 830, a (4-2) DNN 822, a (4-3) DNN 823, a correction unit 840, and an inverse frequency conversion unit 850.
[0171] Frequency conversion unit 810 can convert two-dimensional audio signal A In_2D(t, 2) performs frequency conversion to obtain a two-dimensional audio signal s(f, τ, 2) in the frequency domain. However, as described above, when receiving the two-dimensional audio signal s(f, τ, 2) in the frequency domain from the frequency conversion unit 125, the frequency conversion unit 810 may not be included.
[0172] The (4-1)DNN 821 can be a DNN used to generate audio feature information s(f, τ, N) from a two-dimensional audio signal s(f, τ, 2) in the frequency domain. In this case, the audio feature information can be N one-dimensional audio feature information.
[0173] The audio / image feature integration unit 830 can integrate image feature information V inf (f, τ, M') is integrated with audio feature information s(f, τ, N) to generate audio / image integrated feature information s(f, τ, N+M'). For example, since image feature information is the same as audio feature information in terms of the size of frequency window and frame window components, audio / image feature integration unit 830 can generate audio / image integrated feature information by superimposing image feature mapping information on audio feature information, but is not limited thereto.
[0174] The (4-2)DNN 822 can be used to generate a frequency domain three-dimensional audio signal s(f, τ, N+M') from audio / image integrated feature information s(f, τ, N). 3D A DNN of ) in this case, N 3D It can represent the number of channels in three-dimensional audio.
[0175] The (4-3)DNN 823 can integrate audio / image feature information s(f, τ, N+M') and correction information α. inf (f, τ, N+M') obtains the correction mapping information c(f, τ, N) 3D ).
[0176] The correction unit 840 can be based on the frequency domain three-dimensional audio signal s(f, τ, N) 3D ) and correction mapping information c(f, τ, N) 3D Obtain the corrected frequency domain three-dimensional audio signal Cs(f, τ, N) 3D For example, the correction unit 840 can use the correction mapping information c(f, τ, N) 3D The sample values of ) are added to the frequency domain three-dimensional audio signal s(f, τ, N) 3D The corrected frequency domain three-dimensional audio signal Cs(f, τ, N) is obtained by using the sample values of ) 3DBy correcting the uncertain height component corresponding to the motion of the sound source in the frequency domain three-dimensional audio signal through the correction unit 840 (matching the image domain with the audio domain), the output frequency domain three-dimensional audio signal can have a more definite height component in the frequency domain three-dimensional audio signal.
[0177] The inverse frequency conversion unit 850 can convert the corrected frequency domain three-dimensional audio signal Cs(f, τ, N) into frequency domain signals. 3D Perform inverse frequency conversion to output a three-dimensional audio signal A Pred_B (t,N 3D ).
[0178] Figure 9 This is a diagram used to describe the (4-1) DNN 900 according to the embodiment.
[0179] The (4-1)DNN 900 may include at least one convolutional layer. The convolutional layer processes the input data with filters of a predetermined size to obtain audio feature data. The parameters of the filters in the convolutional layer can be optimized through the training process described below.
[0180] refer to Figure 9 The frequency-domain two-dimensional audio signal 901 is input to the (4-1)th DNN 900. The frequency-domain two-dimensional audio signal 901 can include samples divided into input channels, frame windows, and frequency windows. In other words, the frequency-domain two-dimensional audio signal 901 can be three-dimensional data of samples. Each sample in the frequency-domain two-dimensional audio signal 901 can be a frequency-domain two-dimensional audio signal value. The input channels of the frequency-domain two-dimensional audio signal 901 can be 2 channels, but are not limited to this.
[0181] Figure 9 The magnitude of the frequency domain two-dimensional audio signal 901 is shown to be (f, τ, 2), where the time length (number of frame windows) of the frequency domain two-dimensional audio signal 901 can be τ, the number of input channels can be 2, and the number of frequency windows can be f. According to the embodiment, the magnitude of the frequency domain two-dimensional audio signal 901, the magnitude of the signal input to each layer, and the magnitude of the signal output from each layer can be modified differently.
[0182] The first convolutional layer 910 processes the frequency domain two-dimensional audio signal 901 using c filters of size a×b. For example, as a result of the processing of the first convolutional layer 910, a first intermediate signal of size (f, τ, 32) can be obtained.
[0183] The second convolutional layer 920 processes the first intermediate signal 902 using e filters of size c×d. For example, as a result of the processing of the first convolutional layer, a second intermediate signal 903 of size (f, τ, 32) can be obtained.
[0184] In this configuration, the second convolutional layer 920 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained together. The first and second layers may be the same layer. However, this disclosure is not limited thereto, and the second layer may be a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer is PReLU, and the parameters of the activation function can be trained together.
[0185] The third convolutional layer 930 processes the input second intermediate signal 903 using N filters of size e×f. As a result of the processing by the third convolutional layer 930, audio feature information 904 of (f, τ, N) can be obtained.
[0186] According to the (3-1)th DNN 900 of this embodiment, an audio feature signal 904 corresponding to the horizontal motion of the audio (sound source) is obtained from the frequency domain two-dimensional audio signal 901. In other words, although Figure 9 The example shown is of a (3-1)DNN900 comprising three convolutional layers. However, this is merely an example, and the number of convolutional layers included in the frequency-domain two-dimensional audio signal 901 can be modified in various ways when an audio feature signal 904 comprising N audio features can be obtained from the frequency-domain two-dimensional audio signal 901. Similarly, the number and size of filters used in each convolutional layer can be modified in various ways, as can the connection order and method between each layer.
[0187] Figure 10 This is a diagram used to describe the (4-2) DNN 1000 according to the embodiment.
[0188] The (4-2)DNN 1000 may include at least one convolutional layer. The convolutional layer processes the input data with filters of a predetermined size to obtain audio feature data. The parameters of the filters in the convolutional layer can be optimized through the training process described below.
[0189] refer to Figure 10 The audio / image integrated feature information 1001 is input into the (4-2)DNN 1000. The audio / image integrated feature information 1001 includes the number of features, time (frame window), and frequency window. In other words, the audio / image integrated feature information 1001 can be three-dimensional data about the samples. In other words, each sample in the audio / image integrated feature information 1001 can be an audio / image integrated feature value.
[0190] Figure 10The size of the audio / image integrated feature information 1001 is shown to be (f, τ, N+M'), where the time length (frame window) of the audio / image integrated feature information 1001 can be τ, the number of features corresponding to the frame window and frequency window can be N+M', and the number of frequency windows can be f. According to embodiments, the size of the audio / image integrated feature information 1001, the size of the signal input to each layer, and the size of the signal output from each layer can be modified differently.
[0191] The first convolutional layer 1010 processes the integrated audio / image feature information 1001 using c filters of size a×b. For example, as a result of the processing of the first convolutional layer 1010, a first intermediate signal 1002 of size (f, τ, c) can be obtained.
[0192] The second convolutional layer 1020 processes the first intermediate signal 1002 using e filters of size c×d. For example, as a result of the processing of the second convolutional layer 1020, a second intermediate signal 1003 of size (f, τ, e) can be obtained.
[0193] In this case, the second convolutional layer 1020 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained together. The first and second layers can be the same layer. However, this disclosure is not limited thereto, and the second layer can be a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer is PReLU, and the parameters of the activation function can be trained together.
[0194] The third convolutional layer 1030 uses N with a size of e×f. 3D A filter processes the input signal. As a result of the processing of the third convolutional layer 1030, a filter of size (f, τ, N) can be obtained. 3D The frequency domain three-dimensional audio signal 1004.
[0195] According to the (4-2)th DNN 1000 of this embodiment, a frequency domain three-dimensional audio signal 1004 is obtained from audio / image integrated feature information 1001. In other words, although Figure 10 The example shown is of a (4-2)DNN 1000 comprising three convolutional layers, but this is merely an example, and the number of convolutional layers included in the (4-2)DNN 1000 can be modified in various ways when a frequency-domain three-dimensional audio signal 1004 can be obtained from the audio / image integrated feature information 1001. Similarly, the number and size of filters used in each convolutional layer can be modified in various ways, as can the connection order and method between each layer.
[0196] Figure 11This is a diagram used to describe the (4-3) DNN 1100 according to the embodiment.
[0197] The (4-3)DNN 1100 may include at least one convolutional layer. The convolutional layer processes the input data with filters of a predetermined size to obtain audio feature data. The parameters of the filters in the convolutional layer can be optimized through the training process described below.
[0198] refer to Figure 11 A first intermediate signal 1103 of a new dimension can be obtained by concatenating the audio / image integrated feature information 1101 and the correction information 1102. The audio / image integrated feature information 1001 includes samples divided into feature quantity, time (frame window), and frequency window. In other words, the audio / image integrated feature information 1001 can be three-dimensional data. Each sample in the audio / image integrated feature information 1001 can be an audio / image integrated feature value. The correction information 1102 includes samples divided into feature quantity, time (frame window), and frequency window. In other words, the correction information 1102 can be three-dimensional data. Each sample in the correction information 1102 can be a correction-related feature value.
[0199] Figure 11 The sizes of the audio / image integrated feature information 1101 and correction information 1102 are shown to be (f, τ, N+M'), where the time length (number of frame windows) of the audio / image integrated feature information 1101 and correction information 1102 can be τ, the number of features corresponding to the frame windows and frequency windows can be N+M', and the number of frequency windows can be f. According to embodiments, the sizes of the audio / image integrated feature information 1101 and correction information 1102, the size of the signal input to each layer, and the size of the signal output from each layer can be modified differently.
[0200] The first convolutional layer 1120 processes the first intermediate signal 1103 using c filters of size a×b. For example, as a result of the processing of the first convolutional layer 1120, a second intermediate signal 1104 of size (f, τ, c) can be obtained. In other words, as a result of the processing of the first convolutional layer 1120, a second intermediate signal 1104 of size (f, τ, M”) can be obtained. Here, M” can be 2×(N+M’), but is not limited to this.
[0201] The second convolutional layer 1130 processes the second intermediate signal 1104 using e filters of size c×d. For example, as a result of the processing of the second convolutional layer 1130, a third intermediate signal 1105 of size (f, τ, e) can be obtained. In other words, as a result of the processing of the second convolutional layer 1130, a third intermediate signal 1105 of size (f, t, M”) can be obtained. Here, M” can be 2×(N+M’), but is not limited to this.
[0202] In this case, the second convolutional layer 1130 may include multiple convolutional layers, and the input of the first layer and the output of the second layer can be concatenated and trained together. The first and second layers can be the same layer. However, this disclosure is not limited thereto, and the second layer can be a layer following the first layer. When the second layer is a layer following the first layer, the activation function of the first layer can be PReLU, and the parameters of the activation function can be trained together.
[0203] The third convolutional layer 1140 uses an N array of size e×f. 3D A filter processes the input signal. As a result of the processing of the third convolutional layer 1140, a filter of size (f, τ, N) can be obtained. 3D The correction mapping information 1106.
[0204] According to the (4-3)th DNN 1100 of this disclosure embodiment, correction mapping information 1106 is obtained from audio / image integrated feature information 1101 and correction information 1102. In other words, although Figure 11 The example shown is of a (4-3)DNN 1100 comprising three convolutional layers. However, this is merely an example, and the number of convolutional layers included in the (4-3)DNN 1100 can be modified in various ways when the correction mapping information 1106 can be obtained from the audio / image integrated feature information 1101 and the correction information 1102. Similarly, the number and size of filters used in each convolutional layer can be modified in various ways, as can the connection order and method between each layer.
[0205] Figure 12 This is a diagram used to describe the training methods for the first DNN, second DNN, third DNN, and fourth DNN.
[0206] exist Figure 12 In this context, the first training two-dimensional audio signal 1202 corresponds to the two-dimensional audio signal 102, and the first training image signal 1201 corresponds to the video signal 101. Similarly, each of the training signals corresponds to the reference signal mentioned above. Figure 2 , Figure 4 and Figure 8 The described signals / information.
[0207] The first training image signal 1201 is input into the first DNN 220. The first DNN 220 processes the first training image signal 1201 and obtains the first training image feature signal 1203 according to preset parameters.
[0208] Regarding the first training two-dimensional audio signal 1202, a first frequency domain training two-dimensional audio signal 1204 is obtained through the frequency conversion unit 1220, and the first frequency domain training two-dimensional audio signal 1204 is input to the (4-1)DNN 821. The (4-1)DNN 821 processes the first frequency domain training two-dimensional audio signal 1204 and obtains the first training audio feature signal 1205 according to preset parameters. By processing the first training audio feature signal 1205 and the first training image feature signal 1203 through the audio / image feature integration unit 1220, the first training audio / image integrated feature signal 1206 can be obtained.
[0209] The first training image signal 1201 and the first training two-dimensional audio signal 1202 are input into the second DNN and the third DNN 420. The second DNN and the third DNN 420 (including the (2-1) DNN 421, the (2-2) DNN 422 and the third DNN 423) process the first training two-dimensional audio signal 1202 according to preset parameters and obtain the first training correction signal 1208.
[0210] The first training audio / image integrated feature signal 1206 is input into the (4-2)DNN 822. The (4-2)DNN 822 processes the first training audio / image integrated feature signal 1206 and obtains the first frequency domain training reconstructed three-dimensional audio signal 1207 according to preset parameters.
[0211] The first training correction signal 1207 and the first audio / image integrated feature signal 1206 are input into the (4-3)DNN823.
[0212] The (4-3)DNN 823 processes the first training correction signal 1208 and the first training audio / image integrated feature signal 1206, and obtains the first training frequency correction signal 1209.
[0213] The audio correction unit 1230 can correct the first frequency domain training reconstructed three-dimensional audio signal 1207 based on the first training frequency correction signal 1209, and output the corrected first frequency domain training reconstructed three-dimensional audio signal 1211.
[0214] Meanwhile, relative to the first training three-dimensional audio signal 1212, the first frequency domain training three-dimensional audio signal 1213 is obtained through the frequency conversion unit 1210.
[0215] Based on the comparison between the corrected first frequency domain training 3D audio signal 1213 and the corrected first frequency domain training reconstructed 3D audio signal 1211, generation loss information 1214 is obtained. Generation loss information 1214 may include at least one of the following: L1 norm value, L2 norm value, structural similarity (SSIM) value, peak signal-to-noise ratio-human visual system (PSNR-HVS) value, multi-scale SSIM (MS-SSIM) value, variance inflation factor (VIF) value, and video multi-method evaluation fusion (VMAF) value between the corrected first frequency domain training 3D audio signal 1213 and the corrected first frequency domain training 3D audio signal 1211. For example, loss information 1214 can be expressed as shown in [Mathematical Formula 2].
[0216] [Mathematical Formula 2]
[0217]
[0218] In [Mathematical Formula 2], F() represents the frequency conversion of the frequency conversion unit 1210, and Cs represents the corrected first frequency domain training reconstructed three-dimensional audio signal 1211.
[0219] The generated loss information 1214 indicates the similarity between the corrected first frequency domain training reconstructed three-dimensional audio signal 1211 obtained by processing the first training two-dimensional audio signal 1202 by the first DNN 220 and the first frequency domain training three-dimensional audio signal 1212 obtained by the frequency conversion unit 1210.
[0220] The first DNN 220, the second DNN, the third DNN 420, and the fourth DNN 820 can update parameters to reduce or minimize the generated loss information 1214. The training of the first DNN 220, the second DNN, the third DNN 420, and the fourth DNN 820 can be represented by the mathematical formula shown below.
[0221] [Mathematical Formula 3]
[0222]
[0223] In [Mathematical Formula 3], This represents the parameter set of the first DNN 220, the second DNN, the third DNN 420, and the fourth DNN 820. The first DNN 220, the second DNN, the third DNN 420, and the third DNN 820 are trained to obtain the parameter set used to minimize the generation loss information 1214.
[0224] Figure 13This is a diagram used to describe the training methods of the first DNN, second DNN, third DNN, and fourth DNN that take into account user parameter signals.
[0225] refer to Figure 13 ,and Figure 12 Unlike other DNNs, the correction signal correction unit 1340 exists between the second and third DNNs 420 and the (4-3)DNN 823. The correction signal modification unit 1340 can correct the first training correction signal 1308 of the second and third DNNs 420 using user parameters 1316, and can input the corrected first training correction signal 1315 into the (4-3)DNN 823. For example, the correction signal correction unit 1340 can perform a step of multiplying the value of the first training correction signal 1308 by user parameters C. user The arithmetic operations are used to obtain the first training correction signal for correction, but are not limited to this. In other words, the user parameters are parameters used to adjust the degree of correction of the three-dimensional audio signal through the audio correction unit 1330, and the user (the producer of the three-dimensional audio) can directly input the user parameters, so that the three-dimensional audio signal can be appropriately corrected and reconstructed according to the user's intention.
[0226] Also in Figure 13 In, as referenced Figure 12 As will be understood by those skilled in the art, the parameters of the first DNN 220, the second DNN, the third DNN 420, and the fourth DNN 820 can be trained based on the comparison results between the corrected first frequency domain training reconstructed three-dimensional audio signal 1311 and the first frequency domain training three-dimensional audio signal 1313.
[0227] Figure 14 It is a flowchart used to describe the training process of training device 1400 on the first DNN, second DNN, third DNN and fourth DNN.
[0228] Reference can be performed by training device 1400 Figure 13 The training of the first DNN, second DNN, third DNN, and fourth DNN is described. Training device 1400 may include the first DNN 220, the second DNN and third DNN 420, and the fourth DNN 820. Training device 1400 may be, for example, a video processing device 100 or an additional server.
[0229] The training device 1400 initially sets the parameters of the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823 (S1405).
[0230] The training device 1400 inputs the first training image signal 1201 into the first DNN 220 (S1410).
[0231] The training device 1400 inputs the first training image signal 1201 and the first training two-dimensional audio signal 1202 into the second DNN and the third DNN 420 (S1415).
[0232] The training device 1400 inputs the first frequency domain training two-dimensional audio signal 1204 obtained by the frequency conversion unit 1210 into the (4-1) DNN 821 (S1420).
[0233] The first DNN 220 can output the first training image feature signal 1203 to the audio / image feature integration unit 1410 (S1425).
[0234] The (4-1)DNN 821 can output the first training audio feature signal 1205 to the audio / image feature integration unit 1410 (S1430).
[0235] The audio / image feature integration unit 1410 can output the first trained audio / image integrated feature signal 1206 to the (4-2)DNN 822 and the (4-3)DNN 823 (S1435).
[0236] The (4-2) DNN 822 can output the first training three-dimensional audio signal to the correction unit 1420 (S1440).
[0237] The training device 1400 can input the first training two-dimensional audio signal 1202 and the first frequency domain training two-dimensional audio signal 1204 into the second DNN and the third DNN 420 (S1445).
[0238] The second and third DNNs 420 can output the first training correction signal 1208 to the (4-3)DNN 823 (S1450).
[0239] The (4-3)DNN 823 can output the first training frequency correction signal 1209 to the correction unit 1420 (S1455).
[0240] The correction unit 1420 can output the corrected first frequency domain training reconstructed three-dimensional audio signal 1211 to the training device 1400 (S1460).
[0241] The training device 1400 calculates and generates loss information 1214 by comparing the corrected first frequency domain training reconstructed three-dimensional audio signal 1211 with the first frequency domain training three-dimensional audio signal 1213 obtained through frequency conversion (S1465). Furthermore, the first DNN 220, the second DNN, the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822, and the (4-3) DNN 823 update their parameters based on the generated loss information 1214.
[0242] The training device 1400 can repeat the above operations S1410 to S1490 until the parameters of the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823 are optimized.
[0243] Figure 15 It is a flowchart describing the training process of training device 1500 on the first DNN, second DNN, third DNN and fourth DNN considering user parameters.
[0244] Reference can be performed by training device 1500 Figure 14 The training of the first DNN, second DNN, third DNN, and fourth DNN is described. The training device 1500 may include a first DNN 220, a second DNN, a third DNN 420, and a fourth DNN 820. The training device 1500 may be, for example, a video processing apparatus 100 or an auxiliary server. When training is performed on the auxiliary server, parameter information related to the first DNN, second DNN, third DNN, and fourth DNN can be sent to the video processing apparatus 100, and the video processing apparatus 100 can store the parameter information related to the first DNN, second DNN, third DNN, and fourth DNN. To generate a three-dimensional audio signal from a two-dimensional audio signal, the video processing apparatus 100 can update the parameters of the first DNN, second DNN, third DNN, and fourth DNN based on the parameter information related to the first DNN, second DNN, third DNN, and fourth DNN, and generate and output the three-dimensional audio signal by using the updated first DNN, second DNN, third DNN, and fourth DNN.
[0245] exist Figure 15 In, and reference Figure 14 The description may differ, and may also include a correction signal correction unit 1530, and the correction signal correction unit 1530 may be added to correct the first training correction signal 1308 using user parameters 1316 and output the corrected first training correction signal 1315 to the processing of the (4-3)DNN 823. Therefore, in Figure 15 In, with Figure 14Unlike other methods, training is performed by taking user parameters into account, and therefore, a three-dimensional audio signal can be generated and output that is corrected to further reflect the user's intent.
[0246] Figure 16 This is a diagram used to describe the process by which a user collects data for training using a user terminal 1610.
[0247] exist Figure 16 In this process, user 1600 can obtain the first training two-dimensional audio signal and the first training image signal by using the microphone and camera of user terminal 1610. At the same time, user 1600 can obtain the first training three-dimensional audio signal by separately installing a surround sound microphone 1620 on user terminal 1610 or by using the surround sound microphone 1620 included in user terminal 1620.
[0248] In this scenario, user terminal 1610 is an example of video processing apparatus 100, and user terminal 1610 can train a first DNN 220, a second DNN, and a third DNN 420 (including (2-1) DNN 421, (2-2) DNN 422, and a third DNN 423), (4-1) DNN 821, (4-2) DNN 822, and (4-3) DNN 823) based on training data (such as the obtained first training two-dimensional audio signal, first training image signal, and first training three-dimensional audio signal). Optionally, user terminal 1610 can send the training data to a device connected to user terminal 1610, such as an attached server. Examples of corresponding devices are training devices 1400 and 1500, and the first DNN 220, the second DNN, the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822, and the (4-3) DNN can be trained based on the training data. 823. The system can obtain parameter information for the trained first DNN 220, second DNN, third DNN 420, (4-1) DNN 821, (4-2) DNN 822, and (4-3) DNN 823, and can send this parameter information to the user terminal 1610. The user terminal 1610 can obtain the parameter information for the first DNN 220, second DNN, third DNN 420, (4-1) DNN 821, (4-2) DNN 822, and (4-3) DNN 823, and store the parameter information for the first DNN 220, second DNN, third DNN 420, and (4-1) DNN 821, (4-2) DNN 822, and (4-3) DNN 823. 821. Parameter information of DNN 822 and DNN 823 (4-2).
[0249] After this, the user terminal 1610 can obtain two-dimensional audio signals and image signals. User terminal 1610 can obtain pre-stored parameter information of the first DNN 220, second DNN, third DNN 420, (4-1)DNN 821, (4-2)DNN 822, and (4-3)DNN 823, update the parameters of the first DNN 220, second DNN, third DNN 420, (4-1)DNN 821, (4-2)DNN 822, and (4-3)DNN 823, obtain the updated parameter information of the first DNN 220, second DNN, third DNN 420, (4-1)DNN 821, (4-2)DNN 822, and (4-3)DNN 823, and use the first DNN 220, second DNN, third DNN 420, (4-1)DNN 821, (4-2)DNN 823, and (4-3)DNN 823. 822 and (4-3)DNN 823 generate and output three-dimensional audio signals from two-dimensional audio signals and image signals.
[0250] However, this disclosure is not limited thereto, and the user terminal 1610 is merely a simple training information collection device, and training data can be sent to the device, such as an additional server connected to the user terminal 1610 via a network. In this case, the corresponding devices could be examples of training devices 1400 and 1500 and video processing apparatus 100.
[0251] The corresponding device can obtain parameter information of the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823 based on training data, and train the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823. The parameter information of the trained first DNN 220, second DNN, third DNN 420, (4-1) DNN 821, (4-2) DNN 822, and (4-3) DNN 823 can be obtained. This parameter information can be sent to the user terminal 1610. Alternatively, the parameter information of the first DNN 220, second DNN, third DNN 420, (4-1) DNN 821, (4-2) DNN 822, and (4-3) DNN 823 can be obtained. The parameter information of DNN 822 and DNN 823, and the parameter information of DNN 220, DNN 220, DNN 420, DNN 821, DNN 822 and DNN 823 are stored in the user terminal 1610, or may be stored in the corresponding device or an attached database connected thereto, corresponding to the identifier of the user terminal 1610.
[0252] Subsequently, the user terminal 1610 can obtain two-dimensional audio signals and image signals. The user terminal 1610 can obtain the parameter information of the pre-stored first DNN 220, second DNN and third DNN 420, (4-1)DNN 821, (4-2)DNN 822 and (4-3)DNN 823, and can send the two-dimensional audio signals and image signals together with the parameter information of the first DNN 220, second DNN and third DNN 420, (4-1)DNN 821, (4-2)DNN 822 and (4-3)DNN 823 to the corresponding devices. The corresponding device can obtain the parameter information of the first DNN 220, the second DNN, the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822, and the (4-3) DNN 823 received from the user terminal 1610, update the parameters of the first DNN 220, the second DNN, the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822, and the (4-3) DNN 823, obtain the parameter information of the first DNN 220, the second DNN, the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822, and the (4-3) DNN 823, and, by using the first DNN 220, the second DNN, the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 823, and the (4-3) DNN 823, and... DNN 822 and DNN 823 (4-3) obtain a three-dimensional audio signal from the two-dimensional audio signal and image signal received from the user terminal 1610. Optionally, the user terminal 1610 can send the two-dimensional audio signal and image signal to a corresponding device. The corresponding device can obtain the parameter information of the pre-stored first DNN 220, second DNN and third DNN 420, DNN 821 (4-1), DNN 822 (4-2), and DNN 823 (4-3) to correspond to the identifier of the user terminal 1610, and obtain the three-dimensional audio signal from the two-dimensional audio signal and image signal received from the user terminal 1610 by using the first DNN 220, second DNN and third DNN 420, DNN 821 (4-1), DNN 822 (4-2), and DNN 823 (4-3).
[0253] Meanwhile, the training devices 1400 and 1500, which are connected to the user terminal 1610 via a network, can exist separately from the video processing device 100.
[0254] In this scenario, user terminal 1610 can send training data to training devices 1400 and 1500, and obtain parameter information for the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823. It can also obtain parameter information for the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823, and the parameter information for the first DNN 220, the second DNN and the third DNN 420, the (4-1) DNN 821, the (4-2) DNN 822 and the (4-3) DNN 823 previously obtained along with the two-dimensional audio signal and image signal. Furthermore, it can transfer the parameter information of the first DNN 220, the second DNN and the third DNN... The parameter information of the (4-1)DNN 821, (4-2)DNN 822 and (4-3)DNN 823 is sent to the video processing device 1000, thereby receiving the three-dimensional audio signal from the video processing device 100.
[0255] Figure 17 This is a flowchart describing a video processing method according to an embodiment.
[0256] In operation S1710, the video processing device 100 can generate multiple feature information for each time and frequency based on a first DNN by analyzing a video signal including multiple images.
[0257] During operation S1720, the video processing device 100 can extract a first height component and a first planar component corresponding to the motion of an object in the video from the video signal based on the second DNN.
[0258] During operation S1730, the video processing device 100 can extract a second planar component corresponding to the motion of the sound source in the audio from a first audio signal that does not have a height component, based on a third DNN.
[0259] During operation S1740, the video processing apparatus 100 can generate a second height component from the first height component, the first planar component, and the second planar component. In this case, the generated second height component can be the second height component itself, but is not limited to this, and can be information related to the second height component.
[0260] In operation S1750, the video processing apparatus 100 may output a second audio signal including a second height component based on feature information. This disclosure is not limited thereto, and the video processing apparatus 100 may output a second audio signal including a second height component based on feature information and information related to the second height component.
[0261] During operation S1760, the video processing device 100 can synchronize the second audio signal with the video signal and output the second audio signal.
[0262] Furthermore, the disclosed embodiments can be written into a computer-executable program, and the prepared program can be stored in a medium.
[0263] The medium can continuously store programs executable by a computer, or it can temporarily store programs for execution or download. Furthermore, the medium can be a single hardware device or a combination of several hardware devices, and is not limited to media directly connected to a specific computer system, and can exist in a distributed manner on a network. Examples of media can include: magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical recording media, such as CD-ROMs and DVDs; magneto-optical media, such as optical floppy disks; and media configured to store program instructions via ROM, RAM, flash memory, etc. Additionally, examples of other media can include recording or storage media managed by application stores that distribute applications, sites that provide or distribute various other software, and servers.
[0264] Although the technical concept of this disclosure has been described in detail above with reference to preferred embodiments, the technical concept of this disclosure is not limited to the above embodiments, and those skilled in the art can make various modifications and changes within the scope of the technical concept of this disclosure.
Claims
1. A video processing apparatus, comprising: A memory that stores at least one instruction; as well as At least one processor configured to execute the at least one instruction to: By analyzing video signals comprising multiple images based on a first deep neural network (DNN), multiple feature information regarding time and frequency is generated; Based on a second DNN, a first height component and a first planar component corresponding to the motion of an object in the video are extracted from the video signal. The second DNN includes a (2-1) DNN and a (2-2) DNN. The (2-1) DNN is used to extract N+M feature mapping information corresponding to the motion of an object in the horizontal direction relative to time in the video, where N and M are integers greater than or equal to 1. The (2-2) DNN is used to extract N+M feature mapping information corresponding to the motion of an object in the vertical direction relative to time in the video. Based on the third DNN, the second planar component corresponding to the motion of the sound source in the audio is extracted from the first audio signal; A second height component is generated based on the first height component, the first planar component, and the second planar component; Based on the multiple feature information, a second audio signal including the second height component is output; as well as The second audio signal is synchronized with the video signal, and the synchronized second audio signal and video signal are output.
2. The video processing apparatus of claim 1, wherein the at least one processor is further configured to execute the at least one instruction to: Synchronize the video signal with the first audio signal; By using a first DNN, M one-dimensional image feature mappings corresponding to the motion of objects in the video signal are generated, where M is an integer greater than or equal to 1; and By performing frequency-related tiling on the M one-dimensional image feature mapping information, multiple feature information related to time and frequency is generated, including the M image feature mapping information related to time and frequency.
3. The video processing apparatus of claim 1, wherein the at least one processor is further configured to execute the at least one instruction to: Synchronize the video signal with the first audio signal: By using a third DNN, N+M feature mapping information corresponding to the motion of the sound source in the horizontal direction in the audio is extracted from the first audio signal; Based on the N+M feature mapping information corresponding to the motion of the object in the horizontal direction in the video, the N+M feature mapping information corresponding to the motion of the object in the vertical direction in the video, and the N+M feature mapping information corresponding to the motion of the sound source in the horizontal direction in the audio, N+M correction mapping information with respect to time corresponding to the second height component is generated. as well as By performing frequency-related tiling on the N+M correction mapping information with respect to time, N+M correction mapping information with respect to time and frequency corresponding to the second height component is generated.
4. The video processing apparatus of claim 1, wherein the at least one processor is further configured to execute the at least one instruction to: Time and frequency information for the two channels is generated by performing a frequency conversion operation on the first audio signal; By using the (4-1)th DNN, N audio feature mapping information about time and frequency is generated from the time and frequency information used for the two channels, where N is an integer greater than or equal to 1; Based on the M image feature mapping information about time and frequency included in the multiple feature information about time and frequency, and the N audio feature mapping information about time and frequency, N+M audio and image integrated feature mapping information are generated; By using the (4-2)th DNN, a second frequency domain audio signal for the n-channel is generated from the N+M audio and image integrated feature mapping information, wherein, n is an integer greater than 2; By using the (4-3)th DNN, N+M correction mapping information corresponding to the N+M audio and image integrated feature mapping information and the second height component in terms of time and frequency are generated to produce the audio correction mapping information for the n channels; A corrected frequency-domain second audio signal for the n channels is generated by performing correction on the frequency-domain second audio signal for the n channels based on the audio correction mapping information for the n channels; and The second audio signal for the n-channel is output by performing an inverse frequency conversion on the frequency domain second audio signal used for the correction.
5. The video processing apparatus as claimed in claim 1, wherein, The first DNN is a DNN used to generate the multiple feature information of time and frequency; the second DNN is a DNN used to extract the first height component and the first planar component; and the third DNN is a DNN used to extract the second planar component. The at least one processor is further configured to execute the at least one instruction to train the first DNN, the second DNN, the third DNN, and the fourth DNN based on a comparison result between the first training two-dimensional audio signal and the first frequency domain training reconstructed three-dimensional audio signal reconstructed based on the first training image signal and the first frequency domain training three-dimensional audio signal obtained by frequency conversion of the first training three-dimensional audio signal, wherein the fourth DNN includes the (4-1) DNN, the (4-2) DNN, and the (4-3) DNN.
6. The video processing apparatus of claim 5, wherein the at least one processor is further configured to execute the at least one instruction to: The generation loss information is obtained by comparing the reconstructed 3D audio signal from the first frequency domain training with the 3D audio signal from the first frequency domain training. The parameters of the first DNN, second DNN, third DNN and fourth DNN are updated based on the generated loss information.
7. The video processing apparatus as claimed in claim 1, wherein, The first DNN is a DNN used to generate the multiple feature information of time and frequency; the second DNN is a DNN used to extract the first height component and the first planar component; and the third DNN is a DNN used to extract the second planar component. The at least one processor is further configured to execute the at least one instruction to train the first DNN, the second DNN, the third DNN, and the fourth DNN based on a comparison result between the first training two-dimensional audio signal, the first training image signal, and the frequency domain training reconstructed three-dimensional audio signal reconstructed based on user input parameter information and the first frequency domain training three-dimensional audio signal obtained by frequency conversion of the first training three-dimensional audio signal, wherein the fourth DNN includes the (4-1) DNN, the (4-2) DNN, and the (4-3) DNN.
8. The video processing apparatus of claim 7, wherein the at least one processor is further configured to execute the at least one instruction to: The generation loss information is obtained by comparing the reconstructed 3D audio signal trained in the frequency domain with the 3D audio signal trained in the first frequency domain. The parameters of the first DNN, second DNN, third DNN and fourth DNN are updated based on the generated loss information.
9. The video processing apparatus of claim 5, wherein the first training two-dimensional audio signal and the first training image signal are obtained from a portable terminal, and in, The first training 3D audio signal is obtained from the surround sound microphone of the portable terminal.
10. The video processing apparatus of claim 5, wherein parameter information of the first DNN, the second DNN, the third DNN, and the fourth DNN obtained as training results of the first DNN, the second DNN, the third DNN, and the fourth DNN is stored in the video processing apparatus or received from a terminal connected to the video processing apparatus.
11. A video processing method using a video processing apparatus, the video processing method comprising: By analyzing video signals comprising multiple images based on a first deep neural network (DNN), multiple feature information regarding time and frequency is generated; Extracting a first height component and a first planar component corresponding to the motion of an object in the video from the video signal based on the second DNN, the extraction of the first height component and the first planar component based on the second DNN includes: extracting N+M feature mapping information corresponding to the motion of an object in the horizontal direction relative to time in the video from the video signal by using the (2-1) DNN, where N and M are integers greater than or equal to 1; and extracting N+M feature mapping information corresponding to the motion of an object in the vertical direction relative to time in the video from the video signal by using the (2-2) DNN. Based on the third DNN, the second planar component corresponding to the motion of the sound source in the audio is extracted from the first audio signal; A second height component is generated based on the first height component, the first planar component, and the second planar component; Based on the multiple feature information, a second audio signal including the second height component is output; and The second audio signal is synchronized with the video signal, and the synchronized second audio signal and the video signal are output.
12. The video processing method of claim 11, wherein generating the plurality of feature information regarding time and frequency comprises: Synchronize the video signal with the first audio signal; By using the first DNN, M one-dimensional image feature mapping information corresponding to the motion of objects in the video from the video signal are generated, where M is an integer greater than or equal to 1; as well as By performing frequency-related tiling on the M one-dimensional image feature mapping information, multiple feature information related to time and frequency is generated, including the M image feature mapping information related to time and frequency.
13. The video processing method of claim 11, wherein extracting the second planar component based on the third DNN comprises: Synchronize the video signal with the first audio signal; By using the third DNN, N+M feature mapping information corresponding to the motion of the sound source in the horizontal direction in the audio is extracted from the first audio signal; as well as Generating the second height component based on the first height component, the first planar component, and the second planar component includes: Based on the N+M feature mapping information corresponding to the motion of objects in the horizontal direction in the video, the N+M feature mapping information corresponding to the motion of objects in the vertical direction, and the N+M feature mapping information corresponding to the motion of sound sources in the horizontal direction in the audio, N+M time-related correction mapping information corresponding to the second height component is generated; and By performing frequency-related tiling on the N+M correction mapping information with respect to time, N+M correction mapping information with respect to time and frequency corresponding to the second height component is generated.
14. The video processing method of claim 11, wherein outputting the second audio signal including the second height component based on the plurality of feature information comprises: Time and frequency information for the two channels is obtained by performing a frequency conversion operation on the first audio signal; By using the (4-1)th DNN, N audio feature mapping information about time and frequency is generated from the time and frequency information used for the two channels, where N is an integer greater than or equal to 1; Based on the M image feature mapping information about time and frequency included in the multiple feature information about time and frequency, and the N audio feature mapping information about time and frequency, N+M audio and image integrated feature mapping information are generated; By using the (4-2)th DNN, a second frequency domain audio signal for the n-channel is generated from the N+M audio and image integrated feature mapping information, where n is an integer greater than 2; By using the (4-3)th DNN, audio correction mapping information about the n channels corresponding to the second height component is generated from the N+M audio and image integrated feature mapping information; By performing correction on the frequency domain second audio signal for the n channels based on audio correction mapping information for the n channels, a corrected frequency domain second audio signal for the n channels is generated; and By performing an inverse frequency conversion on the corrected frequency domain second audio signal, a second audio signal for the n channels is output.
15. A computer-readable recording medium having a program recorded thereon, wherein the program is executed by a processor to perform the method of claim 11.
Citation Information
Patent Citations
Generating spatial audio using a predictive model
US20190306451A1