Signal processing device, signal processing method, and program

The signal processing device uses an AI model to process both acoustic and image signals, allowing for precise emphasis, suppression, or extraction of specific sounds from individual sources, addressing the challenge of handling multiple sound sources in the same frequency band.

WO2025134778A1PCT designated stage expired Publication Date: 2025-06-26SONY SEMICON SOLUTIONS CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/042996
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-05
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing signal processing technologies struggle to effectively emphasize or suppress specific sounds originating from individual sound sources, such as people, in a mixed acoustic environment.

Method used

A signal processing device and method that utilize an AI model to perform acoustic processing on both acoustic and image signals. This allows for the identification and separation of sound sources, enabling precise emphasis, suppression, or extraction of specific sounds from individual sources.

Benefits of technology

The solution enables high-quality noise removal and sound extraction by accurately identifying and processing sound sources, even when multiple sources emit sounds in the same frequency band.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024042996_26062025_PF_FP_ABST
    Figure JP2024042996_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A signal processing device according to the present technology is provided with an acoustic processing unit to which an acoustic signal obtained by sound pickup in a prescribed space and an image signal obtained by imaging in the prescribed space are input as data to be processed, and which performs acoustic processing in units of sound sources using an AI model.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing device, signal processing method, and program

[0001] The present technology relates to a signal processing device, a signal processing method, and a program for performing acoustic processing on an acoustic signal obtained by collecting sound.

[0002] Examples of acoustic processing for an audio signal include enhancement processing that emphasizes specific sounds, suppression processing that suppresses specific sounds, etc. Such processing is realized, for example, by using a filter that emphasizes or suppresses specific frequency bands (see Patent Document 1 below).

[0003] Japanese Patent Application Laid-Open No. 2006-39267

[0004] According to the technology of Patent Document 1, it is possible to emphasize and extract audio components, but it is difficult to emphasize or suppress audio for each sound source, specifically for each person.

[0005] The present technology has been made in consideration of the above circumstances, and aims to perform emphasis or suppression for each sound source.

[0006] The signal processing device according to the present technology includes an acoustic processing unit that receives, as processing target data, an acoustic signal obtained by collecting sound in a predetermined space and an image signal obtained by imaging the predetermined space, and performs acoustic processing for each sound source using an AI model. For example, by inputting not only the acoustic signal but also the image signal to the AI ​​model, inference processing as acoustic processing is performed.

[0007] The signal processing method according to the present technology involves a signal processing device performing acoustic processing on a sound source basis using an AI model with acoustic signals obtained by collecting sound in a specified space and image signals obtained by imaging the specified space as processing target data.

[0008] The program according to the present technology causes a processing device to execute a sound processing function for each sound source using an AI model, with an acoustic signal obtained by collecting sound in a predetermined space and an image signal obtained by imaging the predetermined space as processing target data. The signal processing device described above can also be realized by such a signal processing method or program.

[0009] 1 is a block diagram showing a configuration example of a signal processing device according to the present embodiment; FIG. 1 is a block diagram showing a first configuration example of a signal processing device; FIG. 2 is an explanatory diagram of learning of an AI model used in the first configuration example; FIG. 3 is a diagram of a modified example in which a part of the first configuration example is changed; FIG. 4 is a block diagram showing another configuration example of a signal processing device; FIG. 5 is a block diagram showing a fifth configuration example of a signal processing device; FIG. 6 is a block diagram showing a sixth configuration example of a signal processing device; FIG. 7 is a flowchart of a first processing example of a signal processing device; FIG. 8 is a flowchart of a second processing example of a signal processing device; FIG. 9 is a flowchart of a third processing example of a signal processing device; FIG. 10 is a flowchart of a fourth processing example of a signal processing device; FIG. 11 is a block diagram showing a configuration example of a signal processing device according to a second embodiment; FIG. 12 is a flowchart showing an example of processing executed by the signal processing device according to the second embodiment; FIG. 13 is a block diagram showing a configuration example of a signal processing device according to a third embodiment; FIG. 14 is a block diagram showing a configuration example of an unfolding device; 10A and 10B are flowcharts illustrating an example of processing executed by a signal processing device and an unfolding device in the third embodiment;

[0010] Hereinafter, with reference to the accompanying drawings, embodiments according to the present technology will be described in the following order: <1. Configuration of signal processing device> <1-1. First configuration example> <1-2. Second configuration example> <1-3. Third configuration example> <1-4. Fourth configuration example> <1-5. Fifth configuration example> <1-6. Sixth configuration example> <2. Processing examples> <2-1. First processing example> <2-2. Second processing example> <2-3. Third processing example> <2-4. Fourth processing example> <3. Application example> <4. Second embodiment> <5. Third embodiment> <6. Others> <7. Summary> <8. The present technology>

[0011] <1. Configuration of the signal processing device> The signal processing device 1 in this embodiment is a device that obtains desired acoustic data by performing predetermined acoustic processing on an acoustic signal Sa output from a sound collection unit 2 such as a microphone and an image signal Sg output from an imaging unit 3 such as a camera, as processing target data.

[0012] 1 shows an example of the configuration of a signal processing device 1. The signal processing device 1 includes a microphone array 2A as a sound collection unit 2 and a camera unit 3A as an imaging unit 3.

[0013] The signal processing device 1 further includes an acoustic processing unit 4 , an image detection and estimation processing unit 5 , an acoustic output unit 6 , and a display processing unit 7 .

[0014] The microphone array 2A is configured by arranging multiple microphones in an array and is capable of outputting a stereoscopically picked-up sound signal Sa. By arranging multiple microphones at different positions in a predetermined space, the microphone array 2A is capable of detecting differences in sound arrival time and amplitude between the microphones. By comparing the sound signals Sa output from each microphone, it is possible to identify the position and distance of the sound source.

[0015] The camera unit 3A is configured to include, for example, an image sensor for capturing a color image and is capable of outputting a color image signal Sg. The image sensor is configured to include, for example, a light receiving unit configured by a two-dimensional array of pixels sensitive to R (red), G (green), and B (blue) light.

[0016] The camera unit 3A includes an image signal processing unit (not shown), which performs various signal processing to output an image signal Sg based on luminance information of the received light.

[0017] The sound processing unit 4 performs various processes on the sound signal Sa to generate desired sound data. The desired sound data may be, for example, sound data in which the specific sound Ss is emphasized, sound data in which the specific sound Ss is suppressed, or sound data in which only the specific sound Ss is extracted.

[0018] Here, the specific sound Ss targeted for emphasis, suppression, or extraction by the sound processing unit 4 is a sound for each sound source. For example, if sound is emphasized for each frequency, speech sounds other than those of the specific person will also be emphasized or suppressed.

[0019] On the other hand, in the signal processing device 1 of this embodiment, emphasis, suppression, and extraction are performed on a sound source basis, so even if the speech sounds are in the same frequency band, it is possible to emphasize the speech sound of person A and suppress the speech sound of person B.

[0020] For this purpose, the sound processing unit 4 uses the image signal Sg itself acquired by the camera unit 3A, or information obtained as a result of applying a predetermined process to the image signal Sg.

[0021] The acoustic processing section 4 includes an array signal processing section 41 , an acoustic detection and estimation processing section 42 , and an enhancement and suppression processing section 43 .

[0022] The array signal processing unit 41 performs preparatory processing for subsequent processing, such as gain adjustment and A / D conversion processing, on each signal obtained from the multiple microphones included in the microphone array 2A.

[0023] The sound detection and estimation processing unit 42 realizes a sound source separation function F1, a sound event detection function F2, an abnormal sound / predetermined sound detection function F3, etc., using the sound signals Sa for each microphone supplied from the array signal processing unit 41. Only a part of each of these functions may function, or a plurality of functions may function in combination.

[0024] The sound source separation function F1 is a function for separating an acoustic signal Sa emitted from a specific sound source. By using this function, it is possible to emphasize, suppress, or extract only the speech of a specific person.

[0025] The acoustic event detection function F2 is a function for detecting an acoustic event based on the acoustic signal Sa. The acoustic event detection function F2 detects, for example, the occurrence of a stationary noise, such as the occurrence of an acoustic event such as "hand washing" or "road construction."

[0026] The abnormal sound / predetermined sound detection function F3 is a function for detecting predetermined sounds, abnormal sounds, and the like.

[0027] The detection and estimation results from the image detection and estimation processing unit 5 are used in the detection and estimation processes performed by the sound source separation function F1, the acoustic event detection function F2, and the abnormal sound / predetermined sound detection function F3.

[0028] Furthermore, the detection results and estimation results of the sound source separation function F1, the acoustic event detection function F2, and the abnormal sound / predetermined sound detection function F3 are supplied to the suppression processing unit 43.

[0029] The suppression processing unit 43 performs emphasis processing, suppression processing, extraction processing, etc. on the specific sound Ss using the detection results and estimation results of the sound detection / estimation processing unit 42. Note that the extraction processing of the specific sound Ss may be realized by the sound source separation function F1 of the sound detection / estimation processing unit 42. The sound signal Sa after being processed by the suppression processing unit 43 is output to the sound output unit 6 and the display processing unit 7.

[0030] The image detection and estimation processing unit 5 uses the image signal Sg supplied from the camera unit 3A to realize a person / object detection function F4 and a scene estimation function F5.

[0031] The person / object detection function F4 is a function that detects people and objects captured in an image based on the image signal Sg by image processing. Information about the detected people and objects is supplied to the sound detection / estimation processing unit 42 and used in the sound source separation function F1, the sound event detection function F2, and the abnormal sound / predetermined sound detection function F3.

[0032] The scene estimation function F5 is a function that estimates the scene in the space in which the microphone array 2A and the camera unit 3A are installed, i.e., the scene in which recording is taking place, based on the image signal Sg. Information about the estimated scene (scene specification information Ds) is provided to the sound detection / estimation processing unit 42 and is used in the sound source separation function F1, the sound event detection function F2, and the abnormal sound / predetermined sound detection function F3.

[0033] The scene specification information Ds, which is information about a scene, may be, for example, ID (Identification) information that specifies a scene, or other information such as a feature amount about the scene.

[0034] The acoustic output unit 6 outputs the acoustic signal Sa, which is the acoustic signal Sa supplied from the acoustic processing unit 4, with the specific sound Ss emphasized, suppressed, or extracted. The acoustic output unit 6 may be provided as a sound output unit such as a speaker, or as a data transmission unit that transmits the acoustic signal Sa to an external device.

[0035] The display processing unit 7 performs processing to generate a display image based on the image signal Sg output from the camera unit 3 A. The display processing unit 7 also performs highlighting and displaying of the sound source of the specific sound Ss that has been emphasized, suppressed, or extracted by the suppression processing unit 43. For example, when the speech of person A is emphasized, image processing is performed to visually indicate that the speech emitted by person A has been emphasized.

[0036] The display processing unit 7, for example, generates or selects an emphasis image to be superimposed on the subject that is the target of the acoustic processing, and further performs processing to superimpose the emphasis image on the subject on the display image.

[0037] In addition, the acoustic processing unit 4 uses an AI model to generate an acoustic signal Sa in which the specific sound Ss is emphasized, an acoustic signal Sa in which the specific sound Ss is suppressed, or an acoustic signal Sa in which only the specific sound Ss is extracted.

[0038] There are various possible configurations of AI models, input data, and learning methods for AI models. Some configuration examples are shown below.

[0039] <1-1. First Configuration Example> A first configuration example of the signal processing device 1 is shown in Fig. 2. Note that the display processing unit 7 is not shown in Fig. 2.

[0040] In the signal processing device 1 in the first configuration example, an image signal Sg output from a camera unit 3A is supplied to a scene estimation processing unit 5A. The scene estimation processing unit 5A is one aspect of a scene estimation function F5, which is one function of the image detection / estimation processing unit 5 shown in FIG.

[0041] The estimation result by the scene estimation processing section 5A, that is, the scene identification information Ds, is input to the sound processing section 4A. In the following description, a scene ID will be taken as an example of the scene identification information Ds.

[0042] The acoustic processing unit 4A includes a coefficient selection unit 44, a coefficient DB 45, and an acoustic AI processing unit 46A.

[0043] The coefficient selection unit 44 receives a scene ID from the scene estimation processing unit 5A, and selects a coefficient group corresponding to the scene ID from a plurality of coefficient groups stored in the coefficient DB 45.

[0044] Each coefficient set stored in the coefficient DB 45 is a collection of coefficients of the AI ​​model MA used in the acoustic AI processing unit 46A. For example, the first coefficient set stored in the coefficient DB 45 is a coefficient set of the AI ​​model MA1 that performs acoustic processing on an acoustic signal Sa picked up in a forest. The second coefficient set stored in the coefficient DB 45 is a coefficient set of the AI ​​model MA2 that performs acoustic processing on an acoustic signal Sa picked up in a crowd. The third coefficient set stored in the coefficient DB 45 is a coefficient set of the AI ​​model MA3 that performs acoustic processing on an acoustic signal Sa picked up in a factory.

[0045] Note that the scene ID may be assigned taking into consideration not only the location but also the season, time, etc. For example, "in the winter forest" and "in the summer forest", or "in the morning forest" and "in the night forest", may be assigned different scene IDs.

[0046] The coefficient selection unit 44 selects one group of coefficients from the coefficient DB 45, thereby selecting the AI ​​model MA to be used in the acoustic AI processing unit 46A.

[0047] The acoustic AI processing unit 46A performs acoustic processing using the coefficient group selected by the coefficient selection unit 44. In the following description, acoustic processing using an AI model will be referred to as "AI acoustic processing."

[0048] The acoustic AI processing unit 46A inputs to the selected AI model MA the acoustic signal Sa obtained by sound collection by the microphone array 2A, information for selecting the subject to be processed, and information for selecting the processing content (emphasis, suppression, extraction, etc.).

[0049] The AI ​​model MA outputs the audio signal Sa desired by the user based on the input information, and the output audio signal Sa is supplied to the audio output unit 6 and the like.

[0050] Although FIG. 2 shows an example in which the acoustic processing unit 4A includes the coefficient selection unit 44 and the coefficient DB 45, the coefficient selection unit 44 and the coefficient DB 45 may be provided in the scene estimation processing unit 5A.

[0051] The learning of such AI models MA (MA1, MA2, MA3, etc.) is shown in FIG.

[0052] In training the AI ​​model MA, the feature quantities of the sound signal Sa picked up in a specific scene are input. In addition, when the AI ​​model MA is to have multiple functions such as emphasizing, suppressing, or extracting a specific sound Ss, or when the AI ​​model MA is to have a function for selecting a sound source to be processed, instruction information regarding the processing content, information for identifying the sound source, etc. are input to the AI ​​model MA for training. This is also the case in training other AI models described later.

[0053] During learning, the acoustic signal Sa output from the AI ​​model MA is compared with the acoustic signal Sa as training data. Learning of the AI ​​model MA is performed using a loss function and an error function based on the comparison results.

[0054] The AI ​​model MA can be obtained as AI models MA1, MA2, MA3, etc. specialized for each scene by specializing the features of the acoustic signal Sa input during learning for each scene.

[0055] Since the AI ​​model MA is specialized for a particular scene, it is relatively small in size, which makes it possible to reduce the amount of calculation required for the inference process and the size of the storage area required to deploy the AI ​​model MA.

[0056] 3, the number of nodes in each layer is an example. Regarding the acoustic signal Sa input to the AI ​​model MA, information for each frequency band may be input to different nodes in the input layer, information divided by time may be input to different nodes in the input layer, or the acoustic signal Sa for each microphone in the microphone array 2A may be input to different nodes in the input layer.

[0057] Next, a description will be given of a modified example of configuration example 1. In configuration example 1, an example has been shown in which an audio signal Sa from which a specific sound Ss has been removed, an audio signal Sa from which a specific sound Ss has been emphasized, or an audio signal Sa from which a specific sound Ss has been extracted is output from the AI ​​model MA.

[0058] Other examples of the method for suppressing the specific sound Ss are also conceivable, which will be specifically described with reference to FIG.

[0059] The acoustic processing unit 4A includes an adder 47 in addition to a coefficient selection unit 44, a coefficient DB 45, and an acoustic AI processing unit 46A.

[0060] The acoustic AI processing unit 46A performs AI acoustic processing using the coefficient group selected by the coefficient selection unit 44. The acoustic signal Sa output from the acoustic AI processing unit 46A is an acoustic signal Sa emitted from a sound source to be suppressed.

[0061] The adder 47 subtracts the acoustic signal Sa generated by the acoustic AI processing unit 46A from the acoustic signal Sa input from the microphone array 2A, thereby generating an acoustic signal Sa from which the specific sound Ss has been removed.

[0062] The acoustic processing unit 4A shown in FIG. 4 generates the specific sound Ss using the coefficient (AI model MA) selected by the coefficient selection unit 44, but as shown in FIG. 5, the specific sound Ss may also be generated using the acoustic signal Sa output from the microphone array 2A.

[0063] For example, when generating an audio signal Sa from which the "voice of cicadas" has been removed from the audio signal Sa output from the microphone array 2A, in the example shown in Figure 4, the voice of a certain type of cicada is generated using the AI ​​model MA selected by the coefficient selection unit 44.

[0064] 5, the sounds of a specific type of cicada can be generated by further inputting the sound signal Sa output from the microphone array 2A to the AI ​​model MA selected by the coefficient selection unit 44. This makes it possible to generate a sound signal Sa from which the sounds of the cicada species that are actually heard have been removed.

[0065] 5, the acoustic AI processing unit 46A may execute AI acoustic processing to extract a specific sound Ss from the acoustic signal Sa supplied from the microphone array 2A using the AI ​​model MA selected by the coefficient selection unit 44. The addition unit 47 generates a target acoustic signal Sa by subtracting the specific sound Ss extracted by the AI ​​acoustic processing from the original acoustic signal Sa.

[0066] <1-2. Second Configuration Example> A second configuration example of the signal processing device 1 is shown in Fig. 6. Note that, similar to Fig. 2, the display processing unit 7 is not shown in Fig. 6.

[0067] In the signal processing device 1 in the second configuration example, an image signal Sg output from a camera unit 3A is supplied to a scene estimation processing unit 5A. The scene estimation processing unit 5A represents a scene estimation function F5, which is one function of the image detection / estimation processing unit 5 shown in FIG.

[0068] The scene estimation processing unit 5A estimates a scene through image processing and obtains a scene ID. The scene ID for the scene estimated by the scene estimation processing unit 5A is input to the audio processing unit 4B. In other words, the scene ID is not used to select a coefficient set for the AI ​​model.

[0069] The acoustic processing unit 4B includes an acoustic AI processing unit 46B that is capable of outputting an acoustic signal Sa in which the specific sound Ss is emphasized, an acoustic signal Sa in which the specific sound Ss is suppressed, or an acoustic signal Sa in which the specific sound Ss is extracted, without switching the AI ​​model based on the scene ID.

[0070] Here, the AI ​​model used in the acoustic AI processing unit 46B is referred to as the AI ​​model MB. Learning of the AI ​​model MB is shown in FIG.

[0071] In training the AI ​​model MB, the features of the audio signal Sa picked up in a specific scene and the scene ID for the specific scene are input. That is, the features of the audio signal Sa and the scene ID are input to the input layer of the AI ​​model MB.

[0072] Then, by learning combinations of specific scenes and scene IDs for the number of scenes, it is possible to generate an AI model MB corresponding to the acoustic signals Sa picked up in various scenes.

[0073] In addition, if it is desired to be able to select multiple processes, such as emphasis, suppression, and extraction, as the processing content for a specific sound Ss, or if it is desired to be able to select the sound source to be processed, then, as in the first configuration example, instruction information regarding the processing content and information identifying the sound source can be input to the AI ​​model MB and it can be made to learn.

[0074] The AI ​​model MB is trained using, for example, a loss function or an error function.

[0075] <1-3. Third Configuration Example> A third configuration example of the signal processing device 1 is shown in Fig. 8. Note that in Fig. 8, the display processing unit 7 is omitted from the illustration, as in Fig. 2.

[0076] The signal processing device 1 in the third configuration example includes a scene estimation AI processing unit 5B. The scene estimation AI processing unit 5B inputs the image signal Sg output from the camera unit 3A to an AI model MC1, performs scene estimation, and outputs a scene ID. The scene estimation AI processing unit 5B is one aspect of the scene estimation function F5, which is one function of the image detection / estimation processing unit 5 shown in FIG. 1, and performs processing using an AI model.

[0077] The scene ID obtained by the estimation process of the scene estimation AI processing unit 5B is supplied to the audio AI processing unit 46B of the audio processing unit 4B, and is used for audio processing.

[0078] The learning of the AI ​​model MC1 is shown in Fig. 9. Note that Fig. 9 shows both the AI ​​model MB and the AI ​​model MC1.

[0079] First, in the training of the AI ​​model MC1, input information for scene estimation, specifically the feature quantities of the image signal Sg, is input to the input layer.

[0080] The output layer of the AI ​​model MC1 outputs, for example, a scene ID. The AI ​​model MC1 learns by comparing the scene ID output from the output layer with the correct scene ID as training data.

[0081] In training the AI ​​model MB, unlike the example shown in Fig. 7, the features of the audio signal Sa that does not specify a scene are input to the input layer. In the example shown in Fig. 7, the audio features input to the input layer are paired with scene information. Specifically, when a scene ID indicating "in the forest" is input to the input layer, the features of the audio signal Sa picked up in "in the forest" are input to the input layer.

[0082] 9, there is no need to input the feature quantities of the audio signal Sa that matches the scene ID to the input layer of the AI ​​model MB. That is, feature quantities based on the image signal Sg acquired in a specified space are input to the input layer of the AI ​​model MC1 as "scene estimation input information," and feature quantities based on the audio signal Sa acquired in the same space during the same period are input to the input layer of the AI ​​model MB. In this way, the AI ​​models MB and MC1 are trained.

[0083] It should be noted that AI model MC1 and AI model MB may be linked together to conduct learning as a single AI model (see Figure 9). In this case, the output of AI model MC1 is directly linked to the input layer of AI model MB, and the two AI models MC1 and MB perform inference processing as a single unit.

[0084] By comparing the acoustic signal Sa output from the AI ​​model MB with the acoustic signal Sa as teacher data, the learning of the AI ​​model MC1 and the AI ​​model MB proceeds simultaneously.

[0085] In addition, if it is desired to be able to select the processing content for a specific sound Ss from among emphasis, suppression, extraction, etc., or if it is desired to be able to select the sound source to be processed, it is sufficient to input instruction information regarding the processing content and information identifying the sound source into the AI ​​model MB and have it learn, as in the first configuration example.

[0086] 1-4. Fourth Configuration Example A fourth configuration example of the signal processing device 1 is shown in Fig. 10. Note that, similar to Fig. 2, the display processing unit 7 is not shown in Fig. 10.

[0087] The signal processing device 1 in the fourth configuration example differs from the third configuration example in that both the acoustic signal Sa from the microphone array 2A and the image signal Sg from the camera unit 3A are input to the scene estimation AI processing unit 5B.

[0088] The scene estimation AI processing unit 5B inputs the features of the acoustic signal Sa and the image signal Sg into the AI ​​model MC2 and performs inference processing to obtain a scene ID.

[0089] The scene ID obtained by the AI ​​model MC2 is supplied to the audio AI processing unit 46B of the audio processing unit 4B and is used for the AI ​​audio processing of the AI ​​model MB.

[0090] The learning of the AI ​​model MC2 is similar to that already described with reference to Fig. 9. However, the input layer of the AI ​​model MC2 receives not only the feature quantities of the image signal Sg but also the feature quantities of the audio signal Sa as input information for scene estimation.

[0091] <1-5. Fifth Configuration Example> The fifth configuration example of the signal processing device 1 is an example in which the image detection and estimation processing unit 5 does not perform person and object detection processing or scene estimation processing.

[0092] The signal processing device 1 according to the fifth configuration example is shown in Fig. 11. Fig. 11 corresponds to Fig. 1.

[0093] In this configuration example, a person / object detection function F4 and a scene estimation function F5 are provided in the sound processing unit 4. An image signal Sg output from the camera unit 3A is supplied to a sound detection / estimation processing unit 42 and a display processing unit 7.

[0094] A configuration for realizing such a signal processing device 1 is shown in Fig. 12. Note that, similar to Fig. 2, the display processing unit 7 is not shown in Fig. 12.

[0095] In the signal processing device 1 in the fifth configuration example, both the acoustic signal Sa from the microphone array 2A and the image signal Sg from the camera unit 3A are input to an acoustic AI processing unit 46C of the acoustic processing unit 4C.

[0096] The acoustic AI processing unit 46C performs AI acoustic processing using the AI ​​model MD to obtain the desired acoustic signal Sa.

[0097] FIG. 13 shows the learning of the AI ​​model MD used in the acoustic AI processing unit 46C.

[0098] As shown in FIG. 13, the AI ​​model MD is configured by connecting an intermediate layer in which features about a scene are extracted and an intermediate layer for performing acoustic processing.

[0099] In training the AI ​​model MD, the features of an audio signal Sa picked up in a specific space are input to an input layer for AI audio processing, and the scene ID of a specific scene from an image signal Sg captured in the same space is input to the input layer to extract features for the scene.

[0100] Then, by inputting the sound signals Sa collected in various scenes and the image signals Sg captured in various scenes into the input layer, an AI model MD corresponding to various scenes can be generated.

[0101] In addition, if it is desired to be able to select multiple processes, such as emphasis, suppression, and extraction, as the processing content for a specific sound Ss, or if it is desired to be able to select the sound source to be processed, then, as in the first configuration example, instruction information regarding the processing content and information identifying the sound source can be input to the AI ​​model MD and the AI ​​model MD can be made to learn.

[0102] <1-6. Sixth Configuration Example> The sixth configuration example of the signal processing device 1 is an example of a smart speaker that detects the user's line of sight, performs voice recognition in accordance with the line of sight detection result, and responds by voice. Therefore, the signal processing device 1 does not need to be provided with the display processing unit 7.

[0103] As shown in FIG. 14, the signal processing device 1 includes an acoustic processing unit 4A having an acoustic AI processing unit 46A that performs AI acoustic processing using an acoustic signal Sa output from a microphone array 2A.

[0104] The signal processing device 1 includes a voice recognition processing unit 8 and a response generation processing unit 9. The voice recognition processing unit 8 receives an audio signal Sa in which a specific sound Ss has been emphasized, suppressed, or extracted by the audio processing unit 4A.

[0105] An image signal Sg is input from the camera unit 3A to the image detection and estimation processing unit 5C.

[0106] The image detection and estimation processing section 5C includes a person and object detection processing section 5Ca, a scene estimation processing section 5Cb, and a gaze detection section 5Cc.

[0107] The person / object detection processing unit 5Ca is a processing unit that realizes the function of the person / object detection function F4 shown in FIG. 1 using the image signal Sg.

[0108] The scene ID for the scene identified by the person / object detection processing unit 5Ca is supplied to the coefficient selection unit 44, which selects the optimal set of coefficients for the AI ​​model and realizes appropriate AI acoustic processing in the acoustic AI processing unit 46A.

[0109] The scene estimation processing unit 5Cb is a processing unit that realizes the function of the scene estimation function F5 shown in FIG. 1 using the image signal Sg.

[0110] The gaze detection unit 5Cc performs image processing using the image signal Sg and the person detection result of the scene estimation processing unit 5Cb, and performs processing to detect a user directing their gaze toward the signal processing device 1 serving as a smart speaker.

[0111] The gaze information in the gaze detection unit 5Cc or the user information indicating that the gaze is directed toward the signal processing device 1 is supplied to the voice recognition processing unit 8.

[0112] The speech recognition processing unit 8 recognizes a predetermined wake-up word (WuW: Wake up Word) from the acoustic signal Sa in which noise has been suppressed by the acoustic AI processing unit 46A, and recognizes the subsequent voice instruction from the user. The recognition result by the speech recognition processing unit 8 is provided to the response generation processing unit 9.

[0113] Here, the voice recognition processing unit 8 lowers the threshold for the voice of the user who is directing his / her gaze, based on the user information provided by the gaze detection unit 5Cc. Specifically, by lowering the threshold for recognizing a word in the spoken voice as the activation word only for that user, the spoken voice of that user is more likely to be recognized as the activation word.

[0114] The response generation processing unit 9 receives information in which the user's speech following the activation word has been converted into text from the speech recognition processing unit 8, identifies the context (meaning) of the content of the text by performing natural language processing or the like, and generates a response sentence in accordance with the context. At this time, the response generation processing unit 9 may obtain a response sentence (proposed answer) in response to the user's speech via a communication network.

[0115] The generated or acquired response sentence is output as a mechanical voice from a sound output unit (not shown).

[0116] 14 , the audio processing unit 4A may perform AI audio processing based on the detection result of gaze detection by the gaze detection unit 5Cc. For example, when person A and person B are present in the same space, if the detection result of the gaze detection unit 5Cc indicates that person A is speaking with his or her gaze directed toward the signal processing device 1, the audio AI processing unit 46A of the audio processing unit 4A may perform processing to emphasize or extract the speech of person A, or processing to suppress the speech of person B, noise, etc. This can further increase the likelihood that the speech of the person whose gaze has been detected will be recognized as the activation word.

[0117] Here, an example has been shown in which the signal processing device 1 is equipped with a person / object detection processing unit 5Ca and a gaze detection unit 5Cc in addition to the configuration shown in Figure 2, but the signal processing device 1 shown in Figures 4, 5, 6, 8, 10, 12, etc. may also be equipped with the person / object detection processing unit 5Ca and the gaze detection unit 5Cc.

[0118] 2. Processing Example An example of processing executed by the signal processing device 1 in the various configuration examples described above will be described.

[0119] <2-1. First Processing Example> The first processing example is shown in Fig. 15. Note that the example shown in Fig. 15 is an example of processing executed by the signal processing device 1 having the configuration shown in Fig. 2. Furthermore, the processing examples shown in Fig. 15 and the subsequent figures are, for example, processing executed in one frame period in imaging.

[0120] In step S101, the microphone array 2A and the camera unit 3A of the signal processing device 1 perform sound collection processing and image capture processing.

[0121] In step S102, the scene estimation processing unit 5A of the signal processing device 1 performs scene estimation processing. This processing is realized by processing an image obtained based on the image signal Sg output from the camera unit 3A, for example.

[0122] In step S103, the coefficient selection unit 44 of the signal processing device 1 selects one of the coefficient groups from the coefficient DB 45 according to the scene estimation information.

[0123] In step S104, the audio AI processing unit 46A of the signal processing device 1 performs AI audio processing, thereby outputting an audio signal Sa in which the specific sound Ss is emphasized, an audio signal Sa in which the specific sound Ss is suppressed, or an extracted audio signal Sa.

[0124] In step S105, the sound output unit 6 and the display processing unit 7 of the signal processing device 1 perform sound output processing and display processing. As a result, the sound signal Sa obtained as a result of the AI ​​sound processing is output from the sound output unit (not shown). Furthermore, an image based on the image signal Sg is displayed on the display unit (not shown), and an emphasis image is superimposed on the sound source that is the source of the specific sound Ss.

[0125] The signal processing device 1 shown in FIG. 6 can be realized by executing the processes shown in FIG. 15 except for step S103.

[0126] Moreover, the signal processing device 1 shown in FIG. 8 can be realized by executing the process of step S102 as AI processing and omitting the process of step S103 among the processes shown in FIG.

[0127] Furthermore, the signal processing device 1 shown in FIG. 10 can be realized by performing scene estimation using not only the image signal Sg but also the acoustic signal Sa in step S102.

[0128] <2-2. Second Processing Example> A second processing example is shown in Fig. 16. Note that the example shown in Fig. 16 is an example of processing executed by the signal processing device 1 having the configuration shown in Fig. 4 or 5.

[0129] Note that the same steps as those shown in FIG. 15 are given the same step numbers and their explanations are omitted where appropriate.

[0130] The acoustic processing unit 4A of the signal processing device 1 performs AI acoustic processing in step S104, which involves generating an acoustic signal Sa of the specific sound Ss using the AI ​​model MA to which the coefficients selected in step S103 have been applied.

[0131] In step S111, the adder 47 of the signal processing device 1 performs a process of subtracting the sound signal Sa for the specific sound Ss generated in step S104 from the sound signal Sa acquired by the microphone array 2A.

[0132] As a result, in the sound output process of step S105, sound is output based on the acoustic signal Sa from which the specific sound Ss has been removed or suppressed.

[0133] The signal processing device 1 shown in FIG. 5 can be realized by using the acoustic signal Sa acquired by the microphone array 2A when generating the specific sound Ss in the AI ​​acoustic processing in step S104.

[0134] <2-3. Third Processing Example> A third processing example is shown in Fig. 17. Note that the example shown in Fig. 17 is an example of processing executed by the signal processing device 1 having the configuration shown in Fig. 12.

[0135] In step S101, the signal processing device 1 performs a sound collection process using the microphone array 2A and an image capture process using the camera unit 3A.

[0136] In step S104, the signal processing device 1 inputs both the acoustic signal Sa and the image signal Sg to the AI ​​model MD to obtain the desired acoustic signal Sa.

[0137] The acoustic signal Sa obtained here is provided to the acoustic output unit 6, which outputs the sound in step S105.

[0138] <2-4. Fourth Processing Example> A fourth processing example is shown in Fig. 18. Note that the example shown in Fig. 18 is an example of processing executed by the signal processing device 1 having the configuration shown in Fig. 14.

[0139] Note that the same steps as those shown in FIG. 15 are given the same step numbers and their explanations are omitted where appropriate.

[0140] In step S121, the person / object detection processing unit 5Ca of the signal processing device 1 performs person detection processing.

[0141] The gaze detection unit 5Cc of the signal processing device 1 performs gaze detection processing in step S122, thereby detecting a user directing their gaze toward the signal processing device 1. The detection result is supplied to, for example, the voice recognition processing unit 8 or the sound processing unit 4A.

[0142] The sound processing unit 4A of the signal processing device 1 performs AI sound processing in step S104. In this processing, the result of gaze detection by the gaze detection unit 5Cc may be used.

[0143] In step S123, the voice recognition processing unit 8 of the signal processing device 1 performs a threshold adjustment process, which relaxes the voice recognition threshold for a user who is directing their gaze, making it easier to recognize, for example, an activation word.

[0144] In step S124, the voice recognition processing unit 8 of the signal processing device 1 performs voice recognition processing on the user's voice, thereby converting the user's voice into text.

[0145] In step S125, the response generation processing unit 9 of the signal processing device 1 analyzes the text data received from the voice recognition processing unit 8 and performs processing to identify the context.

[0146] In step S126, the response generation processing unit 9 of the signal processing device 1 performs processing to generate a response to the identified context. The response sentence acquired in step S126 may be obtained via a communication network.

[0147] In step S127, the sound output unit of the signal processing device 1 outputs the response sentence generated by the response generation processing unit 9 in mechanical voice. Note that the process of step S127 may be replaced by a process of displaying the response sentence on a screen.

[0148] 3. Application Examples Application examples of the various configurations described above will be described.

[0149] For example, the signal processing device 1 may be a camera device having a zoom microphone function. The zoom microphone function is a function that, when a zoom-in operation is performed using the zoom function of the camera device, emphasizes or extracts and collects sound such as speech from a subject captured within the angle of view, for example, the subject captured largest in the image, as the sound source.

[0150] This allows for sound capture processing that makes it easier to hear the speech of the person being filmed, so when the recorded video data is played back, the playback sound has a high affinity with the situation of the subject within the field of view. Also, since the emphasis is on sounds from sound sources located within the field of view rather than sounds from sources outside the field of view, it is advantageous in that it makes it easier to understand what is happening in the video later.

[0151] The signal processing device 1 may also be a camera / microphone device used in a video conference. For example, when two teams in remote locations hold a video conference, one signal processing device 1 is placed in one of the teams. In this case, multiple people will appear in the image captured by one camera unit 3A and transmitted.

[0152] The signal processing device 1 superimposes the above-mentioned emphasized image on a person who is speaking among multiple people appearing in an image, and also performs AI audio processing to suppress the voices of other people and noise so that the speaking voice of the person in question can be more easily heard, thereby improving the quality of the audio in a video conference.

[0153] The signal processing device 1 may be used to emphasize a specific speech sound among multiple speech sounds emitted from a television receiver. For example, the signal processing device 1 generates an audio signal Sa obtained by performing emphasis processing, suppression processing, etc. on a sound source (performer) based on image analysis of the subject (performer) displayed on the television receiver, as playback data to be played back from headphones or earphones serving as an audio playback unit. By playing back such audio signal Sa through headphones or the like, audio playback can be performed in which the speech sound of a specific person is emphasized.

[0154] Furthermore, the signal processing device 1 may be used to improve the performance of a hearing aid. For example, the signal processing device 1 is configured to include a sound collection and output device that is worn on the ear and has a microphone array 2A that collects sound and an acoustic output unit 6 that outputs sound, and eyewear that has a camera unit 3A that captures images.

[0155] The sound processing unit 4 in the sound collection and output device performs AI sound processing using the results of identifying the target person the user is gazing at through the eyewear, making it possible to emphasize or extract only the speech of the person being watched. This can improve the performance of the hearing aid. Note that this can also be achieved by replacing the eyewear with a mobile terminal device such as a smartphone.

[0156] The signal processing device 1 may also be used as various diagnostic devices, such as a device used for water leakage diagnosis and an electronic stethoscope.

[0157] In the case of the signal processing device 1 used for water leakage diagnosis, a process for removing noise in the same frequency band as the sound of water and a process for enhancing the sound of water are executed as AI sound processing in the sound processing unit 4. In this case, by identifying the scene based on the image captured by the camera unit 3A, it is possible to perform effective noise removal.

[0158] Furthermore, with an electronic stethoscope, noise caused by body movement can be detected as an acoustic event from the patient's body movements captured by the camera unit 3A using the acoustic event detection function F2, and input into AI acoustic processing, thereby effectively removing noise caused by body movement.

[0159] 4. Second Embodiment A signal processing device 1A according to a second embodiment will be described. The signal processing device 1A according to the second embodiment is capable of accepting an operation to designate a subject as a sound source that emits a specific sound Ss.

[0160] 1, the signal processing device 1A includes an operation unit 10 (see FIG. 19) in addition to the units shown in FIG.

[0161] Based on the detection result of the selection operation of the subject, the suppression processing unit 43 performs emphasis processing, suppression processing, or extraction processing on the sound signal Sa having the subject as the sound source. To this end, the operation unit 10 may detect not only the selection operation of the subject but also the selection operation for selecting the processing content of the sound processing.

[0162] Various selection operations by the user can be considered using the operation unit 10. For example, in the signal processing device 1A, a subject may be selected by tapping on the subject captured on the image. In this case, the operation unit 10 may be realized by providing a touch sensor on a display unit (not shown).

[0163] Furthermore, in the signal processing device 1A, a subject may be selected based on the line of sight of a user. In this case, for example, in a mode in which a user wears a display unit such as AR (Augmented Reality) goggles or AR glasses, the line of sight of the user may be detected by a camera further provided in the device such as the AR goggles.

[0164] Furthermore, in the signal processing device 1A, a subject may be selected by an instruction given by a user's voice, in which case the operation unit 10 may be realized as a microphone.

[0165] Alternatively, in the signal processing device 1A, a subject to be subjected to acoustic processing may be selected by a user's gesture operation. In this case, the operation unit 10 may be realized, for example, as a camera that detects the user's operation. The user performs a gesture operation toward a camera that is different from the camera unit 3A and is oriented so as to capture an image of the user, and the signal processing device 1A detects the gesture operation by performing image processing on the image captured by the camera.

[0166] Although FIG. 19 shows an example in which the operation unit 10 is provided in the configuration shown in FIG. 1, the operation unit 10 may be provided in the configuration shown in FIG.

[0167] An example of processing executed by the signal processing device 1A in the second embodiment is shown in Fig. 20. Note that the processing example shown in Fig. 20 is a combination of the processing example shown in Fig. 15 and the features of this embodiment. However, the present invention is not limited to this, and the features of this embodiment may also be combined with the processing examples shown in Fig. 16, Fig. 17, and Fig. 18.

[0168] Furthermore, the same steps as those shown in FIG. 15 are given the same step numbers and the description thereof will be omitted as appropriate.

[0169] In step S101, the microphone array 2A and the camera unit 3A of the signal processing device 1A perform sound collection processing and image capture processing.

[0170] In step S102, the image detection / estimation processing unit 5 (for example, the scene estimation processing unit 5A) of the signal processing device 1A performs a scene estimation process.

[0171] In step S103, the sound detection / estimation processing unit 42 or the suppression processing unit 43 of the signal processing device 1A selects one coefficient group from the coefficient DB 45 according to the scene estimation information.

[0172] The operation unit 10 of the signal processing device 1A determines whether an operation to select a specific subject has been detected in step S131. If it is determined that an operation to select a specific subject has been detected (step S131: Yes), the operation unit 10 supplies the detection result to the suppression processing unit 43 (for example, the acoustic AI processing unit 46A).

[0173] In this case, the suppression processing unit 43 (acoustic AI processing unit 46A) of the signal processing device 1A performs AI acoustic processing taking into account the selection result in step S132, specifically, outputs an acoustic signal Sa in which the specific sound Ss whose sound source is the subject being operated by the user is emphasized, a suppressed acoustic signal Sa, or an extracted acoustic signal Sa.

[0174] On the other hand, if it is determined that an operation to select a specific subject has not been detected (step S131: No), the suppression processing unit 43 (acoustic AI processing unit 46A) of the signal processing device 1A performs AI acoustic processing in step S104.

[0175] In this way, the processing content in the suppression processing unit 43 differs depending on the result of the determination processing in step S131.

[0176] 5. Third Embodiment A signal processing device 1B in the third embodiment is regarded as a recording device that records acoustic data and image data about a subject.

[0177] The acoustic data and image data recorded by the signal processing device 1B are considered to be objectified (packaged) as data for each sound source and used by a device downstream of the signal processing device 1B. In the following description, the acoustic data and image data collected for each sound source will be referred to as "sound source unit data Dss."

[0178] The acoustic data in the sound source unit data Dss is, for example, an acoustic signal Sa emitted from the sound source extracted from the acoustic signal Sa picked up by the microphone array 2A.

[0179] Furthermore, the sound source unit data Dss may include not only data for each subject such as a person or an object, but also object data, in other words, scene data Dsc, which is a set of the sound signal Sa remaining after removing all sound signals Sa whose sound source is a subject detected in the recording environment, and a background image from which the subject has been removed. That is, if subjects A, B, and C are present in the recording environment, the recording process by the signal processing device 1B generates sound source unit data Dss for subject A, sound source unit data Dss for subject B, sound source unit data Dss for subject C, and scene data Dsc derived from the environment in which a sound source cannot be identified.

[0180] The signal processing device 1B may output the recorded sound source unit data Dss in a streaming manner or as a file.

[0181] An example configuration of a signal processing device 1B is shown in Fig. 21. Note that blocks that have already been explained using Fig. 1 etc. are given the same reference numerals and explanations thereof will be omitted where appropriate. 1B includes a microphone array 2A as a sound collection unit 2, a camera unit 3A as an imaging unit 3, an acoustic processing unit 4, an image detection / estimation processing unit 5, and a multiplexing processing unit 11.

[0182] The acoustic processing unit 4 includes an array signal processing unit 41 and an acoustic detection and estimation processing unit 42. The acoustic detection and estimation processing unit 42 mainly generates or separates an acoustic signal Sa for each sound source by a sound source separation function F1.

[0183] The multiplexing processing unit 11 performs multiplexing processing (objectification processing) to combine, for each sound source, the acoustic signal Sa for each sound source provided from the acoustic detection / estimation processing unit 42 and the image signal Sg for each sound source provided from the image detection / estimation processing unit 5.

[0184] The multiplexing process by the multiplexing processor 11 generates the sound source unit data Dss (including the scene data Dsc) described above.

[0185] The multiplexing processing unit 11 further performs processing such as converting the generated sound source unit data Dss into a data format for streaming distribution and encoding processing for file output.

[0186] The sound source unit data Dss may further include metadata in addition to the acoustic data and image data. The metadata is, for example, analysis information about the sound source, and may include information on the results of semantic segmentation of the sound source, and if the sound source is a person, estimated information about the gender, age, height, and posture of the person.

[0187] The sound source unit data Dss generated in the signal processing device 1B is used in a subsequent device, for example, the development device 100.

[0188] The expansion device 100 is a device that generates audio data and image data to be played back on a device such as VR (Virtual Reality) goggles, rearranges audio images and images based on the respective sound source unit data Dss in a virtual space, applies audio processing according to the user's viewing position, and then generates data for viewing.

[0189] The playback data, such as audio data and image data, generated by the unfolding device 100 is supplied to an edge terminal, such as a VR goggle, downstream of the unfolding device 100. The VR goggles perform playback processing of the supplied playback data, thereby outputting audio and displaying images to the user.

[0190] The unfolding device 100 may also function as a playback device such as VR goggles. The unfolding device 100 shown in Fig. 22 is an example that also functions as a playback device.

[0191] The unfolding device 100 includes a separation processing unit 101, a sound source selection unit 102, an emphasis / suppression processing unit 103, an audio output unit 104, and a display processing unit 105. The unfolding device 100 may further include an operation unit 106.

[0192] The separation processing unit 101 receives sound source unit data Dss, for example, packet data or file data, generated by the signal processing device 1B as a recording device, and performs processing to separate the sound data, image data, and metadata.

[0193] The sound source selection unit 102 passes the sound data separated by the separation processing unit 101 to the emphasis / suppression processing unit 103. The emphasis / suppression processing unit 103 performs sound processing on the sound data received from the sound source selection unit 102 according to the viewing position in the virtual space, and also performs emphasis processing or suppression processing, for example, to emphasize the sound of other users who are communicating with the user.

[0194] The sound source selection unit 102 supplies image data separated for each sound source to 105, thereby allowing the user to visually recognize the image data.

[0195] In addition, the sound source selection unit 102 does not have to expand all of the received sound source unit data Dss into the virtual space, and in that case, it may be configured to expand only the acoustic data and image data for the sound source specified by the user via the operation unit 106.

[0196] The sound signal Sa obtained by performing predetermined sound processing by the emphasis / suppression processing unit 103 is supplied to the sound output unit 104, and is output to the user as sound.

[0197] The display processing unit 105 displays the image data supplied by the sound source selection unit 102. Furthermore, the display processing unit 105 performs processing such as superimposing an emphasis image on an image of a subject, for example, so that the subject emphasized by the emphasis / suppression processing unit 103 can be recognized.

[0198] In this way, the sound source unit data Dss acquired by the signal processing device 1B as a recording device can be expanded again in a virtual space or the like, thereby enabling each subject to be positioned in a different positional relationship from that in the recording environment. Then, the emphasis / suppression processing unit 103 applies appropriate sound processing according to the expanded positional relationship, thereby enabling the user to hear appropriate sound according to their viewing position, thereby providing a user with a natural and immersive experience according to their viewing position.

[0199] Although the sound source unit data Dss includes the scene data Dsc, the unfolding device 100 and the VR goggles may play back only the scene data Dsc, thereby making it possible to provide only the atmosphere of the recording environment.

[0200] An example of processing executed by the signal processing device 1B in this embodiment is shown in Fig. 23. Note that the same steps as those shown in Fig. 15 are given the same step numbers, and descriptions thereof will be omitted where appropriate.

[0201] In step S101, the microphone array 2A and the camera unit 3A of the signal processing device 1B perform sound collection and image capture processing.

[0202] In step S102, the image detection and estimation processing unit 5 of the signal processing device 1B performs a process of estimating a scene.

[0203] In step S103, the sound detection and estimation processing unit 42 of the signal processing device 1B performs coefficient selection and selects an AI model.

[0204] In step S141, the sound detection / estimation processing unit 42 of the signal processing device 1B performs, for example, a process of separating sound sources as AI sound processing. This process is realized, for example, by the sound source separation function F1 of the sound detection / estimation processing unit 42 performing AI sound processing using an AI model.

[0205] In step S142, the multiplexing processing unit 11 of the signal processing device 1B performs multiplexing processing, whereby the acoustic signals Sa and the image signals Sg are grouped together for each sound source.

[0206] The multiplexing processing unit 11 of the signal processing device 1B or a processing unit thereafter performs output processing in step S143. The output processing is, for example, processing for streaming output of the sound source unit data Dss objectified for each sound source, or processing for file output.

[0207] 24 shows an example of processing executed by the decompression device 100. In step S201, the separation processing unit 101 of the decompression device 100 performs separation processing. In this separation processing, acoustic data, image data, and metadata are separated from the objectified sound source unit data Dss.

[0208] The sound source selection unit 102 of the unpacking apparatus 100 selects one sound source in step S202.

[0209] Next, the enhancement / suppression processing unit 103 of the unfolding device 100 performs processing to enhance or suppress the acoustic signal Sa for the selected sound source, depending on, for example, the user's viewing position in the virtual space.

[0210] In step S204, the unpacking device 100 determines whether all sound sources have been selected, in other words, whether the processing of step S203 has been performed on the acoustic signals Sa relating to all sound sources.

[0211] If it is determined that an unselected sound source exists (step S204: No), the sound source selection unit 102 and the emphasis / suppression processing unit 103 execute the processes of steps S202 and S203, respectively.

[0212] On the other hand, if it is determined that all sound sources have been selected (step S204: Yes), the sound output unit 104 and display processing unit 105 of the unpacking device 100 each perform output processing in step S205. Specifically, the sound output unit 104 performs sound output processing based on the sound signal Sa in step S205, and the display processing unit 105 displays the image data supplied by the sound source selection unit 102 in step S205.

[0213] <6. Others> In the above example, the microphone array 2A is given as an example of the sound collection unit 2, but other embodiments are also possible.

[0214] For example, the sound collection unit 2 may be a single microphone. The sound collection unit 2 may also collect sound by bone conduction, by ultrasonic waves, or by using an acceleration sensor.

[0215] Furthermore, the camera unit 3A as one embodiment of the imaging unit 3 may be other than an RGB camera that outputs an image signal Sg for a color image. For example, it may be a monochrome camera that outputs a monochrome image signal Sg, an IR camera that outputs an IR (Infrared) image, or a distance measuring camera that outputs an image signal Sg related to distance. Furthermore, the camera unit 3A may be a radar camera that outputs an image signal Sg using radar, or a LiDAR camera that performs sensing using LiDAR (Light Detection and Ranging).

[0216] The embodiments of the sound pickup unit 2 and the camera unit 3A are merely examples, and various other embodiments are possible.

[0217] 7. Summary As described in the above examples, the signal processing device 1 (1A, 1B) includes an acoustic processing unit 4 (4A, 4B, 4C) that receives, as processing target data, an acoustic signal Sa obtained by collecting sound in a predetermined space and an image signal Sg obtained by capturing an image of the predetermined space, and performs acoustic processing on a sound source basis using an AI model MA (MB, MC, MD). For example, inference processing as acoustic processing is performed by inputting not only the acoustic signal Sa but also the image signal Sg to the AI ​​model. That is, by using the image signal Sg, it is possible to separate the acoustic signal Sa for each person or object, in other words, for each sound source, even when multiple objects are generating sounds in the same frequency band. If emphasis or suppression is performed for each frequency using only the acoustic signal Sa, the emphasis or suppression is performed uniformly regardless of the sound source, and it is not possible to preferably emphasize or suppress only the target sound. In contrast, with this configuration, by identifying the source of the sound using the image signal Sg, it is possible to emphasize or suppress the acoustic signal Sa for each sound source, thereby enabling high-quality noise removal and extraction of the acoustic signal Sa.

[0218] 1 to 3, the acoustic processing in the signal processing device 1 (1A, 1B) may be enhancement processing that enhances the specific sound Ss. This makes it possible to enhance only the voice of a specific person, for example, in a situation where multiple groups are talking simultaneously in close proximity.

[0219] 1 to 3, the acoustic processing in the signal processing device 1 (1A, 1B) may be suppression processing that suppresses the specific sound Ss. This makes it possible to suppress only the sound emitted from an unintended subject when, for example, sounds in the same frequency band are emitted from multiple subjects.

[0220] 5 and the like, the suppression processing in the signal processing device 1 (1A, 1B) may be processing for extracting and suppressing a specific sound Ss from the sound signal Sa. For example, by extracting a sound signal Sa for a specific sound source from the sound signal Sa obtained by sound collection and removing the extracted sound signal Sa from the original sound signal Sa, it is possible to generate a sound signal Sa in which the specific sound Ss is suppressed.

[0221] As described with reference to Figures 4 and 5, the suppression process in the signal processing device 1 (1A, 1B) may be a process of generating a specific sound Ss and removing the generated specific sound Ss from the sound signal Sa. For example, a subject generating the specific sound Ss is estimated from the image signal Sg, and a sound signal Sa emitted from the subject is generated using an AI model. By removing the generated sound signal Sa from the original sound signal Sa, it is possible to generate a sound signal Sa from which the sound generated from the specific sound source has been removed.

[0222] As described with reference to FIG. 5 and other figures, the acoustic processing unit 4A in the signal processing device 1 (1A, 1B) may use the acoustic signal Sa to generate the specific sound Ss. For example, the sounds emitted by cicadas vary depending on the type of cicada. In this case, by using the acoustic signal Sa collected to estimate the cicada sounds, it is possible to generate the sounds of a specific type of cicada as the acoustic signal Sa and remove the generated acoustic signal Sa from the original acoustic signal Sa. Therefore, it is possible to generate an acoustic signal Sa from which the sound of a specific sound source has been removed with high accuracy.

[0223] As described with reference to Figures 1 to 3, the audio signal Sa in the signal processing device 1 may be subjected to extraction processing to extract a specific sound Ss. For example, it is possible to extract only the audio signal Sa for a specific sound source from the audio signal Sa obtained by sound collection. This makes it possible to effectively remove noise, such as by reproducing only the voice of a specific person.

[0224] As described with reference to each of FIGS. 11 to 13 , the acoustic processing unit 4C in the signal processing device 1 may perform acoustic processing by inputting both the acoustic signal Sa and the image signal Sg to the AI ​​model MD. For example, the AI ​​model MD performs inference processing that takes into account both the features of the acoustic signal Sa and the features of the image signal Sg. As a result of the inference processing, an acoustic signal Sa that emphasizes a specific sound source, an acoustic signal Sa that extracts only sounds emitted from a specific sound source, or an acoustic signal Sa that suppresses a specific sound source can be obtained. The output of these acoustic signals Sa is achieved, for example, by the user specifying a sound source. This makes it possible to provide the user with acoustic data that has been subjected to appropriate acoustic processing.

[0225] 1 to 13, the acoustic processing unit 4 (4A, 4B) in the signal processing device 1 (1A, 1B) may perform acoustic processing based on the scene estimation result. By performing acoustic processing according to the estimated scene, more suitable acoustic processing can be realized.

[0226] As described with reference to each of Figures 8 to 10, the acoustic processing unit 4B in the signal processing device 1 (1A, 1B) may execute inference processing using scene estimation AI models (AI models MC1, MC2) that perform scene estimation, and inference processing using an acoustic processing AI model (AI model MB) that performs acoustic processing using a scene estimation result from the scene estimation AI model. That is, the acoustic processing unit 4B performs scene estimation and acoustic processing using multiple AI models. By separating the AI ​​models, the AI ​​models deployed in the signal processing device 1 can be made smaller, thereby reducing the amount of calculation and the storage area required in the signal processing device 1.

[0227] 8 to 10, the scene estimation result in the signal processing device 1 (1A, 1B) may be a result obtained based on the image signal Sg. By using the image signal Sg in the scene estimation, it is possible to improve the accuracy of the scene estimation.

[0228] 9 and 10 , the scene estimation result in the signal processing device 1 (1A, 1B) may be a result obtained further based on the acoustic signal Sa. By using the acoustic signal Sa in addition to the image signal Sg in the scene estimation, the accuracy of the scene estimation result can be further improved.

[0229] As described with reference to Figures 2 and 3, the acoustic processing unit 4A in the signal processing device 1 may switch the AI ​​model MA (MA1, MA2, MA3, ...) used for acoustic processing based on the scene estimation result. That is, different AI models MA1, MA2, MA3, ... are prepared for each scene. Each of these AI models MA is a small AI model MA. Therefore, it is possible to reduce the amount of calculation for AI processing as acoustic processing, and the processing time and processing load in the signal processing device 1 can be reduced.

[0230] As described with reference to FIG. 21 and the like, the acoustic processing unit 4 in the signal processing device 1B may generate acoustic data for each sound source. The signal processing device 1B may further include a data generation unit (multiplexing processing unit 11) that generates sound source unit data Dss using the acoustic data for each sound source. Generating meaningful data for each sound source allows subsequent processing to be performed favorably. Examples of subsequent processing include arranging sound sources in a metaverse space using the sound source unit data Dss, and applying a new acoustic signal Sa to the sound emitted from the sound source for each user present in the metaverse space to deliver a natural sound.

[0231] As described with reference to FIG. 21 etc., the data generation unit (multiplexing processing unit 11) in the signal processing device 1B may integrate sound data and image data on a sound source basis to generate sound source unit data Dss. By integrating sound data and image data to generate sound source unit data Dss, for example, it is possible to display images and output sound on a sound source basis in the metaverse space. This makes it possible to display images and output sound appropriately for each position of a user participating in the metaverse space. Furthermore, when playing back captured video at a later date, by using the image data and sound data compiled as sound source unit data Dss, it is possible to specify image data for a specific subject and preferably perform emphasis processing or suppression processing of only the sound of the subject associated with the image data.

[0232] As described with reference to Fig. 21 etc., in the signal processing device 1B, the sound source unit data Dss may include metadata related to the sound source. Examples of the metadata include information for identifying a scene and information about the sound source. The information about the sound source may be, for example, information indicating the posture of a person when the sound source is a person, or information indicating the operating state of an object such as a vehicle when the sound source is an object.

[0233] As described with reference to FIG. 14 and other figures, the signal processing device 1 may include an image processing unit (image detection / estimation processing unit 5C) that performs image processing on the image signal Sg and detects the gaze of the subject, and a voice recognition processing unit 8 that performs voice recognition processing on the subject using information about the subject whose gaze has been detected and the results of the acoustic processing. For example, as a result of the acoustic processing, voice information for each subject is separated and output. The voice recognition processing unit 8 performs voice recognition processing to understand the content of instructions based on the voice from each subject. In this case, when information about a subject looking toward the smart speaker is received, even if the voice about the subject is somewhat unclear, it may be assumed to be an instruction to the smart speaker, and the voice recognition threshold may be variable. For example, when a dialogue between a user and a smart speaker is initiated upon detecting the utterance of a specific activation word, it is conceivable to ease the dialogue start by relaxing the threshold for recognizing the activation word for a subject looking toward the smart speaker. In this way, by changing the threshold for each subject depending on the direction of the gaze, it is possible to understand the user's intention and start an appropriate dialogue.

[0234] 1 and the like, the acoustic signal Sa acquired by the signal processing device 1 may be a signal obtained by three-dimensional sound pickup using the microphone array 2A. As a result, the acoustic signal Sa may be a stereo signal. Therefore, for example, when a sound source is placed in the Metaverse space, more three-dimensional sound reproduction becomes possible.

[0235] The signal processing method of the present technology is executed by a signal processing device 1 (1A, 1B), and the signal processing device executes acoustic processing on a sound source basis using an AI model MA (MB, MC, MD) with an acoustic signal Sa obtained by collecting sound in a specified space and an image signal Sg obtained by imaging in the specified space as processing target data.

[0236] The program of this technology causes a computing device to execute a sound processing function for each sound source using an AI model MA (MB, MC, MD), with an acoustic signal Sa obtained by collecting sound in a specified space and an image signal Sg obtained by imaging in the specified space as processing target data.

[0237] Such a signal processing method and program can also provide the various functions and effects described above.

[0238] The above-mentioned expansion device 100 is equipped with a separation processing unit 101 that receives sound source unit data Dss, which is a compilation of sound data and image data for the sound source obtained by performing sound processing on a sound source unit basis using an AI model MA (MB, MC, MD) using an acoustic signal Sa obtained by collecting sound in a specified space and an image signal Sg obtained by imaging in a specified space as processing target data, and re-separates the data into sound data and image data, and an acoustic processing unit (such as an emphasis / suppression processing unit 103) that generates playback data using the sound data and image data for each sound source separated by the separation processing unit 101.

[0239] <8. The Present Technology> The present technology may also have the following configurations. (1) A signal processing device including an acoustic processing unit that receives as processing target data an acoustic signal obtained by collecting sound in a predetermined space and an image signal obtained by imaging the predetermined space, and performs acoustic processing on a sound source basis using an AI model. (2) The signal processing device according to (1) above, in which the acoustic processing is enhancement processing that enhances a specific sound. (3) The signal processing device according to any of (1) to (2) above, in which the acoustic processing is suppression processing that suppresses a specific sound. (4) The signal processing device according to (3) above, in which the suppression processing is processing that extracts and suppresses the specific sound from the acoustic signal. (5) The signal processing device according to (3) above, in which the suppression processing is processing that generates the specific sound and removes the generated specific sound from the acoustic signal. (6) The signal processing device according to (5) above, in which the acoustic processing unit uses the acoustic signal to generate the specific sound. (7) The signal processing device according to any one of (1) to (6), wherein the acoustic signal has been subjected to extraction processing to extract a specific sound. (8) The signal processing device according to any one of (1) to (7), wherein the acoustic processing unit performs the acoustic processing by inputting both the acoustic signal and the image signal to the AI ​​model. (9) The signal processing device according to any one of (1) to (8), wherein the acoustic processing unit performs the acoustic processing based on a scene estimation result. (10) The signal processing device according to (9), wherein the acoustic processing unit executes inference processing using a scene estimation AI model that performs scene estimation, and inference processing using an acoustic processing AI model that performs the acoustic processing using the scene estimation result by the scene estimation AI model. (11) The signal processing device according to any one of (9) to (10), wherein the scene estimation result is a result obtained based on the image signal. (12) The signal processing device according to (11), wherein the scene estimation result is a result obtained further based on the acoustic signal. (13) The signal processing device according to any one of (9) to (12), wherein the acoustic processing unit switches the AI ​​model used for the acoustic processing based on the scene estimation result.(14) The signal processing device according to any of (1) to (13), wherein the acoustic processing unit generates acoustic data for each sound source, and includes a data generation unit that generates sound source unit data using the acoustic data for each sound source. (15) The signal processing device according to (14), wherein the data generation unit generates the sound source unit data by integrating the acoustic data and image data for each sound source. (16) The sound source unit data includes metadata related to a sound source. (17) The signal processing device according to any of (1) to (16), including: an image processing unit that performs image processing on the image signal to detect a line of sight; and a voice recognition processing unit that performs voice recognition processing on the subject using information about the subject whose line of sight has been detected and a result of the acoustic processing. (18) The signal processing device according to any of (1) to (17), wherein the acoustic signal is a signal obtained by stereophonic sound collection by a microphone array. (19) A signal processing method in which a signal processing device executes acoustic processing for each sound source using an AI model with an acoustic signal obtained by collecting sound in a predetermined space and an image signal obtained by imaging the predetermined space as processing target data. (20) A program that causes a processing device to execute an acoustic processing function for each sound source using an AI model with an acoustic signal obtained by collecting sound in a predetermined space and an image signal obtained by imaging the predetermined space as processing target data.

[0240] 1 Signal processing device 1A Signal processing device 1B Signal processing device 2A Microphone array 4 Acoustic processing unit 4A Acoustic processing unit 4B Acoustic processing unit 4C Acoustic processing unit 5C Image detection and estimation processing unit (image processing unit) 8 Speech recognition processing unit 11 Multiplexing processing unit (data generation unit) Dss Sound source unit data MA AI model MA1 AI model MA2 AI model MA3 AI model MB AI model MC1 AI model MC2 AI model MD AI model Sa Acoustic signal Sg Image signal Ss Specific sound

Claims

1. A signal processing device having an acoustic processing unit that receives an acoustic signal obtained by collecting sound in a specified space and an image signal obtained by imaging the specified space as data to be processed and performs acoustic processing on a sound source basis using an AI model.

2. The signal processing device according to claim 1, wherein the acoustic processing is an enhancement process for enhancing a specific sound.

3. The signal processing device according to claim 1, wherein the acoustic processing is suppression processing for suppressing a specific sound.

4. The signal processing device according to claim 3, wherein the suppression process is a process of extracting and suppressing the specific sound from the acoustic signal.

5. The signal processing device according to claim 3, wherein the suppression process is a process of generating the specific sound and removing the generated specific sound from the acoustic signal.

6. The signal processing device according to claim 5, wherein the audio processing unit uses the audio signal to generate the specific sound.

7. The signal processing device according to claim 1, wherein the audio signal is subjected to extraction processing for extracting a specific sound.

8. The signal processing device according to claim 1, wherein the acoustic processing unit performs the acoustic processing by inputting both the acoustic signal and the image signal to the AI ​​model.

9. The signal processing device according to claim 1, wherein the audio processing unit performs the audio processing based on a scene estimation result.

10. The signal processing device according to claim 9, wherein the acoustic processing unit executes inference processing using a scene estimation AI model that performs scene estimation, and inference processing using an acoustic processing AI model that performs the acoustic processing using the scene estimation result by the scene estimation AI model.

11. The signal processing device according to claim 9, wherein the scene estimation result is a result obtained based on the image signal.

12. The signal processing device according to claim 11, wherein the scene estimation result is a result obtained further based on the acoustic signal.

13. The signal processing device according to claim 9, wherein the acoustic processing unit switches the AI ​​model used for the acoustic processing based on the scene estimation result.

14. The signal processing device according to claim 1, wherein the acoustic processing unit generates acoustic data for each sound source, and further comprises a data generating unit that generates sound source unit data using the acoustic data for each sound source.

15. The signal processing device according to claim 14, wherein the data generating section generates the sound source unit data by integrating the acoustic data and image data for each sound source.

16. The signal processing device according to claim 14, wherein the sound source unit data includes metadata relating to the sound source.

17. A signal processing device as described in claim 1, comprising: an image processing unit that performs image processing on the image signal to detect a gaze; and a voice recognition processing unit that performs voice recognition processing on the subject using information about the subject about which the gaze has been detected and the result of the acoustic processing.

18. The signal processing device according to claim 1, wherein the acoustic signal is a signal picked up in three dimensions by a microphone array.

19. A signal processing method in which a signal processing device performs acoustic processing on a sound source basis using an AI model with an acoustic signal obtained by collecting sound in a specified space and an image signal obtained by imaging the specified space as processing target data.

20. A program that causes a processor to execute a sound processing function for each sound source using an AI model, with the sound signal obtained by collecting sound in a specified space and the image signal obtained by imaging the specified space as the processing target data.

Citation Information

Patent Citations

  • Computer and voice processing method

    JP2022063080A

  • On-line speaker sequential discrimination method, on-line speaker sequential discrimination device and on-line speaker sequential discrimination system

    JP2023092836A

  • Information processing device, voice recognition method, utterance section detection method, and program

    JP2023107552A