Information processing method, information processing system, and program

The method and system address the challenge of personalized wind noise cancellation by recognizing user face shapes and generating tailored filters, enhancing audio quality by effectively reducing wind noise.

WO2026034187A1PCT designated stage Publication Date: 2026-02-12SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/026025
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-08
Filing Date
2025-07-23
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing audio devices struggle to effectively cancel wind noise due to its complex nature, which varies with individual user face shapes, making it difficult to create personalized noise cancellation filters.

Method used

An information processing method and system that recognizes user face shapes through image analysis, calculates contributions using dynamic mode decomposition, and generates personalized noise cancellation filters to counteract wind noise based on these shapes.

Benefits of technology

This approach allows for reliable noise cancellation tailored to individual user face shapes, improving audio quality by reducing wind noise in various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025026025_12022026_PF_FP_ABST
    Figure JP2025026025_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing method, an information processing system, and a program that make it possible to more reliably obtain the effect of noise canceling. This information processing method comprises: performing shape recognition processing for recognizing a user's facial shape influencing the generation of wind noise on the basis of an input facial image; calculating a contribution degree indicating the degree of contribution of a facial shape of supervised data with respect to the user's facial shape; calculating, according to the contribution degree, an estimated value representing wind noise estimated to be generated according to the user's facial shape, by combining weights for distances in a feature amount space between feature amounts of dynamic mode decomposition in facial shapes of respective supervised data, with the feature amounts of the respective supervised data; and generating, on the basis of the estimated value, a corresponding filter for canceling out the influence of the wind noise generated according to the user's facial shape. The present technology can be applied to, for example, noise canceling in acoustic equipment such as a headphone and an earphone.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing system, and program

[0001] The present disclosure relates to an information processing method, an information processing system, and a program, and more particularly to an information processing method, an information processing system, and a program that enable a noise canceling effect to be obtained more reliably.

[0002] Conventionally, audio devices such as headphones and earphones have been able to provide high-quality audio with the influence of external sounds suppressed by applying noise cancellation processing using a filter that is compatible with external sounds.

[0003] However, when using audio equipment in outdoor environments or under air conditioning, it has been difficult to apply noise cancellation processing to noise (hereinafter referred to as wind noise) generated by wind hitting the human body, audio equipment, etc. In other words, wind noise is a mixture of sound sources (sound waves) generated by the collision of fluids and pseudo sound sources (non-sound waves) caused by pressure fluctuations under turbulent flow, making it difficult to identify the cause of wind noise and making it impossible to generate an appropriate filter to deal with wind noise.

[0004] Patent Document 1 discloses a system that obtains the state of flow of objects in a scene using a moving image of the flow captured by a camera as an input.

[0005] Japanese Patent Application Laid-Open No. 2018-113026

[0006] Recently, it has become possible to confirm how wind noise is generated by reproducing the generation of wind noise using a low-noise wind tunnel for audio equipment and comparing it with a fluid simulation. As a result, it has become clear that wind noise caused by the shape of a user's face varies from person to person, and that in order to achieve noise cancellation effects, it is necessary to optimize filters for various user face shapes. However, it has not been easy to create filters that accommodate individual differences in wind noise.

[0007] The present disclosure has been made in view of such circumstances, and aims to make it possible to obtain the effect of noise canceling more reliably.

[0008] An information processing method and program according to one aspect of the present disclosure includes performing a shape recognition process to recognize, based on an input image, the shapes of a user's body parts that affect the generation of wind noise; determining a relationship between the shapes of the user's body parts and the shapes of body parts used to obtain a plurality of pre-calculated training data sets and calculating a contribution indicating the degree to which the shapes of the body parts of the training data contribute to the shape of the user's body parts; calculating an estimate representing wind noise that is estimated to occur in accordance with the shape of the user's body parts by combining, in accordance with the contribution, weights for the distances in a feature space between features of dynamic mode decomposition for the shapes of the body parts of each of the training data sets with the features of each of the training data sets; and generating, based on the estimate, a corresponding filter that cancels out the influence of wind noise that occurs in accordance with the shape of the user's body parts.

[0009] An information processing system according to one aspect of the present disclosure includes an information processing device having: a shape recognition processing unit that performs shape recognition processing to recognize the shape of a user's body part that affects the generation of wind noise based on an input image; a contribution calculation unit that determines the relationship between the shape of the user's body part and the shapes of the body parts used to obtain a plurality of pre-calculated teacher data, and calculates a contribution degree indicating the degree to which the shape of the body part of the teacher data contributes to the shape of the user's body part; an estimate calculation unit that calculates an estimate representing wind noise that is estimated to occur in accordance with the shape of the user's body part by combining, in accordance with the contribution degree, a weight for the distance in feature space between features of dynamic mode decomposition in the shape of the body part for each of the teacher data with the feature for each of the teacher data; and a correspondence filter generation unit that generates, based on the estimate, a correspondence filter that cancels the effect of wind noise that occurs in accordance with the shape of the user's body part; and an audio device having a noise cancellation processing unit that supplies to a speaker an audio signal that has been subjected to noise cancellation processing using the correspondence filter based on an external sound signal including wind noise picked up by a microphone.

[0010] In one aspect of the present disclosure, a shape recognition process is performed to recognize the shapes of a user's body parts that affect the generation of wind noise based on an input image, and a relationship between the shapes of the user's body parts and the shapes of the body parts used to obtain multiple pieces of pre-calculated training data is determined. A contribution indicating the degree to which the shapes of the body parts of the training data contribute to the shape of the user's body parts is calculated. According to the contribution, a weight for the distance in feature space between features of dynamic mode decomposition for the shapes of the body parts of each piece of training data is combined with the features of each piece of training data to calculate an estimated value representing wind noise that is estimated to occur depending on the shape of the user's body parts, and a corresponding filter that cancels out the effect of wind noise that occurs depending on the shape of the user's body parts is generated based on the estimated value.

[0011] FIG. 1 is a block diagram showing an example configuration of an embodiment of an audio system to which the present technology is applied; FIG. 2 is a diagram showing an example of a front face image and a side face image; FIG. 3 is a diagram showing an example of wind noise obtained from three types of face shapes; FIG. 4 is a diagram explaining an example of classification by machine learning; FIG. 5 is a flowchart explaining filter generation processing; and FIG. 6 is a block diagram showing an example configuration of an embodiment of a computer to which the present technology is applied.

[0012] Hereinafter, specific embodiments to which the present technology is applied will be described in detail with reference to the drawings.

[0013] <Configuration Example of Sound System> FIG. 1 is a block diagram showing a configuration example of an embodiment of a sound system to which the present technology is applied.

[0014] 1 is configured to include a filter generation processing unit 21 and an acoustic device 22. For example, the acoustic system 11 generates a filter (hereinafter referred to as a corresponding filter) corresponding to individual differences in wind noise due to the shape of the face of the user using the acoustic device 22 in the filter generation processing unit 21, and installs the corresponding filter in the acoustic device 22, thereby more reliably achieving a noise canceling effect against wind noise.

[0015] The filter generation processing unit 21 includes a face image acquisition unit 31, a face shape recognition processing unit 32, a teacher data storage unit 33, a contribution calculation unit 34, an estimated value calculation unit 35, and a corresponding filter generation unit 36.

[0016] The face image acquisition unit 31 acquires a face image obtained by photographing the face of a user using the audio device 22, and supplies the face image to the face shape recognition processing unit 32. For example, the face image acquisition unit 31 acquires a front face image obtained by photographing the user's face from the front as shown in A of Fig. 2, and a side face image obtained by photographing the user's face from the side as shown in B of Fig. 2. The user can, for example, use the camera function of a smartphone to photograph the front face image and the side face image of the user himself / herself, and transmit them to the audio system 11.

[0017] The facial shape recognition processing unit 32 performs a facial shape recognition process to recognize the user's facial shape (e.g., facial contours, ear position, nose height, cheek shape, etc.) based on the frontal face image and side face image supplied from the facial image acquisition unit 31. Specifically, the facial shape recognition processing unit 32 recognizes the dimensions of each part of the face, such as the width of the face, the distance between the ears, the distance between the eyes, the distance between the right eye and the nose, the distance between the left eye and the nose, shoulder width, and the difference in height between the left and right ears, based on the frontal face image shown in FIG. 2A. Similarly, the facial shape recognition processing unit 32 recognizes the dimensions of each part of the face, such as the thickness of the face, the distance between the eyes and the ears, the distance between the ears and the nose, shoulder thickness, and facial height, based on the side face image shown in FIG. 2B. Note that the facial shape recognition processing unit 32 can appropriately recognize the user's facial shape even if the user is wearing a mask, glasses, etc. Then, the facial shape recognition processing unit 32 supplies facial shape data indicating the user's facial shape obtained by the facial shape recognition process (e.g., the dimensions of each part of the face as described above) to the contribution calculation unit 34.

[0018] The training data storage unit 33 stores training data obtained by performing dynamic mode decomposition on analysis results (e.g., animation showing the process of time-varying distribution data of physical quantities (velocity, pressure, etc.) in an unsteady flow) obtained by performing fluid analysis on a plurality of representative facial shapes, extracting dynamic mode decomposition features (characteristic changes in the fluid) for each facial shape, and calculating in advance the distances between these feature values ​​in feature space. For example, the training data includes facial shape data showing the plurality of facial shapes used to obtain the training data, feature value data showing the dynamic mode decomposition features for each facial shape, and distance data showing the distances between these feature values.

[0019] For example, Fig. 3 shows an example of wind noise calculated from each face shape when fluid analysis is performed using three types of face shapes as training data under the same boundary conditions such as wind speed. As shown in Fig. 3, wind noise varies from person to person depending on the face shape.

[0020] The contribution degree calculation unit 34 determines the relationship between the user's facial shape data supplied from the facial shape recognition processing unit 32 and the facial shape data for each piece of teacher data stored in the teacher data storage unit 33. Then, the contribution degree calculation unit 34 calculates a contribution degree indicating the degree to which the facial shape of the teacher data contributes to the user's facial shape from the similarity between the user's facial shape and the facial shape of each piece of teacher data, and supplies the contribution degree and the teacher data to the estimated value calculation unit 35. Note that if it is difficult to determine the similarity of the facial shapes themselves, the contribution degree calculation unit 34 may determine the contribution degree by a simple calculation.

[0021] For example, when the three types of face shapes shown in FIG. 3 are used as training data, the contribution calculation unit 34 calculates the contribution of Type 1 face shape to the user's face shape, the contribution of Type 2 face shape to the user's face shape, and the contribution of Type 3 face shape to the user's face shape.

[0022] The estimate calculation unit 35 calculates weights for the distances between the features of the dynamic mode decomposition for the face shape of each piece of training data, in accordance with the contributions supplied from the contribution calculation unit 34. The estimate calculation unit 35 then combines the weights with the feature data for each piece of training data to calculate an estimate representing wind noise that is estimated to occur depending on the shape of the user's face, and supplies the estimate to the corresponding filter generation unit 36.

[0023] For example, the estimated value calculation unit 35 can use feature data for each training data previously obtained from three types of face shapes as shown in Figure 3 to calculate an estimated value that changes over time from a user's face shape (still information) that is different from those three face shapes.

[0024] The corresponding filter generation unit 36 ​​generates a corresponding filter that cancels out the effects of wind noise that occurs depending on the shape of the user's face based on the estimated value supplied from the estimated value calculation unit 35, and sets (installs) it in the noise cancellation processing unit 41 of the audio equipment 22.

[0025] The acoustic device 22 includes a noise cancellation processing unit 41 , a microphone 42 , a speaker 43 , and a notification unit 44 .

[0026] In the audio device 22, an external sound signal including wind noise picked up by the microphone 42 is supplied to the noise cancellation processing unit 41, and the noise cancellation processing unit 41 applies noise cancellation processing to the external sound signal using a corresponding filter (wind noise transfer function), and supplies the resulting audio signal to the speaker 43. This allows the audio device 22 to output from the speaker 43 sound that has been subjected to noise cancellation so as to accommodate wind noise that varies from person to person depending on the shape of the user's face, and can provide high-quality audio in which the effects of external sound including wind noise are suppressed.

[0027] At this time, the acoustic device 22 can notify the user (for example, by displaying a message on the display or turning on a lamp) via the notification unit 44 that noise cancellation processing using the corresponding filter has been applied.

[0028] The acoustic system 11 configured as described above can generate a filter optimized for the face shape of each user using the acoustic device 22, and can more reliably achieve a noise canceling effect against wind noise. Therefore, the acoustic system 11 can reduce the effects of wind noise even in outdoor environments or under air conditioning, improving the user's auditory experience.

[0029] Furthermore, the acoustic system 11 can control noise canceling more efficiently by increasing the conditions for calculating the training data. For example, by including the wind direction dependency that generates wind noise as a condition, noise cancellation processing using a corresponding filter can be applied only when wind from a specific direction is detected, thereby improving the control efficiency of noise canceling.

[0030] In addition, the acoustic system 11 can improve the efficiency of noise canceling control by reducing the amount of data by reducing the amount of cutoff for training data with high probability when extracting features of dynamic mode decomposition.

[0031] <Classification by Machine Learning> The acoustic system 11 can improve the accuracy of noise canceling by using classification by machine learning.

[0032] FIG. 4 shows an example of training data that has been clustered by performing classification using machine learning.

[0033] For example, in the acoustic system 11, if a sufficient amount of training data is available, the training data may be pre-classified (clustered) using a support vector machine according to similarities such as major classifications and intermediate classifications, thereby improving the accuracy of noise canceling.

[0034] For example, in Fig. 4, the teacher data is classified into regions similar to face shapes of Type 1, Type 2, and Type 3 according to the major classifications separated by thick dashed lines. In addition, the teacher data is classified into intermediate classifications surrounded by thin dashed lines.

[0035] The acoustic system 11 then refers to the distance data from each training data to represent the user's facial shape based on the facial image input to the filter generation processing unit 21, and can refine the classification by expressing it as Type 1: 15%, Type 2: 75%, and Type 3: 10%. For example, the user's facial shape is classified into the position of the star shown in Figure 4.

[0036] Furthermore, the sound system 11 can manage users of the sound system 11 by user ID, and associate each user ID with a feature of dynamic mode decomposition. This allows the user ID to be associated with a corresponding filter optimized for the face shape of that user.

[0037] The sound system 11 can input multiple facial images taken from two or more locations using a camera with two or more lenses. For example, input using a 360-degree camera, which combines multiple images, may be adopted. This allows input to be completed with a single shot, reducing the burden on the user.

[0038] The sound system 11 can input a three-dimensional shape estimated based on the shading of the user's face shape from a facial image obtained by photographing the user from the front. This allows input to be completed with a single photograph, reducing the burden on the user.

[0039] The sound system 11 can input multiple images obtained by continuously photographing the user or a video obtained by photographing the user. The system can then create a facial shape of the user from the multiple images or video and use an estimation method for dynamic quantities. This allows input to be completed with a single photograph, reducing the burden on the user.

[0040] The acoustic system 11 can input distance data corresponding to the shape of the user's face obtained using a TOF (Time of Flight) sensor. The user's face shape can then be created from the distance data, and an estimation method for dynamic quantities can be used. This allows input to be completed with a single image capture, reducing the burden on the user and improving estimation accuracy by improving the accuracy of the user's face shape.

[0041] The sound system 11 can input multiple facial images obtained by simultaneously capturing the user's face using two or more cameras. The user's facial shape can then be created from the multiple facial images, and an estimation method for dynamic quantities can be used. This improves the accuracy of estimating the three-dimensional shape of the user's face, thereby improving estimation accuracy.

[0042] The acoustic system 11 can input numerical data obtained by actually measuring the shape of the user's face. Based on the numerical data, an estimation method for dynamic quantities can be used. This allows the acoustic system 11 to handle situations where, for example, no device for capturing a facial image is available.

[0043] The sound system 11 can use a user's facial image as input to generate a three-dimensional shape of the user's face using artificial intelligence (AI). Based on the three-dimensional shape, an estimation method for dynamic quantities may be used. This reduces the user's burden in taking photographs and improves convenience.

[0044] The filter generation processing unit 21 can estimate dynamic quantities by inputting, instead of the above-mentioned face image, the distribution of physical quantities (for example, fluid, magnetic field, sound field, velocity field, pressure field, or a combination thereof) obtained by numerical calculation, and calculate a transfer function equivalent to a filter corresponding to wind noise. This makes it possible to generate filters corresponding to phenomena corresponding to each physical quantity.

[0045] By simplifying the deep learning set, the acoustic system 11 can generate corresponding filters in an offline environment, such as on a user's smartphone or PC, without requiring a cloud environment, thereby improving convenience.

[0046] If a face image contains obstacles such as a mask, glasses, or hat that prevent the recognition of the user's face shape, the face shape recognition processing unit 32 can remove these obstacles at the time the face image is supplied, thereby enabling the face shape recognition processing unit 32 to include facial components that cause wind noise. This makes it possible to generate a corresponding filter from a face image from which the mask, glasses, hat, etc. have been removed, thereby improving accuracy in terms of individual differences between users.

[0047] In the acoustic system 11, the accuracy of noise canceling can be improved by increasing the amount of data for comparing the user's face shape with the face shapes of the training data by making the facial image input to the filter generation processing unit 21 more detailed. For example, the user's face may be photographed from three directions, or a method may be used that automatically generates face shapes in multiple directions from shading when photographing the user's face (for example, an automatic generation method such as Gaussian splatting).

[0048] Furthermore, the sound system 11 can use a smartphone application to guide the user to capture a facial image that makes the user's facial shape easier to recognize. For example, a user interface can be displayed on the capture screen that allows the user to efficiently capture the main facial features that affect wind noise (such as the head, chest, eyes, nose, and mouth).

[0049] For example, when capturing a frontal face image as shown in A of Fig. 2, a user interface indicating positions at which the main facial features that affect wind noise will be captured is displayed on the capture screen so that at least both ears, nose, and contour are captured, and more preferably, the relative positions of the ears, nose, eyes, and mouth, as well as the width of the forehead and lips, are also captured. Furthermore, when the facial shape recognition processing unit 32 cannot recognize the shape of the user's face and recaptures the frontal face image, an error is output and an instruction is displayed on the capture screen urging the user to capture the image so that the main facial features that affect wind noise are included.

[0050] 2B, a user interface is displayed on the shooting screen indicating the positions at which the main facial features that affect wind noise will be captured, so that at least the outline and the distance between the ears and nose are captured, and more preferably, the height of the cheekbones, the width of the lips and forehead, the shape of the chin, etc. Furthermore, when the facial shape recognition processing unit 32 cannot recognize the shape of the user's face and re-captures the side face image, an error is output and an instruction is displayed on the shooting screen urging the user to capture the image so that the main facial features that affect wind noise are included.

[0051] <Processing Example of Filter Generation Processing> The filter generation processing executed by the filter generation processing unit 21 will be described with reference to the flowchart shown in FIG.

[0052] In step S11 , the face image acquisition unit 31 acquires face images (front face image and side face image as shown in FIG. 2 above) photographed by the user, and supplies them to the face shape recognition processing unit 32 .

[0053] In step S12, the facial shape recognition processing unit 32 performs a facial shape recognition process to recognize the shape of the user's face based on the facial image supplied from the facial image acquisition unit 31 in step S11.

[0054] In step S13, the facial shape recognition processing unit 32 determines whether or not the shape of the user's face has been recognized in the facial shape recognition processing in step S12.

[0055] If the facial shape recognition processing unit 32 determines in step S13 that it was unable to recognize the shape of the user's face, the process returns to step S11, and facial images are repeatedly acquired. At this time, as described above, the smartphone application can provide guidance to capture a facial image that makes the shape of the user's face easier to recognize.

[0056] On the other hand, if the facial shape recognition processing unit 32 determines in step S13 that it has recognized the shape of the user's face, it supplies the user's facial shape data obtained by the facial shape recognition processing in step S12 to the contribution calculation unit 34, and processing proceeds to step S14.

[0057] In step S14, the contribution calculation unit 34 calculates the contribution based on the user's facial shape data supplied from the facial shape recognition processing unit 32 in step S13 and the teacher data stored in the teacher data memory unit 33, and supplies it to the estimated value calculation unit 35.

[0058] In step S15, the estimate calculation unit 35 calculates an estimate representing wind noise that is estimated to occur depending on the shape of the user's face, in accordance with the contribution supplied from the contribution calculation unit 34 in step S14, and supplies the estimate to the corresponding filter generation unit 36.

[0059] In step S16, the corresponding filter generation unit 36 ​​generates a corresponding filter that cancels out the influence of wind noise caused by the shape of the user's face, based on the estimated value supplied in step S15 from the estimated value calculation unit 35. Then, after the corresponding filter generation unit 36 ​​sets (installs) the corresponding filter in the noise cancellation processing unit 41 of the acoustic device 22, the filter generation process ends.

[0060] By performing the filter generation process described above, it is possible to generate a corresponding filter that is optimized for each user's face shape and can more reliably achieve a noise canceling effect that corresponds to wind noise.

[0061] In this embodiment, the shape of the user's face is used as input, but a body part of the user that affects the generation of wind noise, for example, the shape of the user's upper body, may also be used as input.

[0062] Furthermore, although this embodiment applies the present technology to noise canceling in the acoustic device 22, the present technology can also be applied to various other time-varying phenomena, such as changes in fluid, changes in electromagnetic field distribution, changes in vibration, and moving image processing. That is, the present technology can be applied to estimating the amount of change over time in, for example, structural deformation, fluid change, and electromagnetic field distribution. Furthermore, the present technology makes it possible to automatically and quickly calculate and visualize estimable solutions for deformation or transition phenomena that have been difficult to analyze in the past.

[0063] <Example of Computer Configuration> Next, the above-described series of processes (information processing method) can be performed by hardware or software. When the series of processes is performed by software, a program constituting the software is installed in a general-purpose computer or the like.

[0064] FIG. 6 is a block diagram showing an example of the configuration of an embodiment of a computer in which a program for executing the above-described series of processes is installed.

[0065] The program can be recorded in advance on the hard disk 105 or ROM 103 as a recording medium built into the computer.

[0066] Alternatively, the program can be stored (recorded) on a removable recording medium 111 driven by the drive 109. Such a removable recording medium 111 can be provided as a so-called package software. Here, examples of the removable recording medium 111 include a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disk, a DVD (Digital Versatile Disc), a magnetic disk, and a semiconductor memory.

[0067] The program can be installed into the computer from the removable recording medium 111 as described above, or can be downloaded to the computer via a communication network or a broadcasting network and installed on the built-in hard disk 105. That is, the program can be transferred to the computer wirelessly from a download site via an artificial satellite for digital satellite broadcasting, or transferred to the computer by wire via a network such as a LAN (Local Area Network) or the Internet.

[0068] The computer includes a CPU (Central Processing Unit) 102 , to which an input / output interface 110 is connected via a bus 101 .

[0069] When a user inputs a command via the input / output interface 110 by operating the input unit 107, the CPU 102 executes a program stored in a read-only memory (ROM) 103 in accordance with the command. Alternatively, the CPU 102 loads a program stored on a hard disk 105 into a random access memory (RAM) 104 and executes the program.

[0070] As a result, the CPU 102 performs processing according to the flowchart described above or processing performed by the configuration of the block diagram described above. Then, the CPU 102 outputs the processing results from the output unit 106 via the input / output interface 110, transmits them from the communication unit 108, or records them on the hard disk 105, as necessary.

[0071] The input unit 107 is made up of a keyboard, a mouse, a microphone, etc. The output unit 106 is made up of an LCD (Liquid Crystal Display), a speaker, etc.

[0072] In this specification, the processing performed by a computer according to a program does not necessarily have to be performed in chronological order according to the order described in the flowchart. In other words, the processing performed by a computer according to a program also includes processing that is executed in parallel or individually (for example, parallel processing or object-based processing).

[0073] The program may be processed by a single computer (processor), or may be distributed among multiple computers. Furthermore, the program may be transferred to and executed on a remote computer.

[0074] Furthermore, in this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0075] Also, for example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0076] Furthermore, for example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0077] Furthermore, for example, the above-described program can be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.

[0078] Also, for example, each step described in the above flowchart can be executed by one device or can be shared and executed by multiple devices. Furthermore, if one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.

[0079] In addition, the processing of the steps of a program executed by a computer may be executed in chronological order according to the order described in this specification, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the processing of each step may be executed in an order different from the order described above. Furthermore, the processing of the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0080] It should be noted that the present technologies described in this specification can be implemented independently and singly, unless a contradiction arises. Of course, any two or more of the present technologies can also be implemented in combination. For example, part or all of the present technologies described in any embodiment can be implemented in combination with part or all of the present technologies described in other embodiments. Furthermore, part or all of any of the present technologies described above can also be implemented in combination with other technologies not described above.

[0081] <Example Combinations of Configurations> The present technology can also be configured as follows: (1) An information processing method including: performing a shape recognition process to recognize, based on an input image, the shapes of a user's body parts that affect the generation of wind noise; determining a relationship between the shapes of the user's body parts and the shapes of body parts used to obtain multiple pieces of pre-calculated training data, and calculating a contribution indicating the degree to which the shapes of the body parts of the training data contribute to the shapes of the user's body parts; calculating an estimated value representing wind noise estimated to be generated in accordance with the shapes of the user's body parts by combining, in accordance with the contribution, weights for distances in feature space between features of dynamic mode decomposition for the shapes of the body parts of each piece of training data with the features of each piece of training data; and generating, based on the estimated value, a correspondence filter that cancels out the influence of wind noise generated in accordance with the shapes of the user's body parts. (2) The information processing method described in (1), in which the training data is obtained by performing a fluid analysis on the shapes of multiple representative body parts, performing dynamic mode decomposition on analysis results obtained, extracting features of dynamic mode decomposition for the shapes of each body part, and pre-calculating the distances in feature space between these features. (3) The information processing method according to (1) or (2), wherein an image obtained by photographing the user's face from the front and an image obtained by photographing the user's face from the side are input. (4) The information processing method according to any of (1) to (3), wherein a plurality of facial images photographed from two or more positions by a camera having two or more lenses are input. (5) The information processing method according to any of (1) to (4), wherein a plurality of images obtained by continuously photographing the user or a video image obtained by photographing the user are input. (6) The information processing method according to any of (1) to (5), wherein distance data according to the shape of the user's face obtained by using a TOF (Time of Flight) sensor is input. (7) The information processing method according to any of (1) to (6), wherein a plurality of images obtained by simultaneously photographing body parts of the user using two or more cameras are input.(8) The information processing method according to any one of (1) to (7), wherein numerical data obtained by actually measuring the user's body part is input. (9) The information processing method according to any one of (1) to (8), wherein an image of the user's body part is input, and a three-dimensional shape of the user's body part is generated by a generation AI (artificial intelligence). (10) The information processing method according to any one of (1) to (9), wherein a distribution of physical quantities obtained by numerical calculation is input instead of the image, and a transfer function corresponding to the correspondence filter is calculated. (11) The information processing method according to any one of (1) to (10), wherein the correspondence filter is generated in an offline environment. (12) The information processing method according to any one of (1) to (11), wherein, if obstacles that prevent recognition of the shape of the user's body part are captured in the image, the obstacles are removed to generate the correspondence filter. (13) An information processing system including: an information processing device having: a shape recognition processing unit that performs shape recognition processing to recognize the shapes of body parts of a user that affect the generation of wind noise based on an input image; a contribution calculation unit that determines a relationship between the shapes of the user's body parts and the shapes of body parts used to obtain multiple pieces of pre-calculated teacher data, and calculates a contribution degree indicating the degree to which the shapes of the body parts of the teacher data contribute to the shape of the user's body part; an estimate calculation unit that calculates an estimate representing wind noise that is estimated to be generated in accordance with the shape of the user's body part by combining, in accordance with the contribution degree, weights for distances in a feature space between features of dynamic mode decomposition in the shapes of the body parts of each piece of teacher data with the features of each piece of teacher data; and a correspondence filter generation unit that generates, based on the estimate, a correspondence filter that cancels the effect of wind noise that is generated in accordance with the shape of the user's body part; and an audio device having a noise cancellation processing unit that supplies to a speaker an audio signal that has been noise canceled using the correspondence filter based on an external sound signal including wind noise picked up by a microphone.(14) A program for causing a computer of an information processing device to execute information processing, including: performing shape recognition processing to recognize, based on an input image, the shapes of a user's body parts that affect the generation of wind noise; determining a relationship between the shapes of the user's body parts and the shapes of body parts used to obtain multiple pieces of pre-calculated teacher data, and calculating a contribution indicating the degree to which the shapes of the body parts of the teacher data contribute to the shape of the user's body parts; calculating an estimated value representing wind noise that is estimated to occur depending on the shape of the user's body parts by combining, according to the contribution, weights for distances in a feature space between features of dynamic mode decomposition in the shapes of the body parts of each piece of teacher data with the features of each piece of teacher data; and generating, based on the estimated value, a corresponding filter that cancels out the influence of wind noise that occurs depending on the shape of the user's body parts.

[0082] It should be noted that the present embodiment is not limited to the above-described embodiment, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, the effects described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained.

[0083] REFERENCE SIGNS LIST 11 Acoustic system, 21 Filter generation processing unit, 22 Acoustic equipment, 31 Face image acquisition unit, 32 Face shape recognition processing unit, 33 Teacher data storage unit, 34 Contribution degree calculation unit, 35 Estimated value calculation unit, 36 Corresponding filter generation unit, 41 Noise cancellation processing unit, 42 Microphone, 43 Speaker, 44 Notification unit

Claims

1. An information processing method comprising: performing a shape recognition process to recognize the shapes of a user's body parts that affect the generation of wind noise based on an input image; determining a relationship between the shapes of the user's body parts and the shapes of body parts used to obtain multiple pieces of pre-calculated training data, and calculating a contribution indicating the degree to which the shapes of the body parts of the training data contribute to the shape of the user's body parts; calculating an estimate representing wind noise that is estimated to occur in accordance with the shape of the user's body parts by combining, in accordance with the contribution, weights for the distance in feature space between features of dynamic mode decomposition in the shapes of the body parts of each piece of training data with the features of each piece of training data; and generating, based on the estimate, a corresponding filter that cancels out the influence of wind noise that occurs in accordance with the shape of the user's body parts.

2. The information processing method of claim 1, wherein the training data is obtained by performing fluid analysis on the shapes of a number of representative body parts, applying dynamic mode decomposition to the analysis results, extracting dynamic mode decomposition features for the shapes of each body part, and calculating in advance the distances between these features in feature space.

3. The information processing method according to claim 1, wherein an image obtained by photographing the user's face from the front and an image obtained by photographing the user's face from the side are input.

4. The information processing method according to claim 1, wherein a plurality of facial images taken from two or more locations by a camera having two or more lenses are input.

5. The information processing method according to claim 1, wherein a plurality of images obtained by continuously photographing the user or a moving image obtained by photographing the user is input.

6. The information processing method according to claim 1, wherein distance data corresponding to the shape of the user's face obtained using a TOF (Time of Flight) sensor is used as input.

7. The information processing method according to claim 1, wherein a plurality of images obtained by simultaneously photographing the user's body parts using two or more cameras are input.

8. The information processing method according to claim 1, wherein numerical data obtained by actually measuring the user's body parts is input.

9. The information processing method according to claim 1, wherein an image of the user's body part is used as input and a three-dimensional shape of the user's body part is generated by a generating AI (artificial intelligence).

10. The information processing method according to claim 1, wherein a distribution of physical quantities obtained by numerical calculation is input instead of the image, and a transfer function corresponding to the corresponding filter is calculated.

11. The information processing method according to claim 1, wherein the corresponding filter is generated in an offline environment.

12. The information processing method according to claim 1, wherein if the image contains obstacles that prevent the recognition of the shape of the user's body part, the corresponding filter is generated by removing those obstacles.

13. An information processing system comprising: an information processing device having: a shape recognition processing unit that performs shape recognition processing to recognize the shapes of user's body parts that affect the generation of wind noise based on an input image; a contribution calculation unit that determines the relationship between the shapes of the user's body parts and the shapes of body parts used to obtain multiple pieces of pre-calculated teacher data, and calculates a contribution indicating the degree to which the shapes of the body parts of the teacher data contribute to the shape of the user's body parts; an estimate calculation unit that calculates an estimate representing wind noise that is estimated to occur in accordance with the shape of the user's body parts by combining, in accordance with the contribution, weights for distances in feature space between features of dynamic mode decomposition in the shapes of the body parts of each piece of teacher data with the features of each piece of teacher data; and a correspondence filter generation unit that generates, based on the estimate, a correspondence filter that cancels the effect of wind noise that occurs in accordance with the shape of the user's body parts; and an audio device having a noise cancellation processing unit that applies noise cancellation processing using the correspondence filter to a speaker based on an external sound signal including wind noise picked up by a microphone.

14. A program for causing a computer of an information processing device to execute information processing, including: performing shape recognition processing to recognize, based on an input image, the shapes of a user's body parts that affect the generation of wind noise; determining the relationship between the shapes of the user's body parts and the shapes of body parts used to obtain multiple pieces of pre-calculated training data, and calculating a contribution indicating the degree to which the shapes of the body parts of the training data contribute to the shape of the user's body parts; calculating an estimated value representing wind noise that is estimated to occur in accordance with the shape of the user's body parts by combining, according to the contribution, weights for the distance in feature space between features of dynamic mode decomposition in the shapes of the body parts of each piece of training data with the features of each piece of training data; and generating, based on the estimated value, a corresponding filter that cancels out the influence of wind noise that occurs in accordance with the shape of the user's body parts.

Citation Information

Patent Citations

  • Signal processing device, signal processing program, and signal processing method

    WO2021225100A1