Information processing apparatus, information processing method, and information processing program
The information processing apparatus efficiently calculates personalized HRTF by estimating ear parameters from user images, addressing the inconvenience and computational burden of traditional HRTF methods.
Patent Information
- Application Number
- JP2024085576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-10
- Filing Date
- 2024-05-27
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2039-10-03
AI Technical Summary
Existing methods for calculating user-specific head-related transfer functions (HRTF) are cumbersome due to high computational load and require extensive measurement, making them inconvenient for individual use.
An information processing apparatus that uses a learned model to estimate ear parameters from an image of a user's ear, bypassing the need for 3D model generation and acoustic simulation, thereby calculating personalized HRTF efficiently.
Enables rapid calculation of user-specific HRTF without the burden of physical measurements or lengthy simulations, enhancing convenience and reducing processing time.
Smart Images

Figure 0007715249000014 
Figure 0007715249000015 
Figure 0007715249000016
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program. Specifically, it relates to a calculation process of a head-related transfer function.
Background Art
[0002] A technique of stereoscopically reproducing sound images in headphones or the like is used by using a head-related transfer function (hereinafter, may be referred to as HRTF (Head-Related Transfer Function)) that mathematically represents the way sound reaches the ears from a sound source.
[0003] Since the head-related transfer function has large individual differences, it is desirable to use a head-related transfer function generated for each individual when using it. For example, a technique is known in which a three-dimensional digital model of the head (hereinafter, referred to as a "3D model") is generated based on an image of a user's auricle, and the head-related transfer function of the user is calculated from the generated 3D model.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] According to the prior art, since the head-related transfer function calculated individually for each user can be used for information processing, the sense of sound image localization can be enhanced.
[0006] However, in the above prior art, since a 3D digital model is generated based on an image captured by a user and a head-related transfer function is calculated from the generated model, the processing load of the calculation becomes relatively large. That is, in the above prior art, it is presumed that it takes a long time to provide the head-related transfer function to the user who transmitted the image, so it is difficult to say that the convenience is high.
[0007] Therefore, in the present disclosure, an information processing apparatus, an information processing method, and an information processing program that can improve the convenience of the user in processing related to the head-related transfer function are proposed.
Means for Solving the Problem
[0008] In order to solve the above problems, an information processing apparatus according to one aspect of the present disclosure includes an acquisition unit that acquires a first image including an image of a user's ear, and uses a learned model that is learned to output a head-related transfer function corresponding to the ear when an image including the image of the ear is input, to acquire the head-related transfer function corresponding to the user based on the first image. The learned model is a learned model including an ear parameter estimation model that estimates ear parameters by learning the relationship between a second ear image and a plurality of ear images obtained by changing the texture of the three-dimensional data of the ear.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each of the following embodiments, the same parts are denoted by the same reference numerals, and redundant explanations are omitted.
[0011] (1. First Embodiment) [1-1. Overview of Information Processing According to the First Embodiment] First, using FIG. 1, the configuration of the information processing system 1 according to the present disclosure and the outline of the information processing executed by the information processing system 1 will be described. FIG. 1 is a diagram showing the outline of the information processing according to the first embodiment of the present disclosure. The information processing according to the first embodiment of the present disclosure is realized by the information processing system 1 shown in FIG. 1. The information processing system 1 includes an information processing apparatus 100 and a user terminal 10. The information processing apparatus 100 and the user terminal 10 communicate with each other using a wired or wireless network (not shown). Note that the number of each device constituting the information processing system 1 is not limited to that shown in the figure.
[0012] The information processing apparatus 100 is an example of the information processing apparatus according to the present disclosure, calculates the HRTF (Head Related Transfer Function) corresponding to each user, and provides the calculated HRTF. The information processing apparatus 100 is realized by, for example, a server apparatus or the like.
[0013] The user terminal 10 is an information processing terminal used by a user who intends to receive the provision of the HRTF. The user terminal 10 is realized by, for example, a smartphone having a photographing function or the like. In the example of FIG. 1, the user terminal 10 is used by a user U01 who is an example of a user.
[0014] The HRTF expresses, as a transfer function, the change in sound caused by surrounding objects including the human pinna (auricle) and the shape of the head. Generally, measurement data for obtaining the HRTF is acquired by measuring a measurement acoustic signal using a microphone or a dummy head microphone or the like worn inside the human auricle.
[0015] For example, the HRTF used in technologies such as 3D audio is often calculated using measurement data obtained by a dummy head microphone or the like, or the average value of measurement data obtained from a large number of humans. However, since there are large individual differences in the HRTF, in order to realize a more effective acoustic rendering effect, it is desirable to use the user's own HRTF. That is, by replacing the general HRTF with the user's own HRTF, a more immersive acoustic experience can be provided to the user.
[0016] However, there are various problems in individually measuring the HRTF of a user. For example, in order to obtain an HRTF that provides excellent acoustic effects, relatively high-density measurement data is required. To acquire high-density measurement data, measurement data of acoustic signals output to the user from various angles surrounding the user is necessary. Such measurements take a long time and impose a large physical burden on the user. Also, accurate measurements require measurements in an anechoic chamber or the like, resulting in a large cost burden. For this reason, in calculating the HRTF, reducing the burden on the user and reducing the measurement cost are issues.
[0017] Regarding the above problems, there is a technique for performing pseudo-measurement by representing the user's ears and head as a 3D model and performing acoustic simulation on the 3D model. According to such a technique, the user can cause the HRTF to be calculated without performing actual measurement in a measurement room by providing head scan data or an image of the head.
[0018] However, the generation process of the 3D model and the acoustic simulation for the 3D model have a very large computational processing load. For this reason, even if one attempts to incorporate user-specific HRTF into, for example, software that uses 3D sound using the above technique, there is a risk of a time lag of several tens of minutes or several hours. This is hardly convenient for the user. That is, from the perspective of enabling the user to utilize the HRTF, there is also an issue of improving the processing speed in calculating the HRTF.
[0019] As described above, there are various problems in obtaining HRTF corresponding to an individual user. The information processing apparatus 100 according to the present disclosure solves the above problems by the information processing of the present disclosure.
[0020] Specifically, when an image including an ear video is input, the information processing apparatus 100 calculates the HRTF corresponding to the user by using a learned model (hereinafter simply referred to as "model") that has been learned to output the HRTF corresponding to the ear. For example, when the information processing apparatus 100 acquires an image including a video of the ear of user U 01 from the user terminal 10, the information processing apparatus 100 inputs the image to the model and calculates the HRTF unique to user U 01. That is, the information processing apparatus 100 calculates the HRTF without going through processes such as generating a 3D model based on the image of user U 01 and performing acoustic simulation.
[0021] Thereby, the information processing apparatus 100 can calculate the HRTF in an extremely short time as compared with the case where acoustic simulation is executed. Hereinafter, the outline of the information processing executed by the information processing apparatus 100 according to the present disclosure will be described along the flow with reference to FIG. 1.
[0022] As shown in FIG. 1, user U 01 takes a picture of himself / herself from the side of the head in order to acquire an image including a video of his / her own ear (step S1). For example, user U 01 takes a picture of his / her own head using the camera provided in user terminal 10. In the present disclosure, the ear image is not limited to a general two-dimensional color image that can be captured by user terminal 10 or the like, and may be a monochrome image, a depth image including depth information, or an arbitrary combination thereof. Also, the image used in the processing of the present disclosure may be not one image but a plurality of images.
[0023] The user terminal 10 performs preprocessing for transmitting the image 5 obtained in step S1 to the information processing apparatus 100 (step S2). Specifically, as preprocessing, the user terminal 10 detects the video of the ear of user U 01 included in the image 5 and performs a process of cutting out the range of the detected ear video from the image 5. Details of the preprocessing such as ear detection will be described later.
[0024] The user terminal 10 creates an image 6 including an image of the ear of user U01 through preprocessing. Then, the user terminal 10 transmits the image 6 to the information processing device 100 (step S3). Note that a series of processes such as the process of creating the image 6 from the image 5 obtained by shooting and the process of transmitting the image 6 are executed by a program provided by the information processing device 100 (for example, an application for a smartphone), for example. In this way, in the information processing according to the present disclosure, only the image 6 obtained by cutting out only the image of the ear from the image 5 is transmitted on the network, and the image 5 in which an individual may be identified is not transmitted, so that a highly secure process can be realized. Further, the information processing device 100 can avoid the risk of handling personal information by acquiring only the image 6 without acquiring the image 5. Note that the information processing device 100 may acquire the image 5 from the user terminal 10 and perform a process of creating the image 6 on the information processing device 100 side. This configuration will be described later as a second embodiment.
[0025] The information processing device 100 acquires the image 6 transmitted from the user terminal 10. Then, the information processing device 100 inputs the image 6 into the model stored in the storage unit 120 (step S4). This model is a model learned to output the HRTF corresponding to the ear when a two-dimensional image including an image of the ear is input. That is, the information processing device 100 calculates the HRTF corresponding to the ear (in other words, user U01) included in the image 6 by inputting the image 6 into the model.
[0026] Then, the information processing device 100 provides the calculated HRTF to the user terminal 10 (step S5). In this way, user U01 can obtain his / her own unique HRTF without going through actual measurement processing, acoustic simulation of a 3D model, etc., as long as he / she takes a picture of his / her profile and prepares only the image 5. That is, the information processing device 100 can provide the HRTF to user U01 in an extremely short time without imposing a measurement burden on user U01. As a result, the information processing device 100 can improve the convenience of the user in the process related to the HRTF.
[0027] As described above, in the information processing according to the present disclosure, the processing speed is increased by calculating the HRTF using the model created through the learning process. In FIG. 1, the outline of the process of providing the HRTF to the user U01 among the information processing according to the present disclosure is shown. In FIGS. 2 and below, a series of information processing by the information processing apparatus 100 including the model learning process will be described in detail. Although the details will be described in FIGS. 2 and below, the "model" shown in FIG. 1 does not necessarily indicate a single model, and may be a combination of a plurality of models that output various numerical values.
[0028] [1-2. Overall Flow of Information Processing According to the Present Disclosure] Prior to the detailed description of the configuration and the like of the information processing apparatus 100, FIG. 2 shows the overall flow of the information processing executed by the information processing apparatus 100 according to the present disclosure. FIG. 2 is a conceptual diagram showing the overall flow of the information processing according to the present disclosure.
[0029] First, the information processing apparatus 100 collects data regarding the ear shapes of a plurality of persons, and creates an ear model based on the collected ear shapes (step S11). Note that the ear shape is not necessarily limited to a mold of a person's ear with plaster or the like, and any information may be used as long as it indicates the shape of a person's ear. In the present disclosure, the ear model is a model that outputs the corresponding ear shape when parameters indicating the characteristics of the ear (hereinafter referred to as "ear parameters") are input. The ear parameters can be obtained, for example, by performing principal component analysis on the ear shape based on data regarding the ear shape (for example, data obtained by subjecting the collected ear shape to a CT (Computed Tomography) scan). As a result, when the information processing apparatus 100 obtains the ear parameters, it can obtain data on the shape of the ear corresponding to the ear parameters (in other words, a 3D model simulating the ear).
[0030] After that, the information processing apparatus 100 generates an ear parameter estimation model based on the ear model (step S12). The information processing apparatus 100 can generate a large number of ear images by inputting ear parameters into the ear model generated in step S11. The ear parameters may be input randomly, or the ear parameters may be automatically generated according to an arbitrary rule (for example, if it is found that there is a specific tendency in the shape of the ear for each specific ethnic group, a rule may be derived based on this fact), and the generated values may be input. Therefore, the information processing apparatus 100 can generate a model that outputs the ear parameters corresponding to the ear when an image including the ear is input by learning the relationship between the generated ear image and the ear parameters that are the generation sources. Such a model is an ear parameter estimation model. As a result, if the information processing apparatus 100 obtains a two-dimensional image including an ear video, it can obtain the ear parameters corresponding to the ear. Then, if the information processing apparatus 100 obtains the ear parameters, it can obtain a 3D model of the ear included in the image using the ear model generated in step S11. Regarding the above learning, the relationship between an image of the ear of the person whose ear shape is digitized and the digitized ear converted into ear parameters may be learned. In this case, since learning is performed using actual photographed images instead of CG (Computer Graphics) images, it is assumed that the accuracy of the generated ear parameter estimation model can be further improved.
[0031] The information processing apparatus 100 performs acoustic simulation on the 3D model generated using the ear parameter estimation model, and calculates the unique HRTF corresponding to the 3D model (hereinafter, the HRTF generated corresponding to such an individual ear shape is referred to as "personalized HRTF") (step S13). That is, through the process from step S11 to step S13, the information processing apparatus 100 can realize a series of processes for calculating the personalized HRTF by acoustic simulation from an image including an ear.
[0032] Furthermore, the information processing apparatus 100 generates a large number of 3D models from randomly or regularly generated ear parameters, and repeats the process of performing acoustic simulation on the generated 3D models, thereby learning the relationship between the ear parameters and the personalized HRTF. That is, the information processing apparatus 100 generates an HRTF learning model based on the calculated personalized HRTF (step S14).
[0033] In the present disclosure, the HRTF learning model is a model that outputs a personalized HRTF corresponding to the ear parameters when the ear parameters are input. Thereby, when the ear parameters are obtained by the information processing apparatus 100, the personalized HRTF corresponding to the ear parameters can be obtained.
[0034] After that, when the information processing apparatus 100 acquires an image from the user, the information processing apparatus 100 calculates the personalized HRTF of the user by inputting the image (more precisely, the ear parameters of the ear included in the image) into the HRTF learning model (step S15). The process shown in step S15 corresponds to the series of processes shown in FIG. 1.
[0035] As described above, the information processing apparatus 100 can generate a plurality of models and perform information processing using the generated models, thereby calculating the personalized HRTF from the image acquired from the user. Note that the processes shown in FIG. 2 do not necessarily need to be executed in the order of steps S11 to S15. For example, after generating each model required for the process, the information processing apparatus 100 may only perform the process of step S15.
[0036] The overall flow of the information processing according to the present disclosure has been described above. In FIGS. 3 and below, the configurations of the information processing apparatus 100 and the user terminal 10 will be described, and the details of various learning processes will be described in order.
[0037] [1-3. Configuration of Information Processing Apparatus According to First Embodiment] Using FIG. 3, the configuration of the information processing apparatus 100 according to the first embodiment will be described. FIG. 3 is a diagram showing a configuration example of the information processing apparatus 100 according to the first embodiment of the present disclosure.
[0038] As shown in FIG. 3, the information processing apparatus 100 includes a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing apparatus 100 may have an input unit (for example, a keyboard, a mouse, etc.) that receives various operations from an administrator or the like who manages the information processing apparatus 100, and a display unit (for example, a liquid crystal display, etc.) for displaying various information.
[0039] The communication unit 110 is realized by, for example, a NIC (Network Interface Card) or the like. The communication unit 110 is connected to the network N (the Internet or the like) by wire or wirelessly, and transmits and receives information to and from the user terminal 10 or the like via the network N.
[0040] The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 120 stores various data used for the learning process and models generated by the learning process.
[0041] As shown in FIG. 3, the storage unit 120 includes an ear shape information storage unit 121, an ear model storage unit 122, an ear image storage unit 123, an ear parameter estimation model storage unit 124, an HRTF processing model storage unit 125, an HRTF learning data storage unit 126, and an HRTF learning model storage unit 127.
[0042] The ear shape information storage unit 121 stores information obtained by converting the ear shape actually collected from a human body into 3D model data (that is, information regarding the shape of the ear). Specifically, the ear shape information storage unit 121 stores data (such as 3D polygons) indicating a three-dimensional shape obtained by CT scanning the collected ear shape.
[0043] The ear model storage unit 122 stores the ear model according to the present disclosure. The ear model is a model that outputs the shape of a corresponding ear when ear parameters indicating ear features are input.
[0044] The ear parameters are obtained by performing principal component analysis on data indicating the shape of the ear type stored in the ear type information storage unit 121. That is, the ear parameters are obtained by statistically analyzing (principal component analysis) the 3D polygon of the ear, and quantifying the parts with large variations (characterizing the ear shape) in the ear. The ear parameters according to the present disclosure are represented by, for example, a combination of 10 numerical values, and each numerical value is represented by a numerical value from -10 to +10, for example. For example, ear parameters where all numerical values are "0" correspond to an ear having an average shape of the learning data (the collected ear types). Note that the information processing apparatus 100 may appropriately apply known techniques used, for example, in the generation process of a human face for the generation process of a model indicating the shape of the ear by principal component analysis. Further, the information processing apparatus 100 may generate ear parameters by appropriately using known analysis methods such as independent component analysis and other non-linear models, not limited to principal component analysis. Here, the ear parameters are not limited to those obtained by quantifying the parts with large variations in the ear, and may be, for example, parameters obtained by parameterizing features related to the shape of the ear such that the influence on HRTF becomes large.
[0045] The ear image storage unit 123 stores an image including a video of the ear. For example, the ear image storage unit 123 stores, as an ear image, a CG image obtained by rendering the shape of the ear (3D model of the ear) generated by the ear model. Further, the ear image storage unit 123 may store, as an ear image, an image including a video of the ear transmitted from the user.
[0046] Here, FIG. 4 shows an example of the ear image storage unit 123 according to the present disclosure. FIG. 4 is a diagram showing an example of the ear image storage unit 123 of the present disclosure. In the example shown in FIG. 4, the ear image storage unit 123 has items such as "ear parameters", "ear 3D model data", "head 3D model data", "ear image ID", and "image generation parameters". Further, the "image generation parameters" have sub-items such as "texture", "camera angle", "resolution", and "brightness".
[0047] The "ear parameters" are parameters indicating the characteristics of the shape of the ear. For example, the ear parameters are represented by numerical values in 10 dimensions or the like. The "ear 3D model data" is data indicating the three-dimensional shape of the ear reconstructed based on the ear parameters. The "head 3D model data" is data indicating the three-dimensional shape of the head to which the ear 3D model data is synthesized when reconstructing a 3D model of a person.
[0048] The "ear image ID" indicates identification information for identifying an ear image obtained by rendering a 3D model. As shown in FIG. 4, a plurality of ear images are generated from one 3D model by variously changing the parameters (image generation parameters) set during rendering.
[0049] The "image generation parameters" indicate the setting parameters in rendering for generating an image. The "texture" indicates the setting of the texture of the CG. The "camera angle" indicates the pseudo camera shooting angle when rendering a 3D model to obtain a two-dimensional image. The "resolution" indicates the resolution during rendering. The "brightness" indicates the brightness during rendering. The item of brightness may include setting data such as the angle of light (incident light) in rendering.
[0050] Note that in FIG. 4, the data of each item is conceptually described as "A01" or "B01", but actually, specific data corresponding to each item is stored in the data of each item. For example, in the item of "ear parameters", a list of 10 specific numerical values is stored. Similarly for other items, various numerical values and information corresponding to each item are stored.
[0051] That is, in the example shown in FIG. 4, the ear 3D model data generated by the ear parameter "A01" is "B01", and it is shown that the head 3D model data that constitutes the 3D model of the person in combination with the ear 3D model data is "C01". Also, it is shown that the ear images obtained from the generated 3D model of the person are a plurality of ear images identified by ear image IDs "D01", "D02", "D03", etc. Further, the ear image identified by the ear image ID "D01" has a texture of "E01", a camera angle of "F01", a resolution of "G01", and a brightness of "H01" as image generation parameters during rendering.
[0052] Returning to FIG. 3 and continuing the explanation. The ear parameter estimation model storage unit 124 stores an ear parameter estimation model. The ear parameter estimation model is a model that outputs an ear parameter corresponding to an ear when a two-dimensional image including an image of the ear is input.
[0053] The HRTF processing model storage unit 125 stores an HRTF processing model. Although details will be described later, the HRTF processing model performs a process of compressing the amount of information of the HRTF calculated by acoustic simulation or the like. In the following description, the HRTF compressed by the HRTF processing model may be referred to as an HRTF parameter.
[0054] The HRTF learning data storage unit 126 stores learning data for generating a model (an HRTF learning model described later) for calculating the HRTF from an image including an image of the ear. Specifically, the HRTF learning data storage unit 126 stores, as learning data, data in which an ear parameter indicating the shape of the ear and an HRTF corresponding to the shape of the ear specified based on the ear parameter are combined.
[0055] The HRTF learning model storage unit 127 stores the HRTF learning model. The HRTF learning model is a model that outputs the HRTF corresponding to an ear when an image including an ear video is input. For example, when the HRTF learning model acquires an image including an ear video, it causes the ear parameter estimation model to output the ear parameters corresponding to the ear, and further outputs the HRTF corresponding to the ear parameters.
[0056] The control unit 130 is realized, for example, by a program (for example, the information processing program according to the present disclosure) stored inside the information processing apparatus 100 being executed using a RAM (Random Access Memory) or the like as a work area by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like. Further, the control unit 130 is a controller and may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0057] As shown in FIG. 3, the control unit 130 includes a learning unit 131 and an estimation unit 140. The learning unit 131 includes a reception unit 132, an ear model learning unit 133, an image generation unit 134, an ear parameter learning unit 135, and an HRTF learning unit 136, and realizes or executes the functions and operations of information processing described below. The estimation unit 140 includes an acquisition unit 141, a calculation unit 142, and a provision unit 143, and realizes or executes the functions and operations of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 3, and may be other configurations as long as they perform the information processing described later.
[0058] First, the learning unit 131 will be described. The learning unit 131 performs learning processing on various data and generates various models used by the estimation unit 140.
[0059] Note that the learning unit 131 performs learning to generate a model based on various data. However, the learning process described below is just an example, and the type of learning process executed by the learning unit 131 is not specified to any particular type. For example, the learning unit 131 may generate a model using various learning algorithms such as neural networks, support vector machines, clustering, reinforcement learning, etc.
[0060] The reception unit 132 receives various information. For example, the reception unit 132 receives ear-shaped CT scan data collected from a human body. The reception unit 132 stores the received data in the ear-shaped information storage unit 121.
[0061] The ear model learning unit 133 performs learning processing related to the ear model and generates an ear model. The ear model learning unit 133 stores the generated ear model in the ear model storage unit 122.
[0062] Here, with reference to FIG. 5, an example of the learning process executed by the reception unit 132 and the ear model learning unit 133 will be described. FIG. 5 is a diagram showing an example of the learning process related to the ear model according to the present disclosure.
[0063] As shown in FIG. 5, the reception unit 132 receives the data collected and scanned from the ear shape, stores the received data in the ear-shaped information storage unit 121. Also, the reception unit 132 sends the received data to the ear model learning unit 133 (step S16).
[0064] The ear model learning unit 133 homologizes the acquired ear shape data to generate homologous data of the ear shape (step S17). Here, homologization means unifying the number of vertices of the 3D model and the composition of the polygons to be the same as those of a reference 3D model. In this case, care is taken to ensure that the shape does not change before and after homologization. Further, the ear model learning unit 133 performs principal component analysis on the homologous data (step S18). Thereby, the ear model learning unit 133 generates a model (ear model) that calculates ear parameters indicating the shape of the ear from the shape of the ear. The ear model learning unit 133 stores the generated ear model in the ear model storage unit 122.
[0065] Returning to FIG. 3 and continuing the explanation. The image generation unit 134 generates an image including an ear video. For example, the image generation unit 134 randomly generates ear parameters, inputs the generated ear parameters into the ear model to generate a 3D model of the ear. Further, the image generation unit 134 randomly generates parameters such as the texture of the generated 3D model (e.g., skin color), rendering quality (image quality, etc.), and camera angle in CG rendering (hereinafter referred to as "image generation parameters"). Then, the image generation unit 134 performs rendering by appropriately combining the generated 3D model and a plurality of image generation parameters to generate a CG image in which the shape and skin color of the ear vary diversely.
[0066] In the estimation process described later, the image transmitted from the user is used for processing. However, in the image transmitted from the user, it is highly assumed that the skin color of the user and the angle of the ear at the time of shooting may vary widely. Therefore, in such processing, there is a problem of accurately recognizing the ear video in any image transmitted from the user. The image generation unit 134 improves the accuracy of image recognition and solves the above problem by generating a large number of images corresponding to various situations as described above.
[0067] The ear parameter learning unit 135 generates an ear parameter estimation model by learning the relationship between an image including an ear image and ear parameters. The ear parameter learning unit 135 corresponds to the first learning unit according to the present disclosure. The image including the ear image may be an image actually capturing a person's ear, or may be a CG image generated based on ear parameters as described later.
[0068] For example, the ear parameter learning unit 135 generates an ear parameter estimation model by learning the relationship between an ear image obtained by rendering 3D data obtained by synthesizing 3D data of an ear generated based on ear parameters and 3D data of a head, and the ear parameters. Specifically, the ear parameter learning unit 135 learns the relationship between the CG image generated by the image generation unit 134 and the ear parameters. As described above, since the image generation unit 134 generates a CG image based on randomly or regularly set ear parameters, the ear parameters are uniquely determined for the CG image. Therefore, the ear parameter learning unit 135 can generate a model that outputs ear parameters corresponding to the ear image included in the image when a certain image is input by learning the relationship between the input CG image and the ear parameters. Note that the ear parameter learning unit 135 does not necessarily have to use, for learning, an ear image obtained by rendering synthesized 3D data of a head. That is, the ear parameter learning unit 135 may generate an ear parameter estimation model by learning the relationship between an ear image obtained by rendering only 3D data of an ear generated based on ear parameters and the ear parameters.
[0069] In addition, the ear parameter learning unit 135 generates an ear parameter estimation model by learning the relationship between a plurality of ear images in which the texture of the three-dimensional data of the ear or head, the camera angle in rendering, the brightness in rendering, etc. are changed, and the ear parameters common to the plurality of ear images. In this way, the ear parameter learning unit 135 learns using ear images in various forms, so that it can accurately output ear parameters regardless of what kind of image is input (for example, it can perform stable estimation against any change other than the ear parameters included in the input image), and can generate a stable and robust model.
[0070] Here, with reference to FIG. 6, an example of the learning process executed by the image generation unit 134 and the ear parameter learning unit 135 will be described. FIG. 6 is a diagram showing an example of the learning process related to the ear parameter estimation model according to the present disclosure.
[0071] As shown in FIG. 6, the image generation unit 134 refers to the ear model storage unit 122 (step S21) and acquires an ear model. In addition, the image generation unit 134 generates a random number corresponding to the ear parameter, a random number corresponding to the texture of the CG, the camera angle of the rendering, etc. (step S22). That is, the image generation unit 134 generates various parameters (image generation parameters) for generating an ear image.
[0072] Then, the image generation unit 134 acquires the ear parameter among the image generation parameters (step S23), inputs the acquired ear parameter into the ear model, and reconstructs the 3D model of the ear (step S24).
[0073] Subsequently, the image generation unit 134 acquires parameters such as CG textures among the image generation parameters (step S25), inputs the acquired parameters, and performs 3D CG rendering on the 3D model (step S26). Regarding the head used for rendering, for example, an average head of a plurality of persons (for example, a plurality of persons from whom ear shapes are collected) or a 3D model of a head used as a sample is used. Here, similar to the ear model, the 3D model of the head may be generated by homologating 3D data obtained by 3D scanning the heads of a plurality of persons. In this case, the image generation unit 134 can randomly generate a head 3D model by generating parameters using random numbers. Note that the image generation unit 134 can also generate various textures by using random numbers by creating a texture model generated by a similar method for the texture.
[0074] Thereby, the image generation unit 134 generates an image (ear image) including an image of the ear. Note that the image generation unit 134 can generate a plurality of ear images from one 3D model by varying parameters such as textures in various ways.
[0075] Here, an example of an ear image generated by the image generation unit 134 is shown with reference to FIG. 7. FIG. 7 is a diagram showing an example of the ear image generation process according to the present disclosure.
[0076] The image generation unit 134 generates a 3D model showing the three-dimensional shape of the ear using the randomly generated ear parameters (step S41). Further, the image generation unit 134 acquires a head 3D model generated based on data showing the three-dimensional shape of the average head of a plurality of persons (step S42). Then, the image generation unit 134 synthesizes the ear generated in step S41 and the head 3D model acquired in step S42 to generate a 3D model of a pseudo-person.
[0077] Subsequently, the image generation unit 134 performs a pseudo-photographing on the generated 3D model and performs a process (rendering) of generating a two-dimensional image from the 3D model. For example, the image generation unit 134 sets the front angle of the ear of the 3D model as a pseudo-photographing angle and generates an image in which the image of the ear is substantially in the center.
[0078] Here, the image generation unit 134 randomly inputs image generation parameters to the 3D model, thereby variously changing the CG texture (specifically, skin color, etc.), the quality of rendering (resolution, etc.), the position of the ear relative to the head, and the like. As a result, the image generation unit 134 can generate a large number of images with different skin colors, etc. (step S43).
[0079] The image group 20 shows a plurality of ear images generated by the image generation unit 134. In this way, the image generation unit 134 can improve the recognition accuracy of the ear images described later by generating a large number of diverse ear images.
[0080] Returning to FIG. 6 and continuing the description, the image generation unit 134 stores the generated ear image in the ear image storage unit 123 (step S27). Note that the image generation unit 134 stores the image generation parameters at the time of generating each image in the ear image storage unit 123 in association with the generated ear image (step S28). As a result, the image generation unit 134 can hold a large amount of ear images associated with ear parameters as learning data. For example, the image generation unit 134 can hold a large amount of ear images associated with ear parameters as learning data.
[0081] Subsequently, the ear parameter learning unit 135 refers to the ear image storage unit 123 (step S29) and acquires an ear image and ear parameters. Then, the ear parameter learning unit 135 learns the relationship between the ear image and the ear parameters and generates an ear parameter estimation model. The ear parameter learning unit 135 stores the generated ear parameter estimation model in the ear parameter estimation model storage unit 124 (step S30).
[0082] The ear parameter estimation model is generated using, for example, a convolutional neural network (CNN) that is useful for extracting feature amounts of an image. Note that a cost formula (cost function) in learning is represented by, for example, the following formula (1).
[0083] [Number]
[0084] In formula (1), “α true ” indicates the true value of the ear parameter. Also, “α est ” indicates the estimated value of the ear parameter. Also, “A ear ” indicates an ear model by principal component analysis. Also, the distance function on the right side indicates the L2 norm (Euclidean distance). Note that as the true value of the ear model parameter, for example, parameters indicating the ears of a person measured at the time of collecting the ear shape can be used. That is, the ear parameter used when generating the ear image is the true value, and the value output when the ear image is input to the ear parameter estimation model during learning becomes the estimated value. The information processing apparatus 100 updates the coefficient so as to minimize the value of the cost formula with respect to the current estimated value as learning processing.
[0085] Here, with reference to FIG. 8, the ear parameter estimation model generated by the ear parameter learning unit 135 will be described. FIG. 8 is a diagram for explaining the ear parameter estimation model according to the present disclosure.
[0086] When the information processing apparatus 100 acquires the ear image 30, the acquired ear image 30 is input to the ear parameter estimation model. The ear parameter estimation model has, for example, a convolutional neural network structure, and obtains a feature amount indicating the ear image 30 while dividing the input ear image 30 into rectangular portions for every several pixels. Finally, the ear parameter estimation model outputs an ear parameter corresponding to the video of the ear included in the ear image 30 as a feature amount indicating the ear image 30 (step S45).<s
[0087] Note that the information processing apparatus 100 can reconstruct an ear shape (3D model) corresponding to the ear included in the ear image 30 by inputting the output ear parameters into the ear model. The 3D model 40 shown in FIG. 8 shows a 3D model obtained by reconstructing the ear included in the ear image 30 in CG based on the ear parameters.
[0088] Returning to FIG. 3 and continuing the description, the HRTF learning unit 136 generates various models related to the HRTF by learning the relationship between the information on the ear shape and the HRTF. For example, the HRTF learning unit 136 generates a learned model for calculating the HRTF by learning the relationship between an image including an ear video and the HRTF corresponding to the ear. The HRTF learning unit 136 corresponds to the second learning unit according to the present disclosure.
[0089] For example, the HRTF learning unit 136 performs acoustic simulation on the 3D data obtained by synthesizing the 3D data of the ear generated based on the ear parameters and the 3D data of the head, and generates a learned model by learning the relationship between the HRTF obtained by the acoustic simulation and the ear parameters.
[0090] Alternatively, the HRTF learning unit 136 may compress the amount of information of the HRTF obtained by the acoustic simulation and generate a learned model by learning the relationship between the compressed HRTF and the ear parameters.
[0091] Alternatively, the HRTF learning unit 136 may set a listening point for the 3D data of the ear generated based on the ear parameters and perform acoustic simulation using the set listening point. The listening point is a virtual setting of the position where a human is assumed to listen to sound. For example, the position of the listening point corresponds to the position where the microphone is installed in a dummy head microphone (such as the entrance of the external auditory canal in the dummy head).
[0092] Regarding each process of the above-described HRTF learning unit 136, FIG. 9 shows the flow of generation processing of various models related to HRTF. FIG. 9 is a diagram showing an overview of the flow of generation processing of various models related to HRTF.
[0093] In FIG. 9, an example is shown in which the HRTF learning unit 136 performs a predetermined learning process based on an image transmitted from the user. In this case, the user uses the user terminal 10 to photograph his / her ear (accurately, the head including the ear) (step S51). Then, the user terminal 10 performs preprocessing of specifying the range including the ear image from the photographed photo, cutting out the specified range, and acquiring an ear image (step S52).
[0094] After that, the HRTF learning unit 136 calculates the ear parameters of the ear included in the ear image transmitted from the user using the ear parameter estimation model (step S53). Further, the HRTF learning unit 136 reconstructs a 3D model of the ear based on the ear parameters, and further combines the reconstructed ear with the head 3D model to generate a 3D model of a person (step S54).
[0095] Subsequently, the HRTF learning unit 136 performs acoustic simulation on the generated 3D model and obtains the personalized HRTF of the 3D model (step S55). Thereby, the HRTF learning unit 136 can obtain learning data in which the ear included in the ear image transmitted from the user is associated with the personalized HRTF.
[0096] Note that, in the example of FIG. 9, an example of generating learning data in which the personalized HRTF obtained by acoustic simulation is associated with ear data is shown, but the HRTF learning unit 136 does not necessarily need to obtain the personalized HRTF by acoustic simulation in some cases. For example, when the personalized HRTF of a person from whom the ear shape has been collected (the HRTF obtained using a measuring device in an anechoic chamber or the like) is obtained, the HRTF learning unit 136 may acquire learning data in which the actually measured personalized HRTF is associated with the ear shape (ear parameters) of the person.
[0097] The HRTF learning unit 136 automatically sets the listening point in the 3D model of a person during acoustic simulation. This will be described with reference to FIG. 10. FIG. 10 is a diagram for explaining the reconstruction of the 3D model according to the present disclosure.
[0098] The HRTF learning unit 136 reconstructs the ear 3D model from randomly generated ear parameters (step S71). Subsequently, the HRTF learning unit 136 generates a 3D model of a person by combining the ear 3D model with the head 3D model (step S72). Further, the HRTF learning unit 136 sets the listening point 60 of the sound source based on the shape of the ear in the 3D model (step S73). For example, the HRTF learning unit 136 can automatically set the listening point by learning information associating the shape of the ear with the position of the listening point of the sound source in advance. For example, when generating the 3D model, the HRTF learning unit 136 estimates the position of the listening point in the 3D model and automatically sets the listening point. The listening point corresponds to, for example, the ear canal of the ear, and generally, the position can be estimated from the shape of the ear.
[0099] Thereafter, the HRTF learning unit 136 remeshes the generated 3D model so as to satisfy the computational constraints of the 3D model in acoustic simulation (step S74). This is because, in the simulation of the 3D model, for example, the upper limit on the total number of polygons and the length of the edges connecting the vertices may be determined by the simulation conditions. That is, before subjecting the generated 3D model to simulation, the HRTF learning unit 136 appropriately remeshes the 3D model so as to satisfy the computational constraints and performs processing so that the simulation can be appropriately performed. Then, the HRTF learning unit 136 performs acoustic simulation on the generated 3D model and the set listening point 60, and calculates the personalized HRTF (step S75).
[0100] Next, with reference to FIG. 11, a detailed flow of the process for generating the model related to HRTF will be described. FIG. 11 is a diagram for explaining the details of the process for generating the model related to HRTF.
[0101] As shown in FIG. 10, after performing head synthesis (step S81), the HRTF learning unit 136 performs acoustic simulation (step S82). The HRTF learning unit 136 analyzes the measurement data obtained by the acoustic simulation (HRTF post-processing) and calculates numerical values indicating personalized HRTF (step S83). Note that the HRTF post-processing means, for example, obtaining the HRTF by performing a Fourier transform on the HRIF (Head-Related Impulse Response) obtained by the acoustic simulation.
[0102] Here, the HRTF learning unit 136 refers to the HRTF processing model storage unit 125 (step S84) and inputs the calculated HRTF into a model (HRTF processing model) for processing the HRTF. Thereby, the HRTF learning unit 136 obtains an HRTF with reduced dimensions (step S85). That is, the HRTF learning unit 136 outputs HRTF parameters, which are HRTF with reduced dimensions, from the HRTF processing model (step S86).
[0103] In this way, the HRTF learning unit 136 does not use the HRTF obtained by the acoustic simulation as it is for processing, but performs processing using the HRTF parameters with reduced dimensions. This is because the HRTF is a function with a very large number of dimensions, and when directly performing the model generation process or calculation process, the load of the calculation process becomes large.
[0104] The HRTF learning unit 136 associates the data related to the head for which the acoustic simulation was performed (data that is the basis for head synthesis, such as ear parameters, etc.) with the calculated HRTF parameters and stores them in the HRTF learning data storage unit 126 (step S87).
[0105] After that, the HRTF learning unit 136 newly and randomly generates different ear parameters (step S88), and performs head synthesis using the newly generated ear parameters (step S89). By repeating steps S81 to S89, the HRTF learning unit 136 collects learning data required for learning.
[0106] After that, when sufficient learning data has been accumulated, the HRTF learning unit 136 refers to the HRTF learning data storage unit 126 (step S90), and learns the relationship between the ear parameters and the HRTF (specifically, the HRTF parameters) (step S91). Through such learning, the HRTF learning unit 136 generates an HRTF learning model for directly obtaining the HRTF from the ear parameters, and stores the generated HRTF learning model in the HRTF learning model storage unit 127.
[0107] Next, with reference to FIG. 12, the relationship between the HRTF and the HRTF parameters will be described. FIG. 12 is a diagram for explaining the compression and restoration of the HRTF according to the present disclosure.
[0108] As shown in FIG. 12, the HRTF learning unit 136 performs FFT (Fast Fourier Transform) on the HRTF obtained by acoustic simulation (in the example of FIG. 12, assuming 1000 directions × 500 taps) (step S101). Through such processing, the HRTF learning unit 136 extracts the amplitude characteristics (step S102), and performs a process of thinning out, for example, frequency components with low sensitivity in terms of auditory sensation (step S103). Specifically, the HRTF is represented by a function HRTF(θ, φ, f) with respect to the angle (θ, φ) and the frequency (f). At this time, if the number of frequency bins is k, the frequency f input to the function is f k = f0, f1, f2, ···, f k-1 That is, the HRTF has complex k dimensions for one direction and one ear. Here, the HRTF after the Nyquist frequency (f k / 2 ) is the frequency f k / 2Since it is a folding of the previous complex conjugate, in information processing, as frequency bins, only (k / 2)+1 bins from f0 = 0 to the Nyquist frequency (f k / 2 ) can be used. Also, for at least one or more frequency bins, the absolute value can be used. For example, if all frequencies from f0 to f k / 2 are converted to absolute values, the converted function H2 is represented by the following equation (2).
[0109]
Equation
[0110] That is, the HRTF learning unit 136 can compress the dimension of the original HRTF to a real number (k / 2)+1 dimension. Further, the HRTF learning unit 136 can perform frequency compression on H2 of the above equation (2) and further reduce it to a dimension less than (k / 2)+1. There are various known methods for dimension compression. For example, the HRTF learning unit 136 uses a method such as performing cepstrum transformation on the function and obtaining only one or more frequency bins less than (k / 2)+1. As an example, the HRTF learning unit 136 obtains the average value of a plurality of frequency bins and reduces the dimension based on the average value. For example, when the frequency bin is represented by the following equation (3) (a l , L, l are integers of 0 or more respectively), for l that satisfies f al ≦f´ l <f al+1 , the new function H3 is represented by the following equation (4).
[0111]
Equation
[0112]
Equation
[0113] As a result, the HRTF learning unit 136 can reduce the function H2, which was represented in (K / 2)+1 dimensions, to L dimensions. Note that the method for obtaining the average value is not limited to the above, and for example, it may be obtained using the root mean square or weighted average. As a result, the HRTF is reduced to, for example, about 1000 directions × 50 dimensions. When restoring the dimensions reduced by the function H3 (in steps S110 etc. described later), the HRTF learning unit 136 can restore using various methods such as linear interpolation or spline interpolation. The restored function H'2 is expected to have smoother characteristics than the function H2, but by devising the selection method of a l , H'2(θ, φ, k) with less influence on the auditory sensation can be obtained. As an example, for higher frequencies, the frequency interval between f al and f al+1 is increased by selecting a l . Such a device may be available.
[0114] The HRTF learning unit 136 further performs spherical harmonic fitting on the HRTF with reduced dimensions, whereby the amount of information is compressed to about 50 coefficients × 50 dimensions (step S104). Here, spherical harmonic fitting means performing fitting in the spatial direction for each compressed frequency using spherical harmonic functions. The relationship between the HRTF and the spherical harmonic function is shown by the following formula (5).
[0115]
Equation
[0116] As shown in the above formula (5), the spherical harmonic function Y is represented by the coefficient h nm (f). By truncating the number of dimensions n at this time with a certain finite N, the dimension of the coefficient h nm (f) can be made smaller than the number of dimensions (number of directions) of the original HRTF. This means ignoring spatially too fine amplitudes that are unnecessary for human perception and obtaining only a smooth shape. Note that the vector h of the coefficient h nm = (h 00 , h 1-1, ···) T To obtain it, for example, the least squares method or the like is used.
[0117]
Number
[0118] That is, when Y in the above formula (6) is a matrix of spherical harmonic functions and H is a matrix of spherical harmonic functions, h that minimizes E on the left side is obtained. Note that since the second term on the right side of the above formula (6) is a regularization term, any value of λ can be selected (for example, λ = 0 may also be used). Then, the above h is represented by the following formula (7).
[0119]
Number
[0120] By using the above formula (7), the HRTF learning unit 136 can obtain each h corresponding to the required frequency. Further, the HRTF learning unit 136 compresses the amount of information of the HRTF so that it can be expressed in approximately several hundred dimensions by performing dimensional compression by principal component analysis (step S105). Such information becomes the HRTF parameter (step S106).
[0121] Note that when spherical harmonic fitting is performed after frequency decimation, the value of the above f becomes the representative frequency after decimation. Further, the HRTF learning unit 136 may perform frequency decimation after spherical harmonic fitting. Also, the method of spatially compressing the dimension is not limited to linear combinations such as spherical harmonic functions and principal component analysis, and any method may be used. For example, the HRTF learning unit 136 may use a non-linear method such as kernel principal component analysis. Further, the HRTF learning unit 136 may change the truncation order N of the spherical harmonic function according to the frequency f and use a value such as N(f). Also, coefficients h not used in the number of dimensions or order from 0 to N nmThere may also be. Further, the HRTF learning unit 136 may obtain each of the left and right HRTFs, or may obtain them after converting them into the sum and difference of the left and right, etc. Further, the HRTF to be fitted may be one that has undergone various conversions such as the absolute value of the amplitude or its logarithmic expression.
[0122] Subsequently, the HRTF learning unit 136 can decode the HRTF by performing processing in the reverse flow from step S101 to step S106. First, the HRTF learning unit 136 acquires the HRTF parameters (step S107) and performs restoration of dimensional compression by principal component analysis (step S108). Further, the HRTF learning unit 136 performs spherical harmonic reconstruction processing (step S109) and performs frequency interpolation (step S110). Further, the HRTF learning unit 136 obtains the amplitude characteristics (step S111) and performs minimum phase restoration (step S112). For minimum phase restoration, various known methods may be used. For example, the HRTF learning unit 136 performs an inverse Fourier transform (IFFT (Inverse Fast Fourier Transform)) on the logarithm of the function H´1(θ, φ, k), which is a function obtained by folding back the function H´2 after the Nyquist frequency and restoring it, and takes the real part thereof. Further, appropriate window processing is performed in this region, and the exponential function is inverse Fourier transformed and the real part is taken to perform minimum phase restoration. For example, the following relational expressions (8) hold respectively.
[0123]
Equation
[0124] Note that the HRTF learning unit 136 may add the estimated ITD (Interaural Time Difference) or the previously prepared ITD to the left and right HRIRs (h m ) that have been minimum phase restored. The ITD is obtained based on the difference in group delay of the left and right HRIRs, for example, by the following formulas (9) and (10).
[0125]
Number
[0126]
Number
[0127] Alternatively, the ITD may be calculated by obtaining the cross-correlation on the time axis for the left and right sides and defining the time at which the correlation coefficient is maximized as the ITD. In this case, the ITD is obtained, for example, by the following expressions (11) and (12).
[0128]
Number
[0129]
Number
[0130] For example, when the HRTF learning unit 136 delays the left HRIR by d samples more than the right for the left and right HRIRs, a relational expression such as the following expression (13) is used.
[0131]
Number
[0132] At this time, h in the above expression (13) L is an impulse response that is d longer than h m、L but the length is h m、LIn order to make it the same, the latter part of the above formula (13) is deleted. At this time, the HRTF learning unit 136 may perform processing such as an arbitrary window, a rectangular window, or a Hann window. Note that the HRTF learning unit 136 may not only add an ITD relatively for each direction, but also add a delay including the relative time difference between directions throughout the space. In that case, the HRTF learning unit 136 acquires information indicating not only the ITD but also the relative time difference between directions. Further, when the ITD is a function of frequency, the HRTF learning unit 136 may add the ITD in the frequency domain, or may obtain a representative value or an average value and add the ITD. Then, after obtaining the HRIR in the original format, the HRTF learning unit 136 obtains the HRTF by performing an inverse Fourier transform.
[0133] In this way, the HRTF learning unit 136 may perform compression on the HRTF parameters with less information amount than the original HRTF, and perform the generation process of the HRTF learning model and the calculation process of the HRTF described later in the compressed format. Further, as described above, the compression of the HRTF performs dimensionality reduction using auditory characteristics, for example, by utilizing the fact that the human auditory sense is not sensitive to phase changes, or by preferentially thinning out frequencies that hardly affect the auditory sense. Thereby, the HRTF learning unit 136 can speed up information processing without impairing the sense of localization in the auditory sense, which is a characteristic of the HRTF.
[0134] Returning to FIG. 3, the description will be continued. The estimation unit 140 performs an estimation process of the HRTF corresponding to the user based on the image transmitted from the user.
[0135] The acquisition unit 141 acquires an image including an image of the user's ear. For example, the acquisition unit 141 acquires an ear image in which only the periphery of the user's ear is cut out from the image taken by the user terminal 10.
[0136] Further, the acquisition unit 141 may acquire ear parameters indicating the characteristics of the ear included in the image by inputting the acquired ear image into the ear parameter estimation model.
[0137] When an image including an ear image is input, the calculation unit 142 calculates a HRTF corresponding to the user (personalized HRTF) based on the image acquired by the acquisition unit 141, using a learned model (HRTF learning model) trained to output a HRTF corresponding to the ear when the image including the ear image is input.
[0138] Specifically, the calculation unit 142 calculates a personalized HRTF corresponding to the user by inputting the ear parameters acquired by the acquisition unit 141 into the HRTF learning model.
[0139] Note that when calculating the personalized HRTF, the calculation unit 142 may first calculate HRTF parameters and then decode the calculated HRTF parameters to calculate the HRTF. In this way, by performing a series of processes in a state where the information amount of the HRTF is compressed, the calculation unit 142 can speed up the process. In addition, since the calculation unit 142 can avoid outputting strange HRTFs that are not represented by the HRTF reduction model, stable output can be performed.
[0140] The providing unit 143 provides the HRTF calculated by the calculation unit 142 to the user via the network N.
[0141] Here, with reference to FIG. 13, the flow of the process of estimating the HRTF from the image will be described. FIG. 13 is a diagram showing the flow of the HRTF estimation process according to the present disclosure.
[0142] FIG. 13 shows an example in which the estimation unit 140 performs an estimation process of the HRTF corresponding to the ear included in the image based on the image transmitted from the user. In this case, the user uses the user terminal 10 to photograph his / her own ear (accurately, the head including the ear) (step S131). Then, the user terminal 10 performs preprocessing of specifying the range including the ear image from the photographed photo, cutting out the specified range, and acquiring the ear image (step S132).
[0143] When the acquisition unit 141 acquires an ear image transmitted from a user, the acquired ear image is input into a learned model. Specifically, the acquisition unit 141 inputs the ear image into an ear parameter estimation model. The ear parameter estimation model outputs an ear parameter corresponding to the video of the ear included in the ear image as a feature amount indicating the ear image. Thereby, the acquisition unit 141 acquires an ear parameter corresponding to the image (step S133).
[0144] The calculation unit 142 inputs the acquired ear parameter into an HRTF learning model and calculates a personalized HRTF corresponding to the ear image (step S134). The providing unit 143 provides (transmits) the calculated personalized HRTF to the user terminal 10 which is the transmission source of the image.
[0145] As described above, when various models are generated by the learning unit 131, the information processing apparatus 100 can perform a series of processes from the acquisition of the ear image to the provision of the personalized HRTF. Thereby, the information processing apparatus 100 can improve the convenience of the user regarding the provision of the HRTF.
[0146] In the example of FIG. 13, as an example of the learned model, a combination of an ear parameter estimation model and an HRTF learning model is shown, but the combination of the learned models is not limited to this example. Further, the learned model may be a combination of an ear parameter estimation model and an HRTF learning model individually, or may be configured as one model that performs processes corresponding to the ear parameter estimation model and the HRTF learning model.
[0147] [1-4. Configuration of User Terminal According to First Embodiment] As shown in FIG. 13, in the first embodiment, the user terminal 10 performs photographing of the user's profile and generation of ear images. Here, the configuration of the user terminal 10 according to the first embodiment will be described. FIG. 14 is a diagram showing a configuration example of the user terminal 10 according to the first embodiment of the present disclosure. As shown in FIG. 14, the user terminal 10 includes a communication unit 11, an input unit 12, a display unit 13, a detection unit 14, a storage unit 15, and a control unit 16.
[0148] The communication unit 11 is realized by, for example, a NIC or the like. Such a communication unit 11 is connected to the network N by wire or wirelessly, and transmits and receives information to and from an information processing apparatus 100 or the like via the network N.
[0149] The input unit 12 is an input device that receives various operations from the user. For example, the input unit 12 is realized by operation keys or the like provided in the user terminal 10. The display unit 13 is a display device for displaying various information. For example, the display unit 13 is realized by a liquid crystal display or the like. When a touch panel is adopted in the user terminal 10, a part of the input unit 12 and the display unit 13 are integrated.
[0150] The detection unit 14 is a general term for various sensors, and detects various information related to the user terminal 10. Specifically, the detection unit 14 detects operations of the user on the user terminal 10, location information where the user terminal 10 is located, information related to devices connected to the user terminal 10, the environment in the user terminal 10, and the like.
[0151] Further, as an example of the sensor, the detection unit 14 has a lens and an image sensor for performing photographing. That is, when the user activates an application that operates the photographing function, for example, the detection unit 14 exhibits a function as a camera.
[0152] The storage unit 15 stores various types of information. The storage unit 15 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 15 stores, for example, images taken by a user.
[0153] The control unit 16 is a controller and is realized, for example, by various programs stored in a storage device inside the user terminal 10 being executed with the RAM as a work area by a CPU, an MPU, or the like. Also, the control unit 16 is a controller and is realized, for example, by an integrated circuit such as an ASIC or an FPGA.
[0154] As shown in FIG. 14, the control unit 16 includes an acquisition unit 161, a preprocessing unit 162, a transmission unit 164, and a reception unit 165, and realizes or executes the information processing functions and operations described below. Also, the preprocessing unit 162 includes a posture detection unit 163A and an ear detection unit 163B. Note that the internal configuration of the control unit 16 is not limited to the configuration shown in FIG. 14, and other configurations may be used as long as they perform the information processing described later.
[0155] The acquisition unit 161 acquires various types of information. For example, the acquisition unit 161 acquires an image taken by the detection unit 14.
[0156] The posture detection unit 163A reads the image acquired by the acquisition unit 161 and detects the posture of the user included in the image.
[0157] The ear detection unit 163B detects the range (ear video) including the user's ear in the image based on the posture of the user detected by the posture detection unit 163A. Specifically, the ear detection unit 163B identifies the user's ear video from an image including the entire video of the user's head and detects the identified range as an ear image.
[0158] For example, the ear detection unit 163B identifies the range including the ear video based on the relationship between the feature points of the user's head included in the entire image and the posture of the user.
[0159] Further, when the posture detection unit 163A or the ear detection unit 163B cannot specify the range including the image of the ear based on the relationship between the feature points of the user's head included in the entire image and the user's posture, an image different from the entire image and including the image of the entire user's head may be newly requested from the user. Specifically, the posture detection unit 163A or the ear detection unit 163B displays, on the display unit 13, a message indicating that there is a possibility that the information processing according to the present disclosure cannot be appropriately performed in the image of the user's profile taken by the user, and prompts the user to retake the image. Note that the posture detection unit 163A or the ear detection unit 163B is not limited to the case where the range including the image of the ear cannot be specified. For example, even when the camera angle used during the learning of the ear parameter estimation model and the user's posture are separated beyond a certain threshold, the user may be prompted to retake the image. Further, the posture detection unit 163A or the ear detection unit 163B may generate correction information for correcting the posture and position of the user in the image instead of detecting the user's ear image as preprocessing. The correction information is, for example, information indicating the amount of rotation for rotating the range including the image of the ear according to the inclination and rotation of the feature points of the user's head. Such information is generated based on the user's posture, the positional relationship between the user's profile and the detected position of the ear, etc., as will be described later. In this case, the posture detection unit 163A or the ear detection unit 163B may correct the rotation of the entire image based on the correction information to specify the image of the user's ear, and detect the specified range as an ear image. Further, the posture detection unit 163A or the ear detection unit 163B may transmit the entire image to the information processing apparatus 100 together with the generated correction information. In this case, the information processing apparatus 100 corrects the rotation of the entire image based on the correction information transmitted together with the entire image to specify the image of the user's ear, and performs, by itself, the preprocessing of detecting the specified range as an ear image.
[0160] Here, with reference to FIG. 15, the flow of the preprocessing executed by the preprocessing unit 162 (posture detection unit 163A and ear detection unit 163B) will be described. FIG. 15 is a diagram showing the flow of the detection process according to the present disclosure.
[0161] As shown in FIG. 15, when the user's profile is photographed by the user, the acquisition unit 161 acquires the entire image 50 (step S141).
[0162] The posture detection unit 163A detects the user's profile in the acquired entire image 50 (step S142). For example, the posture detection unit 163A uses a known technique such as a face detection process of a person to specify a range in the entire image 50 that includes the video of the user's profile.
[0163] Here, as shown in the image 51, the posture detection unit 163A detects feature points included in the user's profile. For example, the posture detection unit 163A detects feature points such as a portion protruding in the horizontal direction in the profile (specifically, the apex of the user's nose), the apex of the head, the position of the mouth, the position of the jaw, etc. Further, the posture detection unit 163A detects the position of the user's ear, bangs, etc. from information such as the boundary between the hair and the skin. Further, the posture detection unit 163A detects the position of the user's eyes, etc. from the color information of the profile video.
[0164] Then, the posture detection unit 163A detects the user's posture based on the detected feature points (step S143). For example, the posture detection unit 163A detects the posture of the user's head from the three-dimensional arrangement of the feature points as shown in the image 54.
[0165] Such a posture detection process is a process for ensuring that the posture in the ear image transmitted by the user does not deviate significantly from the posture of the 3D model used during learning. That is, when an image with a significantly different posture from the 3D model is transmitted from the user terminal 10, there is a possibility that the information processing device 100 cannot appropriately perform ear image recognition due to the discrepancy between the learning data and the transmitted ear image.
[0166] Therefore, the posture detection unit 163A determines whether the average value of the angles at the time of rendering in the head 3D model 55 used for learning and the angle obtained from the image 54 are within a predetermined threshold value, and determines whether the user has taken an appropriate shot (step S144). For example, during the learning of the ear parameter estimation model, it is assumed that the angle φ formed by the direction of the camera at the time of rendering the head 3D model 55 and the line segment connecting the head vertex and a predetermined position of the ear (for example, the entrance of the external auditory canal, etc.) is within a predetermined numerical value. Similarly, during the learning of the ear parameter estimation model, it is assumed that the angle θ formed by the direction of the camera and the line segment connecting the nose vertex and the predetermined position of the ear is within a predetermined numerical value. This is because, in order to improve the image recognition accuracy, the ear images used for learning do not deviate significantly from the images showing the human profile. That is, the posture detection unit 163A determines whether the image transmitted from the user maintains an angle that can be recognized as an image showing the human profile, similar to the image during learning.
[0167] When the posture detection unit 163A determines that the user has not taken an appropriate shot (for example, when the tip of the nose is facing downward beyond the predetermined threshold value on the user's face, etc.), it performs processing such as displaying a message instructing the user to retake the shot, and acquires a newly taken image (step S145).
[0168] On the other hand, when it is determined that the user has taken an appropriate shot (step S146), the ear detection unit 163B identifies the range 57 in which the ear video is included from the image 56, and cuts out the range 57 (step S147). Thereby, the ear detection unit 163B acquires the ear image 58.
[0169] By performing the detection process shown in FIG. 15, the information processing apparatus 100 can determine whether the ear is tilted due to a poor shooting state or whether the actual angle of the user's ear is tilted, and calculate the HRTF.
[0170] Further, as described above, the user terminal 10 can cut out the ear image from the entire image of the side face, and instead of transmitting the entire image including the user's face, only transmit the ear image for processing. Thereby, the user terminal 10 can prevent the leakage of personal information and enhance the security of information processing. Note that the user terminal 10 is not limited to the above detection method, and may perform a process of cutting out the ear image from the entire image of the side face by using an image recognition technique based on machine learning or the like to detect the user's ear included in the image.
[0171] Returning to FIG. 14 and continuing the description. The transmission unit 164 transmits the ear image generated based on the range detected by the ear detection unit 163B to the information processing apparatus 100.
[0172] The reception unit 165 receives the personalized HRTF provided from the information processing apparatus 100. For example, the reception unit 165 can realize 3D sound optimized for the user individual by convolving the received personalized HRTF with music or sound in a voice playback application or the like.
[0173] (2. Second Embodiment) Next, the second embodiment will be described. In the first embodiment described above, an example in which the user terminal 10 cuts out only the video of the ear from the image taken by the user to generate an ear image was shown. The information processing apparatus 100A according to the second embodiment performs a process of cutting out only the video of the ear by itself instead of the user terminal 10.
[0174] The configuration of the information processing apparatus 100A according to the second embodiment will be described with reference to FIG. 16. FIG. 16 is a diagram showing a configuration example of the information processing apparatus 100A according to the second embodiment of the present disclosure. As shown in FIG. 16, the information processing apparatus 100A further includes a preprocessing unit 144 (a posture detection unit 145A and an ear detection unit 145B) as compared with the first embodiment.
[0175] The posture detection unit 145A performs the same processing as the posture detection unit 163A according to the first embodiment. Also, the ear detection unit 145B performs the same processing as the ear detection unit 163B according to the first embodiment. That is, the information processing apparatus 100A according to the second embodiment executes the preprocessing executed by the user terminal 10 according to the first embodiment on its own device.
[0176] In the second embodiment, the acquisition unit 141 acquires the entire image taken by the user of a side face from the user terminal 10. Then, the posture detection unit 145A and the ear detection unit 145B perform the same processing as the processing described with reference to FIG. 15, and generate an ear image from the entire image. The calculation unit 142 calculates the personalized HRTF from the ear image generated by the posture detection unit 145A and the ear detection unit 145B.
[0177] As described above, according to the information processing apparatus 100A according to the second embodiment, the user can receive the provision of the personalized HRTF only by taking and transmitting an image. Also, according to the configuration of the second embodiment, since it is not necessary to execute the preprocessing in the user terminal 10, for example, the processing load on the user terminal 10 can be reduced. Also, generally, since it is assumed that the server apparatus (information processing apparatus 100) is faster in processing than the user terminal 10, according to the configuration of the second embodiment, the overall speed of the information processing according to the present disclosure can be improved. Note that when correction information is transmitted together with the entire image, the posture detection unit 145A and the ear detection unit 145B may correct the rotation of the entire image based on the correction information included in the entire image to identify the video of the user's ear, and detect the identified range as an ear image.
[0178] (3. Other Embodiments) The processing according to each of the above-described embodiments may be implemented in various different forms other than the above-described embodiments.
[0179] In addition, among the processes described in each of the above embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.
[0180] In addition, each component of each device shown in the drawings is a functional concept, and it is not necessarily physically configured as shown in the drawings. That is, the specific form of the distribution and integration of each device is not limited to that shown in the drawings, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc.
[0181] In addition, the above-described embodiments and modifications can be appropriately combined within a range that does not conflict with the processing content.
[0182] In addition, the effects described in this specification are merely examples and are not limiting, and there may be other effects.
[0183] (4. Hardware Configuration) Information devices such as the information processing apparatus 100 and the user terminal 10 according to the above-described embodiments are realized by a computer 1000 having a configuration as shown in FIG. 17, for example. Hereinafter, the information processing apparatus 100 according to the first embodiment will be described as an example. FIG. 17 is a hardware configuration diagram showing an example of a computer 1000 that realizes the functions of the information processing apparatus 100. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. Each part of the computer 1000 is connected by a bus 1050.
[0184] The CPU 1100 operates based on programs stored in the ROM 1300 or the HDD 1400 and controls each part. For example, the CPU 1100 expands the programs stored in the ROM 1300 or the HDD 1400 to the RAM 1200 and executes processes corresponding to various programs.
[0185] The ROM 1300 stores a boot program such as the BIOS (Basic Input Output System) executed by the CPU 1100 when the computer 1000 is started up, and programs dependent on the hardware of the computer 1000.
[0186] The HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of the program data 1450.
[0187] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550 (for example, the Internet). For example, the CPU 1100 receives data from other devices or transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0188] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard and a mouse via the input / output interface 1600. Also, the CPU 1100 transmits data to output devices such as a display, a speaker, and a printer via the input / output interface 1600. Further, the input / output interface 1600 may function as a media interface for reading a program or the like recorded on a predetermined recording medium (media). The media is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory or the like.
[0189] For example, when the computer 1000 functions as the information processing apparatus 100 according to the first embodiment, the CPU 1100 of the computer 1000 realizes functions such as the control unit 130 by executing an information processing program loaded on the RAM 1200. Further, the HDD 1400 stores the information processing program according to the present disclosure and the data in the storage unit 120. Note that the CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, these programs may be acquired from other devices via the external network 1550.
[0190] Note that the present technology can also adopt the following configurations. (1) An acquisition unit that acquires a first image including an image of the user's ear, A calculation unit that calculates a head transfer function corresponding to the user based on the first image acquired by the acquisition unit, using a learned model that is learned to output a head transfer function corresponding to the ear when an image including an image of the ear is input. An information processing apparatus comprising: (2) The acquisition unit Obtain an ear parameter, which is a variable indicating the characteristics of the ear included in the first image, The calculation unit By inputting the ear parameter into the learned model, calculate the head transfer function corresponding to the user The information processing apparatus according to (1) above. (3) The acquisition unit Using an ear parameter estimation model learned to output an ear parameter corresponding to an ear when an image including an ear video is input, obtain the ear parameter of the ear included in the first image The information processing apparatus according to (2) above. (4) A first learning unit that generates the ear parameter estimation model by learning the relationship between an image including an ear video and the ear parameter of the ear The information processing apparatus according to (3) above, further comprising (5) The first learning unit Generate the ear parameter estimation model by learning the relationship between an ear image obtained by rendering three-dimensional data of the ear generated based on the ear parameter and the ear parameter The information processing apparatus according to (4) above. (6) The first learning unit Generate the ear parameter estimation model by learning the relationship between a plurality of ear images with the texture of the three-dimensional data of the ear or head, the camera angle in rendering, or the brightness in rendering changed and the ear parameter common to the plurality of ear images The information processing apparatus according to (5) above. (7) A second learning unit that generates the learned model by learning the relationship between an image including an ear video and the head transfer function corresponding to the ear The information processing apparatus according to any one of (1) to (6) above, further comprising (8) The second learning unit Performing acoustic simulation on the 3D data obtained by synthesizing the 3D data of the ear generated based on the ear parameters and the 3D data of the head, and learning the relationship between the head transfer function obtained by the acoustic simulation and the ear parameters to generate the learned model The information processing apparatus according to (7) above (9) The second learning unit Compressing the amount of information of the head transfer function obtained by the acoustic simulation, and learning the relationship between the compressed head transfer function and the ear parameters to generate the learned model The information processing apparatus according to (8) above (10) The second learning unit Setting a listening point of the 3D data of the ear generated based on the ear parameters, and performing the acoustic simulation using the set listening point The information processing apparatus according to (8) or (9) above (11) Further comprising a preprocessing unit that identifies the image of the user's ear from the second image including the image of the entire head of the user, and detects the identified range as the first image The acquisition unit Acquiring the first image detected by the preprocessing unit The information processing apparatus according to any one of (1) to (10) above (12) The preprocessing unit Identifying the range based on the relationship between the feature points of the user's head included in the second image and the posture of the user The information processing apparatus according to (11) above (13) The preprocessing unit When the range cannot be identified based on the relationship between the feature points of the user's head included in the second image and the posture of the user, a new request for acquiring an image different from the second image and including the image of the entire head of the user is made The information processing apparatus according to (12) above (14) The preprocessing unit identifies the video of the user's ear by correcting the rotation of the second image based on the correction information included in the second image, and detects the identified range as the first image The information processing apparatus according to any one of (11) to (13). (15) A computer acquires a first image including a video of the user's ear, calculates a head-related transfer function corresponding to the user based on the acquired first image, using a learned model trained to output a head-related transfer function corresponding to the ear when an image including a video of the ear is input An information processing method. (1..) A computer is caused to function as an acquisition unit that acquires a first image including a video of the user's ear, and a calculation unit that calculates a head-related transfer function corresponding to the user based on the first image acquired by the acquisition unit, using a learned model trained to output a head-related transfer function corresponding to the ear when an image including a video of the ear is input An information processing program for causing the computer to function as such. (17) An information processing system including an information processing apparatus and a user terminal, wherein the user terminal includes a preprocessing unit that identifies the video of the user's ear from a second image including a video of the entire head of the user and detects the identified range as the first image, and a transmission unit that transmits the first image detected by the preprocessing unit to the information processing apparatus, and the processing apparatus includes an acquisition unit that acquires the first image including the video of the user's ear, and a calculation unit that calculates a head-related transfer function corresponding to the user based on the first image acquired by the acquisition unit, using a learned model trained to output a head-related transfer function corresponding to the ear when an image including a video of the ear is input The information processing system including the above.
Description of Symbols
[0191] 1 Information Processing System 10 User Terminal 100 Information Processing Device 110 Communication Unit 120 Storage Unit 130 Control Unit 131 Learning Unit 132 Reception Unit 133 Ear Model Learning Unit 134 Image Generation Unit 135 Ear Parameter Learning Unit 136 HRTF Learning Unit 140 Estimation Unit 141 Acquisition Unit 142 Calculation Unit 143 Provision Unit 144 Preprocessing Unit 145A Posture Detection Unit 145B Ear Detection Unit 161 Acquisition Unit 162 Preprocessing Unit 163A Posture Detection Unit 163B Ear Detection Unit 164 Transmission Unit 165 Reception Unit
Claims
A first learning unit that generates an ear parameter estimation model trained to output ear parameters by learning the relationship between an image including an ear video obtained by rendering three-dimensional ear data generated based on ear parameters, which are variables indicating ear characteristics, and the ear parameters; An acquisition unit that acquires a first image including a video of a user's ear; A calculation unit that calculates a head-related transfer function corresponding to the user from the first image acquired by the acquisition unit, using a learned model trained to output a head-related transfer function corresponding to the ear based on the relationship between the ear parameters output when an image including an ear video is input to the ear parameter estimation model and the head-related transfer function corresponding to the ear parameters; Comprising; The first learning unit Generates the ear parameter estimation model by learning the relationship between an image obtained by rendering a plurality of the three-dimensional data obtained by randomly changing the texture including the skin color of the ear or head or the luminance during rendering and the ear parameters uniquely determined for the plurality of the three-dimensional data; An information processing apparatus.
2. The acquisition unit Acquires ear parameters, which are variables indicating the characteristics of the ear included in the first image, The calculation unit Calculates the head-related transfer function corresponding to the user by inputting the ear parameters into the learned model The information processing apparatus according to claim 1.
3. The acquisition unit Acquires the ear parameters of the ear included in the first image using the ear parameter estimation model The information processing apparatus according to claim 2.
4. The first learning unit Generates the ear parameter estimation model by learning the relationship between a plurality of ear images with the camera angle changed in the rendering and the ear parameters uniquely determined for the plurality of ear images; The information processing apparatus according to claim 1.
5. A second learning unit that generates the learned model by learning the relationship between the ear parameters output when an image including an ear video is input to the ear parameter estimation model and the head-related transfer function corresponding to the ear parameters; The information processing apparatus according to claim 1, further comprising.
6. The second learning unit Performing acoustic simulation on the 3D data obtained by synthesizing the 3D data of the ear generated based on the ear parameters and the 3D data of the head, and learning the relationship between the head transfer function obtained by the acoustic simulation and the ear parameters, thereby generating the learned model. The information processing apparatus according to claim 5.
7. The second learning unit Compressing the amount of information of the head transfer function obtained by the acoustic simulation, and learning the relationship between the compressed head transfer function and the ear parameters, thereby generating the learned model. The information processing apparatus according to claim 6.
8. The second learning unit Setting a listening point of the 3D data of the ear generated based on the ear parameters, and performing the acoustic simulation using the set listening point. The information processing apparatus according to claim 6.
9. Further comprising a preprocessing unit that identifies the image of the user's ear from the second image including the image of the entire head of the user, and detects the identified range as the first image. The acquisition unit Acquiring the first image detected by the preprocessing unit. The information processing apparatus according to claim 1.
10. The preprocessing unit Identifying the range based on the relationship between the feature points of the user's head included in the second image and the posture of the user. The information processing apparatus according to claim 9.
11. The preprocessing unit When the range cannot be identified based on the relationship between the feature points of the user's head included in the second image and the posture of the user, a new request is made to acquire an image different from the second image and including the image of the entire head of the user. The information processing apparatus according to claim 10.
12. The preprocessing unit Identifying the image of the user's ear by correcting the rotation of the second image based on the correction information included in the second image, and detecting the identified range as the first image. The information processing apparatus according to claim 9.
13. A computer Generating an ear parameter estimation model learned to output ear parameters by learning the relationship between an image including an image of the ear obtained by rendering 3D data of the ear generated based on ear parameters, which are variables indicating ear characteristics, and the ear parameters. Acquiring a first image including an image of the user's ear. Based on the relationship between the ear parameters output when an image including an ear video is input to the ear parameter estimation model and the head-related transfer function corresponding to the ear parameters, a learned model that is trained to output the head-related transfer function corresponding to the ear is used to calculate the head-related transfer function corresponding to the user from the acquired first image. An information processing method, Furthermore, a plurality of the 3D data obtained by randomly changing the texture including the color of the ear or the skin of the head or the luminance during rendering is rendered, and the relationship between the obtained image and the ear parameters uniquely determined for the plurality of the 3D data is learned to generate the ear parameter estimation model. An information processing method.
14. A computer, A first learning unit that generates an ear parameter estimation model trained to output ear parameters by learning the relationship between an image including an ear video obtained by rendering 3D data of an ear generated based on ear parameters, which are variables indicating ear characteristics, and the ear parameters. An acquisition unit that acquires a first image including a video of the user's ear. Based on the relationship between the ear parameters output when an image including an ear video is input to the ear parameter estimation model and the head-related transfer function corresponding to the ear parameters, a calculation unit that calculates the head-related transfer function corresponding to the user from the first image acquired by the acquisition unit using a learned model trained to output the head-related transfer function corresponding to the ear. An information processing program for causing the above functions, The first learning unit, Renders a plurality of the 3D data obtained by randomly changing the texture including the color of the ear or the skin of the head or the luminance during rendering, and learns the relationship between the obtained image and the ear parameters uniquely determined for the plurality of the 3D data to generate the ear parameter estimation model. An information processing program.
Citation Information
Patent Citations
Head-related transfer function (HRTF) personalization based on captured images of user
US10038966B1
Estimation of head-related transfer functions for spatial sound representation
US20060067548A1
Customized head-related transfer functions
US9544706B1
Ear shape analysis method, ear shape analysis device, and method for generating ear shape model
WO2017047309A1