Information processing apparatus, information processing method, and program

The information processing device addresses inefficiencies in sound search systems by displaying a visual map of sound data relevance, allowing users to efficiently find desired sound data without exhaustive listening.

JP2026015322APending Publication Date: 2026-01-29KDDI AGILE DEV CENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025142110
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing sound search systems require users to exhaustively listen to multiple sound results, making it inefficient to find the desired sound data associated with a character string input.

Method used

An information processing device that receives a search string, identifies relevant sound data, generates a map of similarity between sounds, and displays a list of sound data with a visual representation of their relevance, reducing the need for users to listen to all results.

Benefits of technology

Facilitates easier identification of desired sound data by providing a visual map of similarity and relevance, minimizing the effort required to find relevant sound data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015322000001_ABST
    Figure 2026015322000001_ABST
Patent Text Reader

Abstract

To reduce labor for finding sound data related to a character string designated by a user.SOLUTION: An information processing device 1 includes a receiver 131 that receives a search string designated by a user, a searcher 135 that specifies a plurality of pieces of audio data in which a degree of association with the search string satisfies a predetermined condition, the degree of association indicating a relationship between a string and a sound, a map generator 136 that generates a map corresponding to a degree of similarity indicating a similarity between sounds of the plurality of pieces of audio data, and an outputter 137 that causes an information terminal to display a second area including a map corresponding to the degree of similarity together with a first area including a list of the plurality of pieces of audio data corresponding to the degree of association.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program for searching for sound data. [Background technology]

[0002] Patent Document 1 describes a system that accepts a character string input by a user and searches a database for sound effects associated with the input character string. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 9-146580 Summary of the Invention [Problem to be solved by the invention]

[0004] In the system described in Patent Document 1, the sound associated with the character string entered by the user is not necessarily the sound the user is looking for. In order to find the sound the user is looking for, the user must exhaustively listen to one or more sounds that are search results corresponding to the entered character string. This poses a problem in that the user must repeatedly search and listen to the sounds, which requires a great deal of effort.

[0005] The present invention has been made in consideration of these points, and has as its object to reduce the effort required for a user to find sound data related to a character string specified by the user. [Means for solving the problem]

[0006] An information processing device of a first aspect of the present invention includes a receiving unit that receives a search string specified by a user, a search unit that identifies multiple sound data whose relevance to the search string satisfies a predetermined condition, where the relevance indicates the relevance between the string and a sound, a map generation unit that generates a map corresponding to the similarity indicating the similarity between the sounds of the multiple sound data, and an output unit that displays on an information terminal a first area including a list of the multiple sound data corresponding to the relevance, and a second area including the map corresponding to the similarity.

[0007] The search unit may identify the plurality of sound data whose relevance is within a predetermined range, or a predetermined number of the plurality of sound data in order of increasing relevance, and the map generation unit may generate the map in which images associated with each of the plurality of sound data identified by the search unit are arranged at a distance corresponding to the similarity.

[0008] The similarity may be a value that indicates the similarity between the sounds indicated by each of the plurality of sound data, without indicating the similarity between the search string and the sounds indicated by each of the plurality of sound data.

[0009] The information processing device may further include an acquisition unit that acquires features generated by performing an encoding process on a combination of a sound and a character string associated with the sound, which is pre-stored in a storage device, and a feature generation unit that generates search features by performing the encoding process on the search character string, and the search unit may identify the degree of association that indicates the similarity between the features acquired by the acquisition unit and the search features generated by the feature generation unit.

[0010] The acquisition unit may acquire the feature generated by performing the encoding process on a combination of a generated sound generated by performing a predetermined sound generation process on a generation character string and the generation character string.

[0011] The receiving unit may receive a selection of one of the plurality of sound data from the user in the first area, and the output unit may cause the display mode in the second area of ​​the sound data selected by the user from the plurality of sound data to be different from the display mode in the second area of ​​sound data not selected by the user from the plurality of sound data.

[0012] The receiving unit may receive a selection of one of the plurality of sound data from the user in the second area, and the output unit may cause the display mode in the first area of ​​the sound data selected by the user from the plurality of sound data to be different from the display mode in the first area of ​​sound data not selected by the user from the plurality of sound data.

[0013] The receiving unit may receive a selection of one of the plurality of sound data from the user in the first area or the second area, and the search unit may identify the plurality of sound data by using a character string associated with the sound data selected by the user as the search character string.

[0014] The receiving unit may receive a selection of one of the plurality of sound data from the user in the first area or the second area, and the information processing device may further have a sound generation unit that generates a generated sound by performing a predetermined sound generation process on a character string associated with the sound data selected by the user, and the output unit may cause the information terminal to output the generated sound.

[0015] The sound generation unit may generate a plurality of the generated sounds by performing the sound generation process a plurality of times using a plurality of parameters that are different from each other on a character string associated with the sound data selected by the user.

[0016] The receiving unit may receive a selection of at least two of the plurality of sound data from the user in the first area or the second area, the information processing device may further have a sound generating unit that generates a generated sound by synthesizing the at least two sound data, and the output unit may cause the information terminal to output the generated sound.

[0017] The information processing device may further include an acquisition unit that acquires, from pre-stored storage devices, a first feature generated by performing a first encoding process on a character string and a second feature generated by performing a second encoding process, different from the first encoding process, on a sound associated with the character string; and a feature generation unit that generates a third feature by performing the first encoding process on the search character string and generates a fourth feature by performing the second encoding process on the search character string, and the search unit may identify the degree of association that indicates a similarity between a combination of the first feature and the second feature acquired by the acquisition unit and a combination of the third feature and the fourth feature generated by the feature generation unit.

[0018] The receiving unit may receive from the user a search sound in addition to the search string, and the information processing device may further include an acquisition unit that acquires, from pre-stored storage devices, a first feature generated by performing a first encoding process on the string and a second feature generated by performing a second encoding process different from the first encoding process on a sound associated with the string, and a feature generation unit that generates a third feature by performing the first encoding process on the search string and generates a fourth feature by performing the second encoding process on the search sound, and the search unit may identify the degree of association that indicates a similarity between a combination of the first feature and the second feature acquired by the acquisition unit and a combination of the third feature and the fourth feature generated by the feature generation unit.

[0019] An information processing method of a second aspect of the present invention includes the steps of: receiving a search string specified by a user, executed by a processor; identifying a plurality of sound data whose relevance indicates the relevance between the string and a sound, and whose relevance to the search string satisfies a predetermined condition; generating a map corresponding to similarities indicating the similarities between the sounds of the plurality of sound data; and displaying on an information terminal a first area including a list of the plurality of sound data corresponding to the relevance, and a second area including the map corresponding to the similarities.

[0020] A third aspect of the program of the present invention causes a processor to execute the following steps: accepting a search string specified by a user; identifying multiple sound data whose relevance indicates the relevance between the string and a sound, and whose relevance to the search string satisfies a predetermined condition; generating a map corresponding to similarities indicating the similarities between the sounds of the multiple sound data; and displaying on an information terminal a first area containing a list of the multiple sound data corresponding to the relevance, and a second area containing the map corresponding to the similarity. [Effects of the Invention]

[0021] The present invention has the effect of reducing the effort required to find sound data related to a character string specified by a user. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 1 is a schematic diagram of an information processing system. [Figure 2] FIG. 1 is a block diagram of an information processing system. [Figure 3] FIG. 3 is a schematic diagram for explaining a feature generation process in the first embodiment. [Figure 4] FIG. 3 is a schematic diagram for explaining a search process for sound data in the first embodiment. [Figure 5] FIG. 10 is a schematic diagram of an information terminal displaying search results. [Figure 6]FIG. 2 is a schematic diagram for explaining a synthesis process of sound data. [Figure 7] FIG. 1 is a flowchart illustrating an exemplary information processing method executed by an information processing system. [Figure 8] FIG. 10 is a schematic diagram for explaining a feature generation process in the second embodiment. [Figure 9] FIG. 2 is a schematic diagram of exemplary feature amounts stored in a storage device. [Figure 10] FIG. 10 is a schematic diagram illustrating a first search process in which a search is performed using a character string specified by a user. [Figure 11] FIG. 10 is a schematic diagram illustrating a first search process in which a search is performed using a character string and a sound specified by a user. DETAILED DESCRIPTION OF THE INVENTION

[0023] First Embodiment [Outline of Information Processing System S] 1 is a schematic diagram of an information processing system S according to this embodiment. The information processing system S includes an information processing device 1, an information terminal 2, and a storage device 3. The information processing system S may also include other devices such as a server and a terminal.

[0024] The information processing device 1 is a computer that manages sound data. The sound data is, for example, data representing sound effects (SE) generated by a computer. The sound data may also represent other sounds such as recorded natural sounds or voices.

[0025] The information processing device 1, for example, executes a process of generating a sound based on a character string input by the user and storing sound data indicating the generated sound in the storage device 3. The information processing device 1 also executes a process of searching the storage device 3 for sound data corresponding to the character string input by the user. The information processing device 1 transmits and receives data between the information terminal 2 and the storage device 3, for example, via wireless communication or wired communication.

[0026] The information terminal 2 is a computer used by a user. The user is a person who handles sound data using the information terminal 2. The information terminal 2 is, for example, a smartphone, a tablet terminal, or a personal computer. The information terminal 2 has an operation unit such as a touch panel or a keyboard for receiving operations, a display unit such as a liquid crystal display for displaying information, and a sound output unit such as a speaker for outputting sound. The information terminal 2 is associated in advance with a user by setting user identification information (Identifier: ID) for identifying the user who uses the information terminal 2.

[0027] The storage device 3 is a device that stores multiple pieces of sound data. The storage device 3 includes storage media such as a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk drive, and an SSD (Solid State Drive). The information processing device 1 may function as the storage device 3 by storing sound data in a storage unit of the information processing device 1.

[0028] An overview of the processing executed by the information processing system S according to this embodiment will be described below. The information processing device 1 receives, from the user at the information terminal 2, a specification of a character string for generation to be used to generate sound data ((1) in FIG. 1). The information processing device 1 performs a predetermined sound generation process on the specified character string for generation, thereby generating a generated sound corresponding to the character string for generation. The information processing device 1 associates sound data representing the generated sound with the character string for generation used to generate the generated sound and stores them in the storage device 3 ((2) in FIG. 1).

[0029] The information processing device 1 receives, from the information terminal 2, a specification of a search string used to search for sound data from the user ((3) in FIG. 1). In the storage device 3, the information processing device 1 identifies a plurality of pieces of sound data whose relevance to the search string satisfies a predetermined condition ((4) in FIG. 1).

[0030] The information processing device 1 causes the information terminal 2 to display search results corresponding to a plurality of sound data ((5) in FIG. 1). The search results include, for example, a first area including a list of a plurality of sound data identified based on the degree of association indicating the association between a character string and a sound, and a second area including a map generated based on the degree of similarity indicating the similarity between sounds.

[0031] In this way, the information processing system S allows the user to check a list of multiple sound data related to the search string in the first area, and sequentially check similar or dissimilar sound data from among the multiple sound data in the second area. This makes it easier for the user to find the desired sound data without having to listen to multiple sound data in a brute-force manner, and reduces the effort required to find sound data related to the search string.

[0032] [Configuration of Information Processing System S] FIG. 2 is a block diagram of an information processing system S according to this embodiment. In FIG. 2, arrows indicate main data flows, and data flows other than those shown in FIG. 2 may also exist. In FIG. 2, each block indicates a functional configuration rather than a hardware (device) configuration. Therefore, the blocks shown in FIG. 2 may be implemented in a single device, or may be implemented separately in multiple devices. Data may be exchanged between blocks via any means, such as a data bus, a network, or a portable storage medium.

[0033] The information processing device 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The information processing device 1 may be configured by connecting two or more physically separate devices via wired or wireless connections. The information processing device 1 may also be configured by a cloud, which is a collection of computer resources.

[0034] The communication unit 11 has a communication controller for transmitting and receiving data to and from the information terminal 2 via the network. The communication unit 11 notifies the control unit 13 of data received from the information terminal 2 via the network. The communication unit 11 also transmits data output from the control unit 13 to the information terminal 2 via the network.

[0035] The storage unit 12 is a storage medium including a ROM, a RAM, a hard disk drive, an SSD, etc. The storage unit 12 stores in advance a program to be executed by the control unit 13. When the information processing device 1 functions as the storage device 3, the storage unit 12 stores a plurality of pieces of sound data. The storage unit 12 may be provided outside the information processing device 1, in which case data may be exchanged between the storage unit 12 and the control unit 13 via a network.

[0036] The control unit 13 has a reception unit 131, a sound generation unit 132, a feature generation unit 133, an acquisition unit 134, a search unit 135, a map generation unit 136, and an output unit 137. The control unit 13 is a processor such as a CPU (Central Processing Unit), and functions as the reception unit 131, the sound generation unit 132, the feature generation unit 133, the acquisition unit 134, the search unit 135, the map generation unit 136, and the output unit 137 by executing a program stored in the storage unit 12.

[0037] [Feature generation process] The following describes in detail the processing executed by the information processing system S. First, a description will be given of the generation processing in this embodiment in which the information processing device 1 generates features of sound data stored in the storage device 3. Fig. 3 is a schematic diagram for explaining the feature generation processing in this embodiment.

[0038] The reception unit 131 receives the specification of a generation string from the user in the information terminal 2. The generation string is a string used by the sound generation unit 132 to generate a sound, and is, for example, a string that specifies the content of the sound the user wants to generate (such as "sound of the wind" or "sound of the river"). The user inputs the generation string using the operation unit of the information terminal 2. The information terminal 2 transmits the input generation string to the information processing device 1. In the information processing device 1, the reception unit 131 receives the generation string transmitted by the information terminal 2.

[0039] The sound generation unit 132 generates a generated sound by performing a predetermined sound generation process on the generation character string received by the reception unit 131. The sound generation unit 132 inputs the generation character string to a machine learning model that outputs a sound corresponding to the input character string, and that has been generated by executing a known machine learning process using a combination of a character string and a sound as training data. The sound generation unit 132 generates the sound output by the machine learning model as a generated sound corresponding to the generation character string.

[0040] The feature generation unit 133 performs a predetermined encoding process on the combination of the generated sound and the character string for generation, thereby generating features that indicate the characteristics of the generated sound and the character string for generation. The features are, for example, vectors in a multidimensional space (e.g., 512 dimensions).

[0041] The feature generation unit 133 acquires, for example, a machine learning model that receives input of a sound or a character string and outputs a feature of the sound or the character string, which is stored in advance in the storage unit 12. The machine learning model that outputs the feature is generated in advance by executing a known machine learning process using a combination of a sound and a character string associated with the sound as training data, so that the feature of the sound and the feature of the character string are similar (for example, the distance in a multidimensional space is small). The machine learning model that outputs the feature is, for example, CLAP (Contrastive Language-Audio Pretraining).

[0042] The machine learning model that outputs features includes a sound encoder (encoder) that outputs features of input sounds, and a string encoder that outputs features of input strings. The features output by the sound encoder and the features output by the string encoder are each a vector represented in a common multidimensional space. The dimensions of the feature vectors are defined in advance in the information processing system S or specified by the user.

[0043] In a situation where a sound and a character string are associated, the distance between the feature output by inputting the sound into a sound encoder and the feature output by inputting the character string into a character string encoder is relatively small. On the other hand, in a situation where a sound and a character string are not associated, the distance between the feature output by inputting the sound into a sound encoder and the feature output by inputting the character string into a character string encoder is relatively large. Therefore, in the feature space, the distance between the sound feature and the character string feature indicates the association between the character string and the sound.

[0044] As an encoding process, the feature generation unit 133 generates, as a feature of the generated sound, a feature output by inputting the generated sound into a sound encoder included in the machine learning model, and generates, as a feature of the generated string, a feature output by inputting the generation string into a string encoder included in the machine learning model. The feature generation unit 133 associates sound data indicating the generated sound, the generation string used to generate the generated sound (i.e., a string indicating the content of the sound), the feature of the generated sound, and the feature of the generation string, and stores them in the storage device 3.

[0045] Alternatively, the feature generation unit 133 may perform the above-described encoding process on recorded sound such as natural sound or voice instead of generated sound, thereby generating features indicating the characteristics of the recorded sound. In this case, the feature generation unit 133 receives a character string (such as "sound of the wind" or "sound of the river") indicating the content of the recorded sound from the user, and performs the above-described encoding process on the specified character string to generate features indicating the characteristics of the character string. The feature generation unit 133 associates sound data indicating the recorded sound, character strings indicating the content of the recorded sound, feature values ​​of the recorded sound, and feature values ​​of the character strings indicating the content of the recorded sound, and stores them in the storage device 3.

[0046] [Sound data search processing] Next, a description will be given of a search process in this embodiment in which the information processing device 1 searches for sound data stored in the storage device 3. Fig. 4 is a schematic diagram for explaining the search process for sound data in this embodiment.

[0047] The accepting unit 131 accepts a search string specification from a user at the information terminal 2. The user searching for sound data may be the same as or different from the user who generated the sound data. The information terminal 2 used to specify the search string may be the same device as or a different device from the information terminal 2 used to specify the generation string.

[0048] The search string is a string used by the search unit 135 to search for sound data, and is, for example, a string that specifies the content of the sound that the user wants to search for (such as "sound of the wind" or "sound of the river"). The user inputs the search string using the operation unit of the information terminal 2. The information terminal 2 transmits the input search string to the information processing device 1. In the information processing device 1, the reception unit 131 receives the search string transmitted by the information terminal 2.

[0049] The acquisition unit 134 acquires, from the storage device 3, the feature amounts of each of the multiple sounds generated by the feature generation unit 133. The feature amounts of the sounds acquired by the acquisition unit 134 are feature amounts generated by performing the above-mentioned encoding process on generated sounds generated by the above-mentioned sound generation process. Alternatively, the feature amounts of the sounds acquired by the acquisition unit 134 may be feature amounts generated by performing the above-mentioned encoding process on recorded sounds such as natural sounds or speech.

[0050] The feature generation unit 133 generates search features indicating features of the search string by performing the above-mentioned encoding process on the search string received by the receiving unit 131. That is, the feature generation unit 133 generates the features of the search string by the encoding process (machine learning model) used to generate the features of the sound acquired by the acquisition unit 134.

[0051] The feature generation unit 133 acquires, for example, a machine learning model that outputs features and that has been used to generate features of a sound, which is stored in advance in the storage unit 12. As described above, the machine learning model that outputs features is generated in advance so that the features of a sound and the features of a character string associated with the sound are similar (for example, so that the distance in a multidimensional space is small).

[0052] As an encoding process, the feature generation unit 133 generates, as a search feature for the search string, a feature output by inputting a search string into the string encoder, one of the string encoder and sound encoder included in the machine learning model.

[0053] The search unit 135 uses the search features generated by the feature generation unit 133 to identify the degree of association indicating the association between the search string and each of the plurality of sound data stored in the storage unit 12. The search unit 135 calculates, for example, a distance (for example, Euclidean distance in a multidimensional space) that is an index indicating the similarity between the sound features indicated by each of the plurality of sound data acquired by the acquisition unit 134 and the search features of the search string generated by the feature generation unit 133, and identifies the calculated distance as the degree of association.

[0054] The search unit 135 identifies, as search results, a plurality of pieces of sound data whose relevance to the search string satisfies a predetermined condition from among the plurality of pieces of sound data stored in the storage unit 12. The search unit 135 identifies, for example, a plurality of pieces of sound data whose relevance is within a predetermined range (for example, the distance is equal to or less than a predetermined reference value), or a predetermined number of pieces of sound data in order of increasing relevance (for example, in order of decreasing distance), as search results.

[0055] The information processing device 1 outputs the search results to the information terminal 2. Fig. 5 is a schematic diagram of the information terminal 2 displaying the search results. Fig. 5 shows an example search result when the user specifies "sound of water" as the search string.

[0056] The map generating unit 136 generates a map corresponding to the similarity indicating the similarity between sounds of the plurality of sound data that are the search results. The map generating unit 136 reduces the dimension of the feature amount of each of the plurality of sound data to two dimensions, for example, by performing principal component analysis on the feature amount (vector in a multidimensional space) of each of the plurality of sound data.

[0057] The map generation unit 136 generates a two-dimensional map by arranging images (icons, symbols, etc.) associated with the sound data using the two-dimensional feature amounts of each of the multiple sound data as coordinates. In the map generated by the map generation unit 136, for each combination of two sound data from the multiple sound data, the closer the feature amounts of the sounds are, the smaller the distance becomes, and the farther the feature amounts of the sounds are, the larger the distance becomes. Therefore, the map generated by the map generation unit 136 represents the similarity (distance) that indicates the similarity between the sounds of the multiple sound data.

[0058] The map generating unit 136 may generate a three-dimensional map by reducing the dimensions of the feature amounts of each of the plurality of sound data to three dimensions.

[0059] The output unit 137 causes the information terminal 2 to display a first area A1 including a list of multiple sound data corresponding to the relevance determined by the search unit 135, as well as a second area A2 including a map corresponding to the similarity generated by the map generation unit 136.

[0060] The first area A1 displays character strings (character strings indicating the content of the sound) associated with multiple sound data identified as search results by the search unit 135 in order of their relevance to the search character string (for example, in order of shortest distance).

[0061] The second area A2 displays images (circular images in the example of FIG. 5) associated with the plurality of sound data identified as search results by the search unit 135, arranged according to the similarity (distance) indicating the similarity between the sounds. That is, for the plurality of sound data displayed in the first area A1, the second area A2 displays images associated with each sound data so that the closer the sound feature values ​​are, the smaller the distance, and the farther the sound feature values ​​are, the larger the distance. A character string (character string indicating the content of the sound) associated with the sound data is displayed near the image of each of the plurality of sound data.

[0062] The output unit 137 may change the scale of the map included in the second area A2 in response to a predetermined operation (for example, pinching in or pinching out) performed by the user on the map. For example, the output unit 137 reduces the map (widens the coordinate range of the map) when the user performs a pinch-in operation, and enlarges the map (narrows the coordinate range of the map) when the user performs a pinch-out operation.

[0063] The receiving unit 131 receives a selection of any sound data from the user in the first area A1 or the second area A2. The output unit 137 acquires the selected sound data from the storage device 3 and causes the sound output unit of the information terminal 2 to output (play) the sound indicated by the acquired sound data.

[0064] This allows the user to view a list of multiple sound data related to the search string in the first area A1, and to sequentially view similar or dissimilar sound data from among the multiple sound data in the second area A2, allowing the user to search for the desired sound data without having to listen to all of the multiple sound data in a brute force manner.

[0065] The output unit 137 may display, on the map included in the second area A2, an image associated with one or more pieces of sound data that have been previously evaluated by one or more users (the user who performed the search or other users) out of the plurality of pieces of sound data identified as search results by the search unit 135. In this case, the receiving unit 131 receives an evaluation (for example, "good" or "bad") of the sound data from the user, and stores evaluation information that associates the sound data with the evaluation in the storage device 3.

[0066] The output unit 137 tally up the number of predetermined ratings (e.g., "good") for each of the multiple sound data based on the rating information stored in the storage device 3. The output unit 137 displays on the map an image associated with sound data for which the number of ratings is equal to or greater than a reference value, and does not display on the map an image associated with sound data for which the number of ratings is less than the reference value. The reference value is defined in advance in the information processing system S or specified by the user. This enables the information processing system S to narrow down the number of multiple sound data to be displayed on the map according to the user's ratings, making it easier for the user to decide which sound data to preview.

[0067] The output unit 137 may cause the display mode (color, pattern, size, etc.) of sound data selected by the user in the first area A1 from the plurality of sound data in the second area A2 to differ from the display mode of sound data not selected by the user from the plurality of sound data in the second area A2. Furthermore, the output unit 137 may cause the display mode of sound data selected by the user in the second area A2 from the display mode of sound data not selected by the user from the plurality of sound data in the first area A1 to differ from the display mode of sound data not selected by the user from the plurality of sound data in the first area A1. This enables the information processing system S to easily allow the user to find sound data selected in one of the first area A1 and the second area A2 in the other of the first area A1 and the second area A2.

[0068] The output unit 137 may change the display mode in the second area A2 of one or more pieces of sound data whose features are close to those of the sound data selected by the user in the first area A1 (for example, the distance between the features is equal to or less than a reference value) from the display mode in the second area A2 of sound data other than the one or more pieces of sound data among the plurality of sound data. This allows the information processing system S to make it easier for the user to recognize one or more pieces of sound data similar to the sound data selected from the list of the plurality of sound data included in the first area A1 in the map included in the second area A2.

[0069] The search unit 135 may perform the above-described search process using, as a search string, a character string (a character string indicating the content of the sound) associated with the sound data selected by the user in the first area A1 or the second area A2, thereby identifying multiple pieces of sound data as new search results. The output unit 137 causes the information terminal 2 to display the first area A1 and the second area A2 using the multiple pieces of sound data that are the new search results. This allows the information processing system S to reduce the effort required for the user to newly search for sound data related to the sound data selected in the first area A1 or the second area A2.

[0070] The sound generation unit 132 may generate a generated sound by performing the above-described sound generation process using, as a generation string, a character string (a character string indicating the content of the sound) associated with sound data selected by the user in the first area A1 or the second area A2. Here, the sound generation unit 132 may generate a plurality of generated sounds by performing the sound generation process multiple times on the generation string using a plurality of mutually different parameters. The output unit 137 outputs (plays) one or more generated sounds generated by the sound generation unit 132 to the sound output unit of the information terminal 2. In this way, even if the sound data desired by the user is not found among the plurality of sound data in the search results, the information processing system S can provide sound data that is closer to the user's request by generating a new generated sound using a character string associated with the sound data selected by the user.

[0071] The sound generation unit 132 may perform synthesis processing to generate a new generated sound by synthesizing at least two pieces of sound data selected by the user. Fig. 6 is a schematic diagram for explaining the synthesis processing of sound data.

[0072] The receiving unit 131 receives a selection of at least two pieces of sound data from the user in the first area A1 or the second area A2. The sound generating unit 132 generates a generated sound by synthesizing the at least two pieces of selected sound data. The sound generating unit 132 may generate a plurality of generated sounds by synthesizing at least two pieces of sound data at a plurality of different synthesis ratios. The sound generating unit 132 may also generate a generated sound by synthesizing at least two pieces of sound data at a synthesis ratio specified by the user on the information terminal 2, for example.

[0073] The output unit 137 outputs (plays) one or more sounds generated by the sound generation unit 132 to the sound output unit of the information terminal 2. As a result, even if the sound data desired by the user is not found among the multiple sound data in the search results, the information processing system S can provide sound data that is closer to the user's request by synthesizing at least two sound data selected by the user.

[0074] [Information processing method flow] Fig. 7 is a diagram showing a flowchart of an exemplary information processing method executed by the information processing system S according to this embodiment. In the flowchart shown in Fig. 7, the search process is performed following the generation process, but the generation process and the search process may be performed independently.

[0075] The reception unit 131 receives a specification of a character string for generation from a user at the information terminal 2 (S11). The sound generation unit 132 generates a generated sound by performing a predetermined sound generation process on the character string for generation received by the reception unit 131 (S12).

[0076] The feature generation unit 133 generates a feature output by inputting a generated sound to a sound encoder included in the machine learning model as a feature of the generated sound, and generates a feature output by inputting a generation string to a string encoder included in the machine learning model as a feature of the generation string (S13). The feature generation unit 133 associates sound data indicating the generated sound, the generation string used to generate the generated sound (i.e., a string indicating the content of the sound), the feature of the generated sound, and the feature of the generation string, and stores them in the storage device 3 (S14).

[0077] The reception unit 131 receives a search string specification from the user at the information terminal 2 (S15). The acquisition unit 134 acquires, from the storage device 3, the features of each of the multiple sounds generated by the feature generation unit 133 (S16). The feature generation unit 133 performs a predetermined encoding process on the search string received by the reception unit 131, thereby generating search features that indicate the features of the search string (S17).

[0078] The search unit 135 uses the search features generated by the feature generation unit 133 to identify the degree of association indicating the association between the search string and each of the plurality of sound data stored in the storage unit 12 (S18). The search unit 135 identifies, as search results, a plurality of sound data from the plurality of sound data stored in the storage unit 12 whose degree of association with the search string satisfies a predetermined condition (S19).

[0079] The map generation unit 136 generates a map corresponding to the similarity indicating the similarity between sounds of the plurality of sound data that are the search results (S20). The output unit 137 causes the information terminal 2 to display a first area A1 that includes a list of the plurality of sound data that correspond to the relevance identified by the search unit 135, as well as a second area A2 that includes the map corresponding to the similarity generated by the map generation unit 136 (S21).

[0080] [Effects of the embodiment] According to the information processing system S of this embodiment, the information processing device 1 allows the user to check a list of multiple sound data related to a search string in the first area A1, and allows the user to sequentially check similar or dissimilar sound data among the multiple sound data in the second area A2. This makes it easier for the user to find desired sound data without having to listen to multiple sound data in a brute-force manner, and reduces the effort required to find sound data related to a search string.

[0081] Second Embodiment In this embodiment, two different types of encoding processes are used to generate features and search for sound data. The following mainly describes the differences from the first embodiment.

[0082] [Feature generation process] The following describes a generation process in which the information processing device 1 generates features of sound data stored in the storage device 3. Fig. 8 is a schematic diagram for explaining the feature generation process in this embodiment. The reception unit 131 receives a specification of a character string for generation from a user at the information terminal 2. The sound generation unit 132 generates a generated sound by performing a predetermined sound generation process on the character string for generation received by the reception unit 131.

[0083] The feature generation unit 133 performs two different types of encoding processing on the combination of the generated sound and the character string for generation, thereby generating a feature that indicates the characteristics of the generated sound and a feature that indicates the characteristics of the character string for generation.

[0084] The feature generation unit 133 generates a first feature indicating a feature of the character string to be generated by performing a first encoding process on the character string to be generated, and generates a second feature indicating a feature of the generated sound by performing a second encoding process on the generated sound that is different from the first encoding process.

[0085] As a first encoding process, the feature generation unit 133 inputs a string to be generated into a first string encoder that outputs features of the input string, and generates the output features as first features indicating the features of the string to be generated. The first string encoder is an encoder that outputs string features from the string alone, without involving machine learning processing for combinations of sounds and strings, as in the machine learning model in the first embodiment. The first string encoder is, for example, an encoder that outputs a vector indicating the score of the input string in the BM25 algorithm. The first features output by the first string encoder are, for example, a sparse vector indicating the frequency of occurrence of each of multiple words in the string.

[0086] As the second encoding process, the feature generation unit 133 acquires a machine learning model similar to that of the first embodiment from the storage unit 12. The machine learning model that outputs the feature is generated in advance by executing a known machine learning process using a combination of a sound and a character string associated with the sound as training data, so that the feature of the sound and the feature of the character string are similar (for example, the distance in multidimensional space is small). The machine learning model that outputs the feature includes a sound encoder that outputs the feature of an input sound, and a second character string encoder that outputs the feature of an input character string.

[0087] The feature generator 133 generates a feature output by inputting the generated sound to a sound encoder included in the machine learning model as a second feature indicating the feature of the generated sound. The second feature output by the sound encoder is, for example, a dense vector indicating the feature of the generated sound in a multidimensional space (e.g., 512 dimensions).

[0088] The feature generation unit 133 associates sound data indicating the generated sound, a generating string used to generate the generated sound (i.e., a string indicating the content of the sound), a first feature indicating the characteristics of the generating string, and a second feature indicating the characteristics of the generated sound, and stores them in the storage device 3.

[0089] Alternatively, the feature generation unit 133 may perform the second encoding process described above on recorded sound, such as natural sound or voice, instead of generated sound, to generate second features indicating characteristics of the recorded sound. In this case, the feature generation unit 133 receives a character string (such as "sound of the wind" or "sound of the river") indicating the content of the recorded sound from the user, and performs the first encoding process described above on the specified character string to generate first features indicating characteristics of the character string. The feature generation unit 133 associates sound data indicating the recorded sound, the character string indicating the content of the recorded sound, the first features indicating characteristics of the character string indicating the content of the recorded sound, and the second features indicating characteristics of the recorded sound, and stores them in the storage device 3.

[0090] 9 is a schematic diagram of exemplary features stored in the storage device 3. The storage device 3 stores feature information that associates identification information of sound data ("No." in FIG. 9), a first feature amount ("sparse vector" in FIG. 9), a second feature amount ("dense vector" in FIG. 9), and a character string indicating the content of the sound ("character string" in FIG. 9). The storage device 3 also stores sound data in association with identification information included in each of the multiple pieces of feature information.

[0091] [Sound data search processing] Next, a description will be given of a search process in which the information processing device 1 in this embodiment searches for sound data stored in the storage device 3. In this embodiment, the information processing device 1 switches between a first search process in which a search is performed using a character string specified by the user, and a second search process in which a search is performed using a combination of a character string and a sound specified by the user.

[0092] 10 is a schematic diagram for explaining a first search process in which a search is performed using a character string designated by a user. In the information terminal 2, the accepting unit 131 accepts a search character string designated by the user.

[0093] The acquisition unit 134 acquires, from the storage device 3, a first feature generated by performing a first encoding process on a character string, and a second feature generated by performing a second encoding process, which is different from the first encoding process, on a sound associated with the character string.

[0094] The feature generation unit 133 generates a third feature indicating a feature of the search string by performing the above-described first encoding process on the search string received by the receiving unit 131. That is, the feature generation unit 133 generates a third feature indicating a feature of the search string by the first encoding process (first string encoder) used to generate the first feature.

[0095] The feature generation unit 133 generates a fourth feature indicating a feature of the search string by performing the above-described second encoding process on the search string received by the receiving unit 131. That is, the feature generation unit 133 generates a fourth feature indicating a feature of the search string by the second encoding process (a second character string encoder corresponding to the sound encoder) used to generate the second feature.

[0096] The search unit 135 identifies a degree of association that indicates the similarity between the combination of the first feature amount and the second feature amount acquired by the acquisition unit 134 and the combination of the third feature amount and the fourth feature amount generated by the feature amount generation unit 133. The search unit 135 calculates, for example, a distance (for example, a Euclidean distance in a multidimensional space) that is an index indicating the similarity between a vector combining the first feature amount and the second feature amount and a vector combining the third feature amount and the fourth feature amount, and identifies the calculated distance as the degree of association.

[0097] The search unit 135 identifies, as search results, a plurality of pieces of sound data whose identified relevance degree satisfies a predetermined condition from among the plurality of pieces of sound data stored in the storage unit 12. The search unit 135 identifies, for example, a plurality of pieces of sound data whose relevance degree is within a predetermined range (for example, the distance is equal to or less than a predetermined reference value), or a predetermined number of pieces of sound data in order of increasing relevance degree (for example, in order of decreasing distance), as search results.

[0098] This allows the search unit 135 to perform a search using a combination of the features of the string itself and the features indicating the association between the string and the sound, thereby making it possible to identify sound data that may be related to the search string specified by the user.

[0099] 11 is a schematic diagram illustrating a first search process in which a search is performed using a character string and sound specified by the user. The reception unit 131 receives a search sound from the user in addition to a search character string in the information terminal 2. The search sound is, for example, a sound indicated by sound data selected by the user on the screen illustrated in FIG.

[0100] The acquisition unit 134 acquires, from the storage device 3, a first feature generated by performing a first encoding process on a character string, and a second feature generated by performing a second encoding process, which is different from the first encoding process, on a sound associated with the character string.

[0101] The feature generation unit 133 generates a third feature indicating a feature of the search string by performing the above-described first encoding process on the search string received by the receiving unit 131. That is, the feature generation unit 133 generates a third feature indicating a feature of the search string by the first encoding process (first string encoder) used to generate the first feature.

[0102] The feature generation unit 133 generates a fourth feature that indicates the feature of the search sound by performing the above-mentioned second encoding process on the search sound received by the receiving unit 131. That is, the feature generation unit 133 generates a fourth feature that indicates the feature of the search sound by the second encoding process (sound encoder) that was used to generate the second feature.

[0103] The search unit 135 identifies a degree of association that indicates the similarity between the combination of the first feature amount and the second feature amount acquired by the acquisition unit 134 and the combination of the third feature amount and the fourth feature amount generated by the feature amount generation unit 133. The search unit 135 calculates, for example, a distance (for example, a Euclidean distance in a multidimensional space) that is an index indicating the similarity between a vector combining the first feature amount and the second feature amount and a vector combining the third feature amount and the fourth feature amount, and identifies the calculated distance as the degree of association.

[0104] The search unit 135 identifies, as search results, a plurality of pieces of sound data whose identified relevance degree satisfies a predetermined condition from among the plurality of pieces of sound data stored in the storage unit 12. The search unit 135 identifies, for example, a plurality of pieces of sound data whose relevance degree is within a predetermined range (for example, the distance is equal to or less than a predetermined reference value), or a predetermined number of pieces of sound data in order of increasing relevance degree (for example, in order of decreasing distance), as search results.

[0105] This allows the search unit 135 to perform a search using a combination of character string features and sound features, thereby identifying sound data that may be related to the search character string and search sound specified by the user.

[0106] As in the first embodiment, the output unit 137 causes the information terminal 2 to display a first area A1 including a list of multiple sound data corresponding to the relevance determined by the search unit 135, as well as a second area A2 including a map corresponding to the similarity generated by the map generation unit 136.

[0107] [Effects of the embodiment] According to the information processing system S of this embodiment, the information processing device 1 searches for sound data using sound features generated using two different types of encoding processes and character string features associated with the sound. This allows the information processing system S to present to the user search results that reflect the combination of sound features and character string features.

[0108] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), which is "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and resilience."

[0109] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]

[0110] S Information Processing System 1. Information processing equipment 11 Communications Department 12 Storage section 13 Control Unit 131 Reception 132 Sound generation section 133 Feature Generation Unit 134 Acquisition Department 135 Search Department 136 Map Generation Unit 137 Output section 2. Information terminal 3 Storage device

Claims

1. a reception unit that receives a search string specified by a user; a search unit that identifies a plurality of pieces of sound data whose relevance to the search string satisfies a predetermined condition, the relevance indicating the relevance between the character string and the sound; a map generating unit that generates a map corresponding to a similarity indicating a similarity between sounds of the plurality of sound data; an output unit that displays, on the information terminal, a first area including a list of the plurality of sound data corresponding to the relevance level and a second area including the map corresponding to the similarity level; An information processing device having the above.

2. the search unit identifies the plurality of sound data whose relevance is within a predetermined range, or a predetermined number of the plurality of sound data whose relevance is high; the map generation unit generates the map in which images associated with each of the plurality of sound data identified by the search unit are arranged at a distance corresponding to the degree of similarity. The information processing device according to claim 1 .

3. the similarity is a value indicating the similarity between sounds indicated by each of the plurality of sound data, without indicating the similarity between the search string and the sound indicated by each of the plurality of sound data; 3. The information processing device according to claim 1 or 2.

4. an acquisition unit that acquires features generated by performing an encoding process on a combination of a sound and a character string associated with the sound, the combination being stored in advance in a storage device; a feature generating unit that generates search features by performing the encoding process on the search string; and the search unit identifies the degree of association indicating a similarity between the feature acquired by the acquisition unit and the feature for search generated by the feature generation unit.

3. The information processing device according to claim 1 or 2.

5. the acquisition unit acquires the feature quantity generated by performing the encoding process on a combination of a generated sound generated by performing a predetermined sound generation process on a generation character string and the generation character string. The information processing device according to claim 4 .

6. the receiving unit receives, in the first area, a selection of any one of the plurality of sound data from the user; the output unit causes a display mode in the second area of ​​sound data selected by the user from among the plurality of sound data to differ from a display mode in the second area of ​​sound data not selected by the user from among the plurality of sound data.

3. The information processing device according to claim 1 or 2.

7. the receiving unit receives, in the second area, a selection of any one of the plurality of sound data from the user; the output unit causes a display mode in the first area of ​​sound data selected by the user from among the plurality of sound data to differ from a display mode in the first area of ​​sound data not selected by the user from among the plurality of sound data.

3. The information processing device according to claim 1 or 2.

8. the receiving unit receives a selection of any one of the plurality of sound data from the user in the first area or the second area; the search unit identifies the plurality of sound data by using a character string associated with the sound data selected by the user as the search character string.

3. The information processing device according to claim 1 or 2.

9. the receiving unit receives a selection of any one of the plurality of sound data from the user in the first area or the second area; a sound generation unit that generates a generated sound by performing a predetermined sound generation process on a character string associated with the sound data selected by the user; the output unit causes the information terminal to output the generated sound.

3. The information processing device according to claim 1 or 2.

10. the sound generation unit generates a plurality of the generated sounds by performing the sound generation process a plurality of times using a plurality of parameters different from each other on the character string associated with the sound data selected by the user. The information processing device according to claim 9 .

11. the receiving unit receives a selection of at least two pieces of sound data from the user in the first area or the second area, and a sound generator that generates a generated sound by synthesizing the at least two pieces of sound data; the output unit causes the information terminal to output the generated sound.

3. The information processing device according to claim 1 or 2.

12. an acquisition unit that acquires a first feature quantity generated by performing a first encoding process on a character string, which are stored in advance in a storage device, and a second feature quantity generated by performing a second encoding process different from the first encoding process on a sound associated with the character string; a feature generating unit that generates a third feature by performing the first encoding process on the search string and generates a fourth feature by performing the second encoding process on the search string; and the search unit identifies the degree of association indicating a similarity between the combination of the first feature amount and the second feature amount acquired by the acquisition unit and the combination of the third feature amount and the fourth feature amount generated by the feature amount generation unit.

3. The information processing device according to claim 1 or 2.

13. the receiving unit receives a designation of a search sound in addition to the search character string from the user; an acquisition unit that acquires a first feature quantity generated by performing a first encoding process on a character string, which are stored in advance in a storage device, and a second feature quantity generated by performing a second encoding process different from the first encoding process on a sound associated with the character string; a feature generating unit that generates a third feature by performing the first encoding process on the search character string and generates a fourth feature by performing the second encoding process on the search sound; and the search unit identifies the degree of association indicating a similarity between the combination of the first feature amount and the second feature amount acquired by the acquisition unit and the combination of the third feature amount and the fourth feature amount generated by the feature amount generation unit.

3. The information processing device according to claim 1 or 2.

14. The processor executes accepting a user-specified search string; a step of identifying a plurality of pieces of sound data whose relevance indicating the relevance between a character string and a sound and whose relevance to the search character string satisfies a predetermined condition; generating a map corresponding to a similarity indicating a similarity between sounds of the plurality of sound data; displaying, on the information terminal, a first area including a list of the plurality of sound data corresponding to the relevance level and a second area including the map corresponding to the similarity level; An information processing method comprising:

15. The processor accepting a user-specified search string; a step of identifying a plurality of pieces of sound data whose relevance indicating the relevance between a character string and a sound and whose relevance to the search character string satisfies a predetermined condition; generating a map corresponding to a similarity indicating a similarity between sounds of the plurality of sound data; displaying, on the information terminal, a first area including a list of the plurality of sound data corresponding to the relevance level and a second area including the map corresponding to the similarity level; A program that executes.

Citation Information

Patent Citations

  • Effect sound retrieving device

    JP1997146580A