Sound recognition system, and sound recognition method

The sound recognition system enhances accuracy by using dual sound collection units with insulation and phoneme recognition, addressing environmental condition impacts on sound propagation through different media.

JP2025107053APending Publication Date: 2025-07-17RYOYO ELECTRO CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024000773
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Conventional sound recognition systems face challenges in maintaining recognition accuracy due to varying external environmental conditions, particularly when sounds propagate through different media.

Method used

A sound recognition system comprising first and second sound collection units that collect sounds through different media, with a sound insulation unit between them, and an evaluation unit that derives recognition results using phoneme recognition, considering the time difference and position of sound collection.

Benefits of technology

Improves recognition accuracy by considering external environmental conditions and suppressing noise from non-target sources, enabling precise identification of sound generation regions and sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025107053000001_ABST
    Figure 2025107053000001_ABST
Patent Text Reader

Abstract

To provide a sound recognition system and a sound recognition method that can improve recognition accuracy.SOLUTION: A sound recognition system 100 evaluates a physical sound, and comprises: a first sound pickup unit 1a and a second sound pickup unit 1b that pick up the physical sounds propagating through media in states different from each other; an acquisition unit that acquires evaluation information based on the first physical sound picked up by the first sound pickup unit 1a and the second physical sound picked up by the second sound pickup unit 1b; and an evaluation unit that derives a recognition result corresponding to the evaluation information by using phoneme recognition. For example, a sound insulation unit 15 is provided between the first sound pickup unit 1a and the second sound pickup unit 1b, and has the first sound pickup unit 1a and the second sound pickup unit 1b arranged thereon.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound recognition system and a sound recognition method.

Background Art

[0002] Conventionally, as a technique for evaluating physical sounds, for example, a sound recognition system disclosed in Patent Document 1 has been proposed.

[0003] In Patent Document 1, there is disclosed a sound recognition system using phoneme recognition, the sound recognition device including: an acquisition unit that acquires physical sound information generated based on a physical sound propagating through a medium; a storage unit that stores a plurality of learning models constructed using reference physical sound information acquired in advance and reference phonemes associated therewith, as well as a plurality of databases associated with the plurality of learning models, the databases being constructed using recognition phonemes acquired in advance and recognition information associated therewith; and a derivation unit that derives a recognition result obtained by recognizing the content of the physical sound information.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Here, there is a concern that the recognition accuracy of a physical sound propagating through a specific medium may decrease due to the conditions of the external environment. In this regard, the disclosed technique of Patent Document 1 does not describe or suggest such a concern. Therefore, a method for improving the recognition accuracy is required.

[0006] Therefore, the present invention has been devised in view of the above - described problems, and an object thereof is to provide a sound recognition system and a sound recognition method capable of improving the recognition accuracy.

Means for Solving the Problems

[0007] The sound recognition system according to the first invention is a sound recognition system for evaluating physical sounds, comprising: a first sound collection unit that collects the physical sounds propagating through media in different states; a second sound collection unit; an acquisition unit that acquires evaluation information based on a first physical sound collected by the first sound collection unit and a second physical sound collected by the second sound collection unit; and an evaluation unit that derives a recognition result corresponding to the evaluation information using phoneme recognition.

[0008] The sound recognition system according to the second invention is, in the first invention, further characterized by comprising a sound insulation unit provided between the first sound collection unit and the second sound collection unit, where the first sound collection unit and the second sound collection unit are arranged.

[0009] The sound recognition system according to the third invention is, in the second invention, characterized in that the first medium through which the first physical sound propagates represents a gas, the second medium through which the second physical sound propagates represents a liquid, and the sound insulation unit floats in the second medium.

[0010] The sound recognition system according to the fourth invention is, in the third invention, further characterized by comprising a position identification unit that acquires position information indicating the floating position of the sound insulation unit, and the evaluation unit specifies the floating position at the time when the evaluation information is acquired based on the position information.

[0011] The sound recognition system according to the fifth invention is, in the first invention, characterized in that the evaluation unit specifies the generation region of the physical sound based on the difference between the time when the first physical sound is collected and the time when the second physical sound is collected.

[0012] The sound recognition system according to the sixth invention further includes a storage unit in which a plurality of recognition conditions for phoneme recognition set in advance are stored in any one of the first to fifth inventions, and the evaluation unit includes: a generation unit that generates a plurality of recognition histories corresponding to the evaluation information based on different recognition conditions; and a derivation unit that derives the recognition result based on the plurality of recognition histories.

[0013] The sound recognition system according to the seventh invention is the same as the sixth invention, wherein the derivation unit specifies feature amounts included in each of the plurality of recognition histories, and includes deriving the recognition result using the plurality of feature amounts.

[0014] The sound recognition method according to the eighth invention is a sound recognition method for evaluating physical sounds, and includes an acquisition step of acquiring evaluation information based on a first physical sound and a second physical sound that are picked up while propagating through media in different states, and an evaluation step of deriving a recognition result corresponding to the evaluation information using phoneme recognition.

Advantages of the Invention

[0015] According to the first to seventh inventions, the first sound collection unit and the second sound collection unit pick up physical sounds that propagate through media in different states. Therefore, by using the two types of physical sounds picked up by each sound collection unit, it is easy to derive a recognition result considering the conditions of the external environment. Thereby, it is possible to improve the recognition accuracy.

[0016] In particular, according to the second invention, the sound insulation unit is provided between the first sound collection unit and the second sound collection unit, and the first sound collection unit and the second sound collection unit are arranged. Therefore, each sound collection unit can suppress the collection of sounds propagated from sources other than the target medium. Thereby, it is possible to further improve the recognition accuracy.

[0017] In particular, according to the third invention, the sound insulation unit floats in the second medium. Therefore, each sound collection unit can easily pick up different physical sounds with the sound insulation unit as a boundary. Thereby, it is possible to improve the convenience.

[0018] In particular, according to the fourth invention, the evaluation unit includes identifying the floating position at the time of obtaining the evaluation information based on the position information. Therefore, even when the floating position of the sound insulation unit changes over time, it is possible to easily identify the generation source of the physical sound that has been evaluated. As a result, it is possible to further improve convenience.

[0019] In particular, according to the fifth invention, the evaluation unit includes identifying the generation region of the physical sound based on the difference between the time when the first physical sound was picked up and the time when the second physical sound was picked up. Therefore, compared with the case of using only the physical sound propagated from one medium, it is possible to easily identify the generation region of the physical sound. As a result, it is possible to improve the accuracy of identifying the generation position of the physical sound.

[0020] In particular, according to the sixth invention, the generation unit generates a plurality of recognition histories corresponding to the evaluation information based on different recognition conditions respectively. Further, the derivation unit derives a recognition result based on the plurality of recognition histories. That is, compared with the case of deriving a recognition result using only one recognition condition, it is possible to easily capture the characteristics of the evaluation information. Therefore, the recognition accuracy for the evaluation information can be improved. As a result, it is possible to further improve the recognition accuracy.

[0021] In particular, according to the seventh invention, the derivation unit identifies the feature amounts included in each of the plurality of recognition histories, and derives a recognition result using the plurality of feature amounts. Therefore, by identifying at least a part of the process of recognizing the evaluation information for each of the plurality of recognition conditions as a feature amount, a comprehensive recognition result can be derived. As a result, it is possible to further improve the recognition accuracy.

[0022] According to the eighth invention, the acquisition step acquires evaluation information based on the first physical sound and the second physical sound that are picked up while propagating through media in different states respectively. Therefore, by using two types of physical sounds with different propagation paths, it is possible to easily derive a recognition result considering the conditions of the external environment. As a result, it is possible to improve the recognition accuracy.

Brief Description of the Drawings

[0023]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0024] Hereinafter, an example of the sound recognition system and sound recognition method in the embodiments of the present invention will be described with reference to the drawings.

[0025] (First Embodiment: Sound Recognition System 100 and Sound Recognition Method) With reference to FIGS. 1(a) and 1(b), an example of the configuration of the sound recognition system 100 in the present embodiment will be described. FIG. 1(a) is a schematic diagram showing an example of the configuration of the sound recognition system 100 in the present embodiment, and FIG. 1(b) is a schematic diagram showing another example of the sound recognition system 100 in the present embodiment.

[0026] The sound recognition system 100 is used to evaluate physical sounds. In particular, the sound recognition system 100 can pick up a plurality of physical sounds propagating through media 5 in different states, and derive a recognition result based on each physical sound.

[0027] The sound recognition system 100 evaluates physical sounds propagating through, for example, the materials (solids) and spaces (gases) of a building, and physical sounds propagating through the sea (liquid) and the sea surface (gas). Physical sounds indicate sounds generated from artificial objects such as hammering sounds and machine sounds, as well as sounds generated from natural phenomena such as rain and wind. Any sound can be used according to the application.

[0028] The medium 5 includes a first medium 51 indicating any one of the states of gas, liquid, and solid, and a second medium 52 indicating a state different from that of the first medium 51. The substance used as the medium 5 can be arbitrarily selected according to the application. Each of the media 51 and 52 has a different sound propagation speed. Therefore, by evaluating each physical sound in consideration of the physical properties of each of the media 51 and 52, it is possible to easily derive a recognition result considering the conditions of the external environment.

[0029] The sound recognition system 100 includes a sound collection unit 1 and a sound recognition device 2. The sound collection unit 1 includes a first sound collection unit 1a and a second sound collection unit 1b. Each of the sound collection units 1a and 1b picks up physical sounds propagating through media 51 and 52 in different states. Each of the sound collection units 1a and 1b is connected via a wire or a communication network 3 described later so that various information such as the picked-up physical sound can be transmitted to the sound recognition device 2. Note that the sound collection unit 1 may include, for example, a plurality of each of the sound collection units 1a and 1b.

[0030] For example, as shown in FIG. 1(a), the first sound collection unit 1a is provided in a first medium 51 that represents a gas (e.g., air), and collects a physical sound (first physical sound) that propagates through the first medium 51. The second sound collection unit 1b is provided in a second medium 52 that represents a solid (e.g., concrete), and collects a physical sound (second physical sound) that propagates through the second medium 52. Note that the first physical sound and the second physical sound may represent the same physical sound that has propagated through different media 5, or may represent physical sounds with different sound sources or different generation timings, for example.

[0031] The sound recognition system 100 may include, for example, a sound insulation unit 15. The sound insulation unit 15 is provided between the first sound collection unit 1a and the second sound collection unit 1b, and the first sound collection unit 1a and the second sound collection unit 1b are arranged therein. The sound insulation unit 15 is provided, for example, at the interface of each medium 51, 52, and blocks the physical sound passing through the interface. Therefore, by providing the sound insulation unit 15, for example, it is possible to suppress the first sound collection unit 1a from collecting the second physical sound, and to suppress the second sound collection unit 1b from collecting the first physical sound.

[0032] For example, as shown in FIG. 1(b), the first sound collection unit 1a is provided in a first medium 51 that represents a gas (e.g., air), and collects a physical sound (first physical sound) that propagates through the first medium 51. The second sound collection unit 1b is provided in a second medium 52 that represents a liquid (e.g., water), and collects a physical sound (second physical sound) that propagates through the second medium 52. In this case, by using a material floating in the second medium 52 as the sound insulation unit 15, it is possible to make it easier to classify the types of physical sounds collected by each sound collection unit 1a, 1b. Note that, for example, the sound insulation unit 15 may house the sound recognition device 2.

[0033] Note that the sound recognition system 100 may include a plurality of, for example, the first sound collection unit 1a, the second sound collection unit 1b, the sound insulation unit 15, and the sound recognition device 2. In this case, for example, various information such as the collected physical sound may be transmitted and received for each sound recognition device 2.

[0034] The voice recognition system 100 may include a server 4 connected via a known communication network 3, as shown in FIG. 2, for example. The voice recognition system 100 may include a sensor for acquiring state information described later, for example.

[0035] As the sound collection unit 1, a known sound collection device such as a microphone that collects physical sound is used. The physical sound is a sound wave (for example, an elastic wave) that propagates through a medium 5 indicating any one of a gas, a liquid, and a solid. The sound wave indicates a human audible frequency band (for example, about 20 Hz to 20,000 Hz) such as an electronic device sound, a structure sound, a metal sound, etc., and may also indicate a frequency band (for example, about 1 Hz to several GHz) including an ultra-low frequency lower than the audible frequency and an ultrasonic wave higher than the audible frequency. In the voice recognition system 100, the frequency of the physical sound to be recognized can be arbitrarily set according to the application.

[0036] The voice recognition device 2 indicates a device for recognizing the physical sound collected via the sound collection unit 1. A known program for recognizing physical sound using phoneme recognition technology is stored in the voice recognition device 2. As the voice recognition device 2, a single-board computer such as Raspberry Pi (registered trademark) is used, and a known electronic device such as a personal computer (PC) may also be used.

[0037] The voice recognition device 2 is housed in the soundproof section 15 and can be provided at an arbitrary position according to the application. The voice recognition device 2 may include a plurality of devices 2a, 2b, as shown in FIG. 2, for example. In this case, for example, each of the devices 2a, 2b may be connected via the communication network 3.

[0038] FIG. 3 is a schematic diagram showing an example of the voice recognition method in the present embodiment. The voice recognition method includes an acquisition step S110 and an evaluation step S120. Note that the voice recognition method can be implemented using the voice recognition system 100.

[0039] <Acquisition step S110> The acquisition step S110 acquires evaluation information based on a first physical sound and a second physical sound that are picked up while propagating through media 51 and 52 in different states respectively. In the acquisition step S110, for example, the sound recognition device 2 acquires evaluation information based on the first physical sound picked up by the first sound pickup unit 1a and the second physical sound picked up by the second sound pickup unit 1b.

[0040] The evaluation information indicates data generated by digitally processing two types of physical sounds picked up using the first sound pickup unit 1a and the second sound pickup unit 1b. The evaluation information includes information on phonemes indicating the characteristics of the physical sound. Note that the digital processing of the physical sound may be performed by the sound pickup unit 1 or by the sound recognition device 2, and can be performed using any device according to the application.

[0041] In the acquisition step S110, for example, the evaluation information is acquired according to the timing at which each of the sound pickup units 1a and 1b picks up the physical sound. Alternatively, for example, the evaluation information pre-stored in the server 4 or the like may be acquired. The evaluation information is stored, for example, using a known file format. The evaluation information indicates two pieces of data generated based on each physical sound picked up by each of the sound pickup units 1a and 1b, and may alternatively indicate one piece of data generated based on a combination of two types of physical sounds picked up by each of the sound pickup units 1a and 1b.

[0042] <Evaluation step S120> The evaluation step S120 derives a recognition result corresponding to the evaluation information using phoneme recognition. The recognition result indicates, for example, the result of recognizing the characteristics of the physical sound. The recognition result includes information indicating normal or abnormal conditions at the sound source of the physical sound, and may also include information indicating the generation region of a physical sound that is different from normal. Note that the "generation region" indicates, for example, the position or component estimated to be the location where the physical sound is generated. In addition to the above, for example, in the case of a physical sound having a natural phenomenon such as rain or wind as the sound source, the intensity (degree) of the natural phenomenon such as the amount of rainfall or wind speed may be used as the recognition result.

[0043] The evaluation step S120 may derive a recognition result based on, for example, the result of comparing the magnitude of the first physical sound with the magnitude of the second physical sound. In this case, the evaluation step S120 can identify the generation region of the physical sound from the difference in the magnitudes of the respective physical sounds and derive it as the recognition result.

[0044] The evaluation step S120 may identify the generation region of the physical sound based on, for example, the difference between the time when the first physical sound was picked up and the time when the second physical sound was picked up. That is, in the evaluation step S120, by identifying the sound pickup units 1a and 1b that picked up the physical sound at an earlier time in light of the propagation speed of the physical sound in each medium 51, 52, the generation region of the physical sound can be identified and derived as the recognition result.

[0045] The evaluation step S120 may derive a recognition result corresponding to the evaluation information using, for example, a known phoneme recognition technique. The evaluation step S120 may refer to a learning model constructed using, for example, known machine learning and derive a recognition result corresponding to the evaluation information.

[0046] The learning model can be constructed by machine learning using a plurality of learning data, for example, with a pair of learning data being the reference evaluation information acquired in advance and the reference recognition result associated with the reference evaluation information. The reference evaluation information can be generated, for example, by picking up a physical sound emitted from an arbitrary sound source and performing digital processing. At this time, the reference recognition result indicates at least one of information regarding the sound source (for example, the generation region such as a component or a position) of the physical sound picked up when generating the reference evaluation information and information regarding the state such as normal or abnormal of the physical sound.

[0047] In the evaluation step S120, for example, the sound recognition device 2 may derive a recognition result corresponding to the evaluation information based on the preset recognition conditions for phoneme recognition. The recognition conditions indicate known conditions used in phoneme recognition technology. For example, they include the type of learning model used when extracting phonemes, the database (dictionary) used when evaluating phonemes, the threshold used when evaluating the extracted phonemes, and conditions that can affect the evaluation of the evaluation information, such as the environmental state when the physical sound is recorded. The recognition conditions indicate, for example, the characteristics of the physical sound generated from the parts of the machine, and by comparing with the evaluation information, the location where the physical sound associated with the evaluation information is generated can be specified. The recognition conditions are, for example, pre-stored in the sound recognition device 2.

[0048] Thereby, the sound recognition method in this embodiment ends. Note that each of the above-described steps S110 and S120 may be performed multiple times.

[0049] According to the sound recognition method in this embodiment, in the acquisition step S110, the evaluation information based on the first physical sound and the second physical sound collected by propagating through the media 5 in different states is acquired. Therefore, by using two types of physical sounds with different propagation paths, it is easy to derive a recognition result considering the conditions of the external environment. Thereby, it is possible to improve the recognition accuracy.

[0050] Hereinafter, the details regarding each configuration will be described.

[0051] <Sound collection unit 1> As the sound collection unit 1, for example, a microphone that can collect sounds in the audible frequency band of humans is used. In addition, for example, an AE sensor that can collect sounds in a frequency band higher than the audible frequency band of humans may be used, and a known device that can collect sound waves in a frequency band suitable for the application can be used.

[0052] For example, when collecting a physical sound propagating through a medium 5 indicating a solid state, the sound collection unit 1 is provided in contact with the solid and, for example, embedded. For example, when collecting a physical sound propagating through a medium 5 indicating a liquid state, the sound collection unit 1 is provided in contact with the liquid.

[0053] <Voice recognition device 2> Figure 4(a) is a schematic diagram showing an example of the configuration of the voice recognition device 2. The voice recognition device 2 includes, for example, a housing 20, a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, a storage unit 204, and I / Fs 205 to 207. Each component 201 to 207 is connected by an internal bus 210.

[0054] The CPU 201 controls the entire voice recognition device 2. The ROM 202 stores the operation code of the CPU 201. The RAM 203 is a work area used during the operation of the CPU 201.

[0055] The storage unit 204 stores various information such as recognition conditions used for voice recognition. A plurality of preset recognition conditions may be stored in the storage unit 204, for example. As the storage unit 204, a known data storage medium such as a memory card, an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc. is used.

[0056] The I / F 205 is a known interface for transmitting and receiving various information with the sound collection unit 1, the communication network 3, etc. connected according to the application. A plurality of I / Fs 205 may be provided, for example.

[0057] The I / F 206 is a known interface for transmitting and receiving various information with the input unit 208 connected according to the application. As the input unit 208, for example, a keyboard is used, and an administrator or the like who manages the voice recognition system 100, etc. inputs or selects various information or control commands of the voice recognition device 2 via the input unit 208.

[0058] I / F207 is a known interface for transmitting and receiving various information to and from a display unit 209 connected according to the application. The display unit 209 outputs various information stored in the storage unit 204, the processing status of the voice recognition device 2, and the like. As the display unit 209, for example, a display is used, and it may be a touch panel type, for example. In this case, the display unit 209 may be configured to include the input unit 208.

[0059] Note that the same I / Fs may be used for I / F205 to I / F207, for example, or a plurality of each of I / F205 to I / F207 may be used, for example. Further, at least one of the sound collection unit 1, the input unit 208, and the display unit 209 may be removed according to the situation.

[0060] FIG. 4(b) is a schematic diagram showing an example of the functions of the voice recognition device 2. The voice recognition device 2 includes an acquisition unit 21 and an evaluation unit 22. The voice recognition device 2 may include at least one of, for example, a position specifying unit 23, an output unit 24, and a storage unit 25. Each function shown in FIG. 4(b) is realized by the CPU 201 executing a program stored in the storage unit 204 and the like using the RAM 203 as a work area.

[0061] <<Acquisition Unit 21>> The acquisition unit 21 acquires evaluation information based on the physical sound collected via the sound collection unit 1. The acquisition unit 21 is used, for example, when performing the acquisition step S110 described above.

[0062] The acquisition unit 21 may acquire evaluation information, for example, by digitally processing the first physical sound received from the first sound collection unit 1a and the second physical sound received from the second sound collection unit 1b. The timing at which the acquisition unit 21 acquires the physical sound or the evaluation information from each of the sound collection units 1a and 1b can be arbitrarily set. The acquisition unit 21 may acquire the evaluation information stored in the storage unit 204 via the storage unit 25, for example.

[0063] The acquisition unit 21 may be used as a function of the device 2a provided, for example, inside the sound insulation unit 15 as shown in FIG. 2. In this case, the acquisition unit 21 may transmit the acquired evaluation information to the device 2b.

[0064] The acquisition unit 21 may include, in the evaluation information, for example, the time when the physical sound is picked up by the sound collection unit 1. In this case, for example, the difference between the time when the first physical sound is picked up by the first sound collection unit 1a and the time when the second physical sound is picked up by the second sound collection unit 1b may be included in the evaluation information.

[0065] <<Evaluation unit 22>> The evaluation unit 22 uses phoneme recognition to derive a recognition result corresponding to the evaluation information. The evaluation unit 22 is used, for example, when performing the above-described evaluation step S120. The evaluation unit 22 uses, for example, Julius or the like to recognize the phonemes included in the evaluation information, and derives a recognition result based on the recognized phonemes.

[0066] The evaluation unit 22 may specify, for example, the generation area of the physical sound associated with the evaluation information and derive it as a recognition result. In this case, the evaluation unit 22 specifies the generation area of the physical sound based on the difference between the time when the first physical sound is picked up by the first sound collection unit 1a and the time when the second physical sound is picked up by the second sound collection unit 1b. As a method for specifying the generation area of the physical sound, for example, an analysis method using a beamforming method can be mentioned.

[0067] <<Position specifying unit 23>> The position specifying unit 23 acquires, for example, position information indicating the existence position (for example, floating position) of the sound insulation unit 15. The position information indicates information obtained using, for example, a known GPS (Global Positioning System, Global Positioning Satellite). When the position information is acquired, for example, the evaluation unit 22 may specify the existence position of the sound insulation unit 15 at the time when the evaluation information is acquired based on the position information, and derive a recognition result including the specified existence position.

[0068] <<Output unit 24>> The output unit 24 outputs various types of information such as recognition results to the display unit 209 and the like. The output unit 24 may output various types of information such as recognition results to an electronic terminal held by, for example, an administrator of the speech recognition system 100. At this time, the output unit 24 may output various types of information via the communication network 3, and the output method is arbitrary.

[0069] <<Memory unit 25>> The memory unit 25 stores various types of information in the storage unit 204 or retrieves various types of information from the storage unit 204. The memory unit 25 performs storage or retrieval of various types of information, for example, according to the processing contents of the acquisition unit 21, the evaluation unit 22, the position specifying unit 23, and the output unit 24.

[0070] According to this embodiment, the first sound collection unit 1a and the second sound collection unit 1b collect physical sounds propagating through media 51 and 52 in different states, respectively. Therefore, by using the two types of physical sounds collected by each of the sound collection units 1a and 1b, it is possible to easily derive a recognition result considering the conditions of the external environment. Thereby, it is possible to improve the recognition accuracy.

[0071] Further, according to this embodiment, the sound insulation unit 15 is provided between the first sound collection unit 1a and the second sound collection unit 1b, and the first sound collection unit 1a and the second sound collection unit 1b are arranged therein. Therefore, each of the sound collection units 1a and 1b can suppress the collection of sounds propagated from sources other than the target medium 5. Thereby, it is possible to further improve the recognition accuracy.

[0072] Further, according to this embodiment, the sound insulation unit 15 floats in the second medium 52. Therefore, each of the sound collection units 1a and 1b can easily collect different physical sounds with the sound insulation unit 15 as a boundary. Thereby, it is possible to improve the convenience.

[0073] Further, according to the present embodiment, the evaluation unit 22 includes identifying the existence position (for example, the floating position) at the time when the evaluation information is acquired based on the position information. Therefore, even when the existence position of the sound insulation unit 15 changes over time, it is possible to easily identify the generation source of the evaluated physical sound. As a result, it is possible to further improve the convenience.

[0074] Further, according to the present embodiment, the evaluation unit 22 includes identifying the generation region of the physical sound based on the difference between the time when the first physical sound is picked up and the time when the second physical sound is picked up. Therefore, compared with the case where only the physical sound propagating from one medium 5 is used, it is possible to easily identify the generation region of the physical sound. As a result, it is possible to improve the accuracy of identifying the generation position of the physical sound.

[0075] (First Embodiment: Modification Example of Evaluation Unit 22) Next, a modification example of the evaluation unit 22 in the present embodiment will be described. The difference between the above-described embodiment and this modification example is that the evaluation unit 22 includes a generation unit and a derivation unit. Note that the description of the same content as that of the above-described embodiment will be omitted.

[0076] <<Generation Unit>> The generation unit generates a plurality of recognition histories corresponding to the evaluation information based on different recognition conditions. The generation unit generates a plurality of recognition histories using, for example, a known phoneme recognition technique such as Julius. In this case, for example, as the recognition conditions, known parameters that can affect the sound recognition can be used.

[0077] The generation unit acquires, for example, at least two recognition conditions stored in the storage unit 204, and generates recognition histories corresponding to the number of the acquired recognition conditions. The generation unit generates a plurality of recognition histories based on, for example, the preset types of recognition conditions and the number of recognition conditions. Note that the types of recognition conditions and the number of recognition conditions can be arbitrarily set according to the application.

[0078] The generation unit refers to, for example, a learning model associated with preset recognition conditions, and extracts phoneme information from the evaluation information. Further, the generation unit refers to, for example, a database associated with preset recognition conditions, and derives an evaluation result corresponding to the phoneme information. In this case, as shown in FIG. 5 for example, the generation unit generates a recognition history including the extracted phoneme information, the information of the learning model referred to, the information of the database referred to, and the derived evaluation result.

[0079] [Recognition history] The recognition history includes, for example, at least any one of phoneme information, recognition condition information, and an evaluation result, as shown in FIG. 5 for example. The recognition history includes, for example, one or more recognition candidates indicating a combination of phoneme information, recognition condition information, and an evaluation result. The phoneme information indicates phonemes extracted from the evaluation information. The recognition condition information indicates conditions (such as a learning model, a database, a threshold value, etc.) used for extracting phonemes and evaluating phonemes from the evaluation information. The evaluation result indicates the result of evaluating the phonemes extracted from the evaluation information.

[0080] [Phoneme information] The phoneme information includes, for example, at least any one of voice phoneme information and environmental phoneme information. In the voice recognition system 100, for example, voice phoneme information or environmental phoneme information corresponding to physical sounds can be set.

[0081] The voice phoneme information includes one or more phonemes uttered by a human (such as "k / a / j / i / d / e / th / u", etc.). The voice phoneme information includes phonemes corresponding to a known language including vowels and consonants. When the voice phoneme information includes two or more voice phonemes, the voice phoneme information may include information regarding the arrangement of each voice phoneme.

[0082] Environmental phoneme information includes one or more phonemes different from phonetic phonemes (for example, environmental phonemes: "@ / ? / # / @ / ? / #", etc.). Environmental phoneme information includes, for example, only phonemes different from vowels and consonants. As the phonemes included in the environmental phoneme information, for example, amplitudes with characteristics different from natural languages are used, and codes different from natural languages are assigned. Note that, as environmental phoneme information, for example, phonetic phonemes may be partially included. Further, when the environmental phoneme information includes two or more phonemes, the environmental phoneme information may include information regarding the arrangement of each sound.

[0083] The phonetic phoneme information and the environmental phoneme information may include, for example, at least one of a silent interval indicating the start of each sound (for example, a start silent interval indicated by "silB", etc.) and a silent interval indicating the end of each sound (for example, an end silent interval indicated by "silE", etc.). The start silent interval and the end silent interval can be extracted by known phoneme recognition techniques.

[0084] The phonetic phoneme information and the environmental phoneme information may include, for example, a pause interval. The pause interval indicates an interval shorter than the start silent interval and the end silent interval, and indicates, for example, an interval (length) comparable to the interval of a phoneme. The pause interval can be extracted by known phoneme recognition techniques.

[0085] In particular, since the environmental phoneme information includes a pause interval, a slight silent state caused by resonance associated with a combination of a plurality of environmental sounds can be extracted as the environmental phoneme information. Thereby, it becomes possible to improve the accuracy of environmental sound recognition.

[0086] [Learning model] The learning model is constructed, for example, using reference evaluation information acquired in advance and reference phonemes associated with the reference evaluation information, and is stored in a plurality of storage units such as the storage unit 204, for example.

[0087] The learning model may show, for example, a relational database as shown in FIG. 5, or may show, for example, a neural network, a function, etc. constructed by known machine learning. Note that the learning model may include, for example, a threshold value set for each reference phoneme. When constructing a plurality of learning models, at least any one of, for example, the type of reference evaluation information, the type of reference phoneme, and the linking method between the reference evaluation information and the reference phoneme is different for each.

[0088] The reference evaluation information represents digital data of the same type as the evaluation information, and represents, for example, data characterized by the degree of amplitude over a certain period. Note that the period of the amplitude represented by the reference evaluation information can be arbitrarily set according to the linked reference phoneme.

[0089] The reference phoneme represents information of the same type as the phoneme information. The reference phoneme includes one or more phonemes and includes, for example, information regarding an array of a plurality of phonemes.

[0090] The generation unit refers to, for example, any one of a plurality of learning models and extracts phoneme information corresponding to the evaluation information using phoneme recognition. The phoneme information includes one or more phonemes and represents, for example, an array of a plurality of phonemes.

[0091] [Database] The database is constructed using the recognition phonemes acquired in advance and the reference recognition results linked to the recognition phonemes, and is stored, for example, in a plurality in the storage unit 204.

[0092] The database may show, for example, a relational database as shown in FIG. 5, or may show, for example, a neural network, a function, etc. constructed by known machine learning. Note that the database may include, for example, a threshold value set for each reference recognition result. A plurality of databases are linked to a plurality of learning models and are stored, for example, in the storage unit 204 etc. in a state where one learning model is linked to one database.

[0093] Note that the learning model and the database may be associated with any number, such as one-to-many or many-to-many. In particular, when the learning model and the database are associated one-to-one, misrecognition can be reduced.

[0094] The recognition phoneme indicates information of the same type as the phoneme information and the reference phoneme. The recognition phoneme includes, for example, the same phoneme as the reference phoneme of the associated learning model.

[0095] The reference recognition result indicates information for specifying the content of the evaluation information. The reference recognition result includes, for example, information for deriving the recognition result.

[0096] [Status Information] The recognition condition information includes, for example, status information. The status information indicates features of the environment when the physical sound is picked up, generation conditions when the evaluation information is generated, and the like. The features of the environment include, for example, features such as temperature, temperature of the physical sound source, and pressure in the sound pickup space, and are measured using, for example, known sensors. The generation conditions include, for example, features such as a noise coefficient, a sound pressure coefficient, and a volume upper limit threshold value, and include, for example, known parameters when the physical sound is digitized.

[0097] [Evaluation Result] The evaluation result refers to the database and shows the result of evaluating the phoneme information. The evaluation result indicates, for example, the accuracy with which the phoneme information corresponds to a specific reference recognition result. The evaluation result may indicate, for example, "Yes" or "No" determined based on a threshold value, or may indicate a confidence score CMV such as "0.95" or "0.66". The evaluation result may include, for example, a reference recognition result that may correspond to the phoneme information.

[0098] The generation unit refers to the database associated with the learning model referred to when extracting the phoneme information, for example, and derives an evaluation result corresponding to the extracted phoneme information.

[0099] <<Derivation Unit>> The derivation unit derives a recognition result corresponding to the evaluation information based on a plurality of recognition histories. For example, the derivation unit derives the recognition result based on two or more recognition histories among the plurality of recognition histories generated by the generation unit.

[0100] The derivation unit, for example, specifies at least one of a plurality of pieces of information included in the recognition history as a feature amount. For example, as shown in FIG. 6, the derivation unit specifies the confidence measure CMV included in each of the plurality of recognition histories, and derives the recognition result using the plurality of confidence measures CMV. The derivation unit derives the recognition result based on, for example, the recognition tendency when the confidence measures CMV are arranged for each of the plurality of recognition histories.

[0101] The derivation unit may calculate, for example, the recognition tendency from the degree of deviation between the result using a known approximation formula for a plurality of feature amounts such as the arranged plurality of confidence measures CMV and the reference data stored in advance in the storage unit 204 or the like, and derive the recognition result. In addition to the above, for example, the recognition result may be derived using a known technique such as a Lucas sequence, a chi-square test, a t-test, an analysis of variance, or a Gray code for the plurality of feature amounts.

[0102] Note that the reference data stored in advance in the storage unit 204 or the like not only indicates data for comparison with a plurality of feature amounts, but may also indicate, for example, a function that can calculate the recognition tendency with the plurality of feature amounts as variables and the solution as the recognition tendency. In this case, reference data for deriving a recognition result according to the result of the function is stored in the storage unit 204 or the like. Further, the derivation unit may derive the recognition result from the plurality of feature amounts using, for example, a known classification technique or determination technique.

[0103] For example, as shown in FIG. 7, the derivation unit may derive a recognition result based on the results of calculating the recognition tendency of feature amounts for each group of conditions (the first condition group and the second condition group in FIG. 7) where some conditions are the same among a plurality of recognition histories. In this case, the derivation unit integrates the results of calculating the recognition tendency of feature amounts for each of the plurality of condition groups, and derives a recognition result based on the difference between the recognition tendencies (region R in FIG. 7). Note that the recognition result based on the above-described region R may be derived, for example, from the degree of deviation between the reference data stored in advance in the storage unit 204 or the like and the region R. In this case, as the recognition result, character strings such as "normal sound of ○○" when the degree of deviation is less than a preset threshold value and "abnormal sound of ○○" when the degree of deviation is greater than or equal to the threshold value may be derived.

[0104] For example, the condition groups may be classified for each learning model. In this case, it is possible to calculate the recognition tendency focusing on different phoneme information for each condition group. In addition to the above, for example, the condition groups may be classified for each state information (for example, for each of the sound collection units 1a and 1b that collect physical sounds when generating evaluation information). In this case, it is possible to calculate the recognition tendency focusing on different generation conditions for each condition group.

[0105] For example, the condition groups may be classified based on sequences, parameters, etc. used for phoneme recognition. The condition groups may be classified, for example, based on the transition history associated with dynamic allocation parameters and the directionality.

[0106] Note that in FIG. 6 and the like, an example of deriving a recognition result in two dimensions (the number of recognition histories and the confidence level CMV) is shown. However, for example, the derivation unit may derive a recognition result in multiple dimensions with a plurality of types of feature amounts specified from the recognition histories (for example, the confidence level CMV, information on the learning model, temperature, humidity, time, etc.), and the dimensions can be arbitrarily set according to the application.

[0107] For example, the above-described output unit 24 may output image data or the like in which the features of the extracted feature amounts are visible, such as a plot diagram of the feature amounts (for example, a graph showing the relationship between the number of recognition histories and the confidence level CMV in FIG. 6).

[0108] According to this modification example, the generation unit generates a plurality of recognition histories corresponding to the evaluation information based on different recognition conditions respectively. Further, the derivation unit derives a recognition result based on the plurality of recognition histories. That is, compared with the case of deriving a recognition result using only one recognition condition, it is easier to capture the characteristics of the evaluation information. For this reason, the recognition accuracy for the evaluation information can be improved. Thereby, it becomes possible to further improve the recognition accuracy.

[0109] Here, in the conventional speech recognition method using phoneme recognition, a recognition result corresponding to one physical sound is derived using one recognition condition. In particular, when a plurality of recognition candidates are extracted for one recognition condition, one recognition candidate specified using an index such as a confidence score CMV is derived as the recognition result. That is, the conventional speech recognition method does not assume reflecting the history at the time of deriving the recognition result in the recognition result. For this reason, when the evaluation information generated based on a combination of a plurality of types of physical sounds is the recognition target, it is difficult to recognize. Thereby, there is a concern that the types of physical sounds that can be recognized are limited.

[0110] On the other hand, according to the speech recognition system 100 in the present embodiment, the generation unit generates a plurality of recognition histories corresponding to the evaluation information based on different recognition conditions respectively. That is, different from the conventional speech recognition method, it is premised on reflecting the recognition history at the time of extracting the recognition candidate in the recognition result. For this reason, even when recognizing the evaluation information generated based on a plurality of types of physical sounds, it can be made easier to recognize compared with the conventional speech recognition method. Thereby, it becomes possible to expand the types of physical sounds that can be recognized.

[0111] Further, according to the present embodiment, the derivation unit specifies the feature amounts included in each of the plurality of recognition histories, and derives a recognition result using the plurality of feature amounts. For this reason, by specifying at least a part of the process of recognizing the evaluation information for each of the plurality of recognition conditions as a feature amount, a comprehensive recognition result can be derived. Thereby, it becomes possible to further improve the recognition accuracy.

[0112] Further, according to the present embodiment, for example, the recognition history may include phoneme information, recognition condition information, and an evaluation result. In this case, based on the differences in each information that varies for each recognition condition, it is possible to easily realize the derivation of the recognition result. Thereby, it becomes possible to further improve the recognition accuracy for the evaluation information.

[0113] Further, according to the present embodiment, for example, the recognition condition information includes information for specifying a learning model used for phoneme recognition. In this case, it is possible to easily identify the difference in evaluation that may occur due to the difference in the learning model. Thereby, it becomes possible to expand the range of extractable phonemes while suppressing a decrease in recognition accuracy.

[0114] Further, according to the present embodiment, for example, the recognition condition information includes information for specifying a database used for phoneme recognition. In this case, it is possible to easily identify the difference in evaluation that may occur due to the difference in the database. Thereby, it becomes possible to expand the range of recognizable phonemes while suppressing a decrease in recognition accuracy.

[0115] Further, according to the present embodiment, for example, the generation unit may refer to any one of a plurality of learning models, extract phoneme information from the physical sound information, refer to a database associated with the referred learning model, and derive evaluation information corresponding to the phoneme information. In this case, compared to the case where only one learning model is implemented, it is possible to expand the possibility of selecting a learning model suitable for the physical sound information. Thereby, it becomes possible to further improve the recognition accuracy for the physical sound.

[0116] (Second Embodiment: Voice Recognition System 100) Next, an example of the voice recognition system 100 in the second embodiment will be described. The difference between the above-described embodiment and the present embodiment is that the recognition conditions used for subsequent processing are selected based on the generated recognition history. Note that the description of the same content as the above-described embodiment will be omitted.

[0117] As shown in, for example, FIG. 8, the generation unit generates a first recognition history corresponding to the evaluation information based on the first recognition condition. After that, the generation unit selects a second recognition condition different from the first recognition condition based on the first recognition history. Note that, as a method of selecting the second recognition condition based on the first recognition history, for example, a table or the like in which the features of a plurality of recognition histories are associated with a plurality of recognition conditions can be used, and a known technique can be used according to the application.

[0118] After that, the generation unit generates a second recognition history corresponding to the evaluation information based on the second recognition condition. The generation unit selects a new recognition condition based on the recognition history, for example, according to the preset number of generated recognition histories.

[0119] For example, the generation unit may select a new recognition condition based on the recognition history until a recognition history having preset features is generated. As this “feature”, for example, a threshold value of the degree of change of the recognition histories generated in order may be used, and when the degree of change is equal to or less than the threshold value, it may be set to end without selecting a new recognition condition.

[0120] For example, the generation unit may perform selection of a new recognition condition and generation of a new recognition history based on the recognition history until the feature amount (for example, the confidence level CMV) included in the recognition history reaches a preset shrinkage value or dispersion value. In addition to the above, for example, the generation unit may perform selection of a new recognition condition and generation of a new recognition history based on the recognition history until the feature amount of the recognition history is included in a preset allowable range or deviates from the allowable range.

[0121] For example, the generation unit may perform selection of a new recognition condition and generation of a new recognition history based on the recognition history until specific information is included in the state information included in the recognition history. In this case, for example, it can be arbitrarily set according to the application, such as continuing to generate the recognition history until the temperature in the sound collection environment of the physical sound exceeds a certain value.

[0122] Next, the derivation unit derives a recognition result based on a plurality of recognition histories including at least the first recognition history and the second recognition history. Note that the method for deriving the recognition result is the same as that in the above-described embodiment.

[0123] According to this embodiment, in addition to the above-described embodiment, the generation unit selects a second recognition condition based on the first recognition history and generates a second recognition history based on the second recognition condition. Therefore, based on the results of the recognition history, it is possible to easily select the recognition conditions necessary for improving the recognition accuracy. As a result, it is possible to further improve the recognition accuracy for the evaluation information.

[0124] (Third Embodiment: Voice Recognition System 100) Next, an example of the voice recognition system 100 in the third embodiment will be described. The difference between the above-described embodiment and this embodiment is that the generation unit includes a plurality of recognition units. Note that descriptions of the same content as in the above-described embodiment will be omitted.

[0125] The generation unit includes, for example, as shown in FIG. 9, a first recognition unit and a second recognition unit. The first recognition unit and the second recognition unit generate recognition histories using different recognition conditions (the first recognition condition and the second recognition condition in FIG. 9) (the first recognition history and the second recognition history in FIG. 9). Note that the generation unit may include, for example, three or more recognition units. By including a plurality of recognition units in the generation unit, for example, for each feature of the physical sounds picked up from the first sound collection unit 1a and the second sound collection unit 1b among one piece of evaluation information, a plurality of recognition histories can be generated by parallel processing. As a result, it is possible to realize real-time response processing. Note that the generation unit may generate a plurality of recognition histories for, for example, a plurality of pieces of evaluation information. In this case, the generation unit generates recognition histories using different recognition units for each of the plurality of pieces of evaluation information, and may also generate recognition histories for each of the plurality of recognition units for one evaluation information group obtained by aggregating the plurality of pieces of evaluation information.

[0126] The generation unit may select recognition conditions used for the first recognition unit and the second recognition unit based on, for example, the first recognition history and the second recognition history, and may select different recognition conditions for the first recognition unit and the second recognition unit, respectively. In this case, the generation unit generates a third recognition history via the first recognition unit and a fourth recognition history via the second recognition unit based on the selected recognition conditions. The generation unit selects recognition conditions based on the recognition history according to the preset number of generated recognition histories, and generates a new recognition history via each recognition unit.

[0127] Next, the derivation unit derives a recognition result based on a plurality of recognition histories including at least the first recognition history, the second recognition history, the third recognition history, and the fourth recognition history. Note that the method for deriving the recognition result is the same as that in the above-described embodiment.

[0128] According to the present embodiment, in addition to the above-described embodiment, the generation unit includes a first recognition unit and a second recognition unit that generate recognition histories using different recognition conditions. Therefore, compared with the case of repeatedly generating recognition histories using the same generation unit, the processing time can be shortened. Thereby, it is possible to improve the reaction speed of speech recognition.

[0129] Also, according to the present embodiment, the generation unit selects recognition conditions in the first recognition unit and the second recognition unit based on the first recognition history and the second recognition history, and generates a third recognition history and a fourth recognition history based on the selected recognition conditions. Therefore, it is possible to easily select recognition conditions necessary for improving the recognition accuracy based on the results of the recognition history. Thereby, it is possible to further improve the recognition accuracy for the evaluation information.

[0130] Note that according to each of the above-described embodiments, for example, the plurality of learning models may include a phoneme model constructed using phonemes and an environmental phoneme model constructed using environmental phonemes. That is, when recognizing the evaluation information, different types of phonemes can be used for recognition. Therefore, it is possible to increase the possibility of recognizing physical sounds that cannot be recognized by phonemes alone. Thereby, it is possible to expand the recognition target.

[0131] Also, according to each of the above-described embodiments, for example, the environmental phoneme model may be constructed using only phonemes different from vowels and consonants. In this case, when recognizing environmental sound information with reference to the environmental sound model, misrecognition with voice information can be suppressed. Thereby, it becomes possible to further improve the accuracy of environmental sound recognition.

[0132] Also, according to each of the above-described embodiments, for example, the generation unit can refer to two or more of a plurality of learning models, extract a plurality of phoneme information from physical sound information, refer to databases associated with each of the referred learning models, and generate a recognition history based on each of the plurality of phoneme information. In this case, for example, even if information on a plurality of sound waves having different characteristics such as frequency is included as evaluation information, by using a plurality of learning models, phoneme information can be extracted for each of the plurality of sound wave information. Thereby, it becomes possible to derive a recognition result based on a plurality of phoneme information.

[0133] Although embodiments of the present invention have been described, the above-described embodiments are presented as examples and are not intended to limit the scope of the invention. The above-described novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. The above-described embodiments and their modifications are included in the scope and gist of the invention and are included in the invention described in the claims and its equivalent scope.

Explanation of Reference Numerals

[0134] 1: Sound collection unit 1a: First sound collection unit 1b: Second sound collection unit 15: Sound insulation unit 2: Sound recognition device 2a: Device 2b: Device 20: Housing 21: Acquisition unit 22: Evaluation unit 23: Position specifying unit 24: Output unit 25: Memory unit 3: Communication network 4: Server 5: Medium 51: First medium 52: Second medium 100: Speech recognition system 201: CPU 202: ROM 203: RAM 204: Storage unit 205: I / F 206: I / F 207: I / F 208: Input unit 209: Display unit 210: Internal bus S110: Acquisition step S120: Evaluation step

Claims

**Claim 1** A sound recognition system for evaluating physical sounds, comprising: a first sound collection unit and a second sound collection unit for collecting the physical sounds propagating through media in different states respectively; an acquisition unit for acquiring evaluation information based on a first physical sound collected by the first sound collection unit and a second physical sound collected by the second sound collection unit; an evaluation unit for deriving a recognition result corresponding to the evaluation information using phoneme recognition; characterized by comprising a sound recognition system. **Claim 2** further comprising a sound insulation part provided between the first sound collection unit and the second sound collection unit, where the first sound collection unit and the second sound collection unit are arranged; characterized by the sound recognition system according to claim 1. **Claim 3** the first medium through which the first physical sound propagates is a gas; the second medium through which the second physical sound propagates is a liquid; the sound insulation part floats in the second medium; characterized by the sound recognition system according to claim 2. **Claim 4** further comprising a position identification unit for acquiring position information indicating the floating position of the sound insulation part; the evaluation unit includes identifying the floating position at the time when the evaluation information is acquired based on the position information; characterized by the sound recognition system according to claim 3. **Claim 5** the evaluation unit includes identifying the generation area of the physical sound based on the difference between the time when the first physical sound is collected and the time when the second physical sound is collected; characterized by the sound recognition system according to claim 1. **Claim 6** further comprising a storage unit in which a plurality of preset phoneme recognition recognition conditions are stored; the evaluation unit includes a generation unit for generating a plurality of recognition histories corresponding to the evaluation information based on different recognition conditions respectively; a derivation unit for deriving the recognition result based on the plurality of recognition histories; characterized by including the sound recognition system according to any one of claims 1 to 5. **Claim 7** the derivation unit includes identifying feature amounts included in each of the plurality of recognition histories; deriving the recognition result using the plurality of feature amounts; characterized by including the sound recognition system according to claim 6. **Claim 8** A sound recognition method for evaluating physical sounds, comprising: an acquisition step of acquiring evaluation information based on a first physical sound and a second physical sound collected by propagating through media in different states respectively; an evaluation step of deriving a recognition result corresponding to the evaluation information using phoneme recognition; characterized by comprising a sound recognition method.

Citation Information

Patent Citations

  • Sound recognition system and sound recognition method

    JP2023108586A