Aerosol volume estimation method, aerosol volume estimation device, and program

The method accurately estimates aerosol production by calculating acoustic features and speaker identity, addressing the inaccuracy of existing methods and enabling risk evaluation and mitigation.

JP7839788B2Active Publication Date: 2026-04-02PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods fail to accurately estimate the amount of aerosols generated during speech, relying solely on sound pressure level measurements.

Method used

An aerosol amount estimation method that calculates acoustic features from speech data, determines speaker identity using a trained model, and estimates aerosol production based on the similarity between speech at rest and speaking states.

Benefits of technology

Accurately estimates aerosol generation during speech, enabling effective evaluation of infection risk and prompting measures to reduce aerosol exposure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839788000001
    Figure 0007839788000001
  • Figure 0007839788000002
    Figure 0007839788000002
  • Figure 0007839788000003
    Figure 0007839788000003
Patent Text Reader

Abstract

This aerosol quantity estimation method comprises: determining whether the sound pressure level of an utterance of a speaker is higher than a predetermined sound pressure level (S103); when the sound pressure level is higher than the predetermined sound pressure level, calculating an acoustic feature quantity from utterance data produced from the utterance of the speaker (S104); calculating, from the acoustic feature quantity, a first speaker feature quantity indicating the speaker individuality of the utterance data using a trained model (S105); calculating the degree of similarity between a second speaker feature quantity that is the speaker feature quantity of the speaker when calm and the first speaker feature quantity (S106); and estimating an aerosol quantity generated from the speaker according to the degree of similarity (S107).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an aerosol amount estimation method, an aerosol amount estimation device, and a program.

Background Art

[0002] Patent Document 1 discloses an alarm device that measures the magnitude of conversation volume and notifies the risk of droplet infection.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, it is difficult to accurately estimate the amount of aerosol generated from a speaker during speech by only measuring the volume.

[0005] [[ID=':38]]<o000028>An object of the present disclosure is to provide an aerosol amount estimation method, an aerosol amount estimation device, and a program that can accurately estimate the amount of aerosol generated from a speaker during speech.

Means for Solving the Problems

[0006] An aerosol amount estimation method according to an aspect of the present disclosure determines whether a sound pressure level in a speaker's speech is greater than a predetermined sound pressure level. When the sound pressure level is greater than the predetermined sound pressure level, an acoustic feature amount is calculated from speech data of the speaker's speech, and a first speaker feature amount representing the speaker property of the speech data is calculated from the acoustic feature amount using a learned model. A similarity between a second speaker feature amount that is the speaker feature amount of the speaker at rest and the first speaker feature amount is calculated, and the amount of aerosol generated from the speaker is estimated according to the similarity.

Effects of the Invention

[0007] This disclosure provides an aerosol volume estimation method, an aerosol volume estimation device, and a program that can accurately estimate the amount of aerosol generated by a speaker during speech. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is a block diagram of an aerosol volume estimation device according to an embodiment. [Figure 2] Figure 2 is a block diagram of the speaker feature calculation unit according to the embodiment. [Figure 3] Figure 3 is a flowchart of the aerosol volume estimation process according to the embodiment. [Figure 4] Figure 4 is a graph showing the relationship between the sound pressure level of speech and the amount of aerosol. [Figure 5] Figure 5 is a graph showing the correlation between aerosol volume and similarity. [Modes for carrying out the invention]

[0009] (Knowledge that forms the basis of this disclosure) In technologies like that described in Patent Document 1, the level of danger is indicated in stages by changing the color of the light emitted or altering the volume, based on the level of conversational volume. However, the amount of aerosols generated by the person speaking (speaker) is not estimated.

[0010] Figure 4 is a graph showing the relationship between the sound pressure level of speech and the amount of aerosol produced. As shown in the graph in Figure 4, when the sound pressure level exceeds a certain level, there is variation in the amount of aerosol produced, making it difficult to accurately estimate the amount of aerosol by measuring only the sound pressure level.

[0011] The inventors discovered that the amount of aerosol produced when a speaker is speaking is correlated with the similarity between the speaker's speechiness during that speech and their speechiness when they are at rest, as shown in Figure 5. Based on this, the inventors have found an aerosol production method that can accurately estimate the amount of aerosol produced by a speaker during speech by using this similarity as an indicator for estimating aerosol production.

[0012] An aerosol quantity estimation method according to one aspect of the present disclosure determines whether the sound pressure level of a speaker's utterance is greater than a predetermined sound pressure level, calculates acoustic features from the utterance data of the speaker if the sound pressure level is greater than the predetermined sound pressure level, calculates a first speaker feature representing the speaker of the utterance data from the acoustic features using a trained model, calculates the similarity between a second speaker feature, which is the speaker feature of the speaker when he is at rest, and the first speaker feature, and estimates the amount of aerosol generated by the speaker according to the similarity.

[0013] According to this, the aerosol volume estimation method calculates a first speaker feature using a trained model to identify the speaker when the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level, and estimates the aerosol volume according to the similarity between this first speaker feature and the second speaker feature when the speaker is at rest. The aerosol volume estimation method can accurately estimate the amount of aerosol generated by the speaker by calculating the similarity, taking advantage of the correlation between the similarity between the speaker's speechiness when the speaker is uttering and the speaker's speechiness when the speaker is at rest.

[0014] Furthermore, in the estimation described above, the amount of aerosol generated may be estimated according to the degree of similarity by using a correlation where the amount of aerosol generated increases as the degree of similarity decreases.

[0015] In addition, the estimation may involve estimating the amount of aerosols generated in predetermined time units at predetermined time units and calculating the cumulative value of the aerosol amounts obtained from the start of the estimation.

[0016] Therefore, the total amount of aerosol from the start of estimation can be estimated, and the infection risk due to the amount of aerosol can be effectively evaluated.

[0017] Furthermore, it may be determined whether the integrated value is greater than a predetermined aerosol amount, and a warning may be given when the integrated value is greater than the predetermined aerosol amount.

[0018] Therefore, a warning can be given when it is determined that the infection risk is high, and it is possible to prompt the user to take measures to reduce the amount of aerosol.

[0019] Furthermore, it may be determined whether the integrated value is greater than a predetermined aerosol amount, and when the integrated value is greater than the predetermined aerosol amount, the ventilation device or air purifier arranged in the space where the speaker is located may be operated.

[0020] Therefore, when it is determined that the infection risk is high, the ventilation device or air purifier can be operated, and the amount of aerosol can be effectively reduced.

[0021] Also, the second speaker feature amount may represent the speaker property of the speech data obtained by the speaker reading aloud a predetermined sentence.

[0022] An aerosol amount estimation device according to an aspect of the present disclosure includes a sound pressure level determination unit that determines whether a sound pressure level in a speaker's speech is greater than a predetermined sound pressure level, an acoustic feature amount calculation unit that calculates an acoustic feature amount from speech data of the speaker's speech when the sound pressure level is greater than the predetermined sound pressure level, a speaker feature amount calculation unit that calculates a similarity of a first speaker feature amount representing the speaker property of the speech data from the acoustic feature amount using a learned model, a second speaker feature amount that is the speaker feature amount of the speaker at rest, a similarity calculation unit that calculates a similarity between the first speaker feature amount and the second speaker feature amount, and an estimation unit that estimates an aerosol amount generated from the speaker according to the similarity.

[0023] According to this, the aerosol volume estimation method calculates a first speaker feature using a trained model to identify the speaker when the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level, and estimates the aerosol volume according to the similarity between this first speaker feature and the second speaker feature when the speaker is at rest. The aerosol volume estimation method can accurately estimate the amount of aerosol generated by the speaker by calculating the similarity, taking advantage of the correlation between the similarity between the speaker's speechiness when the speaker is uttering and the speaker's speechiness when the speaker is at rest.

[0024] A program according to one aspect of this disclosure causes a computer to execute the aerosol quantity estimation method.

[0025] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0026] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.

[0027] (Embodiment 1) Figure 1 is a block diagram showing the configuration of the aerosol volume estimation device 100 according to this embodiment. The aerosol volume estimation device 100 estimates the amount of aerosol generated by the speaker (user). Specifically, the aerosol volume is the number of fine particles of liquid containing saliva that are expelled from the speaker into the space where the speaker is located while the speaker is speaking. For example, the aerosol volume estimation device 100 is included in a terminal device such as a smartphone or tablet. The functions of the aerosol volume estimation device 100 may be realized by a single device or by multiple devices. For example, some functions of the aerosol volume estimation device 100 may be realized by a terminal device, and other functions may be realized by a server or the like that can communicate with the terminal device.

[0028] As shown in Figure 1, the aerosol quantity estimation device comprises a sound acquisition unit 101, a sound pressure level determination unit 102, an acoustic feature calculation unit 103, a speaker feature calculation unit 104, a storage unit 105, a similarity calculation unit 106, an aerosol quantity estimation unit 107, and an output unit 108.

[0029] The voice acquisition unit 101 acquires speech data, which is the audio data of the speaker's utterances. For example, the voice acquisition unit 101 is a microphone and generates speech data by converting the acquired audio into an audio signal. The voice acquisition unit 101 may also acquire speech data generated outside the aerosol volume estimation device 100.

[0030] The sound pressure level determination unit 102 measures the sound pressure level of the spoken speech from the speech data and determines whether the measured sound pressure level is greater than a predetermined sound pressure level. The sound pressure level of the spoken speech may, for example, be the amplitude at the peak of the speech waveform during a predetermined period of the speech data. If there are multiple peaks during the predetermined period, the sound pressure level of the spoken speech may be the maximum value of the amplitudes of the multiple peaks, or the average value of the amplitudes of the multiple peaks. This predetermined period may be, for example, the period from a time earlier than the current time (latest time) to the current time. The first time interval may be, for example, 100 seconds or less. Furthermore, the loudness of the sound indicated by the sound data may be the amplitude at the current time of the envelope connecting the tangents of the peaks of the speech waveform of the sound data, the maximum value of the envelope during a predetermined period, or the average value of the envelope during a predetermined period.

[0031] The acoustic feature calculation unit 103 calculates acoustic features of the speech from the speech data when the sound pressure level determination unit 102 determines that the sound pressure level of the speaker's speech is greater than a predetermined sound pressure level. For example, the acoustic feature calculation unit 103 calculates MFCC (Mel Frequency Cepstral Coefficient), which is a feature of the speech, as an acoustic feature from the speech data. MFCC is a feature that represents the characteristics of the speaker's vocal tract and is commonly used in speech recognition. More specifically, MFCC is an acoustic feature obtained by analyzing the frequency spectrum of speech based on human auditory characteristics. The acoustic feature calculation unit 103 may also calculate the acoustic feature from the speech data by applying a Mel filter bank to the speech signal, or it may calculate the acoustic feature from the spectrogram of the speech signal.

[0032] The speaker feature calculation unit 104 extracts a first speaker feature from the acoustic features calculated from the speech data to identify the speaker of the utterance indicated by the speech data. In other words, the first speaker feature represents the speaker identity of the speech data. More specifically, the speaker feature calculation unit 104 extracts the first speaker feature from the acoustic features using a trained DNN.

[0033] For example, the speaker feature calculation unit 104 extracts the first speaker feature using the x-vector method. Here, the x-vector method is a method for calculating speaker features, which are speaker-specific features called x-vectors. Figure 2 is a block diagram showing an example configuration of the speaker feature calculation unit 104. As shown in Figure 2, for example, the speaker feature calculation unit 104 includes a frame connection processing unit 201 and a DNN 202.

[0034] The frame concatenation processing unit 201 connects multiple acoustic features and outputs the resulting acoustic features to the DNN202. For example, the frame concatenation processing unit 201 connects multiple frames of MFCC, which are acoustic features, and outputs them to the input layer of the DNN202. For example, the frame concatenation processing unit 201 generates a 1200-dimensional vector by connecting 50 frames of MFCC parameters, each consisting of 24-dimensional features, and outputs the generated vector to the input layer of the DNN202.

[0035] DNN202 is a pre-trained machine learning model that outputs a first speaker feature corresponding to the input acoustic features. In the example shown in Figure 2, DNN202 is a neural network consisting of an input layer, multiple hidden layers, and an output layer. Furthermore, DNN202 is pre-generated by machine learning using multiple training data 203. Each of the multiple training data 203 is data that links speaker identification information with speaker utterance data. In other words, DNN202 is a pre-trained model that takes utterance data as input and outputs speaker identification information (speaker label) for the utterance data, but in this embodiment, DNN202 outputs a first speaker feature generated as intermediate data. Note that a pre-trained machine learning model trained using machine learning other than Deep Neural Network may be used instead of DNN202.

[0036] Specifically, the output layer consists of nodes that output speaker labels equal to the number of speakers included in the training data 203. The multiple hidden layers consist of, for example, 2 to 3 hidden layers, and each hidden layer calculates the first speaker feature. The hidden layer that calculates the first speaker feature outputs the calculated first speaker feature as the output of DNN202.

[0037] The memory unit 105 is composed of a rewritable non-volatile memory, such as a hard disk drive or a solid-state drive. The memory unit 105 stores a second speaker feature, which is the first speaker feature of the speaker when they are healthy. For example, the second speaker feature is a speaker feature obtained in advance from speech data of utterances made by the speaker when they are at rest. Speech data of utterances made by the speaker when they are at rest is, for example, speech data obtained by having the speaker read a predetermined sentence aloud. Alternatively, the speech data of utterances made by the speaker when they are at rest may be speech data from utterances made when the speaker is presumed to be at rest (calm) based on biological information such as the speaker's body movements, heart rate, body temperature, sweating, voice, and facial expressions. The second speaker feature may be calculated from multiple first speaker features obtained in multiple past aerosol quantity estimation processes. For example, it may be the average or median of multiple first speaker features obtained in multiple past aerosol quantity estimation processes.

[0038] The similarity calculation unit 106 calculates the similarity between the first speaker feature output from the speaker feature calculation unit 104 and the second speaker feature stored in the memory unit 105. For example, the similarity calculation unit 106 calculates the cosine distance (also called cosine similarity), which represents the vector angle between the first speaker feature and the second speaker feature, by calculating the cosine using the inner product in the vector space model. In this case, a larger value for the vector angle indicates lower similarity. Alternatively, the similarity calculation unit 106 may calculate the cosine distance, which takes a value between -1 and 1, using the inner product of the vector representing the first speaker feature and the vector representing the second speaker feature. In this case, a larger value for the cosine distance indicates higher similarity. A larger similarity indicates that the first speaker feature and the second speaker feature are similar, while a smaller similarity indicates that the first speaker feature and the second speaker feature are not similar.

[0039] The aerosol amount estimation unit 107 estimates the amount of aerosol generated by the speaker based on the similarity calculated by the similarity calculation unit 106. Specifically, the aerosol amount estimation unit 107 estimates the amount of aerosol according to the similarity using a correlation relationship, as shown in Figure 5, where the amount of aerosol generated increases as the similarity decreases. This correlation relationship may be a correlation between the similarity and the amount of aerosol generated in a predetermined time unit.

[0040] Here, the processing by the voice acquisition unit 101, the sound pressure level determination unit 102, the acoustic feature calculation unit 103, the speaker feature calculation unit 104, and the similarity calculation unit 106 may be repeated at predetermined intervals. In this case, the aerosol amount estimation unit 107 estimates the aerosol amount for a predetermined time unit when the sound pressure level determination unit 102 determines that the sound pressure level of the utterance is greater than a predetermined sound pressure level, and calculates the cumulative value of the aerosol amount obtained from the start of the estimation. Therefore, the total amount of aerosols from the start of the estimation can be estimated, and the risk of infection due to the amount of aerosols can be effectively evaluated.

[0041] Furthermore, the aerosol volume estimation unit 107 may determine whether the calculated cumulative value is greater than a predetermined aerosol volume, and if the cumulative value is greater than the predetermined aerosol volume, it may determine that the risk of infection is high. Note that the aerosol volume estimation unit 107 may determine the risk of infection rather than determining whether the risk is high or not. For example, the aerosol volume estimation unit 107 may determine that the lower the similarity, the higher the risk of infection. The result of the determination may be shown in a multi-level classification such as "risk of infection," "high risk of infection," or "very high risk of infection," or it may be shown as a numerical value indicating the risk of infection.

[0042] The output unit 108 notifies the speaker of the determination result obtained by the aerosol volume estimation unit 107. For example, the output unit 108 is a display or speaker provided by the terminal device, and notifies the speaker of the determination result by display or sound. The output unit 108 may also output the determination result to an external device. The output unit 108 may also notify the speaker of the determination result (for example, a warning indicating a high risk of infection) only if it is determined that the risk of infection is high. This allows for a warning to be issued when a high risk of infection is determined, prompting the user to take measures to reduce the amount of aerosols.

[0043] Furthermore, the output unit 108 may activate a ventilation device or air purifier located in the space where the speaker is located if it determines that the risk of infection is high. Specifically, the output unit 108 may activate the ventilation device or air purifier by sending a control signal to the ventilation device or air purifier to activate it if it determines that the risk of infection is high. In this way, the ventilation device or air purifier can be activated when it is determined that the risk of infection is high, and the amount of aerosol can be effectively reduced.

[0044] The following describes the aerosol volume estimation process using the aerosol volume estimation device 100. Figure 3 is a flowchart of the aerosol volume estimation process using the aerosol volume estimation device 100. Note that this explanation assumes that one speaker is pre-registered in the aerosol volume estimation device 100.

[0045] First, in the aerosol volume estimation device 100, the voice acquisition unit 101 acquires speech data, which is the voice data of the speaker's utterance (S101).

[0046] Next, the sound pressure level determination unit 102 measures the sound pressure level of the spoken voice from the speech data (S102) and determines whether the measured sound pressure level is greater than a predetermined sound pressure level (S103).

[0047] Next, if the sound pressure level determination unit 102 determines that the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level (Yes in S103), the acoustic feature calculation unit 103 calculates acoustic features of the utterance from the utterance data (S104). If the sound pressure level determination unit 102 determines that the sound pressure level of the speaker's utterance is less than or equal to a predetermined sound pressure level (No in S103), step S101 is executed.

[0048] Next, the speaker feature calculation unit 104 calculates a first speaker feature to identify the speaker of the utterance indicated by the utterance data from the acoustic features calculated from the utterance data (S105). Specifically, the speaker feature calculation unit 104 outputs a first speaker feature corresponding to the input acoustic features.

[0049] Next, the similarity calculation unit 106 calculates the similarity between the first speaker feature output from the speaker feature calculation unit 104 and the second speaker feature stored in the memory unit 105 (S106).

[0050] Next, the aerosol amount estimation unit 107 estimates the amount of aerosol generated by the speaker based on the similarity calculated by the similarity calculation unit 106 (S107). Specifically, the aerosol amount estimation unit 107 uses a correlation where the amount of aerosol generated increases as the similarity decreases to estimate the amount of aerosol generated in a predetermined time unit according to the similarity. The calculated amount of aerosol generated in a predetermined time unit may be stored in the storage unit 105.

[0051] Next, the aerosol quantity estimation unit 107 calculates the cumulative value of the aerosol quantities obtained from the start of the estimation (S108). Specifically, the aerosol quantity estimation unit 107 calculates the cumulative value by summing up the one or more aerosol quantities obtained from the start of the estimation that are stored in the memory unit 105.

[0052] Next, the aerosol volume estimation unit 107 determines whether the calculated cumulative value is greater than a predetermined aerosol volume (S109).

[0053] If the cumulative value is greater than a predetermined aerosol amount (Yes in S109), the output unit 108 notifies the speaker of the determination result obtained by the aerosol amount estimation unit 107 (S110). The output unit 108 may also send a control signal to the ventilation device or air purifier to operate it if the cumulative value is greater than a predetermined aerosol amount (Yes in S109). If the cumulative value is less than or equal to the predetermined aerosol amount (No in S109), step S101 is executed.

[0054] In the above explanation, an example was shown where one speaker is pre-registered, but multiple speakers may be registered. In this case, the second speaker feature quantity for each speaker is stored in the memory unit 105. In addition, information identifying the speaker is input to the aerosol quantity estimation device 100, and the above processing is performed using the second speaker feature quantity of the identified speaker.

[0055] As described above, the aerosol quantity estimation device 100 determines whether the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level. If the sound pressure level is greater than the predetermined sound pressure level, the aerosol quantity estimation device 100 calculates acoustic features from the speech data of the speaker's utterance. The aerosol quantity estimation device 100 uses a trained DNN (Deep Neural Network) to calculate a first speaker feature representing the speaker identity of the speech data from the acoustic features. The aerosol quantity estimation device 100 calculates the similarity between the second speaker feature, which is the speaker feature of the speaker when they are at rest, and the first speaker feature. The aerosol quantity estimation device 100 estimates the amount of aerosol generated by the speaker according to the similarity.

[0056] In other words, the aerosol volume estimation device 100 calculates a first speaker feature using a trained DNN (Deep Neural Network) to identify the speaker when the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level, and estimates the aerosol volume according to the similarity between this first speaker feature and the second speaker feature when the speaker is at rest. The aerosol volume estimation method can accurately estimate the amount of aerosol generated by the speaker by calculating the similarity, which is obtained by utilizing the correlation between the similarity between the speaker's speechiness when the speaker is uttering and the speaker's speechiness when the speaker is at rest.

[0057] The speaker feature calculation unit 104 is not limited to a configuration comprising a frame connection processing unit 201 and a DNN 202. The speaker feature calculation unit 104 calculates speech physical quantities from the speech signal of the utterance. In this embodiment, the speaker feature calculation unit 104 calculates MFCC (Mel-Frequency Cepstrum Coefficients), which are speech feature quantities, from the speech signal of the utterance. MFCC is a feature quantity that represents the vocal tract characteristics of the speaker. The speaker feature calculation unit 104 is not limited to calculating MFCC as the speech physical quantity of the utterance; it may also calculate the speech signal after applying a Mel filter bank, or it may calculate the spectrogram of the speech signal. Furthermore, the speaker feature calculation unit 104 may use a DNN (Deep Neural Network) to calculate speech feature quantities as the speech physical quantities of the utterance from the speech signal.

[0058] The aerosol volume estimation device according to the embodiments of this disclosure has been described above, but this disclosure is not limited to these embodiments.

[0059] Furthermore, each processing unit included in the aerosol volume estimation device according to the above embodiment is typically implemented as an LSI (Large-Scale Integrated Circuit). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0060] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.

[0061] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0062] Furthermore, this disclosure may be implemented as an aerosol volume estimation method, etc., performed by an aerosol volume estimation device, etc.

[0063] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0064] Furthermore, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0065] The above describes aerosol volume estimation and the like in one or more embodiments based on embodiments, but this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art can conceive of are applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments. [Industrial applicability]

[0066] This disclosure is useful as an aerosol volume estimation method, aerosol volume estimation device, and program that can accurately estimate the amount of aerosol generated by a speaker during speech. [Explanation of Symbols]

[0067] 100 Aerosol volume estimation device 101 Voice acquisition unit 102 Sound pressure level determination unit 103 Acoustic Feature Calculation Unit 104 Speaker Feature Calculation Unit 105 Storage section 106 Similarity calculation unit 107 Aerosol volume estimation unit 108 Output section 201 Frame connection processing unit 202 DNN 203 Training Data

Claims

1. A method for estimating aerosol volume performed by a computer, Determine whether the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level. If the sound pressure level is greater than the predetermined sound pressure level, acoustic features are calculated from the speech data of the speaker's utterance. Using the trained model, a first speaker feature representing the speaker identity of the speech data is calculated from the acoustic features. The similarity between the second speaker feature, which is the speaker feature of the speaker in a calm state, and the first speaker feature is calculated. The amount of aerosol generated by the speaker is estimated according to the similarity. Method for estimating aerosol volume.

2. In the above estimation, the amount of aerosol generated increases as the similarity decreases, and the amount of aerosol generated is estimated according to the similarity. The method for estimating the amount of aerosol according to claim 1.

3. In the estimation described above, the amount of aerosol generated in a predetermined time unit is estimated for each predetermined time unit, and the cumulative value of the aerosol amount obtained from the start of the estimation is calculated. A method for estimating the amount of aerosol according to claim 1 or 2.

4. moreover, Determine whether the cumulative value is greater than a predetermined aerosol amount. If the cumulative value is greater than the predetermined aerosol amount, a warning is issued. The method for estimating the amount of aerosol according to claim 3.

5. moreover, Determine whether the cumulative value is greater than a predetermined aerosol amount. If the cumulative value is greater than the predetermined aerosol amount, the ventilation device or air purifier located in the space where the speaker is present is activated. The method for estimating the amount of aerosol according to claim 3.

6. The second speaker feature represents the speaker identity of the speech data obtained by the speaker reading a predetermined text aloud. A method for estimating the amount of aerosol according to claim 1 or 2.

7. A sound pressure level determination unit that determines whether the sound pressure level of the speaker's utterance is greater than a predetermined sound pressure level, When the sound pressure level is greater than the predetermined sound pressure level, an acoustic feature calculation unit calculates acoustic features from the speech data of the speaker's utterance, A speaker feature calculation unit calculates a first speaker feature representing the speaker identity of the speech data from the acoustic features using a trained model, A similarity calculation unit calculates the similarity between a second speaker feature, which is the speaker feature of the speaker in a calm state, and the first speaker feature. The system includes an estimation unit that estimates the amount of aerosol generated from the speaker according to the similarity. Aerosol volume estimation device.

8. The aerosol volume estimation method described in claim 1 is to be executed by a computer. program.

Citation Information

Patent Citations

  • Cough detector, and method and program for detecting cough

    JP2021003181A

  • Volume-sensing droplet infection prevention alarm device

    JP3230254U