Voice providing device, voice providing method and program

The voice providing system addresses the lack of growth process understanding in existing devices by recognizing and recording voice data to chronologically display changes in a child's speech and voice quality, facilitating the tracking of developmental milestones.

JP2025145158APending Publication Date: 2025-10-03CASIO COMPUTER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024045195
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing devices for measuring the tendency to use baby words do not store input speech sounds or provide a digest of the changes in voice quality, preventing an understanding of a child's growth process.

Method used

A voice providing system comprising a recording device and a terminal device that performs voice recognition, detects general terms and misspellings, records them with voice data, and outputs voice data based on predetermined conditions to chronologically display the child's growth process.

Benefits of technology

Enables understanding of a child's growth process through the analysis of spoken words and voice quality changes over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025145158000001_ABST
    Figure 2025145158000001_ABST
Patent Text Reader

Abstract

To grasp a child's growing process from change of words and vocal quality which the child has uttered.SOLUTION: A CPU 21 of a terminal device 20: acquires voice data corresponding to adult language or a growth step term which satisfies an output condition from an utterance voice database 233, on the basis of a predetermined output condition; and outputs voice when the adult language or the growth step term which satisfies the output condition has been uttered by a predetermined speaker from a sound output unit 26 in time series, on the basis of the voice data.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a voice providing device, a voice providing method, and a program. [Background technology]

[0002] A device for measuring the tendency to use baby words has been disclosed that detects characteristics specific to baby words (for example, "woof woof" for dog and "boo-boo" for car) from input speech and calculates a score representing the tendency to use baby words based on the characteristics specific to the detected baby words (see Patent Document 1 below). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-129849 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the device disclosed in Patent Document 1 only measures the tendency to use baby words, and does not store input speech sounds or provide a digest of the changes in the voice when the baby words are spoken. For this reason, the device cannot grasp the child's growth process from the changes in the words and voice quality spoken by the child.

[0005] The present invention has been made in view of the above problems, and has as its object to make it possible to understand a child's growth process from the words spoken by the child and the changes in voice quality. [Means for solving the problem]

[0006] In order to solve the above problem, the voice providing device of the present invention is characterized by comprising: recognition means for performing voice recognition for a predetermined speaker based on voice data; detection means for comparing a term obtained based on the recognition result by the recognition means with predetermined general terms and with misspellings terms linked to the general terms, and detecting general terms or misspellings that match the term; recording control means for linking the general terms or misspellings detected by the detection means with voice data relating to a voice in which the general terms or misspellings are spoken by the speaker and the timing at which the voice related to the voice data was spoken, and recording them in a predetermined database; and output control means for obtaining from the database the voice data corresponding to the general terms or misspellings that satisfy predetermined output conditions, based on the output conditions, and causing an output unit to chronologically output the voice in which the general terms or misspellings are spoken by the speaker based on the voice data. [Effects of the Invention]

[0007] According to the present invention, it is possible to understand a child's growth process from the words spoken by the child and the changes in voice quality. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram showing a voice providing system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing the functional configuration of the recording device. [Figure 3] FIG. 2 is a block diagram showing the functional configuration of the terminal device. [Figure 4] FIG. 10 is a diagram showing an example of the contents of a growth process dictionary. [Figure 5] FIG. 2 is a diagram showing an example of the contents of a speech voice database. [Figure 6] FIG. 10 is a diagram illustrating a control procedure for developmental stage term detection processing. [Figure 7] FIG. 10 is a diagram illustrating a control procedure for a digest generation process. [Figure 8]FIG. 10 is a diagram illustrating an example of a digest generation reception screen. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0010] As shown in Fig. 1, the voice provision system 1 includes a recording device 10 and a terminal device 20. The voice provision system 1 is a system in which the terminal device 20 stores voice data relating to a child's voice recorded by the recording device 10, generates a voice digest based on the voice data, allowing the user to look back on the child's growth process, and provides the digest to a user (for example, the child's parent).

[0011] The recording device 10 is a small recording device that is removably attached to, for example, a stuffed toy that a child usually plays with. The recording device 10 may also be removably attached to a child's clothing, a baby bed, or the like.

[0012] The terminal device 20 is, for example, a terminal device that can be connected wirelessly or via a wired connection to the recording device 10. In this embodiment, the terminal device 20 is a notebook PC (Personal Computer), but is not limited to this and may be, for example, a desktop PC, a tablet PC, a smartphone, a television, etc.

[0013] 2, the recording device 10 includes a CPU (Central Processing Unit) 11, a RAM (Random Access Memory) 12, a storage unit 13, an operation unit 14, a sound input unit 15, a sensor unit 16, and a communication unit 17. The components of the recording device 10 are connected via a bus 18.

[0014] The CPU 11 is a processor that reads and executes a program 131 stored in the storage unit 13 and performs various arithmetic operations to control the operation of each unit of the recording device 10. The CPU 11 includes a clock circuit (not shown) and acquires the current date and time measured by the clock circuit. The RAM 12 provides the CPU 11 with a working memory space and stores temporary data. The storage unit 13 stores the program 131 executed by the CPU 11 and various data. The program 131 is stored in the storage unit 13 in the form of computer-readable program code. Examples of data stored in the storage unit 13 include recording information 132 that stores audio data relating to a child's speech in association with the date and time of the speech (recording date and time). Note that the audio data stored in the recording information 132 may be audio data containing only the audio of a speech section uttered by a target speaker (e.g., the user's child). A method for generating this audio data includes first detecting an utterance section in which speech is occurring based on audio data sequentially acquired via the sound input unit 15. Then, it is determined whether the speaker in the detected speech section is the target speaker (for example, the user's child). If it is determined that the speaker is the target speaker as a result of this determination, the voice data of the corresponding speech section is stored in the recording information 132. In such cases, since it is necessary to recognize the target speaker, it is assumed that the voice of the target speaker and identification information for identifying the target speaker are stored in advance in the storage unit 13.

[0015] The operation unit 14 includes a power button (not shown) for switching the power on / off, a start / stop button (not shown) for instructing the start / stop of recording, etc., and the CPU 11 controls each unit based on instructions from the operation unit 14. The sound input unit 15, not shown, includes a microphone and an A / D converter for converting external sound captured by the microphone into digital data. The sensor unit 16 includes an acceleration sensor, an angular velocity sensor, a GPS receiver, etc., and outputs measurement results to the CPU 11. Note that the sensor unit 16 may further include sensors other than those described above. The communication unit 17 is, for example, a communication unit that employs a wireless standard such as Bluetooth (registered trademark), or a wired communication unit such as a USB terminal.

[0016] 3, the terminal device 20 includes a CPU 21, a RAM 22, a storage unit 23, a display unit 24, an operation unit 25, a sound output unit 26, and a communication unit 27. The components of the terminal device 20 are connected via a bus 28.

[0017] The CPU 21 (recognition means, detection means, recording control means, acceptance means, output control means) is a processor that reads and executes a program 231 stored in the storage unit 23 and performs various arithmetic processing, thereby controlling the operation of each unit of the terminal device 20. Note that while a single CPU 21 is illustrated in FIG. 3, this is not limiting. Two or more processors such as CPUs may be provided, and the processing executed by the CPU 21 of this embodiment may be shared and executed by these two or more processors. The RAM 22 provides a working memory space for the CPU 21 and stores temporary data. The storage unit 23 stores the program 231 executed by the CPU 21, various data, and the like. The program 231 is stored in the storage unit 23 in the form of computer-readable program code. Examples of data stored in the storage unit 23 include a growth process dictionary 232 and a speech voice database 233.

[0018] As shown in FIG. 4, the developmental stage dictionary 232 is a dictionary in which, for each adult word, the adult word is associated with a developmental stage term corresponding to the adult word. Here, an adult word refers to a word commonly used by adults (general term). A developmental stage term refers to a word used by children as they grow up. The developmental stage terms include misspellings that are used by children in their final developmental stage, and the adult word itself. Specifically, in the developmental stage dictionary 232, the adult word "fire engine" is associated with the developmental stage terms "boobu (misspelling term)," "chibosha (misspelling term)," "choubousha (misspelling term)," and "shobosha (adult word)." For this reason, hereinafter, adult words may also be referred to as developmental stage terms. In addition, the developmental stage dictionary 232 sets the average age and standard deviation of children who speak the developmental stage term, as well as the occurrence probability, for each developmental stage term. The occurrence probability may be the occurrence probability for each age (or age in months) of the child who utters the corresponding developmental stage term. The occurrence probability may be generated arbitrarily or may be generated statistically from actual utterances. The growth process dictionary 232 may also be configured to allow unregistered developmental stage terms to be additionally registered based on a user operation.

[0019] The spoken voice database 233 is a database for storing voices (voice data) of developmental stage terms actually spoken by a subject (e.g., a child of the user) for whom a voice digest is to be generated. As shown in FIG. 5, the spoken voice database 233 stores information on the following items for each developmental stage term actually spoken by a subject for whom a voice digest is to be generated: "utterance date and time," "adult words," "voice data," and "occurrence probability," linked to each other. "Utterance date and time" indicates the date on which the corresponding developmental stage term was uttered. Note that "utterance date and time" may indicate the hour, minute, and second in addition to the date. "Adult words" indicate words that a child will eventually use in relation to the corresponding developmental stage term as they grow up. "Voice data" indicates voice data related to the voice when the corresponding developmental stage term was uttered. "Occurrence probability" indicates the probability of the corresponding developmental stage term appearing. Specifically, in the spoken speech database 233, the developmental stage term "boob" actually uttered by the subject for whom the audio digest is to be generated is associated with and stored in the "utterance date and time" field of "2017 / 2 / 11," the "adult words" field of "fire engine / car / police car," the "audio data" field of "SoundFile01," and the "occurrence probability" field of "0.8." As shown in FIG. 4, the adult words "fire engine," "car," and "police car" all contain the developmental stage term "boob." Therefore, the "adult words" field lists "fire engine," "car," and "police car," each of which contains the developmental stage term "boob." Note that the spoken speech database 233 may further associate and store information in the "likelihood of speech recognition result" or "reliability of speech recognition result" for each developmental stage term actually uttered by the subject for whom the audio digest is to be generated. The "reliability of the speech recognition result" is an index value obtained by multiplying the information in the above-mentioned "probability of occurrence" item by the likelihood of the speech recognition result.

[0020] Returning to FIG. 3 , the display unit 24 is composed of an LCD (Liquid Crystal Display), an EL (Electro Luminescence) display, or the like, and displays various types of information in accordance with display information instructed by the CPU 21. The operation unit 25 has at least one of a touch panel overlaid on the display screen of the display unit 24, physical buttons, a pointing device such as a mouse, and an input device such as a keyboard, and outputs operation information to the CPU 21 in accordance with input operations on the input device. The sound output unit 26 is composed of a DA converter, an amplifier, a speaker, and the like. When outputting sound, the sound output unit 26 converts sound data into analog audio data and outputs it from the speaker. The communication unit 27 is, for example, a communication unit that employs a wireless standard such as Bluetooth (registered trademark) or a wired communication unit such as a USB terminal.

[0021] Next, a description will be given of the operation of the speech providing system 1. Specifically, a developmental stage terminology detection process (see FIG. 6) and a digest generation process (see FIG. 7) executed by the terminal device 20 of the speech providing system 1 will be described.

[0022] First, the developmental stage term detection process will be described. This developmental stage term detection process is triggered when the terminal device 20 acquires (receives) recorded information 132, which has been recorded by the recording device 10 and stored in the storage unit 13, from the recording device 10 via the communication unit 27. Here, the terminal device 20 acquires the recorded information 132 from the recording device 10 at a predetermined time every day, for example. The timing at which the terminal device 20 acquires the recorded information 132 is not particularly limited, and may be, for example, every hour, every half day, every week, every month, or every year. Alternatively, the recorded information 132 may be acquired from the recording device 10 when recording begins and ends with the recording device 10.

[0023] As shown in Fig. 6, when the developmental stage term detection process is started, first, the CPU 21 of the terminal device 20 acquires one piece of speech data from the recording information 132 acquired from the recording device 10 (step S1). Next, the CPU 21 uses a known speaker recognition technique to extract speech from the speech data acquired in step S1, the speech of a target speaker (e.g., the user's child) (step S2). Next, the CPU 21 performs speech recognition on the speech extracted in step S2 (step S3). Here, the speech recognition performed in step S3 is preferably a method capable of recognizing the speech of a growing child (e.g., OpenAI's voice recognition model whisper, etc.).

[0024] Next, the CPU 21 compares the result of the speech recognition in step S3 with each developmental stage term in the growth process dictionary 232 (see FIG. 4) (step S4). Here, the result of the speech recognition means a term (speech recognition term) acquired based on the result of the speech recognition. Furthermore, as defined in paragraph 0018, the developmental stage term includes adult words that a child will eventually use in the process of growing up (e.g., "fire engine," "car," "Patrol car," etc.; see FIG. 4). Next, the CPU 21 determines whether there is a developmental stage term that matches the result of the speech recognition in step S3 (step S5). If it is determined in step S5 that there is a developmental stage term that matches the result of the speech recognition in step S3 (step S5; YES), the CPU 21 stores the spoken voice of the developmental stage term that matches the speech recognition result in the spoken voice database 233 (step S6). Specifically, when the developmental stage term matching the result of the speech recognition in step S3 is "chibosha," as shown in FIG. 5, CPU 21 associates the developmental stage term "chibosha" with information in the "utterance date and time" field of "2018 / 1 / 15," information in the "adult language" field of "fire engine," information in the "audio data" field of "SoundFile02," and information in the "occurrence probability" field of "0.9," and stores the associated information in the speech database 233. Then, CPU 21 proceeds to step S7. Note that in step S6, the likelihood of the speech recognition result in step S3 may be further associated and stored in the speech database 233. Furthermore, when storing the speech of the developmental stage term matching the result of the speech recognition in step S3 in the speech database 233, if there are multiple speeches with the same developmental stage term, it is possible to store only one of the speeches rather than storing all of the speeches (audio data). At this time, the one uttered speech to be saved may be determined taking into consideration the occurrence probability. Also, the one uttered speech to be saved may be determined taking into consideration the likelihood of the result of the speech recognition in step S3. Also, the one uttered speech to be saved may be determined taking into consideration the reliability of the result of the speech recognition obtained by multiplying the occurrence probability by the likelihood of the result of the speech recognition in step S3. In this case, the speech can be saved with high accuracy according to the growth process of the target speaker.

[0025] If it is determined in step S5 that there is no developmental stage term that matches the result of the speech recognition in step S3 (step S5; NO), CPU 21 skips step S6 and proceeds to step S7. Next, CPU 21 determines whether or not the series of processes (detection processes) from step S2 to step S6 have been completed for all the speech data acquired from recording device 10 when starting the developmental stage term detection process (step S7).

[0026] If it is determined in step S7 that the series of processes (detection processes) from step S2 to step S6 have been completed for all of the audio data acquired from the recording device 10 (step S7; YES), the CPU 21 ends the developmental stage term detection process. If it is determined in step S7 that the series of processes (detection processes) from step S2 to step S6 have not been completed for all of the audio data acquired from the recording device 10 (step S7; NO), the CPU 21 acquires unprocessed audio data from the recording information 132 acquired from the recording device 10 (step S8). Then, the CPU 21 returns the process to step S2 and repeats the processes from step S2 onwards. Note that if the processes from step S2 onwards have been repeated,

[0027] Next, the digest generation process will be described. This digest generation process is executed when a digest generation request is made by the user via the operation unit 25.

[0028] As shown in FIG. 7, when the digest generation process is started, first, the CPU 21 of the terminal device 20 displays a digest generation reception screen G1 on the display unit 24 (step S11).

[0029] As shown in FIG. 8, the digest generation reception screen G1 displays a message G11 prompting the user to select a digest to be generated, and provides radio buttons G12 for selecting the digest to be generated. The digest generation reception screen G1 also displays a message G13 prompting the user to select words to be output (audio output) as a digest, and provides a pull-down menu G14 for selecting the words to be output. The words (adult words) displayed in the pull-down menu G14 are words (adult words) whose audio data is stored in the speech database 233. For example, if the audio data of the adult word "patrol car" is not stored in the speech database 233, "patrol car" will not be displayed in the pull-down menu G14. The digest generation reception screen G1 also includes a start generation button G15 used to start digest generation and an end button G16 used to close the digest generation reception screen G1.

[0030] Returning to FIG. 7, next, the CPU 21 determines whether or not the generation start button G15 (see FIG. 8) has been operated by the operation unit 25 (step S12). If it is determined in step S12 that the generation start button G15 has been operated (step S12; YES), the CPU 21 determines whether or not the target for generating a digest is a "language growth process digest" (step S13). Here, the determination of whether or not the target for generating a digest is a "language growth process digest" is made based on the selection state of the radio button G12 (see FIG. 8). Specifically, if the "language growth process digest" is selected by the radio button G12, it is determined that the target for generating a digest is a "language growth process digest." On the other hand, if the "voice quality growth process digest" is selected by the radio button G12, it is determined that the target for generating a digest is not a "language growth process digest."

[0031] If it is determined in step S13 that the digest to be generated is a "language growth process digest" (step S13; YES), the CPU 21 acquires from the speech database 233 (see FIGS. 4 and 5) speech data of developmental stage terms linked to the adult word selected in the pull-down menu G14 (see FIG. 8) (step S14). Specifically, if "fire engine" is selected in the pull-down menu G14, speech data of "boob," "chibosha," "choubousha," and "shobosha," which are linked to "fire engine," is acquired from the speech database 233. Note that, for example, if multiple pieces of speech data for "shobosha" are stored in the speech database 233, one piece of speech data may be selected and acquired instead of acquiring all of the speech data. Furthermore, when selecting and acquiring one piece of speech data, the one piece of speech data to be acquired may be selected taking into consideration its occurrence probability.

[0032] Next, based on the voice data acquired in step S14, CPU 21 outputs the voices of the developmental stage terms from sound output unit 26 in chronological order of the utterance date and time (step S15). Specifically, when voice data "SoundFile01" for "boob," voice data "SoundFile02" for "chibosha," voice data "SoundFile03" for "chobosha," and voice data "SoundFile04" for "fire engine" are acquired from speech voice database 233 in step S14, the voices are output from sound output unit 26 in the order of "boob" ⇒ "chibosha" ⇒ "chobosha" ⇒ "fire engine" based on the utterance date and time linked to each voice data. When each voice is output from sound output unit 26, the date and time of each utterance may be displayed on display unit 24, or the age of the speaker (for example, the user's child) at the time of the utterance may be displayed on display unit 24. When the speaker's age is displayed, it is assumed that the speaker's date of birth is recorded in the storage unit 13 in advance.

[0033] Next, the CPU 21 determines whether or not the end button G16 (see FIG. 8) has been operated by the operation unit 25 (step S18). If it is determined in step S18 that the end button G16 has been operated (step S18; YES), the CPU 21 ends the digest generation process. If it is determined in step S18 that the end button G16 has not been operated (step S18; NO), the CPU 21 returns the process to step S12 and performs the subsequent processes.

[0034] Furthermore, in step S13, if it is determined that the target for generating a digest is not a "language growth process digest" (step S13; NO), that is, if it is determined that the target for generating a digest is a "voice quality growth process digest," the CPU 21 acquires voice data of the adult word (the developmental term finally used in the growth process) selected in the pull-down menu G14 (see FIG. 8) from the speech voice database 233 (see FIGS. 4 and 5) (step S16). Specifically, if "fire engine" is selected in the pull-down menu G14, the CPU 21 acquires voice data of "fire engine (the developmental term finally used in the growth process)" from the speech voice database 233.

[0035] Next, based on the voice data acquired in step S16, the CPU 21 outputs the voices of adult language (developmental terms that are finally used in the growth process) from the sound output unit 26 in chronological order of the utterance date and time (step S17). Specifically, if multiple voice data of "fire truck" are acquired in step S16, the voice data are grouped by the year in which the voice was uttered (e.g., 2022, 2023, 2024) based on the utterance date and time associated with each voice data, and one voice data is selected from each group. Then, each utterance of "fire truck" is output from the sound output unit 26 in chronological order of the year in which the voice was uttered. When each utterance of voice is output from the sound output unit 26, the display unit 24 may display the utterance date and time of each utterance, or the age of the speaker (e.g., the user's child) at the time of the utterance may be displayed on the display unit 24. When the speaker's age is displayed, the speaker's date of birth is assumed to be recorded in advance in the storage unit 13. After executing the process of step S17, the CPU 21 advances the process to step S18 and performs the subsequent processes. Also, if it is determined in step S12 that the generation start button G15 has not been operated (step S12; NO), the CPU 21 advances the process to step S18 and performs the subsequent processes.

[0036] As described above, the CPU 21 of the terminal device 20 performs speech recognition for a predetermined speaker based on the speech data (recorded information 132). The CPU 21 compares terms acquired based on the speech recognition results with predetermined adult words (common terms) and with developmental stage terms (mispronunciation terms) associated with the adult words, and detects adult words or developmental stage terms that match the terms. The CPU 21 associates the detected adult words or developmental stage terms with speech data relating to speech in which the adult words or developmental stage terms are spoken by the speaker and the timing of the speech associated with the speech data, and records the associated data in the speech database 233. The CPU 21 acquires speech data corresponding to adult words or developmental stage terms that satisfy predetermined output conditions from the speech database 233, and controls the sound output unit 26 to output, in time series, speech in which the adult words or developmental stage terms that satisfy the output conditions are spoken by the speaker based on the speech data. Therefore, according to the terminal device 20, based on predetermined output conditions, it is possible to provide a digest of the changes in voice when developmental stage terms including adult words are spoken, so that it becomes possible to understand the growth process of the speaker (for example, the user's child) from the changes in the words and voice quality spoken by the speaker.

[0037] Furthermore, the CPU 21 accepts, based on a user operation, designation of the output content of the speech related to adult words or developmental terms (mispronunciation terms) recorded in the speech database 233. When the designation of the output content of the speech related to adult words or developmental terms is accepted, the CPU 21 acquires, as the output condition described above, speech data corresponding to the adult words or developmental terms that satisfy the output condition from the speech database 233, and based on the speech data, causes the sound output unit 26 to output in chronological order the speech of the adult words or developmental terms that satisfy the output condition spoken by the speaker. Therefore, the terminal device 20 can provide a digest of the transition of speech when developmental terms including adult words are spoken, triggered by the acceptance of the designation of the output content of the speech related to adult words or developmental terms, thereby enabling the user to receive the digest at a timing desired by the user.

[0038] Furthermore, the CPU 21 accepts a designation to output, in chronological order, the speech of a desired adult word and a developmental stage term (a misspelling term) related to the desired adult word spoken by the speaker. When the designation is accepted, the CPU 21 acquires speech data corresponding to the desired adult word and the developmental stage term related to the desired adult word from the speech speech database 233, and based on the speech data, causes the sound output unit 26 to output, in chronological order, the speech of the desired adult word and the developmental stage term related to the desired adult word spoken by the speaker. Therefore, the terminal device 20 can provide a digest of the changes in speech (e.g., "boo-bu" ⇒ "chibosha" ⇒ "choubousha" ⇒ "shoubousha") when the desired adult word and the developmental stage term related to the desired adult word are spoken, thereby enabling the growth process of the speaker (e.g., the user's child) to be understood from the changes in words spoken by the speaker.

[0039] Furthermore, the CPU 21 accepts a designation to output the voice of the desired adult word spoken by the speaker in chronological order. When the designation is accepted, the CPU 21 acquires voice data corresponding to the desired adult word from the speech voice database 233, and based on the voice data, causes the voice of the desired adult word spoken by the speaker to be output in chronological order from the sound output unit 26. Therefore, the terminal device 20 can provide a digest of the voice transition when the adult word is spoken (for example, "fire engine (2022)" ⇒ "fire engine (2023)" ⇒ "fire engine (2024)"), making it possible to understand the growth process of the speaker (for example, the user's child) from the transition of the voice quality spoken by the speaker.

[0040] Furthermore, when a designation to output in chronological order the voice of a desired adult word spoken by the speaker is accepted, the CPU 21 classifies the voice data corresponding to the desired adult word acquired from the speech voice database 233 by the year in which the voice related to the voice data was spoken, selects one voice data from each classified group, and outputs in chronological order the voice based on the voice data selected in each group from the sound output unit 26. Therefore, the terminal device 20 can appropriately generate and output a digest that makes it easy to understand the development process of voice quality.

[0041] Although the present invention has been specifically described above based on the embodiments, the present invention is not limited to the above embodiments and can be modified within the scope of the invention. For example, in the above embodiment, the terminal device 20 executes the developmental stage term detection process (see FIG. 6) and the digest generation process (see FIG. 7), but the recording device 10 may execute these processes.

[0042] In the above embodiment, in step S15 of the digest generation process (see FIG. 7), the speech of the developmental stage terms is output from sound output unit 26 in chronological order of the date and time of utterance based on the speech data acquired in step S14, i.e., a word growth process digest is output from sound output unit 26. However, the word growth process digest may be transferred from terminal device 20 to an external device via communication unit 27 and output from the sound output unit of the external device. Similarly, in step S17 of the digest generation process (see FIG. 7), the speech of adult words (developmental stage terms ultimately used in the growth process) is output from sound output unit 26 in chronological order of the date and time of utterance based on the speech data acquired in step S16, i.e., a voice quality growth process digest is output from sound output unit 26. However, the voice quality growth process digest may be transferred from terminal device 20 to an external device via communication unit 27 and output from the sound output unit of the external device.

[0043] Furthermore, in the above embodiment, when a specification of an adult word that a child will eventually use as he or she grows up is accepted as a desired developmental term from a group of developmental terms recorded in the speech database 233, and a specification to output the desired adult word and a developmental term used before the desired adult word are used, spoken by the speaker, in chronological order (a language developmental process digest), is accepted, the desired adult word and the developmental term used before the desired adult word are used are output in chronological order from the sound output unit 26 based on the speech data relating to the speech of the desired adult word and the developmental term used before the desired adult word are used, spoken by the speaker. However, for example, in step S6 of the developmental term detection process (see FIG. 6 ), when a speech of an adult word that matches the speech recognition result is stored in the speech database 233, the speech of the adult word and the developmental term used before the adult word are used may be output in chronological order from the sound output unit 26 based on the speech data relating to the speech of the adult word and the developmental term used before the adult word are used, In addition, in step S5 of the developmental term detection process (see FIG. 6), when it is determined that there is a developmental term that matches the speech recognition result, the developmental term is an adult word that a child will eventually use in the process of growing up, and the likelihood of the speech recognition result is equal to or greater than a threshold, triggering the audio output unit 26 to output the adult word and a developmental term used before the adult word were used in time series based on audio data related to audio uttered by the speaker. In addition, in step S5 of the developmental term detection process (see FIG. 6), when it is determined that there is a developmental term that matches the speech recognition result, the developmental term is an adult word that a child will eventually use in the process of growing up, and the likelihood of the speech recognition result is equal to or greater than a threshold, triggering the audio output unit 26 to display the adult word in the pull-down menu G14 (see FIG. 8).Furthermore, in step S6 of the developmental stage term detection process (see FIG. 6), when a predetermined number or more of speech sounds of developmental stage terms (including adult words) linked to the same adult word are stored in the speech sound database 233, the adult words and the developmental stage terms used before the adult word are used may be output in chronological order from the sound output unit 26 based on speech data related to speech uttered by the speaker. Furthermore, in step S6 of the developmental stage term detection process (see FIG. 6), when a predetermined number or more of speech sounds of developmental stage terms (including adult words) linked to the same adult word are stored in the speech sound database 233, the adult words may be displayed in the pull-down menu G14 (see FIG. 8).

[0044] Furthermore, in the above embodiment, in step S6 of the developmental stage term detection process (see Figure 6), when the speech of the developmental stage term that matches the result of the speech recognition in step S3 is stored in the speech database 233, if there are multiple speeches with the same developmental stage term, not all of the speeches (speech data) are stored, but one of the speeches is stored.However, for example, one speech may be selected and stored taking into account the likelihood of the result of the speech recognition in step S3 and the probability of occurrence according to the age of the target speaker. [Explanation of symbols]

[0045] 1 Voice providing system, 10 Recording device, 132 Recording information, 20 Terminal device, 21 CPU, 232 Growth process dictionary, 233 Speech voice database, 24 Display unit, 26 Sound output unit

Claims

1. recognition means for performing speech recognition for a predetermined speaker based on the speech data; a detection means for comparing a term acquired based on the recognition result by the recognition means with a predetermined general term and with a misspelling term associated with the general term, and detecting a general term or a misspelling term that matches the term; a recording control means for linking the general term or the misspelling term detected by the detection means, audio data relating to the audio in which the general term or the misspelling term is spoken by the speaker, and the timing at which the audio relating to the audio data was spoken, and recording the linked data in a predetermined database; an output control means for acquiring, based on a predetermined output condition, from the database, the speech data corresponding to the general term or the misspelling term that satisfies the output condition, and for outputting, based on the speech data, a speech in which the general term or the misspelling term that satisfies the output condition, is spoken by the speaker, in time series from an output unit; A sound providing device comprising:

2. a receiving means for receiving, based on a user operation, a designation of speech output content related to the general term or the misspoken term recorded in the database; The output control means the receiving means receives a designation of the output content of the voice related to the general term or the misspoken term as an output condition, the voice data corresponding to the general term or the misspoken term that satisfies the output condition is obtained from the database, and the voice of the general term or the misspoken term that satisfies the output condition is output in chronological order from an output unit based on the voice data.

2. The audio providing device according to claim 1.

3. The output control means the detection means detects the general term that matches the term, and the likelihood of the recognition result for the term is equal to or greater than a threshold, the speech data corresponding to the general term and the misspellings associated with the general term that satisfy the output condition is obtained from the database, and an output unit outputs in time series the speech of the general term and the misspellings associated with the general term that satisfy the output condition, spoken by the speaker, based on the speech data; 2. The audio providing device according to claim 1.

4. The receiving means Accepting a designation to output in time series the speech of the desired general term and the misspelling term related to the desired general term uttered by the speaker; The output control means When the specification is accepted by the accepting means, the speech data corresponding to the desired general term and the misspellings related to the desired general term are acquired from the database, and based on the speech data, speech of the desired general term and the misspellings related to the desired general term spoken by the speaker is output in time series from an output unit.

3. The audio providing device according to claim 2.

5. The receiving means Accepting a designation to output a time-series audio of the desired general term uttered by the speaker; The output control means When the specification is accepted by the accepting means, the voice data corresponding to the desired general term is acquired from the database, and based on the voice data, the voice of the desired general term spoken by the speaker is output in time series from an output unit.

3. The audio providing device according to claim 2.

6. The output control means When the specification is accepted by the accepting means, the voice data corresponding to the desired general term acquired from the database is classified by the year in which the voice related to the voice data was spoken, one voice data is selected from each classified group, and the voice based on the voice data selected in each group is output in chronological order from the output unit.

6. The audio providing device according to claim 5.

7. The recording control means When the detection means detects a general term that matches the term, when recording the general term in the database, the age of the speaker at the time the general term was uttered is further linked and recorded; The output control means When the specification is accepted by the accepting means, the voice data corresponding to the desired general term acquired from the database is classified according to the age of the speaker associated with the voice data, one voice data is selected from each classified group, and voices based on the voice data selected in each group are output in chronological order from the output unit.

6. The audio providing device according to claim 5.

8. A voice providing method executed by a computer of a voice providing device, a recognition step of performing speech recognition for a predetermined speaker based on the speech data; a detection step of comparing a term obtained based on the recognition result of the recognition step with predetermined general terms and with misspellings associated with the general terms, and detecting general terms or misspellings that match the term; a recording control step of linking the general term or the misspelling term detected by the detection step, audio data relating to the audio in which the general term or the misspelling term is spoken by the speaker, and the timing at which the audio relating to the audio data was spoken, and recording them in a predetermined database; an output control step of retrieving, based on predetermined output conditions, the voice data corresponding to the general term or the misspelling term that satisfies the output conditions from the database, and outputting, based on the voice data, the voice of the general term or the misspelling term that satisfies the output conditions, spoken by the speaker, from an output unit in time series; A method for providing audio, comprising:

9. The computer of the audio providing device, recognition means for performing speech recognition for a predetermined speaker based on the speech data; a detection means for comparing a term acquired based on the recognition result by the recognition means with predetermined general terms and with misspellings associated with the general terms, and detecting general terms or misspellings that match the term; a recording control means for linking the general term or the misspelling term detected by the detection means, audio data relating to the audio in which the general term or the misspelling term is spoken by the speaker, and the timing at which the audio relating to the audio data was spoken, and recording these in a predetermined database; an output control means for retrieving, based on predetermined output conditions, the speech data corresponding to the general term or the misspelling term that satisfies the output conditions from the database, and for outputting, based on the speech data, speech in which the general term or the misspelling term that satisfies the output conditions is spoken by the speaker from an output unit in chronological order; A program characterized by functioning as

Citation Information

Patent Citations

  • Childcare word use tendency measuring apparatus, method, and program

    JP2015129849A