Information processing device and information processing method
The information processing device addresses the challenge of accurately recognizing technical terms by identifying error patterns and correcting misreadings, enhancing voice recognition systems' performance for technical terms across different recognition types.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-07
- Publication Date
- 2026-04-16
AI Technical Summary
Conventional voice recognition technologies struggle with accurately recognizing technical terms due to the absence of mechanisms for pre-registering such terms and the issue of multiple possible pronunciations, leading to misrecognition.
An information processing device that includes an error tendency information acquisition unit to identify error patterns in voice recognition, a reading expansion information acquisition unit to map correct readings for technical terms, and a control unit to replace misreadings with correct pronunciations, utilizing a list of technical terms and their correct readings.
Enables accurate and instant recognition of technical terms in voice recognition systems by correcting multiple pronunciations, applicable to both DNN-HMM and End-to-End type voice recognition methods.
Smart Images

Figure JP2024035812_16042026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus and Information Processing Method
[0001] The present disclosure relates to an information processing apparatus and an information processing method. Here, the "technical term" widely means a term that is not a general term, and includes, for example, a technical term used in a specific field, an in-house term used within a specific company, etc.
[0002] Voice recognition technology that recognizes human voice and outputs the recognition result as a character string is known (Patent Documents 1, 2, etc. below). As conventional voice recognition technologies, there are a first type of voice recognition (for example, DNN (Deep Neural Network)-HMM (Hidden Markov Model) type voice recognition) that outputs a voice recognition result (character string) through a plurality of processing stages using a plurality of models (an acoustic model that maps a voice signal to a phoneme, a word dictionary that maps a phoneme to a word, and a language model for deriving a character string from a word), and a second type of voice recognition (End-to-End type voice recognition) that directly recognizes a character string from a voice using a single voice recognition model without going through a plurality of processing stages and outputs the recognition result (character string). Among these, the second type of voice recognition is expected to have higher knowledge-based performance and improved recognition accuracy compared to the first type of voice recognition.
[0003] Japanese Patent Laid-Open No. 8-171396, Japanese Patent Laid-Open No. 2016-102866
[0004] However, in the second type of voice recognition, since the word dictionary in the first type of voice recognition is not used, there is no means to pre-register technical terms that are not general terms, and as a result, technical terms cannot be properly recognized. Therefore, for the second type of voice recognition, a new mechanism for making technical terms the recognition target is required. On the other hand, in the first type of voice recognition, since it is necessary to register technical terms in the word dictionary, a new mechanism for making technical terms the recognition target without performing such work is eagerly awaited.
[0005] Furthermore, a common issue with the above-mentioned speech recognition methods is that various terms, including technical terms, often have multiple possible pronunciations. For example, if part of a term is the letters "SI," then "shi" and "es-i" are possible pronunciations for this part. A technology that can perform accurate speech recognition while taking into account such multiple possible pronunciations is highly anticipated.
[0006] Therefore, this project aims to instantly and accurately recognize speech containing technical terms, taking into account that multiple pronunciations are possible in various speech recognition methods.
[0007] The information processing device relating to this disclosure includes: a reading expansion information acquisition unit that acquires reading expansion information, which is information showing the correspondence between the correct reading of a technical term and the reading when the technical term is misread, based on error trend information, which is information regarding the error tendency of the correct reading in a string recognized from speech that reads the correct reading, and a list of technical terms, which includes the notation and correct reading of technical terms that are not general terms; and a control unit that replaces one of the multiple readings of a technical term with the correct reading, based on multiple readings of a technical term extracted from a string recognized from the target speech and the reading expansion information.
[0008] According to this disclosure, various speech recognition methods can instantly and accurately recognize speech containing technical terms, taking into account that multiple pronunciations are possible.
[0009] This is a diagram illustrating the configuration of a system including an information processing device. This diagram explains the function of the error tendency information acquisition unit. This diagram explains the function of the reading expansion information acquisition unit. This diagram explains the process when the reading expansion information acquisition unit acquires reading expansion information for terms that have been divided from technical terms. This diagram explains the process when acquiring reading expansion information for terms that include long vowels. This diagram explains the function of the control unit. This diagram explains the process of extracting multiple readings using LLM. This is a flowchart of the process executed by the information processing device. This is a flowchart showing the processing details of step S1 in Figure 8. This is a flowchart showing the processing details of step S2 in Figure 8. This is a flowchart showing the processing details of step S3 in Figure 8. This is a diagram showing a modified configuration of the configuration in Figure 1. This is a diagram showing an example of the hardware configuration of the information processing device.
[0010] Hereinafter, an embodiment of the information processing apparatus and information processing method relating to this disclosure will be described with reference to the drawings.
[0011] [Configuration of the system including the information processing device] Figure 1 shows a configuration diagram of system 1 including the information processing device 10 according to this disclosure. As shown in Figure 1, system 1 comprises the information processing device 10 and the speech recognition device 20.
[0012] The speech recognition device 20 is a device that recognizes human speech using speech recognition technology and outputs the recognition result as a string to the information processing device 10. The speech recognition method used by the speech recognition device 20 may be a first type of speech recognition (e.g., DNN-HMM type speech recognition) that outputs the speech recognition result (string) after going through multiple processing stages using multiple models (acoustic model that maps speech signals to phonemes, word dictionary that maps phonemes to words, and language model that derives strings from words), or it may be a second type of speech recognition (e.g., End-to-End type speech recognition) that directly recognizes a string from speech and outputs the recognition result (string) using a single speech recognition model without going through multiple processing stages.
[0013] The information processing device 10 includes an error tendency information acquisition unit 11, a read / unfold information acquisition unit 12, and a control unit 13 in order to realize the functions related to this disclosure.
[0014] Of these, the error tendency information acquisition unit 11 is a functional unit that acquires error tendency information for correct readings based on the string recognized from the audio that reads the correct reading and the correct reading. Specifically, it acquires an error tendency list as error tendency information, which associates the error tendency at the character level with the probability of occurrence of that error tendency. To perform the above function, the error tendency information acquisition unit 11 includes a string correspondence unit 11A that associates the speech recognition result (string) recognized by the speech recognition device 20 from audio data with a transcript with the transcript (correct reading), a reading extraction unit 11B that extracts the reading of the incorrect part in the speech recognition result (string) under predetermined conditions, and a tendency acquisition unit 11C that acquires an error tendency list L that represents the error tendency in speech recognition. In order to improve processing efficiency, the reading extraction unit 11B may limit the target of reading extraction to the readings of words of predetermined parts of speech (for example, parts of speech other than verbs and particles). The acquired error trend list L is stored by the trend acquisition unit 11C and referenced by the word replacement rule acquisition unit 12B, which will be described later.
[0015] To explain the function of the error tendency information acquisition unit 11 using the example in Figure 2, suppose that the speech recognition device 20 recognizes the audio data with a transcript, "Correct reading: Aramaki to Anko o Okuru tame atena o kaku," and the speech recognition result is obtained as the string "Haramaki to Wanko o Okuru tame hatena o kaku." The string correspondence unit 11A associates the speech recognition result (string) with the transcript (correct reading), and the reading extraction unit 11B extracts the error "ha" in the correct reading "a," the error "wa" in the correct reading "a," and the error "ha" in the correct reading "a," as shown by the three dashed arrows in Figure 2. By aggregating the correct / incorrect correspondence information extracted in this way, the tendency of recognition errors that the speech recognition device 20 is likely to make is obtained. The correct / incorrect correspondence information is obtained at non-consecutive syllable units or non-consecutive phoneme units, or at consecutive syllable units or consecutive phoneme units. Figure 2 shows an example where the correspondence information for correctness was obtained at non-contiguous phoneme units.
[0016] The trend acquisition unit 11C stores the correspondence information between correct and incorrect at the locations of the speech recognition errors described above, and obtains an error trend list L representing the error trend in speech recognition from the stored information. Specifically, the trend acquisition unit 11C obtains the correspondence information between correct and incorrect as described above, and calculates the probability of each error occurring based on past error history information, and stores the correspondence information between correct and incorrect and the probability of occurrence in the error trend list L. For example, an error trend list L is obtained and stored that includes information such as: - Probability that the correct reading "A" is incorrectly recognized as "Ha" 30% - Probability that the correct reading "A" is incorrectly recognized as "Wa" 10% - Probability that the correct reading "A" is incorrectly recognized as "Ka" 1% Note that Figure 2 shows the above three pieces of information as examples, but the error trend list L includes information that associates "correct reading" and "incorrectly recognized result" for all combinations of the Japanese syllabary.
[0017] Returning to Figure 1, the reading expansion information acquisition unit 12 is a functional unit that acquires reading expansion information, which is information showing the correspondence between the correct reading of a technical term and the reading when it is misread, based on the error trend information (error trend list L) acquired by the error trend information acquisition unit 11 and a technical term list that includes the notation and correct reading of technical terms that are not general terms. The reading expansion information acquisition unit 12 also has the function of acquiring reading expansion information for technical terms obtained by dividing a technical term. In order to perform the above function, the reading expansion information acquisition unit 12 includes a list acquisition unit 12A that acquires a technical term list that includes the notation and correct reading of technical terms from an external server, and a word replacement rule acquisition unit 12B that performs reading expansion on the readings of technical terms in the error trend list and acquires and stores word replacement rules R as reading expansion information.
[0018] Using the example in Figure 3, the function of the reading and expansion information acquisition unit 12 is explained as follows: The list acquisition unit 12A acquires a list of technical terms from an external server, which includes the spelling and correct reading of the technical terms to be recognized, as shown below: Technical term (set of spelling and reading) Ashiro (spelling): Ashiro (reading)
[0019] The word replacement rule acquisition unit 12B, for the reading of the technical term "Ashiiro" in the technical term list, performs a reading expansion based on the error tendency list L acquired by the error tendency information acquisition unit 11 to determine what reading is most likely to be output in the speech recognition result. When expanding the reading, it decides whether or not to expand based on the probability value of error occurrence held in the error tendency list L. At this time, the word replacement rule acquisition unit 12B assumes various strings formed by all possible combinations of errors when an error occurs at the character level, and calculates the probability of error occurrence for each assumed string by multiplying the error occurrence probability of each character that makes up the assumed string, as follows: Hashiiro (0.3 × 1 × 1 × 1 = 0.3) Haiiro (0.3 × 0.25 × 1 × 1 = 0.075) Haido (0.3 × 0.25 × 1 × 0.4 = 0.03) ... Kaniso (0.01 × 0.05 × 1 × 0.03 = 0.000015)
[0020] The word replacement rule acquisition unit 12B then adds any strings whose calculated error probability exceeds a predetermined threshold to the word replacement rule R as strings that are likely to be misrecognized for the technical term reading "ashiiro". For example, if the threshold is set to "0.1", in the example in Figure 3, the string "hashiiro", which exceeds the threshold, is added to the word replacement rule R as a string that is likely to be misrecognized for the technical term reading "ashiiro". In this case, the word replacement rule "hashiiro, washiiro, kashiiro → ashiiro" is acquired, meaning that "hashiiro, washiiro, kashiiro" are strings that are likely to be misrecognized for the technical term reading "ashiiro", and is stored as the word replacement rule R.
[0021] The following example illustrates how the reading expansion information acquisition unit 12 acquires reading expansion information for technical terms obtained by dividing them, using Figure 4. As shown in Figure 4, suppose the list acquisition unit 12A has acquired a list of technical terms that includes the spelling and correct reading of a technical term containing "Kabushiki Kaisha" (corporation), as follows: Technical term (set of spelling and reading) Kabushiki Kaisha Ashiro (spelling): Kabushiki Gaisha Ashiro (reading) In this case, the list acquisition unit 12A divides the above technical term into the following two technical terms (set of spelling and reading) in order to separate "Kabushiki Kaisha" (corporation), which is widely used in company names, from the rest of the term. Technical term (set of spelling and reading) Kabushiki Kaisha (spelling): Kabushiki Gaisha (reading) Ashiro (spelling): Ashiro (reading)
[0022] Then, the list acquisition unit 12A, for example, by referring to a word dictionary used in the DNN-HMM type model, determines that "Kabushiki Kaisha" (corporation) in the word dictionary is a general term, and determines that "Ashiro" (Ashiro) which is not in the word dictionary is a technical term. Of these, the technical term "Ashiro" is determined to be a technical term, and the technical term (set of spelling and reading) Ashiro (spelling): Ashiro (reading) is acquired as a technical term list. After that, as in the example in Figure 3, reading expansion is performed, and the word replacement rule "Hashiiro, Washiro, Kashiiro → Ashiro" is acquired and stored as word replacement rule R.
[0023] Figure 5 shows an example of obtaining word substitution rules when vowel extensions occur. In such cases, the list acquisition unit 12A converts the vowel extensions into katakana, as shown in Figure 5, such as "Aa → Aaa", "Ee → Ii", ... "Nn → Nn". Specifically, it converts the vowel extensions into katakana that represents the vowel of the preceding character. For "Kaa", the vowel extension is converted into "Kaa" by converting it into the katakana "A" which represents the vowel of the preceding character "Ka". For "Kii", the vowel extension is converted into "Kii" by converting it into the katakana "I" which represents the vowel of the preceding character "Ki". However, for "Nn", the vowel extension is converted into "Nn" by converting it into the preceding character "N". For example, if a user registers the technical term (a set of spelling and pronunciation) Ashiro (spelling): Ashiro (pronunciation), the list acquisition unit 12A converts the vowel extension "shi → shii," and this conversion transforms the registered information into the technical term (a set of spelling and pronunciation) Ashiro (spelling): Ashiro (pronunciation). Thereafter, similar to the example in Figure 3, the pronunciation expansion is performed on the pronunciation "Ashiiro" with a tendency for errors to occur in non-consecutive syllable units, and the word replacement rule "Hashiiro, Washiro, Kashiiro → Ashiro" is acquired and stored as word replacement rule R.
[0024] Returning to Figure 1, the control unit 13 is a functional unit that replaces one of the multiple readings of a technical term with the correct reading, based on the multiple readings of the technical term extracted from the string recognized from the target speech and the reading expansion information. In order to perform the above function, the control unit 13 includes a string acquisition unit 13A that acquires the speech recognition result (string) of the target speech, a reading extraction unit 13B that extracts multiple readings from the string, and a replacement unit 13C that compares the extracted multiple readings with word replacement rules, replaces one of the multiple readings that matches with the corresponding technical term (correct reading), and outputs the corrected (replaced) string.
[0025] As illustrated in the example in Figure 6, the string acquisition unit 13A acquires the speech recognition result (string) of the target speech from the speech recognition device 20 via the route shown by the dashed line in Figure 1, and the reading extraction unit 13B extracts multiple readings from the acquired speech recognition result (string). For example, as shown in Figure 6, if the speech recognition result "I visited Hashiiro" is acquired, the reading extraction unit 13B extracts multiple readings "Hashiiro wo houmon shimashita" and "Hashishoku wo houmon shimashita" from the speech recognition result using conventional morphological analysis. Alternatively, this extraction may be performed by using an external large-scale language model (LLM: Large Language Model), not shown in Figure 1, to extract the above multiple readings from the speech recognition result. When using an LLM, the prompt illustrated in Figure 7 is "Please list multiple possible readings for the following notation, referring to the example." For example, you can input the following into LLM: 'I work at Mizuno Transport' (reading: Mizuno Unso de hataraiteimasu), 'I visited Hashiiro' (reading: Hashiiro o houmon shimashita), and receive the answer from LLM: 'Reading: Hashiiro o houmon shimashita' (reading: Hashishoku o houmon shimashita).
[0026] The replacement unit 13C then searches for the word replacement rule "Hashiiro, Washiro, Kashiiro → Ashiiro" shown in Figure 6 for the multiple extracted readings. As a result, the word replacement rule "Hashiiro" matches the "Hashiiro" in the first reading "Hashiiro wo houmon shimashita". Therefore, the speech recognition result "Hashiiro" for the matching part "Hashiiro" is replaced with the corresponding technical term "Ashiiro" (reading "Ashiiro": notation "Ashiiro"). This results in the corrected string "I visited Ashiiro" being obtained and output from the information processing device 10. Note that "output" in this case can take various forms, such as display output, print output, or data transmission to an external source of the information processing device 10.
[0027] [Regarding the processing performed in the information processing device] The processing performed in the information processing device 10 (processing related to the information processing method of this disclosure) will be described below in accordance with the flowcharts in Figures 8 to 11. For example, the processing in Figure 8 is started when triggered by the input of a start command by the operator of the information processing device 10.
[0028] The error trend information acquisition unit 11 in the information processing device 10 acquires an error trend list as a preliminary step (step S1 in Figure 8). Details of this step S1 are shown in Figure 9.
[0029] As shown in Figure 9, the string matching unit 11A in the error tendency information acquisition unit 11 acquires the speech recognition result (string) recognized by the speech recognition device 20 from the audio data with a transcription (step S11), and associates the speech recognition result (string) with the transcription (correct reading) (step S12). The reading extraction unit 11B determines whether the part of speech of each erroneous part in the speech recognition result (string) is a verb or a particle (step S13), and extracts the reading of the erroneous part only if the part of speech of the erroneous part is neither a verb nor a particle (step S14). In other words, if the part of speech of the erroneous part is a verb or a particle, the reading of that erroneous part is excluded from the extraction target.
[0030] Once the processes in steps S13 to S14 have been completed for all erroneous locations, the trend acquisition unit 11C acquires an error trend list L representing the error trend in speech recognition (step S16). The trend acquisition unit 11C stores the acquired error trend list L.
[0031] Returning to Figure 8, in the next step S2, the reading expansion information acquisition unit 12 acquires word substitution rules as reading expansion information. Details of this step S2 are shown in Figure 10.
[0032] As shown in Figure 10, the list acquisition unit 12A in the reading expansion information acquisition unit 12 acquires a list of technical terms, including their spelling and correct pronunciation, from an external server (step S21), and the word replacement rule acquisition unit 12B performs reading expansion on the readings of technical terms in the error tendency list and acquires word replacement rules as reading expansion information (step S22).
[0033] Returning to Figure 8, in the next step S3, the control unit 13 replaces one of the multiple readings of the technical term with the correct reading and outputs the corrected (replaced) reading. Details of this step S3 are shown in Figure 11.
[0034] As shown in Figure 11, the string acquisition unit 13A in the control unit 13 acquires the speech recognition result (string) of the target speech from the speech recognition device 20 (step S31), and the reading extraction unit 13B extracts multiple readings from the string by conventional morphological analysis or by querying the LLM described above (step S32). The replacement unit 13C compares the extracted multiple readings with word replacement rules, replaces the string of one reading that matches with the corresponding technical term (step S33), and outputs the corrected (replaced) string (step S34).
[0035] As described above, the embodiment allows for the appropriate correction of speech recognition results, including those for technical terms, to the correct pronunciation, taking into account that multiple pronunciations are possible. Since the information processing device 10 performs the correction to the correct pronunciation as a post-processing step after the recognition process by the speech recognition device 20, the above invention is universally applicable to any speech recognition method, whether it is the first type of speech recognition (e.g., DNN-HMM type speech recognition) or the second type of speech recognition (e.g., End-to-End type speech recognition), and the same effects can be obtained.
[0036] Furthermore, as shown in Figure 4, by the process described above, reading expansion information (word substitution rules) can be obtained by further targeting the technical terms obtained by dividing the technical terms. Therefore, the above invention is also applicable to the technical terms obtained by dividing them, and the same effect can be obtained. In addition, even when vowel extensions occur in the reading, the process described in Figure 5 can be performed in the same way as when vowel extensions do not occur in the reading. Therefore, the above invention is also applicable when vowel extensions occur in the reading, and the same effect can be obtained.
[0037] It should be noted that the configuration of the information processing device 10 is not limited to the configuration shown in Figure 1, and it is not essential that the information processing device 10 includes an error tendency information acquisition unit 11. For example, as shown in the system 1S in Figure 12, an external server 11S may be provided with the various functional units of the error tendency information acquisition unit 11 shown in Figure 1, and the information processing device 10 may be provided with a read / expand information acquisition unit 12 and a control unit 13. In this case, the word replacement rule acquisition unit 12B can perform read / expand by appropriately querying the error tendency list L in the external server 11S and acquire word replacement rules.
[0038] The gist of this disclosure is found in the following [1] to [6].
[0039] [1] An information processing device comprising: an error trend information acquisition unit that acquires reading expansion information that shows the correspondence between the correct reading of a technical term and the reading when the technical term is misread, based on error trend information, which is information regarding the error trend of the correct reading in a string recognized from speech that reads the correct reading, and a list of technical terms that includes notation and correct reading of technical terms that are not general terms; and a control unit that replaces one of the multiple readings of a technical term with the correct reading, based on multiple readings of a technical term extracted from a string recognized from the target speech and the reading expansion information.
[0040] [2] The information processing apparatus according to [1], wherein the reading and expansion information acquisition unit further targets the technical terms obtained by dividing the technical terms and acquires the reading and expansion information.
[0041] [3] An error tendency information acquisition unit that acquires error tendency information of the correct reading based on the character string recognized from the voice that reads the correct reading and the correct reading, and further includes the information processing apparatus according to [1] or [2].
[0042] [4] The error tendency information acquisition unit limits the reading for which the error tendency information is to be acquired to the reading of words of a predetermined part of speech. The information processing apparatus according to [3].
[0043] [5] The error tendency information acquisition unit acquires, as the error tendency information, an error tendency list in which the error tendency at the character level is associated with the occurrence probability of the error tendency. The information processing apparatus according to [3] or [4].
[0044] [6] A step of acquiring reading expansion information, which is information indicating the correspondence between the correct reading of the technical term and the reading when the technical term is misread, based on the error tendency information, which is information regarding the error tendency of the correct reading in the character string recognized from the voice that reads the correct reading, and a technical term list including notations and correct readings regarding technical terms that are not general terms; and a step of replacing one of the plurality of readings of the technical term with the correct reading based on the plurality of readings of the technical term extracted from the character string recognized from the target voice and the reading expansion information. An information processing method.
[0045] [Explanation of terms, explanation of hardware configuration (FIG. 13), etc.] Note that the block diagrams used in the above description of the embodiments show blocks of functional units. These functional blocks (components) are realized by an arbitrary combination of at least one of hardware and software. Also, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one physically or logically combined device, or two or more physically or logically separated devices may be directly or indirectly connected (for example, using wired, wireless, etc.), and realized using these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.
[0046] Functions include, but are not limited to, judgment, decision-making, determination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, solution, selection, selection determination, establishment, comparison, assumption, expectation, regarded as, broadcasting, notification, communication, forwarding, configuration, reconfiguration, allocation (allocating, mapping), assignment, etc. For example, a functional block (component) that enables transmission is called a transmitting unit or a transmitter. As described above, the implementation method is not particularly limited.
[0047] For example, the information processing apparatus in one embodiment of the present disclosure may function as a computer that executes the processing of the present disclosure. FIG. 13 is a diagram showing an example of the hardware configuration of an information processing apparatus 10 according to one embodiment of the present disclosure. Physically, the above-described information processing apparatus 10 may be configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.
[0048] In the following description, the term "apparatus" can be read as a circuit, device, unit, etc. The hardware configuration of the information processing apparatus 10 may be configured to include one or more of each apparatus shown in the figure, or may be configured without including some apparatuses.
[0049] Each function in the information processing apparatus 10 is realized by causing the processor 1001 to perform operations by loading a predetermined software (program) onto hardware such as the processor 1001 and the memory 1002, controlling communication by the communication device 1004, or controlling at least one of reading and writing data in the memory 1002 and the storage 1003.
[0050] The processor 1001 controls the entire computer, for example, by running an operating system. The processor 1001 may be composed of a central processing unit (CPU) that includes interfaces with peripheral devices, control units, arithmetic units, registers, etc.
[0051] Furthermore, the processor 1001 reads programs (program code), software modules, data, etc., from at least one of the storage 1003 and the communication device 1004 into the memory 1002, and executes various processes accordingly. The program used is one that causes the computer to execute at least a part of the operations described in the above embodiment. Although it has been explained that the various processes are executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The program may also be transmitted from a network via a telecommunications line.
[0052] The memory 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may also be called a register, cache, main memory, etc. The memory 1002 can store executable programs (program code), software modules, etc., for carrying out a wireless communication method according to one embodiment of the present disclosure.
[0053] The storage 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of the memory 1002 and the storage 1003.
[0054] The communication device 1004 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc. The communication device 1004 may be configured to include, for example, a high-frequency switch, duplexer, filter, frequency synthesizer, etc., in order to implement at least one of frequency division duplex (FDD) and time division duplex (TDD).
[0055] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).
[0056] Furthermore, each device, such as the processor 1001 and memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or different buses may be configured for each device.
[0057] Furthermore, the information processing device 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the processor 1001 may be implemented using at least one of these hardware components.
[0058] The notification of information is not limited to the embodiments described herein and may be carried out by other means. For example, the notification of information may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.
[0059] Each aspect / embodiment described in this disclosure refers to LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), 6th generation mobile communication system (6G), xth generation mobile communication system (xG) (xG (where x is, for example, an integer or decimal)), FRA (Future Radio Access), NR (new Radio), New radio access (NX), Future generation radio access (FX), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, may apply to at least one system utilizing UWB (Ultra-WideBand), Bluetooth®, or other appropriate systems, and to next-generation systems extended, modified, generated, or defined based thereon. Alternatively, multiple systems may be applied in combination (e.g., a combination of at least one of LTE and LTE-A with 5G).
[0060] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described in this disclosure may be reordered, provided they do not contradict each other. For example, the methods described in this disclosure present various step elements using exemplary order and are not limited to the specific order presented.
[0061] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be transmitted to other devices.
[0062] The determination may be made by a value represented by one bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, by comparing with a predetermined value).
[0063] Each aspect / embodiment described in this disclosure may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).
[0064] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.
[0065] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name.
[0066] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0067] The information, signals, etc. described in this disclosure may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0068] In addition, terms used in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of the channel and symbol may be a signal (signaling). Also, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, cell, frequency carrier, etc.
[0069] The terms “system” and “network” as used in this disclosure are interchangeable.
[0070] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values from a given value, or other corresponding information. For example, wireless resources may be indicated by an index.
[0071] The names used for the parameters described above are not restrictive in any way. Furthermore, the formulas and other expressions using these parameters may differ from those expressly disclosed in this disclosure. Various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable name, and therefore, the various names assigned to these various channels and information elements are not restrictive in any way.
[0072] As used in this disclosure, the terms “determining” and “determining” may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, or inquiring (e.g., searching in a table, database, or other data structure), or ascertaining. “Determining” may also include receiving (e.g., receiving information), transmitting (e.g., sending information), inputting, outputting, or accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."
[0073] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."
[0074] Any reference to elements using the designations “first,” “second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to the first and second elements do not imply that only two elements may be employed, or that the first element must precede the second element in any way.
[0075] Where the terms “include,” “including,” and variations thereof are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.
[0076] In this disclosure, if articles are added through translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.
[0077] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."
[0078] 1...System, 10...Information processing device, 11...Error tendency information acquisition unit, 11A...String mapping unit, 11B...Reading extraction unit, 11C...Tendency acquisition unit, L...Error tendency list, 12...Reading expansion information acquisition unit, 12A...List acquisition unit, 12B...Word replacement rule acquisition unit, R...Word replacement rule, 13...Control unit, 13A...String acquisition unit, 13B...Reading extraction unit, 13C...Replacement unit, 20...Speech recognition device, 1001...Processor, 1002...Memory, 1003...Storage, 1004...Communication device, 1005...Input device, 1006...Output device, 1007...Bus.
Claims
1. An information processing device comprising: an error trend information acquisition unit that acquires reading expansion information, which is information indicating the correspondence between the correct reading of a technical term and the reading of the technical term when it is misread, based on error trend information, which is information regarding the error trend of the correct reading in a string recognized from speech that reads the correct reading, and a list of technical terms that include notation and correct reading of technical terms that are not general terms; and a control unit that replaces one of the multiple readings of the technical term with the correct reading, based on multiple readings of the technical term extracted from a string recognized from the target speech and the reading expansion information.
2. The information processing apparatus according to claim 1, wherein the reading and expansion information acquisition unit further acquires the reading and expansion information by targeting the technical terms obtained by dividing the technical terms.
3. An information processing apparatus according to claim 1, further comprising: an error trend information acquisition unit that acquires error trend information of the correct reading based on the string recognized from the voice that reads out the correct reading and the correct reading.
4. The information processing apparatus according to claim 3, wherein the error tendency information acquisition unit limits the readings for which the error tendency information is to be acquired to the readings of words of predetermined parts of speech.
5. The information processing apparatus according to claim 3, wherein the error trend information acquisition unit acquires an error trend list as the error trend information, which associates error trends at the character level with the probability of occurrence of said error trends.
6. An information processing method comprising: an information processing device obtaining reading expansion information, which is information showing the correspondence between the correct reading of a technical term and the incorrect reading of the technical term, based on error trend information, which is information about the tendency of errors in the correct reading of a string recognized from speech that reads the correct reading, and a list of technical terms, which includes the notation and correct reading of technical terms that are not general terms; and an information processing device replacing one of the multiple readings of the technical term with the correct reading, based on multiple readings of the technical term extracted from a string recognized from the target speech and the reading expansion information.
Citation Information
Patent Citations
Voice processor
JP1984121400A
Dictionary update device and program
JP2018040906A
Recognition result correction device, recognition result correction method, and program
JP2020030379A