Measurement recording device, measurement recording program and measurement recording method
The measurement recording device addresses the challenges of accurate voice input for measurement results by using a device with voice recognition, misrecognition correction, and extraction units to ensure accurate recording of measurement content, even from low-skill users.
Patent Information
- Application Number
- JP2023205104
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-17
AI Technical Summary
Existing voice input methods for recording measurement results require accurate voice recognition, which can be skill-dependent and prone to misrecognition, especially when fillers or pauses are present.
A measurement recording device equipped with a voice recognition unit, a misrecognition correction unit, an extraction unit, and a recording unit, which corrects misrecognized text using a correction database and extracts text corresponding to registered expressions to accurately record measurement results.
The solution enables accurate extraction and recording of measurement content voice data even from low-skill users, reducing input workload and improving measurement result recording accuracy.
Smart Images

Figure 2025090097000001_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed in this specification relates to a measurement recording technology by voice input.
Background Art
[0002] When measuring multiple measurement points at a construction site or the like, the measurer records the measurement results for each measurement point. Examples of methods for recording measurement results include a method of inputting and recording the content of the measurement results on a terminal carried by the measurer, or a method of recording a photographed image of the measurement point together with the measurer's comments.
[0003] Another method for recording measurement results is a method in which the measurer speaks the content of the measurement results, inputs the voice into a terminal, converts the voice into text by the voice recognition function of the terminal, and records the text information (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the method of converting the voice of the measurement content into text and directly recording the text information, accurate voice input by the measurer is required, and the accuracy of the recording may also vary depending on the skill of the measurer. In addition, if the measurer's speech contains pauses (fillers) or the like, misrecognition may occur in voice recognition.
[0006] The technology disclosed in the present specification has been made in view of the problems described above, and is a technology for facilitating voice input for measurement results even by a measurer with low skill level and reducing the work load of input.
Means for Solving the Problems
[0007] A measurement recording device which is a first aspect of the technology disclosed in the present specification includes a voice recognition unit for voice-recognizing voice data including measurement information and outputting a voice recognition text, a misrecognition correction unit for correcting a misrecognition text which is text generated by misrecognition in the voice recognition text, an extraction unit for extracting, as an extraction text, text corresponding to a registered expression which is an expression registered in advance according to the measurement content in the voice recognition text, and a recording unit for recording the extraction text in association with each corresponding measurement content, wherein the misrecognition correction unit corrects the misrecognition text based on a correction database in which the misrecognition text and a correction text which is a corrected text corresponding to the misrecognition text are associated and recorded, and the extraction unit extracts the extraction text from the voice recognition text in which the misrecognition text has been corrected. A measurement recording device which is a second aspect of the technology disclosed in the present specification is related to the measurement recording device which is the first aspect, in the correction database, a plurality of the misrecognition texts and the number of times of the misrecognition thereof are associated and recorded with respect to the correction text, and the misrecognition correction unit corrects the misrecognition text to the correction text which is most frequently misrecognized for the misrecognition text based on the number of times of the misrecognition. A measurement recording device which is a third aspect of the technology disclosed in the present specification is related to the measurement recording device which is the second aspect, in the correction database, the correction text, the misrecognition text and the number of times of the misrecognition thereof are recorded in json format. The measurement recording device according to the fourth aspect of the technology disclosed in the present specification is related to the measurement recording device according to any one of the first to third aspects, and further includes a filler correction unit for correcting a filler text which is text generated by a filler in the speech recognition text, and the extraction unit extracts the extraction text from the speech recognition text in which the filler text has been corrected. The measurement recording device according to the fifth aspect of the technology disclosed in the present specification is related to the measurement recording device according to the fourth aspect, and the filler correction unit corrects the filler text in the speech recognition text using a learned model obtained by learning the speech recognition text including the filler text with the speech recognition text in which the filler text has been removed as teacher data. The measurement recording device according to the sixth aspect of the technology disclosed in the present specification is related to the measurement recording device according to the fourth or fifth aspect, and the functions of the filler correction unit and the extraction unit are realized using bidirectional encoder representation from transformers (BERT). The measurement recording device according to the seventh aspect of the technology disclosed in the present specification is related to the measurement recording device according to any one of the fourth to sixth aspects, and when the measurement information included in the speech recognition text does not correspond to the pre-registered measurement content, the function of the filler correction unit is realized using bidirectional and auto-regressive transformers (BART). The measurement recording program according to the eighth aspect of the technology disclosed in the present specification is a measurement recording program having a plurality of computer-executable instructions to be executed by one or more processors. By the plurality of instructions executed by the processor, the computer is caused to perform speech recognition on speech data including measurement information and output a speech recognition text, cause the computer to correct a misrecognition text which is text generated by misrecognition in the speech recognition text, cause the computer to extract, as an extraction text, text corresponding to a registered expression which is an expression registered in advance according to the measurement content in the speech recognition text, cause the computer to record the extraction text in association with the corresponding measurement content, and correcting the misrecognition text is based on a correction database in which the misrecognition text and a corrected text which is the corrected text corresponding to the misrecognition text are associated and recorded, and correcting the misrecognition text, and extracting the extraction text is extracting the extraction text from the speech recognition text in which the misrecognition text has been corrected. The measurement recording method according to the ninth aspect of the technology disclosed in the present specification includes a step of performing speech recognition on speech data including measurement information and outputting a speech recognition text, a step of correcting a misrecognition text which is text generated by misrecognition in the speech recognition text, a step of extracting, as an extraction text, text corresponding to a registered expression which is an expression registered in advance according to the measurement content in the speech recognition text, and a step of recording the extraction text in association with the corresponding measurement content. The step of correcting the misrecognition text is a step of correcting the misrecognition text based on a correction database in which the misrecognition text and a corrected text which is the corrected text corresponding to the misrecognition text are associated and recorded, and the step of extracting the extraction text is a step of extracting the extraction text from the speech recognition text in which the misrecognition text has been corrected.
Advantages of the Invention
[0008] According to at least the 1st, 8th, and 9th aspects of the technology disclosed in the present specification, by extracting the text corresponding to the registered expression and recording it in association with the measurement content, even for voice data by a measurer with low skill level, the voice regarding the measurement content can be accurately extracted and recorded in association with the corresponding measurement content. Therefore, voice input regarding the measurement result becomes easy, and the work burden of the input can be reduced.
[0009] In addition, the objects, features, aspects, and advantages related to the technology disclosed in the present specification will become even clearer by the following detailed description and the attached drawings.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments will be described with reference to the accompanying drawings. In the following embodiments, detailed features and the like are also shown for the purpose of explaining the technology, but these are examples, and not all of them are necessarily essential features for the embodiments to be practicable.
[0012] Note that the drawings are shown schematically, and for convenience of explanation, omissions or simplifications of the configuration are made in the drawings as appropriate. Also, the mutual relationships of the sizes and positions of the configurations shown in different drawings are not necessarily accurately described and can be changed as appropriate. In addition, in drawings such as a plan view that is not a cross-sectional view, hatching may be added to facilitate understanding of the content of the embodiment.
[0013] Also, in the descriptions shown below, the same reference numerals are given to and illustrated for the same components, and their names and functions are also assumed to be the same. Therefore, detailed descriptions thereof may be omitted to avoid duplication.
[0014] Also, in the descriptions described in the present specification, when a certain component is described as "comprising", "including", or "having", etc., it is not an exclusive expression excluding the existence of other components unless otherwise specified.
[0015] Also, in the descriptions described in the present specification, even when ordinal numbers such as "first" or "second" are used, these terms are used for convenience to facilitate understanding of the content of the embodiment, and the content of the embodiment is not limited to the order or the like that may be caused by these ordinal numbers.
[0016] <Embodiment> Hereinafter, a measurement recording apparatus, a measurement recording program, and a measurement recording method according to the present embodiment will be described.
[0017] Note that basic data processing technologies (e.g., communication technology, transmission technology, data acquisition technology, data recording technology, data processing technology, data analysis technology, speech recognition technology, natural language processing technology, image processing technology, or visualization technology, etc.) for realizing the content of this embodiment are well-known technologies, and thus detailed descriptions thereof are omitted.
[0018] <Regarding the configuration of the measurement recording device> FIG. 1 is a diagram conceptually showing an example of the configuration of a measurement recording device according to this embodiment. As shown in the example of FIG. 1, the measurement recording device 10 includes a speech recognition unit 42, a misrecognition correction unit 44, an extraction unit 46, and a recording unit 48. Further, the measurement recording device 10 can include a filler correction unit 50.
[0019] Voice data 100 including measurement information is input to the speech recognition unit 42. Here, the measurement information is information including measurement content, measurement values, evaluation results regarding the measurement, and the like. Then, the speech recognition unit 42 outputs a speech recognition text by applying a known speech recognition technology to the input voice data 100.
[0020] The misrecognition correction unit 44 corrects misrecognition text, which is text generated by misrecognition in the speech recognition text (misrecognition correction process). The misrecognition correction process is performed using a correction rule dictionary described later.
[0021] The extraction unit 46 extracts, as extraction text, text corresponding to a registered expression, which is an expression registered in advance according to the measurement content, in the speech recognition text.
[0022] The recording unit 48 records the extraction text in association with each corresponding measurement content.
[0023] The filler correction unit 50 corrects filler text, which is text generated by a filler in the speech recognition text.
[0024] FIG. 2 is a diagram schematically showing an example of the configuration of a measurement recording system according to the present embodiment. As shown in the example of FIG. 2, the measurement recording system 1 includes a measurement recording device 10 and a terminal 12 of a measurer.
[0025] The measurement recording device 10 and the terminal 12 are connected to each other in a communicable state via a communication network N such as the Internet or a mobile communication line.
[0026] The terminal 12 is an information processing terminal used by a measurer, and is configured by, for example, a smartphone, a mobile phone, a tablet terminal, a laptop personal computer, or a wearable terminal. The terminal 12 includes, for example, a display including a display, a recording device (sound collecting device) including a microphone, and a photographing device including a camera. Further, the terminal 12 transmits and receives data to and from the measurement recording device 10 through the communication network N, and information transmitted from the measurement recording device 10 can be displayed on a display screen configured by the display.
[0027] Note that the display may be a display provided in the terminal 12 itself, or a display connected to the terminal 12 in a wired or wireless form. In addition to a general stationary display, the display connected to the terminal 12 may include an HMD (Head Mounted Display) such as a VR goggles.
[0028] An application program for measurement (hereinafter, measurement application) may be installed in the terminal 12. By starting the measurement application on the terminal 12, the measurer can specify measurement information and check a plurality of measurement locations where measurement regarding the specified measurement information is performed. The measurement information is information indicating a site where measurement is performed (for example, a construction site), a process to be measured, or an area where measurement is performed.
[0029] Also, after the measurement is performed, the measurer can speak the content corresponding to the measurement result and input (record) the voice into the terminal 12 in the state where the measurement application is activated. The terminal 12 converts the input voice into data to generate voice data. Then, the terminal 12 can transmit the generated voice data to the measurement recording device 10 via the communication network N.
[0030] The measurement recording device 10 is composed of a computer and is configured by, for example, a personal computer (PC), a workstation, or a server computer. The measurement recording device 10 may be composed of a single computer or may be composed of a plurality of computers that are parallel-distributed.
[0031] Also, when the computer constituting the measurement recording device 10 is a server computer, it may be a server computer for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). In this case, when necessary information is input on the client terminal, the above server computer performs various processes (calculations) based on the input information, and the calculation result is output on the client terminal side. That is, the function of the server computer that is the measurement recording device 10 can be utilized on the client terminal side.
[0032] As shown in the example of FIG. 2, the computer that constitutes the measurement recording device 10 includes a processor 21, a memory 22, a storage 23, a communication interface 24, an input device 25, and an output device 26. The processor 21 is constituted by, for example, a CPU (Central Processing Unit), MPU (Micro-Processing Unit), MCU (Micro Controller Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), TPU (Tensor Processing Unit), or ASIC (Application Specific Integrated Circuit).
[0033] The memory 22 is constituted by, for example, semiconductor memories such as ROM (Read Only Memory) and RAM (Random Access Memory).
[0034] The storage 23 is constituted by, for example, a flash memory, HDD (Hard Disc Drive), SSD (Solid State Drive), FD (Flexible Disc), MO disk (Magneto-Optical disc), CD (CompactDisc), DVD (Digital Versatile Disc), SD card (Secure Digital card), or USB memory (Universal Serial Bus memory). The storage 23 may be built into the computer main body that constitutes the measurement recording device 10, or may be attached to the computer main body in an external form. Further, the storage 23 may be constituted by a NAS (Network Attached Storage) or the like, or may be an external device that can communicate with one computer that constitutes the measurement recording device 10 through a communication network N, for example, an online storage or a database server. The storage 23 corresponds to the recording unit 48 in FIG. 1.
[0035] The communication interface 24 may be constituted by, for example, a network interface card or a communication interface board. The computer constituting the measurement recording device 10 can perform data communication with other devices connected to the communication network N via the communication interface 24.
[0036] The input device 25 is constituted by, for example, a keyboard, a mouse, or a touch panel. The output device 26 is constituted by, for example, a display and a speaker.
[0037] Also, a program for an operating system (OS) and a program for measurement are installed as software in the computer constituting the measurement recording device 10. By these programs being read out and executed by the processor 21, the computer constituting the measurement recording device 10 exhibits its functions. Specifically, in cooperation with the terminal 12, it executes a series of data processing related to measurement. Specifically, it realizes the functions of the voice recognition unit 42, the misrecognition correction unit 44, the extraction unit 46, and the filler correction unit 50 in FIG. 1.
[0038] Also, information related to measurement, measurement result information, etc. are stored in the storage 23 to construct a database. The information related to measurement includes the name of the related construction work, the period of the work, information about the specifications or construction location of the house to be constructed in the case of construction work, and also includes an image of the house to be constructed. The image of the house may be a drawing of the house (design drawing or construction drawing) drawn by BIM (Building Information Modeling) or CAD (Computer Aided Design), or may be a photographed image of an existing house having the same structure as the house planned for construction.
[0039] Among the information stored in the storage 23, various information necessary for measurement includes information regarding a plurality of measurement locations, information necessary for the assignment of labels described later, a speech recognition model, and a natural language processing model. The speech recognition model and the natural language processing model are mathematical models constructed by machine learning with or without teacher data, and specifically, they are constituted by AI (Artificial Intelligence).
[0040] The measurement result information is recorded in association with the measurement location and is information indicating the content of the measurement result corresponding to the registered expression. Also, when measurement is performed from two or more directions for one measurement location, the measurement result information indicates the measurement results for each direction. Information on the measurement execution date and the information of the measurer may be further associated with the measurement result information.
[0041] Next, the functions of the measurement recording device 10 will be described.
[0042] The functions of the measurement recording device 10 (that is, the functions realized by the speech recognition unit 42, the misrecognition correction unit 44, the extraction unit 46, the filler correction unit 50, etc. shown in FIG. 1) are realized by the processor 21 shown in FIG. 2 executing a measurement program. In other words, the processor 21 executes various processes defined in the measurement program in order to exhibit the functions of the measurement recording device 10.
[0043] Specifically, when the measurer performs measurement on each of a plurality of measurement locations and inputs the measurement results, the processor 21 repeatedly executes display processing, identification processing, reception processing, text conversion processing, filler correction processing, misrecognition correction processing, assignment processing, extraction processing, and recording processing for each measurement location in a corresponding manner.
[0044] Hereinafter, each process executed by the processor 21 will be described with reference to FIG. 3. Here, FIG. 3 is a flowchart showing an example of the operation of the measurement recording device 10 according to the present embodiment.
[0045] Among the above, the display process is a process of displaying an image indicating the positions of a plurality of measurement points on the screen of the terminal 12 of the measurer. Note that the measurement targets are not limited to the measurement (inspection) at the construction site. The display process is executed based on the measurement information sent from the terminal 12. Specifically, the processor 21 identifies the site, process, and measurement area input by the measurer from the measurement information received from the terminal 12. Then, based on these identified pieces of information, the processor 21 selects the target image from the images stored in the storage 23.
[0046] Also, in the display process, as will be described later, the processor 21 may display information corresponding to the measurement result information recorded in association with the measurement point in the area corresponding to the measured measurement point among the images to be displayed.
[0047] Note that the information corresponding to the measurement result information may include characters or numerical values representing the content indicated by the measurement result information, figures or symbols corresponding to the content indicated by the measurement result information, and information related to the content indicated by the measurement result information (for example, an arrow or a sign indicating the measurement point associated with the measurement result information).
[0048] Among the above, the identification process is a process of identifying the designated measurement point designated by the measurer from among a plurality of measurement points. The identification process is executed based on the information of the designated measurement point sent from the terminal 12. Specifically, the processor 21 identifies the designated measurement point designated by the measurer from among a plurality of measurement points (strictly speaking, measurement points where measurement has not been performed) based on the information of the designated measurement point received from the terminal 12.
[0049] Among the above, the reception process is a process of receiving, by the voice recognition unit 42, the voice emitted by the measurer, specifically, the input of voice data when the measurer speaks the content corresponding to the measurement result. That is, it is a process of receiving the voice data generated by the speech in step ST12 in FIG. 3. In the reception process, the processor 21 receives the voice data recorded by the terminal 12 from the terminal 12 via the communication network N.
[0050] Among the above, the text conversion process is a process of recognizing the voice data received in the reception process, specifically, the voice data when the measurer speaks the content corresponding to the measurement result, and converting the voice data into text. That is, it is a process of generating voice recognition text by voice recognition in step ST13 in FIG. 3. In the text conversion process, the processor 21 uses a learned model (AI) for voice recognition to convert voice data into text. That is, when the voice data is input into the learned model for voice recognition, the above voice recognition text is obtained as the output. The voice recognition text is generated in the same language as the language used by the measurer. For example, when the measurer speaks in Japanese, a Japanese voice recognition text is generated.
[0051] Among the above, the filler correction process is a process of correcting the voice recognition text by deleting or converting the filler text in the voice recognition text. That is, it is a process of determining whether to perform the filler correction process in step ST14 in FIG. 3, and generating the voice recognition text after filler correction by the filler correction process in step ST15. Here, the filler text is text generated by hesitation (filler).
[0052] The filler correction process is performed using a learned model for filler removal generated by machine learning with the voice recognition text including fillers (mainly text generated by character recognition from voice data) and the voice recognition text with fillers removed as teacher data. The filler correction process is performed using, for example, bidirectional and auto-regressive transformers (BART). Note that even if the filler text remains in the voice recognition text as described later, if the appropriate exclusion label is assigned, it will be appropriately excluded when recording and no problem will occur. However, by deleting or converting the filler text in advance in the filler correction process, the accuracy of recording can be improved.
[0053] Whether to perform the filler correction process is determined by whether the measurement content to be targeted is the measurement content registered in advance (step ST14). If the measurement content to be targeted is the registered measurement content, the process proceeds to step ST15 to perform the filler correction process. On the other hand, if the measurement content to be targeted is not the registered measurement content, the process proceeds to step ST16 to avoid the filler correction process.
[0054] Whether the measurement content to be targeted is the registered measurement is determined based on, for example, the measurement information included in the speech recognition text. If the measurement content to be targeted is not the registered measurement content, the process proceeds to step ST16 without going through step ST15.
[0055] Note that along with the filler correction process using BART, the correction of the misrecognition text (based on the filler) described later is also performed.
[0056] The misrecognition correction process among the above is a process of correcting the speech recognition text by deleting or converting the misrecognition text in the speech recognition text. That is, it is the process corresponding to the misrecognition correction process in step ST16 in FIG. 3. Here, the misrecognition text is the text generated by misrecognizing the speech data.
[0057] The misrecognition correction process is a process of converting the misrecognition text into the correct text by referring to the correction rule dictionary created in advance.
[0058] The correction rule dictionary is a database recorded in a json format. FIG. 4 is a diagram showing an example of text recorded in the correction rule dictionary. As shown in the example in FIG. 4, in the correction rule dictionary, misrecognized similar-sounding words and the number of times of misrecognition are recorded in association with a correctly recognized reference text (correction text). In the example shown in FIG. 4, misrecognized similar-sounding words "live" and "daibu" are recorded for the reference text "nai" and "daibu", respectively, and "108" and "6" are recorded as the number of times of misrecognition. Similarly, misrecognized similar-sounding word "yomi" is recorded for the reference text "yomi" and "yomi", respectively, and "1" is recorded as the number of times of misrecognition.
[0059] The correction rule dictionary records the text from which differences are extracted by comparing the speech recognition text with the text generated by transcribing the speech data (correction text that serves as teacher data). By extracting differences from text specialized for a specific measurement (test), it is possible to create a correction rule dictionary that reflects the tendency of misrecognition in a specific measurement.
[0060] In the misrecognition correction process, when misrecognized text (misrecognized text) recorded in the correction rule dictionary is found in the speech-recognized text, the misrecognized text is converted to the reference text (correction text) that is recorded in association with the misrecognized text and that has been misrecognized the most times.
[0061] When the voice recognition software used for voice recognition is changed, it is desirable to reconstruct the correction rule dictionary so as to reflect the recognition tendency of the new voice recognition software.
[0062] In this embodiment, the misrecognition correction process is performed prior to the extraction process described below, but the misrecognition correction process may be performed after the extraction process.
[0063] Among the above, the assignment process is a process of assigning a label corresponding to each of two or more words included in the speech recognition text from among a plurality of pre-set labels. In the assignment process, the processor 21 first divides the text information into two or more words using a pre-trained model (AI) for natural language processing. Thereafter, the processor 21 assigns a label corresponding to the meaning of each of the two or more words included in the speech recognition text by means of a technique of named entity recognition (NER).
[0064] The plurality of labels are a plurality of labels including an exclusion label, specifically, BIO (Begin Inside Outside) tags. In the present embodiment, the correspondence between words and labels (that is, the label assignment rule) is determined in advance and is recorded in advance in the storage 23 or the like as table data such as that shown in FIG. 5. Note that FIG. 5 is a diagram showing an example of table data for recording the correspondence between words and labels.
[0065] The processor 21 refers to this table data and assigns a label to each of two or more words included in the text information.
[0066] Here, the label assignment procedure will be specifically described by taking as an example the case where speech recognition text of "X direction +1 mm, Y direction 0 mm" is obtained. The words included in the above speech recognition text are "X (ex), "direction", "1 (one)", "mm (millimeter)", "Y (why)", "direction", "0 (zero)", "mm (millimeter)".
[0067] On the one hand, in the table data shown in FIG. 5, for the words "X" and "Y", the label "B-RESULT", which is the label of the direction of the measurement result, is associated. Also, for the word "direction", the label "I-DIRECTION", which indicates that it is the direction of the measurement result, is associated. Further, for words representing numerical values such as "one" or "zero", the label "I-RESULT", which is the label of the numerical value of the measurement result, is associated. Also, for the word "milli", the "I-UNIT", which indicates that it is the unit of the measurement result, is associated.
[0068] From the above relationships, the following labels are assigned to each word included in the text information "X direction +1 mm, Y direction 0 mm".
[0069] That is, B-DIRECTION is assigned to "X (ex)", I-DIRECTION is assigned to "direction", B-RESULT is assigned to "1 (one)", B-UNIT is assigned to "mm (milli)", B-DIRECTION is assigned to "Y (why)", I-DIRECTION is assigned to "direction", B-RESULT is assigned to "0 (zero)", and I-UNIT is assigned to "mm (milli)".
[0070] Also, when the measurer utters unnecessary languages such as "ah" and "eh" during voice input, that is, so-called fillers, the voice recognition text obtained by converting the voice data into text includes words corresponding to the fillers. On the other hand, in the table data shown in FIG. 5, for words other than those associated with the above four types of labels (B tags, I tags) including fillers, an exclusion label "O (Outside)" is associated. Therefore, when the text information includes a word corresponding to a filler, the exclusion label "O" is assigned to that word. Note that the exclusion label may be assigned to words other than fillers, specifically, words not related to the measurement result.
[0071] Note that the label is not limited to the BIO tag, and labels other than the BIO tag, such as the BILUO tag, may be used. Also, the assignment of labels is not limited to the case of referring to the table data shown in FIG. 5, and a learning model that assigns appropriate labels based on the meaning of words, that is, AI, may be used to assign labels to each word.
[0072] FIG. 6 is a diagram showing another example of the speech recognition text to which the assignment process has been performed. The parentheses in FIG. 6 are the labels assigned to the respective texts.
[0073] Among the above, the extraction process is a process of extracting, as extraction text, a phrase (text) corresponding to the registered expression from the speech recognition text when the registered expression is included in the speech recognition text. That is, it is a process corresponding to the extraction process in step ST17 in FIG. 3.
[0074] Here, the registered expression is an expression (word) that directly or indirectly represents the measurement content, is determined for each measurement content, and is registered in advance before the measurement is performed. The registered expression may be an expression (word) corresponding to a numerical value such as "1" or "2", or an expression (word) corresponding to a symbol such as "A" or "X". That is, in order to simplify the registered expression, when an identification numerical value or symbol is assigned to the registered expression, the identification numerical value or symbol may be used as the registered expression.
[0075] When a plurality of registered expressions are included in the speech recognition text, the processor 21 extracts, in the extraction process, the phrases corresponding to the registered expressions from the speech recognition text for each registered expression. More specifically, when measurements are performed from two or more directions for one measurement location and the speech recognition text includes two or more registered expressions indicating the directions in which the measurements are performed, the processor 21 extracts, in the extraction process, the phrases corresponding to the measurement directions from the speech recognition text for each direction.
[0076] In the extraction process, the processor 21 extracts a word (hereinafter referred to as the word with the corresponding label) to which a label corresponding to a registered expression is assigned among two or more words included in the speech recognition text, and which has a predetermined order relationship, and a word to which a label other than the exclusion label is assigned.
[0077] The label corresponding to the registered expression (corresponding label) is the label assigned to the word constituting the registered expression. For example, when the registered expression is "X direction", the label "B-DIRECTION" associated with the word "X (ex)" and the label "I-DIRECTION" associated with the word "direction" correspond to the label corresponding to the registered expression. Note that the label corresponding to the registered expression (corresponding label) may be one or two or more.
[0078] The word having a predetermined order relationship with the word of the corresponding label is, for example, the word following the word of the corresponding label in the speech recognition text. This is because when inputting voice, it is generally assumed that the measurer speaks the result of the management item indicated by the registered expression, that is, information corresponding to the measurement result, immediately after the registered expression. The word to which a label other than the exclusion label is assigned is a word to which a label other than the label "O" is assigned, and in short, it is a word other than a word not related to the measurement including the filler.
[0079] Then, a word (hereinafter referred to as the applicable word) having a predetermined order relationship with the word of the corresponding label and to which a label other than the exclusion label is assigned is extracted as a phrase corresponding to the registered expression. Also, when there are two or more applicable words, the applicable words having a certain meaning and coherence (strictly speaking, the group of applicable words) are extracted as a phrase corresponding to the registered expression.
[0080] The extraction process is performed using, for example, bidirectional encoder representation from transformers (BERT). By extracting the words and phrases (extracted text) corresponding to the registered representation using BERT, extracted text with a high degree of similarity to the registered representation can be extracted, enabling flexible handling even of speech recognition text generated from speech data that does not strictly follow the input format.
[0081] Note that filler correction is also performed in conjunction with the extraction process using BERT. If filler correction processing has already been performed, filler correction is performed with even higher accuracy.
[0082] The procedure for extracting the words and phrases corresponding to the registered representation will be specifically described with reference to FIG. 7, taking the speech recognition text "1 mm in the X direction, 0 mm in the Y direction" as an example. Here, FIG. 7 is a diagram showing examples of registered representations, labels, and words and phrases corresponding to the registered representations.
[0083] When the registered representations are "X direction" and "Y direction", the words of the corresponding labels are "X (ex) " and "direction (houkou)", and "Y (wai)" and "direction (houkou)".
[0084] In this case, the words "1 (ichi)" and "mm (milli)", which are in a predetermined order with the words "X (ex)" and "direction (houkou)" that are the words of the corresponding label, are as shown in the example in FIG. 7. Since these are words to which labels other than the exclusion labels are assigned, they correspond to the relevant words. Also, "1mm (ichimilli)", which is a combination of "1 (ichi)" and "mm (milli)", corresponds to a group of relevant words having a certain meaning and coherence. Therefore, in the extraction process, as shown in the example in FIG. 7, "1mm" is extracted as the word and phrase corresponding to the registered representation "X direction".
[0085] Similarly, the words "Y (wai)" and "direction (houkou)" which are the words of the corresponding label and the words in a predetermined order are, as shown in the example in FIG. 7, "0 (zero)" and "mm (milli)", and since these are the words to which labels other than the exclusion label are assigned, they correspond to the corresponding words. Further, "0mm (zero milli)" which is a combination of "0 (zero)" and "mm (milli)" corresponds to the corresponding group of words having a certain meaning and coherence. Therefore, in the extraction process, as shown in the example in FIG. 7, "0mm" is extracted as the word corresponding to the registered expression "Y direction".
[0086] Among the above, the recording process is a process of recording the measurement result information based on the words extracted in the extraction process, that is, the words corresponding to the registered expression (extracted text), in association with the measurement location for each measurement content in the recording unit 48. That is, it is a process corresponding to the recording process of step ST18 in FIG. 3.
[0087] The measurement result information based on the extracted words (text) is information representing the extracted words (specifically, text such as a numerical value or a determination result) as they are, information of a calculated value obtained by substituting the numerical value indicated by the extracted words into a predetermined arithmetic expression, information indicating the magnitude relationship between the numerical value indicated by the extracted words and a reference value, or information indicating whether the numerical value indicated by the extracted words is within a reference range.
[0088] FIG. 8 is a diagram showing an example of the recorded measurement result information. As shown in the example in FIG. 8, the words (text) corresponding to the registered expression are associated with measurement items (installation status), measurement locations (overall height, fence height), etc. and recorded in a table format.
[0089] FIG. 9 is a diagram showing an example of the recorded measurement result information. As shown in the example in FIG. 9, the words (text) corresponding to the registered expression are recorded in association with each measurement location of the measurement object (for example, a fence member) represented in the figure. For example, the corresponding measurement numerical values are described at positions near each measurement location.
[0090] In the recording process, when the processor 21 records the measurement result information by associating it with the measurement location for each measurement content, the processor 21 records the measurement result information in association with the designated measurement location specified in the most recently executed specific process.
[0091] Also, when the speech recognition text contains a plurality of registered expressions, in the extraction process, in order to extract the phrases corresponding to the registered expressions from the speech recognition text for each registered expression, accordingly, in the recording process, the processor 21 records the measurement result information based on the phrases extracted for each registered expression in association with the measurement location (designated measurement location). More specifically, when the speech recognition text contains two or more registered expressions indicating the direction in which the measurement is performed, in the extraction process, the phrases corresponding to the measurement execution direction are extracted from the speech recognition text for each direction, and in the recording process, the measurement result information based on the phrases extracted for each direction is recorded in association with the measurement location (designated measurement location).
[0092] Next, an information processing flow using the measurement recording device 10 will be described with reference to FIG. 10. Here, FIG. 10 is a flowchart showing an example of the operation (information processing flow) of the measurement recording device 10 according to the present embodiment.
[0093] Note that the flow shown in FIG. 10 is merely an example, and unnecessary steps may be deleted, new steps may be added, or the execution order of the steps may be changed without departing from the spirit of the present embodiment.
[0094] First, when the measurer operates the terminal 12 to start the measurement application, the operation is started using this as a trigger. In each step (process), the processor 21 of the computer constituting the measurement recording device 10 executes the process corresponding to each step.
[0095] Next, the processor 21 receives measurement information from the terminal 12 held by the measurer (step ST001). After that, based on the received measurement information, the processor 21 selects an image to be displayed from the images of houses and the like stored in the storage 23, and causes the selected image to be displayed on the screen of the display provided in the terminal 12 (step ST002).
[0096] In step ST002, the positions of a plurality of measurement locations are shown in the image of a house or the like displayed on the screen. By looking at the image displayed on the screen, the measurer can grasp the positions of the plurality of measurement locations shown in the image. Also, in step ST002, each measurement location is displayed in a display mode in which it is possible to identify whether it has been measured or not. Therefore, the measurer can easily recognize which measurement locations have been measured.
[0097] Then, among the plurality of measurement locations, the measurer designates any one of the unmeasured measurement locations as the designated measurement location, and the terminal 12 transmits the information of the designated measurement location designated by the measurer to the measurement recording device 10. On the measurement recording device 10 side, the processor 21 receives the information of the designated measurement location sent from the terminal 12 (step ST003), and specifies the designated measurement location designated by the measurer based on the received information (step ST004).
[0098] After that, the measurer performs a measurement on the designated measurement location, speaks the content corresponding to the measurement result, and inputs the voice with the terminal 12. The input voice data is transmitted from the terminal 12 to the measurement recording device 10. On the measurement recording device 10 side, the processor 21 receives the above voice data, thereby accepting the input of the voice data (step ST005).
[0099] Next, the processor 21 performs voice recognition on the voice data accepted in step ST005 to convert the voice into text (step ST006). Also, filler correction processing and misrecognition correction processing are appropriately applied to the voice recognition text.
[0100] When the text information is obtained in step ST006, the processor 21 divides the text information into two or more words, and assigns a label corresponding to each word from among a plurality of labels to each word (step ST007).
[0101] And, when the registered expression is included in the text information obtained in step ST006, the processor 21 extracts, from the text information, a phrase corresponding to the registered expression based on the labels assigned to the respective words in step ST007 (step ST008).
[0102] Specifically, among two or more words included in the text information, the corresponding word that has a predetermined order relationship with the word of the corresponding label and to which a label other than the exclusion label is assigned is extracted as a phrase corresponding to the registered expression. When there are two or more corresponding words, a group of corresponding words having a certain meaning and coherence is extracted as a phrase corresponding to the registered expression.
[0103] Also, when two or more registered expressions are included in the text information, in step ST008, the processor 21 extracts, from the text information, phrases corresponding to the registered expressions for each registered expression.
[0104] Next, the processor 21 records the measurement result information based on the phrases extracted in step ST008 in association with the measurement location (step ST009). In step ST009, the measurement result information is recorded in association with the most recently specified designated measurement location.
[0105] Also, when, in step ST008, phrases corresponding to the registered expressions are extracted for each registered expression from the text information, in step ST009, the measurement result information based on the phrases extracted for each registered expression is recorded in association with the measurement location (strictly speaking, the designated measurement location).
[0106] Among the above steps, steps ST003 to ST009 are repeatedly executed until measurement result information is recorded for all of the plurality of measurement locations (step ST010). Then, when the measurement result information has been recorded for all of the plurality of measurement locations, the operation ends.
[0107] <Regarding the effects produced by the embodiments described above> Next, examples of the effects produced by the embodiments described above are shown. In the following description, although the effects are described based on the specific configurations shown in the embodiments described above, within the range where the same effects are produced, they may be replaced with other specific configurations shown in the present specification. That is, hereinafter, for convenience, only one of the corresponding specific configurations may be representatively described, but the representatively described specific configuration may be replaced with other corresponding specific configurations.
[0108] According to the embodiments described above, the measurement recording device includes a voice recognition unit 42, a misrecognition correction unit 44, an extraction unit 46, and a recording unit 48. The voice recognition unit 42 performs voice recognition on voice data 100 including measurement information and outputs a voice recognition text. The misrecognition correction unit 44 corrects a misrecognition text, which is text generated by misrecognition in the voice recognition text. The extraction unit 46 extracts, as extraction text, text corresponding to a registered expression, which is an expression registered in advance according to the measurement content, in the voice recognition text. The recording unit 48 records the extraction text in association with each corresponding measurement content. The misrecognition correction unit 44 corrects the misrecognition text based on a correction database in which the misrecognition text and a corrected text, which is the corrected text corresponding to the misrecognition text, are recorded in association with each other. The extraction unit 46 extracts the extraction text from the voice recognition text in which the misrecognition text has been corrected.
[0109] According to such a configuration, by extracting the text corresponding to the registered expression and recording it in association with the measurement content, even for voice data by a measurer with low skill level, the voice related to the measurement content can be accurately extracted and recorded in association with the corresponding measurement content. Therefore, voice input regarding the measurement result becomes easy, and the work load of the input can be reduced.
[0110] When uttering the content of the measurement result, the measurer usually utters the registered expression and then utters the content of the measurement result in sequence. As a result, the voice recognition text obtained by converting the measurer's voice into text includes the registered expression.
[0111] Then, the registered expression included in the voice recognition text is specified, and the text (phrase) corresponding to the registered expression, for example, the text (phrase) immediately after the registered expression, is extracted from the voice recognition text. The text (phrase) extracted in this way is likely to be the phrase corresponding to the content of the measurement result. Thereby, information regarding the measurement result can be accurately obtained from the voice recognition text of the voice data input by the measurer, and the measurement result can be accurately recorded.
[0112] In addition, according to the present embodiment, the versatility regarding the input of the measurement result is improved, and it can flexibly cope with cases where the place where the measurement is carried out is different, the measurement content is different, and the measurement location is different. In any case, the measurement result can be appropriately voice-input and extracted from the voice recognition text.
[0113] In addition, in the present embodiment, when extracting the phrase corresponding to the registered expression from the voice recognition text, for each of two or more words included in the text information, a label corresponding to each word is assigned from among a plurality of labels. Then, based on the labels assigned to the respective words, the phrase corresponding to the registered expression is extracted from the text information. Thereby, the phrase corresponding to the registered expression can be appropriately and easily extracted from the text information.
[0114] Even if the voice spoken by the measurer does not exactly match the registered expression, by assigning a label corresponding to the registered expression to the voice, the phrase corresponding to the text can be appropriately extracted as an indication of the detection result.
[0115] Furthermore, even if two or more measurement results are spoken continuously, by using the position of the registered expression (the order of speech) as a clue, the phrases indicating each measurement result can be appropriately extracted.
[0116] Also, among the plurality of labels, there is an exclusion label. Even if inappropriate words (specifically, fillers or words not related to measurement) are included in the text information as the phrases to be extracted, the exclusion label is assigned to such words, so that the situation where the above words are extracted can be avoided. As a result, the phrases corresponding to the registered expressions can be more appropriately extracted from the text information.
[0117] In addition, when other configurations exemplified in the specification of the present application are appropriately added to the above configuration, that is, even when other configurations in the specification of the present application not mentioned as the above configuration are appropriately added, the same effects can be achieved.
[0118] Also, according to the embodiment described above, in the correction database, a plurality of misrecognized texts and the number of times of such misrecognition are associated and recorded for the correction text.
[0119] The misrecognition correction unit 44 corrects the misrecognized text to the correction text that is most frequently misrecognized in the misrecognized text based on the number of times of misrecognition. According to such a configuration, since the misrecognized text in the speech recognition text is converted to the correction text with the most frequent misrecognition, the speech recognition text can be corrected with high accuracy while reflecting the tendency of misrecognition in a specific measurement.
[0120] Also, according to the embodiments described above, in the correction database, the corrected text, the misrecognized text, and the number of times of misrecognition are recorded in json format. According to such a configuration, while maintaining high extensibility, the relationship between the corrected text and the misrecognized text can be recorded together with the number of times of misrecognition.
[0121] Also, according to the embodiments described above, the apparatus further includes a filler correction unit 50 for correcting filler text, which is text generated by a filler in the speech recognition text, and the extraction unit 46 extracts extraction text from the speech recognition text in which the filler text has been corrected. According to such a configuration, extraction text can be extracted with high accuracy from the speech recognition text from which the filler has been appropriately removed.
[0122] Also, according to the embodiments described above, the filler correction unit 50 corrects the filler text in the speech recognition text by using a learned model obtained by learning the speech recognition text including the filler text with the speech recognition text from which the filler text has been removed as teacher data. According to such a configuration, by performing machine learning using the speech recognition text including the filler and the speech recognition text from which the filler has been removed as teacher data, the filler text can be corrected with high accuracy.
[0123] Also, according to the embodiments described above, the functions of the filler correction unit 50 and the extraction unit 46 are realized by using bidirectional encoder representation from transformers (BERT). According to such a configuration, the accuracy of the assignment process, the extraction process, and the filler correction process can be improved by learning with a bidirectional transformer.
[0124] Also, according to the embodiments described above, when the measurement information included in the speech recognition text does not correspond to the pre-registered measurement content, the function of the filler correction unit 50 is realized using bidirectional and auto-regressive transformers (BART). With such a configuration, it is possible to correct the filler text with high accuracy while avoiding problems such as deleting speech recognition text that is not registered in the filler correction process.
[0125] According to the embodiments described above, the measurement recording program is a measurement recording program having a plurality of computer-executable instructions for being executed by one or more processors. In the measurement recording program, the computer is caused to perform the following operations by a plurality of instructions executed by the processor. First, the computer is caused to perform speech recognition on the speech data 100 including measurement information and output a speech recognition text. Then, the computer is caused to correct the misrecognition text, which is the text generated by misrecognition in the speech recognition text. Then, the computer is caused to extract, as extraction text, the text corresponding to the registered expression, which is the expression pre-registered according to the measurement content, in the speech recognition text. Then, the computer is caused to record the extraction text in association with each corresponding measurement content. Here, correcting the misrecognition text means correcting the misrecognition text based on the correction database in which the misrecognition text and the correction text, which is the corrected text corresponding to the misrecognition text, are associated and recorded. Also, extracting the extraction text means extracting the extraction text from the speech recognition text in which the misrecognition text has been corrected.
[0126] With such a configuration, by extracting the text corresponding to the registered expression and recording it in association with the measurement content, it is possible to accurately extract the speech regarding the measurement content even from the speech data by a measurer with low skill level and record it in association with the corresponding measurement content. Therefore, voice input regarding the measurement result becomes easy, and the work burden of the input can be reduced.
[0127] In addition, when at least one of the other configurations exemplified in the specification of the present application is appropriately added to the above configuration, that is, even when other configurations exemplified in the specification of the present application that are not mentioned as the above configuration are appropriately added, the same effects can be obtained.
[0128] Further, the above program may be recorded on a computer-readable portable recording medium (non-transitory recording medium) such as a magnetic disk, a flexible disk, an optical disk, a compact disk, a Blu-ray disk (registered trademark), or a DVD. And a portable recording medium on which the program for realizing the above functions is recorded may be commercially distributed.
[0129] According to the embodiment described above, in the measurement recording method, voice data 100 including measurement information is voice-recognized to output a voice recognition text. Then, a misrecognition text, which is text generated by misrecognition in the voice recognition text, is corrected. And text corresponding to a registered expression, which is an expression registered in advance according to the measurement content, in the voice recognition text is extracted as extraction text. And the extraction text is associated and recorded for each corresponding measurement content. Here, the step of correcting the misrecognition text is a step of correcting the misrecognition text based on a correction database in which the misrecognition text and a correction text, which is the corrected text corresponding to the misrecognition text, are associated and recorded. Also, the step of extracting the extraction text is a step of extracting the extraction text from the voice recognition text in which the misrecognition text has been corrected.
[0130] According to such a configuration, by extracting text corresponding to the registered expression and recording it in association with the measurement content, even for voice data by a measurer with low skill level, the voice regarding the measurement content can be accurately extracted and recorded in association with the corresponding measurement content. Therefore, voice input regarding the measurement result becomes easy, and the work burden of the input can be reduced.
[0131] In addition, when there are no special restrictions, the order in which each process is performed can be changed.
[0132] Also, when other configurations exemplified in the present specification are appropriately added to the above configuration, that is, even when other configurations in the present specification that are not mentioned as the above configuration are appropriately added, the same effects can be achieved.
[0133] <Regarding the modifications of the embodiments described above> In the embodiments described above, the dimensions, shapes, relative arrangement relationships, or implementation conditions of each component may be described, but these are all examples in all aspects and are not limiting.
[0134] Therefore, countless modifications and equivalents that are not exemplified are assumed to be within the scope of the technology disclosed in the present specification. For example, it is assumed to include cases where at least one component is modified, added, or omitted.
Explanation of Reference Numerals
[0135] 10 Measurement recording device 21 Processor 42 Voice recognition unit 44 Misrecognition correction unit 46 Extraction unit 48 Recording unit 50 Filler correction unit 100 Voice data
Claims
1. A voice recognition unit that performs voice recognition on voice data including measurement information and outputs a voice recognition text; A misrecognition correction unit for correcting misrecognition text, which is text generated by misrecognition in the voice recognition text; An extraction unit for extracting, as extraction text, text corresponding to a registered expression that is a pre-registered expression according to the measurement content in the voice recognition text; And a recording unit for associating and recording the extraction text for each corresponding measurement content; The misrecognition correction unit corrects the misrecognition text based on a correction database in which the misrecognition text and correction text, which is the corrected text corresponding to the misrecognition text, are associated and recorded; The extraction unit extracts the extraction text from the voice recognition text in which the misrecognition text has been corrected; A measurement recording device.
2. The measurement recording device according to claim 1, In the correction database, a plurality of the misrecognition texts and the number of times of misrecognition thereof are associated and recorded with respect to the correction text; The misrecognition correction unit corrects the misrecognition text to the correction text that is most frequently misrecognized for the misrecognition text based on the number of times of misrecognition; A measurement recording device.
3. The measurement recording device according to claim 2, In the correction database, the correction text, the misrecognition text, and the number of times of misrecognition thereof are recorded in json format; A measurement recording device.
4. The measurement recording device according to any one of claims 1 to 3, Further comprising a filler correction unit for correcting filler text, which is text generated by a filler in the voice recognition text; The extraction unit extracts the extraction text from the speech recognition text in which the filler text has been corrected. Measuring and recording device.
5. The measuring and recording device according to claim 4, The filler correction unit corrects the filler text in the speech recognition text by using a learned model obtained by learning the speech recognition text including the filler text with the speech recognition text from which the filler text has been removed as teacher data. Measuring and recording device.
6. The measuring and recording device according to claim 4, The functions of the filler correction unit and the extraction unit are realized by using bidirectional encoder representation from transformers (BERT). Measuring and recording device.
7. The measuring and recording device according to claim 4, When the measurement information included in the speech recognition text does not correspond to the pre-registered measurement content, the function of the filler correction unit is realized by using bidirectional and auto-regressive transformers (BART). Measuring and recording device.
8. A measurement recording program having a plurality of computer-executable instructions for being executed by one or more processors, By the plurality of instructions executed by the processor, cause the computer to perform speech recognition on speech data including measurement information and output a speech recognition text; cause the computer to correct a misrecognition text that is text generated by misrecognition in the speech recognition text; Cause the computer to extract, as extraction text, text corresponding to a registered expression that is a pre-registered expression according to the measurement content in the speech recognition text. Cause the computer to record the extraction text in association with each corresponding measurement content. Correcting the misrecognized text is to correct the misrecognized text based on a correction database in which the misrecognized text and a corrected text that is the corrected text corresponding to the misrecognized text are associated and recorded. Extracting the extraction text is to extract the extraction text from the speech recognition text in which the misrecognized text has been corrected. Measurement recording program.
9. A step of performing speech recognition on speech data including measurement information and outputting speech recognition text; A step of correcting misrecognized text that is text generated by misrecognition in the speech recognition text; A step of extracting, as extraction text, text corresponding to a registered expression that is a pre-registered expression according to the measurement content in the speech recognition text; And a step of recording the extraction text in association with each corresponding measurement content, The step of correcting the misrecognized text is a step of correcting the misrecognized text based on a correction database in which the misrecognized text and a corrected text that is the corrected text corresponding to the misrecognized text are associated and recorded; The step of extracting the extraction text is a step of extracting the extraction text from the speech recognition text in which the misrecognized text has been corrected, Measurement recording method.
Citation Information
Patent Citations
Information processing system, terminal device, server, information processing method and program
JP2017084050A