Data output device and data output method

The data output device addresses the challenge of skill transfer in inspections by converting language-based input into non-language sensory data using a machine learning model and user-specific attributes, enhancing the ability of unskilled workers to recognize abnormalities.

WO2025109884A1PCT designated stage expired Publication Date: 2025-05-30HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/035748
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-10-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing technologies face challenges in effectively transferring skills from skilled workers to unskilled workers in visual and sound inspections, as unskilled workers struggle to interpret language-based information indicating abnormalities in equipment or facilities.

Method used

A data output device and method that generate and output non-language information data, such as sounds or images, based on language input information, using a machine learning model and user-specific attributes to ensure accurate representation of desired sensory data.

Benefits of technology

Enables efficient generation and output of data corresponding to language input, allowing unskilled workers to better understand and recognize abnormalities, thereby facilitating skill transfer and improving inspection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024035748_30052025_PF_FP_ABST
    Figure JP2024035748_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A data output device (100) comprises: an input unit (input device (191)) that acquires input information input by a user as information indicated in a language; a generation unit (113) that generates, as non-language information data on the basis of a user attribute, data corresponding to the input information acquired by the input unit and indicated in the language; and an output unit (output device (192)) that outputs the generated data so as to be sensed by at least one of the five human senses. The data output device (100) may further comprise: an evaluation result acquisition unit (114) that acquires an evaluation result provided by the user with respect to the output result that has been output; and an update unit (115) that repeatedly instructs the generation unit (113) to generate data until the evaluation result with respect to the output result of data corresponding to the input information reflecting the evaluation result becomes an evaluation result indicating that the output result is desired by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Data output device and data output method

[0001] The present invention relates to a data output device and a data output method for generating and outputting data corresponding to input information expressed in language.

[0002] With the declining birthrate and aging population, labor shortages and the transfer of skills from skilled to unskilled workers are becoming an issue. For example, when inspecting equipment and facilities through visual inspections or tapping tests, abnormal conditions or signs of abnormalities may be indicated through language (words, linguistic information, linguistic expressions) such as letters or speech. However, it is difficult for unskilled workers who have little experience witnessing abnormalities in equipment / facilities to determine whether an abnormality is present in practice based on linguistic information.

[0003] An elevator inspection system described in Patent Document 1 is a technology related to inspection work using sound. This elevator inspection system inspects elevators and includes: a data processing unit that classifies the type of sound generated in the elevator while the inspection is being performed; an input unit that accepts input of a period during which the sound occurred, information about the elevator, and onomatopoeia that expresses the sound; a search unit that searches a search database for maintenance information that indicates the need for maintenance of the elevator based on the period during which the sound occurred, the information about the elevator, the type of sound, and the onomatopoeia; and a display unit that displays the maintenance information.

[0004] JP 2022-128764 A

[0005] An elevator inspection system searches for and displays maintenance information based on onomatopoeia that indicate abnormal sounds. However, even for the same abnormal sound, the onomatopoeia that indicates it varies depending on the inspector, and appropriate maintenance information may not be displayed. While image data and sound data that serve as samples of abnormal conditions or signs of abnormalities can be useful, there is little accumulated data of this kind, and it is difficult to collect them. It is desirable to be able to refer to sound and visual data that indicate or are signs of abnormalities in some way. This applies not only to sound and visual appearance, but also to smell, taste, touch, vibration, and other elements that are perceived by at least one of the five human senses. The present invention has been made in light of this background, and aims to provide a data output device and a data output method that enable the generation of non-verbal information data desired by a user based on linguistic information.

[0006] In order to solve the above-mentioned problems, the data output device of the present invention comprises an input unit that acquires input information input by a user as information expressed in language, a generation unit that generates data corresponding to the input information acquired by the input unit and expressed in language as non-verbal information data based on user attributes that are attributes of the user, and an output unit that outputs the generated data so that it can be perceived by at least one of the five human senses.

[0007] According to the present invention, it is possible to provide a data output device and a data output method that enable the generation of non-verbal information data desired by a user based on linguistic information. Problems, configurations, and effects other than those described above will become clear from the description of the following embodiments.

[0008] FIG. 1 is a flowchart of a procedure for using a data output device according to the first embodiment. FIG. 2 is a functional block diagram of a data output device according to the first embodiment. FIG. 3 is a diagram for explaining data output processing according to the first embodiment. FIG. 4 is a flowchart of data output processing according to the first embodiment. FIG. 5 is a functional block diagram of a data output device according to the second embodiment. FIG. 6 is a diagram for explaining data output processing according to the second embodiment. FIG. 7 is a functional block diagram of a data output device according to the third embodiment. FIG. 8 is a diagram for explaining user attribute update processing according to the third embodiment. FIG. 9 is a flowchart of user attribute update processing according to the third embodiment. FIG. 10 is a functional block diagram of a data output device according to the fourth embodiment. FIG. 11 is a functional block diagram of a data output device according to the fifth embodiment. FIG. 12 is a flowchart of inspection work support processing according to the fifth embodiment. FIG. 13 is a diagram for explaining a procedure for generating an image of a rusted handle according to the sixth embodiment. FIG. 14 is a functional block diagram of a data output device according to the sixth embodiment. FIG. 15 is a diagram for explaining a data generation model according to a modified example of the sixth embodiment. FIG. 16 is a hardware configuration diagram showing an example of a computer that realizes the functions of the data output device according to the above-mentioned embodiments.

[0009] Overview of Data Output Device A data output device according to a mode (embodiment) for carrying out the present invention will be described below. The data output device uses a machine learning model to generate and output data corresponding to input information expressed in a language including onomatopoeia and mimetic words. In other words, the data output device acquires linguistic information, converts it into non-linguistic information based on user attributes, and outputs it. The data output device receives a user's evaluation result regarding the output result in which data corresponding to the input information is output, generates data corresponding to the input information that reflects the evaluation result, and outputs it again. The data output device repeats the process of receiving the evaluation result and outputting the data until the data desired by the user is obtained. Next, the data output device adjusts the parameters of the machine learning model so that the data corresponding to the initial input information becomes the data desired by the user. The adjusted parameters (user attributes) will differ for each user.

[0010] With such a data output device, for example, even if the input information and evaluation results differ for each user (inspector), it is possible to generate sound data indicating an abnormality or a sign of an abnormality that the user desires. By using a data output device with parameters adjusted for each user, it becomes possible to efficiently generate data with the sound and appearance that the user desires (reducing the repetition of data output and output evaluation) in response to the language (input information) indicated by the user. By using the data output device, an expert can generate data on sounds that actually indicate an abnormality or sounds that are signs of an abnormality, enabling the transfer of skills from an expert to a non-expert.

[0011] The following describes a data output device related to the inspection of facilities and equipment. A user inputs sounds indicating abnormalities or signs of abnormalities (sounds from a hammering test or sounds during operation) as text or voice via a keyboard or other device into a microphone. The data output device generates and outputs sound data corresponding to the input information, which is linguistic information (text or voice), or image data indicating the appearance (discoloration, rust, deformation, etc.). Data output devices that generate data in output formats (modalities) such as smell, taste, touch, and vibration, or data output devices that generate data in multiple modalities, can also be implemented in a similar manner. Note that text input can also be performed by converting voice input via a microphone into text.

[0012] <First Embodiment: Procedure for Using Data Output Device> Before describing the configuration of the data output device 100 according to the first embodiment (see FIG. 2 described later), we will explain the procedure for using the data output device 100. Fig. 1 is a flowchart of the procedure for using the data output device 100 according to the first embodiment.

[0013] In step S11, the user (inspector) inputs a linguistic expression of a desired sound to the data output device 100. For example, the user may input a sound indicating an abnormality in a facility / equipment as a "clanging sound." In response to this input, the data output device 100 outputs a sound corresponding to the input. The input here is linguistic information, such as text information manually entered by the user or audio information uttered by the user. The output here is non-linguistic information, such as a sound from a speaker that can be perceived by at least one of the five human senses.

[0014] In step S12, the user evaluates the output sound. For example, if the output sound is low, the user evaluates that the desired sound is higher pitched. If the output sound is sufficiently similar to the desired sound, the user evaluates it as OK. In step S13, if the evaluation result of step S12 is OK (step S13 → YES), the user proceeds to step S14; if not (step S13 → NO), the user proceeds to step S15.

[0015] In step S14, the user inputs "OK" to the data output device 100. In step S15, the user inputs the evaluation result (e.g., "higher sound") to the data output device 100, and the process returns to step S12. The user repeats the evaluation of the sound output by the data output device 100 (see step S12) until the desired sound is output.

[0016] First Embodiment: Configuration of Data Output Device Fig. 2 is a functional block diagram of a data output device 100 according to the first embodiment. The data output device 100 is a computer, and includes a control unit 110, a storage unit 120, and an input / output unit 180. An input device 191 (input unit) such as a keyboard, mouse, or microphone, and an output device 192 (output unit) such as a display or speaker are connected to the input / output unit 180. The input / output unit 180 may include a communication device, enabling data transmission and reception with other input devices or output devices. Note that the data output device 100 may include the input device 191 or the output device 192.

[0017] First Embodiment: Storage Unit The storage unit 120 is configured to include storage devices such as a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD), etc. The storage unit 120 stores a data generation model 130, a user attribute database 140, a user information database 150, and a program 128. The program 128 includes a description of a data output process (see FIG. 4 ), which will be described later.

[0018] First Embodiment: Data Generation Model The data generation model 130 is a machine learning model, such as a neural network or a deep learning model. The data generation model 130 uses input information represented (expressed in language, for example, text) as an explanatory variable and sound data corresponding to the input information as a target variable. Such a data generation model 130 can be generated by training (learning) using learning data in which, for example, text is used as an explanatory variable and sound data corresponding to the text is used as a correct answer label. Alternatively, the data generation model 130 may be a machine learning model using a generation AI technology that uses prompt information, which is text, as input and generates sound data or image data corresponding to the prompt information.

[0019] First Embodiment: User Attribute Database / User Information Database The user attribute database 140 stores the user (inspector) of the data output device 100 and the user attributes of the user in association with each other. Note that the user attributes are the correspondence between the sound desired by the user and the linguistic expression that represents the sound. In the first embodiment, the user attributes are parameters of the data generation model 130 (see parameters 621 and 622 in FIG. 3 described below), such as the weights of a neural network.

[0020] The user information database 150 stores users and their user information in association with each other. Examples of user information include gender, age, years of inspection experience, and equipment / devices with inspection experience. User information may also include the results of sensory tests of the five senses, such as hearing diagnostic test results and visual test results (e.g., color vision test results).

[0021] <<First Embodiment: Control Unit>> The control unit 110 is configured to include a CPU (Central Processing Unit) and is equipped with an input acquisition unit 111, an attribute acquisition unit 112, a generation unit 113, an evaluation result acquisition unit 114, and an update unit 115. The control unit 110 may be configured to include a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), etc. FIG. 3 is a diagram for explaining data output processing according to the first embodiment. The processing of each functional unit will be explained below with reference to FIG. 3.

[0022] First Embodiment: Input Acquisition Unit The input acquisition unit 111 acquires, as input information 611, a linguistic expression (linguistic information) input via the input device 191. For example, the input acquisition unit 111 acquires, as the input information 611, text input using a keyboard or a "clanging sound" that is a voice input (uttered by a user) using a microphone.

[0023] First Embodiment: Attribute Acquisition Unit The attribute acquisition unit 112 acquires parameters 621 of the data generation model 130, which are user attributes related to the user, from the user attribute database 140. In the case of a new user, the specified parameters 621 of the data generation model 130 are acquired as the user attributes.

[0024] First Embodiment: Generation Unit The generation unit 113 generates sound data 631 (non-verbal information) corresponding to input information 611 using the data generation model 130. Although the data generation model 130 itself is data, it may be described as a functional unit. For example, it may be described as the data generation model 130 outputting the sound data 631 corresponding to the input information 611.

[0025] As described above, the data output device 100 includes an input unit (see input device 191) that acquires input information input by a user as information expressed in language. The data output device 100 also includes a generation unit 113 that generates data (see sound data 631) corresponding to the input information 611 acquired by the input unit and expressed in language as non-verbal information data (see sound data 631) based on user attributes (see parameters 621) that are attributes of the user. The data output device 100 also includes an output unit (see output device 192) that outputs the generated data in a manner that can be perceived by at least one of the five human senses. The generation unit 113 uses the input information 611 expressed in language as an explanatory variable and generates data corresponding to the input information 611 using a machine learning model (see data generation model 130) that has been trained to output data corresponding to the input information 611. The user attributes are parameters 621 and 622 of the machine learning model.

[0026] First Embodiment: Evaluation Result Acquisition Unit The evaluation result acquisition unit 114 acquires a user's evaluation result 612 for the sound data 631 output from the output device 192. If the sound the user desires is a clang clang sound that is higher in pitch than the output sound, the evaluation result acquisition unit 114 acquires the evaluation result 612 of "higher pitch" via the input device 191.

[0027] First Embodiment: Update Unit The update unit 115 updates parameters 621, which are user attributes, in the following manner and stores the updated parameters in the user attribute database. First, the update unit 115 generates input information 613, which is linguistic information, by reflecting an evaluation result 612 on initial input information 611 input from the input device 191. For example, the update unit 115 generates input information 613, which is "a higher-pitched clang sound," by reflecting "a higher pitch" in "a clang sound."

[0028] Next, the update unit 115 instructs the generation unit 113 to acquire sound data 632 corresponding to the input information 613 and output it from the output device 192. Subsequently, the update unit 115 instructs the evaluation result acquisition unit 114 to acquire the user's evaluation result 614. The update unit 115 repeats the generation of the input information 613, the acquisition and output of the sound data 632, and the acquisition of the evaluation result 614 until sound data corresponding to the first input information 611 desired by the user is generated (see steps S12, S13, and S15 in FIG. 1 ).

[0029] The following description will be given assuming that sound data 632 is sound data (e.g., a hammering test sound indicating an abnormality) corresponding to the initial input information 611 desired by the user. The update unit 115 adjusts the parameter 621 so that sound data 633 corresponding to the initial input information 611 generated by the generation unit 113 becomes sound data 632, and stores the adjusted parameter 622 in the user attribute database 140. Here, "sound data 633 becomes sound data 632" means that the difference between the sound data 632 and 633 as sound data is equal to or less than a predetermined value. An example of a method for adjusting the parameter 622 is a gradient method for an optimization problem, which searches for a minimum value of an objective function that uses the parameter 622 as an input parameter and calculates the difference between the sound data 632 and 633 as a function value. The update unit 115 may also adjust the parameter 622 using other methods.

[0030] As described above, the data output device 100 includes an evaluation result acquisition unit 114 that acquires user evaluation results 612, 614 for the output result output by the output unit (see output device 192). The data output device 100 also includes an update unit 115 that repeatedly instructs the generation unit 113 to generate data until the evaluation result 614 of the output result of data corresponding to input information 613, which reflects the evaluation result 612, becomes the user's desired output result. The update unit 115 updates the user attributes (see parameters 622) so that the data (see sound data 633) corresponding to the input information 611 acquired by the input unit (see input device 191) becomes data corresponding to the desired output result (see sound data 632).

[0031] <First Embodiment: Data Output Processing> Figure 4 is a flowchart of the data output processing according to the first embodiment. With reference to Figure 4, the process in which the data output device 100 acquires a linguistic expression of a sound desired by a user and outputs corresponding sound data will be described. In step S21, the input acquisition unit 111 acquires the user's linguistic expression of the data to be output (e.g., "clang clang sound"). Note that step S21 corresponds to step S11 in Figure 1.

[0032] In step S22, the generation unit 113 generates and outputs sound data corresponding to the linguistic expression using the data generation model 130. If the output sound is not the desired sound, the user inputs what kind of sound the desired sound is to the data output device 100 via the input device 191 (see step S15 in FIG. 1). If the output sound is the desired sound, the user inputs "OK" (see step S14). In step S23, the evaluation result acquisition unit 114 acquires the user's evaluation result for the sound data output in step S22. In step S24, if the evaluation result is the desired sound data 632 (step S24 → YES), the update unit 115 proceeds to step S25; if not (step S24 → NO), the update unit 115 proceeds to step S26.

[0033] In step S25, the update unit 115 adjusts the parameters 621 so that the sound data 633 corresponding to the linguistic expression acquired in step S21 (see input information 611 in FIG. 3 ) generated by the generation unit 113 becomes the desired sound data 632. Next, the update unit 115 stores the adjusted parameters 622 in the user attribute database 140 as a user attribute corresponding to the user. In step S26, the update unit 115 generates linguistic information (e.g., "a higher-pitched clang sound") that reflects the evaluation result 612, and returns to step S22.

[0034] <<First Embodiment: Features of Data Output Device>> The data output device 100 acquires an initial linguistic expression (input information 611) of a sound desired by the user and an evaluation result 612 for the output sound, and acquires desired sound data 632. Next, the data output device 100 adjusts parameters 621 of the data generation model 130 so that the desired sound data 632 can be output from the initial linguistic expression, and acquires parameters 622 that are the adjustment results.

[0035] Even for the same sound, the linguistic expression for that sound varies from user to user. By using the data output device 100 in which the parameters 622 (user attributes) have been adjusted for each user, it becomes possible to efficiently generate sound data according to the linguistic expression of each individual user (reducing the repetition of steps S12, S13, and S15 shown in FIG. 1).

[0036] Second Embodiment The data output device 100 of the first embodiment adjusts the parameters of the data generation model 130 as user attributes for each user (see step S25 in FIG. 4). In a data generation model 130A of the second embodiment (see FIGS. 5 and 6 described below), latent variables 661 and 662, which are inputs (explanatory variables) instead of parameters, are used as user attributes.

[0037] FIG. 5 is a functional block diagram of a data output device 100A according to the second embodiment. Compared to the data output device 100 according to the first embodiment (see FIG. 2), the update unit 115A and the data generation model 130A are different. The input (explanatory variables) of the data generation model 130A are the linguistic expression of the desired sound and latent variables indicating user attributes, which are correspondences between the sound desired by the user and the linguistic expression indicating the sound. The input of the data generation model 130A may be a distribution of latent variables instead of latent variables. The update unit 115A updates the input latent variables, rather than the parameters 621 of the data generation model 130 (see FIG. 3).

[0038] 6 is a diagram for explaining the data output process according to the second embodiment. The latent variable 661, which is input when the generation unit 113 generates sound data 671 corresponding to initial input information 651, is a default latent variable. In other words, the same sound data 671 is generated for the same input information 651 regardless of the user. The initial latent variable is not limited to a default latent variable, and a random latent variable 661 according to a predetermined probability distribution may also be used.

[0039] As in the first embodiment, the update unit 115A repeats the generation of input information 653, the acquisition and output of sound data 672, and the acquisition of evaluation result 654 until sound data (non-linguistic information) corresponding to the initial input information 651 (linguistic information) desired by the user is generated. When the sound desired by the user is obtained, the update unit 115A adjusts the latent variable 661 so that sound data 673 corresponding to the initial input information 651 generated by the generation unit 113 becomes sound data 672, and acquires latent variable 662 as the adjustment result.

[0040] As described above, the generation unit 113 uses input information 651 expressed in language and latent variables 661, 662 indicating user attributes or the distribution of latent variables as explanatory variables, and generates data corresponding to the input information 651 using a machine learning model (see data generation model 130A) that has been trained to output data corresponding to the input information 651 (see sound data 671).

[0041] Second Embodiment: Features of Data Output Device The data output device 100A may adjust the latent variables 662 that are input to the data generation model 130A, rather than adjusting the parameters 622 of the data generation model 130 for each user. This allows the data output device 100A to efficiently generate sound data that corresponds to the linguistic expressions of each individual user.

[0042] Third Embodiment The data output device 100, 100A first acquires a linguistic expression of a sound desired by the user (see step S21), and then repeatedly evaluates the output sound (see steps S22 to S24, S26). After the desired sound is obtained (see step S24 → YES), the data output device 100, 100A adjusts parameters 621 and latent variables 661, which are user attributes, and acquires parameters 622 and latent variables 662 as the adjustment results (see step S25). Instead, a data output device 100B according to a third embodiment (see FIG. 7, described below) receives the evaluation results for a sample sound and adjusts the user attributes so that the sample sound is output when the evaluation results are input.

[0043] 7 is a functional block diagram of a data output device 100B according to the third embodiment. Compared to the data output device 100 according to the first embodiment, an update unit 115B is different, and sample data 121 is added to the storage unit 120. The sample data 121 is sound data 731, 733, 735 (see FIG. 8 described later) used to adjust user attributes (parameters of the data generation model 130). The user inputs evaluation results 711, 713, 715 that express the output sound of the sound data in language.

[0044] FIG. 8 is a diagram illustrating a user attribute update process according to the third embodiment. The processing details of the update unit 115B will be described with reference to FIG. 8 . Regarding a user's evaluation result 711 for the output sound of sound data 731, the update unit 115B adjusts parameters 721 (user attributes) of the data generation model 130 so that sound data 732 generated by the generation unit 113 when the evaluation result 711 is used as input information 712 becomes sound data 731. For example, if the evaluation result 711 for the output sound of sound data 731 is "clang," the parameters 721 are adjusted so that sound data 732 becomes sound data 731 when "clang" is used as input information 712 (so that the difference as sound data is equal to or less than a predetermined value).

[0045] Thereafter, the updating unit 115B repeats this process. With respect to a user's evaluation result 713 for the output sound of sound data 733, the updating unit 115B adjusts the parameter 722 so that sound data 734 generated by the generation unit 113 when the evaluation result 713 is used as input information 714 becomes sound data 733. Furthermore, with respect to a user's evaluation result 715 for the output sound of sound data 735, the updating unit 115B adjusts the parameter 723 so that sound data 736 generated by the generation unit 113 when the evaluation result 715 is used as input information 716 becomes sound data 735. In this way, the updating unit 115B repeats adjustment of the parameters 721, 722, and 723, which are user attributes. Note that although the updating unit 115B adjusts the parameters of the data generation model 130, it may also adjust latent variables, which are input to the data generation model 130A according to the second embodiment.

[0046] 9 is a flowchart of the user attribute update process according to the third embodiment. In step S31, the update unit 115B starts the process of repeating steps S32 to S34 for each sample data 121. Hereinafter, the target of the repeated process will be referred to as the target sample data (sound data). In step S32, the update unit 115B outputs the target sample data to the output device 192.

[0047] In step S33, the evaluation result acquisition unit 114 acquires the user's evaluation result for the output (output sound) of step S32. In step S34, the update unit 115B adjusts the parameters of the data generation model 130 so that the sound data generated by the generation unit 113 using the evaluation result as input becomes the sample data to be processed (so that the difference as sound data is equal to or less than a predetermined value).

[0048] As described above, the evaluation result acquisition unit 114 acquires information (see evaluation results 711, 713, and 715) in a language corresponding to the output result of the sample data (see sound data 731, 733, and 735). The update unit 115B updates the user attributes (parameters 721, 722, and 723) so that the data corresponding to the information (see sound data 732, 734, and 736) becomes the sample data (see sound data 731, 733, and 735).

[0049] <<Third Embodiment: Features of the Data Output Device>> The data output device 100B adjusts the parameters of the data generation model 130, which are user attributes, without the user having to repeatedly input a linguistic expression of a desired sound or evaluate the output sound as shown in Fig. 1. By adjusting the user attributes using sample data 121 before generating sound data of the desired sound, it is expected that sound data of the desired sound will be generated efficiently.

[0050] Fourth Embodiment A data output device 100B according to the third embodiment adjusts the user attributes of a new user using sample data 121 (see FIG. 7). The user attributes of a new user may be adjusted based on user information. FIG. 10 is a functional block diagram of a data output device 100C according to the fourth embodiment. Compared to the data output device 100 according to the first embodiment, an update unit 115C is different.

[0051] More specifically, the update unit 115C acquires a user whose user attributes have been adjusted and stored in the user attribute database 140. Next, the update unit 115C acquires the user information of the user from the user information database 150. Next, the update unit 115C sets the user attributes of the user corresponding to the user information that is most similar to the user information of the new user as the user attributes of the new user.

[0052] The update unit 115C may set the average value of the user attributes of users corresponding to a predetermined number of pieces of user information similar to the user information of the new user as the user attributes of the new user. Alternatively, the update unit 115C may set the average value of the user attributes weighted according to the similarity as the user attributes of the new user.

[0053] Furthermore, the update unit 115C may perform the user attribute update process (see FIG. 9) according to the third embodiment, using the user attributes thus acquired (parameters of the data generation model 130 and latent variables that are inputs to the data generation model 130A) as initial values. Also, instead of the parameters 621 (see FIG. 3) or the latent variables 661 (see FIG. 6), user attributes based on the user attributes of existing users may be used as initial values ​​of the user attributes in the data output process (see FIG. 4).

[0054] As described above, the update unit 115C calculates the user attributes of a new user based on the user attributes of a user whose user information is similar to the user information of the new user and whose user attributes have been updated (stored in the user attribute database 140).

[0055] <Fourth embodiment: Features of data output device> By using the user attributes of a user who has similar gender, age, years of inspection work experience, and sensory test results of the five senses as the user attributes of a new user, it is expected that the desired sound data can be obtained more efficiently (by reducing the number of evaluations of the output sound) than when using default user attributes.

[0056] Fifth Embodiment Data output devices 100, 100A, 100B, and 100C will be described below as inspection support devices. Fig. 11 is a functional block diagram of a data output device 100D according to the fifth embodiment. Compared to the data output device 100 according to the first embodiment (see Fig. 2), an inspection work support unit 116 is added to the control unit 110, and an inspection work database 160 is added to the storage unit 120.

[0057] The inspection work database 160 stores information related to inspection work on facilities and equipment, such as the name of the facility or equipment and its identification information, the name of the part to be inspected, its location, the inspection content, whether or not there is an abnormality, the content of the abnormality, the inspection date and time, and the inspection worker.

[0058] When an abnormality is found during an inspection, the inspection work support unit 116 stores data indicating the abnormality in the inspection work database 160. The data indicating the abnormality includes abnormal sounds from facilities or equipment, abnormal sounds during a hammering test, image data of an area where an abnormality was found, and the like.

[0059] If no abnormality is found during the inspection, the inspection work support unit 116 receives instructions from the user, the inspection worker, to generate data indicating the abnormality and store the data in the inspection work database 160. For example, the inspection work support unit 116 first accepts "a clanging sound" as a linguistic expression of an abnormal sound in a hammering test, and after repeatedly receiving evaluations from the user, stores sound data of the desired abnormal sound in the inspection work database 160 as an example of an abnormal sound.

[0060] As described above, the input information indicates an abnormality or a symptom of an abnormality in an object (e.g., equipment or machinery, or a building / construction as described below). The data output device 100D includes an inspection work support unit 116 that stores the generated data as abnormality data related to the object.

[0061] 12 is a flowchart of the inspection work support process according to the fifth embodiment. The inspection work support process is a process that is repeated for each inspection work target, such as a facility, a device, or a part thereof. In step S51, the inspection work support unit 116 displays inspection instructions including the name and location of the inspection work target, the inspection content, etc. The user, who is an inspection worker, performs the inspection work according to the inspection instructions and inputs the inspection results.

[0062] In step S52, the inspection work support unit 116 accepts the inspection results entered by the user. In step S53, if the inspection results indicate an abnormality (step S53 → YES), the inspection work support unit 116 proceeds to step S54, and if no abnormality is found (step S53 → NO), the inspection work support unit 116 proceeds to step S55. In step S54, the inspection work support unit 116 stores data indicating the abnormality and the details of the abnormality in the inspection work database 160, and ends the inspection work support process.

[0063] In step S55, the inspection work support unit 116 stores data indicating normality in the inspection work database 160. For example, the inspection work support unit 116 stores image data of a portion of equipment or machinery that is not abnormal. In step S56, the inspection work support unit 116 asks the user whether or not to create an abnormality sample, and if the user replies that they will (step S56 → YES), the process proceeds to step S57. If the user replies that they will not create an abnormality sample (step S56 → NO), the inspection work support unit 116 ends the inspection work support process.

[0064] In step S57, the inspection work support unit 116 generates sample data indicating an abnormality (abnormal data) and stores it in the inspection work database 160. Examples of samples indicating an abnormality include abnormal sounds from equipment or devices, abnormal sounds during a hammering test, and the appearance of an area where an abnormality was found. The inspection work support unit 116 instructs the user to express the sample indicating an abnormality in language. The inspection work support unit 116 executes data output processing (see FIG. 4 ) to obtain the abnormality data desired by the user and stores it in the inspection work database 160.

[0065] Fifth Embodiment: Features of the Data Output Device According to the data output device 100D, for facilities, equipment, and parts thereof for which no abnormalities were found during inspection work, sound and visual data indicating an abnormality are stored in the inspection work database 160. This makes it possible to convey inspection know-how to unskilled personnel who have little experience and have never witnessed an abnormality. Note that, although the inspection work support process (see FIG. 12 ) generates abnormality data during inspection work, this is not limited to this. The data output device 100D may also generate abnormality data for inspection work targets specified by the user after the inspection work is completed.

[0066] Fifth Embodiment: Modification: Data Stability By using the data output device 100D, abnormal data from multiple users (experts) is stored in the inspection work database 160. The data output device 100D may evaluate the quality or stability of the abnormal data as a sample using the difference between the abnormal data for the same abnormality. For example, the inspection work support unit 116 may determine that the smaller the variance of the generated abnormal data, the higher the quality or stability, or that the greater the number of abnormal data, the higher the quality or stability.

[0067] As described above, the inspection operation support unit 116 calculates the quality or stability based on the variance of abnormal data from multiple users.

[0068] Fifth Embodiment: Modification: Acquisition of Abnormality Data The inspection work support unit 116 may search the inspection work database 160 to acquire and output abnormality data for an inspection target designated by the user.

[0069] As described above, the inspection work support unit 116 outputs abnormality data related to a designated object (for example, a facility or machine, or a building / construction as described below).

[0070] Sixth Embodiment In the fifth embodiment, the data output device 100D generates sound and appearance data indicating an abnormality. Instead of the sound and appearance data themselves, related data may also be generated. For example, if rust on a steering wheel is considered abnormal, an image of the rusted steering wheel may be generated and stored instead of generating an image of the rust and storing it in the inspection work database 160.

[0071] 13 is a diagram illustrating the procedure for generating an image 754 of a rusted steering wheel according to the sixth embodiment. When an image of a rusted steering wheel is desired, the user instructs the data output device 100E (see FIG. 14 described later) to generate an image 751 of the rust. Next, the user photographs the steering wheel to obtain an image 752 (auxiliary data). The user then designates an area 753 of the steering wheel where the rust will occur and instructs the image 752 to be combined with the image 751, thereby obtaining an image 754 of the rusted steering wheel.

[0072] FIG. 14 is a functional block diagram of a data output device 100E according to the sixth embodiment. Compared to the data output device 100D (see FIG. 11 ), the data output device 100E differs in that an inspection support unit 116E is included, and a synthesis unit 117 is added to the control unit 110. The process of generating a rust image 751 is similar to the data output process according to the first embodiment (see FIG. 4 ). The data output device 100E generates the rust image 751 desired by the user based on, for example, the initial input information, such as "red rust," and the evaluation result, "black spots mixed in here and there." Note that "red rust" and "black spots mixed in here and there" are information entered by the user as text via a keyboard or information (linguistic information) spoken into a microphone (see evaluation results 612 and 614 in FIG. 3 ).

[0073] Next, the inspection work support unit 116E acquires and displays an image 752 of the handle photographed by the user, and further receives a user-specified designation of a rusted handle region 753. Next, the composition unit 117 pastes the rust image 751 into the handle region 753 to generate an image 754 of the rusted handle (non-verbal information), and stores this in the inspection work database 160.

[0074] Sixth Embodiment: Features of the Data Output Device When an example of an abnormality is a rusty steering wheel, it is considered easier for a non-expert to understand if an image of a rusted steering wheel that is closer to the actual object is used rather than simply using an image of rust. By having an expert user generate such data using the data output device 100E, the transmission of inspection know-how becomes more efficient. While an image is used as an example here, data in other output formats (modalities) may also be used. For abnormal sounds that occasionally or periodically mix with normal operating sounds (surrounding environment), the data output device 100E may generate and acquire the abnormal sound and combine it with the recorded normal operating sounds.

[0075] <<Modification: Data Generation Using Auxiliary Data>> In the sixth embodiment, the data output device 100E synthesizes the rust image 751 and the handle image 752 (auxiliary data) to generate the rusted handle image 754. The generation unit 113F may directly generate the rusted handle image using a data generation model 130F (see FIG. 15 , described later) that has a synthesis function.

[0076] FIG. 15 is a diagram illustrating a data generation model 130F according to a modified example of the sixth embodiment. The input of the data generation model 130F includes input information 761, which is the linguistic expression to be generated, as well as auxiliary data corresponding to a portion of the input information 761. In FIG. 15, the auxiliary data includes an image 762 corresponding to the "rusty steering wheel" in the input information 761, and an area 763 corresponding to the rusted portion. In response to this input, the data generation model 130F generates an image 765 of the steering wheel with rust in the area 763. This data generation model 130F can be configured by combining the data generation model 130 and a machine learning model corresponding to the processing of the synthesis unit 117. The data generation model 130F can also be configured using inpainting technology, which is an image generation AI technology.

[0077] As described above, the inspection work support unit 116 stores the generated data (see image 751) combined with auxiliary data (see image 754), which is at least one of data indicating the object (see image 752) and data indicating the environment around the object (e.g., normal operating sounds of facilities / equipment), as abnormality data related to the object.

[0078] <<Modification: Designation of Output Format>> In the above-described embodiments, the input to the data output devices 100, 100A, 100B, 100C, 100D, and 100E is linguistic information (e.g., "clanging sound") indicating an output that is non-linguistic information (e.g., sound). Input other than the input information 611, 651, and 761 may also include an output format (modality) such as sound or smell.

[0079] The output format is not limited to one, and may be multiple. For example, an image and a tactile sensation may be output for input information such as "rusty steering wheel." When multiple output formats are specified, the generation unit 113 may use multiple data generation models 130 that generate data for each output format to generate data for each output format, or may use a data generation model that generates data in multiple output formats (multimodal).

[0080] <Other Modifications> Although several embodiments of the present invention have been described above, these embodiments are merely illustrative and do not limit the technical scope of the present invention. For example, in addition to sound data and image data of sounds and appearances that indicate / are signs of abnormalities in facilities or devices, data of other modalities such as smell, taste, touch, and vibration can also be generated in a similar manner. Furthermore, the present invention is not limited to sound and image data referenced during inspection, but can also be used to generate sound and image data referenced during inspection before product shipment. The objects of inspection and testing are not limited to facilities and devices, but may also be objects including buildings, bridges, tunnels, and other structures.

[0081] The present invention can take on various other embodiments, and various modifications such as omissions and substitutions can be made without departing from the spirit of the present invention. These embodiments and modifications are included in the scope and spirit of the invention described in this specification, etc., and are also included in the invention described in the claims and their equivalents.

[0082] <Hardware Configuration> The data output devices 100, 100A, 100B, 100C, 100D, and 100E according to the above-described embodiments are realized by a computer 900 having a configuration such as that shown in FIG. 16 . FIG. 16 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the data output devices 100, 100A, 100B, 100C, 100D, and 100E according to the above-described embodiments. The computer 900 includes a CPU 901, a ROM 902, a RAM 903, an SSD 904, an input / output interface 905 (referred to as an input / output I / F (Interface) in FIG. 16 ), a communication interface 906 (referred to as a communication I / F in FIG. 16 ), and a media interface 907 (referred to as a media I / F in FIG. 16 ). The computer 900 may include a hard disk drive (HDD) instead of the SSD 904, or may include an HDD in addition to the SSD 904.

[0083] The CPU 901 operates based on programs stored in the ROM 902 or the SSD 904, and performs control by the control unit 110 in Fig. 2. The ROM 902 stores a boot program executed by the CPU 901 when the computer 900 starts up, programs related to the hardware of the computer 900, and the like. The CPU 901 controls an input device 910 such as a mouse or keyboard, and an output device 911 such as a display or printer, via an input / output interface 905. The CPU 901 acquires data from the input device 910 and outputs generated data to the output device 911 via the input / output interface 905.

[0084] The SSD 904 stores programs executed by the CPU 901 and data used by the programs. The communication interface 906 receives data from other devices (not shown) via a communication network and outputs the data to the CPU 901, and also transmits data generated by the CPU 901 to other devices via the communication network. The media interface 907 reads programs or data stored on a recording medium 912 and outputs the programs to the CPU 901 via the RAM 903. The CPU 901 loads the programs from the recording medium 912 onto the RAM 903 via the media interface 907 and executes the loaded programs. The recording medium 912 may be an optical recording medium such as a DVD (Digital Versatile Disk), a magneto-optical recording medium such as an MO (Magneto Optical Disk), a magnetic recording medium, a conductive memory tape medium, or a semiconductor memory.

[0085] For example, when the computer 900 functions as the data output devices 100, 100A, 100B, 100C, 100D, and 100E according to the above-described embodiments, the CPU 901 of the computer 900 executes a program 128 (see FIG. 2 ) loaded onto the RAM 903 to realize the functions of the data output devices 100, 100A, 100B, 100C, 100D, and 100E. The CPU 901 reads the program from a recording medium 912 and executes it. Alternatively, the CPU 901 may read the program from another device via a communication network, or may install the program 128 from the recording medium 912 onto the SSD 904 and execute it.

[0086] 100, 100A, 100B, 100C, 100D, 100E Data output device 111 Input acquisition unit 112 Attribute acquisition unit 113, 113F Generation unit 114 Evaluation result acquisition unit 115, 115A, 115B, 115C Update unit 116, 116E Inspection work support unit 117 Synthesis unit 121 Sample data 130, 130A, 130F Data generation model (machine learning model) 140 User attribute database 150 User information database 160 Inspection work database 191 Input device (input unit) 192 Output device (output unit) 611, 613, 651, 653, 712, 714, 716, 761 Input information 612, 614, 652, 654, 711, 713, 715 Evaluation results 621, 622 Parameters (user attributes) 631, 632, 633, 671, 672, 673, 731 to 736 Sound data (data) 661, 662 Latent variables (user attributes) 751, 765 Images (data)

Claims

1. A data output device comprising: an input unit that acquires input information input by a user as information expressed in language; a generation unit that generates data corresponding to the input information acquired by the input unit and expressed in language as non-verbal information data based on user attributes that are attributes of the user; and an output unit that outputs the generated data so that it can be perceived by at least one of the five human senses.

2. A data output device as described in claim 1, comprising: an evaluation result acquisition unit that acquires a user's evaluation result regarding the output result output by the output unit; and an update unit that repeatedly instructs the generation unit to generate data until the evaluation result of the output result of data corresponding to the input information reflecting the evaluation result becomes the output result desired by the user.

3. The data output device according to claim 2, wherein the update unit updates the user attributes so that data corresponding to the input information acquired by the input unit becomes data corresponding to the desired output result.

4. The data output device of claim 3, wherein the generation unit uses input information expressed in the language as an explanatory variable and generates data corresponding to the input information using a machine learning model that has been trained to output data corresponding to the input information, and the user attributes are parameters of the machine learning model.

5. The data output device according to claim 3, wherein the generation unit generates data corresponding to the input information using a machine learning model that has been trained to output data corresponding to the input information, using input information expressed in the language and latent variables or a distribution of latent variables indicating the user attributes as explanatory variables.

6. The data output device according to claim 3, wherein the evaluation result acquisition unit acquires information in a language corresponding to the output result of the sample data, and the update unit updates the user attributes so that the data corresponding to the information becomes the sample data.

7. The data output device according to claim 3, wherein the update unit calculates the user attributes of the new user based on the user attributes of a user whose user information is similar to the user information of the new user and who has updated user attributes.

8. The data output device according to claim 2, further comprising an inspection work support unit that stores the generated data, wherein the input information indicates an abnormality in the object or indicates a symptom of an abnormality, as abnormality data relating to the object.

9. The data output device according to claim 8, wherein the inspection work support unit stores data obtained by combining the generated data with auxiliary data, which is at least one of data indicating the object and data indicating the environment around the object, as abnormality data related to the object.

10. The data output device according to claim 8, wherein the inspection work support unit calculates quality or stability based on a variance of the abnormal data of a plurality of users.

11. The data output device according to claim 8, wherein the inspection work support unit outputs abnormality data related to the specified object.

12. A data output method in which a data output device executes the steps of: acquiring input information input by a user as information expressed in language; generating data corresponding to the acquired input information expressed in language as non-verbal information data based on user attributes that are attributes of the user; and outputting the generated data so as to be perceivable by at least one of the five human senses.

Citation Information

Patent Citations

  • Language processing system and language processing method as well as program and recording medium

    JP2002318593A

  • Voice recognition conversation device and voice recognition conversation processing method

    JP2004029804A

  • Electronic comic viewer device, electronic comic browsing system, viewer program and recording medium recording viewer program

    JP2012133662A