Information processing device, input / output device, information processing method, input / output method, program, model, and information processing system
The information processing system generates text information for floor map images based on user intentions, addressing the limitations of existing technologies and enhancing BIM/CIM system utilization.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RICOH CO LTD
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-22
Smart Images

Figure 2026084995000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an input / output apparatus, an information processing method, an input / output method, a program, a model, and an information processing system.
Background Art
[0002] Patent Document 1 discloses an office design support system that enables estimation of the required area of an office and further enables office design corresponding to the latest office form. Patent Document 2 also discloses an office environment analysis apparatus that pays attention to points other than physical aspects and can evaluate an office environment.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, Patent Documents 1 and 2 do not disclose a technique for storing a model generated by executing learning processing using a floor map image and text as learning data, and generating text information based on the floor map image displayed on a display unit and the model.
[0004] The present invention has been made in view of the above, and an object thereof is to provide an information processing apparatus, an input / output apparatus, an information processing method, an input / output method, a program, a model, and an information processing system that can generate text information for a floor map image according to a user's intention.
Means for Solving the Problems
[0005] In order to solve the above-described problems and achieve the object, the present invention includes storage means for storing a model generated by executing learning processing using a floor map image and text as learning data, and text information generation means for generating text information based on a floor map target image displayed on a display unit and the model.
Effects of the Invention
[0006] According to the present invention, it is possible to generate text information for floor map images that aligns with the user's intentions. [Brief explanation of the drawing]
[0007] [Figure 1] Figure 1 is an overall diagram of the information processing system according to this embodiment. [Figure 2] Figure 2 is a hardware configuration diagram of the terminal device and server according to this embodiment. [Figure 3] Figure 3 is a functional block diagram of the information processing system according to this embodiment. [Figure 4] Figure 4 is a conceptual diagram showing an example of a user information management table and an image management table according to this embodiment. [Figure 5] Figure 5 is a sequence diagram showing an example of the text information generation process according to this embodiment. [Figure 6] Figure 6 is a sequence diagram showing an example of the model update process according to this embodiment. [Figure 7] Figure 7 is a flowchart showing an example of various processes according to this embodiment. [Figure 8] Figure 8 is an explanatory diagram of the model update process and text information generation process according to this embodiment. [Figure 9] Figure 9 is another explanatory diagram of the model update process and text information generation process according to this embodiment. [Figure 10] Figure 10 is a diagram illustrating an example of the display process of target image and text information in the information processing system according to this embodiment. [Modes for carrying out the invention]
[0008] The following describes in detail embodiments of the information processing device, input / output device, information processing method, input / output method, program, model, and information processing system with reference to the attached drawings.
[0009] In industries such as civil engineering and construction, BIM is being introduced with the aim of addressing issues such as the declining birthrate and aging population, and improving labor productivity. The implementation of CIM is progressing.
[0010] BIM stands for Building Information Modeling, and it is a solution for utilizing information in all stages of a building project, from design and construction to maintenance and management, by adding attribute data such as cost, finishes, and management information to a three-dimensional digital model of a building created on a computer (hereinafter referred to as the 3D model).
[0011] CIM stands for Construction Information Modeling and is a solution for the civil engineering sector (including infrastructure in general, such as roads, power, gas, and water) that was proposed following the example of BIM, which was being developed in the architectural field. Similar to BIM, it aims to improve the efficiency and sophistication of the entire construction production system by sharing information among stakeholders, primarily using 3D models.
[0012] A crucial aspect of promoting BIM / CIM implementation is how to effectively utilize the BIM / CIM systems that have been constructed.
[0013] Specifically, 3D spaces restored using BIM / CIM can be used not only for design and construction purposes, but also for other tasks such as maintenance and site surveys. In other words, they can be used not only as design drawings, but also for other purposes such as recording information in 3D space and sharing it with others.
[0014] Furthermore, since tasks performed digitally can be recorded as logs, if tacit knowledge can be extracted from these logs, it can be effectively used for transferring skills from experienced personnel to younger ones. This is expected to lead to front-loading of operations and improved human resource development.
[0015] Here, focusing on the transmission of tacit knowledge, not limited to 3D space, even for omnidirectional images, planar images, etc., as described above, it can be said that the transmission of tacit knowledge between different workpieces or between users with different levels of proficiency is an issue.
[0016] In view of the above, the purpose of this embodiment is to achieve the transmission of tacit knowledge regarding images between different workpieces or between users with different levels of proficiency.
[0017] FIG. 1 is an overall configuration diagram of the information processing system according to this embodiment. The information processing system 1 of this embodiment is constructed by a terminal device 10, which is an example of an input / output device, and a server 40.
[0018] The server 40 is an example of an information processing device connected to an input / output device via a network. The terminal device 10 may be a glass device or a wearable device, and a plurality of terminal devices 10 may be provided in the information processing system 1.
[0019] The terminal device 10 and the server 40 can communicate via a communication network 100. The communication network 100 is constructed by the Internet, a mobile communication network, a LAN (Local Area Network), etc. The communication network 100 may include not only wired communication but also networks based on wireless communication such as 3G (3rd Generation), WiMAX (Worldwide Interoperability for Microwave Access), LTE (Long Term Evolution). Further, the terminal device 10 can communicate by means of a short-range communication technology such as NFC (Near Field Communication) (registered trademark).
[0020] FIG. 2 is a hardware configuration diagram of the terminal device and the server according to this embodiment. The hardware configurations of each terminal device 10 are the same, and are indicated by the reference numerals in the 100 series as the terminal device 10. The hardware configurations of each management device are indicated by the reference numerals in the 400 series.
[0021] The hardware configuration of terminal device 10 will be described below, but the hardware configuration of server 40 is the same and will therefore not be described.
[0022] The terminal device 10 is built using a computer and, as shown in Figure 2, includes a CPU (Central Processing Unit) 101, ROM (Read Only Memory) 102, RAM (Random Access Memory) 103, HD (Hard Disk) 104, HDD (Hard Disk Drive) controller 105, display I / F 106, and communication I / F 107.
[0023] Of these components, the CPU 101 controls the overall operation of the terminal device 10. The ROM 102 stores programs used to drive the CPU 101, such as the IPL (Initial Program Loader). The RAM 103 is used as the work area for the CPU 101.
[0024] HD104 stores various data such as programs. HDD controller 105 The read / write operation of various data to HD104 is controlled according to the control of CPU101.
[0025] The display I / F 106 is a circuit that displays images on the display 106a. The display 106a is a type of display unit such as a liquid crystal or organic EL (electroluminescence) that displays various information such as cursors, menus, windows, characters, or images. The communication I / F 107 is an interface used for communication with other devices.
[0026] If the terminal device 10 is a glass device, the terminal device 10 may use a circuit that displays an image on a lens or the like, which is a transmissive reflective material, instead of the display I / F 106.
[0027] Communication I / F107 is, for example, a NIC (Network Interface Card) that supports TCP (Transmission Control Protocol) / IP (Internet Protocol).
[0028] Furthermore, the terminal device 10 is equipped with a sensor I / F 108, an audio input / output I / F 109, an input I / F 110, a media I / F 111, and a DVD-RW (Digital Versatile Disk Rewritable) drive 112.
[0029] The sensor interface 108 is an interface for receiving detection information from various sensors. The sound input / output interface 109 is a circuit that processes the input and output of sound signals between the speaker 109a and the microphone 109b according to the control of the CPU 101. The input interface 110 is an interface for connecting a predetermined input means to the terminal device 10.
[0030] Keyboard 110a is a type of input device equipped with multiple keys for inputting characters, numbers, various instructions, etc. Mouse 110b is a type of input device used for selecting and executing various instructions, selecting processing targets, moving the cursor, and operating on the display screen, etc.
[0031] The media interface 111 controls the reading or writing (storage) of data to or from a recording medium 111a such as flash memory. The DVD-RW drive 112 controls the reading or writing of various types of data to or from a DVD-RW 112a, which is an example of a removable recording medium. Note that it is not limited to DVD-RW; it may also be a DVD-R or other type of disc. Furthermore, the DVD-RW drive 112 may be a Blu-ray drive that controls the reading or writing of various types of data to or from a Blu-ray Disc (registered trademark).
[0032] Furthermore, the terminal device 10 is equipped with a bus line 113. The bus line 113 is an address bus, data bus, etc., for electrically connecting each component such as the CPU 101.
[0033] Furthermore, recording media such as HDs and CD-ROMs on which the above programs are stored can be provided domestically or internationally as program products. The terminal device 10 realizes the information processing method according to the present invention by executing, for example, the program according to the present invention.
[0034] Figure 3 is a functional block diagram of the information processing system according to this embodiment.
[0035] As shown in Figure 3, the terminal device 10 includes a transmitting / receiving unit 11, an input receiving unit 12, a display control unit 13, an audio control unit 14, a processing unit 15, and a storage / reading unit 19. Each of these units is a function or means of functioning, realized by any of the components shown in Figure 2 operating according to instructions from the CPU 101 that follow a program deployed from the HD 104 onto the RAM 103. The terminal device 10 also has a storage unit 1000 constructed from the RAM 103 and HD 104 shown in Figure 2.
[0036] (Functional configuration of each terminal device) Next, we will describe each component of the terminal device 10.
[0037] The transmitting / receiving unit 11 is an example of a transmission means and is implemented by commands from the CPU 101 shown in Figure 2, as well as the communication interface 107, and transmits and receives various data (or information) with other terminals, devices, or systems via the communication network 100.
[0038] The input receiving unit 12 is an example of an input receiving means, and is mainly implemented by commands from the CPU 101 shown in Figure 2, as well as the input I / F 110 and the sound input / output I / F 109, and accepts various inputs from the user using the microphone 109b, keyboard 110a and mouse 110b. Specifically, the input receiving unit 12 is an example of an input receiving means that accepts voice, text or operation to identify the floor map target image on the display screen of a display unit such as the display 106a.
[0039] The display control unit 13 is an example of a display control means and output means, and is implemented by instructions from the CPU 101 shown in Figure 2 and the display I / F 106, and causes various images and screens to be displayed on the display 106a, which is an example of a display unit. Specifically, the display control unit 13 is an example of an output means that outputs text information generated by the text generation unit 45 based on the tacit knowledge model 4004 generated by performing a learning process using the floor map image and text as learning data, and the floor map target image.
[0040] If the terminal device 10 is a glass device, the display control unit 13 displays a virtual image on a lens or other transmissive reflective material instead of the display I / F 106.
[0041] The audio control unit 14 is an example of an audio control means and output means, and is implemented by commands from the CPU 101 shown in Figure 2 and the audio input / output I / F 109, causing the speaker 109a, which is an example of an audio playback unit, to play sound.
[0042] The processing unit 15 is an example of a processing means, and is implemented by instructions from the CPU 101 shown in Figure 2. It performs processes such as converting text information to audio information and converting audio information to text information.
[0043] The storage / reading unit 19 is an example of a storage control means and is executed by instructions from the CPU 101 shown in Figure 2, as well as by the HD 104, media I / F 111, and DVD-RW drive 112. It performs processing such as storing various data in the storage unit 1000, recording media 111a, and DVD-RW 112a, and reading various data from the storage unit 1000, recording media 111a, and DVD-RW 112a.
[0044] <Server Functional Configuration> Server 40 includes a transmitting / receiving unit 41, a screen generation unit 42, a determination unit 43, a specification unit 44, a text generation unit 45, an update unit 46, and a storage / reading unit 49. Each of these units is a function or means of functioning, realized by any of the components shown in Figure 2 operating according to instructions from the CPU 401 following a program deployed from the HDD 404 onto the RAM 403. Server 40 also has a storage unit 4000 constructed from the HDD 404 shown in Figure 2. The storage unit 4000 is an example of a storage means.
[0045] (Configuration of each function of the management server) Next, we will describe the components of Server 40. Server 40 may be configured to distribute its functions across multiple computers. Furthermore, although we will describe Server 40 as a server computer located in a cloud environment, it may also be a server located in an on-premises environment.
[0046] The transmitting / receiving unit 41 is an example of a transmission means and is implemented by commands from the CPU 401 shown in Figure 2, as well as the communication I / F 407, and transmits and receives various data (or information) with other terminals, devices, or systems via the communication network 100.
[0047] The screen generation unit 42 is an example of a screen generation means and is implemented by instructions from the CPU 401 shown in Figure 2, and generates various screens.
[0048] The decision unit 43 is an example of a decision-making mechanism, and is implemented by instructions from the CPU 401 shown in Figure 2, and performs various decisions as described later.
[0049] The identification unit 44 is an example of an identification means, and is implemented by instructions from the CPU 401 shown in Figure 2, and identifies the target image.
[0050] The text generation unit 45 is an example of a text information generation means, and is implemented by instructions from the CPU 401 shown in Figure 2, and generates text information. In other words, the text generation unit 45 is an example of a text information generation means that generates text information such as tacit knowledge comments based on the floor map target image, which is a floor map image displayed on the display 106a (an example of a display unit) of the terminal device 10, and the tacit knowledge model 4004, which will be described later.
[0051] Here, the floor map target image may include images of multiple rooms. In this case, the text generation unit 45 may generate text information such as tacit knowledge comments for each image of each of the multiple rooms. The text generation unit 45 also generates text information such as tacit knowledge comments based on at least one of the floor identification information and numerical information stored in the image management DB 4002. Here, the floor identification information is an example of floor identification information that identifies the type of floor shown in the floor map target image. Here, the numerical information is an example of numerical information relating to the layout of the floor shown in the floor map target image.
[0052] Furthermore, the text generation unit 45 may generate text information such as tacit knowledge text based on the tacit knowledge model 4004, the floor map target image, and the input information. Here, the input information is an example of input information received by the input reception unit 12 (an example of input reception means) of the terminal device 10 when the display screen is displayed on the display 106a of the terminal device 10. In addition, if the input information includes a question, the text generation unit 45 may generate text information that includes the answer to that question.
[0053] The update unit 46 is an example of an update means, implemented by instructions from the CPU 401 shown in Figure 2, and performs various settings and decisions as described later. The update unit 46 is an example of a model update means that updates the tacit knowledge model 1004 using the floor map target image and text data based on at least one of the audio information and character information received by the input reception unit 12 as training data.
[0054] For example, the update unit 46 updates the tacit knowledge model 4004 based on at least one of floor identification information and numerical information. Alternatively, for example, the update unit 46 updates the tacit knowledge model 4004 based on the degree of relevance between the floor map target image and the text data. Alternatively, for example, the update unit 46 may update the tacit knowledge model 4004 based on the degree of relevance between the caption comment describing the content of the floor map target image and the text data.
[0055] The storage / reading unit 49 is an example of a storage control means and is executed by instructions from the CPU 401 shown in Figure 2, as well as by the HDD 404, media I / F 411, and DVD-RW drive 412. It performs processing such as storing various data in the storage unit 4000, recording media 411a, and DVD-RW 412a, and reading various data from the storage unit 4000, recording media 411a, and DVD-RW 412a. The storage unit 4000, recording media 411a, and DVD-RW 412a are examples of storage means.
[0056] The memory unit 4000 contains a user information management DB 4001 composed of user information management tables, an image management DB 4002 composed of image information management tables, a caption model 4003, an implicit knowledge model 4004, and a large-scale language model 4005.
[0057] User information management DB4001 stores and manages various types of information, while image management DB4002 stores and manages various types of images, such as three-dimensional point clouds, three-dimensional images of three-dimensional models, and 360-degree spherical images. Specifically, image management DB4002 may also store floor map target images associated with floor identification information and numerical information.
[0058] The caption model 4003 is a model that is generated by performing a training process using image and caption comment combinations as training data, and then instructs a computer to output caption comments based on the image.
[0059] Here, a caption comment is text data, and is a comment that describes an image, whether it is spoken or written.
[0060] The tacit knowledge model 4004 is generated by performing a learning process using combinations of images and tacit knowledge comments for those images as training data, and is a model that causes a computer to function to output tacit knowledge comments based on images. In other words, the tacit knowledge model 4004 is an example of a model generated by performing a learning process using images such as floor map images and text such as tacit knowledge comments for those images as training data. The memory unit 4000 is an example of a memory means for storing the tacit knowledge model 4004.
[0061] Here, tacit knowledge comments are text data, and are comments expressed in audio or text, excluding caption comments, that is, comments relating to content not represented in the image.
[0062] The large-scale language model 4005 is a computer language model that is generated by performing a training process using a vast amount of unlabeled text as training data, and consists of an artificial neural network with numerous parameters.
[0063] The large-scale language model 4005 can capture much of the syntax and meaning of human language by being sufficiently trained with context-learning techniques, such as next sentence prediction, which understands context by determining whether sentence 1 and sentence 2 are consecutive, and the masked language model, which understands context by masking words in a sentence and predicting the masked words from the words before and after them.
[0064] Figure 4(a) is a conceptual diagram showing an example of a user information management table according to this embodiment.
[0065] The memory unit 4000 has a user information management DB 4001 constructed, which consists of a user information management table as shown in Figure 4(a). In this user information management table, attribute information indicating the user's attributes and whether or not they have permission to use the text generation unit 45 are associated and managed for each user ID. The attribute information includes information such as planning staff, sales staff, contractors, skilled workers, and unskilled workers.
[0066] Figure 4(b) is a conceptual diagram showing an example of an image management table according to this embodiment.
[0067] The memory unit 4000 has an image management DB 4002 constructed, which consists of an image management table as shown in Figure 4(b).
[0068] In this image management table, each image ID is associated with and managed as follows: image information such as 3D images and 360-degree images (e.g., floor map target images), acquisition information identifying the acquisition scene in which the image information was obtained, usage information identifying the usage scene in which the image information or the display screen containing the image information is used, floor identification information, and numerical information 1 to 3. In other words, floor map target images are associated with floor identification information and at least one of the numerical information.
[0069] Acquired information includes the date the image information was acquired, the name of the property indicated by the image information, and the name of the process in which the image information was acquired. Usage information includes the date the image information was used, the name of the property in which the image information was used, and the name of the process in which the image information was used. Multiple pieces of usage information are managed in the image management table.
[0070] Floor identification information is an example of floor identification information that identifies the type of floor shown in the floor map target image containing the image information (e.g., reception room, negotiation corner, meeting room, entrance, reception). Numerical information 1 to 3 are examples of numerical information related to the floor layout shown in the floor map target image. Specifically, numerical information 1 is the value obtained by dividing the target office area by the target number of people (target number of seats), that is, the area per person (m²).2 The value is per person (seat). Numerical information 2 is the total number of seats for work desks, private rooms (e.g., conference rooms, reception rooms), and open meeting areas, divided by the number of people (number of seats). Numerical information 3 is the total storage capacity of desk carts, side desks, storage units, low cabinets, and discard cabinets, divided by the number of people (number of seats).
[0071] Figure 5 is a sequence diagram showing an example of the text information generation process according to this embodiment.
[0072] The input receiving unit 12 of the terminal device 10 receives input operations related to the user's user ID, image ID, and usage information for the input / output screen displayed on the display 106a (step S1). The transmitting / receiving unit 11 transmits the user ID, image ID, and usage information received in step S1 to the server 40, and the transmitting / receiving unit 41 of the server 40 receives the user ID, image ID, and usage information transmitted from the terminal device 10 (step S2).
[0073] Next, the storage and reading unit 49 of the server 40 stores the usage information received in step S2 in the image management DB 4002 in association with the image ID, and also searches the image management DB 4002 using the image ID received in step S2 as a search key to read out the image information, acquisition information, and usage information corresponding to the image ID (step S3).
[0074] The screen generation unit 42 generates a display screen that includes the image information read by the storage / reading unit 49 (step S4).
[0075] Here, in step S1, multiple image IDs are entered, and in step S3, multiple image information can be retrieved by searching the image management DB4002 using the multiple image IDs as search keys.
[0076] Alternatively, in step S1, instead of an image ID, a document ID that identifies the document containing the desired image may be entered, and the image management DB 4002 may manage one or more image information associated with each document ID.
[0077] As a result, in step S2, the transmitting / receiving unit 11 transmits a material ID instead of an image ID, and in step S3, by searching the image management DB 4002 using the material ID as a search key, one or more image information corresponding to the material ID can be retrieved.
[0078] The transmitting / receiving unit 41 transmits display screen information relating to the display screen generated in step S4 to the terminal device 10, and the transmitting / receiving unit 11 of the terminal device 10 receives the display screen information transmitted from the server 40 (step S5). In an alternative configuration, instead of steps S1 to S5, the transmitting / receiving unit 11 of the terminal device 10 may receive information indicating the captured image transmitted from the camera as display screen information.
[0079] The display control unit 13 of the terminal device 10 displays the display screen received in step S5 on the display 106a (step S6). Alternatively, the display control unit 13 may display an image managed by the terminal device 10 as the display screen on the display 106a instead of the display screen received in step S5. The input receiving unit 12 of the terminal device 10 receives input information to be entered by the user in response to the displayed display screen (step S7).
[0080] This input information includes voice information, text information, and operation information entered by the user, and the processing unit 15 may convert the voice information included in the input information into text information. The voice information, text information, and operation information include identifying information that identifies a target image, such as a floor map target image, on the display screen, and questions regarding the target image.
[0081] The transmitting / receiving unit 11 transmits the input information received by the input receiving unit 12 to the server 40, and the transmitting / receiving unit 41 of the server 40 receives the input information transmitted from the terminal device 10 (step S8). Alternatively, the transmitting / receiving unit 11 may transmit to the server 40, along with the input information, display screen information which is a captured image or an image managed by the terminal device 10, and the transmitting / receiving unit 41 of the server 40 may receive the input information transmitted from the terminal device 10 and the display screen information.
[0082] The memory / reading unit 49 searches the user information management DB 4001 using the user ID received in step S2 as a search key to read out the user attributes and access rights associated with the user ID, and the determination unit 43 determines whether the user has access rights to the text generation unit 45 based on the access rights read from the user information management DB 4001 (step S9).
[0083] If it is determined in step S9 that there is an access right, the identification unit 44 identifies the target image among the displayed screens based on the identification information contained in the input information received in step S8 (step S10).
[0084] The text generation unit 45 uses the target image identified in step S10 to acquire tacit knowledge comments based on the tacit knowledge model 4004, and uses the tacit knowledge comments and the questions extracted from the input information received in step S8 to generate text information based on the large-scale language model 4005 (step S11).
[0085] The text generation unit 45 may convert the audio information contained in the input information into text information, and the text information generated by the text generation unit 45 may be either audio information or text information.
[0086] The text generation unit 45 may generate text information without using a question, or it may generate fixed questions internally within the system and use those fixed questions. In this case, the question text is not visible to the user. Alternatively, the text generation unit 45 may generate fixed questions internally within the system, display the fixed questions on the display unit for the user to select, and use the selected question.
[0087] The transmitting / receiving unit 41 transmits the text information generated in step S11 to the terminal device 10, and the transmitting / receiving unit 11 of the terminal device 10 receives the text information transmitted from the server 40 (step S12).
[0088] If the received text information is character information, the display control unit 13 of the terminal device 10 displays the text information on the display 106a, or the processing unit 15 converts the received text information into audio information, and the audio control unit 14 plays the converted text information through the speaker 109a (step S13). If the received text information is audio information, the text information is played through the speaker 109a, or the processing unit 15 converts the received text information into character information and displays the converted text information on the display 106a.
[0089] In the above, if the information processing system 1 is equipped with multiple terminal devices 10, steps S1 and S2, steps S5 and S6, steps S7 and S8, and steps S12 and S13 may each be executed by a different terminal device, or each may be executed by multiple terminal devices.
[0090] Alternatively, the functions of the server 40 in Figure 3 may be integrated into the terminal device 10, and the processing performed by the server 40 in Figure 5 may be executed by the terminal device 10.
[0091] Figure 6 is a sequence diagram showing an example of the model update process according to this embodiment.
[0092] The input receiving unit 12 of the terminal device 10 receives input operations related to the user's user ID, image ID, and usage information for the input / output screen displayed on the display 106a (step S21). The transmitting / receiving unit 11 transmits the user ID, image ID, and usage information received in step S21 to the server 40, and the transmitting / receiving unit 41 of the server 40 receives the user ID, image ID, and usage information transmitted from the terminal device 10 (step S22).
[0093] Next, the storage and reading unit 49 of the server 40 stores the usage information received in step S22 in the image management DB 4002 in association with the image ID, and also searches the image management DB 4002 using the image ID received in step S22 as a search key to read out the image information, acquisition information, and usage information corresponding to the image ID (step S23). The screen generation unit 42 generates a display screen that includes the image information read out by the storage and reading unit 49 (step S24).
[0094] Here, in step S21, multiple image IDs are entered, and in step S23, multiple image information can be retrieved by searching the image management DB 4002 using the multiple image IDs as search keys.
[0095] Alternatively, in step S21, instead of an image ID, a document ID that identifies the document containing the desired image may be entered, and the image management DB 4002 may manage one or more image information associated with each document ID.
[0096] As a result, in step S22, the transmitting / receiving unit 11 transmits a material ID instead of an image ID, and in step S23, by searching the image management DB 4002 using the material ID as a search key, one or more image information corresponding to the material ID can be retrieved.
[0097] The transmitting / receiving unit 41 transmits display screen information relating to the display screen generated in step S24 to the terminal device 10, and the transmitting / receiving unit 11 of the terminal device 10 receives the display screen information transmitted from the server 40 (step S25). In an alternative configuration, instead of steps S21 to S25, the transmitting / receiving unit 11 of the terminal device 10 may receive information indicating the captured image transmitted from the camera as display screen information.
[0098] The display control unit 13 of the terminal device 10 displays the display screen received in step S25 on the display 106a (step S26). Alternatively, the display control unit 13 may display an image managed by the terminal device 10 as the display screen on the display 106a instead of the display screen received in step S25. The input receiving unit 12 of the terminal device 10 receives input information to be entered by the user in response to the displayed display screen (step S27).
[0099] This input information includes voice information, text information, and operation information entered by the user, and the processing unit 15 converts the voice information included in the input information into text information. The voice information, text information, and operation information include identification information that identifies the target image, such as the floor map target image, on the display screen. In addition, the voice information and text information include caption comments that describe the target image, and tacit knowledge comments that relate to content not expressed in the target image.
[0100] The transmitting / receiving unit 11 transmits input information related to the input operation received by the input receiving unit 12 to the server 40, and the transmitting / receiving unit 41 of the server 40 receives the input information transmitted from the terminal device 10 (step S28). Alternatively, the transmitting / receiving unit 11 may transmit to the server 40, along with the input information, display screen information using a captured image or an image managed by the terminal device 10 as the display screen, and the transmitting / receiving unit 41 of the server 40 may receive the input information transmitted from the terminal device 10 and the display screen information.
[0101] The identification unit 44 identifies the target image on the display screen based on the identification information contained in the input information received in step S28 (step S29).
[0102] The determination unit 43 uses the target image identified in step S29 to obtain a caption comment based on the caption model 4003, and determines the degree of relevance between the caption comment and the comment included in the input information received in step S28 (step S30).
[0103] The determination unit 43 may determine the degree of relevance of the entire comment included in the input information received in step S28 with the acquired caption comment, or it may divide the comment included in the input information received in step S28 into multiple parts and determine the degree of relevance of each divided comment with the acquired caption comment.
[0104] The update unit 46 updates the caption model 4003 as training data, using the comments deemed highly relevant in step S30 as caption comments along with the target image identified in step S29, and updates the tacit knowledge model 4004 as training data, using the comments deemed less relevant in step S30 as tacit knowledge comments along with the target image identified in step S29 (step S31).
[0105] In step S31, the storage / reading unit 49 searches the user information management DB 4001 using the user ID received in step S22 as a search key to read out the user attributes associated with the user ID. The update unit 46 then updates the tacit knowledge model 4004 as training data by including the user attributes and the acquired information and usage information read out in step S23 regarding the target image in the tacit knowledge comment.
[0106] In the above, if the information processing system 1 is equipped with multiple terminal devices 10, steps S21 and S22, steps S25 and S26, and steps S27 and S28 may each be executed by a different terminal device, or by multiple terminal devices, and the terminal device 10 shown in Figure 6 may be different from the terminal device shown in Figure 5.
[0107] Alternatively, the functions of the server 40 in Figure 3 may be integrated into the terminal device 10, and the processing performed by the server 40 in Figure 6 may be executed by the terminal device 10.
[0108] Figure 7 is a flowchart showing an example of various processes according to this embodiment.
[0109] Figure 7(a) shows the process corresponding to step S10 in Figure 5 and step S29 in Figure 6.
[0110] The determination unit 43 determines whether the operation information included in the input information received from the terminal device 10 includes an operation to identify a target image such as a floor map target image (step S41). If an operation to identify is included, the identification unit 44 identifies the target image according to the operation to identify (step S42).
[0111] The determination unit 43 determines whether the audio information or text information included in the input information received from the terminal device 10 contains a comment that identifies the target image (step S43). If a comment that identifies the target image is included, the identification unit 44 identifies the target image according to the comment that identifies the target image (step S44). Comments that identify the target image include, for example, comments that identify the position on the display screen such as right or left, or comments that identify the acquisition information of the image.
[0112] If the operation information does not include an operation to identify the target image, and the audio information and text information do not include a comment to identify the target image, the identification unit 44 identifies the entire display screen as the target image (step S45).
[0113] In step S42, if the operation information includes an operation to identify each of the multiple target images, the identification unit 44 may identify each of the multiple target images according to each operation.
[0114] In step S44, if the audio information or text information includes a comment that identifies each of the multiple target images, the identification unit 44 may identify each of the multiple target images according to each comment.
[0115] Furthermore, if the operation information includes an operation to identify a target image, and the audio information or text information includes a comment to identify another target image, the identification unit 44 may identify the target image according to the operation and the comment, respectively.
[0116] Furthermore, in step S45, if the display screen contains text information, the identification unit 44 may identify the portion of the display screen excluding the text information as the target image.
[0117] Figure 7(b) shows the processes corresponding to steps S30 and S31 in Figure 6.
[0118] The determination unit 43 determines the degree of relevance between the caption comments obtained from the caption model 4003 using the target image and the comments included in the input information received from the terminal device 10 (step S51). For example, the determination unit 43 determines that the comments included in the input information have a high degree of relevance to the caption comments if the proportion of comments that match the caption comments is greater than or equal to a predetermined value. Here, a high degree of relevance to the caption comments means that there is a high probability that the comments describe the content of the target image.
[0119] The determination unit 43 may determine the degree of relevance of the entire comment included in the input information received from the terminal device 10 to the acquired caption comment, or it may divide the comment included in the input information received from the terminal device 10 into multiple parts and determine the degree of relevance of each divided comment to the acquired caption comment.
[0120] The determination unit 43 may, for example, determine an object in the target image through image recognition or object detection, and then determine whether the words related to the caption comment representing that object are included in the comments in the input information, or the proportion of the number of words included, as a degree of relevance.
[0121] The update unit 46 updates the tacit knowledge model 4004 as training data, using the comments that were determined to have low relevance in step S51 as tacit knowledge comments, along with the target image, acquired and used information related to the target image, and user attributes (step S52).
[0122] The update unit 46 updates the caption model 4003 as training data, using the comments that were determined to be highly relevant in step S51 as caption comments, along with the target image (step S53).
[0123] Here, updating the tacit knowledge model 4004 and the caption model 4003 is optional. That is, the update unit 46 may, when certain conditions are met, execute step S52 to update the tacit knowledge model 4004 or execute step S53 to update the caption model 4003.
[0124] Figure 7(c) shows the process corresponding to step S11 in Figure 5.
[0125] The text generation unit 45 uses the target image to acquire tacit knowledge comments based on the tacit knowledge model 4004 (step 61), and uses the tacit knowledge comments, the questions extracted from the input information received in step S8, the acquired information and usage information read in step S3 regarding the target image, and the user attributes read in step S9 to generate text information based on the large-scale language model 4005 (step S62).
[0126] Figure 8 is an explanatory diagram of the model update process and text information generation process according to this embodiment.
[0127] Figure 8(a) is an explanatory diagram of the model update process corresponding to steps S26 and S27 in Figure 6, and Figures 7(a) and (b).
[0128] The display control unit 13 of the terminal device 10 displays the display screen 800 received from the server 40 on the display 106a, and the display screen 800 includes an image 1100 and text 1200.
[0129] The input receiving unit 12 of the terminal device 10 receives audio information from the microphone 109b indicating the conversation Q1, A1, Q2, and A2 between user M1 and user M2 as input information to be entered by the user in response to the displayed screen 800.
[0130] The identification unit 44 identifies the image 1100, which is the portion of the display screen 800 excluding the text 1200, as the target image.
[0131] The judgment unit 43 then uses the target image 1100 to determine the degree of relevance between the caption comments obtained from the caption model 4003 and the dialogues Q1, A1, Q2, and A2.
[0132] The update unit 46 updates the tacit knowledge model 4004 by using comments from the dialogues Q1, A1, Q2, and A2 that are judged to have a low degree of relevance as tacit knowledge comments, along with the target image 1100, as training data, and updates the caption model 4003 by using comments that are judged to have a high degree of relevance as caption comments, along with the target image 1100, as training data.
[0133] Figure 8(b) is an explanatory diagram of the text information generation process corresponding to steps S6, S7, and S13 in Figure 5, and Figures 7(a) and (c).
[0134] The display control unit 13 of the terminal device 10 displays the display screen 800 received from the server 40 on the display 106a, and the display screen 800 includes an image 1100 and text 1200.
[0135] The input receiving unit 12 of the terminal device 10 receives audio information indicating questions Q11 and Q12 from user M3 via the microphone 109b as input information that the user inputs to the displayed screen.
[0136] The identification unit 44 identifies the image 1100, which is the portion of the display screen 800 excluding the text 1200, as the target image.
[0137] The text generation unit 45 uses the target image 1110 to acquire tacit knowledge comments based on the tacit knowledge model 4004, and uses the tacit knowledge comments, questions Q11 and Q12, etc., to generate text information related to the answers A11 and A12 to questions Q11 and Q12, respectively, based on the large-scale language model 4005.
[0138] The display control unit 13 of the terminal device 10 displays the text information related to the responses A11 and A12 received from the server 40 on the display 106a.
[0139] Figure 9 is another explanatory diagram of the model update process and text information generation process according to this embodiment.
[0140] Figure 9(a) is an explanatory diagram of the model update process corresponding to steps S26 and S27 in Figure 6, and Figures 7(a) and (b).
[0141] The display control unit 13 of the terminal device 10 displays the display screen 800 received from the server 40 on the display 106a, and the display screen 800 includes a first image 1100A and a second image 1100B.
[0142] The input receiving unit 12 of the terminal device 10 receives character information from the keyboard 110a, which represents comments C1 to C3 by user M4, as input information to be entered by the user on the displayed screen 800.
[0143] Furthermore, the input receiving unit 12 receives operation information from the mouse 110b as input information to be entered by the user on the displayed screen 800, which indicates an operation by user M4 to identify a partial image 1100B1 in the first image 1100B.
[0144] The identification unit 44 identifies the first image 1100B of the display screen 800 as the target image according to the operation information. The identification unit 44 may also identify a partial image 1100B1 as the target image.
[0145] The judgment unit 43 then uses the target image 1100B to determine the degree of relevance between the caption comments obtained from the caption model 4003 and comments C1 to C3.
[0146] The update unit 46 updates the tacit knowledge model 4004 by using the comments C1 to C3 that are judged to have a low degree of relevance as tacit knowledge comments, along with the target image 1100B, etc., as training data, and updates the caption model 4003 by using the comments that are judged to have a high degree of relevance as caption comments, along with the target image 1100B, as training data.
[0147] Figure 9(b) is an explanatory diagram of the text information generation process corresponding to steps S6, S7, and S13 in Figure 5, and Figures 7(a) and (c).
[0148] The display control unit 13 of the terminal device 10 displays the display screen 800 received from the server 40 on the display 106a, and the display screen 800 includes an image 1100.
[0149] User M5 does not input anything to the displayed screen 800, and the input reception unit 12 does not accept any input information from the user to the displayed screen 800. The identification unit 44 identifies the entire displayed screen 800, which is image 1100, as the target image.
[0150] When user M5 performs an operation to identify a partial image on the display screen 1000B, such as the partial image 1100B1 in Figure 9(a), the input receiving unit 12 receives operation information from the mouse 110b as input information, indicating the operation to identify the partial image. In this case, the identification unit 44 identifies the relevant partial image on the display screen 800 as the target image according to the operation information.
[0151] The text generation unit 45 uses the target image 1110 to acquire tacit knowledge comments based on the tacit knowledge model 4004, and uses the tacit knowledge comments and other information to generate text information related to comments C11 to C14 based on the large-scale language model 4005. The text generation unit 45 may also generate text information using preset standard questions.
[0152] The display control unit 13 of the terminal device 10 displays the text information related to comments C11 to C14 received from the server 40 on the display 106a.
[0153] Figure 10 is a diagram illustrating an example of the display processing of target images and text information in the information processing system according to this embodiment. In this embodiment, the display control unit 13 of the terminal device 10 displays a display screen 800 including a floor map target image such as the first image 1100A and a comment C1 (an example of text information) on the display 106a, as shown in Figure 10(a).
[0154] Furthermore, as shown in Figure 10(b), the display control unit 13 can change the ratio between the display area of the floor map target images, such as the first image 1100A, the second image 1100B, and the partial image 1100B1, and the display area of comments C1 and C3 (an example of text information) on the display screen 800. The display control unit 13 may also change at least one of the number, size, and content of the displayed comments C1 and C3 according to the size of at least one of the display area of the floor map target images and the display area of comments C1 and C3.
[0155] For example, the display control unit 13 increases the number of characters in comments C1 and C3 if the display area is large, and decreases the number of characters if the display area is small. Also, if the display control unit 13 has a large display area for the floor map target image, such as an office drawing, or if there are many rooms in the office and the display area for the floor map target image is large, it reduces the display area of comments C1 and C3 in the comment section.
[0156] Thus, according to the information processing system 1 of this embodiment, it is possible to generate text information for a floor map image that aligns with the user's intentions.
[0157] The programs executed by the terminal device 10 and server 40 in this embodiment are provided pre-installed in ROM or the like. The programs executed by the terminal device 10 and server 40 in this embodiment may also be provided as files in an installable or executable format, recorded on a computer-readable recording medium such as a CD-ROM, flexible disk (FD), CD-R, or DVD (Digital Versatile Disk).
[0158] Furthermore, the programs executed by the terminal device 10 and server 40 of this embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Alternatively, the programs executed by the terminal device 10 and server 40 of this embodiment may be provided or distributed via a network such as the Internet.
[0159] The program executed by the terminal device 10 of this embodiment has a modular configuration that includes the above-mentioned parts (transmitting / receiving unit 11, input receiving unit 12, display control unit 13, voice control unit 14, processing unit 15, and storage / reading unit 19). In actual hardware, a processor such as the CPU 101 reads the program from the ROM and executes it, loading the above-mentioned parts onto the main memory, and generating the transmitting / receiving unit 11, input receiving unit 12, display control unit 13, voice control unit 14, processing unit 15, and storage / reading unit 19 on the main memory.
[0160] The program executed by the server 40 in this embodiment has a modular configuration that includes the above-mentioned parts (transmit / receive unit 41, screen generation unit 42, determination unit 43, identification unit 44, text generation unit 45, update unit 46, and storage / read unit 49). In actual hardware, a processor such as the CPU 401 reads the program from the ROM and executes it, loading the above-mentioned parts onto the main memory, and generating the transmit / receive unit 41, screen generation unit 42, determination unit 43, identification unit 44, text generation unit 45, update unit 46, and storage / read unit 49 on the main memory. [Explanation of symbols]
[0161] 1. Information Processing System 10 Terminal devices 11 Transmitter / Receiver 12 Input reception section 13 Display Control Unit 14. Audio Control Unit 15 Processing Unit 19 Memory / readout section 100 Communication Networks 106a Display 40 servers 41 Transmitter / Receiver 42 Screen generation section 43 Judgment Department 44 Specific part 45 Text generation unit 46 Update section 49 Memory / readout section 1000,4000 storage section 4001 User Information Management Database 4002 Image Management Database 4003 Caption Model 4004 Tacit Knowledge Model 4005 Large-scale language models [Prior art documents] [Patent Documents]
[0162] [Patent Document 1] Patent No. 3858569 [Patent Document 2] Patent No. 4500846
Claims
1. A storage means for storing a model generated by performing a training process using floor map images and text as training data, A text information generation means generates text information based on the floor map target image displayed on the display unit and the model, Equipped with an information processing device.
2. The aforementioned floor map image includes images of multiple rooms, The information processing apparatus according to claim 1, wherein the text information generation means generates the text information for each image of each of the plurality of rooms.
3. The information processing apparatus according to claim 1, wherein the floor map target image is associated with at least one of floor identification information that identifies the type of floor shown in the floor map target image, and numerical information relating to the layout of the floor shown in the floor map target image.
4. The information processing apparatus according to claim 1, wherein the text information generation means further generates the text information based on at least one of floor identification information that identifies the type of floor shown in the floor map target image, and numerical information relating to the layout of the floor shown in the floor map target image.
5. The information processing apparatus according to claim 1, wherein the text information generation means generates the text information based on the model, the floor map target image, and input information received by the input receiving means when the display screen is displayed on the display unit.
6. The information processing apparatus according to claim 5, wherein if the input information includes a question, the text information includes an answer to the question.
7. The information processing apparatus according to claim 1, further comprising: a model update means for updating the model using the floor map target image and text data based on at least one of audio information and character information received by the input receiving means as the learning data.
8. The information processing apparatus according to claim 7, wherein the model updating means updates the model based on at least one of floor identification information that identifies the type of floor shown in the floor map target image, and numerical information relating to the layout of the floor shown in the floor map target image.
9. The information processing apparatus according to claim 8, wherein the model updating means updates the model based on the degree of association between the floor map target image and the text data.
10. The information processing apparatus according to claim 9, wherein the model updating means updates the model based on the degree of relevance between the caption comments describing the content of the floor map target image and the text data.
11. An input receiving means that accepts voice, text, or operation to identify the floor map target image among the display screens displayed on the display unit, A model generated by performing a training process using floor map images and text as training data, and an output means that outputs text information generated by a text information generation means based on the floor map target image, An input / output device equipped with the following features.
12. The input / output device according to claim 11, wherein the output means causes a display screen including the floor map target image and the text information to be displayed on the display unit.
13. The output means can change the ratio between the display area of the floor map target image and the display area of the text information on the display screen. The input / output device according to claim 12, wherein the number, size, and content of the displayed text information are changed according to the size of at least one of the display area of the floor map target image and the display area of the text information.
14. An information processing method performed by an information processing device, A step of generating text information is performed based on a model generated by performing a training process using a floor map target image displayed on the display unit and the floor map image and text as training data. Information processing methods including
15. A program that causes a computer to execute a text information generation step, which generates text information based on a model generated by performing a training process using a floor map target image displayed on the display unit, and the floor map image and text as training data.
16. An input / output method performed in an input / output device, An input receiving step that accepts voice, text, or operation to identify the floor map target image among the display screens shown on the display unit, An output step which outputs a model generated by performing a training process using floor map images and text as training data, and text information generated by a text information generation means based on the floor map target image, Input / output methods including
17. An input receiving step that accepts voice, text, or operation to identify the floor map target image among the display screens shown on the display unit, An output step which outputs a model generated by performing a training process using floor map images and text as training data, and text information generated by a text information generation means based on the floor map target image, A program that causes a computer to execute something.
18. A model is generated by performing a training process using a floor map image and tacit knowledge comments for the floor map image as training data, and causes a computer to function to output tacit knowledge comments based on the floor map image, A model that outputs implicit knowledge comments used for generating text information by a text information generation means, based on a floor map target image displayed on the display unit.
19. An information processing system comprising an input / output device and an information processing device capable of communicating with the input / output device, The aforementioned input / output device is An input receiving means that accepts voice, text, or operation to identify the floor map target image among the display screens displayed on the display unit, The system includes a transmission means for transmitting the floor map target image to the information processing device, The aforementioned information processing device is Receiving means for receiving the floor map target image transmitted from the input / output device, A transmission means for transmitting to the input / output device a model generated by performing a training process using floor map images and text as training data, and text information generated based on the floor map target image. Equipped with, The aforementioned input / output device is Receiving means for receiving the text information from the information processing device, Output means for outputting the aforementioned text information, An information processing system equipped with the following features.