Information processing device, information processing method, and program
The information processing device and method address the challenge of generating suitable prompts for text-based object detection by acquiring and evaluating visual expression texts to enhance detection accuracy.
Patent Information
- Application Number
- JP2024040386
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-29
Smart Images

Figure 2025140802000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] Image recognition techniques for recognizing (detecting) objects in an image are known (for example, see Patent Document 1). Such techniques are required to recognize (detect) objects with high accuracy. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 05-174147 Summary of the Invention [Problem to be solved by the invention]
[0004] Recently, text-based object detection techniques have been developed that train object detectors to associate objects in images with text prompts that describe the objects. Since the accuracy of object detection in such techniques depends on the prompts, it is desirable to generate more suitable prompts. However, generating suitable prompts places a burden on users.
[0005] The present disclosure has been made in view of the above-mentioned problems, and an exemplary purpose thereof is to provide a technology capable of generating a suitable prompt for performing object detection in an image. [Means for solving the problem]
[0006] An information processing device according to an exemplary aspect of the present disclosure includes: a text group acquisition means for acquiring a visual expression text group including a plurality of texts visually expressing the detection target by referring to input data specifying the detection target; a prompt generation means for generating a prompt by referring to the visual expression text group; and a provision means for providing the prompt generated by the prompt generation means to a detection model that receives the prompt and an image as input and detects the detection target specified by the prompt from the image. It is equipped with:
[0007] An information processing device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring input data specifying a detection target, and a provision means for providing a prompt obtained by referring to the input data to a detection model that receives a prompt and an image as input and detects the detection target specified by the prompt from the image, and the prompt provided by the provision means is a prompt generated by a process including a text group acquisition process for acquiring a visual expression text group including a plurality of texts that visually express the detection target, and a prompt generation process for generating a prompt by referring to the visual expression text group.
[0008] An information processing method according to an exemplary aspect of the present disclosure includes: obtaining a visual representation text group including a plurality of texts visually representing a detection target by referring to input data specifying the detection target; generating a prompt by referring to the visual representation text group; and providing the generated prompt to a detection model that receives the prompt and an image as input and detects the detection target specified by the prompt from the image.
[0009] In addition, the information processing device according to each aspect may be realized by a computer. In this case, a program for realizing the information processing device on a computer by causing the computer to operate as each means provided in the information processing device, and a computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention. [Effects of the Invention]
[0010] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that a suitable prompt for performing object detection in an image can be generated. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a block diagram illustrating a configuration of an information processing device according to the present disclosure. [Figure 2] FIG. 1 is a flow diagram showing the flow of an information processing method according to the present disclosure. [Figure 3] 1 is a block diagram illustrating a configuration of an information processing device according to the present disclosure. [Figure 4] FIG. 1 is a flow diagram showing the flow of an information processing method according to the present disclosure. [Figure 5] 1 is a block diagram illustrating a configuration of an information processing system according to the present disclosure. [Figure 6] FIG. 1 is a diagram for explaining processing by an information processing system according to the present disclosure. [Figure 7] FIG. 1 is a diagram for explaining processing by an information processing system according to the present disclosure. [Figure 8] FIG. 1 is a diagram for explaining processing by an information processing system according to the present disclosure. [Figure 9] FIG. 1 is a diagram for explaining processing by an information processing system according to the present disclosure. [Figure 10] FIG. 1 is a diagram for explaining processing by an information processing system according to the present disclosure. [Figure 11] 1 is a block diagram illustrating a configuration of an information processing system according to the present disclosure. [Figure 12] FIG. 1 is a block diagram illustrating a configuration of a computer that functions as an information processing device according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.
[0013] First Exemplary Embodiment A first exemplary embodiment, which is one example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is the basic form of each exemplary embodiment described later. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referred to in describing this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure to the extent that no particular technical obstacles arise.
[0014] (Configuration of information processing device 1) The configuration of an information processing device 1 according to this exemplary embodiment will be described below with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1 according to this exemplary embodiment. As shown in Fig. 1, the information processing device 1 includes a text group acquisition unit 11, a prompt generation unit 12, and a provision unit 13.
[0015] (Text group acquisition unit 11) The text group acquisition unit 11 refers to input data that specifies a detection target and acquires a visual expression text group that includes multiple texts that visually express the detection target. Here, the input data referenced by the text group acquisition unit 11 includes, for example, a statement (text) for identifying the detection target (object) included in the target image. The text group acquisition unit 11 may be configured to refer to input data that has been acquired in advance and stored in a storage unit (not shown), or may be configured to refer to input data acquired from a user via an input unit (not shown).
[0016] Furthermore, the "text visually expressing the detection target" included in the text group acquired by the text group acquisition unit 11 may be, for example, text that specifies the color, shape, movement, type, and environment of the detection target, but these examples do not limit the present exemplary embodiment. Furthermore, the phrase "visually expressing the detection target" does not limit the present exemplary embodiment and may simply be read as "expressing the detection target." Similarly, the phrase "visually expressing text group" does not limit the present exemplary embodiment and may be expressed as a "text group," an "object expressing text group," or the like.
[0017] As one example, the text group acquisition unit 11 may acquire (generate) a plurality of texts visually expressing the detection target by inputting the input data into a language model that has been trained by machine learning in advance. As another example, the text group acquisition unit 11 may acquire (generate) a plurality of texts visually expressing the detection target by comparing the input data with correspondence information created in advance. However, these examples do not limit the present exemplary embodiment.
[0018] (Prompt generation unit 12) The prompt generation unit 12 generates a prompt by referring to the visual expression text group acquired by the text group acquisition unit 11. Here, the prompt generation process by the prompt generation unit 12 includes the following steps: An evaluation process for evaluating the appropriateness of at least one of the texts included in the visual representation text group. A selection process for selecting one or more texts to be used to generate the prompt from the group of visual representation texts with reference to the result of the evaluation process. However, these examples do not limit the present exemplary embodiment.
[0019] (Provider part 13) The providing unit 13 receives a prompt and an image as input, and provides the prompt generated by the prompt generating unit 12 to a detection model that detects a detection target specified by the prompt from the image. The information processing device 1 may further include a configuration that acquires a detection result output by the detection model and generates output information by referring to the detection result, but this does not limit the present exemplary embodiment.
[0020] (Effects of information processing device 1) As described above, the information processing device 1 according to this exemplary embodiment: Referencing input data that specifies a detection target, and acquiring a visual representation text group that includes a plurality of texts that visually represent the detection target; generating a prompt with reference to the visually represented text; Providing the generated prompt to a detection model that receives a prompt and an image as input and detects a target specified by the prompt from the image. According to the above configuration, a visual expression text group including a plurality of texts visually expressing a detection target is acquired, and a prompt is generated by referring to the visual expression text group, so that a suitable prompt can be generated to be provided to a detection model that detects a detection target from an image based on the prompt. In other words, a suitable prompt can be generated for detecting an object in an image.
[0021] (Flow of information processing method S1) Next, the flow of information processing method S1 according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of information processing method S1. As shown in Fig. 2, information processing method S1 includes a process (step, process) S11 of acquiring a visual representation text group, a process (step, process) S12 of generating a prompt, and a process (step, process) S13 of providing the prompt.
[0022] (Step S11) In step S11, the text group acquisition unit 11 refers to input data specifying the detection target and acquires a visual expression text group including multiple texts that visually express the detection target. The specific processing by the text group acquisition unit 11 has been described above, so a description thereof will be omitted here.
[0023] (Step S12) In step S12, the prompt generation unit 12 generates a prompt by referring to the visual expression text group acquired by the text group acquisition unit 11 in step S11. The specific processing by the prompt generation unit 12 has been described above, and therefore will not be described here.
[0024] (Step S13) In step S13, the providing unit 13 receives the prompt and the image as input, and provides the prompt generated by the prompt generating unit 12 in step S12 to a detection model that detects the detection target specified by the prompt from the image.
[0025] (Effect of information processing method S1) As described above, the information processing method S1 according to this exemplary embodiment: Referencing input data that specifies a detection target, and acquiring a visual representation text group that includes a plurality of texts that visually represent the detection target; generating a prompt with reference to the visually represented text; Providing the generated prompt to a detection model that receives a prompt and an image as input and detects a target specified by the prompt from the image. The information processing method S1 including the above-described processes provides the same effects as the information processing device 1 according to this exemplary embodiment.
[0026] (Configuration of information processing device 2) Next, the configuration of the information processing device 2 according to this exemplary embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the information processing device 2 according to this exemplary embodiment. As shown in Fig. 3, the information processing device 2 includes an acquisition unit 21 and a provision unit 22.
[0027] (Acquisition part 21) The acquisition unit 21 acquires input data specifying a detection target. Here, the input data includes, for example, a statement (text) for identifying a detection target (object) included in a target image. The acquisition unit 21 may be configured to acquire the input data stored in a storage unit (not shown), or may be configured to acquire input data received from a user via an input unit (not shown).
[0028] (Provider 22) The providing unit 22 receives a prompt and an image as input, and provides a prompt obtained by referring to the input data to a detection model that detects a detection target specified by the prompt from the image. Here, the prompt provided to the detection model by the providing unit 22 is, for example, a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process for generating a prompt by referring to the visual representation text group; The prompt is generated by a process including the following: Furthermore, the prompt may be generated by the prompt generating unit 12 included in the information processing device 1 described above.
[0029] The information processing device 2 may further include a configuration for acquiring the detection results output by the detection model and generating output information by referring to the detection results, but this does not limit this exemplary embodiment.
[0030] (Effects of information processing device 2) As described above, the information processing device 2 according to this exemplary embodiment: -Acquire input data that specifies the detection target, A prompt and an image are input, and a detection model that detects a detection target specified by the prompt from the image is provided with a prompt obtained by referring to the input data. In this configuration, The prompts provided to the detection model include: a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process for generating a prompt by referring to the visual representation text group; is a prompt generated by a process that includes According to the above configuration, a suitable prompt is generated by referring to a visual expression text group including a plurality of texts that visually express the detection target. Furthermore, the suitable prompt can be used to perform object detection using a detection model.
[0031] (Flow of information processing method S2) Next, the flow of the information processing method S2 according to this exemplary embodiment will be described with reference to Fig. 4. Fig. 4 is a flow diagram showing the flow of the information processing method S2. As shown in Fig. 4, the information processing method S2 includes a process (step, process) S21 of acquiring input data and a process (step, process) S12 of providing a prompt.
[0032] (Step S21) In step S21, the acquisition unit 21 acquires input data specifying a detection target. The specific processing by the acquisition unit 21 has been described above, and therefore will not be described here.
[0033] (Step S22) In step S22, the providing unit 22 receives a prompt and an image as input, and provides a prompt obtained by referring to the input data acquired in step S21 to a detection model that detects a detection target specified by the prompt from the image. Here, the prompt provided to the detection model by the providing unit 22 in this step is, for example, a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process for generating a prompt by referring to the visual representation text group; The prompt is generated by a process including the steps of: In addition, the prompt may be generated by the information processing method S1 described above.
[0034] (Effect of information processing method S2) As described above, the information processing method S2 according to this exemplary embodiment: -Acquire input data that specifies the detection target, A prompt and an image are input, and a detection model that detects a detection target specified by the prompt from the image is provided with a prompt obtained by referring to the input data. In this configuration, The prompts provided to the detection model include: a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process for generating a prompt by referring to the visual representation text group; is a prompt generated by a process that includes The information processing method S2 including the above-described processes provides the same effects as the information processing device 2 according to this exemplary embodiment.
[0035] Second Exemplary Embodiment A second exemplary embodiment, which is one example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be assigned the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in each drawing referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.
[0036] (Configuration of information processing system 100A) The configuration of an information processing system 100A according to this exemplary embodiment will be described with reference to FIG. 5. FIG. 5 is a block diagram showing the configuration of the information processing system 100A. As shown in FIG. 5, the information processing system 100A includes an information processing device 1A and a plurality of servers 51, 52, 53, etc. connected to the information processing device 1A via a network N. The specific configuration of the network N does not limit this exemplary embodiment, and examples include a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination of these networks. Note that it is not essential for the information processing system 100A to include the plurality of servers 51, 52, 53, etc.; the information processing device 1A may have the functions of these servers, and such a configuration is also included in this exemplary embodiment.
[0037] (server) As shown in FIG. 5, the information processing system 100A includes, for example, a plurality of servers 51, 52, 53, and so on. In the example shown in FIG. 5, the server 51 has a first language model LM1, inputs various data provided from the information processing device 1A to the first language model LM1, and provides output data output by the first language model LM1 to the information processing device 1A. Similarly, the server 52 has a second language model LM2, inputs various data provided from the information processing device 1A to the second language model LM2, and provides output data output by the second language model LM2 to the information processing device 1A. Here, the first language model LM1 and the second language model LM2 are different large-scale language models (LLMs), and, for example, are large-scale language models trained with reference to different training data.
[0038] On the other hand, the server 53 has a detection model DM that takes a prompt and an image as input and detects the detection target specified by the prompt from the image, inputs the image and prompt provided from the information processing device 1A into the detection model DM, and provides the output data (detection results) output by the detection model to the information processing device 1A.
[0039] Although the details of the detection model DM do not limit this exemplary embodiment, as an example, the detection model DM is a model trained to associate objects in an image with text prompts that represent the objects. More specifically, the detection model DM is a model trained with reference to training data that includes multiple pairs of images and text that represents one or more objects included in the images.
[0040] The detection model DM is configured to be able to detect a target object from a target image by, for example, undergoing the above-described learning process and referring to a prompt containing a text expression of the target object.
[0041] (Configuration of information processing device 1A) Next, the configuration of the information processing device 1A according to this exemplary embodiment will be described with reference to Fig. 5. As shown in Fig. 5, the information processing device 1A includes a control unit 10A, a storage unit 20A, a communication unit 30, and an input / output unit 40.
[0042] (Communication unit 30) The communication unit 30 communicates with devices external to the information processing device 1A. As an example, the communication unit 30 communicates with a plurality of servers 51, 52, 53, etc. The communication unit 30 transmits data supplied from the control unit 10A to any one of the plurality of servers 51, 52, 53, etc., and supplies data received from any one of the plurality of servers 51, 52, 53, etc. to the control unit 10A.
[0043] The data transmitted from the communication unit 30 to the servers 51 and 52 includes the following: A prompt for generating a visual representation text group (text group generation prompt), which will be described later; and A prompt to evaluate the visual representation text group (text group evaluation prompt) The data that the communication unit 30 receives from the servers 51 and 52 may include: one or more visual representations of texts output by the language models LM1 and LM2 in response to the text group generation prompts; and Evaluation results of one or more visually expressed texts output by language models LM1 and LM2, referring to the text group evaluation prompts. may be included.
[0044] In addition, the data transmitted from the communication unit 30 to the server 53 includes the following: Target image, and - Prompt for detecting the target from the target image (detection prompt) The data received by the communication unit 30 from the server 53 may include: - Detection results by the detection model DM that references the detection prompt may be included.
[0045] (Input / output section 40) The input / output unit 40 is configured to include at least one of input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel. Alternatively, the input / output unit 40 may be configured to have input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel connected to it. In this configuration, the input / output unit 40 accepts various types of information input to the information processing device 1A from the connected input devices. Furthermore, the input / output unit 40 outputs various types of information to the connected output devices under the control of the control unit 10A. An example of the input / output unit 40 is an interface such as a USB (Universal Serial Bus).
[0046] (Storage unit 20A) The storage unit 20A stores various data referenced by the control unit 10A and various data generated by the control unit 10A. Input data IND Visual representation text group TG Prompt Group PR Target image TIM Detection result DR Output information OUT etc. are stored.
[0047] Here, the input data IND includes, for example, a statement (text) for identifying a detection target (object) included in the target image. The input data IND is, for example, input by a user via the input / output unit 40 and stored in the storage unit 20A. Alternatively, the input data IND is, for example, acquired from another device via the communication unit 30 and stored in the storage unit 20A.
[0048] The visual expression text group TG is a text group including one or more texts acquired (generated) by the text group acquisition unit 11, which will be described later. The visual expression text group TG will be described in detail later.
[0049] The prompt group PRG includes prompts generated by the prompt generator 12. For example, the prompt group PRG may include: · Prompts for generating visual representation text clusters (Text Cluster Generation Prompts TGP), · Prompts for evaluating visually presented text groups (Text Group Evaluation Prompts TEP) - Prompt for detecting the target from the target image (detection prompt DP) These prompts are generated, for example, by the prompt generating unit 12. The prompt group PRG will be described in detail later.
[0050] The target image TIM is an image that is the target of the detection process, and is, for example, an image provided to the server 53. The detection model held by the server 53 detects a detection target from the target image TIM by referring to the target image TIM and the detection prompt DP. Details of the target image TIM will be described later.
[0051] The detection result DR is a detection result output by the detection model DM to which the target image TIM and the detection prompt DP are input. The detection result DR will be described in detail later.
[0052] The output information OUT is output information generated by an output information generation unit 16 (described later) with reference to the detection results DR. As an example, the output information OUT is presented to the user via the input / output unit 40. Details of the output information OUT will be described later.
[0053] (Control unit 10A) As shown in FIG. 5, the control unit 10A includes an acquisition unit 14, a text group acquisition unit 11, a prompt generation unit 12, a provision unit 13, and an output information generation unit 15.
[0054] (Acquisition part 14) The acquisition unit 14 acquires input data IND. The acquisition unit 14 may acquire the input data IND input by a user via the input / output unit 40, or may acquire the input data IND stored in the storage unit 20A. As described above, the input data IND includes, for example, a statement (text) for identifying a detection target (object) included in a target image. Specific examples of the input data IND will be described later. In addition, the acquisition unit 14 acquires, via the communication unit 30, output data output by each of the multiple servers 51, 52, 53, ... included in the information processing system 100A.
[0055] (Text group acquisition unit 11) The text group acquisition unit 11 acquires a visual expression text group TG including a plurality of texts that visually express the detection target, with reference to the input data IND acquired by the acquisition unit 14. As an example, the text group acquisition unit 11 acquires (generates) a plurality of texts that visually express the detection target from the wording (text) included in the input data IND. The plurality of texts constitute the visual expression text group TG.
[0056] As in the first exemplary embodiment, the "text visually expressing the detection target" may be, for example, text that specifies the color, shape, movement, type, and environment of the detection target, but these examples do not limit the present exemplary embodiment. Furthermore, the phrase "visually expressing the detection target" does not limit the present exemplary embodiment and may simply be read as "expressing the detection target." Similarly, the phrase "visually expressing text group" does not limit the present exemplary embodiment and may be expressed as "text group," "object expressing text group," or the like.
[0057] As an example, the text group acquisition unit 11 may provide a text group generation prompt TGP obtained by referring to the input data IND to the language model LM1 or LM2 of the server 51 or 52, and use the text group output by the language model LM1 or LM2 as the visual representation text group TG.
[0058] As an example, when the detection target is designated as "roller" in the input data IND, the text group acquisition unit 11 "List several visual features that can be used to detect a roller in an image." The text group acquisition unit 11 may provide a text group generation prompt TGP such as the following to at least one of the language models LM1 and LM2. In addition, the text group acquisition unit 11 may provide a text group generation prompt TGP such as the following from at least one of the language models LM1 and LM2 as a response to the text group generation prompt TGP: "The following visual characteristics are used to detect rollers: - Cone-shaped, handle, on the ground, with a shadow" and each piece of text included in the answer may be included in the visual representation text group TG.
[0059] The processing by the text group acquisition unit 11 is not limited to the above example. As another example, the text group acquisition unit 11 may generate a visual expression text group TG including a plurality of texts visually expressing the detection target by comparing the input data IND with pre-created correspondence information without using the language models LM1 and LM2. More specific processing by the text group acquisition unit 11 will be described later.
[0060] (Prompt generation unit 12) The prompt generation unit 12 generates a prompt by referring to the visual expression text group TG generated by the text group acquisition unit 11. As an example, the prompt generation unit 12 generates a detection prompt DP for detecting a detection target from a target image TIM by referring to the visual expression text group TG. Here, the prompt generation process by the prompt generation unit 12 is as follows: an evaluation process for evaluating the appropriateness of at least one text included in the visually expressed text group TG; and A selection process for selecting one or more texts to be used to generate a detection prompt DP from the visual representation text group TG with reference to the result of the evaluation process. For example, the text group acquisition unit 11 may include: "Please list only those visual features below that are useful in detecting rollers. - Cone-shaped, handle, on the ground, with a shadow" and provide the text group evaluation prompt TEP to at least one of the language models LM1 and LM2. Furthermore, the text group acquisition unit 11 may generate a text group evaluation prompt TEP such as "Cone, handle, on the ground" If the answer is obtained, only the text contained in the answer may be selected as the text to be used to generate the detection prompt DP.
[0061] Furthermore, in the prompt generation process by the prompt generation unit 12, a search process for searching for text other than the one or more texts selected by the selection process as further text to be used to generate the detection prompt DP; The above configuration may be adopted.
[0062] As an example, the text group acquisition unit 11 "The following are some visual features to detect rollers. Any others you can think of? - Cone-shaped, with handle, on the ground and provide the search prompt to at least one of the language models LM1 and LM2. Furthermore, the text group acquisition unit 11 may generate a search prompt such as "With wheels" When an answer such as the above is obtained, the text "with wheels" contained in the answer may be included in the visual representation text group TG.
[0063] When the prompt generation unit 12 generates a prompt (text group generation prompt TGP) for generating a visual expression text group TG with reference to the input data IND, the text group generation prompt TGP includes the following: -Text contained in the input data, and - Instructions to output one or more expressions that visually explain the above text The generated text group generation prompt TGP is used by the text group acquisition unit 11.
[0064] Furthermore, when the prompt generation unit 12 generates a prompt (text group evaluation prompt TEP) for evaluating the visual expression text group TG, the text group evaluation prompt TEP includes the following: one or more texts included in the visually represented text group TG, and A prompt to rate the appropriateness of each of the above texts or texts The generated text group evaluation prompt TEP is used, for example, in the evaluation process by the prompt generation unit 12. More specific prompts generated by the prompt generation unit 12 will be described later with reference to different drawings.
[0065] (Providing part 13) The providing unit 13 receives a prompt and a target image TIM as input, and provides the prompt (detection prompt DP) generated by the prompt generating unit 12 to a detection model DM that detects a detection target specified by the prompt from the target image TIM. As an example, the providing unit 13 provides the target image TIM and the detection prompt DP generated by the prompt generating unit 12 to the server 53 via the communication unit 30. Then, the server 53 inputs the target image TIM and the detection prompt DP to the detection model DM.
[0066] (Output information generation unit 15) The output information generation unit 15 generates output information OUT from the detection result DR output by the detection model DM to which the detection prompt DP and the target image TIM have been input. As an example, the generated output information OUT is visually presented (displayed) to the user via the input / output unit 40. Specific examples of the output information OUT generated by the output information generation unit 15 will be described later.
[0067] (Specific Configuration Example of Information Processing System 100A) A specific example of the configuration of the information processing system 100A will be described below with reference to different drawings. FIG. 6 is a diagram showing a specific example of the configuration of the information processing system 100A according to this exemplary embodiment. Each component will be described below in order of processing. Note that the arrows in FIG. 6 merely show an example of the direction in which data moves, and data may move in the opposite direction, or data may move between components other than those connected by arrows.
[0068] 6, detection target object description information is input as input data IND to the text group acquisition unit 11. Here, the detection target object description information is information that describes an object that is a detection target, and includes, for example, the name and type of the object.
[0069] (Text group acquisition unit 11) As shown in Fig. 6, in this example, the text group acquisition unit 11 to which the detection target object description information is input includes a plurality of (N in the example of Fig. 6) visual expression text group generation units 11-1 to 11-N. The detection target object description information described above is input to each of these visual expression text group generation units. In the following description, each visual expression text group generation unit may be expressed as visual expression text group generation unit 11-i, 11-j, etc. using indexes i, j, etc.
[0070] As an example, each visual representation text group generation unit provides a text group generation prompt TGP obtained by referring to the detection target object description information to a language model possessed by each of the multiple servers, and uses the text group output by the language model as a visual representation text group TG.
[0071] For example, the visual expression text group generation unit 11-1 provides a text group generation prompt TGP obtained by referring to the detection target object description information to a language model LM1 possessed by the server 51, and uses the text group output by the language model LM1 as the visual expression text group TG-1.
[0072] Similarly, the visual expression text group generation unit 11-2 provides the text group generation prompt TGP to the language model LM2 of the server 52, and uses the text group output by the language model LM2 as the visual expression text group TG-2.
[0073] Similarly, the other visual expression text group generators 11-j (3≦j≦N) provide the text group generation prompt TGP to the corresponding language models, and use the text groups output by the corresponding language models as the visual expression text groups TG-j (3≦j≦N). The visual expression text groups TG-1 to TG-N generated by the respective visual expression text group generators 11-1 to 11-N constitute the visual expression text group TG described above.
[0074] In this way, the text group acquisition process by the text group acquisition unit 11 includes the following steps: The process includes generating multiple visually expressed text groups TG-1 to TG-N using multiple different generative models (language models LM1, LM2, ...).
[0075] Fig. 7 shows processing examples 1 and 2 by the visual representation text group generation unit. More specifically, Fig. 7 shows an example of a prompt (text group generation prompt TGP) that the visual representation text group generation unit 11-i (1≦i≦N) according to this example provides to the language model in order to generate the visual representation text group TG-i when the input data IND includes "excavator" as a phrase indicating the detection target, and an example of an answer that the language model outputs by referring to the prompt.
[0076] As shown in the upper part of Figure 7, the text group generation prompt TGP includes the following: - The wording included in the input data ("excavator" in the upper part of Figure 7), and An instruction (query) instructing the system to output one or more expressions that visually explain the above statement ("What are useful visual features for detecting a {excavator} in an image?" in the top row of Figure 7) Contains:
[0077] Then, the visual expression text group generation unit 11-i according to this example generates the answer output by the language model. "- excavator has bucket and clamshell - excavator has tracks and tires - excavator is usually yellow or orange in color" From the above, we extract "bucket," "clamshell," "tracks," "tires," "yellow," and "orange" as visual representation texts, and generate a visual representation text group TG-i that includes these texts.
[0078] Furthermore, the visual representation text group generator 11-i may generate the text group generation prompt TGP using the object names explicitly learned by the detection model DM. Specifically, as shown in processing example 2 in the lower part of Fig. 7, the visual representation text group generator 11-i may generate a prompt including the object names explicitly learned by the detection model DM, such as "person, bicycle, car, sports ball, kite, baseball bat, sandwich, orange, broccoli, carrot, chair, couch, potted plant, bed, ...," as the text group generation prompt TGP.
[0079] Then, the visual expression text group generation unit 11-i according to this example generates the answer output by the language model. “Tracks or wheels: Excavators typically have tracks or wheels, making them visually distinct from stationary objects like chairs, couches, and potted plants. Bucket arm: The excavator features a distinctive bucket arm, which sets it apart from a wide range of objects listed above. …” From the above, we extract "tracks," "wheels," and "bucket arm" as visual representation texts, and generate a visual representation text group TG-i that includes these texts.
[0080] 6, the visual expression text groups TG-1 to TG-N are input to the prompt generation unit 12. As shown in FIG. 6, the prompt generation unit 12 according to this example includes a text group evaluation unit 121, a text selection unit 122, an end determination unit 123, and a detection prompt generation unit 124.
[0081] (Text Group Evaluation Unit 121) The text group evaluation unit 121 executes an evaluation process to evaluate the appropriateness of at least any text included in the visual expression text group TG. By performing this evaluation process, the text group evaluation unit 121 can exclude noise expressions (noise text) that are not suitable for object detection from the visual expression text group TG.
[0082] 6, the text group evaluation unit 121 according to this example includes a plurality of (N in the example of FIG. 6) visual expression text group evaluation units 121-1 to 121-N. The plurality of visual expression text groups TG-1 to TG-N constituting the visual expression text group TG are input to the plurality of visual expression text group evaluation units 121-1 to 121-N.
[0083] As an example, each of the visual representation text group evaluation units 121-1 to 121-N performs the above evaluation process by providing the above-mentioned text group evaluation prompt TEP to the language model possessed by each of the multiple servers and obtaining the evaluation results output by the language model.
[0084] For example, the visual expression text group evaluation unit 121-1 provides the text group evaluation prompt TEP to the language model LM1 of the server 51, and obtains the evaluation result output by the language model LM1.
[0085] Similarly, the visual expression text group evaluation unit 121-2 provides the text group evaluation prompt TEP to the language model LM2 of the server 52, and obtains the evaluation result output by the language model LM2.
[0086] Similarly, the other visual expression text group evaluation units 121-j (3≦j≦N) provide the text group evaluation prompt TEP to the corresponding language model and obtain the evaluation results output by the corresponding language model.
[0087] Here, the above text group evaluation prompt TEP may include, for example: One or more texts included in at least one of the plurality of visual representation text groups TG-1 to TG-N, and A prompt to rate the appropriateness of each of the above texts or texts Includes:
[0088] In this way, the evaluation process executed by the text group evaluation unit 121 includes a process of evaluating the plurality of visual expression text groups TG-1 to TG-N using a plurality of mutually different evaluation models (language models LM1, LM2, ...). The evaluation results by each of the visual expression text group evaluation units 121-1 to 121-N are provided to the text selection unit 122, which will be described later.
[0089] Fig. 8 shows processing examples 1 and 2 by the visual representation text group evaluation unit. More specifically, Fig. 8 shows an example of a prompt (text group evaluation prompt TEP) that the visual representation text group evaluation unit 121-i (1≦i≦N) according to this example provides to the language model to generate an evaluation result when the input data IND includes "excavator" as a phrase indicating the detection target, and an example of an answer that the language model outputs by referring to the prompt.
[0090] As shown in the upper part of Figure 8, Example 1, the text group evaluation prompt TEP includes the following: - A word indicating the detection target contained in the input data ("excavator" in the upper part of Figure 7), One or more texts included in the visual representation text group TG ("yellow", "arm", "bucket" in the top row of Figure 8), and - An instruction to evaluate the appropriateness of each of the above one or more pieces of text ("Evaluate the necessity of the following visual features for detecting an excavator in an image and remove unnecessary visual features." in the upper part of Figure 8) Contains:
[0091] Then, the visual expression text group evaluation unit 121-i according to this example evaluates the answer output by the language model. "1. Hydraulic Arms and Joints: Necessity: Low Rationale: While hydraulic arms and joints are characteristic features, they may not be visible in all images. They are more specific details and may not contribute significantly to initial object detection. 2. Characteristic Shape (Body, Arm, and Bucket): Necessity: High Rationale: The overall shape, which includes the combination of body, arm, and bucket, is crucial for accurate excavator detection. It encompasses the core features that define the object.” Based on this, we give a negative evaluation to "Hydraulic Arms and Joints" as a visual representation text and a positive evaluation to "Body, Arm, and Bucket."
[0092] Furthermore, the visual expression text group evaluation unit 121-i may generate a prompt to ask what the user associates with the text included in the visual expression text group TG as the text group evaluation prompt TEP. As an example, as shown in the processing example 2 in the lower part of FIG. 8, the visual expression text group evaluation unit 121-i may generate a prompt to ask what the user associates with the text included in the visual expression text group TG as the text group evaluation prompt TEP. A question asking what is associated with the text ("bucket" and "clamshell") included in the visual text group TG “Q:What is an object in construction site which has the following visual features? - bucket and clamshell Please enumerate possible objects as much as possible.” A prompt containing the following may be generated as a text group evaluation prompt TEP.
[0093] Then, the visual expression text group evaluation unit 121-i according to this example evaluates the answer output by the language model. 1. Backhoe loader 2. Excavator with clamshell attachment 3. Clamshell bucket crane 4. Material handler with clamshell grab” In this example, the answer output by the language model includes the detection target "excavator" indicated in the input data IND, so the visual representation text group evaluation unit 121-i makes a positive evaluation of the above text ("bucket" and "clamshell") included in the visual representation text group TG.
[0094] Each of the visual expression text group evaluation units 121-1 to 121-N is The configuration may be such that one of the plurality of visual expression text groups TG-1 to TG-N corresponding to the user is evaluated (evaluation process example 1), Among the plurality of visual expression text groups TG-1 to TG-N, one or more visual expression text groups other than the visual expression text group corresponding to itself may be evaluated (evaluation process example 2). All of the plurality of visual expression text groups TG-1 to TG-N may be evaluated (evaluation process example 3).
[0095] For example, when the evaluation of the visual expression text groups TG-1, TG-2, and TG-3 is performed by the visual expression text group evaluation units 121-1, 121-2, and 121-3, The visual expression text group evaluation unit 121-1 evaluates the visual expression text group TG-1. The visual expression text group evaluation unit 121-2 evaluates the visual expression text group TG-2. The visual expression text group evaluation unit 121-3 evaluates the visual expression text group TG-3. (corresponding to the above evaluation processing example 1), The visual expression text group evaluation unit 121-1 evaluates the visual expression text groups TG-2 and TG-3. The visual expression text group evaluation unit 121-2 evaluates the visual expression text groups TG-3 and TG-1. The visual expression text group evaluation unit 121-3 evaluates the visual expression text groups TG-1 and TG-2. The above configuration may be adopted (corresponding to the above evaluation processing example 2). The visual expression text group evaluation unit 121-1 evaluates the visual expression text groups TG-1 to TG-3. The visual expression text group evaluation unit 121-2 also evaluates the visual expression text groups TG-1 to TG-3. The visual expression text group evaluation unit 121-3 also evaluates the visual expression text groups TG-1 to TG-3. The above configuration may be adopted (corresponding to the above evaluation processing example 3).
[0096] When each of the visual expression text group evaluation units 121-1, 121-2, and 121-3 is configured to perform evaluation processing using a different language model, the above evaluation processing examples 2 and 3 include processing (also called mutual evaluation processing) in which a visual expression text group generated by one language model is evaluated by another language model.
[0097] In other words, the text group acquisition process by the text group acquisition unit 11 includes the following steps: generating a first set of text using a first generative model; A process of generating a second text group using a second generation model and is included, In the evaluation process by the text group evaluation unit 121, a process of evaluating the second text group using a first evaluation model including the first generation model, and a process of evaluating the first text group using a second evaluation model including the second generation model are included.
[0098] According to the above configuration, it is possible to obtain more diverse target object expressions (visual expressions) and remove inappropriate expressions caused by the bias of the expressions of individual language models by the text selection unit 122 described later.
[0099] (Text selection unit 122) The text selection unit 122 executes a selection process of selecting one or more texts used for generating the detection prompt DP from the visual expression text group TG by referring to the result of the evaluation process by the text group evaluation unit 121.
[0100] As an example, in the text group evaluation unit 121, when each of the N plurality of visual expression text group evaluation units 121-1 to 121-N uses a different language model for the evaluation process, the text selection unit 122 · Among the plurality of visual expression text groups TG-1 to TG-N, select a visual expression text group (in other words, a visual expression text group whose evaluation value exceeds a predetermined standard) whose evaluation by M (0 < M < N) or more visual expression text group evaluation units is positive as the text to be used for generating the detection prompt DP may be configured as. In other words, the text selection unit 122 · Among the plurality of texts included in the plurality of visual representation text groups TG-1 to TG-N, one or more texts (in other words, one or more texts whose evaluation values exceed a predetermined criterion) for which the evaluation by the M (0 < M < N) or more visual representation text group evaluation units is affirmative A) Select as the text to be used for generating the detection prompt DP It may be configured as follows.
[0101] Also, reliability may be assigned to each of the above different language models, and the text selection unit 122 may execute the following selection process in consideration of the reliability. Here, as an example, the reliability may be configured to be set by the text selection unit 122 by referring to information such as the domain (category, field) of the data used for learning by each language model. Further, the reliability may be configured to be set by the selection unit 122 by further referring to information such as the domain (category, field) of the data used for learning by the detection model DM. For example, the selection unit 122 may perform processing such as setting the reliability of the language model that used data in a domain overlapping with the domain of the data used for learning by the detection model DM to be higher than the reliability of other language models.
[0102] (Example of selection process considering reliability 1) As an example of the selection process considering the above reliability, the text selection unit 122 · Select a plurality of texts among the plurality of texts included in the plurality of visual representation text groups TG-1 to TG-N for which the evaluation by the M (0 < M < N) or more visual representation text group evaluation units is affirmative · Among the above selected plurality of texts, select as the text to be used for generating the detection prompt DP a text for which the total reliability of the language model used by the visual representation text group evaluation unit that gave an affirmative evaluation to the text is greater than or equal to a predetermined threshold It may be configured as follows.
[0103] For example, in a configuration where different language models LM1, LM2, and LM3 are respectively used by the visual representation text group evaluation units 121-1, 121-2, and 121-3 The visual representation text group evaluation units 121-1 and 121-2 give a positive evaluation to a certain text TX1 (in other words, the language models LM1 and LM2 give a positive evaluation), When the visual representation text group evaluation units 121-2 and 121-3 have made a positive evaluation of another text TX2 (in other words, the language models LM2 and LM3 have made a positive evaluation), When the reliability of each of the visual expression text group evaluation units 121-1, 121-2, and 121-3 (in other words, the reliability of each of the language models LM1, LM2, and LM3) is 0.6, 0.3, and 0.1, and the predetermined threshold is 0.8, For example, in this case, For text TX1, the sum of the confidence scores of the language models (LM1, LM2) used by the visual expression text group evaluation unit that gave a positive evaluation to the text is 0.6+0.3=0.9 Since this is equal to or greater than the predetermined threshold, the text TX1 is selected as the text to be used to generate the detection prompt DP.
[0104] on the other hand, For text TX2, the sum of the confidence scores of the language models (LM2, LM3) used by the visual expression text group evaluation unit that gave a positive evaluation to the text is 0.3+0.1=0.4 Since this is less than the predetermined threshold, the text TX2 is not selected as the text to be used to generate the detection prompt DP.
[0105] (Selection processing example 2 taking reliability into account) As another example of the selection process by the text selection unit 122 taking the reliability into consideration, the text selection unit 122 Calculating a linear sum of evaluation values by each of the plurality of visual expression text group evaluation units 121-1 to 121-N for each of the plurality of texts included in the plurality of visual expression text group TG-1 to TG-N, the linear sum being weighted by the reliability of the language model used by the visual expression text group evaluation unit; Texts for which the linear sum is equal to or greater than a predetermined threshold are selected as texts to be used to generate a detection prompt DP. The above configuration may also be used.
[0106] For example, in a configuration in which the visual expression text group evaluation units 121-1, 121-2, and 121-3 use different language models LM1, LM2, and LM3, respectively, and the reliability and predetermined threshold values are given as described above, -As an evaluation value for text TX1, Visual expression text group evaluation unit 121-1 evaluation value: 0.9 Visual expression text group evaluation unit 121-2 evaluation value: 0.9 Visual expression text group evaluation unit 121-3 evaluation value: 0.3 is calculated, -As an evaluation value for text TX2, Visual expression text group evaluation unit 121-1 evaluation value: 0.1 Visual expression text group evaluation unit 121-2 evaluation value: 0.9 Visual expression text group evaluation unit 121-3 evaluation value: 0.9 Let us take the case where the following is calculated. In this case, The linear sum of the evaluation values by the visual expression text group evaluation units 121-1, 121-2, and 121-3 for the text TX1, in which the reliability of the language model used by the visual expression text group evaluation unit is used as a weighting coefficient, is 0.6×0.9+0.3×0.9+0.1×0.3=0.84 Since this is equal to or greater than the predetermined threshold, the text TX1 is selected as the text to be used to generate the detection prompt DP.
[0107] on the other hand, The linear sum of the evaluation values by the visual expression text group evaluation units 121-1, 121-2, and 121-3 for the text TX2, in which the reliability of the language model used by the visual expression text group evaluation unit is used as a weighting coefficient, is 0.6×0.1+0.3×0.9+0.1×0.9=0.42 Since this is less than the predetermined threshold, the text TX2 is not selected as the text to be used to generate the detection prompt DP.
[0108] (End determination unit 123) The end determination unit 123 performing a search process to search for text other than the one or more texts selected by the selection process by the selection unit 122 as further text to be used to generate the detection prompt DP; If no further text is found by the search process, an instruction to generate a detection prompt is issued to the detection prompt generating unit 124, which will be described later.
[0109] As an example, the end determination unit 123 instructs the visual representation text group generation unit 11-i to generate text that visually represents the detection target, other than the one or more texts selected by the selection process by the selection unit 122. Then, as an example, the visual representation text group generation unit 11-i inquires of the language model whether such text exists. If the visual representation text group generation unit 11-i finds such text, the end determination unit 123 adds the text to the visual representation text group TG.
[0110] The termination determination unit 123 repeats the above process, and when no further text is found, it instructs the detection prompt generation unit 124, which will be described later, to generate a detection prompt. By performing the above process, the termination determination unit 123 can improve the comprehensiveness of the text expressions included in the visual expression text group TG, and expand the clues for object detection.
[0111] (Detection prompt generation unit 124) The detection prompt generation unit 124 generates a detection prompt DP by referring to the visual expression text group TG that is generated by the visual expression text group generation unit 11-i and positively evaluated by the visual expression text group evaluation unit 121. The generated detection prompt DP is provided to the detection model DM included in the server 53 together with the target image TIM.
[0112] Then, the detection model DM to which the detection prompt DP and the target image TIM have been input detects the detection target specified by the detection prompt DP from the target image DM, and provides the detection result DR to the information processing device 1A. As an example, the detection result DR is acquired by the acquisition unit 14 via the communication unit 30 and stored in the storage unit 20A. Here, although specific examples of information included in the detection result DR do not limit this exemplary embodiment, as an example, the detection result DR may be, as shown in FIG. Detected object information includes object class name, score, and location information The detection result DR may be configured to include the following: The detection result DR is then referred to by the output information generating unit 15, for example, and is used to generate output information OUT.
[0113] 9 shows an example of a detection prompt DP generated by the detection prompt generation unit 124 and a detection result based on a detection model DM that references the detection prompt DP. In Example 1 shown in the upper part of FIG. 9, an example of a detection prompt DP is “Detect “human. excavator” in the image. - “excavator bucket clamshell” The detection result by the detection model DM that references the detection prompt DP is shown in FIG. 1. Here, the detection prompt DP includes: - Words indicating the detection target contained in the input data ("human" or "excavator"), -Text included in the visual representation text group TG for the detection target "excavator" ("bucket" and "clamshell") In this example, the detection results include the object class names "excavator" and "human" and the bounding boxes surrounding the objects.
[0114] In addition, in Example 2 shown in the lower part of Figure 9, an example of the detection prompt DP is "Human prompt. Roller prompt." and the detection result by the detection model DM that references the detection prompt DP. Here, the "person prompt" in the detection prompt DP is a prompt for identifying the person who is the detection target, and the "compactor prompt" is a prompt for identifying the compactor who is the detection target. At least one of the "person prompt" and the "compactor prompt" is accompanied by text that visually expresses the detection target (text included in the visual expression text group TG related to the detection target).
[0115] The detection prompt generating unit 124 may generate a detection prompt DR for each of the plurality of detection targets included in the target image TIM. In Example 3 shown in FIG. 10, the detection prompt generating unit 124 generates a detection prompt DR for each of the plurality of detection targets included in the target image TIM. Generate "human prompt" and "roller prompt" separately, By providing each prompt and the target image TIM to the detection model DM, the detection result for the person (detection result 1 in Figure 10) and the detection result for the roller (detection result 2 in Figure 10) are obtained separately. In this example, the output information generation unit 15 combines the detection result 1 and the detection result 2 to generate output information OUT that includes the detection results related to the person and the roller.
[0116] The detection prompt DP generated by the detection prompt generating unit 124 and the visual expression text group TG referenced to generate the detection prompt DP are stored in the storage unit 20A.
[0117] (Effects of information processing device 1A) As described above, in the information processing device 1A according to this exemplary embodiment, By referring to input data IND that specifies the detection target, a visual representation text group TG is generated that includes a plurality of texts that visually represent the detection target; Generate a prompt (detection prompt DP) by referring to the visual representation text group TG; The generated prompt (detection prompt DP) is provided to a detection model DM that receives a prompt and an image as input and detects the detection target specified by the prompt from the image. According to the above configuration, a visual expression text group including a plurality of texts visually expressing a detection target is generated, and a prompt is generated by referring to the visual expression text group, so that a suitable prompt can be generated to be provided to a detection model that detects a detection target from an image based on the prompt. In other words, a suitable prompt can be generated for detecting an object in an image.
[0118] Furthermore, since a visual expression text group TG is automatically generated from the detection target (object class name) included in the input data IND, no effort is required from the user. Furthermore, a prompt (detection prompt DP) is generated by referring to the visual expression text group TG and provided to the detection model DM, so that objects that are not easy to detect by the detection model DM (rare objects in the domain of the training data of the detection model DM) can also be detected with high accuracy.
[0119] Furthermore, as described above, by generating and evaluating the visual expression text group TG using a language model in the text group acquisition unit 11 and the text group evaluation unit 121, it is possible to automatically generate prompts that achieve highly accurate object detection.
[0120] Furthermore, as described above, the text group acquisition unit 11 and the text group evaluation unit 121 are configured to perform a process (also called a mutual evaluation process) in which a visual expression text group generated by one language model is evaluated by another language model, thereby making it possible to acquire a wider variety of target object expressions and generate prompts that remove noise expressions (expressions unsuitable for object detection) caused by bias in the expressions of the language models.
[0121] (Prompt Note 1) The rules for generating the detection prompt generated by the detection prompt generator 124 are not limiting to the present exemplary embodiment, but as partly described above, the following rules may be used:
[0122] (Rules for the detection prompt DP to be input into the detection model DM) Enumerate each object prompt: “P0 . P1 . Pi . … . PN” where Pi is the prompt for object i (0 < i < N, N: number of object types). Also, "." is a delimiter defined by the detection model DM. (Prompt generation rules for each object) Rule 1: List only visual representations “Vi0 Vi1 ···” Vij: jth visual representation text of object i (0 < j < Mi, Mi: number of visual representation texts of object i) Rule 2: List the object name and visual representation “Ni Vi0 Vi1 …” where Ni is the class name of object i However, Ni can be inserted between any visual representation text.
[0123] (Prompt Note 2) The prompt generator 12 generates various prompts. Text Swarm Generation Prompt TGP Text Group Evaluation Prompt TEP Prompt DP for detection At least one of the prompts may be presented to the user via the input / output unit 40, a correction instruction from the user may be received, and at least one of the prompts may be corrected based on the correction instruction. The control unit 10A may then use the corrected prompt in subsequent processing. This configuration allows the user's intention to be reflected, thereby generating a more suitable prompt.
[0124] (Application example) The information processing system 100A described in this exemplary embodiment is applicable to a variety of fields without any particular limitations. For example, the information processing system 100A may be used to understand the work of workers in the construction, civil engineering, manufacturing, and other industries, and to provide support to the workers. For example, a real-time video of a construction site may be acquired as the target image TIM, and input data IND may be used that specifies "human" and "excavator" as detection targets. The output information generator 15 may then identify the positional relationship between the person and the excavator in the video, and if a danger is detected, a warning display or warning sound may be generated.
[0125] In other words, the information processing device 1A: Generate a detection prompt DP that specifies "human" and "excavator" as the detection targets, -Provide the detection prompt DP to the detection model DM, The detection result DR output by the detection model DM to which the detection prompt DP and the target image TIM have been input is acquired by the acquisition unit 14 via the communication unit 30; The output information generating unit (warning means) 15 executes a warning process by referring to the detection result DR output by the detection model DM. The information processing device 1A may also be used to detect whether or not a safety protector is being worn at a construction site, or to detect tools used by workers at construction, factory, civil engineering, or other sites.
[0126] Third Exemplary Embodiment A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be assigned the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in each drawing referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.
[0127] (Configuration of information processing system 100B) The configuration of an information processing system 100B according to this exemplary embodiment will be described with reference to Fig. 11. Fig. 11 is a block diagram showing the configuration of the information processing system 100B. As shown in Fig. 11, the information processing system 100B includes an information processing device 1B and a server 53 connected to the information processing device 1B via a network N. On the other hand, the information processing system 100B does not include the servers 52 and 53 that are included in the information processing system 100A according to the second exemplary embodiment.
[0128] 11, the information processing device 1B according to this exemplary embodiment does not include the text group acquisition unit 11 and the prompt generation unit 12 that are included in the information processing device 1A according to exemplary embodiment 2, but does include a prompt selection unit 23. In the following, explanations of matters that overlap with exemplary embodiment 2 will be omitted, and only matters that differ from exemplary embodiment 2 will be explained.
[0129] (Prompt Selection Section 23) The prompt selection unit 23 refers to the wording that specifies the detection target included in the input data IND, and selects a detection prompt DP to be provided to the detection model DM from the plurality of prompts included in the prompt group PRG.
[0130] Here, the prompt group PRG according to this exemplary embodiment includes: A text group acquisition process for acquiring a visual expression text group TG including a plurality of texts visually expressing the detection target; A prompt generation process that generates a prompt (detection prompt DP) by referring to the visual representation text group TG; As an example, the prompt group PRG according to this exemplary embodiment stores a plurality of prompts (detection prompts DP) generated by the prompt generator 12 according to the second exemplary embodiment.
[0131] In this way, the information processing device 1B according to this exemplary embodiment: Acquisition means (acquisition unit 14 (21)) for acquiring input data IND that specifies a detection target; A providing means (providing unit 13 (22)) that receives a prompt and an image as input and provides a prompt (detection prompt DP) obtained by referring to the input data IND to a detection model DM that detects a detection target specified by the prompt from the image; It is equipped with The prompt provided by the providing means comprises: a text group acquisition process for acquiring a visual expression text group TG including a plurality of texts visually expressing the detection target; a prompt generation process for generating a prompt by referring to the visual representation text group TG; This is a prompt generated by a process that includes
[0132] According to the above configuration, a suitable prompt is generated by referring to a visual expression text group including a plurality of texts visually expressing a detection target, and the suitable prompt can be used to perform object detection using a detection model.
[0133] [Software implementation example] Some or all of the functions of the information processing devices 1, 2, 1A, and 1B (hereinafter also referred to as "each of the above devices") may be realized by hardware such as an integrated circuit (IC chip), or by software.
[0134] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 12. Figure 12 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.
[0135] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to operate as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.
[0136] The processor C1 may be, for example, a central processing unit (CPU), a graphic processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0137] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.
[0138] Furthermore, the program P can be recorded on a non-transitory tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0139] [Additional Notes] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0140] (Appendix A1) a text group acquisition means for acquiring a visual expression text group including a plurality of texts visually expressing the detection target by referring to input data specifying the detection target; prompt generation means for generating a prompt by referring to the set of visually represented texts; providing means for providing the prompt generated by the prompt generating means to a detection model that receives a prompt and an image as input and detects a detection target specified by the prompt from the image; An information processing device comprising:
[0141] (Appendix A2) The prompt generation process by the prompt generation means includes: and an evaluation process for evaluating the appropriateness of at least any text included in the visually expressed text group. 10. The information processing device according to claim 1,
[0142] (Appendix A3) The prompt generation process by the prompt generation means includes: a selection process for selecting one or more texts to be used to generate the prompt from the group of visual representation texts with reference to a result of the evaluation process; 10. The information processing device according to claim 9, wherein the information processing device is a
[0143] (Appendix A4) The prompt generation process by the prompt generation means includes: a search process for searching text other than the one or more texts selected by the selection process for further text to use in generating the prompt; 1. The information processing device according to claim 1,
[0144] (Appendix A5) The prompt generation process by the prompt generation means includes: and generating a prompt from the one or more pieces of text selected by the selection process if the search process does not find the additional text. 1. The information processing device according to claim 1,
[0145] (Appendix A6) The text group acquisition process by the text group acquisition means includes: generating a plurality of sets of said visually represented texts using a plurality of different generative models; The evaluation process includes: and evaluating the plurality of sets of visually represented text using a plurality of different evaluation models. An information processing device according to any one of appendices A2 to A5.
[0146] (Appendix A7) The text group acquisition process includes: generating a first set of text using a first generative model; generating a second set of text using a second generative model; evaluating the second collection of text using a first evaluation model that includes the first generative model; evaluating the first set of texts using a second evaluation model that includes the second generative model; Contains 10. The information processing device according to claim 9, wherein the information processing device is an information processing device according to claim 8.
[0147] (Appendix A8) an acquisition means for acquiring a detection result based on the detection model; a warning means for executing a warning process by referring to a detection result obtained by the detection model; Equipped with An information processing device according to any one of appendices A1 to A7.
[0148] (Appendix A9) an acquisition means for acquiring input data specifying a detection target; providing means for providing a prompt obtained by referring to the input data to a detection model that receives a prompt and an image and detects a detection target specified by the prompt from the image; It is equipped with The prompt provided by the providing means comprises: a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process that generates a prompt by referencing the visual representation text group; is a prompt generated by a process that includes Information processing device.
[0149] (Appendix A10) generating a visual representation text group including a plurality of texts visually representing the detection target by referring to input data specifying the detection target; generating a prompt with reference to the visually represented text; providing the generated prompt to a detection model that receives a prompt and an image as input and detects a target object specified by the prompt in the image; An information processing method comprising:
[0150] (Appendix A11) acquiring input data specifying a detection target; a prompt obtained by referring to the input data is provided to a detection model that receives a prompt and an image as input and detects a detection target specified by the prompt from the image; It contains The prompt provided may include: a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process that generates a prompt by referencing the visual representation text group; is a prompt generated by a process that includes Information processing methods.
[0151] (Appendix A12) A program that causes a computer to function as an information processing device, The computer a text group acquisition means for acquiring a visual expression text group including a plurality of texts visually expressing the detection target by referring to input data specifying the detection target; prompt generation means for generating a prompt by referring to the set of visually represented texts; providing means for providing the prompt generated by the prompt generating means to a detection model that receives a prompt and an image as input and detects a detection target specified by the prompt from the image; A program that functions as a
[0152] (Appendix A13) A program that causes a computer to function as an information processing device, The computer an acquisition means for acquiring input data specifying a detection target; providing means for providing a prompt obtained by referring to the input data to a detection model that receives a prompt and an image and detects a detection target specified by the prompt from the image; It functions as The prompt provided by the providing means comprises: a text group acquisition process for generating a visually expressed text group including a plurality of texts visually expressing the detection target; a prompt generation process that generates a prompt by referencing the visual representation text group; is a prompt generated by a process that includes program. [Explanation of symbols]
[0153] 1, 2, 1A, 1B ··· Information processing device 11 Text group acquisition section 12 Prompt Generation 121 Text Group Evaluation Unit 122 Text selection 13...Providing Department 15...Output information generation unit (warning means) 14(21) ···Acquisition part 100A, 100B ···Information processing system
Claims
1. a text group acquisition means for acquiring a visual expression text group including a plurality of texts visually expressing the detection target by referring to input data specifying the detection target; prompt generation means for generating a prompt by referring to the set of visually represented texts; providing means for providing the prompt generated by the prompt generating means to a detection model that receives a prompt and an image as input and detects a detection target specified by the prompt from the image; An information processing device comprising:
2. The prompt generation process by the prompt generation means includes: and an evaluation process for evaluating the appropriateness of at least any text included in the visually expressed text group. The information processing device according to claim 1 .
3. The prompt generation process by the prompt generation means includes: a selection process for selecting one or more texts to be used to generate the prompt from the group of visual representation texts with reference to a result of the evaluation process; The information processing device according to claim 2 .
4. The prompt generation process by the prompt generation means includes: A search process is included that searches for text other than the one or more texts selected by the selection process as additional text to use in generating the prompt. The information processing device according to claim 3 .
5. The prompt generation process by the prompt generation means includes: and generating a prompt from the one or more pieces of text selected by the selection process if the search process does not find the additional text. The information processing device according to claim 4 .
6. The text group acquisition process by the text group acquisition means includes: generating a plurality of sets of said visually represented texts using a plurality of different generative models; The evaluation process includes: and evaluating the plurality of sets of visually represented text using a plurality of different evaluation models. The information processing device according to claim 2 .
7. The text group acquisition process includes: generating a first body of text using a first generative model; generating a second body of text using a second generative model; evaluating the second collection of text using a first evaluation model that includes the first generative model; evaluating the first collection of texts using a second evaluation model that includes the second generative model; and Contains The information processing device according to claim 6 .
8. an acquisition means for acquiring input data specifying a detection target; providing means for providing a prompt obtained by referring to the input data to a detection model that receives a prompt and an image and detects a detection target specified by the prompt from the image; It is equipped with The prompt provided by the providing means comprises: a text group acquisition process for acquiring a visual expression text group including a plurality of texts visually expressing the detection target; a prompt generation process that generates a prompt by referencing the visual representation text group; is a prompt generated by a process that includes Information processing device.
9. referencing input data specifying a detection target, and acquiring a visual expression text group including a plurality of texts visually expressing the detection target; generating a prompt with reference to the visually represented text; providing the generated prompt to a detection model that receives a prompt and an image as input and detects a target object specified by the prompt in the image; An information processing method comprising:
10. A program that causes a computer to function as an information processing device, The computer a text group acquisition means for acquiring a visual expression text group including a plurality of texts visually expressing the detection target by referring to input data specifying the detection target; prompt generation means for generating a prompt by referring to the set of visually represented texts; providing means for providing the prompt generated by the prompt generating means to a detection model that receives a prompt and an image as input and detects a detection target specified by the prompt from the image; A program that functions as a
Citation Information
Patent Citations
Moving image recognition processing system
JP1993174147A