Marking method and program
By detecting and selecting the least frequent character type for markings and dividing images into regions, the method enhances image understanding accuracy in large-scale multimodal models by preventing confusion with original image symbols.
Patent Information
- Application Number
- JP2024002621
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
AI Technical Summary
Existing methods for applying markings to images for use in large-scale multimodal models can lead to confusion between characters and symbols in the original image and the applied markings, which affects the accuracy of image understanding.
Detect characters or symbols from an image, count their types, select the type with the smallest number for marking, divide the image into regions, and assign the selected type as markings to each region, then input the marked image with user text into a large-scale multimodal model.
This approach avoids confusion between original image characters and markings, improving the accuracy of image understanding in large-scale multimodal models.
Smart Images

Figure 2025108996000001_ABST
Abstract
Description
Technical Field
[0001] The disclosed technology relates to a marking method and a marking program.
Background Art
[0002] Conventionally, as preprocessing for an image input to a machine learning model, marking such as characters and symbols has been performed. For example, a document processing apparatus has been proposed that automatically creates a document image (summary or table of contents) by extracting only the locations of desired document elements from a manuscript document image and composing them. This apparatus divides a document image into a plurality of document elements, and assigns an identifier representing the meaning of the document elements such as a title and an author to each of the divided document elements. Then, this apparatus extracts elements having identifiers necessary for summary creation, table of contents creation, etc. from the group of elements to which the identifiers are assigned, and generates an output image based on the partial image corresponding to the extracted elements.
[0003] Also, a technique for improving the accuracy of image understanding in visual prompts has been proposed. This technique uses an interactive segmentation model to divide an image into regions at various granularity levels, assigns markings to each region, and uses the image obtained by superimposing the assigned identifiers on the original image as an input to a large-scale multimodal model.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the prior art of superimposing markings on an image, there is a problem that when recognizing an image in a large-scale multimodal model, confusion may occur between the characters and symbols included in the original image and the applied markings. Although this prior art describes that markings should be applied to avoid competition with the original image content, no specific method is disclosed.
[0007] In addition, since the above-mentioned prior art document processing apparatus has nothing to do with visual prompts using a large-scale multimodal model, it cannot solve the problem that confusion may occur between the characters and symbols included in the original image and the applied markings.
[0008] As one aspect, the disclosed technology aims to apply markings that avoid confusion with the characters and symbols included in the original image.
Means for Solving the Problems
[0009] In one aspect, the disclosed technology detects characters or symbols from a first image, counts the detected characters or symbols by type, and selects the type of character or symbol with the smallest detected number as the type of character or symbol to be used for marking. Further, the disclosed technology divides the first image into regions corresponding to objects, and assigns the selected type of character or symbol as a marking to each of the divided regions of the first image. Then, the disclosed technology inputs a second image in which the region division result and the assigned marking are superimposed on the first image, together with the text input from the user, into a large-scale multimodal model.
Advantages of the Invention
[0010] As one aspect, it has the effect of being able to apply markings that avoid confusion with characters and symbols included in the original image.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Embodiments for Carrying Out the Invention
[0012] Hereinafter, with reference to the drawings, an example of an embodiment according to the disclosed technology will be described.
[0013] Before explaining the details of this embodiment, the visual prompt of the large multimodal model (Large Multimodal Models: hereinafter referred to as "LMMs") to which this embodiment is applied will be described.
[0014] A visual prompt is an input to LMMs, which is a combination of text and an image. LMMs are machine learning models that combine a large language model (Large Language Model: hereinafter referred to as "LLM") and a visual language model (Visual Language Model: hereinafter referred to as "VLM"). In LMMs, visual information is given to the LLM by obtaining an embedded representation of the image by the VLM. However, LMMs do not always understand the content of an image from the embedded representation of the image. Understanding the content of an image means a state in which LMMs can correctly detect objects in the image and the relationships between objects and can appropriately respond to questions (text) about the image.
[0015] Therefore, in order to assist in understanding the image, it is conceivable to use, as the input to LMMs, an image that has been pre-processed by adding markings to each region obtained by semantically dividing the image. However, if the image contains characters or symbols that are the same as or similar to the characters or symbols to be added as markings, there is a possibility of confusion between the characters or symbols contained in the image and the markings to be added.
[0016] For example, as shown in the left figure of Fig. 1, assume that the original image contains numbers (in the example of Fig. 1, "1", "2", "3", "4"). As a specific example, in an image taken inside a factory, buttons with numbers for operating equipment inside the factory are assumed. In such a case, as shown in the right figure of Fig. 1, if numbers are used as markings, confusion may occur between the numbers on the buttons and the numbers used as markings. For example, when the question input as text is "Which button should I press next?", even though the correct answer is the second button, the answer output from the LMMs may be something like "Please press the first button."
[0017] Therefore, in this embodiment, by selecting the type of characters or symbols to be used as markings according to the image content, confusion between the characters or symbols attached to the objects in the image and the markings is prevented, and the accuracy of image understanding of the LMMs is improved. Hereinafter, the marking device according to this embodiment will be described in detail.
[0018] As shown in Fig. 2, the marking device 10 according to this embodiment functionally includes a detection unit 12, a selection unit 14, a division unit 16, an application unit 18, and an input unit 20. Also, an object detection model 32 and a region division model 34 are stored in a predetermined storage area of the marking device 10.
[0019] The detection unit 12 acquires the original image input to the marking device 10 and detects characters or symbols (hereinafter referred to as "character types") from the original image. The original image is an example of the "first image" of the disclosed technology. The character types are, for example, numbers, alphabets, symbols, etc., and a plurality of different characters or symbols form a group to represent one type of notation system.
[0020] Specifically, the detection unit 12 applies an object detection algorithm to the original image to detect an object that matches an object including character types of a predetermined type to be used for marking. The object detection algorithm may be, for example, YOLO (You Only Look Once) or the like.
[0021] More specifically, the detection unit 12 inputs the original image into an object detection model 32 that has been pre-trained in a manner corresponding to an object detection algorithm, using an image of an object including characters as training data, and obtains the output of the object detection model 32 as an object detection result. For example, when it is predetermined in advance to use numbers and alphabets for marking, as training data, images of objects including numbers from 0 to 9 and images of objects including alphabets from A to Z are used. Thereby, the object detection model 32 detects an object including a number and an object including an alphabet from the original image.
[0022] The selection unit 14 counts the detected characters for each type based on the object detection result by the detection unit 12. For example, the selection unit 14 counts the number of detected objects including numbers as the number of detections of the type "number" and the number of detected objects including alphabets as the number of detections of the type "alphabet". The selection unit 14 selects the type of character with the smallest number of detections as the type of character to be used for marking.
[0023] For example, in the example of FIG. 1, since the number of detections of the type "number" is 4 and the number of detections of the type "alphabet" is 0, the selection unit 14 selects the type "alphabet" with the smallest number of detections as the type of character to be used for marking. Also, for example, assume that the types of characters to be used for marking are "number", "alphabet", and "katakana". In this case, assume that the number of detections of the type "number" is 10, the number of detections of the type "alphabet" is 6, and the number of detections of the type "katakana" is 2. In this case, the selection unit 14 selects the type "katakana" with the smallest number of detections as the type of character to be used for marking.
[0024] The segmentation unit 16 divides the original image into regions corresponding to each object. Specifically, the segmentation unit 16 applies a region segmentation model 34 such as SAM (Segment Anything Model) to the original image to semantically segment the original image. Note that the region segmentation method is not limited to SAM, and any method that can segment the image into regions corresponding to each object in the image, such as semantic segmentation, instance segmentation, etc., may be used.
[0025] The segmentation unit 16 superimposes the result of the region segmentation on the original image. For example, the segmentation unit 16 may superimpose different colors for each region with a set transmittance, or may superimpose a boundary line surrounding each region.
[0026] The assignment unit 18 assigns the types of characters selected by the selection unit 14 to each segmented region of the original image as markings. Specifically, the assignment unit 18 selects characters from the selected type so that different markings are assigned to each region, and assigns them to each region. For example, when the type "number" is selected as the marking type, the assignment unit 18 assigns markings to each region in the order of 1, 2, 3, ···. Also, for example, when the type "alphabet" is selected as the marking type, the assignment unit 18 assigns markings to each region in the order of A, B, C, ···. Fig. 3 shows an example in which markings of the type "alphabet" selected by the selection unit 14 are assigned to the same example as Fig. 1.
[0027] The assignment unit 18 generates an input image in which markings are further superimposed on the image in which the result of the region segmentation by the segmentation unit 16 is superimposed on the original image. The left figure in Fig. 3 is an example of a schematic diagram of the original image, and the right figure in Fig. 3 is an example of a schematic diagram of the input image. In the example of the right figure in Fig. 3, the result of the region segmentation is represented by color separation (difference in the density of halftone dots).
[0028] The input unit 20 inputs the input image generated by the granting unit 18, together with the text prompt input by the user, into the LMMs 32. For example, when causing the LMMs 32 to output an answer to a question about the original image as shown in the left figure of FIG. 3, a question such as "Which button should I press next?" is input as the text prompt. Note that the LMMs 32 are pre-trained to be able to recognize the types of characters used for marking.
[0029] The marking device 10 may be realized by, for example, a computer 40 shown in FIG. 4. The computer 40 includes a CPU (Central Processing Unit) 41, a GPU (Graphics Processing Unit) 42, a memory 43 as a temporary storage area, and a non-volatile storage device 44. Further, the computer 40 includes an input / output device 45 such as an input device and a display device, and an R / W (Read / Write) device 46 that controls reading and writing of data to and from the storage medium 49. Further, the computer 40 includes a communication I / F (Interface) 47 connected to a network such as the Internet. The CPU 41, GPU 42, memory 43, storage device 44, input / output device 45, R / W device 46, and communication I / F 47 are connected to each other via a bus 48.
[0030] The storage device 44 is, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, or the like. The storage device 44 as a storage medium stores a marking program 50 for causing the computer 40 to function as the marking device 10. The marking program 50 has a detection process control command 52, a selection process control command 54, a division process control command 56, a granting process control command 58, and an input process control command 60. Further, the storage device 44 has an information storage area 70 in which information constituting each of the object detection model 32 and the region division model 34 is stored.
[0031] The CPU 41 reads the marking program 50 from the storage device 44, expands it in the memory 43, and sequentially executes the control instructions included in the marking program 50. By executing the detection process control instruction 52, the CPU 41 operates as the detection unit 12 shown in FIG. 2. Also, by executing the selection process control instruction 54, the CPU 41 operates as the selection unit 14 shown in FIG. 2. Further, by executing the division process control instruction 56, the CPU 41 operates as the division unit 16 shown in FIG. 2. Additionally, by executing the assignment process control instruction 58, the CPU 41 operates as the assignment unit 18 shown in FIG. 2. Moreover, by executing the input process control instruction 60, the CPU 41 operates as the input unit 20 shown in FIG. 2. Also, the CPU 41 reads information from the information storage area 70 and expands each of the object detection model 32 and the region division model 34 in the memory 43. As a result, the computer 40 that has executed the marking program 50 functions as the marking device 10. Note that the CPU 41 that executes the program is hardware. Also, a part of the program may be executed by the GPU 42.
[0032] Note that the functions realized by the marking program 50 may be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or the like.
[0033] Next, the operation of the marking device 10 according to the present embodiment will be described. When an original image is input to the marking device 10, the marking process shown in FIG. 5 is executed in the marking device 10. Note that the marking process is an example of the marking method of the disclosed technology.
[0034] In step S10, the detection unit 12 acquires the original image input to the marking device 10. Next, in step S12, the detection unit 12 inputs the acquired original image to the object detection model 32 and detects an object that matches an object including a type of character defined in advance as the type of character to be used for marking.
[0035] Next, in step S14, the selection unit 14 counts the number of detected objects for each type of character based on the object detection result in step S12 above. Next, in step S16, the selection unit 14 selects the type of character with the smallest detection count as the type of character to be used for marking. Next, in step S18, the division unit 16 applies the region division model 34 to the original image, semantically divides the original image into regions, and superimposes the result of the region division on the original image.
[0036] Next, in step S20, the assignment unit 18 assigns to each region the character type selected so that different markings are assigned to each region divided by region from the type selected in step S16 above. Then, the assignment unit 18 generates an input image in which the marking is further superimposed on the image in which the result of the region division is superimposed on the original image. Next, in step S22, the input unit 20 acquires the text prompt input from the user, inputs the input image generated in step S20 together with the text prompt into the LMMs 32, and the marking process ends.
[0037] For example, assume that a text prompt such as "Which button should I press next?" is input into the LMMs 32 together with the input image shown in the right figure of FIG. 3. In this case, unlike the example of FIG. 1, confusion between the characters and the markings included in the original image is avoided, and for example, the correct answer such as "Please press the second button." is output from the LMMs 32.
[0038] As described above, the marking device according to the present embodiment detects characters from the original image, counts the detected characters for each type, and selects the type of characters with the smallest detected number as the type of characters to be used for marking. Further, the marking device divides the original image into regions corresponding to objects, and assigns characters of the selected type as markings to each of the divided regions of the original image. Then, the marking device inputs an input image in which the region division result and the applied marking are superimposed on the original image, together with the text input from the user, to the LMMs. Thereby, it is possible to apply a marking that avoids confusion with the characters included in the original image.
[0039] Here, an example of the effect of the present embodiment will be described when compared with a reference example in which the type of marking is not selected according to the characters included in the original image. FIG. 6 shows an example of the marking result in the reference example. As shown in the left diagram of FIG. 6, the original image includes circled numbers indicating each part of the notebook computer. In this reference example, as shown in the right diagram of FIG. 6, it is assumed that circled numbers similar to the circled numbers included in the original image are used as the marking.
[0040] When this input image was input to the LMMs together with the text prompt "What is above the keyboard?", the answer of the LMMs was "Above the numeric keypad, there is a liquid crystal display. The position is the position of '2' in the image." Thus, in the reference example, confusion occurs between the circled numbers included in the original image and the circled numbers applied as the marking in the input image for input to the LMMs, and an improvement in the image understanding accuracy of the LMMs in the visual prompt cannot be expected.
[0041] On the other hand, in the present embodiment, it is assumed that numbers and alphabets are prepared as the types of markings. In this case, when comparing the number of detected numbers and alphabets from the original image as shown in the left diagram of FIG. 7 by the object detection algorithm, since the number of detected alphabets is smaller, the alphabet is selected as the type of marking in the present embodiment.
[0042] Then, as shown in the right figure of FIG. 7, when an input image with alphabet markings superimposed and the same text prompt as above are input into the LMMs, the answer of the LMMs is "Above the keypad, there is a liquid crystal display. The position is the position of 'a' in the image." Thus, in this embodiment, by applying markings that avoid confusion with the characters included in the original image, the accuracy of image understanding of the LMMs in the visual prompt can be improved.
[0043] Note that the color separation (difference in the density of hatching) in the right figures of FIGS. 6 and 7 represents the result of region division.
[0044] Also, in the above embodiment, the marking program is pre-stored (installed) in the storage device, but it is not limited to this. The program according to the disclosed technology may be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, USB memory, etc.
[0045] Regarding the above embodiments, the following additional notes are further disclosed.
[0046] (Additional Note 1) Detect characters or symbols from the first image, Count the detected characters or symbols for each type, and select the type of characters or symbols with the smallest detected number as the type of characters or symbols to be used for marking, Divide the first image into regions corresponding to objects, For each of the divided regions of the first image, assign the selected type of characters or symbols as markings, Input a second image in which the region division result and the assigned markings are superimposed on the first image, together with the text input by the user, into a large-scale multimodal model, A marking method in which a computer executes a process including this.
[0047] (Additional Note 2) Detecting the character or symbol from the first image includes detecting a character or symbol of a type predetermined as the character or symbol to be used for the marking, according to the marking method described in Supplementary Note 1.
[0048] (Supplementary Note 3) Among the characters or symbols of the type with the smallest detected number, the character or symbol of the type with a detection number of 0 among the characters or symbols of the predetermined type is included, according to the marking method described in Supplementary Note 2.
[0049] (Supplementary Note 4) Detecting the character or symbol from the first image includes applying an object detection algorithm to the first image to detect an object that matches an object including the character or symbol of the predetermined type, according to the marking method described in Supplementary Note 2 or Supplementary Note 3.
[0050] (Supplementary Note 5) The large-scale multimodal model is a machine learning model capable of recognizing the type of character or symbol to be used for the marking, according to the marking method described in any one of Supplementary Notes 1 to 4.
[0051] (Supplementary Note 6) Detect a character or symbol from the first image, Count the detected characters or symbols for each type, select the type of character or symbol with the smallest detected number as the type of character or symbol to be used for the marking, Divide the first image into regions corresponding to objects, For each of the divided regions of the first image, assign the selected type of character or symbol as a marking, Input a second image in which the region division result and the applied marking of the first image are superimposed, together with the text input by the user, into a large-scale multimodal model. A marking program for causing a computer to execute a process including this.
[0052] (Supplementary Note 7) Detecting the character or symbol from the first image includes detecting a character or symbol of a type predetermined as the character or symbol to be used for the marking, according to the marking program described in Supplementary Note 6.
[0053] (Supplementary Note 8) Among the characters or symbols of the type with the smallest detected number, the character or symbol of the type with a detection number of 0 among the characters or symbols of the predetermined type is included, according to the marking program described in Supplementary Note 7.
[0054] (Supplementary Note 9) Detecting the character or symbol from the first image includes applying an object detection algorithm to the first image to detect an object that matches an object including the character or symbol of the predetermined type, according to the marking program described in Supplementary Note 7 or Supplementary Note 8.
[0055] (Supplementary Note 10) The large-scale multimodal model is a machine learning model capable of recognizing the type of character or symbol to be used for the marking, according to the marking program described in any one of Supplementary Notes 6 to 9.
[0056] (Supplementary Note 11) A detection unit that detects a character or symbol from a first image, A selection unit that counts the detected characters or symbols for each type and selects the character or symbol of the type with the smallest detected number as the type of character or symbol to be used for the marking, A division unit that divides the first image into regions corresponding to objects, An assignment unit that assigns the selected character or symbol of the type as a marking to each of the divided regions of the first image, An input unit that inputs a second image in which the division result of the region and the assigned marking are superimposed on the first image, together with the text input from the user, into a large-scale multimodal model, A marking device including the above.
[0057] (Supplementary Note 12) The marking device according to Supplementary Note 11, wherein the detection unit detects characters or symbols of a type predetermined as the characters or symbols to be used for the marking.
[0058] (Supplementary Note 13) The marking device according to Supplementary Note 12, wherein the characters or symbols of the type with the smallest detected number include the characters or symbols of the type with a detection number of 0 among the characters or symbols of the type predetermined as the characters or symbols to be used for the marking.
[0059] (Supplementary Note 14) The marking device according to Supplementary Note 12 or Supplementary Note 13, wherein the detection unit applies an object detection algorithm to the first image to detect an object that matches an object including the characters or symbols of the type predetermined as the characters or symbols to be used for the marking.
[0060] (Supplementary Note 15) The marking device according to any one of Supplementary Notes 11 to 14, wherein the large-scale multimodal model is a machine learning model capable of recognizing the characters or symbols of the type to be used for the marking.
[0061] (Supplementary Note 16) Detect characters or symbols from the first image, Count the detected characters or symbols for each type, select the characters or symbols of the type with the smallest detected number as the type of characters or symbols to be used for the marking, Divide the first image into regions corresponding to objects, To each of the divided regions of the first image, assign the selected characters or symbols of the type as a marking, Input a second image in which the region division result and the applied marking of the first image are superimposed, together with the text input from the user, into a large-scale multimodal model, A non-transitory storage medium storing a marking program for causing a computer to execute a process including the above.
Explanation of Signs
[0062] 10 Marking device 12 Detection unit 14 Selection Unit 16 Division Unit 18 Assignment Unit 20 Input Unit 32 Object Detection Model 34 Region Division Model 40 Computer 41 CPU 42 GPU 43 Memory 44 Storage Device 45 Input / Output Device 46 R / W Device 47 Communication I / F 48 Bus 49 Storage Medium 50 Marking Program 52 Detection Process Control Instruction 54 Selection Process Control Instruction 56 Division Process Control Instruction 58 Assignment Process Control Instruction 60 Input Process Control Instruction 70 Information Storage Area
Claims
1. Detect characters or symbols from the first image, count the detected characters or symbols by type, and select the type of characters or symbols with the smallest detected number as the type of characters or symbols to be used for marking, divide the first image into regions corresponding to objects, assign the selected type of characters or symbols as markings to each of the divided regions of the first image, input a second image in which the region division result and the applied markings of the first image are superimposed, together with the text input by the user, into a large-scale multimodal model, A marking method in which a computer executes a process including this.
2. Detecting the characters or symbols from the first image includes detecting characters or symbols of a type predetermined as the characters or symbols to be used for marking. The marking method according to claim 1.
3. Among the characters or symbols of the type with the smallest detected number, the marking method according to claim 2 includes characters or symbols of a type with a detection number of 0 among the characters or symbols of the predetermined type.
4. Detecting the characters or symbols from the first image includes applying an object detection algorithm to the first image to detect an object that matches an object including the characters or symbols of the predetermined type. The marking method according to claim 2 or claim 3.
5. The large-scale multimodal model is a machine learning model capable of recognizing the type of characters or symbols to be used for marking. The marking method according to any one of claims 1 to 3.
6. Detect characters or symbols from the first image, count the detected characters or symbols by type, and select the type of characters or symbols with the smallest detected number as the type of characters or symbols to be used for marking, divide the first image into regions corresponding to objects, assign the selected type of characters or symbols as markings to each of the divided regions of the first image, input a second image in which the region division result and the applied markings of the first image are superimposed, together with the text input by the user, into a large-scale multimodal model, A marking program for causing a computer to execute a process including this.
Citation Information
Patent Citations
Document processor
JP1993342326A