Retrieval device and control method

By analyzing document structure using font and frame line information, the search device accurately identifies relevant sentences, addressing the issue of incorrect paragraph detection in conventional systems and enhancing information retrieval precision.

WO2025109775A1PCT designated stage expired Publication Date: 2025-05-30PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Patent Information

Application Number
PCT/JP2024/013012
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-03-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional instruction manual search devices may incorrectly detect visually separate sentences as different paragraphs due to their distance from each other.

Method used

A search device that analyzes document structure based on font information and frame line information to accurately identify and treat relevant sentences as part of the same paragraph.

Benefits of technology

This approach allows for accurate analysis of document structure, ensuring that distantly separated sentences are correctly identified as relevant, thereby improving the precision of information retrieval from document data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024013012_30052025_PF_FP_ABST
    Figure JP2024013012_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A retrieval device (4) comprises an acquisition unit (14) that acquires document data and an analysis unit (16) that analyzes the document structure of the document data acquired by the acquisition unit (14) on the basis of font information related to the font of characters included in the document data and frame line information related to frame lines included in the document data.
Need to check novelty before this filing date? Find Prior Art

Description

Search device and control method

[0001] The present disclosure relates to a search device for searching document data for information desired by a user, and a control method thereof.

[0002] There is known an instruction manual search device that, when a user inputs a question about an instruction manual for a household electrical appliance or the like into a smartphone or the like, highlights and displays the section of the instruction manual that is the basis for the answer on the smartphone or the like (see, for example, Patent Document 1). This instruction manual search device acquires document data indicating the instruction manual in advance, and analyzes the document structure of the document data by detecting paragraphs based on the coordinates of characters included in the acquired document data.

[0003] Japanese Patent Application Laid-Open No. 2006-48605

[0004] In the conventional instruction manual search device described above, when analyzing the document structure of document data, there is a risk that two sentences that are distantly separated but visually appear to be one paragraph may be mistakenly detected as separate paragraphs.

[0005] Therefore, the present disclosure provides a search device and a control method thereof that can accurately analyze the document structure of document data.

[0006] A search device according to one aspect of the present disclosure is a search device for searching for information desired by a user from document data, and includes an acquisition unit that acquires the document data, and an analysis unit that analyzes the document structure of the document data based on font information regarding the fonts of characters included in the document data acquired by the acquisition unit, and border information regarding border lines included in the document data.

[0007] These comprehensive or specific aspects may be realized by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM (Compact Disc-Read Only Memory), or may be realized by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0008] According to a search device or the like according to an aspect of the present disclosure, the document structure of document data can be analyzed with high accuracy.

[0009] 19 is a conceptual diagram showing an overview of a search system according to an embodiment. 20 is a block diagram showing the functional configuration of a search device according to an embodiment. 21 is a flowchart showing the overall operation of the search device according to an embodiment. 22 is a flowchart showing details of step S1 in the flowchart of FIG. 3. 23 is a flowchart showing details of step S14 in the flowchart of FIG. 4. 24 is a flowchart showing details of step S142 in the flowchart of FIG. 5. 25 is a flowchart showing details of step S143 in the flowchart of FIG. 5. 26 is a flowchart showing details of step S1431 in the flowchart of FIG. 7. 27 is a diagram showing an example of a border line detected in document data. 28 is a flowchart showing details of step S1432 in the flowchart of FIG. 7. 28 is a diagram showing an example of a frame estimated in document data. 29 is a flowchart showing details of step S16 in the flowchart of FIG. 4. 29 is a diagram for explaining the content of step S16 in the flowchart of FIG. 4. 29 is a flowchart showing details of step S17 in the flowchart of FIG. 4. 29 is a diagram for explaining the content of step S17 in the flowchart of FIG. 4. 29 is a flowchart showing details of step S18 in the flowchart of FIG. 4. 29 is a flowchart showing an example of a procedure for creating a dense vector model. 29 is a flowchart showing an example of a procedure for creating a sparse vector model. 29 is a flowchart showing details of step S19 in the flowchart of FIG. 4. 29 is a diagram for explaining the content of the flowchart of FIG. 29. 25 is a flowchart showing details of step S20 in the flowchart of FIG. 4. It is a diagram showing an example of document data of an instruction manual. It is a diagram showing an example of an analysis result of the document data by the analysis unit. It is a flowchart showing details of step S2 in the flowchart of FIG. 3. It is a diagram for explaining the contents ... flowchart showing details of step S26 in the flowchart of FIG. 25. It is a flowchart showing details of step S27 in the flowchart of FIG. 25. It is a flowchart showing details of step S3 in the flowchart of FIG. 3.10A and 10B are diagrams for explaining another example of presentation of search results by the presentation unit, and FIG. 10C are diagrams for explaining another example of presentation of search results by the presentation unit.

[0010] (Technology 1) A search device for searching for information desired by a user from document data, comprising: an acquisition unit that acquires the document data; and an analysis unit that analyzes the document structure of the document data based on font information regarding the fonts of characters included in the document data acquired by the acquisition unit, and border information regarding borders included in the document data.

[0011] According to Technique 1, the analysis unit analyzes the document structure of document data based on font information and border information. As a result, even if two or more sentences are distant from each other, the two or more sentences can be treated as highly related to each other based on the font information and border information. As a result, the document structure of document data can be analyzed with high accuracy.

[0012] (Technology 2) The search device described in Technology 1, wherein the analysis unit (i) estimates, based on the font information, a set of characters that have the same font and whose adjacent distance in a first direction is shorter than the font size as a line in the document data, and (ii) estimates, based on the font information, a set of lines whose adjacent distance in a second direction orthogonal to the first direction is shorter than the font size as a frame in the document data, thereby analyzing the document structure of the document data.

[0013] According to Technique 2, by estimating frames in document data, two or more sentences in a frame can be treated as sentences that are highly related to each other, thereby enabling the document structure of document data to be analyzed with high accuracy.

[0014] (Technology 3) The search device according to Technology 2, wherein the analysis unit (i) detects the plurality of line segments as the frame line by acquiring, as the frame line information, coordinates of each of the plurality of line segments that form a closed area among the plurality of line segments included in the document data, and (ii) estimates, based on the frame line information, that a set of the lines included inside the frame line is the same frame.

[0015] According to Technique 3, by estimating that a set of lines included inside a frame belongs to the same frame, two or more sentences within the frame can be treated as sentences that are highly related to each other. As a result, the document structure of document data can be analyzed with high accuracy.

[0016] (Technology 4) The search device according to Technology 2 or 3, wherein the document data is written horizontally, and when the analysis unit estimates multiple frames, it assigns a first reading order to the multiple frames, which is the order in which the multiple frames are to be read from left to right and from top to bottom.

[0017] According to Technique 4, by assigning a first reading order to a plurality of frames, the semantic connections between the plurality of frames are taken into consideration, and the document structure of the document data can be analyzed with higher accuracy.

[0018] (Technology 5) The search device according to Technology 4, wherein the analysis unit further assigns a second reading order to the lines in the frame, the second reading order being an order in which the lines are read from left to right and from top to bottom, when the analysis unit estimates the lines in the frame.

[0019] According to technique 5, by assigning a second reading order to multiple lines within a frame, the semantic connections between the multiple lines within the frame are taken into account, enabling more accurate analysis of the document structure of the document data.

[0020] (Technology 6) The search device according to any one of Technologies 1 to 5, further comprising: an input unit that accepts a search query input by the user; a first vector generation unit that generates a first vector by vectorizing the document data analyzed by the analysis unit; a second vector generation unit that generates a second vector by vectorizing the search query accepted by the input unit; and a search unit that calculates a similarity between the first vector and the second vector and extracts a specific character string from the document data as a search result for the search query based on the calculated similarity.

[0021] According to the sixth technique, as described above, the document structure of the document data can be analyzed with high accuracy, and therefore the accuracy of the search performed by the search unit can be improved.

[0022] (Technology 7) A search device according to Technology 6, wherein the first vector generation unit generates a first sparse vector and a first dense vector as the first vector, the second vector generation unit generates a second sparse vector and a second dense vector as the second vector, and the search unit (i) calculates a first similarity between the first sparse vector and the second sparse vector as the similarity when the user is still in the middle of inputting the search query, and (ii) calculates a second similarity between the first dense vector and the second dense vector as the similarity when the user has completed inputting the search query.

[0023] According to Technology 7, when a user is in the middle of inputting a search query, a first similarity between a first sparse vector and a second sparse vector is calculated, thereby performing, for example, an exact word match search. This provides a high level of search convincing for users who are familiar with the content of the document data and therefore capable of inputting an appropriate search query. On the other hand, when the user has completed inputting the search query, a second similarity between the first dense vector and the second dense vector is calculated, thereby enabling search results that do not include exact matches. This provides a high level of search convincing for users who are not familiar with the content of the document data and therefore have difficulty inputting an appropriate search query. As described above, by switching between sparse vector search and dense vector search during and after inputting a search query, the convincingness of the search can be increased for all users.

[0024] (Technology 8) The search device according to Technology 6 or 7, further comprising a presentation unit that presents search results of the search unit to the user, the presentation unit highlighting the specific character string extracted by the search unit in a display mode according to the similarity.

[0025] According to technique 8, specific character strings extracted by the search unit are highlighted in a display mode according to the degree of similarity, making it easier for the user to recognize noteworthy search results.

[0026] (Technology 9) The search device according to Technology 8, wherein the presentation unit highlights the specific character string in a display manner in which the color becomes darker as the similarity increases.

[0027] According to Technique 9, by highlighting a specific character string in a display mode in which the color becomes darker as the similarity increases, the user can more easily recognize noteworthy search results.

[0028] (Technology 10) The search device according to Technology 6 or 7, further comprising a presentation unit that presents search results of the search unit to the user, generates a summary of an answer based on the search results of the search unit, and displays the generated summary.

[0029] According to technique 10, by displaying a summary of the answer, the user can easily understand the content of the answer.

[0030] (Technology 11) A control method for a search device for searching for information desired by a user from document data, the control method comprising: (a) a step of acquiring the document data; and (b) a step of analyzing the document structure of the document data based on font information regarding the fonts of characters included in the document data acquired in (a) and border information regarding borders included in the document data.

[0031] According to Technique 11, the document structure of document data is analyzed based on font information and border information. As a result, even if two or more sentences are distant from each other, these two or more sentences can be treated as highly related to each other based on the font information and border information. As a result, the document structure of document data can be analyzed with high accuracy.

[0032] (Technology 12) A program that causes a computer to execute the control method according to Technology 11.

[0033] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, or a recording medium.

[0034] Hereinafter, the embodiments will be specifically described with reference to the drawings.

[0035] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components.

[0036] (Embodiment) [1. Overview of Search System] First, an overview of a search system 2 according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a conceptual diagram showing an overview of a search system 2 according to an embodiment.

[0037] 1, the search system 2 includes a search device 4 and a terminal device 6. The search device 4 and the terminal device 6 are capable of communicating with each other via a network 8 such as the Internet.

[0038] The search device 4 is a device for searching for information desired by the user 12 from document data representing the instruction manual 10, and is configured, for example, as a server. The search device 4 acquires document data representing the instruction manual 10 in advance from an external server or the like (not shown), and analyzes the document structure of the acquired document data. The document data is, for example, data in PDF (Portable Document Format), and in this embodiment is written horizontally. The instruction manual 10 is, for example, a manual for a household electrical appliance used by the user 12. The household electrical appliance is a device used by the user 12, such as a refrigerator / freezer, a liquid crystal television receiver, a washing machine, an air conditioner, a microwave oven, an electric rice cooker, an electric shaver, etc.

[0039] The terminal device 6 is a mobile terminal such as a smartphone or a tablet terminal, and is operated by the user 12. The terminal device 6 may be a desktop or laptop personal computer.

[0040] For example, when a user 12 wants to find out how to use a household electrical appliance, the user 12 operates the terminal device 6 to launch an application dedicated to the search system 2. Next, the user 12 inputs a search query, which is a character string representing a question about the instruction manual 10 for the household electrical appliance, on a search screen in the application. Note that in this specification, a "character string" is a set of characters, and is a concept that includes, for example, a word, a sentence, and a paragraph.

[0041] As a result, the search device 4 searches for the part that serves as the basis for the answer from document data whose document structure has been analyzed in advance, based on the search query entered by the user 12, and presents the search results to the user 12 by displaying them on a smartphone or the like.

[0042] 2. Functional Configuration of Search Device Next, the functional configuration of the search device 4 according to the embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the functional configuration of the search device 4 according to the embodiment.

[0043] As shown in FIG. 2, the search device 4 has, as its functional configuration, an acquisition unit 14, an analysis unit 16, a first vector generation unit 18, a memory unit 20, an input unit 22, a detection unit 24, a second vector generation unit 26, a search unit 28, a presentation content generation unit 30, and a presentation unit 32.

[0044] The acquisition unit 14 acquires document data indicating the instruction manual 10 (see FIG. 1 ) from an external server, etc. The acquisition unit 14 outputs the acquired document data to the analysis unit 16 .

[0045] The analysis unit 16 analyzes the document structure of the document data based on font information relating to the font (typeface) of characters included in the document data and frame information relating to frame lines included in the document data acquired by the acquisition unit 14. The analysis unit 16 analyzes the document structure of the document data to acquire frames, which are units of character strings in the document data.

[0046] The first vector generation unit 18 generates a first vector by vectorizing the document data analyzed by the analysis unit 16. More specifically, the first vector generation unit 18 sparsely vectorizes the document data analyzed by the analysis unit 16, thereby generating a first sparse vector as the first vector. Here, a "sparse vector" is a sparse vector in which the elements corresponding to words included in the document data are almost zero. Furthermore, the first vector generation unit 18 densely vectorizes the document data analyzed by the analysis unit 16, thereby generating a first dense vector as the first vector. Here, a "dense vector" is a dense vector in which the elements corresponding to words included in the document data are almost non-zero, and is generated, for example, by deep learning.

[0047] The storage unit 20 is a memory that stores the document data acquired by the acquisition unit 14 , the analysis results of the analysis unit 16 , and the first vector generated by the first vector generation unit 18 .

[0048] The input unit 22 accepts a search query input by the user 12 (see FIG. 1).

[0049] The detection unit 24 detects the input state of the search query by the user 12 based on the reception result of the input unit 22. The detection unit 24 detects, as the input state of the search query, any of (a) input in progress, (b) input complete, and (c) no input.

[0050] The second vector generation unit 26 generates a second vector by vectorizing the search query received by the input unit 22. More specifically, the second vector generation unit 26 generates a second sparse vector as the second vector by sparsely vectorizing the search query received by the input unit 22. Furthermore, the second vector generation unit 26 generates a second dense vector as the second vector by densely vectorizing the search query received by the input unit 22.

[0051] The search unit 28 searches the document data analyzed by the analysis unit 16 for a portion that serves as the basis for an answer to the search query. Specifically, the search unit 28 calculates the similarity between the first vector generated by the first vector generation unit 18 and the second vector generated by the second vector generation unit 26, and extracts a specific character string from the document data as a search result for the search query based on the calculated similarity. Here, if the input state of the search query detected by the detection unit 24 is "input in progress," the search unit 28 calculates the similarity (an example of a first similarity) between the first sparse vector generated by the first vector generation unit 18 and the second sparse vector generated by the second vector generation unit 26. Furthermore, if the input state of the search query detected by the detection unit 24 is "input completed," the search unit 28 calculates the similarity (an example of a second similarity) between the first dense vector generated by the first vector generation unit 18 and the second dense vector generated by the second vector generation unit 26.

[0052] The presentation content generator 30 generates, as presentation content to be presented to the user 12, the specific character string extracted as a search result by the search unit 28.

[0053] The presentation unit 32 displays the presentation content generated by the presentation content generation unit 30 on the terminal device 6 (see FIG. 1 ), thereby presenting the presentation content to the user 12. Specifically, the presentation unit 32 highlights the specific character string extracted by the search unit 28 in a display mode according to the degree of similarity. More specifically, the presentation unit 32 highlights the specific character string in a display mode in which the color becomes darker as the degree of similarity increases.

[0054] 3. Operation of Search Device Hereinafter, the operation of the search device 4 (method of controlling the search device 4) according to the embodiment will be described with reference to FIGS.

[0055] [3-1. Overall Operation of Search Device] The overall operation of the search device 4 according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the flow of the overall operation of the search device 4 according to the embodiment.

[0056] As shown in Figure 3, first, the analysis unit 16 analyzes the document structure of the document data based on font information regarding the font of the characters included in the document data of the instruction manual 10 and border information regarding the border lines included in the document data (S1).

[0057] Next, the search unit 28 searches the document data analyzed by the analysis unit 16 for a portion that serves as the basis for the answer to the search query entered by the user 12 (S2).

[0058] Next, the presentation unit 32 presents the specific character string extracted as a search result by the search unit 28 to the user 12 (S3).

[0059] [3-2. Details of Step S1] Details of step S1 (analyzing document structure) in the flowchart of Fig. 3 will be described with reference to Fig. 4. Fig. 4 is a flowchart showing details of step S1 in the flowchart of Fig. 3.

[0060] 4, first, the acquisition unit 14 acquires document data of the instruction manual 10 from an external server or the like (S11). The acquisition unit 14 outputs the acquired document data to the analysis unit 16. In this embodiment, it is assumed that the document data includes characters, tables, and figures.

[0061] Next, the analysis unit 16 acquires layout elements from the document data acquired by the acquisition unit 14 (S12). Specifically, the analysis unit 16 acquires characters, tables, figures, and their coordinates as layout elements from the binary data of the document data and pixel data obtained by imaging the document data, using, for example, optical character recognition technology or the like.

[0062] If the layout element is a character ("character" in S13), the analysis unit 16 performs character analysis (S14) to obtain a character string for the character layout element. Next, if analysis processing has not been performed for all layout elements (NO in S15), the process returns to step S13.

[0063] In step S13, if the layout element is a table ("table" in S13), the analysis unit 16 performs table analysis (S16) and obtains the character string for the table layout element. Next, if the analysis process has not been performed for all layout elements (NO in S15), the process returns to step S13.

[0064] In step S13, if the layout element is a diagram ("Diagram" in S13), the analysis unit 16 executes diagram analysis (S17) and obtains a character string for the diagram layout element.

[0065] Next, if the analysis process has been performed on all layout elements (YES in S15), the first vector generation unit 18 generates first vectors (first sparse vectors and first dense vectors) by vectorizing (S18) the character strings for each layout element acquired by the analysis unit 16. The first vector generation unit 18 stores the generated first vectors in the storage unit 20.

[0066] Next, the analysis unit 16 performs a spread determination (S19) to determine whether the page included in the document data acquired by the acquisition unit 14 is a spread page. A spread page is a page where two pages, left and right, are combined into one, and it is expected that it would be desirable to assign a reading order (described later) to each of the left and right pages.

[0067] Next, the analysis unit 16 estimates a reading order, which is the order in which the layout elements are to be read (S20). The analysis unit 16 stores the estimated reading order in the storage unit 20.

[0068] In the flowchart shown in FIG. 4, if the document data does not include tables and figures, steps S16 and S17 may be omitted.

[0069] [3-2-1. Details of Step S14] Details of step S14 (character analysis) in the flowchart of Fig. 4 will be described with reference to Fig. 5. Fig. 5 is a flowchart showing details of step S14 in the flowchart of Fig. 4.

[0070] 5, first, the analysis unit 16 acquires font information related to the font of characters included in the document data, for example, by using a document analysis program or optical character recognition technology (S141). Specifically, the analysis unit 16 acquires, as the font information, the characters, the font type of the characters, the font size of the characters, and the coordinates of the characters.

[0071] Next, the analysis unit 16 estimates, based on the acquired font information, a group of characters that have the same font and whose adjacent characters are spaced apart in the horizontal direction (an example of a first direction) by less than the font size as a line in the document data (S142). Note that if the document data is written vertically, the analysis unit 16 may estimate, based on the acquired font information, a group of characters that have the same font and whose adjacent characters are spaced apart in the vertical direction (an example of a first direction) by less than the font size as a line in the document data.

[0072] Next, the analysis unit 16 estimates a set of lines whose adjacent distance in the vertical direction (an example of a second direction orthogonal to the first direction) is shorter than the font size as a frame in the document data (S143). Note that if the document data is written vertically, the analysis unit 16 may estimate a set of lines whose adjacent distance in the horizontal direction (an example of a second direction orthogonal to the first direction) is shorter than the font size as a frame in the document data.

[0073] In the flowchart of FIG. 5, after step S141 is executed, only one of steps S142 and S143 may be executed.

[0074] [3-2-1-1. Details of Step S142] Details of step S142 (estimating a row) in the flowchart of Fig. 5 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing details of step S142 in the flowchart of Fig. 5.

[0075] 6, the analysis unit 16 first searches for a reference character in the document data (S1421). Here, the reference character is a character that is not included in any previously displayed line and is located at the top left corner of the entire page. The position of the character is based on the coordinates of the top left corner of a rectangle circumscribing the character.

[0076] Next, if the reference character does not exist (NO in S1422), the analysis unit 16 ends the line estimation process.

[0077] On the other hand, if the reference character exists (YES in S1422), the analysis unit 16 searches for the character that is closest to the reference character in the horizontal direction (hereinafter referred to as the "adjacent character") (S1423).

[0078] If an adjacent character exists (YES in S1424), the analysis unit 16 determines whether the adjacent character and the reference character exist on the same line (S1425). Specifically, the analysis unit 16 determines that the adjacent character and the reference character exist on the same line if (a) the adjacent character and the reference character are of the same font type, (b) the adjacent character and the reference character are of the same font size, and (c) the vertical positional deviation between the adjacent character and the reference character is less than a threshold value with respect to the font size of the reference character. Note that the threshold value is, for example, "0.1" in the vertical direction when the font size of the reference character is "1."

[0079] If the adjacent character and the reference character are on the same line (YES in S1425), the process returns to step S1423, and steps S1423 to S1425 are executed again.

[0080] On the other hand, if the adjacent character and the reference character are not on the same line (NO in S1425), the analysis unit 16 updates the line to be estimated (S1426). Then, the process returns to step S1421, and steps S1421 to S1425 are executed for the updated line.

[0081] Returning to step S1424, if there is no adjacent character (NO in S1424), the process proceeds to step S1426 described above.

[0082] [3-2-1-2. Details of Step S143] Details of step S143 (estimating a frame) in the flowchart of Fig. 5 will be described with reference to Figs. 7 to 11. Fig. 7 is a flowchart showing details of step S143 in the flowchart of Fig. 5. Fig. 8 is a flowchart showing details of step S1431 in the flowchart of Fig. 7. Fig. 9 is a diagram showing an example of a frame detected in document data. Fig. 10 is a flowchart showing details of step S1432 in the flowchart of Fig. 7. Fig. 11 is a diagram showing an example of a frame estimated in document data.

[0083] As shown in FIG. 7 , the analysis unit 16 first detects frame lines in the document data (S1431). Specifically, as shown in FIG. 8 , the analysis unit 16 first acquires lines (hereinafter referred to as "line segments") and their coordinates contained in the document data using, for example, a document analysis program (S14311). In this specification, "line segments" includes not only straight lines connecting two points but also curved lines connecting two points. When the analysis unit 16 detects a curved line in the document data, it divides the curve into multiple dividing lines of a length that can be considered as a straight line, and detects the start and end points of each of the multiple dividing lines, thereby treating each of the multiple dividing lines as a line segment. The analysis unit 16 may also acquire boundaries where the background color changes or the boundaries of the document data itself as line segments.

[0084] Next, the analysis unit 16 detects dashed lines in the document data (S14312). Specifically, the analysis unit 16 detects a set of relatively short line segments arranged in a straight line in the document data as dashed lines, and acquires the detected dashed lines as line segments. Note that if two line segments arranged in a straight line are adjacent to each other at a distance equal to or less than a threshold as a percentage of the line segment length, the analysis unit 16 determines that they are part of the same dashed line.

[0085] Next, the analysis unit 16 extends the start point and end point of the line segment by a predetermined value (for example, 1% of the length of the line segment) (S14313).

[0086] Next, if the line segments intersect with each other, the analysis unit 16 divides each line segment at the intersection point (S14314) and treats it as a separate intersection point.

[0087] Next, the analysis unit 16 acquires, as frame lines, a plurality of line segments that form a closed area from among the plurality of line segments included in the document data (S14315). That is, the analysis unit 16 acquires, as frame line information, the coordinates of each of the plurality of line segments that form a closed area from among the plurality of line segments included in the document data, thereby detecting the plurality of line segments as frame lines.

[0088] For example, if there are multiple line segments forming a closed area in the document data, as shown by the thin solid lines in (a) of Figure 9, the analysis unit 16 detects the multiple line segments as frame lines, as shown by the thick solid lines in (b) of Figure 9.

[0089] Returning to Fig. 7, after step S1431, the analysis unit 16 estimates the frame within the frame line (S1432). Specifically, as shown in Fig. 10, the analysis unit 16 first searches for a reference line within the frame of interest in the document data (S14321). Here, the reference line is the line that is not included in any of the previously-discussed frames and is located at the top left corner of the entire page. The position of the line is based on the coordinates of the top left corner of the rectangle circumscribing the line.

[0090] Next, if there is no reference row (NO in S14322), the analysis unit 16 ends the estimation process of the frame within the frame line.

[0091] On the other hand, if the reference row exists (YES in S14322), the analysis unit 16 searches for the row that is closest to the reference row in the vertical direction (hereinafter referred to as the "adjacent row") (S14323).

[0092] If an adjacent line exists (YES in S14324), the analysis unit 16 determines whether the adjacent line and the reference line exist in the same frame (S14325). Specifically, the analysis unit 16 determines that the adjacent line and the reference line exist in the same frame if the vertical positional deviation between the adjacent line and the reference line is less than a threshold value for the font size of the characters that make up the reference line. Note that the threshold value is, for example, "0.1" in the vertical direction when the font size of the characters that make up the reference line is "1."

[0093] If the adjacent row and the reference row are in the same frame (YES in S14325), the process returns to step S14323 and steps S14323 to S14325 are executed again. That is, the analysis unit 16 estimates that the set of rows inside the frame line are in the same frame.

[0094] On the other hand, if the adjacent row and the reference row are not in the same row (NO in S14325), the analysis unit 16 updates the frame to be estimated (S14326). Then, the process returns to step S14321, and steps S14321 to S14325 are executed for the updated frame.

[0095] Returning to step S14324, if there is no adjacent row (NO in S14324), the analysis unit 16 proceeds to step S14326 described above. Note that the analysis unit 16 also estimates an area not surrounded by a frame line as one frame, and executes the same processing as described above.

[0096] When document data includes a character string and a frame surrounding the character string as shown in Fig. 11(a), the analysis unit 16 estimates that the character strings included inside the frame constitute the same frame as shown in Fig. 11(b). Note that the multiple frame lines shown in Fig. 11(a) correspond to the multiple frames shown in Fig. 11(b). For example, the frame surrounding the character string "symptoms" in the upper left corner of Fig. 11(a) corresponds to the frame in Fig. 11(b) with the reading order "1" (described later) in the upper left corner.

[0097] [3-2-2. Details of Step S16] Details of step S16 (table analysis) in the flowchart of Fig. 4 will be described with reference to Fig. 12 and Fig. 13. Fig. 12 is a flowchart showing details of step S16 in the flowchart of Fig. 4. Fig. 13 is a diagram for explaining the contents of step S16 in the flowchart of Fig. 4.

[0098] As shown in FIG. 12, first, the analysis unit 16 acquires a table included in the document data, the coordinates of the table, and the character string included in the table, for example, using a table detection program or image processing (S161).

[0099] Next, the analysis unit 16 converts the character strings obtained from the table into natural sentences (S162). Specifically, the analysis unit 16 converts the character strings obtained from the table into natural sentences by, for example, giving an instruction to the generation AI (Artificial Intelligence) such as "Explain the character strings obtained from the table below in natural language." The analysis unit 16 stores the converted natural sentences in the storage unit 20.

[0100] For example, as shown in Fig. 13(a), when the document data includes a table and character strings included in the table, the analysis unit 16 performs simple character extraction to mechanically extract only the characters from the table, as shown in Fig. 13(b). Next, as shown in Fig. 13(c), the analysis unit 16 converts the characters extracted from the table into natural-sounding sentences such as "The alarm will beep in one minute and the interior light will flash once. In three minutes..."

[0101] [3-2-3. Details of Step S17] Details of step S17 (diagram analysis) in the flowchart of Fig. 4 will be described with reference to Fig. 14 and Fig. 15. Fig. 14 is a flowchart showing details of step S17 in the flowchart of Fig. 4. Fig. 15 is a diagram for explaining the contents of step S17 in the flowchart of Fig. 4.

[0102] As shown in FIG. 14, first, the analysis unit 16 acquires the figure included in the document data, the coordinates of the figure, and the pixel data of the figure, for example, by using image processing or the like (S171).

[0103] Next, the analysis unit 16 converts the acquired pixel data into natural-sounding sentences using, for example, image description AI (S172). The analysis unit 16 stores the converted natural-sounding sentences in the storage unit 20.

[0104] For example, if the document data includes a diagram such as that shown in (a) of Figure 15, the analysis unit 16 converts the diagram into a natural sentence such as "The ice compartment of the freezer is being removed. There is a lot of ice in the ice compartment," as shown in (b) of Figure 15.

[0105] [3-2-4. Details of Step S18] Details of step S18 (vectorization) in the flowchart of Fig. 4 will be described with reference to Fig. 16. Fig. 16 is a flowchart showing details of step S18 in the flowchart of Fig. 4.

[0106] As shown in FIG. 16, first, the first vector generating unit 18 acquires the analysis result of the document data of the instruction manual 10 by the analyzing unit 16 (S181).

[0107] At this time, processing may be performed to divide or combine sentences included in the document data into sentences of an appropriate length suitable for searching. Specifically, for example, (a) processing to divide or combine sentences so as to approach a predetermined number of characters, (b) processing to instruct the generation AI to divide the sentences taking into account semantic coherence, and (c) processing to calculate the similarity between two consecutive sentences using a dense vector or the like and determine whether or not to mark the break in division (i.e., if the similarity with the immediately preceding sentence is low, it is considered to be the beginning of a new section).

[0108] Next, the first vector generation unit 18 tokenizes the character strings included in the document data, that is, divides the character strings included in the document data into tokens (S182).

[0109] Next, the first vector generation unit 18 generates a first sparse vector by converting the document data into a sparse vector using a separately created sparse vector model (S183). The first vector generation unit 18 also generates a first dense vector by converting the document data into a dense vector using a separately created dense vector model (S183). The first vector generation unit 18 stores the generated first sparse vector and first dense vector in the storage unit 20.

[0110] An example of a procedure for creating a dense vector model will now be described with reference to Fig. 17. Fig. 17 is a flowchart showing an example of a procedure for creating a dense vector model.

[0111] 17, first, a pre-training model for generating dense vectors from large-scale document data is created and pre-trained (S41). Note that at this time, a pre-training model based on open source software created by a third party may be used.

[0112] Next, the pre-trained model is additionally trained using a search query or the like (S42). In this manner, a dense vector model is created. Note that step S42 may be omitted in the flowchart shown in FIG.

[0113] Next, an example of a procedure for creating a sparse vector model will be described with reference to Fig. 18. Fig. 18 is a flowchart showing an example of a procedure for generating a sparse vector model.

[0114] As shown in FIG. 18, first, the analysis result of the document data of the instruction manual 10 by the analysis unit 16 is obtained (S51).

[0115] At this time, processing may be performed to divide or combine sentences included in the document data into sentences of an appropriate length suitable for searching. Specifically, for example, (a) processing to divide or combine sentences so as to approach a predetermined number of characters, (b) processing to instruct the generation AI to divide the sentences taking into account semantic coherence, and (c) processing to calculate the similarity between two consecutive sentences using a dense vector or the like and determine whether or not to mark the break in division (i.e., if the similarity with the immediately preceding sentence is low, it is considered to be the beginning of a new section).

[0116] Next, the character strings contained in the document data are tokenized, that is, the character strings contained in the document data are divided into tokens, for example, by a morphological element analyzer (S52).

[0117] Next, a weight representing the importance of each word is calculated based on the frequency of occurrence of the word (S53), and the calculation result is stored. In this manner, a sparse vector model is created. Note that, for example, TF-IDF (Term Frequency-Inverse Document Frequency) or BM25 (Best Match 25) can be used as an index of importance.

[0118] [3-2-5. Details of Step S19] Details of step S19 (spread determination) in the flowchart of Fig. 4 will be described with reference to Fig. 19 to Fig. 21. Fig. 19 is a flowchart showing details of step S19 in the flowchart of Fig. 4. Figs. 20 and 21 are diagrams for explaining the contents of the flowchart of Fig. 19.

[0119] As shown in FIG. 19, first, the analysis unit 16 acquires the size of a page included in the document data acquired by the acquisition unit 14 and the coordinates of a frame included in the page (S191).

[0120] Next, the analysis unit 16 determines whether the page included in the document data is a landscape page (S192).

[0121] As shown in FIG. 20A, if the page included in the document data is a landscape page (YES in S192), the analysis unit 16 determines whether or not there is a frame that spans the center of the page (S193).

[0122] As shown in Fig. 20A, if there is no frame that straddles the middle of the page (indicated by the dashed line) (NO in S193), the analysis unit 16 determines that the page included in the document data is a double-page spread (S194). In this case, as indicated by the arrows in Fig. 20B, the analysis unit 16 estimates the reading order for each of the left and right pages using the procedure described below.

[0123] Returning to step S193, if there is a frame (shown by a dashed line frame) that spans the middle of the page (shown by a dashed line) as shown in (a) of Figure 21 (YES in S193), the analysis unit 16 determines that the page included in the document data is not a double-page spread (S195).

[0124] Returning to step S192, if the page included in the document data is a portrait page (NO in S192), as shown in (b) of Figure 21, the analysis unit 16 determines that the page included in the document data is not a double-page spread (S195).

[0125] [3-2-6. Details of Step S20] Details of step S20 (estimating the reading order) in the flowchart of Fig. 4 will be described with reference to Fig. 22. Fig. 22 is a flowchart showing details of step S20 in the flowchart of Fig. 4.

[0126] As shown in Fig. 22, first, the analysis unit 16 estimates (assigns) a reading order (an example of a first reading order) that is the order in which the frames are to be read from left to right and top to bottom for the frames estimated in step S143 of the flowchart in Fig. 5 (S201). For example, as shown in (b) of Fig. 11 above, when a total of 14 frames are estimated, the analysis unit 16 estimates a reading order of "1" to "14" from left to right and top to bottom. The position of a frame is based on the top left position of a rectangle circumscribing the frame.

[0127] Next, the analysis unit 16 estimates (assigns) a reading order (an example of a second reading order) for the multiple lines in the frame, which is the order in which the multiple lines are read from left to right and top to bottom (S202).

[0128] 23 and 24, an example of the analysis result of document data by the analysis unit 16 will be described. Fig. 23 is a diagram showing an example of document data of an instruction manual. Fig. 24 is a diagram showing an example of the analysis result of document data by the analysis unit 16.

[0129] Analysis unit 16 analyzes the document structure of the document data of the instruction manual shown in Fig. 23 by executing steps S11 to S20 of the flowchart shown in Fig. 4, and obtains the analysis results in the form of a table as shown in Fig. 24. The document data of the instruction manual shown in Fig. 23 is, for example, a page in the instruction manual for a refrigerator / freezer that describes the ice making compartment.

[0130] In the example of the analysis results shown in Figure 24, a string ID (text_id), a string (text), a frame ID (frame_id), a frame type (frame_type), a reading order (read_order), a coordinate (position), a page (page), and a file name (filename) are associated with each other.

[0131] The character string (text) is a character string obtained by the character analysis, table analysis, and diagram analysis described above. The character string ID (text_id) is an ID (Identification) for identifying the character string. The frame ID (frame_id) is an ID for identifying the frame corresponding to the character string. The frame type (frame_type) is the type of layout element of the frame corresponding to the character string. The reading order (read_order) is the reading order assigned to the frame. The coordinates (position) are coordinates indicating the position of the character string. The page (page) is the page of the document data of the instruction manual. The file name (filename) is the name of the document data of the instruction manual.

[0132] [3-3. Details of Step S2] Details of step S2 (search for document data) in the flowchart of Fig. 3 will be described with reference to Fig. 25 to Fig. 27. Fig. 25 is a flowchart showing details of step S2 in the flowchart of Fig. 3. Figs. 26 and 27 are diagrams for explaining the contents of step S2 in the flowchart of Fig. 3.

[0133] 25, first, the input unit 22 accepts the user 12's intention to start a search (S21). Specifically, as shown in (a) of Fig. 26, the user 12 operates the terminal device 6 to start an application dedicated to the search system 2. Then, the user 12 presses a search start button 36 on the top screen 34 of the application.

[0134] Next, the input unit 22 accepts the selection of a product (e.g., a household electrical appliance) to be searched for by the user 12 (S22). Specifically, as shown in (b) of FIG. 26, the user 12 presses a selection button 38 on the top screen 34 to select, for example, a refrigerator-freezer. When the user 12 presses the selection button 38, a cover 40 of document data for the refrigerator-freezer instruction manual 10 is displayed in the upper section of the top screen 34, as shown in (c) of FIG. 26. In addition, a start button 42 is displayed in the lower section of the top screen 34. Note that if a product owned by the user 12 has been registered in the application in advance, a selection button for selecting that product may be displayed at the top.

[0135] Next, the input unit 22 accepts a search screen launch operation by the user 12 and launches the search screen (S23). Specifically, as shown in (c) of FIG. 26, the user 12 presses the launch button 42 on the top screen 34. This transitions from the top screen 34 to a search screen 44, as shown in (a) of FIG. 27. A software keyboard 46 and a search box 48 are displayed at the bottom of the search screen 44. The software keyboard 46 includes an input completion key 50 for inputting an input that indicates completion of input of the search query into the search box 48.

[0136] Next, the input unit 22 accepts input of a search query by the user 12 (S24). Specifically, as shown in (a) of Figure 27, the user 12 operates the software keyboard 46 on the search screen 44 to input the search query into the search box 48.

[0137] Next, the detection unit 24 determines whether or not the user 12 has input a search query (S25). If the user 12 has not performed any operation on the software keyboard 46, the detection unit 24 determines that the user 12 has not input a search query ("No input" in S25).

[0138] Furthermore, if the user 12 operates the software keyboard 46 but does not press the input completion key 50, the detection unit 24 determines that the user 12 is in the middle of inputting a search query ("input in progress" in S25). Specifically, as shown in (a) of FIG. 27, when the user 12 attempts to input the search query "How long does it take to make ice?", the user first inputs only the character "ice," which is part of the search query, into the search box 48. In this case, the search unit 28 executes an in-progress search (S26).

[0139] Furthermore, when the user 12 operates the software keyboard 46 and presses the input completion key 50, the detection unit 24 determines that the user 12 has completed input of the search query ("input completed" in S25). Specifically, as shown in (b) of FIG. 27, the user 12 inputs the search query "How long does it take to make ice?" into the search box 48 and presses the input completion key 50. In this case, the search unit 28 executes an input completion search triggered by the pressing of the input completion key 50 (S27). Note that the search unit 28 may also execute an input completion search triggered by the absence of an input update for a certain period of time after the search query was entered.

[0140] [3-3-1. Details of Step S26] Details of step S26 (mid-input search) in the flowchart of Fig. 25 will be described with reference to Fig. 28. Fig. 28 is a flowchart showing details of step S26 in the flowchart of Fig. 25.

[0141] As shown in FIG. 28, first, the search unit 28 tokenizes a part of the search query that is being input into the search box 48, that is, divides the part of the search query into tokens (S261).

[0142] Next, the second vector generation unit 26 generates a second sparse vector by sparsely vectorizing a part of the search query using the sparse vector model (S262). At this time, it is preferable that the second vector generation unit 26 uses the same sparse vector model as used in step S183 of FIG. 16 .

[0143] Next, the search unit 28 calculates, for example, cosine similarity as the similarity (an example of the first similarity) between the first sparse vector stored in the memory unit 20 and the second sparse vector generated in step S262 (S263).

[0144] Next, the search unit 28 acquires a first sparse vector whose similarity is equal to or greater than a first threshold, and acquires a character string ID (see FIG. 24 ) corresponding to the first sparse vector from the document data. Next, the search unit 28 acquires a specific character string corresponding to the character string ID as a search result, and acquires highlight coordinates for highlighting the specific character string (S264).

[0145] [3-3-2. Details of Step S27] Details of step S27 (input completion search) in the flowchart of Fig. 25 will be described with reference to Fig. 29. Fig. 29 is a flowchart showing details of step S27 in the flowchart of Fig. 25.

[0146] As shown in FIG. 29, first, the search unit 28 tokenizes the search query that has been input into the search box 48, that is, divides the search query into tokens (S271).

[0147] Next, the second vector generation unit 26 generates a second dense vector by converting the search query into a dense vector using a dense vector model (S272). At this time, it is preferable that the second vector generation unit 26 uses the same dense vector model as that used in step S183 of FIG. 16 .

[0148] Next, the search unit 28 calculates, for example, cosine similarity as the similarity (an example of second similarity) between the first dense vector stored in the memory unit 20 and the second dense vector generated in step S272 (S273).

[0149] Next, the search unit 28 acquires a first dense vector whose similarity is equal to or greater than the second threshold, and acquires a string ID (see FIG. 24 ) corresponding to the first dense vector from the document data. Next, the search unit 28 acquires a specific string corresponding to the string ID as a search result, and acquires highlight coordinates for highlighting the specific string (S274).

[0150] Note that the above-described steps S26 and S27 may be executed together. Specifically, the similarity between the first sparse vector and the second sparse vector, and the similarity between the first dense vector and the second dense vector may be calculated, and then the calculation results may be integrated using, for example, Reciprocal Rank Fusion to obtain the final search results.

[0151] [3-4. Details of Step S3] Details of step S3 (presenting search results to the user) in the flowchart of Fig. 3 will be described with reference to Fig. 30. Fig. 30 is a flowchart showing details of step S3 in the flowchart of Fig. 3.

[0152] As shown in FIG. 30, first, the presentation content generator 30 determines the highlighted portion (coordinates) in the document data based on the highlighted coordinates acquired by the search unit 28 (S31).

[0153] Next, the presentation content generator 30 determines the density (opacity) of the highlight color in accordance with the similarity calculated by the search unit 28 (S32). Specifically, the presentation content generator 30 normalizes the similarity within the range of "0" to "1" and associates it with the density of the highlight color, thereby determining the density of the highlight color to be higher as the similarity is higher.

[0154] Next, the presentation unit 32 performs highlighting on the document data based on the content determined by the presentation content generation unit 30 (S33).

[0155] Specifically, as shown in (a) of FIG. 27 described above, for example, when a search query is being entered, a relevant page 52 (a page describing an ice compartment) containing a specific character string from the search results is displayed in the upper section of search screen 44. If there are multiple relevant pages containing the specific character string from the search results, the page containing the most similar specific character string is displayed. Then, the character string 54 for "ice compartment" on relevant page 52 is highlighted as a search result from an in-progress search for the search query "ice" entered in search box 48.

[0156] 27(b) , for example, when the input of a search query is completed, a relevant page 52 containing a specific character string from the search results among the document data is displayed in the upper section of the search screen 44. At this time, if there are multiple relevant pages containing the specific character string from the search results, the relevant page containing the specific character string with the highest similarity is displayed. Then, a character string 56 on the relevant page 52, "Ice can be made in as little as 120 minutes," is highlighted as a search result of the input-completed search for the search query "How long does it take to make ice?" entered in the search box 48.

[0157] Here, another example of presentation of the search results by the presentation unit 32 will be described with reference to Fig. 31 and Fig. 32. Fig. 31 and Fig. 32 are diagrams for explaining another example of presentation of the search results by the presentation unit 32.

[0158] In the example shown in Figure 31(a), the presentation unit 32 generates a link for transitioning to a relevant page of the document data based on the search results of the search unit 28, and displays the link on the terminal device 6. For example, if the user 12 selects link 58 displaying "P6 Not cooling well...", the user is transitioned to the relevant page of the document data that is the link destination of link 58, as shown in Figure 31(b). On the linked page, character string 60 "Not cooling well", which has a relatively high similarity, is highlighted in a relatively dark color, and character string 62 "When first using, wait until the refrigerator is sufficiently cooled...", which has a relatively low similarity, is highlighted in a relatively light color. Note that in this example, the mid-input search described above is not required.

[0159] 32, the presentation unit 32 generates and presents a summary 64 of the answer based on the search results of the search unit 28. At this time, the presentation unit 32 may generate the summary 64 using, for example, a generation AI. Alternatively, the presentation unit 32 may generate a standard phrase such as "The following are the search results" as the summary 64. Note that in this example, the mid-input search described above is not essential.

[0160] In addition to highlighting the relevant portion of the document data, the presentation unit 32 may also display a link to, for example, a FAQ page on the web page of the manufacturer of the household electrical appliance. In this case, the user 12 may be able to select whether or not to display the link.

[0161] [4. Effects] In this embodiment, as described above, the analysis unit 16 analyzes the document structure of document data based on font information and border information. As a result, even if two or more sentences are distant from each other, these two or more sentences can be treated as highly related to each other based on the font information and border information. As a result, the document structure of document data can be analyzed with high accuracy.

[0162] (Other Modifications) While the search device according to one or more aspects has been described above based on the above embodiment, the present disclosure is not limited to the above embodiment. As long as it does not deviate from the spirit of the present disclosure, various modifications conceivable by a person skilled in the art to the above embodiment or configurations constructed by combining components of different embodiments may also be included within the scope of one or more aspects.

[0163] For example, in the above embodiment, the instruction manual 10 is an instruction manual for a household electrical appliance, but it is not limited to this and may be an instruction manual for industrial equipment, furniture, daily necessities, etc.

[0164] Furthermore, for example, in the above embodiment, the search target of the search device 4 is the document data of the instruction manual 10, but this is not limitative and the search target may be document data of various documents other than the instruction manual 10.

[0165] For example, in the above embodiment, the analysis unit 16 analyzed the document structure of the document data based on both font information and border information, but this is not limited to this, and the document structure of the document data may be analyzed based on either font information or border information.

[0166] In the above embodiment, the presentation unit 32 highlights the specific character string extracted by the search unit 28 in a display manner in which the color becomes darker as the similarity increases, but the presenting unit 32 may highlight the specific character string in a display manner in which the color becomes thicker as the similarity increases. Alternatively, the highlighting may be performed in any of the following display manners: (a) a display manner in which the font size increases as the similarity increases, (b) a display manner in which the font color becomes brighter as the similarity increases, (c) a display manner in which the font color changes to an underline, italics, or a highlighted color when the similarity exceeds a threshold, or (d) a display manner in which the font color or font size is made less noticeable in areas other than those where the similarity exceeds a threshold, thereby relatively emphasizing the area.

[0167] In the above-described embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0168] Furthermore, some or all of the functions of the search device according to the above-described embodiment may be realized by a processor such as a CPU executing a program.

[0169] The present disclosure may be a method shown in the control method according to Technology 11. It may also be a computer program for implementing these methods on a computer, or a digital signal comprising the computer program. The present disclosure may also be a computer program or a digital signal recorded on a computer-readable non-transitory recording medium, such as a flexible disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray (registered trademark) Disc), semiconductor memory, or the like.

[0170] The present disclosure is applicable to a search device or the like for searching for information desired by a user from document data showing, for example, an instruction manual.

[0171] 2 Search system 4 Search device 6 Terminal device 8 Network 10 Instruction manual 12 User 14 Acquisition unit 16 Analysis unit 18 First vector generation unit 20 Storage unit 22 Input unit 24 Detection unit 26 Second vector generation unit 28 Search unit 30 Presentation content generation unit 32 Presentation unit 34 Top screen 36 Search start button 38 Selection button 40 Cover 42 Start button 44 Search screen 46 Software keyboard 48 Search box 50 Input completion key 52 Corresponding page 54, 56, 60, 62 Character string 58 Link 64 Abstract

Claims

1. A search device for searching for information desired by a user from document data, comprising: an acquisition unit that acquires the document data; and an analysis unit that analyzes a document structure of the document data based on font information related to the font of characters included in the document data acquired by the acquisition unit, and border information related to borders included in the document data.

2. The search device of claim 1, wherein the analysis unit (i) estimates, based on the font information, a set of characters that have the same font and whose adjacent distance in a first direction is shorter than the font size as a line in the document data, and (ii) estimates a set of lines whose adjacent distance in a second direction perpendicular to the first direction is shorter than the font size as a frame in the document data, thereby analyzing the document structure of the document data.

3. The search device described in claim 2, wherein the analysis unit (i) detects a plurality of line segments as the border line by obtaining the coordinates of each of a plurality of line segments that form a closed area among the plurality of line segments included in the document data as the border line information, and (ii) estimates that a set of lines included inside the border line is the same frame based on the border line information.

4. The search device according to claim 2 or 3, wherein the document data is written horizontally, and when the analysis unit estimates a plurality of frames, it assigns a first reading order to the plurality of frames, which is the order in which the plurality of frames are to be read from left to right and from top to bottom.

5. The search device according to claim 4, wherein the analysis unit further assigns a second reading order to the lines in the frame when the lines are estimated to be present in the frame, the second reading order being an order in which the lines are to be read from left to right and from top to bottom.

6. The search device of claim 1, further comprising: an input unit that accepts input of a search query by the user; a first vector generation unit that generates a first vector by vectorizing the document data analyzed by the analysis unit; a second vector generation unit that generates a second vector by vectorizing the search query accepted by the input unit; and a search unit that calculates a similarity between the first vector and the second vector and extracts a specific character string from the document data as a search result for the search query based on the calculated similarity.

7. The search device described in claim 6, wherein the first vector generation unit generates a first sparse vector and a first dense vector as the first vector, the second vector generation unit generates a second sparse vector and a second dense vector as the second vector, and the search unit (i) calculates a first similarity between the first sparse vector and the second sparse vector as the similarity when the user is in the middle of inputting the search query, and (ii) calculates a second similarity between the first dense vector and the second dense vector as the similarity when the user has completed inputting the search query.

8. The search device according to claim 6 or 7, further comprising a presentation unit that presents search results of the search unit to the user, the presentation unit highlighting the specific character string extracted by the search unit in a display mode according to the degree of similarity.

9. The search device according to claim 8, wherein the presentation unit highlights the specific character string in a display manner in which the higher the similarity, the darker the color.

10. The search device according to claim 6 or 7, further comprising a presentation unit that presents search results of the search unit to the user, generates a summary of the answer based on the search results of the search unit, and displays the generated summary.

11. A control method for a search device for searching for information desired by a user from document data, comprising: (a) a step of acquiring the document data; and (b) a step of analyzing a document structure of the document data based on font information relating to the font of characters included in the document data acquired in (a) and frame information relating to frame lines included in the document data.

12. A program for causing a computer to execute the control method according to claim 11.

Citation Information

Patent Citations

  • Device and method of searching instruction

    JP2006048605A

  • Document retrieval device

    JP2006344010A

  • Information processing apparatus and method

    JP2021149935A

  • Enterance control method based on body recognition

    KR102385891B1

Cited By

  • Context-based multilingual text conversion AI keyboard system and its operation method

    JP7901861B1