Electronic Device being capable Classifying, Searching and Recommending Document
The electronic device enhances document retrieval by analyzing image-related data within documents, addressing the limitations of keyword-based methods and improving the accuracy of document classification, search, and recommendation.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- ELECTRONICS & TELECOMM RES INST
- Filing Date
- 2019-11-07
- Publication Date
- 2026-07-15
AI Technical Summary
Conventional keyword-based document search methods struggle to effectively find documents containing core content in the form of figures, tables, or graphs, as they fail to capture the information conveyed by such data.
An electronic device that classifies, searches, and recommends documents based on user interest indicators by extracting and analyzing image-related data such as pictures, tables, and graphs, and determining features from these elements, in addition to text-based keywords.
Improves the accuracy of document classification, search, and recommendation by utilizing image-related data, enabling more precise retrieval of documents, even when the core content is in graphical form.
Smart Images

Figure 112019114568920-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The various embodiments disclosed in this document relate to document retrieval technology. Background Technology
[0002] As the field of science and technology grows rapidly, documents describing technological achievements are also increasing at a rapid pace. For example, the number of research papers published annually shows a growth rate of about 4–5%, and nearly 3 million research papers in the field of science and technology are released every year.
[0003] Therefore, many researchers can obtain necessary knowledge, determine research directions, or evaluate research outcomes through already published research papers. However, since too many papers are publicly available, users may spend a significant amount of time and effort searching for the specific papers they desire. Furthermore, depending on the gist of the research paper, users may fail to find the documents they need even after spending a great deal of time and effort.
[0004] Recently, with the advancement of big data and artificial intelligence technologies, searching for necessary documents has become more convenient for users. Various search engines can provide optimal search results that reflect individual characteristics (e.g., interests) based on entered keywords and the user's search history. The problem to be solved
[0005] Keyword-based document search methods can find documents that are somewhat similar to the document the user intends to search for, based on the similarity between the keyword and the text within the document. However, keyword-based document search methods have struggled to find documents corresponding to the keyword when the core content of the document consists of figures, tables, or graphs. In fact, in many research fields, the core content of papers is expressed in forms such as figures, tables, and graphs, and information such as table labels and axes plays a crucial role in conveying information. Conventional keyword-based document search methods have limitations in finding the information conveyed by such data, and it can be difficult to search for such documents.
[0006] Various embodiments disclosed in this document may provide an electronic device capable of classifying, searching, or recommending documents based on a user's interest indicators (e.g., indicator documents). means of solving the problem
[0007] An electronic device according to one embodiment disclosed in this document comprises: a memory for executing a plurality of groups of document information and at least one instruction; and a processor, wherein the processor, by executing the at least one instruction, extracts image-related data included in the document, extracts features from the extracted image-related data, and classifies the document based on the similarity between the features related to the plurality of groups of document information and the extracted features.
[0008] Additionally, an electronic device according to one embodiment disclosed in this document includes a memory for storing at least one instruction; and a processor, wherein the processor, by executing the at least one instruction, acquires a keyword entered by a user, searches for documents corresponding to the keyword, identifies an indicator document based on the user's history information, extracts features from image-related data included in the searched documents, selects at least some of the searched documents based on the similarity between the features related to the user's indicator document and the extracted features, and provides the selected at least some of the documents to the user.
[0009] Additionally, an electronic device according to one embodiment disclosed in this document includes a memory for storing at least one instruction; and a processor, wherein the processor, by executing the at least one instruction, searches for documents corresponding to a user's history information, identifies an indicator document based on the user's history information, extracts features from image-related data included in the searched documents, selects at least some of the searched documents based on the similarity between the features related to the user's indicator document and the extracted features, and provides the selected at least some of the documents to the user. Effects of the invention
[0010] According to the various embodiments disclosed in this document, documents can be classified, searched, or recommended more accurately based on the user's interest indicators (e.g., indicator documents). In addition, various effects that can be identified directly or indirectly through this document may be provided. Brief explanation of the drawing
[0011] FIG. 1 shows an implementation environment of an electronic device according to one embodiment. FIG. 2 shows a configuration diagram of an electronic device according to one embodiment. FIG. 3 illustrates a document classification method according to one embodiment. FIG. 4 illustrates a document search method according to one embodiment. FIG. 5 illustrates a document recommendation method according to one embodiment. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Specific details for implementing the invention
[0012] FIG. 1 shows an implementation environment of an electronic device according to one embodiment.
[0013] Referring to FIG. 1, according to one embodiment, an electronic device (100) can perform at least one of document classification and document search.
[0014] According to one embodiment, an electronic device (100) may acquire a document, extract at least one image-related data among a picture, a table, and a graph included in the document from the acquired document, and determine features corresponding to each document (hereinafter referred to as "first features") from the extracted image-related data. The electronic device (100) may classify the document based on the determined features. For example, the electronic device (100) may extract at least one first feature among a caption and an image from a picture, extract at least one first feature among a caption and a classification criterion corresponding to a row or column of said table from a table, or determine at least one first feature among a caption, a graph image, and an axis variable from a graph.
[0015] According to one embodiment, when an electronic device (100) acquires a keyword set by a user, it searches for documents (e.g., web documents) corresponding to the keyword based on text features (which may be referred to as second features below), and can provide at least some of the searched documents to the user.
[0016] The electronic device (100) can provide at least some documents among the searched documents that correspond to the user's indicator documents. For example, the electronic device (100) can extract at least one image-related data among pictures, tables, and graphs from the searched documents and indicator documents, and determine first features corresponding to the searched documents and indicator documents, respectively, based on the extracted image-related data. The electronic device (100) can provide the user with at least some documents among the searched documents that are presumed to be of high interest to the user, based on the similarity between the determined first features and the features related to the indicator documents.
[0017] According to the above-described embodiment, the electronic device (100) can improve the accuracy of document classification, document search, and document recommendation by utilizing not only text included in the document but also image-related data.
[0018] FIG. 2 shows a configuration diagram of an electronic device according to one embodiment.
[0019] Referring to FIG. 2, an electronic device (100) according to one embodiment may include a memory (240) and a processor (250). In one embodiment, the electronic device (100) may omit some components or include additional components. For example, the electronic device (100) may further include a communication circuit (210). Additionally, some of the components of the electronic device (100) may be combined to form a single entity, while performing the same functions as the components prior to combination. The electronic device (100) according to the embodiment of this document may be, for example, a web server that performs at least one of document classification, document search, and document recommendation to an external electronic device through the communication circuit (210).
[0020] The communication circuit (210) can support the establishment of a communication channel or a wireless communication channel between the electronic device (100) and another device (e.g., an external electronic device), and the performance of communication through the established communication channel. The communication channel may be a communication channel of a communication method such as, for example, a LAN (local area network), FTTH (Fiber to the home), xDSL (x-Digital Subscriber Line), WiFi, Wibro, 3G, or 4G.
[0021] The memory (240) can store various data used by at least one component (e.g., processor (250)) of the electronic device (100). The data may include, for example, input data or output data for software and related instructions. For example, the memory (240) may store at least one instruction for performing at least one function among document classification, document search, and document recommendation. The memory (240) may include volatile memory or non-volatile memory.
[0022] The processor (250) can control at least one other component (e.g., hardware or software component) of the electronic device (100) by executing at least one instruction and can perform various data processing or operations. For example, the processor (250) can perform at least one function among document classification, document search, and document recommendation by executing at least one instruction. The processor (250) may include at least one of, for example, a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor, an application processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), and may have multiple cores.
[0023] According to one embodiment, the processor (250) may acquire any document through at least one website and classify the acquired document. For example, the processor (250) may acquire documents through web scrolling from a plurality of websites or a designated database (e.g., a database storing research papers). The document may include at least one of a document disclosed on a website or a document linked to a website.
[0024] When the processor (250) acquires a document, it may extract image-related data included in or related (e.g., tagging) to the acquired document and determine first features of the document based on the image-related data. In one embodiment, the image-related data may include at least one image-related data among a picture, a table, and a graph. For example, if the image-related data is a picture, the processor (250) may extract at least one of a caption and an image from the picture and determine at least one first feature among the caption and the image. The caption includes a descriptive text related to the picture (e.g., a title, a description related to the picture). In another example, if the image-related data is a table, the processor (250) may extract at least one of a caption and a classification criterion corresponding to a row or column of the table from the table and determine at least one first feature among the caption and the classification criterion. The caption includes a descriptive text related to the table (e.g., a title). The classification criterion may be, for example, an item represented by the data included in each row and column. As another example, if the image-related data is a graph, the processor (250) may extract at least one of a caption, a graph image, and an axis variable from the graph and determine a first feature of at least one of the caption, the graph image, and the axis variable. The caption includes a descriptive text (e.g., a title) related to the figure. The axis variable may be, for example, a variable represented by a horizontal axis or a vertical axis (e.g., the horizontal axis represents time and the vertical axis represents current magnitude). Additionally, the processor (250) may extract texts related to the document (e.g., included or tagged) and determine second features of the document based on the extracted texts.
[0025] The processor (250) may apply weights corresponding to the technical field of the document to the first features of the determined document. For example, the processor (250) may identify the technical field to which the document belongs based on text (e.g., title, description) related to the document (e.g., inclusion or tagging) and apply different weights according to the technical field. As another example, the processor (250) may apply different weights to each group to which the document belongs through a machine learning process and an accuracy (or reliability) evaluation process. As yet another example, the processor (250) may optimize the weights through at least one of the machine learning process and the accuracy (or reliability) evaluation process.
[0026] The processor (250) can classify documents based on at least first features. For example, the processor (250) can check the similarity between the first features of documents in multiple groups belonging to multiple groups and the first features of the acquired document, and determine the group with relatively high similarity as the group to which the acquired document belongs. As another example, the processor (250) can determine the group to which the acquired document belongs based on the similarity between the first features and second features of documents in multiple groups and the first features of the acquired document. In this process, the processor (250) can check cosine similarity. Alternatively, the processor (250) can check similarity using deep learning techniques such as RNN (Recurrent neural network), LSTM (Long short-term memory), and CNN (Convolutional Neural Network).
[0027] According to one embodiment, the processor (250) can obtain a keyword input by a user of an external electronic device from an external electronic device through a communication circuit (210) and search for documents corresponding to the keyword. For example, the processor (250) can search for documents corresponding to the keyword from a plurality of websites through the communication circuit (210) based on text features. In this process, the processor (250) can search for documents presumed to be meaningful based on at least one of the reference history information of each document and the user's search history information. For example, the processor (250) can search for documents that have been referenced relatively frequently among the documents corresponding to the keyword based on a page rank algorithm.
[0028] Additionally or generally, the processor (250) can identify user characteristics related to at least one of the user's history, such as the user's search history information, the user's interest document information, information on other users' interest documents that have identified documents similar to the user's interest document, and information on other documents referenced by the user's interest document (e.g., reference information included in the interest document and author network information of the interest document), and search for documents corresponding to keywords and user characteristics. For convenience of explanation, the following description uses an example of an electronic device (100) searching for documents corresponding to keywords and user characteristics.
[0029] The processor (250) searches for documents corresponding to keywords and can select some documents corresponding to user characteristics among the documents corresponding to keywords based on content-related features according to search history information and user interest document information (content-based filtering). For example, the processor (250) can identify documents of interest to the user that include at least one of documents previously searched by the user, documents classified with interest by the user (e.g., scraped or bookmarked), and documents disclosed by the user (e.g., published, posted, announced). The processor (250) can provide at least some documents among the documents corresponding to keywords that have a relatively high similarity to the documents of interest based on the similarity between the first features of the documents of interest to the user and the first features of the keyword. The previously searched documents can be identified based on the user's past search history or information on frequently visited sites.
[0030] The processor (250) can search for documents corresponding to user characteristics among documents corresponding to keywords based on content-related features according to documents of interest of other users having similar patterns to the user (collaborative filtering). For example, the processor (250) can identify at least one document of interest among documents previously searched by the user, documents published by the user, and documents marked as interest, and identify other users who have identified (e.g., searched) at least one document of interest. The processor (250) can identify the documents of interest of other users and provide at least some documents among the searched documents that have relatively high similarity to the document of interest based on the similarity between the features of the searched documents corresponding to keywords and the features of the other users' documents of interest.
[0031] The processor (250) generates a relationship diagram between the user's document of interest and other documents that the document of interest references, and based on the relationship diagram, can search for some documents corresponding to user characteristics among documents corresponding to keywords, for example, by a random walk process (graph-based method). For example, the processor (250) can search for documents that are published or classified by the author of the document of interest among documents corresponding to keywords, based on author information included in the user's document of interest. As another example, the processor (250) can search for documents that the document of interest references among documents corresponding to keywords, based on reference information included in the user's document of interest.
[0032] The processor (250) can extract image-related data related to each document (e.g., inclusion or tagging) from the retrieved documents (documents corresponding to keywords and user characteristics) and determine first features of the retrieved documents based on the extracted image-related data. For example, the processor (250) can determine a first feature related to at least one of a caption and an image from a picture, determine a first feature related to at least one of a distinction criterion corresponding to a row or column of a table from a table, or determine a first feature related to at least one of a caption, a graph image, and an axis variable from a graph.
[0033] The processor (250) can check the user's indicator document and extract at least one image-related data among pictures, tables, and graphs related to the indicator document. The processor (250) can extract first features of the indicator document based on the image-related data related to the indicator document. The indicator document may include, for example, at least one document among documents recently disclosed by the user within a certain period, documents selected as indicators of interest by the user, and stored documents with a similar document recommendation function set. For other examples, the indicator document may be set as some high-importance documents among the indicator documents.
[0034] The processor (250) can apply weights to each of the first features and determine the similarity between documents based on the first features to which the weights are applied. For example, the processor (250) can apply different weights to each group to which the document belongs through a machine learning process and an accuracy (or reliability) evaluation process. As another example, the processor (250) can optimize the weights through at least one of the machine learning process and the accuracy (or reliability) evaluation process.
[0035] The processor (250) can determine the similarity between the first features of the index document and the first features of the searched documents (documents corresponding to keywords and user characteristics). Based on the determined similarity, the processor (250) can adjust the priority of the documents corresponding to keywords and user characteristics and sort the documents corresponding to keywords and user characteristics according to the adjusted priority. The processor (250) can provide at least some of the documents corresponding to keywords and user characteristics sorted according to priority to the user. In this case, the processor (250) can provide at least some of the documents to an external electronic device through the communication circuit (210).
[0036] According to one embodiment, the processor (250) can recommend documents expected to be of high interest to the user based on user characteristics and user indicator documents without acquiring keywords.
[0037] The processor (250) can identify content-related features of at least one document of interest among the user's search history information, documents previously disclosed by the user, or documents classified as interest. Based on the similarity of the identified content-related features, the processor (250) can search for at least one document of interest among other users' documents of interest that have identified documents similar to the user's document of interest, and other documents referenced by the user's document of interest (e.g., reference information and author network information included in the document of interest). Based on the user's search history information, the processor (250) can search for a document of interest that is presumed not to have been previously identified by the first researcher. The document of interest may include, for example, at least one document among documents classified as interest by each user (e.g., scrap or favorite settings) and documents disclosed by each user (e.g., publication, posting, presentation). The processor (250) can extract at least one image-related data among pictures, tables, and graphs from the retrieved documents, determine first features of the retrieved documents corresponding to user characteristics based on the image-related data, and relate the determined first features to the retrieved documents.
[0038] The processor (250) can acquire a user's indicator document and extract at least one image-related data among pictures, tables, and graphs related to the indicator document. The indicator document may include, for example, at least one document among documents recently disclosed by the user within a certain period, documents selected as indicators of interest by the user, and stored documents with a similar document recommendation function set. For other examples, the indicator document may be set as some high-importance documents among the indicator documents. The processor (250) can determine first features of the indicator document based on the image-related data and associate the determined first features with the indicator document.
[0039] The processor (250) can check the similarity between the first features of the index document and the first features of the searched documents, and readjust the priority of documents corresponding to user characteristics based on the confirmed similarity. For example, the processor (250) can increase the priority of documents of interest that have relatively high similarity to the index document.
[0040] The processor (250) can transmit a specified number of documents in order of highest priority to an external electronic device through the communication circuit (210).
[0041] Hereinafter, a specific example is described in which a processor (250) according to one embodiment identifies user characteristics of a first researcher who studies the field of communication and recommends papers that are predicted to be of high interest to the researcher corresponding to the identified user characteristics.
[0042] The processor (250) can search for interest documents related to at least one user's history information among research papers previously disclosed by the first researcher, research papers classified with interest by the first researcher (corresponding to the said interest documents), research papers confirmed with interest by the first researcher and co-authors or authors of the said interest-classified research papers, and interest documents of other researchers in a field of interest similar to that of the first researcher. In this process, the processor (250) can search for interest documents that are similar to the content-related features (text features) (or user features) of the research papers previously disclosed or the research papers classified with interest by the first researcher by a threshold similarity or higher. Based on the user's search history information, the processor (250) can search for interest documents presumed not to have been previously confirmed by the first researcher.
[0043] The processor (250) can verify an indicator document set directly or indirectly by the first researcher. For example, the processor (250) may include at least one document among a paper published by the first researcher within a recent period, a research paper selected as an indicator of interest by the first researcher, and a stored document in which a similar document recommendation function is set by the first researcher.
[0044] The processor (250) can determine the characteristics of each document based on image-related data included in the searched interest documents and indicator documents. For example, the processor (250) can extract at least one image-related data among the pictures, tables, and graphs included in the indicator documents and determine the characteristics of the indicator documents. As another example, the processor (250) can extract at least one image-related data among the pictures, tables, and graphs included in the searched interest documents and determine the characteristics of the interest documents. The processor (250) can check the similarity between the characteristics of the indicator documents and the characteristics of the interest documents, and based on the confirmed similarity, readjust the priority of the interest documents in order of high similarity to the indicator documents. Subsequently, the processor (250) can provide the interest documents to an external electronic device in order of high priority according to the readjusted priority.
[0045] According to various embodiments, the electronic device (100) can classify, search, or recommend documents based on some image-related data included in the documents. For example, the electronic device (100) can classify, search, or recommend documents based on some pages of the documents (e.g., some pages of the first part, pages of the conclusion part).
[0046] According to various embodiments, the electronic device (100) may be a computing device that includes an input circuit (220) and an output circuit (230) and interacts with a user through the input circuit (220) and the output circuit (230). The input circuit (220) may include at least one of a mouse, a keyboard, or a touchpad and may detect or receive user input. The output circuit (230) may include a display that outputs various content, and the display may include, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, or an organic light-emitting diode (OLED) display. In this case, the electronic device (100) may acquire keywords entered by the user through the input circuit (220) and check the user's search history information and documents of interest based on the user's login information. The electronic device (100) may search for documents corresponding to the keywords and provide them to the user through the output circuit (230). Alternatively, the electronic device (100) may recommend documents of interest to the user based on the user's search history information and documents of interest.
[0047] According to the above-described embodiment, the electronic device (100) can improve the accuracy of document classification, search, or recommendation by utilizing image-related data as well as text included in the document to classify, search, and recommend documents.
[0048] In addition, according to the above-described embodiment, the electronic device (100) can search for documents corresponding to keywords or documents of interest when the core content of the document is a picture, table, or graph.
[0049] FIG. 3 illustrates a document classification method according to one embodiment.
[0050] Referring to FIG. 3, in operation 310, the processor (250) can obtain any document A through at least one website. For example, the processor (250) can obtain documents through web scrolling from multiple websites or a designated database (e.g., a database where research papers are stored).
[0051] In operation 320, when the processor (250) acquires a document, it can extract image-related data including pictures, tables, and graphs that are included in or related (e.g., tagged) to the acquired document.
[0052] In operation 330, the processor (250) can extract detailed data from image-related data. For example, if the image-related data is a picture, the processor (250) can extract at least one of a caption and an image from the picture, extract at least one of a caption and a distinction criterion corresponding to a row or column of the table from a table, and extract at least one of a caption, a graph image, and an axis variable from a graph.
[0053] In operation 340, the processor (250) determines first features of a document by characterizing (e.g., quantifying) detailed data and may apply weights to the first features. For example, the processor (250) may apply weights corresponding to the technical field of the document to the determined first features of the document.
[0054] In operation 350, the processor (250) can determine second features related to the content based on any document A.
[0055] In operation 360, the processor (250) can determine the similarity between a plurality of groups of documents and any document A based on the first features and the second features.
[0056] In operation 370, the processor (250) can identify the group with the highest similarity among the multiple groups.
[0057] In operation 380, the processor (250) can classify document A as a document belonging to a verified group.
[0058] FIG. 4 illustrates a document search method according to one embodiment.
[0059] Referring to FIG. 4, in operation 405, when the processor (250) obtains a keyword entered by a user, it can search for documents corresponding to the keyword based on content-related features (e.g., text features). For example, the processor (250) can search for documents that are relatively frequently referenced among the documents corresponding to the keyword based on a page rank algorithm.
[0060] In operation 410, the processor (250) can identify user characteristics including at least one of the user's search history information, the user's interest document information, information on other users' interest documents that have identified documents similar to the user's interest document, and information on other documents that the user's interest document refers to (e.g., reference information included in the interest document and author network information of the interest document).
[0061] In operation 415, the processor (250) can select some documents corresponding to user characteristics among the documents corresponding to the keyword. For example, the processor (250) can search for documents corresponding to the keyword and select some documents corresponding to user characteristics (or user history information) among the documents corresponding to the keyword based on content-related features based on search history information and user interest document information (content-based filtering). As another example, the processor (250) can search for documents corresponding to user characteristics among the documents corresponding to the keyword based on content-related features based on interest documents of other users having a pattern similar to the user (collaborative filtering). As yet another example, the processor (250) can generate a relationship diagram between the user's interest documents and other documents referenced by said interest documents, and based on said relationship diagram, search for some documents corresponding to user characteristics among the documents corresponding to the keyword by, for example, a random walk process (graph-based method).
[0062] In operation 420, the processor (250) can check the user's indicator documents. For example, the processor (250) can check at least one indicator document among documents recently published by the user within a certain period, documents selected as interest indicators by the user, and saved documents with a similar document recommendation function set.
[0063] In operation 425, the processor (250) may extract at least one image-related data among a figure, a table, and a graph related to the index document, and may extract first features of the index document based on the image-related data. For example, the processor (250) may determine a first feature related to at least one of a caption and an image from a figure, determine a first feature related to at least one of a distinction criterion corresponding to a row or column of a table and a caption from a table, or determine a first feature related to at least one of a caption, a graph image, and an axis variable from a graph. In operation 425, the processor (250) may apply weights to each of the first features of the index document.
[0064] In operation 430, image-related data related to (e.g., inclusion or tagging) the searched documents corresponding to keywords and user characteristics can be extracted, and first features of the searched documents can be determined based on the extracted image-related data.
[0065] In operation 440, the processor (250) can determine the similarity between the first features of the indicator document and the first features of the searched documents (documents corresponding to keywords and user characteristics).
[0066] In operation 450, the processor (250) can adjust the priority of documents corresponding to keywords and user characteristics based on the identified similarity, and sort the documents corresponding to keywords and user characteristics according to the adjusted priority.
[0067] In operation 460, the processor (250) can provide the user with at least some of the documents corresponding to keywords and user attributes sorted according to priority.
[0068] FIG. 5 illustrates a document recommendation method according to one embodiment.
[0069] Referring to FIG. 5, in operation 510, the processor (250) can identify user characteristics related to at least one of the user's history (or interest documents) among the user's search history information, documents previously disclosed by the user, or documents classified with interest.
[0070] In operation 520, the processor (250) can search for at least one document of interest among other users' documents of interest that have identified documents similar to the user's document of interest based on the similarity of identified content-related features, and other documents that the user's document of interest references (e.g., reference information and author network information included in the document of interest). The processor (250) can search for a document of interest that is presumed not to have been previously identified by the user based on the user's search history information.
[0071] In operation 530, the processor (250) may obtain a user’s indicator document. The indicator document may include, for example, at least one document among a document recently disclosed by the user within a certain period, a document selected as an indicator of interest by the user, and a stored document with a similar document recommendation function set. For other examples, the indicator document may be set as some high-importance documents among the indicator documents.
[0072] In operation 540, the processor (250) extracts at least one image-related data among a picture, table, and graph related to an index document, determines first features of the index document based on the image-related data, and can associate the determined first features with the index document.
[0073] In operation 550, the processor (250) extracts at least one image-related data among pictures, tables, and graphs from documents corresponding to user characteristics, determines first features of the searched documents corresponding to user characteristics based on the image-related data, and can associate the determined first features with the searched documents.
[0074] In operation 560, the processor (250) can check the similarity between the first features of the indicator document and the first features of the searched documents, and readjust the priority of the documents corresponding to the user characteristics based on the checked similarity.
[0075] In operation 570, the processor (250) can recommend a specified number of documents to the user in order of highest priority.
[0076] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B or C,” “at least one of A, B and C,” and “at least one of A, B, or C” may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another corresponding component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0077] As used herein, the terms “module,” “part,” “unit,” and “means” may include units implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0078] Various embodiments of the present document may be implemented as software (e.g., program) comprising one or more instructions stored in a storage medium (e.g., internal memory or external memory) (memory (240)) that can be read by a machine (e.g., electronic device (100)). For example, a processor (e.g., processor (250)) of a device (e.g., electronic device (100)) may call at least one of one or more instructions stored from a storage medium and execute it. This enables the device to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. A storage medium readable by the device may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.
[0079] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0080] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to the integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
Claim 1 delete Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 In an electronic device, a memory that stores at least one instruction; The system includes a processor, wherein the processor, by executing at least one instruction, obtains a keyword entered by a user, searches for documents corresponding to the keyword, selects from the searched documents a portion of documents that have a relatively high similarity to documents of interest including documents disclosed by the user, documents classified as scraps or favorites, and documents disclosed by the user, and extracts first image-related features from image-related data including pictures, tables, and graphs from the portion of documents, wherein in the picture, features are extracted from the caption and the image; in the table, features are extracted from the caption and the classification criteria corresponding to the column or row of the table; and in the graph, features are extracted from the caption, the graph image, and the axis variable; identifies an indicator document among documents disclosed by the user within a recent period of time, documents set as an interest indicator, and documents disclosed or classified by the author of the document set as an interest indicator based on the user's history information, extracts second image-related features from the identified indicator document, and the second image-related features included in the indicator document and the first image-related features extracted from the selected portion of documents. An electronic device that adjusts the priority of some documents based on similarity between features, provides the some documents to the user according to the adjusted priority, and, if the user is a researcher of a paper, the documents of interest include papers published by the user, co-authors of the published papers, and documents of interest of other researchers in a field of interest similar to that of the user. Claim 8 An electronic device according to claim 7, wherein the processor searches for at least one document among documents previously searched and classified by the user, documents disclosed within a recent set of times, documents set as an interest indicator, and documents disclosed or classified by the author of the documents set as an interest indicator by executing at least one instruction. Claim 9 An electronic device according to claim 7, wherein the indicator document comprises at least one document among a stored document in which a similar document recommendation function is set by the user and a document of relatively high importance among some of the documents set as the indicator document by the user. Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 A method for searching documents using an electronic device comprises: an operation of searching for documents corresponding to a keyword when a keyword entered by a user is obtained; an operation of selecting some documents among the searched documents that have a relatively high similarity to a document of interest, including a document disclosed by the user, a scrap or a document of interest set as a favorite, and a document disclosed by the user; an operation of extracting first image-related features from image-related data including pictures, tables, and graphs in the selected documents; an operation of identifying an index document among a document disclosed by the user within a certain period of time recently, a document set as an index of interest, and a document disclosed or classified by the author of the document set as an index of interest, based on the user's history information; an operation of extracting second image-related features from the identified index document; and an operation of selecting at least some documents among the searched documents based on the similarity between the second image-related features included in the user's index document and the second image-related features extracted from the searched documents. A document search method comprising the operation of providing at least some of the selected documents to the user, wherein the extraction operations include: the operation of extracting first and second image-related features from the caption and image, respectively, in the figure; the operation of extracting first and second image-related features from the caption and the distinction criteria corresponding to the column or row of the table in the table in the table; and the operation of extracting first and second image-related features from the caption, graph image, and axis variable in the graph in the graph, wherein if the user is a researcher of a paper, the documents of interest include papers published by the user, co-authors of the published papers, and documents of interest of other researchers in fields of interest similar to the user. Claim 15 A document search method according to claim 14, wherein the confirming operation includes the operation of searching for at least one document among documents previously searched and classified by the user, documents disclosed within a recent set of times, documents set as interest indicators, and documents disclosed or classified by the author of the documents set as interest indicators. Claim 16 A document search method according to claim 14, wherein the indicator document further comprises at least one document among a stored document in which a similar document recommendation function is set by the user and a document of relatively high importance among some of the documents set as the indicator document by the user. Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete