Information collection apparatus, information collection method, and information collection program

The information collection device uses learning models to infer and generate keywords, addressing the inefficiency in extracting website attributes, thereby reducing the effort in data collection and registration.

JP2025187180AInactive Publication Date: 2025-12-25MITSUBISHI ELECTRIC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024095768
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-12-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently extract characteristic keywords from websites, especially when they are not clearly stated, leading to increased effort in researching and registering websites.

Method used

An information collection device utilizing an attribute extraction model and a keyword generation model to acquire attribute information from websites, even if it is not explicitly stated, by employing learning models like generative AI to infer and generate relevant keywords.

Benefits of technology

Reduces the effort required to investigate and register websites by accurately extracting attribute information, even when it is not clearly stated, enhancing the efficiency of data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025187180000001_ABST
    Figure 2025187180000001_ABST
Patent Text Reader

Abstract

To further reduce labor to investigate and register websites.SOLUTION: A scraping unit 23 inputs information of a website for the website to be an information collection target into an attribute extraction model that is a learning model, and acquires attribute information that is output by the attribute extraction model and is information on specified attributes relating to the website. A data output unit 24 outputs the attribute information acquired by the scraping unit 23 in association with the website.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for collecting website information. [Background technology]

[0002] There are application programs that collect websites published on the Internet into specific units and make them available to users. For example, there is an application program that collects websites related to restaurants and tourist spots around each station and makes them available to users. Businesses that publish data using such application programs must search for websites to publish, extract the necessary information from the websites, and register it in a database. Enhancing high-quality websites is important for improving the value of services. However, researching and registering a large number of websites is time-consuming.

[0003] Patent Document 1 describes a technology for collecting information that indicates the characteristics of various facilities. In Patent Document 1, characteristic keywords are extracted by morphological analysis from websites searched using the facility name as a search keyword. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2024-008408 Summary of the Invention [Problem to be solved by the invention]

[0005] The technology described in Patent Document 1 extracts characteristic keywords from searched websites. By using this technology, it may be possible to reduce the effort required to research and register websites to some extent. However, the technology described in Patent Document 1 uses morphological analysis to extract characteristic keywords, and therefore cannot extract the characteristic keywords unless they are clearly stated on the website. For example, suppose you want to extract the name of the nearest station as a characteristic keyword. In this case, the technology described in Patent Document 1 cannot extract the necessary information unless the station name, XX Station, is clearly stated. As a result, it is not possible to extract the necessary information, and the effort required to research and register websites cannot be sufficiently reduced. The present disclosure aims to make it possible to further reduce the effort required to research and register on websites. [Means for solving the problem]

[0006] The information collection device according to the present disclosure includes: A scraping unit that inputs website information of a website to be collected into an attribute extraction model, which is a learning model, and acquires attribute information output by the attribute extraction model, which is information about specified attributes related to the website; a data output unit that outputs the attribute information acquired by the scraping unit in association with the website; Equipped with. [Effects of the Invention]

[0007] In the present disclosure, attribute information for specified attributes is acquired from a website using an attribute extraction model. As a result, even if the attribute information is not clearly stated on the website, if the attribute information can be estimated from the description on the website, the attribute information can be acquired. This can further reduce the effort required to investigate and register websites. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a configuration diagram of an information collection device 10 according to a first embodiment. [Figure 2] 3 is a flowchart of the overall processing of the information collection device 10 according to the first embodiment. [Figure 3]10 is a flowchart of a keyword generation process according to the first embodiment. [Figure 4] FIG. 3 is an explanatory diagram of a keyword generation process according to the first embodiment. [Figure 5] FIG. 4 is a diagram showing examples of input keywords 41 according to the first embodiment. [Figure 6] 1 is a flowchart of a crawling process according to the first embodiment. [Figure 7] FIG. 2 is an explanatory diagram of a crawling process according to the first embodiment. [Figure 8] 1 is a flowchart of a scraping process according to the first embodiment. [Figure 9] FIG. 2 is an explanatory diagram of a scraping process according to the first embodiment. [Figure 10] FIG. 3 is an explanatory diagram of subcategories according to the first embodiment. [Figure 11] 4 is a flowchart of a data output process according to the first embodiment. [Figure 12] FIG. 3 is an explanatory diagram of a data output process according to the first embodiment. [Figure 13] 10 is a flowchart of a scraping process according to the second embodiment. [Figure 14] 11 is a flowchart of a scraping process according to the third embodiment. [Figure 15] 10 is a flowchart of a scraping process according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Embodiment 1 ***Configuration Description***

[0010] The configuration of an information collection device 10 according to the first embodiment will be described with reference to FIG. The information collection device 10 is a computer. The information collection device 10 includes hardware such as a processor 11, a memory 12, a storage 13, and a communication interface 14. The processor 11 is connected to other hardware via signal lines and controls the other hardware.

[0011] The information collection device 10 includes, as functional components, a keyword generation unit 21, a crawling unit 22, a scraping unit 23, and a data output unit 24. The functions of the functional components of the information collection device 10 are realized by software. The storage 13 stores a program that realizes the function of each functional component of the information collection device 10. The program is read into the memory 12 by the processor 11 and executed by the processor 11. In this way, the function of each functional component of the information collection device 10 is realized.

[0012] The storage 13 also stores an additional phrase list 31 and a station correspondence list 32.

[0013] The information collection device 10 is connected to the Internet 90 via a communication interface 14. The information collection device 10 is also connected to a business operator terminal 91 and various websites 92 via the Internet 90.

[0014] The processor 11 is an IC that performs processing. IC stands for Integrated Circuit. Specific examples of the processor 11 include a CPU, a DSP, and a GPU. CPU stands for Central Processing Unit. DSP stands for Digital Signal Processor. GPU stands for Graphics Processing Unit.

[0015] The memory 12 is a storage device that temporarily stores data. Specific examples of the memory 12 include SRAM and DRAM. SRAM stands for Static Random Access Memory. DRAM stands for Dynamic Random Access Memory.

[0016] The storage 13 is a storage device that stores data. A specific example of the storage 13 is an HDD. HDD is an abbreviation for Hard Disk Drive. The storage 13 may also be a portable recording medium such as an SD (registered trademark) memory card, CompactFlash (registered trademark), NAND flash, a flexible disk, an optical disk, a compact disk, a Blu-ray (registered trademark) disk, or a DVD. SD is an abbreviation for Secure Digital. DVD is an abbreviation for Digital Versatile Disk.

[0017] The communication interface 14 is an interface for communicating with external devices. Specific examples of the communication interface 14 include Ethernet (registered trademark), USB, and HDMI (registered trademark) ports. USB stands for Universal Serial Bus. HDMI stands for High-Definition Multimedia Interface.

[0018] 1 shows only one processor 11. However, there may be a plurality of processors 11, and the plurality of processors 11 may cooperate to execute programs that realize the respective functions.

[0019] ***Explanation of Operation*** The operation of the information collection device 10 according to the first embodiment will be described with reference to FIGS. The operation procedure of the information collecting device 10 according to the embodiment 1 corresponds to the information collecting method according to the embodiment 1. Furthermore, the program that realizes the operation of the information collecting device 10 according to the embodiment 1 corresponds to the information collecting program according to the embodiment 1.

[0020] In the first embodiment, an example will be described in which website information is collected for an application program that collects websites related to restaurants, tourist spots, and the like around each station and makes them available to users.

[0021] The overall processing of the information collection device 10 according to the first embodiment will be described with reference to FIG. (Step S1: Keyword generation process) The keyword generating unit 21 generates a search keyword 42 from an input keyword 41 input by a user.

[0022] (Step S2: Crawling process) The crawling unit 22 searches for websites 92 using the search keywords 42 generated in step S1, and identifies one or more websites 92 that correspond to the search keywords 42 as websites 51.

[0023] (Step S3: Scraping process) The scraping unit 23 acquires attribute information 61, which is information about specified attributes related to each of the one or more websites 51 identified in step S2 as information collection targets.

[0024] (Step S4: Data output process) The data output unit 24 outputs the attribute information 61 acquired in step S3 for each of one or more websites 51. That is, the data output unit 24 outputs the attribute information in association with each website.

[0025] The keyword generation process (step S1 in FIG. 2) according to the first embodiment will be described with reference to FIGS. (Step S11: Keyword reception process) The keyword generating unit 21 receives an input keyword 41 input by a user. Specifically, the user inputs an input keyword 41 for identifying information using the business operator terminal 91. The keyword generation unit 21 receives the input keyword 41 via the communication interface 14.

[0026] Here, a word or phrase that serves as a keyword for each specified element is input as input keyword 41. In FIG. 4, the elements are specified as (1) station name or place name, (2) facility genre, and (3) type of facility. Suppose that "Fujisawa Station" is input for (1), "cafe" for (2), and "stylish" for (3) as input keywords 41. As shown in FIG. 5, possible input keywords 41 are "Kamakura Station" for (1), "sports" for (2), and "family-friendly" for (3). Also possible input keywords are "Shonan" for (1), "hot spring inn" for (2), and "long-established establishment" for (3). It is not necessary to specify one word or sentence for each of the specified elements. Words or sentences may be specified for only some of the specified elements. Furthermore, multiple words or sentences may be specified for at least some of the specified elements.

[0027] (Step S12: Keyword conversion process) The keyword generation unit 21 inputs the input keyword 41 received in step S11 to a learning model, that is, a keyword generation model 43. The keyword generation model 43 is a model that generates one or more synonymous output keywords 44 from the input keyword 41. The keyword generation unit 21 acquires one or more output keywords 44 output by the keyword generation model 43. Then, the keyword generation unit 21 generates search keywords 42 from the acquired one or more output keywords 44. Here, the keyword generation unit 21 sets the input keyword 41 and each of the one or more output keywords 44 as the search keywords 42.

[0028] In FIG. 4, the output keyword 44 is output in which "cafe" in (2) is converted to "coffee shop," and the output keyword 44 is output in which "oshare" in (3) is converted to "stylish."

[0029] Here, the learning model is what is known as generative AI or LLM. AI stands for Artificial Intelligence. LLM stands for Large Language Model. Specific examples of the learning model may be constructed using algorithms such as BERT and GPT. BERT stands for Bidirectional Encoder Representations from Transformers. GPT stands for Generative Pretrained Transformer. The learning model may be constructed by combining multiple algorithms including these algorithms.

[0030] (Step S13: Word addition process) The keyword generation unit 21 generates additional keywords 45 by adding additional phrases set in the additional phrase list 31 stored in the storage 13 to each of the search keywords 42 set in step S12. That is, the keyword generation unit 21 generates additional keywords 45 by adding additional phrases to each of the input keywords 41 and one or more output keywords 44. Then, the keyword generation unit 21 adds the additional keywords 45 to the search keywords 42.

[0031] 4, the additional phrase "popular" is added to each of the input keyword 41 and two output keywords 44 to generate three additional keywords 45. As a result, six search keywords 42 are set.

[0032] The keyword generating unit 21 may set only the additional keywords 45 as the search keywords 42 without setting the input keywords 41 and the output keywords 44 as the search keywords 42 .

[0033] The crawling process (step S2 in FIG. 2) according to the first embodiment will be described with reference to FIGS. (Step S21: Keyword selection process) The crawling unit 22 selects an unselected search keyword 42 from among the search keywords 42 generated in step S1 as a target search keyword 42.

[0034] (Step S22: Search process) The crawling unit 22 searches for websites 92 using the target search keyword 42 selected in step S21, and identifies one or more websites 92 that correspond to the target search keyword 42 as websites 51. Specifically, the crawling unit 22 identifies a top reference number of websites 51 from among the websites 51 searched using the target search keyword 42 as a keyword. Then, as shown in FIG. 7, the crawling unit 22 collects information about the identified websites 51, including the URL, title, text content, and OGP information, as website information 52. URL stands for Uniform Resource Locator. OGP stands for Open Graph Protocol. The OGP information includes information such as a thumbnail and the site name.

[0035] (Step S23: End determination process) The crawling unit 22 determines whether or not all of the search keywords 42 generated in step S1 have been selected as target search keywords 42. If the crawling unit 22 has finished selecting all of the search keywords 42, the process proceeds to step S24. On the other hand, if there are any unselected search keywords 42, the crawling unit 22 returns the process to step S21.

[0036] (Step S24: Integration process) The crawling unit 22 integrates the website information 52 collected using each search keyword 42 in step S22 to generate a crawling result 53. Specifically, when website information 52 about the same website 51 is collected in duplicate using multiple search keywords 42, crawling unit 22 deletes the duplicates to generate crawling result 53. As a result, in crawling result 53, website information 52 about the same website 51 becomes one.

[0037] The scraping process (step S3 in FIG. 2) according to the first embodiment will be described with reference to FIGS. (Step S31: Site selection process) The scraping unit 23 selects an unselected website 51 from the websites 51 identified in step S2 as a target website 51.

[0038] (Step S32: Attribute extraction process) The scraping unit 23 inputs website information 52 about the target website 51 selected in step S31 into the attribute extraction model 62, which is a learning model. Here, the scraping unit 23 may input the URL of the website information 52, or may input the title and text content of the website information 52. The attribute extraction model 62 is a model that extracts attribute information 61, which is information about a specified attribute, from the website 51 indicated by the website information 52. The specified attribute refers to an attribute specified by a business operator in order to extract data from a website that the business operator wishes to provide via an application program. That is, when a URL is input as the website information 52, the attribute extraction model 62 extracts attribute information 61, which is information about the specified attribute, from the website 51 indicated by the URL. Furthermore, when a title and text content are input as the website information 52, the attribute extraction model 62 extracts attribute information 61, which is information about the specified attribute, from the title and text content of the website information 52. The attribute extraction model 62 extracts, as the attribute information 61, not only information specified in the website 51 or the website information 52, but also information that can be inferred from the content described in the website 51 or the website information 52. The scraping unit 23 acquires the attribute information 61 output by the attribute extraction model 62.

[0039] Here, as shown in FIG. 9, the scraping unit 23 acquires station names, facility names, tags, and subcategories as attribute information 61. The station names are the names of the nearest stations to the facilities targeted by the website 51. The facility names are the names of the facilities targeted by the website 51. The tags are words that represent the characteristics of the website 51. The subcategories represent the classifications of the website 51, as shown in FIG. 10. The attribute extraction model 62 identifies one or more classifications that correspond to the website 51 from among the pre-specified classifications as subcategories.

[0040] At this time, the scraping unit 23 may instruct the attribute extraction model 62 to estimate a specific attribute, which is an attribute among the specified attributes, if the specific attribute is not clearly stated on the website 51. Specifically, the scraping unit 23 writes a command to estimate the specific attribute in a prompt that writes instructions to the attribute extraction model 62. For example, in the example of FIG. 9, the specific attribute is a station name. In this case, if the station name is not listed on the website 51, the scraping unit 23 writes in the prompt an instruction to estimate the name of the nearest train station. At this time, it is desirable to clarify that the name to be estimated is the name of the train station, not the name of the bus stop. Alternatively, more specifically, if the website 51 does not list a station name but lists the address of a facility, the scraping unit 23 may write in the prompt an instruction to identify the name of the nearest train station from the address. Furthermore, if the website 51 does not list a station name, the scraping unit 23 may write in the prompt an instruction to first extract the names of restaurants, tourist spots, etc., and then search a search engine using the extracted names as keywords to identify the name of the nearest train station.

[0041] Furthermore, the scraping unit 23 may instruct on a method for identifying at least some of the specified attributes. Specifically, the scraping unit 23 may describe the method for identifying at least some of the attributes in a prompt that describes instructions for the attribute extraction model 62. For example, in the example of FIG. 9, for tags, the scraping unit 23 may write in the prompt an instruction to extract words that appear frequently on the website 51 or words included in the input keywords 41 as tags.

[0042] (Step S33: End determination process) The scraping unit 23 determines whether all of the websites 51 identified in step S2 have been selected as target websites 51. If the scraping unit 23 has finished selecting all the websites 51, the process proceeds to step S34. On the other hand, if there is an unselected website 51, the scraping unit 23 returns the process to step S31.

[0043] (Step S34: Result generation process) The scraping unit 23 sets a pair of the URL and the attribute information 61 acquired in step S32 in the scraping result 63 for each website 51 identified in step S2.

[0044] The data output process (step S4 in FIG. 2) according to the first embodiment will be described with reference to FIGS. (Step S41: Data shaping process) As shown in FIG. 12, the data output unit 24 formats the scraping result 63 into a database format used in the application program to generate output data 64. For example, suppose the database contains an item for station ID, which is identification information for a station. ID is an abbreviation for IDentifier. In this case, the data output unit 24 refers to the station correspondence list 32, which associates station names with station IDs, identifies the corresponding station ID from the station name included in the scraping result 63, and adds it to the scraping result 63. Also, for example, suppose the database has a hash tag item. In this case, the data output unit 24 adds # to the beginning of each tag included in the scraping result 63 to create a hash tag.

[0045] (Step S42: Output process) The data output unit 24 outputs the output data 64 to the business operator terminal 91. The user checks the output data 64 using the business operator terminal 91, and then registers it in a database to be used by the application program.

[0046] ***Effects of the First Embodiment*** As described above, the information collection device 10 according to the first embodiment acquires attribute information 61 for a specified attribute from the website 51 using the attribute extraction model 62. As a result, even if the attribute information 61 is not clearly stated on the website 51, if the attribute information 61 can be estimated from the description on the website 51, the attribute information 61 can be acquired. This reduces the effort required to investigate and register the website 51.

[0047] When a user inputs an input keyword 41, the information collection device 10 according to the first embodiment generates a search keyword 42 using a keyword generation model 43. This makes it possible to identify an appropriate website 51 intended by the user. As a result, the effort required to research and register the website 51 can be reduced.

[0048] The information collection device 10 according to the first embodiment generates search keywords 42 from input keywords 41 using additional phrases set in the additional phrase list 31. This makes it possible to identify appropriate websites 51 that are suited to an application program by registering appropriate additional phrases that are suited to the application program. As a result, the effort required to research and register websites 51 can be reduced.

[0049] ***Other Configurations*** <Variation 1> In the first embodiment, the additional phrase list 31 is set with additional phrases to be added to all input keywords 41 and output keywords 44. An additional phrase may be set for each element that is not input with the input keyword 41. In the example of Figure 4, if the genre of the facility (2) has not been set, "restaurant," "tourist attraction," etc. are set as additional phrases. Also, if the type of facility (3) has not been set, "recommended," "trendy," etc. are set as additional phrases.

[0050] In step S13 of FIG. 3, the keyword generation unit 21 generates additional keywords 45 by adding additional words corresponding to elements that were not input in the input keywords 41. For example, suppose that "Fujisawa Station" is input for (1) and "stylish" is input for (3) as the input keywords 41. In other words, suppose that no input is made for (2). In this case, the keyword generation unit 21 generates, as search keywords 42, "Fujisawa Station, restaurant, stylish" by adding the additional word "restaurant" and "Fujisawa Station, tourist attraction, stylish" by adding the additional word "tourist attraction".

[0051] <Variation 2> In the first embodiment, each functional component is realized by software. However, as a second modification, each functional component may be realized by hardware. The differences between the first embodiment and the second modification will be described below.

[0052] When each functional component is realized by hardware, the information collection device 10 includes an electronic circuit instead of the processor 11, the memory 12, and the storage 13. The electronic circuit is a dedicated circuit for realizing the functions of each functional component, the memory 12, and the storage 13.

[0053] Possible electronic circuits include single circuits, composite circuits, programmed processors, parallel programmed processors, logic ICs, GAs, ASICs, and FPGAs. GA stands for Gate Array. ASIC stands for Application Specific Integrated Circuit. FPGA stands for Field-Programmable Gate Array. Each functional component may be realized by one electronic circuit, or each functional component may be realized by distributing it among a plurality of electronic circuits.

[0054] <Variation 3> As a third modification, some of the functional components may be realized by hardware, and other functional components may be realized by software.

[0055] The processor 11, memory 12, storage 13, and electronic circuitry are collectively referred to as a processing circuit. In other words, the functions of the functional components are realized by the processing circuit.

[0056] Furthermore, the term "unit" in the above description may be read as a "circuit," "step," "procedure," "process," or "processing circuit."

[0057] Embodiment 2 The second embodiment differs from the first embodiment in that if attribute information 61 cannot be acquired for a specific attribute among the specified attributes, a search is performed again to acquire the attribute information 61. In the second embodiment, this difference will be explained, and explanations of the same points will be omitted.

[0058] ***Explanation of Operation*** The scraping process (step S3 in FIG. 2) according to the second embodiment will be described with reference to FIG. The processing from step S31A to step S32A is the same as the processing from step S31 to step S32 in Fig. 8. The processing from step S35A to step S36A is the same as the processing from step S33 to step S34 in Fig. 8.

[0059] (Step S33A: Acquisition determination process) The scraping unit 23 determines whether the attribute information 61 acquired from the target website 51 includes information of a specific attribute. If the information of the specific attribute is included, the scraping unit 23 proceeds to step S35A. On the other hand, if the information of the specific attribute is not included, the scraping unit 23 proceeds to step S34A.

[0060] (Step S34A: Re-search process) The scraping unit 23 searches for a website 92 using, as a keyword, information acquired from the target website 51 about attributes other than the specific attribute among the multiple attributes that are the designated attributes. The website 92 here is not limited to the website 51 identified in step S2, but means any website 92 on the Internet. In this way, the scraping unit 23 acquires information about the specific attribute about the target website 51.

[0061] For example, in the example of Figure 9, the specific attribute is a station name. In this case, in step S33A, if the station name has not been acquired, the scraping unit 23 proceeds to step S34A. In step S34A, the scraping unit 23 performs an Internet search using the facility name or the like as a keyword to acquire the name of the nearest station.

[0062] ***Effects of the Second Embodiment*** As described above, when the information collection device 10 according to the second embodiment is unable to acquire attribute information 61 for a specific attribute, it performs a search again to acquire attribute information 61. This makes it possible to acquire information on a specific attribute based on the description on the website 51, even if the information on the specific attribute is not clearly stated on the website 51. This reduces the effort required to search and register the website 51.

[0063] Embodiment 3 The third embodiment differs from the first and second embodiments in that it checks whether information about a specific attribute among the designated attributes is appropriate. In the third embodiment, this difference will be explained, and explanation of the same points will be omitted. In the third embodiment, a case where a modification is made to the first embodiment will be described. However, it is also possible to make modifications to the second embodiment.

[0064] ***Explanation of Operation*** The scraping process (step S3 in FIG. 2) according to the third embodiment will be described with reference to FIG. The processing from step S31B to step S34B is the same as the processing from step S31 to step S34 in FIG.

[0065] (Step S35B: Information determination process) The scraping unit 23 determines whether the information of the specific attribute acquired in step S32B for each website 51 is appropriate. One possible method of determination is to prepare a list of information that can be taken as the specific attribute and determine whether the acquired information exists in the list. If the acquired information exists in the list, it is determined to be appropriate, and if the acquired information does not exist in the list, it is determined to be inappropriate.

[0066] If the scraping unit 23 determines that the information of the specific attribute is inappropriate, it sets information that the information of the specific attribute is inappropriate for the target website 51 in the scraping result 63. For example, the scraping unit 23 may indicate that the information of the specific attribute is inappropriate by marking the information of the specific attribute with an asterisk (*) or the like. This allows the user to manually or otherwise supplement the information on the specific attribute when checking the output data 64 output in the process of step S42.

[0067] ***Effects of the Third Embodiment*** As described above, the information collecting device 10 according to the third embodiment checks whether the information about the specific attribute is appropriate, and if it is inappropriate, sets information indicating that fact. This makes it possible to reduce the possibility of registering incorrect information.

[0068] ***Other Configurations*** <Variation 4> If it is determined to be inappropriate in step S35B, the scraping unit 23 may perform the re-search process (step S34A in FIG. 13) described in the second embodiment. This may enable appropriate information to be acquired for the specific attribute. When the scraping unit 23 performs the search process again, the scraping unit 23 may perform the information determination process (step S35B in FIG. 14) again to determine whether the information acquired in the search process again is appropriate.

[0069] Alternatively, the information determination process (step S35B in FIG. 14) described in the third embodiment may be simply performed at the end of the scraping process described in the second embodiment. That is, the information determination process (step S35B in FIG. 14) may be performed after the process of step S36A in FIG. 13.

[0070] Embodiment 4 The fourth embodiment differs from the first to third embodiments in that curation sites are excluded from the websites 51 identified in step S2. In the fourth embodiment, this difference will be explained, and explanation of the same points will be omitted. In the fourth embodiment, a case where a modification is made to the first embodiment will be described. However, it is also possible to make modifications to the second embodiment.

[0071] ***Explanation of Operation*** The scraping process (step S3 in FIG. 2) in the fourth embodiment will be described with reference to FIG. The processing from step S33C to step S35C is the same as the processing from step S32 to step S34 in FIG.

[0072] (Step S31C: Site exclusion process) The scraping unit 23 excludes curation sites from the websites 51 identified in step S2. Curation sites are also called aggregation sites. Specifically, the scraping unit 23 identifies websites 51 that have titles or the like containing phrases specific to curation sites, such as "Summary" or "Selection of x" (x can be replaced with a number), as curation sites. Then, the scraping unit 23 excludes websites 51 identified as curation sites.

[0073] (Step S32C: Site Selection Process) The scraping unit 23 selects an unselected website 51 as a target website 51 from among the websites 51 identified in step S2 that are not excluded in step S31C and remain.

[0074] ***Effects of the Fourth Embodiment*** As described above, the information collection device 10 according to the fourth embodiment excludes curation sites. Assume that website data is collected for an application program that collects websites related to restaurants, tourist spots, and the like around each station and makes them available to users. In such a case, one website is linked to one station. Curation sites comprehensively cover popular spots and the like in each region, and are therefore not suitable for the above-described application program format. Therefore, it is effective to exclude curation sites.

[0075] Various aspects of the present disclosure are summarized below as appendices. (Appendix 1) A scraping unit that inputs website information of a website to be collected into an attribute extraction model, which is a learning model, and acquires attribute information output by the attribute extraction model, which is information about specified attributes related to the website; a data output unit that outputs the attribute information acquired by the scraping unit in association with the website; An information collection device comprising: (Appendix 2) The information collection device further a keyword generation unit that inputs input keywords input by a user into a keyword generation model that is a learning model, acquires output keywords output by the keyword generation model, and generates search keywords from the acquired output keywords; a crawling unit that searches websites using the search keywords generated by the keyword generating unit and collects website information of the websites that are the information collection targets; 2. The information collection device according to claim 1, comprising: (Appendix 3) The keyword generation unit sets the input keyword and the output keyword as the search keyword. 10. The information collection device of claim 2. (Appendix 4) The keyword generating unit generates an additional keyword by adding an additional phrase set in an additional phrase list to the input keyword and the output keyword, and sets the additional keyword as the search keyword. 4. The information collection device according to claim 2 or 3. (Appendix 5) The specified attribute is a plurality of attributes, When the attribute information acquired from a specific site that is one of the websites that are the information collection targets does not include information on a specific attribute among the plurality of attributes, the scraping unit acquires information on the specific attribute for the specific site by searching the website using information acquired from the specific site about an attribute other than the specific attribute among the plurality of attributes as a keyword. 5. The information collection device according to any one of appendixes 1 to 4. (Appendix 6) The scraping unit determines whether the attribute information exists in a list. 7. The information collection device according to any one of appendixes 1 to 6. (Appendix 7) The scraping unit acquires the attribute information from websites remaining after excluding curation sites from the websites to be collected. 7. The information collection device according to any one of appendixes 1 to 6. (Appendix 8) The specified attributes include a nearest station name and a facility name, The data output unit outputs the attribute information to a system that manages information for each station. 8. The information collection device according to any one of appendixes 1 to 7. (Appendix 9) When the attribute information acquired from a specific site that is one of the websites that are the information collection targets does not include a nearest station name, the scraping unit acquires the nearest station name for the specific site by searching the website using the facility name acquired from the specific site as a keyword. 9. The information collection device of claim 8. (Appendix 10) a computer inputs website information of a website to be collected into an attribute extraction model, which is a learning model, and acquires attribute information output by the attribute extraction model, the attribute information being information on a designated attribute related to the website; An information gathering method in which a computer outputs the attribute information in association with the website. (Appendix 11) A scraping process in which website information of a website to be collected is input into an attribute extraction model, which is a learning model, and attribute information output by the attribute extraction model, which is information about specified attributes related to the website, is acquired; a data output process of outputting the attribute information acquired by the scraping process in association with the website; An information gathering program that causes a computer to function as an information gathering device.

[0076] The embodiments and modifications of the present disclosure have been described above. Some of these embodiments and modifications may be combined and implemented. Also, one or more of them may be implemented partially. Note that the present disclosure is not limited to the above embodiments and modifications, and various modifications are possible as needed. [Explanation of symbols]

[0077] 10 Information collection device, 11 Processor, 12 Memory, 13 Storage, 14 Communication interface, 21 Keyword generation unit, 22 Crawling unit, 23 Scraping unit, 24 Data output unit, 31 Additional phrase list, 32 Station correspondence list, 41 Input keyword, 42 Search keyword, 43 Keyword generation model, 44 Output keyword, 45 Additional keyword, 51 Website, 52 Website information, 53 Crawling result, 61 Attribute information, 62 Attribute extraction model, 63 Scraping result, 64 Output data, 90 Internet, 91 Operator terminal, 92 Website.

Claims

1. A scraping unit that inputs website information of a website to be collected into an attribute extraction model, which is a learning model, and acquires attribute information output by the attribute extraction model, which is information about specified attributes related to the website; a data output unit that outputs the attribute information acquired by the scraping unit in association with the website; An information collection device comprising:

2. The information collection device further a keyword generation unit that inputs input keywords input by a user into a keyword generation model that is a learning model, acquires output keywords output by the keyword generation model, and generates search keywords from the acquired output keywords; a crawling unit that searches websites using the search keywords generated by the keyword generating unit and collects website information of the websites that are the information collection targets; The information collection device according to claim 1 , comprising:

3. The keyword generation unit sets the input keyword and the output keyword as the search keyword. The information collection device according to claim 2 .

4. The keyword generating unit generates an additional keyword by adding an additional phrase set in an additional phrase list to the input keyword and the output keyword, and sets the additional keyword as the search keyword. The information collection device according to claim 2 .

5. The specified attribute is a plurality of attributes, When the attribute information acquired from a specific site that is one of the websites that are the information collection targets does not include information on a specific attribute among the plurality of attributes, the scraping unit acquires information on the specific attribute for the specific site by searching the website using information acquired from the specific site about an attribute other than the specific attribute among the plurality of attributes as a keyword. The information collection device according to claim 1 .

6. The scraping unit determines whether the attribute information exists in a list. The information collection device according to claim 1 .

7. The scraping unit acquires the attribute information from websites remaining after excluding curation sites from the websites to be collected. The information collection device according to claim 1 .

8. The specified attributes include the name of the nearest station and the name of the facility, The data output unit outputs the attribute information to a system that manages information for each station. The information collection device according to claim 1 .

9. When the attribute information acquired from the specific site that is the website to be collected does not include the name of the nearest station, the scraping unit acquires the name of the nearest station for the specific site by searching the website using the facility name acquired from the specific site as a keyword. The information collection device according to claim 8.

10. a computer inputs website information of a website to be collected into an attribute extraction model, which is a learning model, and acquires attribute information output by the attribute extraction model, the attribute information being information on a specified attribute related to the website; An information gathering method in which a computer outputs the attribute information in association with the website.

11. A scraping process in which website information of a website to be collected is input into an attribute extraction model, which is a learning model, and attribute information output by the attribute extraction model, which is information about specified attributes related to the website, is acquired; a data output process of outputting the attribute information acquired by the scraping process in association with the website; An information gathering program that causes a computer to function as an information gathering device.

Citation Information

Patent Citations

  • Information extraction method and information extraction device

    JP2006155275A

  • Information retrieval apparatus, method, and program

    JP2013045182A

  • Information display device, information display method, and information display program

    JP2015068672A

  • Information management system, information management apparatus, information management method, and information management program

    JP2024008408A