Information retrieval system
The information retrieval system addresses the challenge of accurately using abstract words from social media posts by calculating frequency and co-occurrence relationships to enhance search accuracy.
Patent Information
- Application Number
- JP2024086136
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-28
- Publication Date
- 2025-12-10
AI Technical Summary
Existing information retrieval systems struggle to accurately search for information that matches users' intent, particularly when using abstract words like impressions and experiences from social media posts, as these words are less frequent and difficult to weight effectively.
An information retrieval system that extracts words from social media posts, calculates their frequency and co-occurrence relationships, and weights them to enhance accuracy in information retrieval, allowing for the use of abstract words in searches.
Enables highly accurate information retrieval by using the frequency and co-occurrence relationships of words extracted from social media posts, improving the relevance of search results to users' intentions.
Smart Images

Figure 2025179410000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information retrieval system for performing information retrieval based on messages posted on a network. Regarding the topic. [Background technology]
[0002] Conventionally, there has been a trend of using information on the vast amount of information that exists on computer networks and databases. When searching for information, the user enters a search word (also called a search phrase). and retrieves information on computer networks and databases that matches the search words you entered. However, some of the information that matches the search words may not be relevant to the user. It is assumed that the search results contain a lot of information that is different from the user's search intent. Search for information by setting criteria such as the reliability of the source site, frequency of access, and trend order. However, it has been difficult to properly search for information that matches the user's search intent.
[0003] On the other hand, in recent years, when users post messages on a computer network via their terminals, Both systems provide a way to view posts posted by other users. For example, blogs, SNS (Social Network Service), X (registered trademark), chat, etc. (hereinafter referred to as SNS, etc.). In the above SNS, etc., users actually experience or feel the most For example, JP 2019-149145 A has the advantage of being able to quickly obtain the latest information. The report includes words obtained by morphological analysis of posts posted on social media etc. about the spot. Generate keywords, link them to the spot, store them in the database, and then use the search results entered. A technology has been proposed for performing spot searches using words and generated feature words. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2019-149145 A (paragraphs 0065, 0109-0113) Summary of the Invention [Problem to be solved by the invention]
[0005] Here, as in Patent Document 1, it is possible to generate feature words from posts posted on SNS etc. In this case, in addition to words specific to the spot, such as the name and address, abstract words such as impressions and experiences are also included. The advantage is that you can search for spots using simple words (such as "fun" or "beautiful"). On the other hand, since posts are written down with each individual's thoughts and feelings, all of the content is It is difficult to say that the content is accurate. In particular, the spots with many posts are The number of feature words that can be added will increase, and it is expected that many inappropriate feature words will also be included. Reference 1 also discloses that weighting of feature words is performed based on the frequency of occurrence. In weighting, abstract words such as impressions and experiences appear less frequently than unique words. As a result, the benefit of generating feature words from posted text was lost. It was said.
[0006] The present invention has been made to solve the above-mentioned problems in the past, and Abstract words such as impressions and experiences, using words extracted from submitted posts as keywords It also enables information search by the frequency of occurrence of words extracted from posts. In addition, by weighting each word using the co-occurrence relationship between words, i.e., the correlation between words, the accuracy is improved. The purpose of this invention is to provide an information retrieval system that enables highly sophisticated information retrieval. . [Means for solving the problem]
[0007] In order to achieve the above object, the information retrieval system according to the present invention is It is an information retrieval method that searches for information based on words and acquires posts posted on the network. and for each piece of information to be searched by the information search means, By analyzing the posted text, words contained in the posted text are extracted. At the same time, the frequency with which the extracted words appear in the posted text and the relationship between the extracted words are calculated. A posted text analysis means for identifying co-occurrence relationships, and based on the analysis results of the posted text analysis means, For each piece of information to be searched by the information search means, the frequency of appearance in the posted text is equal to or greater than a threshold. A word is set as a frequently occurring word, and a related word that is a word related to the frequently occurring word is set as a related word. and a word setting means for setting the word based on the co-occurrence relationship, and the information retrieval means Searching for information in which the frequently occurring words or related words that match the search word are set In addition, information retrieval is performed by weighting the frequently occurring words more heavily than the related words. Furthermore, "search for information that contains frequently occurring words or related words that match the search word" "Match" does not necessarily mean that the search word and the frequently occurring word or related word match exactly. This also includes cases where there is only a partial match or where there is a certain percentage of similarity. [Effects of the Invention]
[0008] According to the information retrieval system of the present invention having the above-described configuration, posts posted on SNS etc. By extracting words contained in the sentences and using them as keywords, the words extracted from the posts can be used as keywords. It is possible to search for information using abstract words such as impressions and experiences. For words extracted from a sentence, in addition to the frequency of occurrence, the co-occurrence relationship between words, i.e., the ? between words, is calculated. By using the weighting function to weight each word, information retrieval using the weighting function can be performed with high accuracy. It becomes possible to perform information searches. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a schematic configuration diagram showing an information search system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram schematically illustrating a control system of the information providing server according to the embodiment. [Figure 3] FIG. 10 is a diagram showing an example of posted message information stored in a posted message information DB. [Figure 4] FIG. 10 is a diagram showing an example of information stored in a distribution information DB. [Figure 5] FIG. 10 is a diagram showing an example of information stored in a search setting word DB. [Figure 6] FIG. 2 is a block diagram schematically illustrating a control system of the communication terminal according to the present embodiment. [Figure 7] 10 is a flowchart of a word setting processing program according to the present embodiment. [Figure 8] FIG. 10 is a diagram illustrating co-occurrence analysis using a co-occurrence network. [Figure 9] 10 is a flowchart of an information provision processing program according to the present embodiment. [Figure 10] FIG. 10 is a diagram showing a search word input screen displayed on the communication terminal. [Figure 11] FIG. 10 is a diagram showing a search result screen displayed on the communication terminal. [Figure 12] FIG. 10 is a diagram showing an example of weighting of search words. [Figure 13] FIG. 10 is a diagram illustrating a method for calculating points indicating a matching rate with a search word. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, an information retrieval system according to an embodiment of the present invention will be described with reference to the drawings. First, the general configuration of the information retrieval system 1 according to this embodiment will be described. This will be explained with reference to Fig. 1. Fig. 1 shows a schematic configuration of an information retrieval system 1 according to this embodiment. Figure.
[0011] As shown in FIG. 1, the information retrieval system 1 according to this embodiment includes an information providing center 2. The system basically comprises an information providing server 3 and a communication terminal 5 held by a user 4. The information providing server 3 and the communication terminal 5 transmit and receive electronic data to and from each other via a communication network 6. The communication terminal 5 may be, for example, a mobile phone, a smartphone, a tablet, or the like. Examples include tablet terminals, personal computers, and in-vehicle navigation devices.
[0012] Here, the information providing server 3 sends the following to the communication terminal 5 (i.e., the user who owns the communication terminal 5): The information providing server 3 is a server device that manages the information provided to the communication terminal 5. Information on information providing locations around the country that are the subject of the information provision is stored in the distribution information DB7. In addition, there are no particular restrictions on the genre or size of the information provided locations, for example, restaurants, Commercial facilities such as retail stores, public facilities such as stations and hospitals, as well as accommodation facilities and parking lots, etc. In addition, the location is not limited to a facility, but may be a tourist spot, etc. The information about the information providing point stored in the distribution information DB 7 is as follows: In addition to the name, address, location coordinates, telephone number, business hours, etc., a description of the information providing point This also includes "posted texts" posted to the information source on the network. Then, the information providing server 3 stores the information in the DB in response to a request from the communication terminal 5 as described later. The information to be provided is searched for from the information on the information provision point that has been received, and the information is used to The information is provided (distributed) to the communication terminal 5 via the network 6 .
[0013] The communication terminal 5 is an information terminal that is carried by a user and has a communication function, a navigation function, etc. For example, mobile phones, smartphones, tablet devices, personal computers In particular, the communication terminal 5 is a smartphone or the like. If the device is capable of running applications, the information will be included as part of the application. An application program that provides information about the location is installed. Specifically, when a user enters a word that will be used as a search condition (hereinafter referred to as a search word), Information about the information providing location related to the entered search word is distributed from the information providing server 3. This is a program that is provided by the Ministry of Land, Infrastructure, Transport and Tourism. This function may be part of the navigation function that provides guidance to the destination, or may be separate from the navigation function. It may be executed by an application program other than the one described above. may have a function of setting the provided information provision point as a destination.
[0014] In addition, the communications network 6 is made up of numerous base stations located throughout the country and a network that manages each base station. and the telecommunications companies that control them, and base stations and telecommunications companies that connect them to wired networks (optical fiber, ISDN, etc.) ) or wirelessly connected to each other. The base station has a transceiver (transmitter / receiver) and an antenna for communication. While wireless communication is performed between companies, it also serves as the end of the communication network 6 and is not covered by the radio waves from the base station. It has the role of relaying communications between the communication terminals 5 within its range (cell) and the information providing server 3.
[0015] On the other hand, the information providing server 3 included in the information retrieval system 1 exists on the network. SNS services that provide social networking services (hereinafter referred to as SNS) Server 8 identifies the location associated with a message posted on the network. The post can be obtained along with the location information (hereinafter, the post and location information will be collectively referred to as post information). The posted message information can be obtained via a network. The information may be acquired via a storage medium such as a flash memory.
[0016] Here, the SNS server 8 is a server device that builds an SNS on the network. NS is a community for building connections with others (not only individuals but also corporations). It is a mobile-based service, and service users are clients who use smartphones, tablets, etc. Access the site from a personal computer or other device and post messages, images, etc. By doing so, other service users can view the content. Other service users who view the image can comment on the post. You can quote the content and retweet (repost) it, or follow the poster using the follow function. or if you agree with the content of the post, you can give a positive reaction (for example, "Like" or "Like"). In this embodiment, there is no particular restriction on the users. It is an open SNS that can be used by anyone who is a registered user. The server 8 is equipped with a storage DB 9, which stores messages and image data posted by service users. The data, the date and time of the post, hashtags, and the location where the post was made (even the facility name or location coordinates) (This is also good), and the number of positive reactions and retweets from other service users for each post. The number of users and the number of followers per service user are stored in the storage DB9.
[0017] In this embodiment, the information providing server 3 stores the above data in the storage DB 9. , especially the content of the post (it can be just the message, or if there is an image attached, the image as well) The combination of the post and the location where the post was posted is recorded as post information. and retrieves it from the SNS server 8.
[0018] Next, the configuration of the information providing server 3 in the information retrieval system 1 will be explained with reference to FIG. The information providing server 3 includes a server control unit 11 and a server A posted message information DB 13 and a distribution information DB 7 are connected to the control unit 11 as information recording means. The system includes a map information DB 14, a search setting word DB 15, and a server-side communication device 16.
[0019] The server control unit 11 is a control unit (MCU, MCU) that controls the entire information providing server 3. PU, etc.), and the CPU 21 as a calculation device and a control device, and RAM22 is used as working memory for processing calculations, and In addition to the gram, the word setting processing program (Fig. 7) and information provision processing program (Fig. 9) are also included. ) is recorded in ROM23, and the flash memory stores the program read from ROM23. The server control unit 11 is provided with an internal storage device such as a memory 24. For example, the information search means searches for information based on search words entered by the user. The posted message acquisition means acquires posted messages posted on the network. The posted text analysis means performs the following for each piece of information that is searched by the information search means: By analyzing the posted text, the words contained in the text are extracted and Identify the frequency with which the extracted words appear in posts and the co-occurrence relationships between the extracted words The word setting means sets search targets to be searched by the information search means based on the analysis results of the posted text analysis means. For each piece of information, we set the words that appear in posts with a frequency above a threshold as frequent words. In both cases, related words that are words related to frequently occurring words are set based on co-occurrence relationships.
[0020] The posted message information DB 13 stores posted message information acquired from the SNS server 8. Here, the posted message information includes the message posted on the network as described above. The posted message includes location information that identifies the location linked to the posted message. In this embodiment, the "place where the message was posted" refers to the place where the message was posted (the place where the message was posted). However, if the post contains a place name, it can be used as a hashtag. If a place name is linked to the post, that place name may be used as the location linked to the post. The posted message information stored in the posted message information DB 13 is information about the information providing location. The information is provided to users as part of the information provision point, as described later. It is used to set frequently occurring words and related words that are used as keywords to search for information points. can be.
[0021] FIG. 3 is a diagram showing an example of posted message information stored in the posted message information DB 13. As shown in FIG. The posted message information DB 13 stores the contents of posted messages (message text data) as shown in FIG. ) is stored in association with the location name that indicates the location associated with the posted text. In the example shown, the location names linked to the posts are stored separately, but how do you The classification and storage can be changed as needed. In the example shown in Figure 3, The content is limited to text information, but if an image is attached, it may also be included. In addition, the location name is stored as information indicating the location linked to the posted text. The location coordinates may be used instead of the name. When receiving a message posted on the network by the service user, The location coordinates of the device used in the post (identified by the device's GPS, etc.) will also be included in the post. Therefore, as shown in Figure 3, the location of the posted message is also acquired. When you want to remember the specific location name where you posted, refer to the map information and obtain the location coordinates. It is necessary to identify the location name of the location where the service user is expected to be located. This processing may be performed by the SNS server 8 or the information providing server 3. The date and time of submission may also be stored.
[0022] In addition, as mentioned above, the distribution information DB7 is a database of information provision areas that are the target of information provision across the country. 4 is a storage means for storing various information related to points. FIG. 10 is a diagram showing an example of information to be displayed.
[0023] As shown in FIG. 4, the distribution information DB 7 stores information providing points all over the country. D. Name of the location, coordinates of the location, description, and posted text about the information providing location, etc. The "description" contains more detailed information about the information point, such as The service content, facilities, business hours, services provided, menu, etc. However, the distribution information DB 7 does not necessarily store all of this information. It is not necessary to store the entire "posted text." Only keywords indicating the evaluation or status of the location may be extracted and stored.
[0024] For example, in the distribution information DB7 shown in FIG. 4, “XX Ramen” is located at the position coordinates (x1, y1). " facility information and posted text are stored. Similarly, information about other information providing points is stored. The information is also stored.
[0025] The map information DB 14 is a storage means for storing map information. It consists of various information necessary for route search, route guidance and map display, including the network. For example, link data for roads (links), node data for node points, and each intersection Intersection data for points, location data for facilities, etc., and maps for displaying maps From display data, search data for searching routes, search data for searching locations, etc. become.
[0026] Then, the server control unit 11 responds to a request from the communication terminal 5 by using the map information DB 14. For example, map display data for displaying a map image on the communication terminal 5 may be transmitted, or input. When a location matching the entered search criteria is searched for or a route search request is received, Search for a route from the departure point to the destination using map information stored in the map information DB 14 It is also possible.
[0027] However, if the communication terminal 5 has map information, the communication terminal 5 uses the map information that it has. In this case, the information providing server 3 can perform the above processing. The information DB 14 is not necessarily required.
[0028] In addition, the search setting word DB15 is stored for each information provision location that is the target of information provision across the country. Frequently used words and related words that are used as keywords to search for information points are linked and ranked. The frequently used words and related words are stored in the posted message information DB 13. It is set based on the posted message information. More specifically, as described later, The posted texts for the information providing point are morphologically analyzed, and the frequency of appearances in the posted texts is analyzed. Words with a frequency above a threshold are set as frequent words, and related words are set as related words. The words are set based on co-occurrence relationships (co-occurrence networks), as will be described in detail later.
[0029] Here, FIG. 5 is a diagram showing an example of information stored in the search setting word DB 15. For example, For example, in the example shown in Figure 5, the frequently occurring word for one of the information provision points, “○○ Ramen,” is Four words were set as the target: "ramen," "delicious," "beautiful," and "kind." The four words set are "spicy," "soup," "gyoza," and "fun." Then, the server control unit 11 stores the search setting word DB 15 as described later. The system uses the frequently used words and related words to search for information for the user. The information provided is based on frequently occurring words or related words that match the search words entered by the user. Search for information about the location. Note that "match" means that the search word matches a frequently occurring word or related word. It is not limited to cases where the words match completely, but may match only partially or have a certain degree of similarity. Also, there are cases where multiple words are entered as search words, and If any of the entered words match, information is provided that includes frequently occurring words or related words. Search for information about the location of the service. In addition, frequently occurring words and related words are newly added to the network. It is updated periodically (e.g., monthly) to reflect the information posted on the network.
[0030] On the other hand, the server-side communication device 16 is connected to the communication terminal 5, which is the target of information transmission and reception, and the communication network It is a communication device for communicating via the network 6. In addition to the communication terminal 5, and traffic information centers, such as VICS (Vehicle Information and Communications Information such as congestion information, regulation information, and traffic accident information sent from the Traffic Information System Center, etc. It is also possible to receive traffic information consisting of the above. By sending this information, you can receive weather information for each region of the country, information about events being held across the country, and more. It will also be possible to receive event information, local news, congestion information at locations, etc.
[0031] Next, the schematic configuration of the communication terminal 5 owned by the user will be described with reference to FIG. FIG. 1 is a block diagram showing a control system of a communication terminal 5 according to the present embodiment. The following description will be given taking as an example a case where the communication terminal 5 is a smartphone.
[0032] As shown in FIG. 6, the communication terminal 5 has a CPU 31 and a communication terminal 5 on a data bus BUS. User information (user ID, name, etc.) and application programs, etc. and an interface such as a microphone 33 and a speaker 34. an input / output unit 35, a display 36 configured with a liquid crystal display panel or the like, and a touch panel or An input operation unit 37 including a keyboard, a GPS 38, and a communication network 6 A transmitting / receiving circuit unit (RF) 39 that transmits and receives signals to and from a base station is connected. It is composed of:
[0033] Here, the CPU 31 built in the communication terminal 5 executes the operation program stored in the memory 32. It is a control means of the communication terminal 5 that executes various operations according to the program, and together with the memory 32 The communication terminal control unit 41 is configured. The various processing contents of the communication terminal control unit 41 are The result is displayed on the display 36.
[0034] Furthermore, the communication terminal 5 performs communication via the transmission / reception circuit unit 39, and in addition to making calls, Internet communication, receiving information about charging facilities from the information providing server 3, and traffic information Traffic congestion information transmitted from a center, such as a VICS (registered trademark) center or a probe center. It is also possible to receive traffic information consisting of various information such as road safety information, traffic regulation information, and traffic accident information.
[0035] The memory 32 also stores user information (user ID, Name, etc.), map information, user web browsing history, GPS38 and other sensors The user's movement history is a history of location information detected based on the It is a storage medium that stores search word history, etc. It also stores the information provision processing program ( Various application programs including the program code (see FIG. 9) are also stored in the memory 32. may be configured by a hard disk, a memory card, etc.
[0036] In addition to the voice output of the call, the speaker 34 also outputs the voice of the communication terminal when the navigation function is executed. Guides driving along a guide route (a planned route for the user) based on instructions from the control unit 41 The device outputs voice guidance.
[0037] The display 36 is disposed on one side of the housing and may be a liquid crystal display or an organic electroluminescence display. The various applications installed in the communication terminal 5 are displayed on the LCD screen. The top screen for running applications and the screen related to the executed application ( Internet screen, email screen, navigation screen, etc.), as well as various information such as images and videos, are displayed. In particular, in this embodiment, an application program that provides information on points of interest is started. When the user clicks the search button, an input screen is displayed in which the user can enter search terms. When a search word is entered, information about the information providing point related to the entered search word is displayed. The information is distributed from the information providing server 3 and displayed on the display 36.
[0038] The input operation unit 37 may be a touch panel provided on the front of the display 36 or a display on the housing. The communication terminal control unit 41 is configured by hard buttons and the like arranged on the touch panel. Based on the electrical signals output by pressing the touch panel or hard buttons, The input operation unit 37 controls the number / character input keys, the displayed Various keys such as cursor keys to move the cursor to select content, and a confirmation key to confirm the selection It can also be configured with a key or the like.
[0039] In addition, GPS38 receives radio waves generated by artificial satellites. The current location and current time of the receiving terminal 5 (i.e., the user) can be detected. In addition, other devices (such as a gyro sensor) for detecting the current position and direction of the communication terminal 5 may also be used. It may also be configured to include:
[0040] The transmission / reception circuit unit 39 also communicates with a communication network according to communication standards such as 3G, 4G, and LTE. This is a circuit section for transmitting and receiving signals to and from base stations in the network 6.
[0041] Next, in the information retrieval system 1 having the above configuration, the information providing server 3 executes The word setting processing program will be described with reference to FIG. 7. FIG. 7 shows the word setting processing program according to this embodiment. 1 is a flowchart of a word setting processing program. The message is executed after a predetermined period (for example, one month) has elapsed since the message was first executed, and is stored in the posted message information DB 13. By analyzing the posted information, the information providing point can be searched for for each information providing point. It is a program that sets (generates) frequently occurring words, related words, and explanatory words that are keywords for The program shown in the flowchart of FIG. 7 is provided in the information providing server 3. The program is stored in the RAM 22, ROM 23, etc., and is executed by the CPU 21.
[0042] Here, the word setting processing program processes information provision points across the country, and For each location where information is provided, "frequently occurring words," "related words," and "explanatory words" are specified. However, it is not necessary to execute the process for all information providing points in the country. For example, it is possible to target only specific genres or locations of a certain size or larger. Furthermore, the interval at which the word setting processing program is executed may be changed for each location.
[0043] First, in step (hereinafter abbreviated as S) 1, the CPU 21 checks the posted message information DB 13 Determine whether there is any posted message information linked to the information provision location to be processed this time. In addition, in order to set accurate "frequent words" and "related words," older posts will not be analyzed. It is desirable to exclude posts posted within the last three months or six months. Also, it is determined whether there is at least one piece of posted message information. However, if the number of samples is small, it is difficult to set accurate "frequent words" and "related words." Therefore, it may be determined whether there is a predetermined number (for example, 3) or more pieces of posted message information. .
[0044] As shown in FIG. 3, the posted message information DB 13 is previously acquired from an external SNS server 8. The posted message information is stored, especially the content of the posted message (text data of the message). is stored in association with the location name that indicates the location associated with the post (the place where the post was posted). Therefore, in S1, the information provided point is linked to the name of the point to be processed. It is determined whether or not there is any posted message information.
[0045] Then, the posted message information DB13 contains a post linked to the information providing location to be processed this time. If it is determined that there is sentence information (S1: YES), the process proceeds to S2. The posted message information DB 13 stores posted message information linked to the information providing location to be processed this time. If it is determined that there is no information provided (S1: NO), the information providing point to be processed is not subject to analysis. Since no target post has been posted, we move on to S8.
[0046] In S2, the CPU 21 selects the information provider location to be evaluated this time from the posted message information DB 13. The information on posts linked to the points is extracted and acquired. In order to set the "words", it is desirable to exclude old posts from the analysis. Information on posts posted within the last month or last six months is obtained.
[0047] The following process in S3 is performed for each piece of posted message information acquired in S2. After executing the process of S3 for all the posted message information, the process proceeds to S4.
[0048] In S3, the CPU 21 processes the posted message information to be processed, particularly the text indicating the posted content. Morphological analysis is performed on the data using dedicated software. Perform steps (1) to (3). (1) Format the text data of the post into an analyzable format. (2) Divide the formatted text data into words. (3) Over-split words are combined to form meaningful words, and then the words become noise. Meaningless words (stop words) are removed. For example, particles and auxiliary verbs that have no meaning on their own are removed. Excluding words that have no meaning, 9 parts of speech (nouns, sahen nouns, adjectives, proper nouns, organization names) , names of people, places, verbs, adjectives). As a result of the above process, words contained in the posted message information to be processed are extracted. Then, words are extracted from all the posted message information acquired in S2.
[0049] Thereafter, in S4, the CPU 21 By performing co-occurrence analysis on the data, we can identify the connections between words contained in the posted text information. For example, in co-occurrence analysis using a co-occurrence network, dedicated software is used to identify the degree of association. By using the software and inputting the text data to be processed, a co-occurrence network like that shown in Figure 8 is generated. In the co-occurrence network 45 of FIG. Co-occurrence network for only the words finally extracted by morphological analysis (i.e., 9 parts of speech) 45, but the co-occurrence network also includes particles, auxiliary verbs, and other words that do not have meaning on their own. In addition, if the frequency of appearance in the posted text is more than a certain number, the network 45 may be output. Alternatively, the co-occurrence network 45 may be output by targeting only the words.
[0050] In the co-occurrence network 45 shown in FIG. 8, each circle (node) represents a word, and the lines connecting the circles represent (Link) indicates the relationship between words. Here, the relationship between words can be expressed by, for example, the well-known Jacc ard coefficient (similarity based on the number of simultaneous appearances) and Cos similarity (vectors with similar orientation: close to 1) For example, the Jaccard coefficient can be used to identify the similarity. In this case, the Jaccard coefficient is the product of a combination of words and a union. More specifically, either or both of two specific words were used. The percentage of text that uses both, a value between 0 and 1 The larger the value, the greater the proportion of texts in which both words are used, and This means that the similarity between two sets of texts containing the word is high.
[0051] Furthermore, words with a Jaccard coefficient of 0.1 or more are related to each other (co-occurrence relationship). The thickness of the link is calculated by the Jaccard coefficient. For example, the closer to 1, the thicker the link. The words connected by the link indicate that they are related (they have a co-occurrence relationship), and The thicker the link, the stronger the co-occurrence relationship. According to Q45, the information provided at the location targeted for processing this time is included in the linked posted information. This makes it possible to identify co-occurrence relationships between words.
[0052] Thereafter, in S5, the CPU 21 extracts the morphological information from the posted message information by the morphological analysis in S3. For each word, the frequency of its appearance in the posted message information acquired in S2 is analyzed. Then, only words whose frequency of occurrence is equal to or greater than a threshold are extracted as frequent words. For example, If the average occurrence frequency (number of occurrences) of each word is X, the frequency of occurrence is X+3σ(σ: The words that are more than the standard deviation are extracted as frequent words. This can be changed as needed. For example, if the number of occurrences per posted sentence is less than a threshold (e.g., 0.1 times / sentence), The above condition can also be used as a condition for a frequently occurring word.
[0053] Next, in S6, the CPU 21 refers to the result of the co-occurrence analysis in S4 and Words that are related to the extracted frequently occurring words, i.e., words that have a co-occurrence relationship, are extracted as related words. However, words that have already been recognized as frequently occurring words are excluded from the related words. For example, in the co-occurrence network 45 shown in FIG. 8, If a word is linked to a frequently occurring word by a link thicker than a certain thickness, it can be recognized as having a co-occurrence relationship. As mentioned above, the co-occurrence network 45 In the example, the thicker the link between words, the stronger the co-occurrence relationship. If we use the Jaccard coefficient, for example, we will use words that have a relationship of 0.1 or more with a frequently occurring word. The word may be a word, or a word having a relationship of 0.2 or more with a frequently occurring word may be a related word. In addition, there are cases where multiple related words are extracted for one frequently occurring word. You can extract all related words, or you can extract the words that have the strongest co-occurrence relationship with one frequently occurring word. It is also possible to extract only the related words.
[0054] Then, in S7, the CPU 21 compares the frequently used words extracted in S5 with the frequently used words extracted in S6. The related words are linked to the information provision point that will be the evaluation process this time and stored in the search setting word DB15. As a result, as shown in Figure 5, the information provision points that are the target of information provision across the country are Each location is associated with frequently occurring words and related words that serve as keywords for searching for that information point. The data will be stored in the form.
[0055] Next, in S8, the CPU 21 selects the information to be evaluated from the distribution information DB 7. The description of the location is obtained. The "description" contains detailed information about the location, such as For example, the services provided, facilities, opening hours, services offered, menu, etc. The description can be, for example, the person involved in the relevant information point or the information provider. It is assumed that the information has been input in advance by the operator of the server 3. Then, the CPU 21 acquires The text data showing the explanatory text was analyzed using dedicated software in the same way as in S3. Morphological analysis is performed. As a result, words contained in the explanatory text (hereinafter referred to as explanatory words) are extracted. can be.
[0056] Thereafter, in S9, the CPU 21 uses the explanatory words extracted in S8 in the current evaluation process. The stored explanation is stored in the search setting word DB 15 in association with the information providing point that serves as the explanation. The words are used as keywords to search for information points along with the frequently occurring words and related words mentioned above. become.
[0057] Next, in the information retrieval system 1, the information providing server 3 and the communication terminal 5 execute the information retrieval system. The information provision processing program will be described with reference to FIG. 9. FIG. 9 shows the information provision processing program according to this embodiment. 10 is a flowchart of a processing program. Here, the information provision processing program is executed by a communication terminal. In step 5, a predetermined application program is started to obtain information on the information providing point. It is executed after the search term is entered and provides relevant information points based on the search term entered by the user. This is a program that searches for information about and provides it to the user. The programs shown in the chart are stored in the RAM and It is stored in the ROM and executed by the CPU 21 or CPU 31.
[0058] First, the information provision processing program executed by the CPU 31 of the communication terminal 5 will be explained with reference to FIG. In step S21, the CPU 31 executes a predetermined access command to obtain information about an information providing point. The application program (hereinafter referred to as the information provision application) is started. The app can be a navigation app or a dedicated app that is different from the navigation app. It may be an application program. The information providing application is installed in advance on a web server, etc. It is assumed that the software has been downloaded and installed on the communication terminal 5.
[0059] When the information providing application is started on the communication terminal 5, the following message is first displayed on the display 36: A search word input screen 51 is displayed (S22). 10 is a diagram showing a search word input screen 51. FIG.
[0060] Then, when the search word input screen 51 is displayed, the user operates the input operation unit 37 to When you enter a search word (also called a search phrase), the search results will be The search word is displayed in the search word input section 52 of the search word input screen 51 (S23). A word used to identify the location where the user wants information when searching for a location , that is, search conditions, and as will be described later, the information providing server 3 receives information related to the input search word. (More specifically, at least one of the frequent words, related words, and explanatory words is matched to the search word. Search for information points that match the search criteria. You can also enter multiple words as search terms. For example, in the example shown in Figure 10, the search words are "ramen," "gyoza," and "oi" Five words are entered: "good," "fun," and "energetic."
[0061] In addition to the search words, exclusion words can also be entered on the search word input screen 51. Then, when the search word input screen 51 is displayed, the user operates the input operation unit 37. If you enter an exclusion word (also called an exclusion phrase) in are displayed in the exclusion word input section 53 of the search word input screen 51 (S23). When a user searches for information provision points, the user can select information provision points that the user wants to exclude from the information provision target. The search terms are used to identify the search results, i.e., the exclusion conditions. As will be described later, the information providing server 3 Even if the information provided by the relevant site is related to the entered exclusion words (more specifically, the exclusion words Regarding information points (where at least one of the frequently occurring words, related words, and explanatory words matches the If multiple words are entered as exclusion words, they will be excluded from the information provided (search target). It is also possible to input exclusion words. It is also possible to search by entering only the search word without typing.
[0062] Next, in S24, the CPU 31 inputs the search word and the exclusion word on the search word input screen 51. The search start button 54 is operated with a word (exclusion words may not be entered) entered. Upon receiving this signal, a request signal for information on the information providing location is sent to the information providing server 3. The request signal includes a terminal ID for identifying the communication terminal 5 that is the sender, a search word, and Includes excluded words.
[0063] Thereafter, in S25, the CPU 31 provides information in response to the request signal sent in S24. The information received in step S25 is sent to the user. Therefore, information points related to the entered search words (however, if exclusion words are entered, In this case, information points related to the excluded words are excluded. The information includes the location name, location coordinates, description, and posted text about the location. The information about the information providing point received in step S25 includes the information providing point Priority for providing information is set for each point, and the total points are calculated in S36 (described later). Higher priority is given to information points with higher values.
[0064] Next, in S26, the CPU 31 uses the information regarding the information provision point received in S25. The information is displayed on the display 36. If there is information about multiple information providing points, Regarding the above, information about the information providing point with the highest priority is output first. The information providing points may be displayed in order of priority, or the information with priority above a threshold may be displayed in order of priority. Only the provided points may be displayed.
[0065] Here, FIG. 11 shows the display on which the information about the information providing point is displayed in S26. 11 shows an example of a play 36. As shown in FIG. 11, the display 36 displays a search result screen 55 The search result screen 55 shows the name and description of the information providing point in order of priority, for example. The sentences are displayed in a list. Then, the user selects any of the information displayed on the search result screen 55. When you select a location, more detailed information about the selected location is provided. The screen transitions to the presentation screen 60.
[0066] The information provision screen 60 displays, for example, the name of the information provision point, a photo of the exterior, the location, business hours, and contact details. Posts posted to that location, along with information about location details such as contact details However, the "posted text" does not include the entire text, but rather the evaluation and status of the information providing point. It is also possible to display only keywords that indicate the situation. For example, the post emotion can be either positive or negative. The poster's emotions are judged based on two factors: positive and negative emotions. Whether you are leaning towards either of the two extremes or neutral In addition, the information provision screen 60 has a destination registration button 61. The information is also displayed, and the user can register the information provided by pressing the registration button 61. It is also possible to register the information providing point as a destination.
[0067] It should be noted that the search result screen 55 shown in FIG. 10 and the information provision screen 60 shown in FIG. 11 are merely examples. Any display mode may be used as long as it is possible to provide information about the information providing point. For example, only the names of the information providing points may be displayed in a list in order of priority. When the user visually recognizes the search result screen 55 shown in FIG. 10 or the information provision screen 60 shown in FIG. By doing so, the user can easily search for information related to the search word entered on the search word entry screen 51. You can get information about the location provided. If the user does not like the search word, the user can input the search word into the search word input section 52 or the exclusion word into the exclusion word input section 53. You can change the words in the search, add words, or remove words and search again. It is Noh.
[0068] Next, the information providing processing program executed by the CPU 21 of the information providing server 3 will be explained. The following steps S31 to S40 are executed after receiving the corresponding information from the communication terminal 5. Therefore, the order of execution of each step is not necessarily determined by the step number. The procedures may not necessarily be carried out in the order listed.
[0069] First, in S31, the CPU 21 receives an information request signal transmitted from the communication terminal 5. The request signal includes a terminal ID for identifying the communication terminal 5 that is the sender, a search word, and an exclusion Contains words (but only if excluded words are entered).
[0070] Thereafter, in S32, the CPU 21 executes the search command included in the request signal received in S31. Weighting is performed on words. Specifically, when multiple words are entered as search words, The later a word is input by the user, the heavier the weight is set. If the search word is entered earlier than half (decimal points rounded down), Set the weight to "1 (normal)" and weight words entered later than 1 / 2. For example, as shown in Figure 12, set the search word "Ra" to "2 (heavy)". The five words "men," "gyoza," "delicious," "fun," and "energetic" were entered in order. In this case, the weighting of the first half, "Ramen" and "Gyoza", is "1", and the weighting of the second half, "Oishii" is "1". The weighting for "good", "fun" and "lively" is "2". The more search words are used, the more information that matches the search words will be searched for in the subsequent information search process. and output.
[0071] In the conventional general information search concept, the earlier a word is entered, the more likely it is to be found. The importance of the word is high, and the later the word is entered, the lower its importance is. However, in this embodiment, the later a word is input by the user, the more weight it is assigned. Here, when a user inputs multiple search words, the input The words that are later in the order are not words that specify specific product names or service names for users. It is assumed that the word will have the nuance of "it would be nice if it existed." By giving more weight to words with a nuance of "I wish there was such a thing," This makes it possible to find information that would not be found by conventional search engines. In particular, we use words extracted from posts posted on SNS etc. to search for information providing locations. These keywords are set as follows (S2 to S7), and among these words are selected impressions, experiences, etc. It also includes symbolic words (such as "fun" and "beautiful"). Even if you use a word with a nuance as a search term, you can search for information that matches the search term. do.
[0072] The following steps S33 to S36 are processed at information providing points nationwide. Then, after executing the processes of S33 to S36 for all the information providing points, Moving on to S37.
[0073] First, in S33, the CPU 31 checks the search unit included in the request signal received in S31. The search engine compares the word with the frequently occurring words set for the information providing point to be processed, and finds the word that matches the frequently occurring word. Specify the search word. Note that "match" means that the search word and the frequently occurring word or related word are completely the same. Not just in cases of matching, but also in cases of only partial matching or similarity above a certain percentage Also includes (the same below). And the number and weight of search words and frequently occurring words that are judged to match Points are calculated based on the matching rate. The frequency of appearance in posts about the location is equal to or greater than the threshold. The search setting unit is set in the search processing program (Fig. 7) and linked to each information providing point. These are stored in the word DB15 (Fig. 5).
[0074] For example, as shown in Figure 13, the search words are "ramen," "gyoza," "delicious," Five words, "fun" and "energetic," are input, and the information provision point to be processed is a frequently occurring word. For example, if "Ramen", "Delicious", "Beautiful", and "Gentle" are set as In the example shown in Figure 13, the most frequently used search words are "ramen" and "delicious." Then, for the search words that are determined to match the frequently occurring words, By multiplying the total weighting set in S32 by the weighting of frequent words, "3", The points are calculated. As shown in Figure 12, the weighting of the search word "ramen" is " The weighting of "1" and "delicious" is "2", so the points are (1 + 2) x 3 = 9 In particular, in this embodiment, frequently occurring words are weighted more heavily than related words and explanatory words, which will be described later. do.
[0075] Next, in S34, the CPU 31 executes the search unit included in the request signal received in S31. The search result is compared with the related words set for the information point to be processed, and a search result that matches the related words is Then, the number and weight of the search words and related words that are determined to match are calculated. The points indicating the match rate are calculated based on the results. These are words related (having a co-occurrence relationship) to the frequently occurring words set at the points, and are used in the word setting process described above. The search setting word D is set in the management program (Fig. 7) and linked to each information providing point. It is stored in B15 (Figure 5).
[0076] For example, as shown in Figure 13, the search words are "ramen," "gyoza," "delicious," The five words "fun" and "energetic" are entered, and the relevant words for the information provision point to be processed are For example, if "spicy", "soup", "gyoza", and "fun" are set as the menu items, In the example shown in Figure 13, among the search words, "gyoza" and "fun" match with related words. Then, the search words that are determined to match the related words are used to The points are calculated by multiplying the total weighting of the related words by "2". As shown in Figure 12, the weighting of the search word "gyoza" is "1" and the weighting of "fun" is "2". Since the weighting of "Yes" is "2", the points are (1+2) x 2 = 6.
[0077] Next, in S35, the CPU 31 checks whether the search result included in the request signal received in S31 is correct. The word is compared with the explanatory words set for the information point to be processed, and if it matches the explanatory word, Identify search words. Then, count and weight of search words and explanatory words that are determined to match. The points indicating the matching rate are calculated based on the "descriptive words." These are words contained in the explanatory text that describes the location, and are used in the word setting processing program (Figure 7) described above. The search terms are set in the search keyword database 15 and linked to each information point. do.
[0078] For example, as shown in Figure 13, the search words are "ramen," "gyoza," "delicious," The five words "fun" and "energetic" are entered, and the information provision point to be processed is assigned the explanatory word For example, if "Ramen", "Gyoza", "Lunch Set", etc. are set as In the example shown in Figure 13, among the search words, "ramen" and "gyoza" are explanatory words. Then, the search word determined to match the explanatory word is processed in step S32. The points are calculated by multiplying the sum of the weights set in (1) by the weight of the explanatory word. As shown in Figure 12, the weight of the search word "ramen" is "1", Since the weighting of "Gyoza" is "1", the points are (1+1) x 1 = 2.
[0079] Thereafter, in S36, the CPU 31 calculates the number of points calculated in S33 to S35. The total is calculated as a point indicating the matching rate between the search word and the information providing location being processed. The higher the total points calculated in S36, the more likely it is that the search word and the information providing location to be processed will match. It indicates that the match rate is high, that is, it is an information providing point where users request information. The more frequently used words, related words, and explanatory words that match the search word, the higher the total points. Also, as mentioned above, when multiple words are entered as search words, Therefore, the later a word is entered by the user, the heavier the weight is set. Points are added for information points that match search terms entered later by the user. In addition, the weighting of frequently occurring words, related words, and explanatory words is set individually. This means that more relevant information points are found than those whose description words match the search words entered by the user. The points for information points that match the words will be higher, and the points for information points that match the frequently used words will be higher. The point of delivery will earn more points.
[0080] Next, in S37, the CPU 21 performs the process of determining the location of the information providing points nationwide in S36. The calculated total points are compared, and information provision locations that are subject to information provision to the user are selected. For example, a predetermined number of information providing points may be selected in descending order of total points. Alternatively, an information providing point having a total point of a predetermined number or more may be selected. Priority is also set in descending order of total points.
[0081] Next, in S38, the CPU 21 checks the exclusion unit included in the request signal received in S31. The word, the frequently occurring words, related words, and explanations set for the information provision point selected in S37 The information includes frequently occurring words, related words, and explanatory words that match the excluded words. If there is a location where information is provided, that location is excluded from the selection targets in step S37. The location will be excluded from the search targets for information provision points by the information provision server 3. For information provision locations that the user wants to exclude from the information provision target, Even if the words are included in the information provided, they can be excluded from the information provided. If no input has been made by the above, the process of S38 is omitted.
[0082] Here, in the past, users have been able to post messages on networks such as SNS. The comments can be viewed using information terminals such as smartphones and PCs, but negative comments cannot be viewed. Therefore, negative conditions (e.g. Even if there are some issues (such as congestion or poor transportation), viewing the posts will not affect the experience. It is difficult to find an information providing point that avoids these negative conditions. Then, we use the words extracted from the posts, including negative posts, to search for information sources. Since it is set as a keyword (S2 to S7), for example, it can be avoided as an exclusion word by the user. By entering the negative conditions you want to avoid, you can avoid the negative conditions you want to avoid. It is possible to search for information providing locations.
[0083] Thereafter, in S39, the CPU 21 retrieves the information selected in S37 from the distribution information DB 7. Information about the information providing point (excluding the information providing point excluded in S38) is extracted. As mentioned above, the information DB7 contains the information points that are the subject of information provision across the country. It includes the location name, location coordinates, description, and any posts posted about the information providing location. Various information such as the date, time, and date is stored (Figure 4).
[0084] Next, in S40, the CPU 21 determines whether the communication The information about the information provision point extracted in S39 is transmitted to the terminal 5. The priority set in 7 will also be sent. Instead of the priority, each piece of information provided The total number of points of the location may be transmitted. After that, the communication terminal 5 that received the information As shown above, information on information points is given in order of priority (i.e., information points with higher total points are given). In this embodiment, the information about the location is outputted (priority is given to the information about the location) (FIG. 11). As a result, frequently occurring words are weighted more heavily than related words, so the search term is Information points that match the search term are given higher priority than information points that match related words. It will be output first.
[0085] As described above in detail, the information search system 1, the information providing server 3, and the The communication terminal 5 acquires the posted text posted on the network (S2) and searches for the text. By analyzing the posts posted for each piece of information, The frequency of the extracted words appearing in the posts and the number of extracted words are also calculated. The co-occurrence relationships between words are identified (S3-S5), and words that appear more frequently in posts than a threshold are identified. The words are set as frequent words, and related words that are related to the frequent words are set as co-occurrence related words. Then, when performing an information search, the user inputs the Search for information that contains frequently occurring words or related words that match the search word entered. In addition, information retrieval is performed by weighting frequently occurring words more heavily than related words (S32 to S39). By extracting words contained in posts posted on SNS etc. and using them as keywords, Information based on abstract words such as impressions and experiences using words extracted from manuscripts as keywords In addition to the frequency of occurrence, the words extracted from the posts are also searched for. The co-occurrence relationship between words, that is, the correlation between words, is used to weight each word. By performing information searches, it becomes possible to perform highly accurate information searches. The information to be searched is a location, and the posted text on the network is searched for by the relevant The location linked to the posted text is acquired for each specific location (S2), and the location linked to each location is acquired (S3). We analyzed the words contained in the posted texts (S3-S5) and identified the most frequently used words for each location. Related words are set (S7), and frequently occurring words or words that match the search words entered by the user are selected. Search for locations with related words set, and if multiple locations match, search word Points where frequent words match the search word are given priority over points where related words match the search word. Therefore, when searching for information about a location, it is possible to easily find information posted on social media, etc. By extracting words contained in the posted text and using them as keywords, It becomes possible to search for locations using abstract words such as impressions and experiences as keywords. In addition, multiple words can be input as search words, and multiple words can be input as search words. When a word is input, the later the word is input by the user, the heavier the weight is. Since the information search is performed (S32 to S39), information that would not be found by a normal search engine is retrieved. It becomes possible to search for. In addition to the search words, the user can input exclusion words. Information that contains frequently occurring words or related words that match the search criteria will be excluded from the search results (S 38) So, for example, you can enter negative terms that you want to avoid as exclusion words. This makes it possible to search for information that avoids negative conditions that users want to avoid.
[0086] The present invention is not limited to the above-described embodiment, and the scope of the present invention is not limited to the above-described embodiment. Of course, various improvements and modifications are possible within the scope of the present invention. For example, in this embodiment, frequently occurring words generated from posts posted on the network are In addition to related words, words contained in the description of the information point are used as keywords. The explanatory words are not used as keywords, but are used as frequently occurring words and related words. Alternatively, the name of the information providing point or Words contained in the address may also be added as keywords.
[0087] In this embodiment, after searching for information provision points that match the search word, the exclusion word However, if the information providing point matches the excluded word first, After excluding the search terms, a search is performed for the information points that match the search terms. That's fine.
[0088] In this embodiment, the information search is performed by searching for information on information providing points. The information that is the subject of postings on the network is If the information can be searched based on the words entered by the user, it is related to the information providing location. It can also be applied to searches for information other than the information currently being searched. For example, it is possible to search for event information. .
[0089] In this embodiment, information retrieval is performed by weighting frequently occurring words more heavily than related words. However, the weighting of frequently occurring words and related words can be the same, or vice versa. It is also possible to perform information retrieval by weighting consecutive words more heavily than frequently occurring words.
[0090] In this embodiment, an example in which the communication terminal 5 is applied to a smartphone has been described. If the terminal has a function to output information about the information providing point, it can be used for other types of communication terminals. It can also be applied to mobile phones, tablet terminals, personal computers, etc. It can be applied to car navigation systems and other on-board devices. When applied to devices other than mobile devices, the user may move in situations other than by car, for example, on foot. It can be implemented even in moving situations.
[0091] In this embodiment, the information providing server 3 executes the word setting processing program (FIG. 7). However, part of the processing may be executed by the communication terminal 5. [Explanation of symbols]
[0092] 1...information retrieval system, 2...information provision center, 3...information provision server, 4...user, 5... Communication terminal, 6... communication network, 7... distribution information DB, 8... SNS server, 11... server control unit, 13...posted text information DB, 15...search setting word DB, 36...display, 41 ...communication terminal control unit, 45...co-occurrence network, 52...search word input unit, 53...exclusion word input unit Power section, 55...search results screen, 60...information provision screen
Claims
1. an information search means for performing an information search based on a search word input by a user; a posted message acquisition means for acquiring a posted message posted on the network; For each piece of information to be searched by the information search means, By analyzing the posted text, words contained in the posted text are extracted and the extracted words are The frequency with which the extracted words appear in the posted text and the co-occurrence relationships between the extracted words are identified. A manuscript analysis tool, Based on the analysis result of the posted text analysis means, the information is searched by the information search means. For each piece of information, words that appear in the posted text with a frequency equal to or greater than a threshold are set as frequently occurring words. and setting related words that are related to the frequently occurring words based on the co-occurrence relationships. setting means, The information search means is configured to search for frequently occurring words or related words that match the search word. The set information is searched for, and the frequently occurring words are weighted more heavily than the related words. An information retrieval system that performs information retrieval using the above method.
2. The information to be searched by the information search means is a location, The posted message acquisition means acquires a posted message posted on a network by Obtained at each location, The posted text analysis means analyzes words included in the posted text associated with each location. Conduct an analysis on the subject, the word setting means sets the frequently occurring words and the related words for each location, The information search means is configured to search for frequently occurring words or related words that match the search word. Search for the specified location, and if multiple locations match, search for the search word. The points where the frequently occurring words match are determined based on the points where the related words match for the search word.
2. The information retrieval system according to claim 1, wherein the information retrieval system also outputs the information retrieval result preferentially.
3. A plurality of words can be input as the search word, When a plurality of words are input as the search words, the information search means The later a word is input, the heavier the weight is given to the word to perform information retrieval. Item 3. An information retrieval system according to item 2.
4. A user can input exclusion words separately from the search words, The information search means is configured to search for the frequently occurring words or the related words that match the excluded words.
3. The information search method according to claim 1, wherein the set information is excluded from the search target. system.
Citation Information
Patent Citations
Information search system
JP2019149145A