University information extraction system, university information extraction server, and university information extraction program

The university information extraction system helps applicants find suitable universities by analyzing personal information and university profiles, addressing the complexity of selecting a university based on personal preferences.

JP2026016268APending Publication Date: 2026-02-03新井 伶菜
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024130232
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

The challenge of selecting a suitable university based on personal preferences and future goals is complicated by the vast number of universities worldwide, making it difficult for applicants to determine a good fit.

Method used

A university information extraction system and server that accepts personal information from applicants, calculates similarity with university profile information using morphological analysis and cosine similarity, and transmits suitable university information based on a threshold.

Benefits of technology

Facilitates the extraction of candidate universities that align with applicants' personal characteristics, enhancing the relevance of university selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016268000001_ABST
    Figure 2026016268000001_ABST
Patent Text Reader

Abstract

To provide a university information extraction system, a university information extraction server and a university information extraction program for extracting the candidate of a university considered to be suitable for a university school advance applicant.SOLUTION: A university information extraction system 10 for extracting university information suitable for an applicant for going to a university school includes a university information extraction terminal 11 for receiving an input of personal information of the applicant for going to a university school, and a university information extraction server 12 for extracting university information suitable for the applicant for going to a university school based on the personal information input from the university information extraction terminal and a university information database for storing university profile information about each university. The university information extraction server calculates similarity between the personal information and the university profile information, determines whether or not the calculated similarity is equal to or more than a predetermined threshold, transmits the university information of the university profile information determined to be equal to or more than the predetermined threshold to the university information extraction terminal, and the university information extraction terminal outputs the received university information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a university information extraction system, a university information extraction server, and a university information extraction program that extract candidate universities that suit a student. [Background technology]

[0002] Traditionally, when people hoping to go to university (for example, high school students, also referred to as university hopefuls) chose a university, they often did so based on the university's deviation score or the name of the university. In other words, university hopefuls chose a university because it was one they could go to with their academic ability (deviation score), or because it was a well-known university. Another method being considered for university hopefuls to go to university is to appeal to the university by appealing to them, attracting their interest, and then entering the university through recommendation admission or the like (see Patent Document 1).

[0003] On the other hand, in recent years, an increasing number of people who wish to go to university are choosing a university based on their personality and what they want to do in the future, rather than choosing a university based on the above-mentioned criteria. This is because they believe that studying at a university that suits them is much more important in planning their future than studying at a university they happen to enter without any particular purpose.

[0004] Those who wish to enter a university with this kind of thinking read the university's website or brochure to check the type of person the university is looking for in its students and what they can learn at the university, and then decide whether or not they are suited to studying at that university.In recent years, however, with access to a wide range of information via the Internet, information about universities can be obtained from sources other than websites and brochures. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-50928 Summary of the Invention [Problem to be solved by the invention]

[0006] However, there are a huge number of universities around the world, and it is difficult to research information about each university, determine whether it is a good fit for you, and then select a university.

[0007] The present invention aims to solve the above-mentioned conventional problems by providing a university information extraction system, a university information extraction server, and a university information extraction program that can extract candidate universities that are likely to suit those who wish to go to university. [Means for solving the problem]

[0008] In order to solve the above problem, the invention described in claim 1 provides a university information extraction system that extracts university information suitable for university applicants who wish to go to university, and includes a university information extraction terminal that accepts input of personal information of the university applicant from the university applicant, and a university information extraction server that extracts university information suitable for the university applicant based on the personal information input from the university information extraction terminal and a university information database that stores university profile information for each university, wherein the university information extraction server calculates a similarity between the personal information and the university profile information, determines whether the calculated similarity is above a predetermined threshold, and transmits university information of the university profile information that is determined to be above the predetermined threshold to the university information extraction terminal, and the university information extraction terminal outputs the received university information.

[0009] The invention of claim 2 relates to a university information extraction system for extracting university information suitable for university applicants, the university information extraction server comprising: a university information extraction terminal that accepts input of personal information of university applicants by the applicants who wish to go to university; and a university information extraction server that extracts university information suitable for the applicants based on the personal information input from the university information extraction terminal and a university information database that stores university profile information for each university. The university information extraction server of the university information extraction system for extracting university information suitable for the applicants is characterized by having: a storage unit that stores the personal information of the applicants who wish to go to university input from the university information extraction terminal and the university profile information for each university; a processing unit that calculates the similarity between the personal information and the university profile information; a judgment unit that judges whether the calculated similarity is equal to or greater than a predetermined threshold; and a transmission unit that transmits university information of the university profile information that is judged to be equal to or greater than the predetermined threshold to the university information extraction terminal.

[0010] In the invention described in claim 3, it is a preferred embodiment that the personal information includes at least one of the following: self-introduction information of the university applicant, academic ability information, extracurricular activity information, work experience information, information on areas of expertise, future dreams, information on awards received or selected for, and information on scholarship aspirations.

[0011] In the invention described in claim 4, it is a preferred embodiment that the university profile information includes at least one of information on the type of person the university is looking for, information on the university's strengths, information on academic ability requirements for admission, information on university scholarships, and employment information.

[0012] In the invention described in claim 5, it is a preferred embodiment that when the processing unit calculates the similarity between the personal information and the university profile information, it performs morphological analysis on the text of the personal information and the text of the university profile information, vectorizes the morphologically analyzed text, and calculates the similarity between the personal information and the university profile information as a cosine similarity based on the similarity between the vectors.

[0013] In the invention described in claim 6, it is a preferred embodiment that when the processing unit vectorizes the morphologically analyzed text, it weights the text numerically according to the words broken down by the morphological analysis and vectorizes the text.

[0014] In the invention described in claim 7, when the processing unit calculates the cosine similarity, if the number of columns of particles or auxiliary verbs in the words decomposed by the morphological analysis is greater than a predetermined number, it is a preferred embodiment that the calculated value of the cosine similarity is multiplied by a numerical value corresponding to the number of columns of particles or auxiliary verbs.

[0015] The invention described in claim 8 is a university information extraction program that extracts university information suitable for a university applicant who wishes to go to university, and is characterized by having the steps of calculating a similarity based on personal information of the university applicant and university profile information for each university, determining whether the calculated similarity is above a predetermined threshold, and transmitting university information from the university profile information determined to be above the predetermined threshold to a terminal of the university applicant. [Effects of the Invention]

[0016] The invention of claim 1 provides a university information extraction system that extracts university information suitable for university applicants who wish to go to university, and includes a university information extraction terminal that accepts input of personal information of the university applicant from the university applicant, and a university information extraction server that extracts university information suitable for the university applicant based on the personal information input from the university information extraction terminal and a university information database that stores university profile information for each university. The university information extraction server calculates the similarity between the personal information and the university profile information, determines whether the calculated similarity is above a predetermined threshold, and transmits university information of the university profile information that is determined to be above the predetermined threshold to the university information extraction terminal. The university information extraction terminal outputs the received university information, thereby extracting candidate universities that are likely to be suitable for the university applicant.

[0017] In the invention of claim 2, the university information extraction server of the university information extraction system for extracting university information suitable for university applicants comprises a university information extraction terminal that accepts input of personal information of university applicants by the applicants who wish to go to university, and a university information extraction server that extracts university information suitable for the applicants based on the personal information input from the university information extraction terminal and a university information database that stores university profile information for each university. The university information extraction server of the university information extraction system for extracting university information suitable for the applicants includes a storage unit that stores the personal information of the applicants input from the university information extraction terminal and the university profile information for each university, a processing unit that calculates the similarity between the personal information and the university profile information, a judgment unit that judges whether the calculated similarity is equal to or greater than a predetermined threshold, and a transmission unit that transmits university information from the university profile information that is judged to be equal to or greater than the predetermined threshold to the university information extraction terminal, thereby making it possible to extract candidate universities that are likely to be suitable for the applicants.

[0018] In the invention described in claim 3, the personal information includes at least one of the following: self-introduction information of the university applicant, academic ability information, extracurricular activity information, work experience information, information on areas of expertise, future dreams, information on awards received or selected for, and information on scholarship preferences, thereby making it possible to extract university information that is more suited to the university applicant.

[0019] In the invention described in claim 4, the university profile information includes at least one of the following: information on the type of person the university is looking for, information on the university's strengths, academic ability information for admission, information on university scholarships, and employment information, thereby making it possible to attract the students that the university is looking for.

[0020] In the invention described in claim 5, when the processing unit calculates the similarity between the personal information and the university profile information, it can measure the similarity between the texts by morphologically analyzing the text of the personal information and the text of the university profile information, vectorizing the morphologically analyzed text, and calculating the similarity between the personal information and the university profile information as a cosine similarity based on the similarity between the vectors.

[0021] In the invention described in claim 6, when the processing unit vectorizes the morphologically analyzed text, it weights the text numerically according to the words broken down by the morphological analysis, thereby removing noise in the calculation of the similarity between texts.

[0022] In the invention described in claim 7, when the processing unit calculates the cosine similarity, if the number of columns of particles or auxiliary verbs among the words decomposed by the morphological analysis is greater than a predetermined number, the calculated value of the cosine similarity is multiplied by a numerical value corresponding to the number of columns of particles or auxiliary verbs, thereby making it possible to remove the influence of items that are not relevant to the calculation of the similarity between texts.

[0023] In the invention described in claim 8, a university information extraction program extracts university information suitable for a university applicant who wishes to go to university, and includes the steps of calculating a similarity based on the personal information of the university applicant and university profile information for each university, determining whether the calculated similarity is above a predetermined threshold, and transmitting university information from the university profile information determined to be above the predetermined threshold to the terminal of the university applicant, thereby making it possible to extract candidate universities that are likely to be suitable for the university applicant. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a diagram illustrating an example of a configuration of a university information extraction system according to the present invention. [Figure 2]FIG. 2 is a diagram showing an example of personal information in the university information extraction system according to the present invention. [Figure 3] FIG. 2 is a diagram showing an example of the configuration of a university information extraction terminal according to the present invention. [Figure 4] 10 is a flowchart showing an example of a processing flow for inputting personal information into the university information extraction terminal according to the present invention. [Figure 5] FIG. 2 is a diagram illustrating an example of the configuration of a university information extraction server according to the present invention. [Figure 6] FIG. 2 is a diagram showing an example of university profile information in the university information extraction server according to the present invention. [Figure 7] FIG. 10 is a diagram showing an example of a result of dividing text in the university information extraction server according to the present invention. [Figure 8] FIG. 10 is a diagram showing an example of vectorization of texts A, B, and C using BoW in the university information extraction server according to the present invention. [Figure 9] 10 is a flowchart showing an example of a processing flow of a university information extraction server according to the present invention. [Figure 10] FIG. 2 is a diagram showing an example of the configuration of a university terminal according to the present invention. [Figure 11] 1 is a flowchart showing an example of a processing flow for extracting university information in the university information extraction system according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] Hereinafter, a university information extraction system, a university information extraction server, and a university information extraction program according to the present invention will be described with reference to the drawings.

[0026] First, the configuration of one embodiment will be described. Figures 1 to 11 are diagrams for explaining a university information extraction system according to one embodiment of the present invention.

[0027] As shown in Figure 1, the university information extraction system 10 of this embodiment extracts information about universities that suit university applicants who wish to go to university, and is composed of university information extraction terminals 11 (11a, 11b, 11c), a university information extraction server 12, and university terminals 13 (13a, 13b, 13c). The university information extraction system 10 can provide university applicants with information about the universities they desire by having each terminal and server operate mutually.

[0028] The university information extraction terminal 11, the university information extraction server 12, and the university terminal 13 are connected via an internet line 14, enabling them to send and receive information (data) among themselves. Furthermore, the university information extraction terminal 11 and the university terminal 13 may send and receive information not only via wired communication but also via wireless communication such as Wi-Fi.

[0029] 1, there are three university information extraction terminals 11 (university information extraction terminals 11a, 11b, 11c), one university information extraction server 12, and three university terminals 13 (university terminals 13a, 13b, 13c), but this is just an example and the number of terminals is not limited to this. Below, each terminal and server will be explained in detail using the diagram.

[0030] First, the university information extraction terminal 11 will be described with reference to Fig. 2. As shown in Fig. 2, the university information extraction terminal 11 is a terminal through which a person who wishes to go to university inputs information, and may be a PC (Personal Computer) or a mobile terminal (for example, a smartphone). The information input here is information about the person who wishes to go to university and is called personal information. To support the input of personal information, an input form is displayed on the screen of the university information extraction terminal 11. Fig. 2 shows an enlarged screen 20 of that screen.

[0031] 2, examples of personal information include name, age, nationality, desired faculty, academic ability, extracurricular activities, work experience, field of expertise, future dreams, awards received, and whether or not a scholarship is desired, and this information may be text data, image data, video data, or the like. In the case of text data, the information is input using a keyboard (not shown) or a touch panel (not shown) of the university information extraction terminal 11. This text data may be in Japanese or a foreign language such as English. In the case of image data or video data, the information is input by transmitting previously prepared image data or video data to the university information extraction terminal 11.

[0032] Here, an example of personal information will be specifically described. The personal information here is just an example, and other items may be included, or the items described here may not be included. Note that it is not necessary to enter all of this information, and only predetermined required information may be entered.

[0033] The name is the name of the person hoping to go to university, for example "Sato Taro." The age is the age of the person hoping to go to university at the time of entering the information, for example "18 years old." The nationality is the nationality of the person hoping to go to university, for example "Japanese." If the person has multiple nationalities, information about each nationality may be entered. Information such as the native language (for example, Japanese) and second foreign language (English) may also be entered. The desired faculty is information about the faculty that the person hoping to go to university wants to study, for example "Faculty of Engineering." It may also be possible to enter multiple faculties. This is because there may be multiple faculties with similar learning content or faculties that interest the person.

[0034] Academic ability is a score or evaluation on a specific test, such as "** test score: 44 points." Examples of specific tests include, but are not limited to, the International Baccalaureate (IB) test, the Advanced Placement (AP) test, the Scholastic Assessment Test (SAT), and the TOEFL (TOEFL) test.

[0035] Extracurricular activities are activities both inside and outside of school, such as "activities led by the student council president," "participating in Model United Nations," "volunteer cleaning of local parks," etc. However, the content of extracurricular activities is not limited to these, and university applicants can write down any activities they have undertaken in the past.

[0036] Work experience refers to the types of jobs you have had up until now, including part-time jobs and internships. For example, "part-time work at a convenience store" or "internship planning and developing new products at XX Co., Ltd."

[0037] The specialty is the content of the field in which the university applicant is good at, such as "mathematics" or "programming." The specialty is not limited to things related to learning, but may be something related to sports or one's own strengths. For example, it may be "soccer," "baseball," "good communication skills," or "the ability to complete anything without giving up."

[0038] Future dreams are what those hoping to go to university want to become or do in the future, such as "researching and developing AI (artificial intelligence)" or "starting a financial consulting company."

[0039] The award history is the details of the awards that the university applicant has received up to now, such as "first place in the World Mathematical Olympiad," "winning the All-Japan High School Baseball Championship," or "third place in the World Classical Volleyball Competition." Note that the competitions for which the awards were received may not only be world competitions, but also Japanese competitions, and may be regardless of the scale or content of the competition. On the other hand, the competitions for which the awards were received may be limited, and only the award history for those competitions may be entered.

[0040] The option of whether or not a scholarship is desired indicates whether or not the person intending to go to university wishes to receive a scholarship from the university when going to university, and examples include "scholarship required," "scholarship not required," "scholarship required if possible," "partial scholarship required," etc. If a scholarship is required, the amount required may be entered.

[0041] In other words, personal information includes at least one of the following: self-introduction information of a university applicant (such as name, age, and nationality), academic ability information (such as IB score), extracurricular activity information (such as volunteering), work experience information (such as information on internships), information on areas of expertise (such as JAVA skills), future dreams (such as desired career), information on awards received or selected for (such as winning the Mathematical Olympiad), and information on scholarship hopes (such as whether or not a scholarship is needed).

[0042] Next, an example of the configuration of the university information extraction terminal 11 will be described with reference to Fig. 3. As shown in Fig. 3, the university information extraction terminal 11 is composed of an input unit 30, a processing unit 31, a storage unit 32, a transmission unit 33, a reception unit 34, and an output unit 35.

[0043] The input unit 30, processing unit 31, transmission unit 33, reception unit 34, and output unit 35 correspond to a central processing unit (CPU) on which various software programs are executed, and the storage unit 32 corresponds to memory, which stores various software programs and is used as a work area when the software programs are executed by the CPU. Note that the configuration of the university information extraction terminal 11 is not limited to this, and may include other components. For example, it may include components for acquiring image data and video data.

[0044] The input unit 30 accepts input of personal information from the university applicant, and accepts data entered, for example, via a keyboard (not shown) connected to the university information extraction terminal 11 or a touch panel of the university information extraction terminal 11. In this case, the input unit 30 can accept not only text data, but also image data and video data. For example, the data may include image data of the university applicant himself, image data of certificates of awards he has received, video data introducing the university applicant himself (video data for self-introduction), etc.

[0045] The processing unit 31 performs various processes of the university information extraction terminal 11, for example, storing personal information data received by the input unit 30 in the storage unit 32 described later, and storing data received by the receiving unit 34 described later in the storage unit 32.

[0046] The storage unit 32 stores personal information data and university information received from the university information extraction server 12 .

[0047] The transmitting unit 33 transmits the personal information data to the university information extraction server 12 .

[0048] The receiving unit 34 receives university information from the university information extraction server 12 .

[0049] The output unit 35 displays the personal information input screen and the received university information on the display of the university information extraction terminal 11 .

[0050] Next, an example of the flow of the personal information input process of the university information extraction terminal 11 will be described with reference to FIG. 4. As shown in FIG. 4, the input unit 30 accepts personal information from a university applicant via a keyboard or the like (step S401). The processing unit 31 stores the accepted personal information data in the storage unit 32 (step S402). At this time, the processing unit 31 may assign an ID to each piece of personal information and store the ID in association with the name of the university applicant, etc. In this way, when correcting personal information, by entering an ID, the personal information data associated with that ID can be identified and displayed on a display or the like.

[0051] The transmitting unit 33 transmits the personal information data stored in the storage unit 32 to the university information extraction server 12 and requests the university information extraction server 12 to extract university information (step S403). The receiving unit 34 receives the university information from the university information extraction server 12 in response to the extraction request (step S404). The output unit 35 displays the received university information on a display (step S405). The processing unit 31 may store the received university information in the storage unit 32.

[0052] Next, an example of the configuration of the university information extraction server 12 will be described with reference to Fig. 5. As shown in Fig. 5, the university information extraction server 12 is composed of a receiving unit 50, a processing unit 51, a storage unit 52, a determining unit 53, an extracting unit 54, and a transmitting unit 55.

[0053] The receiving unit 50, processing unit 51, determining unit 53, extracting unit 54, and transmitting unit 55 correspond to a central processing unit (CPU) on which various software is executed, and the storage unit 52 corresponds to memory, which stores various software and is used as a working area when the software is executed by the CPU. Note that the configuration of the university information extraction server 12 is not limited to this, and may include other components.

[0054] The receiving unit 50 receives personal information transmitted from the university information extraction terminal 11, and receives university profile information relating to the university transmitted from each university terminal 13. The specific contents of the university profile information will be described later.

[0055] The processing unit 51 stores the personal information data received by the receiving unit 50 in the storage unit 52, and also stores the university profile information data received by the receiving unit 50 in the storage unit 52. The university profile information may be stored in advance in the storage unit 52, rather than being transmitted from each university terminal 13. This saves the university time and effort, and also allows information on universities that do not transmit university profile information to be extracted, making it possible to provide university information that suits those who wish to enter university from among a large amount of data.

[0056] Furthermore, when storing data in the storage unit 52, the processing unit 51 may modify (translate) the data into a predetermined language and store it. For example, if the predetermined storage languages ​​are Japanese, English, French, German, Chinese, or Spanish, and the received personal information data or university profile information data is only Japanese text data, the Japanese text data may be translated into English, French, German, Chinese, or Spanish text data and stored. This prevents the text similarity from being unable to be calculated when the personal information data is only Japanese and the university profile information data is only English.

[0057] The processing unit 51 also calculates the degree of similarity between the personal information and the university profile information. A specific method for calculating the degree of similarity by the processing unit 51 will be described later.

[0058] The storage unit 52 stores personal information data and university profile information.

[0059] The determination unit 53 determines whether the similarity calculated by the processing unit 51 is equal to or greater than a predetermined threshold value. A specific determination method by the determination unit 53 will be described later.

[0060] The extraction unit 54 extracts university information from the university profile information that is determined by the determination unit 53 to be equal to or greater than the predetermined threshold value.

[0061] The transmitting unit 55 transmits the university information of the university profile information to the university information extraction terminal 11 .

[0062] Before explaining how the processing unit 51 calculates the similarity, an example of university profile information will be described with reference to FIG. 6. The university profile information 60 is information about a university, and may have the content shown in FIG. 6, for example. The university profile information 60 here is only an example, and may include other items, or may not include the items described here. Note that it is not necessary to input all of this information, and only predetermined required information may be input.

[0063] The university profile information 60 includes, for example, the name of the university, the country, the address, the faculty, the type of person the university is looking for (leadership, social contribution, etc.), academic ability (including academic ability information for admission (English proficiency, IB score, SAT score, AP score, etc.) and required subjects), the university's areas of expertise (information on the university's strengths (business, environmental, research, etc.)), whether or not there are scholarships (information on the university's scholarships (including information on tuition fees)), application submission deadline, documents required for the entrance exam, and employment information (company name, etc.).

[0064] This information may be text data, image data, video data, etc. In the case of text data, it may be input using a keyboard (not shown) or a touch panel (not shown) of the university terminal 13. This text data may be in a foreign language such as English, as well as Japanese. This is because the text of the university profile information 60 provided by overseas universities is in a foreign language. In the case of image data or video data, it is input by transmitting pre-prepared image data or video data to the university terminal 13, for example.

[0065] The university name is the name of the university that corresponds to the university profile information 60, such as "XX University" or "YY University."

[0066] The country name is the name of the country in which the university is located, such as "Japan" or "USA."

[0067] The address is the location of the university, for example, "XX City □□ 1234-5."

[0068] The faculty is information about the faculties that exist in the university, such as "Faculty of Law, Faculty of Economics, Faculty of Business Administration, Faculty of Letters, Faculty of Commerce, Faculty of Science, Faculty of Engineering, Faculty of Sports."

[0069] The type of person a university is looking for is a summary of the type of student they want to enroll, such as "someone with leadership skills" or "someone with expertise."

[0070] The academic ability is the academic ability required for university admission, such as "80 points or above on the XX test." It may also include the required language ability, such as native-level English ability or everyday conversational English ability.

[0071] A university's specialty is an academic field that the university can offer with confidence, such as "economics," "environmental studies," "psychology," etc. It can also be a field in which the university places emphasis, such as "developing entrepreneurs" or "developing leaders," in addition to academic fields.

[0072] The availability of scholarships refers to whether the university offers its own scholarships, and if so, what the criteria are for providing them. For example, "There is a scholarship system, but only for the top 5% of students."

[0073] The application deadline is the deadline for submitting an application for university admission, and may be, for example, "must arrive by December 31, 2025" or "postmarked by December 25, 2025."

[0074] The documents required for the examination are documents required when taking the examination, such as a "letter of recommendation from the teacher in charge" or "scores from the XX test."

[0075] Below, we will explain an example of a method for calculating the similarity between personal information and university profile information by the processing unit 51. Here, we will explain an example of a method for calculating the similarity of text (sentences) using natural language processing technology. Natural language refers to languages ​​such as Japanese and English, and natural language processing refers to technology that allows a computer to recognize and process natural language.

[0076] There are three main steps to calculating text similarity: morphological analysis, text vectorization, and text similarity calculation. Each of these steps is explained below.

[0077] Morphological analysis is the process of breaking down text into the smallest units of language that have meaning. In other words, it is the process of breaking down text into words. For example, consider the following three texts. Text A is part of the text information of a person applying to a certain university. Text B is part of the text information of a university profile for one university. Text C is part of the text information of a university profile for another university.

[0078] The first text (Text A) is, "I became student council president, demonstrated leadership, and contributed to the success of school events." The second text (Text B) is, "A person who has leadership skills, takes on new challenges, and contributes to the school." The third text (Text C) is, "A person who is dedicated to one thing and specializes in a specialized field."

[0079] Figure 7 shows an example of the results of segmenting each text. As shown in Figure 7, Text A is segmented into "I," "am," "student," "president," "and," "become," "leadership," "to," "demonstrate," "do," "school," "of," "event," "of," "success," "to," "contribute," and "did." Text B is segmented into "leadership," "to," "have," "anything," "also," "challenge," "do," "school," "to," "contribute," "do," and "person." Text C is segmented into "one," "of," "thing," "to," "dedicate," "expert," "of," "field," "to," "specialize," "did," and "person."

[0080] As described above, the text is broken down into words, and the next step is to vectorize the text. Note that while the text is Japanese here, it can also be in a foreign language such as English. In the case of English, morphological analysis is relatively easy because the text is broken down into words.

[0081] Next, we will explain how to vectorize text. As mentioned above, by dividing text into words and then vectorizing these, it is possible to calculate the similarity between texts. In this case, a vector refers to a string of numbers. There are several existing technologies for vectorization, but here we will explain using a method called Bag-of-Words (BoW). BoW converts text into numerical vectors based on the frequency of word occurrence.

[0082] When texts A, B, and C are vectorized using BoW, the result is as shown in Figure 8. As shown in Figure 8, the words of each text are listed, and a "1" is inserted in the corresponding column for texts that contain the listed words, and a "0" is inserted in the corresponding column for texts that do not contain the listed words. For example, the word "I" in text A is used only in text A, so a "1" is inserted only in the corresponding column for text A, and a "0" is inserted in the columns for other texts. Furthermore, the word "ni" in text A is also used in texts B and C, so a "1" is inserted in the columns for each of texts A, B, and C.

[0083] Next, we will explain how to calculate the similarity of texts. The similarity is calculated using the vectors of each text obtained as described above. The similarity here refers to cosine similarity, but text similarity calculated by other methods other than cosine similarity may also be used.

[0084] Cosine similarity is an index (method) that measures the degree to which two vectors point in the same direction (orientation). The calculated cosine similarity takes a value between -1 and 1, with the value being closer to 1 if the two vectors point in the same direction and closer to -1 if the two vectors point in opposite directions. Here, to calculate the cosine similarity between personal information and university profile information, the cosine similarity between text A and B, and between text A and C, is calculated. Note that the combination of texts with a calculated value closer to 1 is the most similar.

[0085] Specifically, the following formula 1 is used to calculate cosine similarity. The denominator indicates the magnitude (norm) of each vector, and the numerator indicates the dot product of each vector. The number of dimensions (number of columns) of a vector is the number of words that appear in texts A, B, and C. In the example shown in Figure 8, there are 27 columns, so the number of dimensions n is 27.

[0086]

number

[0087] Using Equation 1, the cosine similarity between the vector of text A and the vector of text B, and the vector of text A and the vector of text C are calculated. For example, the cosine similarity between the vector of text A and the vector of text B is calculated to be 0.5, and the cosine similarity between the vector of text A and the vector of text C is calculated to be -0.2.

[0088] In other words, when calculating the similarity between personal information and university profile information, the processing unit 51 of the university information extraction server 12 performs morphological analysis on each of the text of the personal information and the text of the university profile information, vectorizes the morphologically analyzed text, and calculates the similarity between the personal information and the university profile information as cosine similarity based on the similarity between the vectors.

[0089] Note that when vectorizing text, the vectorization may be performed by weighting the values. For example, when the text is broken down into words, only the broken down words that are nouns or adjectives may be quantified (inserting "1"), and those that are particles or auxiliary verbs may be vectorized as "0" instead of "1". Specifically, if the broken down word is a particle such as "wa", a "0" instead of a "1" is inserted in the corresponding field even if the text uses it. This prevents texts from being judged as similar due to the influence of particles or auxiliary verbs that appear in many texts.

[0090] That is, when vectorizing the morphologically analyzed text, the processing unit 51 of the university information extraction server 12 weights the text numerically according to the words resolved by the morphological analysis and vectorizes the text.

[0091] Another weighting method is to multiply the calculated cosine similarity value by a numerical value corresponding to the number of columns of particles or auxiliary verbs when the number of columns (number of dimensions) of particles or auxiliary verbs is greater than a predetermined number. As an example, when the ratio of the number of columns of particles or auxiliary verbs to the total number of columns exceeds half (0.5), the calculated cosine similarity value may be multiplied by a numerical value corresponding to the number of columns of particles or auxiliary verbs.

[0092] For example, if the total number of columns is 30 and the number of columns of particles and auxiliary verbs is 16, the ratio of the number of columns of particles and auxiliary verbs to the total number of columns is 16 / 30, which exceeds 0.5. In this case, the calculated cosine similarity is multiplied by a certain value. The certain value to be multiplied may be calculated as follows. For example, out of the total number of columns of 30, only one column exceeds 0.5, so the certain value to be multiplied may be the ratio of the number of columns that exceeds 0.5 to the total number of columns (1 / 30). Here, the ratio of the number of columns of particles and auxiliary verbs to the total number of columns is set to 0.5, but this is not limited to this and other values ​​may be used. Furthermore, rather than a ratio, it may be determined simply by whether the number of columns of particles and auxiliary verbs exceeds a predetermined value (for example, 10). In this case, for example, if the total number of columns is 30 and the number of columns of particles and auxiliary verbs is 12, the number of columns of particles and auxiliary verbs exceeds the predetermined value of 10, so the value to be multiplied by the cosine similarity may be the proportion (2 / 30) of the total number of columns that exceeds the predetermined value.

[0093] In other words, when calculating the cosine similarity, if the number of columns of particles or auxiliary verbs in a word decomposed by morphological analysis is greater than a predetermined number, the processing unit 51 of the university information extraction server 12 multiplies the calculated cosine similarity value by a numerical value corresponding to the number of columns of particles or auxiliary verbs.

[0094] Next, it is determined whether the texts are similar to each other based on the calculated cosine similarity values. An example of a specific determination method by the determination unit 53 will be described below. The determination unit 53 determines whether the calculated cosine similarity value is equal to or greater than a predetermined value (predetermined threshold value). For example, as described above, consider the case where the cosine similarity between the vector of text A and the vector of text B is calculated to be 0.5, and the cosine similarity between the vector of text A and the vector of text C is calculated to be -0.2.

[0095] If the predetermined threshold is 0.4, the determination unit 53 determines whether each cosine similarity is equal to or greater than the predetermined threshold of 0.4. In this case, the determination unit 53 determines that only the cosine similarity of 0.5 between the vectors of text A and text B is equal to or greater than the predetermined threshold of 0.4. The cosine similarity between the vectors of text A and text C is smaller than the predetermined threshold of 0.4.

[0096] Based on the determination result by the determination unit 53, the extraction unit 54 extracts university information from the university profile information that is determined by the determination unit 53 to be equal to or greater than a predetermined threshold. In this case, the extraction unit 54 extracts university information from the university profile information that includes text B. When extracting, all of the university profile information may be extracted, or only some of the information (for example, university name, country, faculty, etc.) may be extracted. This makes it possible to provide university information that suits those who wish to go to university.

[0097] Next, an example of the processing flow of the university information extraction server will be described with reference to Fig. 9. As shown in Fig. 9, the receiving unit 50 receives personal information transmitted from the university information extraction terminal 11, and also receives university profile information transmitted from each university terminal 13 (step S901). At this time, it is not necessary to receive the personal information and the university profile information simultaneously; there may be a time lag between them. The university profile information may also be stored in advance in the storage unit 52 of the university information extraction server 12.

[0098] The processing unit 51 stores the personal information data received by the receiving unit 50 in the storage unit 52, and stores the university profile information data received by the receiving unit 50 in the storage unit 52 (step S902). At this time, when storing the data in the storage unit 52, the processing unit 51 may modify (translate) the data into a predetermined language and store it. For example, if the predetermined storage languages ​​are Japanese, English, French, German, Chinese, or Spanish, and the received personal information data or university profile information data is only Japanese text data, the Japanese text data may be translated into English, French, German, Chinese, or Spanish text data and stored.

[0099] Then, the processing unit 51 calculates the similarity between the personal information and the university profile information (step S903).

[0100] The determination unit 53 determines whether the similarity calculated by the processing unit 51 is equal to or greater than a predetermined threshold (step S904). If it is determined that the similarity is equal to or greater than the predetermined threshold (YES), the extraction unit 54 extracts university information from the university profile information determined to be equal to or greater than the predetermined threshold (step S905), and the transmission unit 55 transmits the university information from the university profile information to the university information extraction terminal 11 (step S906).

[0101] On the other hand, if it is determined that the number is not equal to or greater than the predetermined threshold (NO), the transmitting unit 55 transmits to the university information extraction terminal 11 a message that there is no target information, that is, that there is no corresponding university profile information (step S907).

[0102] An example of the configuration of the university terminal 13 will now be described with reference to FIG. 10. As shown in FIG. 10, the university terminal 13 is composed of an input unit 100, a processing unit 101, a storage unit 102, a transmission unit 103, a reception unit 104, and an output unit 105. The input unit 100, the processing unit 101, the transmission unit 103, the reception unit 104, and the output unit 105 correspond to a central processing unit (CPU) on which various software programs are executed, and the storage unit 102 corresponds to memory, which stores various software programs and is used as a work area when the CPU executes the software. Note that the configuration of the university terminal 13 is not limited to this, and may include other components. For example, it may include components for acquiring image data and video data.

[0103] The input unit 100 accepts input of university profile information from the university, and accepts data entered, for example, via a keyboard (not shown) connected to the university terminal 13 or a touch panel of the university terminal 13. In this case, the input unit 100 can accept not only text data, but also image data and video data. For example, this may include image data and video data introducing the university.

[0104] The processing unit 101 performs various processes of the university terminal 13, for example, storing university profile information data received by the input unit 100 in the storage unit 102 described later, and storing data received by the receiving unit 104 described later in the storage unit 102.

[0105] The storage unit 102 stores data of university profile information and various information received from the university information extraction server 12 .

[0106] The transmitting unit 103 transmits the data of the university profile information to the university information extraction server 12 .

[0107] The receiving unit 104 receives various information from the university information extraction server 12 .

[0108] The output unit 105 displays the input screen for university profile information and the various received information on the display of the university terminal 13 .

[0109] Next, an example of the processing flow for university information extraction will be described with reference to FIG. 11. Here, the processing flow for university information extraction based on a university information extraction request from a university applicant will be described. As shown in FIG. 11, first, the university information extraction terminal 11 accepts personal information entered by the university applicant via an input form (step S1101). The university information extraction terminal 11 transmits the accepted personal information to the university information extraction server 12 (step S1102). At this time, the university information extraction terminal 11 may store the accepted personal information in its own storage unit 32. This makes it possible, for example, even when the university information extraction terminal 11 is shared, to appropriately disclose the university information sent from the university information extraction server 12 to the university applicant who requested the university information extraction by comparing it with information such as name.

[0110] When the university information extraction server 12 receives the personal information from the university information extraction terminal 11, it calculates the similarity between the personal information and the university profile information (step S1103). Furthermore, the university information extraction server 12 determines whether the calculated similarity is equal to or greater than a predetermined threshold (step S1104). If it is determined that the similarity is equal to or greater than the predetermined threshold (YES), the university information extraction server 12 extracts university information from the university profile information determined to be equal to or greater than the predetermined threshold (step S1105), and transmits the university information from the university profile information to the university information extraction terminal 11 (step S1106).

[0111] On the other hand, if it is determined that the number is not equal to or greater than the predetermined threshold (NO), the university information extraction server 12 transmits to the university information extraction terminal 11 a message that there is no target information, that is, that there is no corresponding university profile information (step S1107).

[0112] The university information extraction terminal 11 displays on its own display or a connected display or the like information that there is no university information or target information received from the university information extraction server 12 (step S1108). At this time, if the university information extraction terminal 11 is used in a shared manner, the university information extraction terminal 11 may check the personal information of the university applicant who requested the university information extraction so that only the applicant can view the information. [Industrial Applicability]

[0113] The university information extraction system, university information extraction server, and university information extraction program according to the present invention can extract candidate universities that are likely to suit those who wish to go to university, and therefore have industrial applicability. [Explanation of symbols]

[0114] 10 University Information Extraction System 11 University information extraction terminal 12 University information extraction server 13 University terminal 14 Internet connection 20 Enlarged Screen 30 Input section 31, 51, 101 Processing section 32, 52, 102 storage compartments 33, 55, 103 Transmitter 34, 50, 104 Receiver 35, 105 Output section 53 Judgment section 54 Extraction part 60 University Profile Information 100 Input section

Claims

1. A university information extraction system that extracts university information suitable for university applicants who wish to go to university, a university information extraction terminal that accepts input of personal information of the university applicant by the university applicant; a university information extraction server that extracts university information suitable for the applicant based on the personal information input from the university information extraction terminal and a university information database that stores university profile information for each university; the university information extraction server calculates a similarity between the personal information and the university profile information, determines whether the calculated similarity is equal to or greater than a predetermined threshold, and transmits university information of the university profile information determined to be equal to or greater than the predetermined threshold to the university information extraction terminal; The university information extraction system is characterized in that the university information extraction terminal outputs the received university information.

2. a university information extraction terminal that accepts input of personal information of university applicants who wish to go to university; a university information extraction server for extracting university information suitable for the applicant based on the personal information input from the university information extraction terminal and a university information database that stores university profile information for each university, the university information extraction server of the university information extraction system extracting university information suitable for the applicant, a storage unit for storing personal information of the university applicants input from the university information extraction terminal and university profile information regarding each university; a processing unit that calculates a similarity between the personal information and the university profile information; a determination unit that determines whether the calculated similarity is equal to or greater than a predetermined threshold; a transmitting unit that transmits university information of the university profile information that is determined to be equal to or greater than the predetermined threshold value to the university information extraction terminal.

3. The university information extraction server described in claim 2, characterized in that the personal information includes at least one of the following: self-introduction information of the university applicant, academic ability information, extracurricular activity information, work experience information, information on areas of expertise, information on future dreams, information on awards won or selected for, and information on scholarship preferences.

4. The university information extraction server described in claim 2 or 3, characterized in that the university profile information includes at least one of the following: information on the type of person the university is looking for, information on the university's strengths, academic ability information for passing, information on university scholarships, and employment information.

5. The university information extraction server described in any one of claims 2 to 4, characterized in that when calculating the similarity between the personal information and the university profile information, the processing unit performs morphological analysis on the text of the personal information and the text of the university profile information, vectorizes the morphologically analyzed text, and calculates the similarity between the personal information and the university profile information as a cosine similarity based on the similarity between the vectors.

6. The university information extraction server of claim 5, characterized in that when the processing unit vectorizes the morphologically analyzed text, it weights the text numerically according to the words decomposed by the morphological analysis and vectorizes the text.

7. The university information extraction server of claim 5 or 6, characterized in that when calculating the cosine similarity, if the number of columns of particles or auxiliary verbs in the words decomposed by the morphological analysis is greater than a predetermined number, the processing unit multiplies the calculated value of the cosine similarity by a number corresponding to the number of columns of particles or auxiliary verbs.

8. A university information extraction program that extracts university information suitable for university applicants who wish to go to university, calculating a similarity based on personal information of the university applicant and university profile information of each university; determining whether the calculated similarity is equal to or greater than a predetermined threshold; transmitting university information of the university profile information determined to be equal to or greater than the predetermined threshold value to a terminal of the university applicant; A university information extraction program characterized by having:

Citation Information

Patent Citations

  • Matching system, server, and storage medium

    JP2003050928A