system
The system automates the extraction and verification of relevant information from website pages using a natural language processing engine, addressing inefficiencies in pricing plan updates by reducing manual labor and external reliance, thereby enhancing operational efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Existing methods for updating pricing plans and service revisions on official websites are time-consuming, costly, and inefficient, particularly when relying on manual checks or external vendors, which can adversely affect business operations.
A system that automatically collects links from an official website, processes HTML data using a natural language processing engine to extract relevant strings, allows user verification and correction, and generates a report in a structured format, reducing the need for manual labor and external outsourcing.
This system streamlines the process of revising pricing plans and service content by enabling efficient, accurate, and rapid responses, minimizing time and cost while improving work efficiency.
Smart Images

Figure 2026062220000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the existing method, when the fee plan or service revision of the official website is carried out, it takes a lot of time and cost to manually check all pages and extract and list the corresponding character strings. Also, the method of entrusting an external vendor is similarly inefficient and not suitable in modern times where quick response is required. As a result, the work efficiency of the persons in charge of relevant departments and the persons in charge of web production decreases, and as a result, there is a risk of adversely affecting the overall business operation.
Means for Solving the Problems
[0005] This invention provides a means for collecting links to all pages on an official website and obtaining the HTML data of those pages. It also includes a means for processing the obtained HTML data with a natural language processing engine to extract specific strings. Furthermore, it provides a means for listing the relevant pages and strings based on the extracted strings and for providing an interface that allows the user to modify and confirm the generated list. Finally, it includes a means for generating a report based on the list, including the user's modifications, and distributing it to the person in charge of the relevant department. This significantly improves the efficiency of work when revising pricing plans and service content, and reduces the workload of the person in charge.
[0006] A "link" refers to a reference or URL to another webpage that exists within a webpage.
[0007] "HTML data" refers to text data written in a markup language used to describe the content of a web page.
[0008] A "natural language processing engine" refers to artificial intelligence technology used to understand, analyze, and generate human language.
[0009] "Extracting a specific string" refers to removing a sequence of characters or words from data based on certain criteria or conditions.
[0010] "Listing" refers to organizing data by specific categories and compiling it into a list format.
[0011] An "interface" refers to the means or screen that a user uses to interact with a system.
[0012] A "user" refers to a person who operates a system and utilizes its functions.
[0013] A "report" refers to a document or data set that organizes, analyzes, and summarizes specific information.
[0014] "The 'person in charge' refers to a person in a position responsible for the management and operation of specific business or projects."
[0015] "Crawling" refers to the process of automatically traversing web pages on the Internet and collecting data."
Brief Description of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment." [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment." [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment." [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment." [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment." [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment." [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment." [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment." [Figure 9] It shows an emotion map to which multiple emotions are mapped." [Figure 10] It shows an emotion map to which multiple emotions are mapped." [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1." [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1." [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. This system enables efficient work without the need to outsource to external vendors.
[0038] This system is implemented using the following devices and programs.
[0039] System Configuration
[0040] 1. Crawling Module
[0041] The server collects links to all pages on the site based on the starting URL (e.g., "https: / / example.com"). This collection process involves parsing HTML "a" tags and recursively retrieving internal links.
[0042] 2. Data Acquisition Module
[0043] The server retrieves the HTML data for each page based on the collected links. The server makes an HTTP request for each URL and saves the HTML content of each page to a local database.
[0044] 3. Natural Language Processing (NLP) Module
[0045] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant strings.
[0046] 4. List creation module
[0047] The server creates a list of relevant pages and strings based on the extraction results output from the NLP engine. This list is saved in a structured data format such as JSON or CSV.
[0048] 5. User Interface Module
[0049] Users will review this list using their device and make corrections as needed. The user interface is designed to display the list through a browser, allowing users to easily identify the relevant sections.
[0050] 6. Report Generation Module
[0051] The server generates a final list reflecting the user's modifications and creates a report based on it. This report is output in PDF or CSV format and distributed to the relevant department heads.
[0052] Examples
[0053] Specific example 1: When introducing a new pricing plan
[0054] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0055] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new pricing plan". For example, if "https: / / example.com / plan1.html" contains the phrase "new pricing plan", that section will be extracted.
[0056] 3. The server creates a list of pages containing the extracted text and the corresponding strings, and saves it in JSON format. This list includes the URL of each page and the corresponding location within that page.
[0057] 4. The user reviews the list generated through the browser and makes corrections as needed. For example, if a string has been incorrectly extracted, the user corrects it.
[0058] 5. The server generates a report based on the final list, including the user's modifications, and distributes it to the relevant department head. This report includes links and content for all relevant pages.
[0059] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[0060] The following describes the processing flow.
[0061] Step 1:
[0062] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. The server parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0063] Step 2:
[0064] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[0065] Step 3:
[0066] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. The server instructs the NLP engine using a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision") to extract the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[0067] Step 4:
[0068] The server lists the corresponding pages and text based on the text extracted from the NLP engine. The server saves the extraction results in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0069] Step 5:
[0070] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0071] Step 6:
[0072] Users provide feedback to the server with the list they have reviewed and corrected. Users send the server the corrected items and their details, and the server updates the list based on that information. For example, if a user finds an error in the "New Pricing Plan" and enters the correct information, the server will reflect that correction.
[0073] Step 7:
[0074] The server generates a final list and creates a report based on the user's corrections. The server outputs the final list in CSV or PDF format and distributes it to the responsible person in charge in the relevant department. For example, the final report includes links to all relevant pages and their contents, and is provided in a format that the responsible person can easily review.
[0075] Thus, the system of the present invention enables efficient extraction and listing of revised pricing plans and service details through a series of steps, and prompt provision of this information to the relevant departments.
[0076] (Example 1)
[0077] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0078] With existing methods, identifying and listing relevant information from every page on the official website when revising pricing plans or service details is extremely time-consuming and laborious, making it difficult to do so efficiently and accurately. Furthermore, outsourcing this task to external vendors increases costs.
[0079] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0080] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, means for saving the data to a local database, and means for generating the report in PDF or CSV format. This makes it possible to efficiently and accurately identify and list the information necessary when revising pricing plans or service content.
[0081] An "official website" refers to a website on the internet operated by a specific organization or company.
[0082] "Means of collecting links" refers to methods or devices for extracting and listing hyperlinks within a specific webpage.
[0083] "HTML data" refers to data written in a standard markup language that defines the structure and content of a web page.
[0084] A "natural language processing engine" refers to a software system that analyzes text data and extracts specific patterns or keywords.
[0085] "Means for extracting specific strings" refers to a method or device for detecting and extracting specified keywords or phrases from data.
[0086] "Means of creating lists" refers to methods or devices for organizing information into a list format.
[0087] A "user-editable and verifiable interface" refers to a user interface that displays information and allows users to modify or verify it as needed.
[0088] "Means for generating the final list" refers to a method or apparatus for creating a final list based on the corrected information.
[0089] "Means of compiling and distributing reports" refers to methods or devices for sending generated information in report format to relevant parties.
[0090] A "local database" refers to a data storage location located locally and accessible without using the internet.
[0091] "Means of generating in PDF or CSV format" refers to a method or device for outputting information as a file in PDF (Portable Document Format) or CSV (Comma-Separated Values) format.
[0092] This invention provides a system that automatically identifies and lists relevant information from all pages of an official website when revising pricing plans or services. This system consists of multiple modules and means and aims to process information efficiently and accurately.
[0093] First, the server crawls all pages on the site starting from the specified start URL. The crawling is performed using the Python libraries BeautifulSoup and Requests. For the start URL (e.g., "https: / / example.com"), Requests is used to retrieve the HTML content, and BeautifulSoup is used to parse the content. During this process, the HTML "a" tags are parsed to collect internal links.
[0094] Next, the server makes an HTTP request to all the collected links and retrieves the HTML data for each page. This data is stored in a local database (such as SQLite or MySQL®) for later processing. For example, it stores the HTML obtained from URLs like "https: / / example.com / page1" and "https: / / example.com / page2".
[0095] The acquired HTML data is passed to a natural language processing (NLP) engine. The server uses libraries such as Python's NLTK and spaCy to analyze keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plan", "service details", "price revision"), it extracts the relevant strings.
[0096] Next, the server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The Python pandas library is used to create a data frame, and the data is then converted to the desired storage format. For example, if "https: / / example.com / page1" contains the phrase "New Pricing Plan," that information is added to the list.
[0097] The user reviews the generated list through their browser. The user interface is built using HTML and JavaScript (registered trademark), allowing the user to check relevant sections and make corrections as needed. For example, the user can correct incorrectly extracted content.
[0098] Finally, the server generates a final list reflecting the user's changes. A report is then created based on this final list. The report is generated in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation. For example, a file titled "Price Revision Report" is generated and sent to the relevant department head via email.
[0099] As a concrete example, the following prompt statement is used:
[0100] 1. "Crawl every page of the official website and collect all links."
[0101] 2. "Retrieve the HTML data of the collected links and save it to the database."
[0102] 3. "Analyze the HTML data and extract the text that matches 'New Pricing Plan'."
[0103] 4. "List the extracted text and URLs and save them in JSON format."
[0104] 5. "Use a browser to check the list and make corrections as needed."
[0105] 6. "Create a report based on the final list and distribute it to the relevant departments."
[0106] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[0107] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0108] Program processing flow
[0109] Step 1: Collect links
[0110] The server makes an HTTP request to the specified start URL (e.g., "https: / / example.com") and retrieves HTML content. It then parses the retrieved HTML content and collects internal links. Specifically, it accesses the start URL using the Python Requests library and parses the HTML "a" tags using BeautifulSoup. This retrieves all internal links as a list.
[0111] Input: Start URL (e.g., "https: / / example.com")
[0112] Output: List of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[0113] Step 2: Get HTML data
[0114] The server makes an individual HTTP request to each collected link and retrieves the HTML content of each page. The retrieved HTML data is saved to a local database (e.g., SQLite or MySQL) for later analysis. The Python Requests library is used to access the collected links, retrieve the HTML data, and save it to the database.
[0115] Input: A list of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[0116] Output: Recording of HTML content corresponding to each URL
[0117] Step 3: Keyword analysis using an NLP engine
[0118] The server passes the stored HTML data to a natural language processing engine. Using an NLP engine (e.g., NLTK or spaCy), it extracts the relevant strings from the HTML data of each page based on a predefined keyword list (e.g., "new pricing plan", "service details", "price revision").
[0119] Input: HTML data stored in a local database, keyword list
[0120] Output: Extracted strings and their corresponding page information
[0121] Step 4: List the extracted results
[0122] The server lists the relevant pages and strings based on the analysis results output from the NLP engine. It then uses the Python pandas library to organize the extracted text data and save it in JSON or CSV format.
[0123] Input: Extracted string and page information
[0124] Output: Structured list (e.g., JSON file)
[0125] Step 5: User reviews and modifies the list.
[0126] Users view the list generated from their browser using their device. The user interface is built using HTML and JavaScript and is designed to allow users to easily check the list contents and modify them as needed.
[0127] Input: List in JSON format
[0128] Output: User-modified list
[0129] Step 6: Generating the final list and creating the report.
[0130] The server generates a final list reflecting the user's modifications. Based on this final list, reports are created in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation.
[0131] Input: User-modified list
[0132] Output: Final list and report files (e.g., PDF format, CSV format)
[0133] By sequentially performing the processes from Step 1 to Step 6 in this way, a system is realized that efficiently collects and analyzes relevant information from all pages of the official website, and then creates and distributes a final report.
[0134] (Application Example 1)
[0135] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0136] Updating pricing plans and service revisions on the official website is a time-consuming and labor-intensive process. Furthermore, there is a lack of a means to quickly and accurately grasp the information on all pages without relying on external vendors. Therefore, there is a need for a system that efficiently and automatically collects information, extracts specific strings, and allows users to easily correct and verify the information.
[0137] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0138] In this invention, the server includes means for collecting links to all pages on the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, and means for running as an application installed on a smartphone that automatically analyzes and extracts specific strings corresponding to specific prompt messages. This makes it possible to streamline and quickly and accurately update information related to price plan and service revisions on the official website.
[0139] A "link" is a reference to another page or resource within a web page.
[0140] "HTML data" refers to data in HTML format, which is a markup language used to describe the structure and content of a web page.
[0141] A "natural language processing engine" is software that analyzes text data and extracts meaningful information.
[0142] A "specific string" refers to a predefined keyword or phrase.
[0143] "Methods for creating lists" refers to methods and processes for structuring acquired data and organizing it into a list.
[0144] An "interface" refers to the means or methods by which a user interacts with a system.
[0145] "Method of compiling and distributing a report" refers to the method of organizing the final data and sending it to relevant parties as a document.
[0146] An "application installed on a smartphone" is a program or software that can be executed on a smartphone.
[0147] A "prompt message" is text that a user uses to input an operation or instruction for the system to perform.
[0148] This invention details a system for efficiently and automatically collecting information and extracting specific strings when pricing plans or services are revised on an official website. This system consists of the following main modules:
[0149] 1. Link Collection Module
[0150] The server collects links to all pages within the official website based on the starting URL. To do this, the server parses the HTML "a" tags and recursively retrieves internal links.
[0151] 2. Data Acquisition Module
[0152] The server retrieves HTML data for each page based on the collected links. Specifically, it makes an HTTP request for each URL, retrieves the HTML content of each page, and saves it to a local database.
[0153] 3. Natural Language Processing Module
[0154] The server passes the acquired HTML data to a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the relevant strings.
[0155] 4. List creation module
[0156] The server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format.
[0157] 5. User Interface Module
[0158] Users can view this list through a smartphone app and make corrections as needed. The user interface is designed to display the list via a browser, allowing users to review and modify the relevant sections.
[0159] 6. Report Generation Module
[0160] The server generates a report based on the final list that reflects the user's modifications, and distributes it to relevant parties in PDF or CSV format.
[0161] Furthermore, the system of the present invention includes means for automatically analyzing and extracting specific strings corresponding to specific prompt statements, as an "application installed on a smartphone." This allows the system to operate based on a specific prompt statement, such as "Automatically identify and list the page on https: / / example.com that contains new pricing plans and campaign information," enabling it to update information quickly and accurately.
[0162] This system includes a series of processes from link collection and data acquisition to natural language processing analysis, listing, user interface verification, final list generation, and report distribution. In this way, it is possible to efficiently manage changes resulting from revisions to pricing plans and services on the official website, saving time and effort.
[0163] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0164] Step 1:
[0165] The server collects links to all pages within the official website based on the starting URL. Specifically, it accesses the starting URL and recursively retrieves internal links by parsing the HTML "a" tags. The input is the starting URL, and the output is a list of collected links.
[0166] Step 2:
[0167] The server retrieves HTML data for each page based on the collected links. It sends an HTTP request to each link to retrieve the page's HTML content. The retrieved HTML data is stored in a local database. The input is a list of links, and the output is the HTML content of each page.
[0168] Step 3:
[0169] The server processes the acquired HTML data using a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the corresponding strings. The input consists of HTML data and a keyword list, and the output is a list of the corresponding strings.
[0170] Step 4:
[0171] The server lists the extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The input is information about the extracted strings and their corresponding pages, and the output is a structured data list.
[0172] Step 5:
[0173] Users review the list generated through the smartphone app and make corrections as needed. The list is displayed via a user interface, allowing users to review and modify relevant sections. The input is the generated list, and the output is the final list with user modifications.
[0174] Step 6:
[0175] The server generates a report based on the final list that reflects the user's modifications. This report is delivered in PDF or CSV format. The input is the final modified list, and the output is the final report.
[0176] Step 7:
[0177] The server operates as an application installed on a smartphone, automatically parsing and extracting specific strings corresponding to specific prompt messages. It takes a prompt message as input and outputs a list of parsed and extracted strings. For example, the parsing is performed based on the prompt message, "Automatically identify and list the pages on https: / / example.com that contain new pricing plans and campaign information."
[0178] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0179] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. Furthermore, this invention has the function of improving the accuracy of user feedback by combining it with an emotion engine that recognizes user emotions. This system enables efficient work to be carried out without outsourcing to external vendors.
[0180] This system is implemented using the following devices and programs.
[0181] System Configuration
[0182] 1. Crawling Module
[0183] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0184] 2. Data Acquisition Module
[0185] The server retrieves HTML data for each page based on the collected links and saves the HTML data obtained for each URL via HTTP requests to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[0186] 3. Natural Language Processing (NLP) Module
[0187] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "New Pricing Plan" from "https: / / example.com / plan1.html", it records that text.
[0188] 4. List creation module
[0189] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0190] 5. User Interface Module
[0191] The user views the list on their device and examines its contents. Through their browser, the user checks the items in the list (page URLs and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0192] 6. Emotion Recognition Module
[0193] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[0194] 7. Feedback and notification module
[0195] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0196] 8. Report Generation Module
[0197] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[0198] Examples
[0199] Specific example 2: When revising the content of a new service
[0200] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0201] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[0202] 3. The server creates a list of pages containing the extracted text and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[0203] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[0204] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[0205] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[0206] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the relevant department head. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the department head.
[0207] The following describes the processing flow.
[0208] Step 1:
[0209] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. It parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0210] Step 2:
[0211] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[0212] Step 3:
[0213] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[0214] Step 4:
[0215] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0216] Step 5:
[0217] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0218] Step 6:
[0219] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[0220] Step 7:
[0221] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0222] Step 8:
[0223] The user provides feedback to the server with the list they have reviewed and corrected. The corrections are sent to the server, which then updates the list in the local database based on them. For example, if a user finds an error in "New Pricing Plans" and corrects it, the server reflects the correction and keeps the database up to date.
[0224] Step 9:
[0225] The server generates the final list and creates a report. The server outputs the final list in CSV or PDF format and distributes it to the responsible personnel in the relevant departments. For example, the final report includes links to all relevant pages and their contents, provided in a format that is easy for the responsible personnel to review.
[0226] In this way, a system is realized that efficiently extracts and lists information on revisions to pricing plans and service details, and improves the accuracy of feedback by combining this with user sentiment recognition.
[0227] (Example 2)
[0228] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0229] When revising pricing plans or service details, manually identifying and listing relevant strings from all pages of the official website is time-consuming and laborious. Therefore, a system that can perform this task efficiently and accurately is needed. Furthermore, improving the accuracy of feedback that takes user sentiment into account is also necessary.
[0230] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0231] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for providing an emotion engine that recognizes the user's emotions in real time, means for improving the accuracy of feedback and providing notifications based on the emotional state recognized by the emotion engine, and means for generating a final list including user modifications and distributing it as a report. This enables quick and accurate responses when changing pricing plans or service content, and allows for improved accuracy of feedback that takes user emotions into consideration.
[0232] An "official website" is an online information page managed by a specific company or organization.
[0233] A "link" is a reference to another page or resource within a web page, and is implemented using the HTML "a" tag.
[0234] "HTML data" is a standard markup language used to describe the structure and content of web pages.
[0235] A "natural language processing engine" is a collection of algorithms and programs designed to analyze and understand text data.
[0236] A "specific string" refers to a predefined keyword or phrase, which is the string that the system targets for extraction.
[0237] "Listing" refers to systematically organizing extracted data and saving it in list format.
[0238] An "interface" refers to the screen or input device that a user uses to interact with a system.
[0239] An "emotion engine" is a collection of software and hardware used to analyze and recognize a user's emotional state.
[0240] "Feedback" refers to collecting opinions and evaluations from users and using them to improve and adjust the system.
[0241] A "report" is a document that summarizes the data and results that have been collected and analyzed.
[0242] This invention is a system that automatically identifies and lists relevant strings from each page of an official website when price plans or service details are revised. Furthermore, it incorporates an emotion engine to improve the accuracy of user feedback. This system provides a means to perform tasks efficiently and can execute processing automatically without outsourcing to external vendors.
[0243] System Configuration
[0244] 1. Crawling module:
[0245] The server sets a starting URL and crawls the entire site from this URL. It collects internal links by parsing HTML "a" tags and recursively retrieves all internal links. For example, the server sets "https: / / example.com" as the starting URL, accesses "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0246] 2. Data Acquisition Module:
[0247] The server retrieves HTML data for each page based on the collected links. It uses HTTP requests to access each link and saves the retrieved HTML data to a local database. For example, the server sends an HTTP request to "https: / / example.com / plan1.html", retrieves its HTML data, and saves it.
[0248] 3. Natural Language Processing (NLP) Module:
[0249] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined keyword list, the NLP engine extracts the relevant text. For example, the server extracts descriptions related to "new pricing plans" from "https: / / example.com / plan1.html".
[0250] 4. List creation module:
[0251] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0252] 5. User Interface Module:
[0253] The user reviews the generated list on their device and scrutinizes its contents. They check the list items through their browser and make corrections as needed. For example, if the user finds incorrect information in the list, they correct it.
[0254] 6. Emotion Recognition Module:
[0255] A built-in emotion engine in the device recognizes the user's emotions in real time as they review and edit lists. It analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate their emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[0256] 7. Feedback and notification module:
[0257] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[0258] 8. Report generation module:
[0259] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[0260] Specific example
[0261] When the new service content is revised
[0262] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0263] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[0264] 3. The server lists the pages containing the extracted text and the text itself, and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[0265] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string has been incorrectly extracted, the user can correct it.
[0266] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[0267] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[0268] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the relevant department head. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the department head.
[0269] Example of a prompt
[0270] Please tell me how to design a system that automatically identifies and lists descriptions related to "new pricing plans" and "service details" from all pages on the official website when revising new service content, and how to improve the accuracy of feedback by recognizing user sentiment in real time.
[0271] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0272] Step 1:
[0273] Link crawling
[0274] The server sets a starting URL and crawls the entire site from this URL. The input is the starting URL, and the server parses the HTML "a" tags and recursively collects all internal links on all pages within the site. The output is a list of all links. For example, if the server sets "https: / / example.com" as the starting URL and accesses "https: / / example.com / index.html", it will discover internal links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" and add them to the list.
[0275] Step 2:
[0276] Retrieving and saving HTML data
[0277] The server retrieves HTML data for each page based on the collected links. The input is a list of collected links; the server accesses each link using an HTTP request and saves the retrieved HTML data to a local database. The output is the HTML data for each page. For example, the server sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[0278] Step 3:
[0279] Text analysis using Natural Language Processing (NLP)
[0280] The server passes the retrieved HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. The input consists of saved HTML data and a predefined keyword list, and the NLP engine extracts the relevant text. The output is the extracted specific string. For example, the server extracts the description of "new pricing plan" from "https: / / example.com / plan1.html".
[0281] Step 4:
[0282] Creation and Saving of a List
[0283] Based on the extracted text, the server creates a list of corresponding pages and text. The input is the extracted text and the corresponding page URL, and the list is saved in a structured data format such as JSON or CSV. The output is structured list data. For example, if the server has "https: / / example.com / plan1.html" and the text related to "New Pricing Plan" included in that page, it adds them to the list.
[0284] Step 5:
[0285] Confirmation and Revision of the List
[0286] The user checks the generated list on the terminal and scrutinizes the content. The input is the generated list, and the user checks the items of the list (page URL and corresponding text) through a browser and makes revisions if necessary. The output is the revised list. For example, if the user discovers incorrect information in the list, they correct it.
[0287] Step 6:
[0288] Emotion Recognition
[0289] The emotion engine incorporated in the terminal recognizes the emotion of the user in real time when the user checks and revises the list. The input is the user's operation data and biometric data (facial expression, tone of voice, typing speed), and the engine analyzes them to evaluate the emotional state. The output is the emotion recognition result. For example, if the user is feeling strong stress, it detects it.
[0290] Step 7:
[0291] Feedback and Notification<00P0921>
[0292] The server uses the user's emotional state, recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the emotion recognition result, which the server uses to provide additional information or notifications. The output is information on improving the accuracy of feedback and notification content. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[0293] Step 8:
[0294] Report generation and distribution
[0295] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. Inputs are the corrected list and sentiment recognition results; the final list is output in CSV or PDF format and distributed to the relevant department heads. The output is the final report. For example, the final report includes links and content for all relevant pages and is provided in a format easily accessible to the department heads.
[0296] (Application Example 2)
[0297] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0298] When revising pricing plans or services on an official website, the process of identifying and listing relevant strings from all pages is time-consuming and often outsourced to external vendors, leading to increased costs. Furthermore, the low accuracy of user feedback is a significant challenge.
[0299] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting links to all pages of the official website, means for obtaining markup language data for each page based on the collected links, means for extracting specific strings by running the obtained markup language data through a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for improving the accuracy of feedback using an emotion engine that recognizes the user's emotions, and means for generating a final list including user modifications and distributing it as a report. This makes it possible to efficiently identify and list specific strings from all pages of the official website and improve the accuracy of user feedback.
[0300] An "official website" is a collection of publicly available information on the internet provided by companies, organizations, etc.
[0301] A "link" refers to a hypertext connection between web pages, an element that allows users to navigate to other web pages by clicking on it.
[0302] "Markup language data" refers to text data that includes a set of tags that define the structure and content of a web page, and is written in formats such as HTML and XML.
[0303] A "natural language processing engine" is a software module designed to understand and analyze human language, and it has the function of extracting specific information from language data.
[0304] A "string" refers to a sequence of consecutive characters and constitutes a part of text data.
[0305] "Listing" refers to the process of organizing data into a list format based on certain rules.
[0306] "Interface" refers to the interface or mechanism for humans and computer systems to exchange information with each other.
[0307] "Emotion Engine" is a software module for real-time recognition of a person's emotional state from the person's facial expressions, voice tones, typing speed of characters, etc.
[0308] "Final List" refers to the finally determined data list after reflecting the modifications made by the user.
[0309] "Report" is a document that summarizes the collected and analyzed information and is for reporting to relevant parties according to specific purposes.
[0310] The present invention relates to a system for automatically collecting information during tariff plan and service revisions on an official website and improving the feedback accuracy of users. This system is configured as follows.
[0311] 1. Crawling Module
[0312] The server sets the start URL of the official website and collects the links of all pages within the site. At this time, the "a" tag of HTML is analyzed to recursively obtain the internal links. For example, the process of accessing "https: / / example.com" and collecting its internal links is executed.
[0313] 2. Data Collection Module
[0314] The server obtains the markup language data (HTML data) of each page based on the collected links. The obtained HTML data is obtained from each URL using an HTTP request. For example, an HTTP request is sent to "https: / / example.com / plan1.html" to obtain the HTML data of that page.
[0315] 3. Natural Language Processing Engine
[0316] The server passes the acquired HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. For example, based on a list of keywords such as "new pricing plan," "service details," and "price revision," the NLP engine extracts the relevant text.
[0317] 4. List creation module
[0318] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format (JSON or CSV). For example, it adds "https: / / example.com / plan1.html" and the text related to "New Pricing Plans" contained on that page to the list.
[0319] 5. User Interface Module
[0320] Users can view and modify the generated list through their browser. For example, if the list contains incorrect information, the user can correct it in their browser.
[0321] 6. Emotion Recognition Module
[0322] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[0323] 7. Feedback and notification module
[0324] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0325] 8. Report Generation Module
[0326] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant departments. For example, the final report includes links and content for all relevant pages and is provided in a format that can be easily reviewed by the relevant departments.
[0327] Specific example:
[0328] This system is used by e-commerce site administrators when revising prices for new products. The administrator crawls the entire site to identify pages and text related to price changes. They then review the list and make corrections, and if the sentiment engine detects stress on the administrator, it automatically provides guidelines and support.
[0329] Example of a prompt:
[0330] "Identify the relevant page and text regarding the price revision for the new product. Recognize user sentiment in real time to improve the accuracy of feedback using an emotion engine."
[0331] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0332] Step 1:
[0333] The server sets the official website's starting URL and collects links to all pages within the site. The input is the starting URL, and the output is a list of links to all pages. Specifically, the server parses HTML "a" tags and recursively retrieves internal links. This process allows the server to understand the overall page structure of the site.
[0334] Step 2:
[0335] The server retrieves the markup language data (HTML data) for each page based on the collected links. The input is a list of links, and the output is the HTML data for each page. Specifically, the server uses HTTP requests to retrieve the linked web pages and stores their HTML data locally.
[0336] Step 3:
[0337] The server processes the acquired HTML data using a natural language processing (NLP) engine to analyze keywords related to pricing plans and services. The input is HTML data, and the output is the corresponding text. Specifically, the NLP engine extracts the relevant text based on a predefined list of keywords (e.g., "new pricing plan," "service details," "price revision").
[0338] Step 4:
[0339] The server creates a list of corresponding pages and texts based on the extracted text. The input is the text and the data of the corresponding page, and the output is a list in a structured data format (JSON or CSV). Specifically, the extracted text and its page URL are combined, added to the list, and saved.
[0340] Step 5:
[0341] The user reviews the generated list through a browser and modifies its contents. The input is a structured data list, and the output is the user-modified list. Specifically, the user checks the list contents using a web browser and corrects any inaccuracies.
[0342] Step 6:
[0343] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Inputs include the user's facial expressions, tone of voice, and typing speed, while output is an evaluation of the user's emotions. Specifically, it monitors the user's emotional state using a camera and microphone and analyzes that state.
[0344] Step 7:
[0345] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the user's emotional evaluation, and the output is suggestions for improving feedback accuracy and notifications. Specifically, if high levels of stress or dissatisfaction are detected, the server provides additional guidance and support to help resolve the issue.
[0346] Step 8:
[0347] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The input is the user-modified list and sentiment recognition results, and the output is the final report. Specifically, the final list is output in CSV or PDF format and distributed to the relevant departments.
[0348] This series of processing steps allows for the efficient identification and listing of specific strings from all pages of the official website, thereby improving the accuracy of user feedback.
[0349] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0350] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0351] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0352] [Second Embodiment]
[0353] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0354] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0355] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0356] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0357] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0358] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0359] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0360] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0361] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0362] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0363] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0364] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0365] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. This system enables efficient work without the need to outsource to external vendors.
[0366] This system is implemented using the following devices and programs.
[0367] System Configuration
[0368] 1. Crawling Module
[0369] The server collects links to all pages on the site based on the starting URL (e.g., "https: / / example.com"). This collection process involves parsing HTML "a" tags and recursively retrieving internal links.
[0370] 2. Data Acquisition Module
[0371] The server retrieves the HTML data for each page based on the collected links. The server makes an HTTP request for each URL and saves the HTML content of each page to a local database.
[0372] 3. Natural Language Processing (NLP) Module
[0373] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant strings.
[0374] 4. List creation module
[0375] The server creates a list of relevant pages and strings based on the extraction results output from the NLP engine. This list is saved in a structured data format such as JSON or CSV.
[0376] 5. User Interface Module
[0377] Users will review this list using their device and make corrections as needed. The user interface is designed to display the list through a browser, allowing users to easily identify the relevant sections.
[0378] 6. Report Generation Module
[0379] The server generates a final list reflecting the user's modifications and creates a report based on it. This report is output in PDF or CSV format and distributed to the relevant department heads.
[0380] Examples
[0381] Specific example 1: When introducing a new pricing plan
[0382] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0383] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new pricing plan". For example, if "https: / / example.com / plan1.html" contains the phrase "new pricing plan", that part will be extracted.
[0384] 3. The server creates a list of pages containing the extracted text and the corresponding strings, and saves it in JSON format. This list includes the URL of each page and the corresponding location within that page.
[0385] 4. The user reviews the list generated through the browser and makes corrections as needed. For example, if a string has been incorrectly extracted, the user corrects it.
[0386] 5. The server generates a report based on the final list, including the user's modifications, and distributes it to the relevant department head. This report includes links and content for all relevant pages.
[0387] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[0388] The following describes the processing flow.
[0389] Step 1:
[0390] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. The server parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0391] Step 2:
[0392] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[0393] Step 3:
[0394] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. The server instructs the NLP engine using a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision") to extract the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[0395] Step 4:
[0396] The server lists the corresponding pages and text based on the text extracted from the NLP engine. The server saves the extraction results in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0397] Step 5:
[0398] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0399] Step 6:
[0400] Users provide feedback to the server with the list they have reviewed and corrected. Users send the server the corrected items and their details, and the server updates the list based on that information. For example, if a user finds an error in the "New Pricing Plan" and enters the correct information, the server will reflect that correction.
[0401] Step 7:
[0402] The server generates a final list and creates a report based on the user's corrections. The server outputs the final list in CSV or PDF format and distributes it to the responsible person in charge in the relevant department. For example, the final report includes links to all relevant pages and their contents, and is provided in a format that the responsible person can easily review.
[0403] Thus, the system of the present invention enables efficient extraction and listing of revised pricing plans and service details through a series of steps, and prompt provision of this information to the relevant departments.
[0404] (Example 1)
[0405] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0406] With existing methods, identifying and listing relevant information from every page on the official website when revising pricing plans or service details is extremely time-consuming and laborious, making it difficult to do so efficiently and accurately. Furthermore, outsourcing this task to external vendors increases costs.
[0407] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0408] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, means for saving the data to a local database, and means for generating the report in PDF or CSV format. This makes it possible to efficiently and accurately identify and list the information necessary when revising pricing plans or service content.
[0409] An "official website" refers to a website on the internet operated by a specific organization or company.
[0410] "Means of collecting links" refers to methods or devices for extracting and listing hyperlinks within a specific webpage.
[0411] "HTML data" refers to data written in a standard markup language that defines the structure and content of a web page.
[0412] A "natural language processing engine" refers to a software system that analyzes text data and extracts specific patterns or keywords.
[0413] "Means for extracting specific strings" refers to a method or device for detecting and extracting specified keywords or phrases from data.
[0414] "Means of creating lists" refers to methods or devices for organizing information into a list format.
[0415] A "user-editable and verifiable interface" refers to a user interface that displays information and allows users to modify or verify it as needed.
[0416] "Means for generating the final list" refers to a method or apparatus for creating a final list based on the corrected information.
[0417] "Means of compiling and distributing reports" refers to methods or devices for sending generated information in report format to relevant parties.
[0418] A "local database" refers to a data storage location located locally and accessible without using the internet.
[0419] "Means of generating in PDF or CSV format" refers to a method or device for outputting information as a file in PDF (Portable Document Format) or CSV (Comma-Separated Values) format.
[0420] This invention provides a system that automatically identifies and lists relevant information from all pages of an official website when revising pricing plans or services. This system consists of multiple modules and means and aims to process information efficiently and accurately.
[0421] First, the server crawls all pages on the site starting from the specified start URL. The crawling is performed using the Python libraries BeautifulSoup and Requests. For the start URL (e.g., "https: / / example.com"), Requests is used to retrieve the HTML content, and BeautifulSoup is used to parse the content. During this process, the HTML "a" tags are parsed to collect internal links.
[0422] Next, the server makes an HTTP request to all the collected links and retrieves the HTML data for each page. This data is stored in a local database (such as SQLite or MySQL) for later processing. For example, it stores the HTML obtained from URLs like "https: / / example.com / page1" and "https: / / example.com / page2".
[0423] The acquired HTML data is passed to a natural language processing (NLP) engine. The server uses libraries such as Python's NLTK and spaCy to analyze keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plan", "service details", "price revision"), it extracts the relevant strings.
[0424] Next, the server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The Python pandas library is used to create a data frame, and the data is then converted to the desired storage format. For example, if "https: / / example.com / page1" contains the phrase "New Pricing Plan," that information is added to the list.
[0425] The user reviews the generated list through their browser. The user interface is built using HTML and JavaScript, allowing the user to check relevant sections and make corrections as needed. For example, the user can correct incorrectly extracted content.
[0426] Finally, the server generates a final list reflecting the user's changes. A report is then created based on this final list. The report is generated in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation. For example, a file titled "Price Revision Report" is generated and sent to the relevant department head via email.
[0427] As a concrete example, the following prompt statement is used:
[0428] 1. "Crawl every page of the official website and collect all links."
[0429] 2. "Retrieve the HTML data of the collected links and save it to the database."
[0430] 3. "Analyze the HTML data and extract the text that matches 'New Pricing Plan'."
[0431] 4. "List the extracted text and URLs and save them in JSON format."
[0432] 5. "Use a browser to check the list and make corrections as needed."
[0433] 6. "Create a report based on the final list and distribute it to the relevant departments."
[0434] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[0435] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0436] Program processing flow
[0437] Step 1: Collect links
[0438] The server makes an HTTP request to the specified start URL (e.g., "https: / / example.com") and retrieves HTML content. It then parses the retrieved HTML content and collects internal links. Specifically, it accesses the start URL using the Python Requests library and parses the HTML "a" tags using BeautifulSoup. This retrieves all internal links as a list.
[0439] Input: Start URL (e.g., "https: / / example.com")
[0440] Output: List of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[0441] Step 2: Get HTML data
[0442] The server makes an individual HTTP request to each collected link and retrieves the HTML content of each page. The retrieved HTML data is saved to a local database (e.g., SQLite or MySQL) for later analysis. The Python Requests library is used to access the collected links, retrieve the HTML data, and save it to the database.
[0443] Input: A list of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[0444] Output: Recording of HTML content corresponding to each URL
[0445] Step 3: Keyword analysis using an NLP engine
[0446] The server passes the stored HTML data to a natural language processing engine. Using an NLP engine (e.g., NLTK or spaCy), it extracts the relevant strings from the HTML data of each page based on a predefined keyword list (e.g., "new pricing plan", "service details", "price revision").
[0447] Input: HTML data stored in a local database, keyword list
[0448] Output: Extracted strings and their corresponding page information
[0449] Step 4: List the extracted results
[0450] The server lists the relevant pages and strings based on the analysis results output from the NLP engine. It then uses the Python pandas library to organize the extracted text data and save it in JSON or CSV format.
[0451] Input: Extracted string and page information
[0452] Output: Structured list (e.g., JSON file)
[0453] Step 5: User reviews and modifies the list.
[0454] Users view the list generated from their browser using their device. The user interface is built using HTML and JavaScript and is designed to allow users to easily check the list contents and modify them as needed.
[0455] Input: List in JSON format
[0456] Output: User-modified list
[0457] Step 6: Generating the final list and creating the report.
[0458] The server generates a final list reflecting the user's modifications. Based on this final list, reports are created in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation.
[0459] Input: User-modified list
[0460] Output: Final list and report files (e.g., PDF format, CSV format)
[0461] By sequentially performing the processes from Step 1 to Step 6, a system is created that efficiently collects and analyzes relevant information from all pages of the official website, and then generates and distributes a final report.
[0462] (Application Example 1)
[0463] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0464] Updating pricing plans and service revisions on the official website is a time-consuming and labor-intensive process. Furthermore, there is a lack of a means to quickly and accurately grasp the information on all pages without relying on external vendors. Therefore, there is a need for a system that efficiently and automatically collects information, extracts specific strings, and allows users to easily correct and verify the information.
[0465] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0466] In this invention, the server includes means for collecting links to all pages on the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, and means for running as an application installed on a smartphone that automatically analyzes and extracts specific strings corresponding to specific prompt messages. This makes it possible to streamline and quickly and accurately update information related to price plan and service revisions on the official website.
[0467] A "link" is a reference to another page or resource within a web page.
[0468] "HTML data" refers to data in HTML format, which is a markup language used to describe the structure and content of a web page.
[0469] A "natural language processing engine" is software that analyzes text data and extracts meaningful information.
[0470] A "specific string" refers to a predefined keyword or phrase.
[0471] "Methods for creating lists" refers to methods and processes for structuring acquired data and organizing it into a list.
[0472] An "interface" refers to the means or methods by which a user interacts with a system.
[0473] "Method of compiling and distributing a report" refers to the method of organizing the final data and sending it to relevant parties as a document.
[0474] An "application installed on a smartphone" is a program or software that can be executed on a smartphone.
[0475] A "prompt message" is text that a user uses to input an operation or instruction for the system to perform.
[0476] This invention details a system for efficiently and automatically collecting information and extracting specific strings when pricing plans or services are revised on an official website. This system consists of the following main modules:
[0477] 1. Link Collection Module
[0478] The server collects links to all pages within the official website based on the starting URL. To do this, the server parses the HTML "a" tags and recursively retrieves internal links.
[0479] 2. Data Acquisition Module
[0480] The server retrieves HTML data for each page based on the collected links. Specifically, it makes an HTTP request for each URL, retrieves the HTML content of each page, and saves it to a local database.
[0481] 3. Natural Language Processing Module
[0482] The server passes the acquired HTML data to a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the relevant strings.
[0483] 4. List creation module
[0484] The server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format.
[0485] 5. User Interface Module
[0486] Users can view this list through a smartphone app and make corrections as needed. The user interface is designed to display the list via a browser, allowing users to review and modify the relevant sections.
[0487] 6. Report Generation Module
[0488] The server generates a report based on the final list that reflects the user's modifications, and distributes it to relevant parties in PDF or CSV format.
[0489] Furthermore, the system of the present invention includes means for automatically analyzing and extracting specific strings corresponding to specific prompt statements, as an "application installed on a smartphone." This allows the system to operate based on a specific prompt statement, such as "Automatically identify and list the page on https: / / example.com that contains new pricing plans and campaign information," enabling it to update information quickly and accurately.
[0490] This system includes a series of processes from link collection and data acquisition to natural language processing analysis, listing, user interface verification, final list generation, and report distribution. In this way, it is possible to efficiently manage changes resulting from revisions to pricing plans and services on the official website, saving time and effort.
[0491] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0492] Step 1:
[0493] The server collects links to all pages within the official website based on the starting URL. Specifically, it accesses the starting URL and recursively retrieves internal links by parsing the HTML "a" tags. The input is the starting URL, and the output is a list of collected links.
[0494] Step 2:
[0495] The server retrieves HTML data for each page based on the collected links. It sends an HTTP request to each link to retrieve the page's HTML content. The retrieved HTML data is stored in a local database. The input is a list of links, and the output is the HTML content of each page.
[0496] Step 3:
[0497] The server processes the acquired HTML data using a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the corresponding strings. The input consists of HTML data and a keyword list, and the output is a list of the corresponding strings.
[0498] Step 4:
[0499] The server lists the extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The input is information about the extracted strings and their corresponding pages, and the output is a structured data list.
[0500] Step 5:
[0501] Users review the list generated through the smartphone app and make corrections as needed. The list is displayed via a user interface, allowing users to review and modify relevant sections. The input is the generated list, and the output is the final list with user modifications.
[0502] Step 6:
[0503] The server generates a report based on the final list that reflects the user's modifications. This report is delivered in PDF or CSV format. The input is the final modified list, and the output is the final report.
[0504] Step 7:
[0505] The server operates as an application installed on a smartphone, automatically parsing and extracting specific strings corresponding to specific prompt messages. It takes a prompt message as input and outputs a list of parsed and extracted strings. For example, the parsing is performed based on the prompt message, "Automatically identify and list the pages on https: / / example.com that contain new pricing plans and campaign information."
[0506] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0507] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. Furthermore, this invention has the function of improving the accuracy of user feedback by combining it with an emotion engine that recognizes user emotions. This system enables efficient work to be carried out without outsourcing to external vendors.
[0508] This system is implemented using the following devices and programs.
[0509] System Configuration
[0510] 1. Crawling Module
[0511] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0512] 2. Data Acquisition Module
[0513] The server retrieves HTML data for each page based on the collected links and saves the HTML data obtained for each URL via HTTP requests to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[0514] 3. Natural Language Processing (NLP) Module
[0515] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "New Pricing Plan" from "https: / / example.com / plan1.html", it records that text.
[0516] 4. List creation module
[0517] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0518] 5. User Interface Module
[0519] The user views the list on their device and examines its contents. Through their browser, the user checks the items in the list (page URLs and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0520] 6. Emotion Recognition Module
[0521] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[0522] 7. Feedback and notification module
[0523] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0524] 8. Report Generation Module
[0525] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[0526] Examples
[0527] Specific example 2: When revising the content of a new service
[0528] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0529] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[0530] 3. The server creates a list of pages containing the extracted text and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[0531] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[0532] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and modify lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[0533] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[0534] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the relevant department head. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the department head.
[0535] The following describes the processing flow.
[0536] Step 1:
[0537] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. It parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0538] Step 2:
[0539] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[0540] Step 3:
[0541] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[0542] Step 4:
[0543] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0544] Step 5:
[0545] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0546] Step 6:
[0547] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[0548] Step 7:
[0549] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0550] Step 8:
[0551] The user provides feedback to the server with the list they have reviewed and corrected. The corrections are sent to the server, which then updates the list in the local database based on them. For example, if a user finds an error in "New Pricing Plans" and corrects it, the server reflects the correction and keeps the database up to date.
[0552] Step 9:
[0553] The server generates the final list and creates a report. The server outputs the final list in CSV or PDF format and distributes it to the responsible personnel in the relevant departments. For example, the final report includes links to all relevant pages and their contents, provided in a format that is easy for the responsible personnel to review.
[0554] In this way, a system is realized that efficiently extracts and lists information on revisions to pricing plans and service details, and improves the accuracy of feedback by combining this with user sentiment recognition.
[0555] (Example 2)
[0556] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0557] When revising pricing plans or service details, manually identifying and listing relevant strings from all pages of the official website is time-consuming and laborious. Therefore, a system that can perform this task efficiently and accurately is needed. Furthermore, improving the accuracy of feedback that takes user sentiment into account is also necessary.
[0558] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0559] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for providing an emotion engine that recognizes the user's emotions in real time, means for improving the accuracy of feedback and providing notifications based on the emotional state recognized by the emotion engine, and means for generating a final list including user modifications and distributing it as a report. This enables quick and accurate responses when changing pricing plans or service content, and allows for improved accuracy of feedback that takes user emotions into consideration.
[0560] An "official website" is an online information page managed by a specific company or organization.
[0561] A "link" is a reference to another page or resource within a web page, and is implemented using the HTML "a" tag.
[0562] "HTML data" is a standard markup language used to describe the structure and content of web pages.
[0563] A "natural language processing engine" is a collection of algorithms and programs designed to analyze and understand text data.
[0564] A "specific string" refers to a predefined keyword or phrase, which is the string that the system targets for extraction.
[0565] "Listing" refers to systematically organizing extracted data and saving it in list format.
[0566] An "interface" refers to the screen or input device that a user uses to interact with a system.
[0567] An "emotion engine" is a collection of software and hardware used to analyze and recognize a user's emotional state.
[0568] "Feedback" refers to collecting opinions and evaluations from users and using them to improve and adjust the system.
[0569] A "report" is a document that summarizes the data and results that have been collected and analyzed.
[0570] This invention is a system that automatically identifies and lists relevant strings from each page of an official website when price plans or service details are revised. Furthermore, it incorporates an emotion engine to improve the accuracy of user feedback. This system provides a means to perform tasks efficiently and can execute processing automatically without outsourcing to external vendors.
[0571] System Configuration
[0572] 1. Crawling module:
[0573] The server sets a starting URL and crawls the entire site from this URL. It collects internal links by parsing HTML "a" tags and recursively retrieves all internal links. For example, the server sets "https: / / example.com" as the starting URL, accesses "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0574] 2. Data Acquisition Module:
[0575] The server retrieves HTML data for each page based on the collected links. It uses HTTP requests to access each link and saves the retrieved HTML data to a local database. For example, the server sends an HTTP request to "https: / / example.com / plan1.html", retrieves its HTML data, and saves it.
[0576] 3. Natural Language Processing (NLP) Module:
[0577] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined keyword list, the NLP engine extracts the relevant text. For example, the server extracts descriptions related to "new pricing plans" from "https: / / example.com / plan1.html".
[0578] 4. List creation module:
[0579] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0580] 5. User Interface Module:
[0581] The user reviews the generated list on their device and scrutinizes its contents. They check the list items through their browser and make corrections as needed. For example, if the user finds incorrect information in the list, they correct it.
[0582] 6. Emotion Recognition Module:
[0583] A built-in emotion engine in the device recognizes the user's emotions in real time as they review and edit lists. It analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate their emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[0584] 7. Feedback and notification module:
[0585] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[0586] 8. Report generation module:
[0587] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[0588] Specific example
[0589] When the new service content is revised
[0590] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0591] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[0592] 3. The server lists the pages containing the extracted text and the text itself, and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[0593] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[0594] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and modify lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[0595] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[0596] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the responsible person in charge of the relevant department. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the responsible person.
[0597] Example of a prompt
[0598] Please tell me how to design a system that automatically identifies and lists descriptions related to "new pricing plans" and "service details" from all pages on the official website when revising new service content, and how to improve the accuracy of feedback by recognizing user sentiment in real time.
[0599] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0600] Step 1:
[0601] Link crawling
[0602] The server sets a starting URL and crawls the entire site from this URL. The input is the starting URL, and the server parses the HTML "a" tags and recursively collects all internal links on all pages within the site. The output is a list of all links. For example, if the server sets "https: / / example.com" as the starting URL and accesses "https: / / example.com / index.html", it will discover internal links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" and add them to the list.
[0603] Step 2:
[0604] Retrieving and saving HTML data
[0605] The server retrieves HTML data for each page based on the collected links. The input is a list of collected links; the server accesses each link using an HTTP request and saves the retrieved HTML data to a local database. The output is the HTML data for each page. For example, the server sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[0606] Step 3:
[0607] Text analysis using Natural Language Processing (NLP)
[0608] The server passes the retrieved HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. The input consists of saved HTML data and a predefined keyword list, and the NLP engine extracts the relevant text. The output is the extracted specific string. For example, the server extracts the description of "new pricing plan" from "https: / / example.com / plan1.html".
[0609] Step 4:
[0610] Creating and saving lists
[0611] The server creates a list of corresponding pages and text based on the extracted text. The input is the extracted text and the corresponding page URL, and the list is saved in a structured data format such as JSON or CSV. The output is structured list data. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0612] Step 5:
[0613] Check and correct the list
[0614] The user reviews the generated list on their device and scrutinizes its contents. The input is the generated list; the user checks the list items (page URLs and corresponding text) through their browser and makes corrections as needed. The output is the corrected list. For example, if the user finds incorrect information in the list, they correct it.
[0615] Step 6:
[0616] emotion recognition
[0617] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Input consists of user operation data and biometric data (facial expressions, voice tone, typing speed), which the engine analyzes and evaluates the emotional state. The output is the emotion recognition result. For example, if the user is experiencing high levels of stress, it will be detected.
[0618] Step 7:
[0619] Feedback and notifications
[0620] The server uses the user's emotional state, recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the emotion recognition result, which the server uses to provide additional information or notifications. The output is information on improving the accuracy of feedback and notification content. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[0621] Step 8:
[0622] Report generation and distribution
[0623] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. Inputs are the corrected list and sentiment recognition results; the final list is output in CSV or PDF format and distributed to the relevant department heads. The output is the final report. For example, the final report includes links and content for all relevant pages and is provided in a format easily accessible to the department heads.
[0624] (Application Example 2)
[0625] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0626] When revising pricing plans or services on an official website, the process of identifying and listing relevant strings from all pages is time-consuming and often outsourced to external vendors, leading to increased costs. Furthermore, the low accuracy of user feedback is a significant challenge.
[0627] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting links to all pages of the official website, means for obtaining markup language data for each page based on the collected links, means for extracting specific strings by running the obtained markup language data through a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for improving the accuracy of feedback using an emotion engine that recognizes the user's emotions, and means for generating a final list including user modifications and distributing it as a report. This makes it possible to efficiently identify and list specific strings from all pages of the official website and improve the accuracy of user feedback.
[0628] An "official website" is a collection of publicly available information on the internet provided by companies, organizations, etc.
[0629] A "link" refers to a hypertext connection between web pages, an element that allows users to navigate to other web pages by clicking on it.
[0630] "Markup language data" refers to text data that includes a set of tags that define the structure and content of a web page, and is written in formats such as HTML and XML.
[0631] A "natural language processing engine" is a software module designed to understand and analyze human language, and it has the function of extracting specific information from language data.
[0632] A "string" refers to a sequence of consecutive characters and constitutes a part of text data.
[0633] "Listing" refers to the process of organizing data into a list format based on certain rules.
[0634] An "interface" is the boundary or mechanism by which humans and computer systems exchange information with each other.
[0635] An "emotion engine" is a software module that recognizes a person's emotional state in real time based on their facial expressions, tone of voice, typing speed, and other factors.
[0636] The "final list" refers to the final, confirmed list of data after incorporating user modifications.
[0637] A "report" is a document that compiles collected and analyzed information, intended to be presented to relevant parties for specific purposes.
[0638] This invention relates to a system for automatically collecting information on pricing plans and service revisions on official websites and for improving the accuracy of user feedback. This system is configured as follows:
[0639] 1. Crawling Module
[0640] The server sets the official website's starting URL and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com" and executes a process to collect its internal links.
[0641] 2. Data Acquisition Module
[0642] The server retrieves the markup language data (HTML data) for each page based on the collected links. The retrieved HTML data is obtained from each URL using HTTP requests. For example, an HTTP request is sent to "https: / / example.com / plan1.html" to retrieve the HTML data for that page.
[0643] 3. Natural Language Processing Engine
[0644] The server passes the acquired HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. For example, based on a list of keywords such as "new pricing plan," "service details," and "price revision," the NLP engine extracts the relevant text.
[0645] 4. List creation module
[0646] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format (JSON or CSV). For example, it adds "https: / / example.com / plan1.html" and the text related to "New Pricing Plans" contained on that page to the list.
[0647] 5. User Interface Module
[0648] Users can view and modify the generated list through their browser. For example, if the list contains incorrect information, the user can correct it in their browser.
[0649] 6. Emotion Recognition Module
[0650] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[0651] 7. Feedback and notification module
[0652] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0653] 8. Report Generation Module
[0654] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant departments. For example, the final report includes links and content for all relevant pages and is provided in a format that can be easily reviewed by the relevant departments.
[0655] Specific example:
[0656] This system is used by e-commerce site administrators when revising prices for new products. The administrator crawls the entire site to identify pages and text related to price changes. They then review the list and make corrections, and if the sentiment engine detects stress on the administrator, it automatically provides guidelines and support.
[0657] Example of a prompt:
[0658] "Identify the relevant page and text regarding the price revision for the new product. Recognize user sentiment in real time to improve the accuracy of feedback using an emotion engine."
[0659] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0660] Step 1:
[0661] The server sets the official website's starting URL and collects links to all pages within the site. The input is the starting URL, and the output is a list of links to all pages. Specifically, the server parses HTML "a" tags and recursively retrieves internal links. This process allows the server to understand the overall page structure of the site.
[0662] Step 2:
[0663] The server retrieves the markup language data (HTML data) for each page based on the collected links. The input is a list of links, and the output is the HTML data for each page. Specifically, the server uses HTTP requests to retrieve the linked web pages and stores their HTML data locally.
[0664] Step 3:
[0665] The server processes the acquired HTML data using a natural language processing (NLP) engine to analyze keywords related to pricing plans and services. The input is HTML data, and the output is the corresponding text. Specifically, the NLP engine extracts the relevant text based on a predefined list of keywords (e.g., "new pricing plan," "service details," "price revision").
[0666] Step 4:
[0667] The server creates a list of corresponding pages and texts based on the extracted text. The input is the text and the data of the corresponding page, and the output is a list in a structured data format (JSON or CSV). Specifically, the extracted text and its page URL are combined, added to the list, and saved.
[0668] Step 5:
[0669] The user reviews the generated list through a browser and modifies its contents. The input is a structured data list, and the output is the user-modified list. Specifically, the user checks the list contents using a web browser and corrects any inaccuracies.
[0670] Step 6:
[0671] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Inputs include the user's facial expressions, tone of voice, and typing speed, while output is an evaluation of the user's emotions. Specifically, it monitors the user's emotional state using a camera and microphone and analyzes that state.
[0672] Step 7:
[0673] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the user's emotional evaluation, and the output is suggestions for improving feedback accuracy and notifications. Specifically, if high levels of stress or dissatisfaction are detected, the server provides additional guidance and support to help resolve the issue.
[0674] Step 8:
[0675] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The input is the user-modified list and sentiment recognition results, and the output is the final report. Specifically, the final list is output in CSV or PDF format and distributed to the relevant departments.
[0676] This series of processing steps allows for the efficient identification and listing of specific strings from all pages of the official website, thereby improving the accuracy of user feedback.
[0677] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0678] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0679] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0680] [Third Embodiment]
[0681] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0682] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0683] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0684] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0685] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0686] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0687] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0688] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0689] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0690] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0691] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0692] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0693] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. This system enables efficient work without the need to outsource to external vendors.
[0694] This system is implemented using the following devices and programs.
[0695] System Configuration
[0696] 1. Crawling Module
[0697] The server collects links to all pages on the site based on the starting URL (e.g., "https: / / example.com"). This collection process involves parsing HTML "a" tags and recursively retrieving internal links.
[0698] 2. Data Acquisition Module
[0699] The server retrieves the HTML data for each page based on the collected links. The server makes an HTTP request for each URL and saves the HTML content of each page to a local database.
[0700] 3. Natural Language Processing (NLP) Module
[0701] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant strings.
[0702] 4. List creation module
[0703] The server creates a list of relevant pages and strings based on the extraction results output from the NLP engine. This list is saved in a structured data format such as JSON or CSV.
[0704] 5. User Interface Module
[0705] Users will review this list using their device and make corrections as needed. The user interface is designed to display the list through a browser, allowing users to easily identify the relevant sections.
[0706] 6. Report Generation Module
[0707] The server generates a final list reflecting the user's modifications and creates a report based on it. This report is output in PDF or CSV format and distributed to the relevant department heads.
[0708] Examples
[0709] Specific example 1: When introducing a new pricing plan
[0710] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0711] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new pricing plan". For example, if "https: / / example.com / plan1.html" contains the phrase "new pricing plan", that part will be extracted.
[0712] 3. The server creates a list of pages containing the extracted text and the corresponding strings, and saves it in JSON format. This list includes the URL of each page and the corresponding location within that page.
[0713] 4. The user reviews the list generated through the browser and makes corrections as needed. For example, if a string has been incorrectly extracted, the user corrects it.
[0714] 5. The server generates a report based on the final list, including the user's modifications, and distributes it to the relevant department head. This report includes links and content for all relevant pages.
[0715] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[0716] The following describes the processing flow.
[0717] Step 1:
[0718] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. The server parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0719] Step 2:
[0720] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[0721] Step 3:
[0722] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. The server instructs the NLP engine using a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision") to extract the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[0723] Step 4:
[0724] The server lists the corresponding pages and text based on the text extracted from the NLP engine. The server saves the extraction results in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0725] Step 5:
[0726] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0727] Step 6:
[0728] Users provide feedback to the server with the list they have reviewed and corrected. Users send the server the corrected items and their details, and the server updates the list based on that information. For example, if a user finds an error in the "New Pricing Plan" and enters the correct information, the server will reflect that correction.
[0729] Step 7:
[0730] The server generates a final list and creates a report based on the user's corrections. The server outputs the final list in CSV or PDF format and distributes it to the responsible person in charge in the relevant department. For example, the final report includes links to all relevant pages and their contents, and is provided in a format that the responsible person can easily review.
[0731] Thus, the system of the present invention enables efficient extraction and listing of revised pricing plans and service details through a series of steps, and prompt provision of this information to the relevant departments.
[0732] (Example 1)
[0733] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0734] With existing methods, identifying and listing relevant information from every page on the official website when revising pricing plans or service details is extremely time-consuming and laborious, making it difficult to do so efficiently and accurately. Furthermore, outsourcing this task to external vendors increases costs.
[0735] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0736] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, means for saving the data to a local database, and means for generating the report in PDF or CSV format. This makes it possible to efficiently and accurately identify and list the information necessary when revising pricing plans or service content.
[0737] An "official website" refers to a website on the internet operated by a specific organization or company.
[0738] "Means of collecting links" refers to methods or devices for extracting and listing hyperlinks within a specific webpage.
[0739] "HTML data" refers to data written in a standard markup language that defines the structure and content of a web page.
[0740] A "natural language processing engine" refers to a software system that analyzes text data and extracts specific patterns or keywords.
[0741] "Means for extracting specific strings" refers to a method or device for detecting and extracting specified keywords or phrases from data.
[0742] "Means of creating lists" refers to methods or devices for organizing information into a list format.
[0743] A "user-editable and verifiable interface" refers to a user interface that displays information and allows users to modify or verify it as needed.
[0744] "Means for generating the final list" refers to a method or apparatus for creating a final list based on the corrected information.
[0745] "Means of compiling and distributing reports" refers to methods or devices for sending generated information in report format to relevant parties.
[0746] A "local database" refers to a data storage location located locally and accessible without using the internet.
[0747] "Means of generating in PDF or CSV format" refers to a method or device for outputting information as a file in PDF (Portable Document Format) or CSV (Comma-Separated Values) format.
[0748] This invention provides a system that automatically identifies and lists relevant information from all pages of an official website when revising pricing plans or services. This system consists of multiple modules and means and aims to process information efficiently and accurately.
[0749] First, the server crawls all pages on the site starting from the specified start URL. The crawling is performed using the Python libraries BeautifulSoup and Requests. For the start URL (e.g., "https: / / example.com"), Requests is used to retrieve the HTML content, and BeautifulSoup is used to parse the content. During this process, the HTML "a" tags are parsed to collect internal links.
[0750] Next, the server makes an HTTP request to all the collected links and retrieves the HTML data for each page. This data is stored in a local database (such as SQLite or MySQL) for later processing. For example, it stores the HTML obtained from URLs like "https: / / example.com / page1" and "https: / / example.com / page2".
[0751] The acquired HTML data is passed to a natural language processing (NLP) engine. The server uses libraries such as Python's NLTK and spaCy to analyze keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plan", "service details", "price revision"), it extracts the relevant strings.
[0752] Next, the server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The Python pandas library is used to create a data frame, and the data is then converted to the desired storage format. For example, if "https: / / example.com / page1" contains the phrase "New Pricing Plan," that information is added to the list.
[0753] The user reviews the generated list through their browser. The user interface is built using HTML and JavaScript, allowing the user to check relevant sections and make corrections as needed. For example, the user can correct incorrectly extracted content.
[0754] Finally, the server generates a final list reflecting the user's changes. A report is then created based on this final list. The report is generated in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation. For example, a file titled "Price Revision Report" is generated and sent to the relevant department head via email.
[0755] As a concrete example, the following prompt statement is used:
[0756] 1. "Crawl every page of the official website and collect all links."
[0757] 2. "Retrieve the HTML data of the collected links and save it to the database."
[0758] 3. "Analyze the HTML data and extract the text that matches 'New Pricing Plan'."
[0759] 4. "List the extracted text and URLs and save them in JSON format."
[0760] 5. "Use a browser to check the list and make corrections as needed."
[0761] 6. "Create a report based on the final list and distribute it to the relevant departments."
[0762] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[0763] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0764] Program processing flow
[0765] Step 1: Collect links
[0766] The server makes an HTTP request to the specified start URL (e.g., "https: / / example.com") and retrieves HTML content. It then parses the retrieved HTML content and collects internal links. Specifically, it accesses the start URL using the Python Requests library and parses the HTML "a" tags using BeautifulSoup. This retrieves all internal links as a list.
[0767] Input: Start URL (e.g., "https: / / example.com")
[0768] Output: List of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[0769] Step 2: Get HTML data
[0770] The server makes an individual HTTP request to each collected link and retrieves the HTML content of each page. The retrieved HTML data is saved to a local database (e.g., SQLite or MySQL) for later analysis. The Python Requests library is used to access the collected links, retrieve the HTML data, and save it to the database.
[0771] Input: A list of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[0772] Output: Recording of HTML content corresponding to each URL
[0773] Step 3: Keyword analysis using an NLP engine
[0774] The server passes the stored HTML data to a natural language processing engine. Using an NLP engine (e.g., NLTK or spaCy), it extracts the relevant strings from the HTML data of each page based on a predefined keyword list (e.g., "new pricing plan", "service details", "price revision").
[0775] Input: HTML data stored in a local database, keyword list
[0776] Output: Extracted strings and their corresponding page information
[0777] Step 4: List the extracted results
[0778] The server lists the relevant pages and strings based on the analysis results output from the NLP engine. It then uses the Python pandas library to organize the extracted text data and save it in JSON or CSV format.
[0779] Input: Extracted string and page information
[0780] Output: Structured list (e.g., JSON file)
[0781] Step 5: User reviews and modifies the list.
[0782] Users view the list generated from their browser using their device. The user interface is built using HTML and JavaScript and is designed to allow users to easily check the list contents and modify them as needed.
[0783] Input: List in JSON format
[0784] Output: User-modified list
[0785] Step 6: Generating the final list and creating the report.
[0786] The server generates a final list reflecting the user's modifications. Based on this final list, reports are created in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation.
[0787] Input: User-modified list
[0788] Output: Final list and report files (e.g., PDF format, CSV format)
[0789] By sequentially performing the processes from Step 1 to Step 6, a system is created that efficiently collects and analyzes relevant information from all pages of the official website, and then generates and distributes a final report.
[0790] (Application Example 1)
[0791] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0792] Updating pricing plans and service revisions on the official website is a time-consuming and labor-intensive process. Furthermore, there is a lack of a means to quickly and accurately grasp the information on all pages without relying on external vendors. Therefore, there is a need for a system that efficiently and automatically collects information, extracts specific strings, and allows users to easily correct and verify the information.
[0793] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0794] In this invention, the server includes means for collecting links to all pages on the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, and means for running as an application installed on a smartphone that automatically analyzes and extracts specific strings corresponding to specific prompt messages. This makes it possible to streamline and quickly and accurately update information related to price plan and service revisions on the official website.
[0795] A "link" is a reference to another page or resource within a web page.
[0796] "HTML data" refers to data in HTML format, which is a markup language used to describe the structure and content of a web page.
[0797] A "natural language processing engine" is software that analyzes text data and extracts meaningful information.
[0798] A "specific string" refers to a predefined keyword or phrase.
[0799] "Methods for creating lists" refers to methods and processes for structuring acquired data and organizing it into a list.
[0800] An "interface" refers to the means or methods by which a user interacts with a system.
[0801] "Method of compiling and distributing a report" refers to the method of organizing the final data and sending it to relevant parties as a document.
[0802] An "application installed on a smartphone" is a program or software that can be executed on a smartphone.
[0803] A "prompt message" is text that a user uses to input an operation or instruction for the system to perform.
[0804] This invention details a system for efficiently and automatically collecting information and extracting specific strings when pricing plans or services are revised on an official website. This system consists of the following main modules:
[0805] 1. Link Collection Module
[0806] The server collects links to all pages within the official website based on the starting URL. To do this, the server parses the HTML "a" tags and recursively retrieves internal links.
[0807] 2. Data Acquisition Module
[0808] The server retrieves HTML data for each page based on the collected links. Specifically, it makes an HTTP request for each URL, retrieves the HTML content of each page, and saves it to a local database.
[0809] 3. Natural Language Processing Module
[0810] The server passes the acquired HTML data to a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the relevant strings.
[0811] 4. List creation module
[0812] The server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format.
[0813] 5. User Interface Module
[0814] Users can view this list through a smartphone app and make corrections as needed. The user interface is designed to display the list via a browser, allowing users to review and modify the relevant sections.
[0815] 6. Report Generation Module
[0816] The server generates a report based on the final list that reflects the user's modifications, and distributes it to relevant parties in PDF or CSV format.
[0817] Furthermore, the system of the present invention includes means for automatically analyzing and extracting specific strings corresponding to specific prompt statements, as an "application installed on a smartphone." This allows the system to operate based on a specific prompt statement, such as "Automatically identify and list the page on https: / / example.com that contains new pricing plans and campaign information," enabling it to update information quickly and accurately.
[0818] This system includes a series of processes from link collection and data acquisition to natural language processing analysis, listing, user interface verification, final list generation, and report distribution. In this way, it is possible to efficiently manage changes resulting from revisions to pricing plans and services on the official website, saving time and effort.
[0819] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0820] Step 1:
[0821] The server collects links to all pages within the official website based on the starting URL. Specifically, it accesses the starting URL and recursively retrieves internal links by parsing the HTML "a" tags. The input is the starting URL, and the output is a list of collected links.
[0822] Step 2:
[0823] The server retrieves HTML data for each page based on the collected links. It sends an HTTP request to each link to retrieve the page's HTML content. The retrieved HTML data is stored in a local database. The input is a list of links, and the output is the HTML content of each page.
[0824] Step 3:
[0825] The server processes the acquired HTML data using a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the corresponding strings. The input consists of HTML data and a keyword list, and the output is a list of the corresponding strings.
[0826] Step 4:
[0827] The server lists the extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The input is information about the extracted strings and their corresponding pages, and the output is a structured data list.
[0828] Step 5:
[0829] Users review the list generated through the smartphone app and make corrections as needed. The list is displayed via a user interface, allowing users to review and modify relevant sections. The input is the generated list, and the output is the final list with user modifications.
[0830] Step 6:
[0831] The server generates a report based on the final list that reflects the user's modifications. This report is delivered in PDF or CSV format. The input is the final modified list, and the output is the final report.
[0832] Step 7:
[0833] The server operates as an application installed on a smartphone, automatically parsing and extracting specific strings corresponding to specific prompt messages. It takes a prompt message as input and outputs a list of parsed and extracted strings. For example, the parsing is performed based on the prompt message, "Automatically identify and list the pages on https: / / example.com that contain new pricing plans and campaign information."
[0834] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0835] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. Furthermore, this invention has the function of improving the accuracy of user feedback by combining it with an emotion engine that recognizes user emotions. This system enables efficient work to be carried out without outsourcing to external vendors.
[0836] This system is implemented using the following devices and programs.
[0837] System Configuration
[0838] 1. Crawling Module
[0839] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0840] 2. Data Acquisition Module
[0841] The server retrieves HTML data for each page based on the collected links and saves the HTML data obtained for each URL via HTTP requests to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[0842] 3. Natural Language Processing (NLP) Module
[0843] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "New Pricing Plan" from "https: / / example.com / plan1.html", it records that text.
[0844] 4. List creation module
[0845] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0846] 5. User Interface Module
[0847] The user views the list on their device and examines its contents. Through their browser, the user checks the items in the list (page URLs and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0848] 6. Emotion Recognition Module
[0849] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[0850] 7. Feedback and notification module
[0851] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0852] 8. Report Generation Module
[0853] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[0854] Examples
[0855] Specific example 2: When revising the content of a new service
[0856] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0857] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[0858] 3. The server creates a list of pages containing the extracted text and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[0859] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[0860] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and modify lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[0861] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[0862] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the relevant department head. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the department head.
[0863] The following describes the processing flow.
[0864] Step 1:
[0865] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. It parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0866] Step 2:
[0867] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[0868] Step 3:
[0869] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[0870] Step 4:
[0871] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0872] Step 5:
[0873] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[0874] Step 6:
[0875] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[0876] Step 7:
[0877] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0878] Step 8:
[0879] The user provides feedback to the server with the list they have reviewed and corrected. The corrections are sent to the server, which then updates the list in the local database based on them. For example, if a user finds an error in "New Pricing Plans" and corrects it, the server reflects the correction and keeps the database up to date.
[0880] Step 9:
[0881] The server generates the final list and creates a report. The server outputs the final list in CSV or PDF format and distributes it to the responsible personnel in the relevant departments. For example, the final report includes links to all relevant pages and their contents, provided in a format that is easy for the responsible personnel to review.
[0882] In this way, a system is realized that efficiently extracts and lists information on revisions to pricing plans and service details, and improves the accuracy of feedback by combining this with user sentiment recognition.
[0883] (Example 2)
[0884] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0885] When revising pricing plans or service details, manually identifying and listing relevant strings from all pages of the official website is time-consuming and laborious. Therefore, a system that can perform this task efficiently and accurately is needed. Furthermore, improving the accuracy of feedback that takes user sentiment into account is also necessary.
[0886] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0887] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for providing an emotion engine that recognizes the user's emotions in real time, means for improving the accuracy of feedback and providing notifications based on the emotional state recognized by the emotion engine, and means for generating a final list including user modifications and distributing it as a report. This enables quick and accurate responses when changing pricing plans or service content, and allows for improved accuracy of feedback that takes user emotions into consideration.
[0888] An "official website" is an online information page managed by a specific company or organization.
[0889] A "link" is a reference to another page or resource within a web page, and is implemented using the HTML "a" tag.
[0890] "HTML data" is a standard markup language used to describe the structure and content of web pages.
[0891] A "natural language processing engine" is a collection of algorithms and programs designed to analyze and understand text data.
[0892] A "specific string" refers to a predefined keyword or phrase, which is the string that the system targets for extraction.
[0893] "Listing" refers to systematically organizing extracted data and saving it in list format.
[0894] An "interface" refers to the screen or input device that a user uses to interact with a system.
[0895] An "emotion engine" is a collection of software and hardware used to analyze and recognize a user's emotional state.
[0896] "Feedback" refers to collecting opinions and evaluations from users and using them to improve and adjust the system.
[0897] A "report" is a document that summarizes the data and results that have been collected and analyzed.
[0898] This invention is a system that automatically identifies and lists relevant strings from each page of an official website when price plans or service details are revised. Furthermore, it incorporates an emotion engine to improve the accuracy of user feedback. This system provides a means to perform tasks efficiently and can execute processing automatically without outsourcing to external vendors.
[0899] System Configuration
[0900] 1. Crawling module:
[0901] The server sets a starting URL and crawls the entire site from this URL. It collects internal links by parsing HTML "a" tags and recursively retrieves all internal links. For example, the server sets "https: / / example.com" as the starting URL, accesses "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[0902] 2. Data Acquisition Module:
[0903] The server retrieves HTML data for each page based on the collected links. It uses HTTP requests to access each link and saves the retrieved HTML data to a local database. For example, the server sends an HTTP request to "https: / / example.com / plan1.html", retrieves its HTML data, and saves it.
[0904] 3. Natural Language Processing (NLP) Module:
[0905] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined keyword list, the NLP engine extracts the relevant text. For example, the server extracts descriptions related to "new pricing plans" from "https: / / example.com / plan1.html".
[0906] 4. List creation module:
[0907] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0908] 5. User Interface Module:
[0909] The user reviews the generated list on their device and scrutinizes its contents. They check the list items through their browser and make corrections as needed. For example, if the user finds incorrect information in the list, they correct it.
[0910] 6. Emotion Recognition Module:
[0911] A built-in emotion engine in the device recognizes the user's emotions in real time as they review and edit lists. It analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate their emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[0912] 7. Feedback and notification module:
[0913] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[0914] 8. Report generation module:
[0915] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[0916] Specific example
[0917] When the new service content is revised
[0918] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[0919] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[0920] 3. The server lists the pages containing the extracted text and the text itself, and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[0921] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[0922] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and modify lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[0923] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[0924] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the responsible person in charge of the relevant department. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the responsible person.
[0925] Example of a prompt
[0926] Please tell me how to design a system that automatically identifies and lists descriptions related to "new pricing plans" and "service details" from all pages on the official website when revising new service content, and how to improve the accuracy of feedback by recognizing user sentiment in real time.
[0927] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0928] Step 1:
[0929] Link crawling
[0930] The server sets a starting URL and crawls the entire site from this URL. The input is the starting URL, and the server parses the HTML "a" tags and recursively collects all internal links on all pages within the site. The output is a list of all links. For example, if the server sets "https: / / example.com" as the starting URL and accesses "https: / / example.com / index.html", it will discover internal links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" and add them to the list.
[0931] Step 2:
[0932] Retrieving and saving HTML data
[0933] The server retrieves HTML data for each page based on the collected links. The input is a list of collected links; the server accesses each link using an HTTP request and saves the retrieved HTML data to a local database. The output is the HTML data for each page. For example, the server sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[0934] Step 3:
[0935] Text analysis using Natural Language Processing (NLP)
[0936] The server passes the retrieved HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. The input consists of saved HTML data and a predefined keyword list, and the NLP engine extracts the relevant text. The output is the extracted specific string. For example, the server extracts the description of "new pricing plan" from "https: / / example.com / plan1.html".
[0937] Step 4:
[0938] Creating and saving lists
[0939] The server creates a list of corresponding pages and text based on the extracted text. The input is the extracted text and the corresponding page URL, and the list is saved in a structured data format such as JSON or CSV. The output is structured list data. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[0940] Step 5:
[0941] Check and correct the list
[0942] The user reviews the generated list on their device and scrutinizes its contents. The input is the generated list; the user checks the list items (page URLs and corresponding text) through their browser and makes corrections as needed. The output is the corrected list. For example, if the user finds incorrect information in the list, they correct it.
[0943] Step 6:
[0944] emotion recognition
[0945] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Input consists of user operation data and biometric data (facial expressions, voice tone, typing speed), which the engine analyzes and evaluates the emotional state. The output is the emotion recognition result. For example, if the user is experiencing high levels of stress, it will be detected.
[0946] Step 7:
[0947] Feedback and notifications
[0948] The server uses the user's emotional state, recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the emotion recognition result, which the server uses to provide additional information or notifications. The output is information on improving the accuracy of feedback and notification content. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[0949] Step 8:
[0950] Report generation and distribution
[0951] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. Inputs are the corrected list and sentiment recognition results; the final list is output in CSV or PDF format and distributed to the relevant department heads. The output is the final report. For example, the final report includes links and content for all relevant pages and is provided in a format easily accessible to the department heads.
[0952] (Application Example 2)
[0953] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0954] When revising pricing plans or services on an official website, the process of identifying and listing relevant strings from all pages is time-consuming and often outsourced to external vendors, leading to increased costs. Furthermore, the low accuracy of user feedback is a significant challenge.
[0955] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting links to all pages of the official website, means for obtaining markup language data for each page based on the collected links, means for extracting specific strings by running the obtained markup language data through a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for improving the accuracy of feedback using an emotion engine that recognizes the user's emotions, and means for generating a final list including user modifications and distributing it as a report. This makes it possible to efficiently identify and list specific strings from all pages of the official website and improve the accuracy of user feedback.
[0956] An "official website" is a collection of publicly available information on the internet provided by companies, organizations, etc.
[0957] A "link" refers to a hypertext connection between web pages, an element that allows users to navigate to other web pages by clicking on it.
[0958] "Markup language data" refers to text data that includes a set of tags that define the structure and content of a web page, and is written in formats such as HTML and XML.
[0959] A "natural language processing engine" is a software module designed to understand and analyze human language, and it has the function of extracting specific information from language data.
[0960] A "string" refers to a sequence of consecutive characters and constitutes a part of text data.
[0961] "Listing" refers to the process of organizing data into a list format based on certain rules.
[0962] An "interface" is the boundary or mechanism by which humans and computer systems exchange information with each other.
[0963] An "emotion engine" is a software module that recognizes a person's emotional state in real time based on their facial expressions, tone of voice, typing speed, and other factors.
[0964] The "final list" refers to the final, confirmed list of data after incorporating user modifications.
[0965] A "report" is a document that compiles collected and analyzed information, intended to be presented to relevant parties for specific purposes.
[0966] This invention relates to a system for automatically collecting information on pricing plans and service revisions on official websites and for improving the accuracy of user feedback. This system is configured as follows:
[0967] 1. Crawling Module
[0968] The server sets the official website's starting URL and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com" and executes a process to collect its internal links.
[0969] 2. Data Acquisition Module
[0970] The server retrieves the markup language data (HTML data) for each page based on the collected links. The retrieved HTML data is obtained from each URL using HTTP requests. For example, an HTTP request is sent to "https: / / example.com / plan1.html" to retrieve the HTML data for that page.
[0971] 3. Natural Language Processing Engine
[0972] The server passes the acquired HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. For example, based on a list of keywords such as "new pricing plan," "service details," and "price revision," the NLP engine extracts the relevant text.
[0973] 4. List creation module
[0974] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format (JSON or CSV). For example, it adds "https: / / example.com / plan1.html" and the text related to "New Pricing Plans" contained on that page to the list.
[0975] 5. User Interface Module
[0976] Users can view and modify the generated list through their browser. For example, if the list contains incorrect information, the user can correct it in their browser.
[0977] 6. Emotion Recognition Module
[0978] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[0979] 7. Feedback and notification module
[0980] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[0981] 8. Report Generation Module
[0982] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant departments. For example, the final report includes links and content for all relevant pages and is provided in a format that can be easily reviewed by the relevant departments.
[0983] Specific example:
[0984] This system is used by e-commerce site administrators when revising prices for new products. The administrator crawls the entire site to identify pages and text related to price changes. They then review the list and make corrections, and if the sentiment engine detects stress on the administrator, it automatically provides guidelines and support.
[0985] Example of a prompt:
[0986] "Identify the relevant page and text regarding the price revision for the new product. Recognize user sentiment in real time to improve the accuracy of feedback using an emotion engine."
[0987] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0988] Step 1:
[0989] The server sets the official website's starting URL and collects links to all pages within the site. The input is the starting URL, and the output is a list of links to all pages. Specifically, the server parses HTML "a" tags and recursively retrieves internal links. This process allows the server to understand the overall page structure of the site.
[0990] Step 2:
[0991] The server retrieves the markup language data (HTML data) for each page based on the collected links. The input is a list of links, and the output is the HTML data for each page. Specifically, the server uses HTTP requests to retrieve the linked web pages and stores their HTML data locally.
[0992] Step 3:
[0993] The server processes the acquired HTML data using a natural language processing (NLP) engine to analyze keywords related to pricing plans and services. The input is HTML data, and the output is the corresponding text. Specifically, the NLP engine extracts the relevant text based on a predefined list of keywords (e.g., "new pricing plan," "service details," "price revision").
[0994] Step 4:
[0995] The server creates a list of corresponding pages and texts based on the extracted text. The input is the text and the data of the corresponding page, and the output is a list in a structured data format (JSON or CSV). Specifically, the extracted text and its page URL are combined, added to the list, and saved.
[0996] Step 5:
[0997] The user reviews the generated list through a browser and modifies its contents. The input is a structured data list, and the output is the user-modified list. Specifically, the user checks the list contents using a web browser and corrects any inaccuracies.
[0998] Step 6:
[0999] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Inputs include the user's facial expressions, tone of voice, and typing speed, while output is an evaluation of the user's emotions. Specifically, it monitors the user's emotional state using a camera and microphone and analyzes that state.
[1000] Step 7:
[1001] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the user's emotional evaluation, and the output is suggestions for improving feedback accuracy and notifications. Specifically, if high levels of stress or dissatisfaction are detected, the server provides additional guidance and support to help resolve the issue.
[1002] Step 8:
[1003] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The input is the user-modified list and sentiment recognition results, and the output is the final report. Specifically, the final list is output in CSV or PDF format and distributed to the relevant departments.
[1004] This series of processing steps allows for the efficient identification and listing of specific strings from all pages of the official website, thereby improving the accuracy of user feedback.
[1005] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1006] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1007] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1008] [Fourth Embodiment]
[1009] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1010] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1011] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1012] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1013] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1014] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1015] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1016] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1017] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1018] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1019] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1020] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1021] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1022] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. This system enables efficient work without the need to outsource to external vendors.
[1023] This system is implemented using the following devices and programs.
[1024] System Configuration
[1025] 1. Crawling Module
[1026] The server collects links to all pages on the site based on the starting URL (e.g., "https: / / example.com"). This collection process involves parsing HTML "a" tags and recursively retrieving internal links.
[1027] 2. Data Acquisition Module
[1028] The server retrieves the HTML data for each page based on the collected links. The server makes an HTTP request for each URL and saves the HTML content of each page to a local database.
[1029] 3. Natural Language Processing (NLP) Module
[1030] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant strings.
[1031] 4. List creation module
[1032] The server creates a list of relevant pages and strings based on the extraction results output from the NLP engine. This list is saved in a structured data format such as JSON or CSV.
[1033] 5. User Interface Module
[1034] Users will review this list using their device and make corrections as needed. The user interface is designed to display the list through a browser, allowing users to easily identify the relevant sections.
[1035] 6. Report Generation Module
[1036] The server generates a final list reflecting the user's modifications and creates a report based on it. This report is output in PDF or CSV format and distributed to the relevant department heads.
[1037] Examples
[1038] Specific example 1: When introducing a new pricing plan
[1039] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[1040] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new pricing plan". For example, if "https: / / example.com / plan1.html" contains the phrase "new pricing plan", that part will be extracted.
[1041] 3. The server creates a list of pages containing the extracted text and the corresponding strings, and saves it in JSON format. This list includes the URL of each page and the corresponding location within that page.
[1042] 4. The user reviews the list generated through the browser and makes corrections as needed. For example, if a string has been incorrectly extracted, the user corrects it.
[1043] 5. The server generates a report based on the final list, including the user's modifications, and distributes it to the relevant department head. This report includes links and content for all relevant pages.
[1044] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[1045] The following describes the processing flow.
[1046] Step 1:
[1047] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. The server parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[1048] Step 2:
[1049] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[1050] Step 3:
[1051] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. The server instructs the NLP engine using a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision") to extract the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[1052] Step 4:
[1053] The server lists the corresponding pages and text based on the text extracted from the NLP engine. The server saves the extraction results in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[1054] Step 5:
[1055] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[1056] Step 6:
[1057] Users provide feedback to the server with the list they have reviewed and corrected. Users send the server the corrected items and their details, and the server updates the list based on that information. For example, if a user finds an error in the "New Pricing Plan" and enters the correct information, the server will reflect that correction.
[1058] Step 7:
[1059] The server generates a final list and creates a report based on the user's corrections. The server outputs the final list in CSV or PDF format and distributes it to the responsible person in charge in the relevant department. For example, the final report includes links to all relevant pages and their contents, and is provided in a format that the responsible person can easily review.
[1060] Thus, the system of the present invention enables efficient extraction and listing of revised pricing plans and service details through a series of steps, and prompt provision of this information to the relevant departments.
[1061] (Example 1)
[1062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1063] With existing methods, identifying and listing relevant information from every page on the official website when revising pricing plans or service details is extremely time-consuming and laborious, making it difficult to do so efficiently and accurately. Furthermore, outsourcing this task to external vendors increases costs.
[1064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1065] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, means for saving the data to a local database, and means for generating the report in PDF or CSV format. This makes it possible to efficiently and accurately identify and list the information necessary when revising pricing plans or service content.
[1066] An "official website" refers to a website on the internet operated by a specific organization or company.
[1067] "Means of collecting links" refers to methods or devices for extracting and listing hyperlinks within a specific webpage.
[1068] "HTML data" refers to data written in a standard markup language that defines the structure and content of a web page.
[1069] A "natural language processing engine" refers to a software system that analyzes text data and extracts specific patterns or keywords.
[1070] "Means for extracting specific strings" refers to a method or device for detecting and extracting specified keywords or phrases from data.
[1071] "Means of creating lists" refers to methods or devices for organizing information into a list format.
[1072] A "user-editable and verifiable interface" refers to a user interface that displays information and allows users to modify or verify it as needed.
[1073] "Means for generating the final list" refers to a method or apparatus for creating a final list based on the corrected information.
[1074] "Means of compiling and distributing reports" refers to methods or devices for sending generated information in report format to relevant parties.
[1075] A "local database" refers to a data storage location located locally and accessible without using the internet.
[1076] "Means of generating in PDF or CSV format" refers to a method or device for outputting information as a file in PDF (Portable Document Format) or CSV (Comma-Separated Values) format.
[1077] This invention provides a system that automatically identifies and lists relevant information from all pages of an official website when revising pricing plans or services. This system consists of multiple modules and means and aims to process information efficiently and accurately.
[1078] First, the server crawls all pages on the site starting from the specified start URL. The crawling is performed using the Python libraries BeautifulSoup and Requests. For the start URL (e.g., "https: / / example.com"), Requests is used to retrieve the HTML content, and BeautifulSoup is used to parse the content. During this process, the HTML "a" tags are parsed to collect internal links.
[1079] Next, the server makes an HTTP request to all the collected links and retrieves the HTML data for each page. This data is stored in a local database (such as SQLite or MySQL) for later processing. For example, it stores the HTML obtained from URLs like "https: / / example.com / page1" and "https: / / example.com / page2".
[1080] The acquired HTML data is passed to a natural language processing (NLP) engine. The server uses libraries such as Python's NLTK and spaCy to analyze keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plan", "service details", "price revision"), it extracts the relevant strings.
[1081] Next, the server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The Python pandas library is used to create a data frame, and the data is then converted to the desired storage format. For example, if "https: / / example.com / page1" contains the phrase "New Pricing Plan," that information is added to the list.
[1082] The user reviews the generated list through their browser. The user interface is built using HTML and JavaScript, allowing the user to check relevant sections and make corrections as needed. For example, the user can correct incorrectly extracted content.
[1083] Finally, the server generates a final list reflecting the user's changes. A report is then created based on this final list. The report is generated in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation. For example, a file titled "Price Revision Report" is generated and sent to the relevant department head via email.
[1084] As a concrete example, the following prompt statement is used:
[1085] 1. "Crawl every page of the official website and collect all links."
[1086] 2. "Retrieve the HTML data of the collected links and save it to the database."
[1087] 3. "Analyze the HTML data and extract the text that matches 'New Pricing Plan'."
[1088] 4. "List the extracted text and URLs and save them in JSON format."
[1089] 5. "Use a browser to check the list and make corrections as needed."
[1090] 6. "Create a report based on the final list and distribute it to the relevant departments."
[1091] In this way, the system of the present invention streamlines the process of revising pricing plans and service content, enabling accurate and rapid responses.
[1092] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1093] Program processing flow
[1094] Step 1: Collect links
[1095] The server makes an HTTP request to the specified start URL (e.g., "https: / / example.com") and retrieves HTML content. It then parses the retrieved HTML content and collects internal links. Specifically, it accesses the start URL using the Python Requests library and parses the HTML "a" tags using BeautifulSoup. This retrieves all internal links as a list.
[1096] Input: Start URL (e.g., "https: / / example.com")
[1097] Output: List of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[1098] Step 2: Get HTML data
[1099] The server makes an individual HTTP request to each collected link and retrieves the HTML content of each page. The retrieved HTML data is saved to a local database (e.g., SQLite or MySQL) for later analysis. The Python Requests library is used to access the collected links, retrieve the HTML data, and save it to the database.
[1100] Input: A list of internal links (e.g., ["https: / / example.com / page1", "https: / / example.com / page2"])
[1101] Output: Recording of HTML content corresponding to each URL
[1102] Step 3: Keyword analysis using an NLP engine
[1103] The server passes the stored HTML data to a natural language processing engine. Using an NLP engine (e.g., NLTK or spaCy), it extracts the relevant strings from the HTML data of each page based on a predefined keyword list (e.g., "new pricing plan", "service details", "price revision").
[1104] Input: HTML data stored in a local database, keyword list
[1105] Output: Extracted strings and their corresponding page information
[1106] Step 4: List the extracted results
[1107] The server lists the relevant pages and strings based on the analysis results output from the NLP engine. It then uses the Python pandas library to organize the extracted text data and save it in JSON or CSV format.
[1108] Input: Extracted string and page information
[1109] Output: Structured list (e.g., JSON file)
[1110] Step 5: User reviews and modifies the list.
[1111] Users view the list generated from their browser using their device. The user interface is built using HTML and JavaScript and is designed to allow users to easily check the list contents and modify them as needed.
[1112] Input: List in JSON format
[1113] Output: User-modified list
[1114] Step 6: Generating the final list and creating the report.
[1115] The server generates a final list reflecting the user's modifications. Based on this final list, reports are created in PDF or CSV format and distributed to the relevant department heads. The Python ReportLab library is used for PDF generation.
[1116] Input: User-modified list
[1117] Output: Final list and report files (e.g., PDF format, CSV format)
[1118] By sequentially performing the processes from Step 1 to Step 6, a system is created that efficiently collects and analyzes relevant information from all pages of the official website, and then generates and distributes a final report.
[1119] (Application Example 1)
[1120] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1121] Updating pricing plans and service revisions on the official website is a time-consuming and labor-intensive process. Furthermore, there is a lack of a means to quickly and accurately grasp the information on all pages without relying on external vendors. Therefore, there is a need for a system that efficiently and automatically collects information, extracts specific strings, and allows users to easily correct and verify the information.
[1122] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1123] In this invention, the server includes means for collecting links to all pages on the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for generating a final list including user modifications and distributing it as a report, and means for running as an application installed on a smartphone that automatically analyzes and extracts specific strings corresponding to specific prompt messages. This makes it possible to streamline and quickly and accurately update information related to price plan and service revisions on the official website.
[1124] A "link" is a reference to another page or resource within a web page.
[1125] "HTML data" refers to data in HTML format, which is a markup language used to describe the structure and content of a web page.
[1126] A "natural language processing engine" is software that analyzes text data and extracts meaningful information.
[1127] A "specific string" refers to a predefined keyword or phrase.
[1128] "Methods for creating lists" refers to methods and processes for structuring acquired data and organizing it into a list.
[1129] An "interface" refers to the means or methods by which a user interacts with a system.
[1130] "Method of compiling and distributing a report" refers to the method of organizing the final data and sending it to relevant parties as a document.
[1131] An "application installed on a smartphone" is a program or software that can be executed on a smartphone.
[1132] A "prompt message" is text that a user uses to input an operation or instruction for the system to perform.
[1133] This invention details a system for efficiently and automatically collecting information and extracting specific strings when pricing plans or services are revised on an official website. This system consists of the following main modules:
[1134] 1. Link Collection Module
[1135] The server collects links to all pages within the official website based on the starting URL. To do this, the server parses the HTML "a" tags and recursively retrieves internal links.
[1136] 2. Data Acquisition Module
[1137] The server retrieves HTML data for each page based on the collected links. Specifically, it makes an HTTP request for each URL, retrieves the HTML content of each page, and saves it to a local database.
[1138] 3. Natural Language Processing Module
[1139] The server passes the acquired HTML data to a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the relevant strings.
[1140] 4. List creation module
[1141] The server creates a list of extracted strings and their corresponding pages. This list is saved in JSON or CSV format.
[1142] 5. User Interface Module
[1143] Users can view this list through a smartphone app and make corrections as needed. The user interface is designed to display the list via a browser, allowing users to review and modify the relevant sections.
[1144] 6. Report Generation Module
[1145] The server generates a report based on the final list that reflects the user's modifications, and distributes it to relevant parties in PDF or CSV format.
[1146] Furthermore, the system of the present invention includes means for automatically analyzing and extracting specific strings corresponding to specific prompt statements, as an "application installed on a smartphone." This allows the system to operate based on a specific prompt statement, such as "Automatically identify and list the page on https: / / example.com that contains new pricing plans and campaign information," enabling it to update information quickly and accurately.
[1147] This system includes a series of processes from link collection and data acquisition to natural language processing analysis, listing, user interface verification, final list generation, and report distribution. In this way, it is possible to efficiently manage changes resulting from revisions to pricing plans and services on the official website, saving time and effort.
[1148] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1149] Step 1:
[1150] The server collects links to all pages within the official website based on the starting URL. Specifically, it accesses the starting URL and recursively retrieves internal links by parsing the HTML "a" tags. The input is the starting URL, and the output is a list of collected links.
[1151] Step 2:
[1152] The server retrieves HTML data for each page based on the collected links. It sends an HTTP request to each link to retrieve the page's HTML content. The retrieved HTML data is stored in a local database. The input is a list of links, and the output is the HTML content of each page.
[1153] Step 3:
[1154] The server processes the acquired HTML data using a natural language processing engine to extract specific strings. Based on a pre-configured keyword list (e.g., "new pricing plan", "sale", "campaign"), the NLP engine analyzes and extracts the corresponding strings. The input consists of HTML data and a keyword list, and the output is a list of the corresponding strings.
[1155] Step 4:
[1156] The server lists the extracted strings and their corresponding pages. This list is saved in JSON or CSV format. The input is information about the extracted strings and their corresponding pages, and the output is a structured data list.
[1157] Step 5:
[1158] Users review the list generated through the smartphone app and make corrections as needed. The list is displayed via a user interface, allowing users to review and modify relevant sections. The input is the generated list, and the output is the final list with user modifications.
[1159] Step 6:
[1160] The server generates a report based on the final list that reflects the user's modifications. This report is delivered in PDF or CSV format. The input is the final modified list, and the output is the final report.
[1161] Step 7:
[1162] The server operates as an application installed on a smartphone, automatically parsing and extracting specific strings corresponding to specific prompt messages. It takes a prompt message as input and outputs a list of parsed and extracted strings. For example, the parsing is performed based on the prompt message, "Automatically identify and list the pages on https: / / example.com that contain new pricing plans and campaign information."
[1163] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1164] This invention relates to a system that automatically identifies and lists relevant strings from all pages on an official website when revising pricing plans or services. Furthermore, this invention has the function of improving the accuracy of user feedback by combining it with an emotion engine that recognizes user emotions. This system enables efficient work to be carried out without outsourcing to external vendors.
[1165] This system is implemented using the following devices and programs.
[1166] System Configuration
[1167] 1. Crawling Module
[1168] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[1169] 2. Data Acquisition Module
[1170] The server retrieves HTML data for each page based on the collected links and saves the HTML data obtained for each URL via HTTP requests to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[1171] 3. Natural Language Processing (NLP) Module
[1172] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "New Pricing Plan", "Service Details", "Price Revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "New Pricing Plan" from "https: / / example.com / plan1.html", it records that text.
[1173] 4. List creation module
[1174] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[1175] 5. User Interface Module
[1176] The user views the list on their device and examines its contents. Through their browser, the user checks the items in the list (page URLs and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[1177] 6. Emotion Recognition Module
[1178] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[1179] 7. Feedback and notification module
[1180] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[1181] 8. Report Generation Module
[1182] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[1183] Examples
[1184] Specific example 2: When revising the content of a new service
[1185] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[1186] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[1187] 3. The server creates a list of pages containing the extracted text and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[1188] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[1189] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and modify lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[1190] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[1191] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the relevant department head. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the department head.
[1192] The following describes the processing flow.
[1193] Step 1:
[1194] The server sets a starting URL (e.g., "https: / / example.com") and collects links to all pages of the website. It parses HTML "a" tags and recursively collects internal links to create a list of all page URLs. For example, it visits "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[1195] Step 2:
[1196] The server retrieves HTML data for each page based on the collected list of URLs. The server sends an HTTP request for each URL and saves the retrieved HTML data to a local database. For example, it sends an HTTP request to "https: / / example.com / plan1.html" to retrieve and save the HTML data for that page.
[1197] Step 3:
[1198] The server passes the collected HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. Based on a predefined list of keywords (e.g., "new pricing plans", "service details", "price revision"), the NLP engine extracts the relevant text. For example, if the NLP engine extracts a description related to "new pricing plans" from "https: / / example.com / plan1.html", that text is recorded.
[1199] Step 4:
[1200] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, it adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[1201] Step 5:
[1202] The user reviews the list generated by the server on their device and examines its contents. Using their device's browser, the user checks each listed item (page URL and corresponding text) and makes corrections as needed. For example, if the list contains incorrect information, the user corrects it.
[1203] Step 6:
[1204] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, it can detect if the user is experiencing high levels of stress.
[1205] Step 7:
[1206] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[1207] Step 8:
[1208] The user provides feedback to the server with the list they have reviewed and corrected. The corrections are sent to the server, which then updates the list in the local database based on them. For example, if a user finds an error in "New Pricing Plans" and corrects it, the server reflects the correction and keeps the database up to date.
[1209] Step 9:
[1210] The server generates the final list and creates a report. The server outputs the final list in CSV or PDF format and distributes it to the responsible personnel in the relevant departments. For example, the final report includes links to all relevant pages and their contents, provided in a format that is easy for the responsible personnel to review.
[1211] In this way, a system is realized that efficiently extracts and lists information on revisions to pricing plans and service details, and improves the accuracy of feedback by combining this with user sentiment recognition.
[1212] (Example 2)
[1213] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1214] When revising pricing plans or service details, manually identifying and listing relevant strings from all pages of the official website is time-consuming and laborious. Therefore, a system that can perform this task efficiently and accurately is needed. Furthermore, improving the accuracy of feedback that takes user sentiment into account is also necessary.
[1215] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1216] In this invention, the server includes means for collecting links to all pages within the official website, means for obtaining HTML data for each page based on the collected links, means for extracting specific strings from the obtained HTML data using a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for providing an emotion engine that recognizes the user's emotions in real time, means for improving the accuracy of feedback and providing notifications based on the emotional state recognized by the emotion engine, and means for generating a final list including user modifications and distributing it as a report. This enables quick and accurate responses when changing pricing plans or service content, and allows for improved accuracy of feedback that takes user emotions into consideration.
[1217] An "official website" is an online information page managed by a specific company or organization.
[1218] A "link" is a reference to another page or resource within a web page, and is implemented using the HTML "a" tag.
[1219] "HTML data" is a standard markup language used to describe the structure and content of web pages.
[1220] A "natural language processing engine" is a collection of algorithms and programs designed to analyze and understand text data.
[1221] A "specific string" refers to a predefined keyword or phrase, which is the string that the system targets for extraction.
[1222] "Listing" refers to systematically organizing extracted data and saving it in list format.
[1223] An "interface" refers to the screen or input device that a user uses to interact with a system.
[1224] An "emotion engine" is a collection of software and hardware used to analyze and recognize a user's emotional state.
[1225] "Feedback" refers to collecting opinions and evaluations from users and using them to improve and adjust the system.
[1226] A "report" is a document that summarizes the data and results that have been collected and analyzed.
[1227] This invention is a system that automatically identifies and lists relevant strings from each page of an official website when price plans or service details are revised. Furthermore, it incorporates an emotion engine to improve the accuracy of user feedback. This system provides a means to perform tasks efficiently and can execute processing automatically without outsourcing to external vendors.
[1228] System Configuration
[1229] 1. Crawling module:
[1230] The server sets a starting URL and crawls the entire site from this URL. It collects internal links by parsing HTML "a" tags and recursively retrieves all internal links. For example, the server sets "https: / / example.com" as the starting URL, accesses "https: / / example.com / index.html", parses the internal links, and adds links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" to the list.
[1231] 2. Data Acquisition Module:
[1232] The server retrieves HTML data for each page based on the collected links. It uses HTTP requests to access each link and saves the retrieved HTML data to a local database. For example, the server sends an HTTP request to "https: / / example.com / plan1.html", retrieves its HTML data, and saves it.
[1233] 3. Natural Language Processing (NLP) Module:
[1234] The server passes the acquired HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. Based on a predefined keyword list, the NLP engine extracts the relevant text. For example, the server extracts descriptions related to "new pricing plans" from "https: / / example.com / plan1.html".
[1235] 4. List creation module:
[1236] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format such as JSON or CSV. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[1237] 5. User Interface Module:
[1238] The user reviews the generated list on their device and scrutinizes its contents. They check the list items through their browser and make corrections as needed. For example, if the user finds incorrect information in the list, they correct it.
[1239] 6. Emotion Recognition Module:
[1240] A built-in emotion engine in the device recognizes the user's emotions in real time as they review and edit lists. It analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate their emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[1241] 7. Feedback and notification module:
[1242] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[1243] 8. Report generation module:
[1244] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant department heads. For example, the final report includes links and content for all relevant pages and is provided in a format that allows the department heads to easily review it.
[1245] Specific example
[1246] When the new service content is revised
[1247] 1. The server crawls all pages of the official website and collects all links within "https: / / example.com". At the same time, it also retrieves the HTML data for each page.
[1248] 2. The server processes the collected HTML data using an NLP engine to extract text that matches keywords such as "new service details." For example, if "https: / / example.com / service1.html" contains the phrase "new service details," that section will be extracted.
[1249] 3. The server lists the pages containing the extracted text and the text itself, and saves it in JSON format. This list includes the URL of each page and the corresponding section within that page.
[1250] 4. The user reviews the generated list through their browser and makes corrections as needed. For example, if a string is incorrectly extracted, the user corrects it.
[1251] 5. The emotion engine built into the device recognizes the user's emotions in real time as they review and modify lists, and notifies the user if they are experiencing high levels of stress or dissatisfaction.
[1252] 6. The server uses the emotion engine's recognition results to help improve the accuracy of user feedback.
[1253] 7. The server generates a final list reflecting the user's modifications, creates a report based on it, and distributes it to the responsible person in charge of the relevant department. This report includes links and content for all relevant pages and is provided in a format that is easily accessible to the responsible person.
[1254] Example of a prompt
[1255] Please tell me how to design a system that automatically identifies and lists descriptions related to "new pricing plans" and "service details" from all pages on the official website when revising new service content, and how to improve the accuracy of feedback by recognizing user sentiment in real time.
[1256] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1257] Step 1:
[1258] Link crawling
[1259] The server sets a starting URL and crawls the entire site from this URL. The input is the starting URL, and the server parses the HTML "a" tags and recursively collects all internal links on all pages within the site. The output is a list of all links. For example, if the server sets "https: / / example.com" as the starting URL and accesses "https: / / example.com / index.html", it will discover internal links such as "https: / / example.com / plan1.html" and "https: / / example.com / plan2.html" and add them to the list.
[1260] Step 2:
[1261] Retrieving and saving HTML data
[1262] The server retrieves HTML data for each page based on the collected links. The input is a list of collected links; the server accesses each link using an HTTP request and saves the retrieved HTML data to a local database. The output is the HTML data for each page. For example, the server sends an HTTP request to "https: / / example.com / plan1.html" and retrieves and saves the HTML data for that page.
[1263] Step 3:
[1264] Text analysis using Natural Language Processing (NLP)
[1265] The server passes the retrieved HTML data to the NLP engine, which analyzes keywords related to pricing plans and services. The input consists of saved HTML data and a predefined keyword list, and the NLP engine extracts the relevant text. The output is the extracted specific string. For example, the server extracts the description of "new pricing plan" from "https: / / example.com / plan1.html".
[1266] Step 4:
[1267] Creating and saving lists
[1268] The server creates a list of corresponding pages and text based on the extracted text. The input is the extracted text and the corresponding page URL, and the list is saved in a structured data format such as JSON or CSV. The output is structured list data. For example, the server adds "https: / / example.com / plan1.html" and the text about the "new pricing plan" contained on that page to the list.
[1269] Step 5:
[1270] Check and correct the list
[1271] The user reviews the generated list on their device and scrutinizes its contents. The input is the generated list; the user checks the list items (page URLs and corresponding text) through their browser and makes corrections as needed. The output is the corrected list. For example, if the user finds incorrect information in the list, they correct it.
[1272] Step 6:
[1273] emotion recognition
[1274] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Input consists of user operation data and biometric data (facial expressions, voice tone, typing speed), which the engine analyzes and evaluates the emotional state. The output is the emotion recognition result. For example, if the user is experiencing high levels of stress, it will be detected.
[1275] Step 7:
[1276] Feedback and notifications
[1277] The server uses the user's emotional state, recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the emotion recognition result, which the server uses to provide additional information or notifications. The output is information on improving the accuracy of feedback and notification content. For example, if a user expresses strong dissatisfaction, the server will provide additional information to improve the accuracy of the correction list.
[1278] Step 8:
[1279] Report generation and distribution
[1280] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. Inputs are the corrected list and sentiment recognition results; the final list is output in CSV or PDF format and distributed to the relevant department heads. The output is the final report. For example, the final report includes links and content for all relevant pages and is provided in a format easily accessible to the department heads.
[1281] (Application Example 2)
[1282] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1283] When revising pricing plans or services on an official website, the process of identifying and listing relevant strings from all pages is time-consuming and often outsourced to external vendors, leading to increased costs. Furthermore, the low accuracy of user feedback is a significant challenge.
[1284] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for collecting links to all pages of the official website, means for obtaining markup language data for each page based on the collected links, means for extracting specific strings by running the obtained markup language data through a natural language processing engine, means for listing corresponding pages and strings based on the extracted strings, means for providing an interface that allows the user to modify and confirm the generated list, means for improving the accuracy of feedback using an emotion engine that recognizes the user's emotions, and means for generating a final list including user modifications and distributing it as a report. This makes it possible to efficiently identify and list specific strings from all pages of the official website and improve the accuracy of user feedback.
[1285] An "official website" is a collection of publicly available information on the internet provided by companies, organizations, etc.
[1286] A "link" refers to a hypertext connection between web pages, an element that allows users to navigate to other web pages by clicking on it.
[1287] "Markup language data" refers to text data that includes a set of tags that define the structure and content of a web page, and is written in formats such as HTML and XML.
[1288] A "natural language processing engine" is a software module designed to understand and analyze human language, and it has the function of extracting specific information from language data.
[1289] A "string" refers to a sequence of consecutive characters and constitutes a part of text data.
[1290] "Listing" refers to the process of organizing data into a list format based on certain rules.
[1291] An "interface" is the boundary or mechanism by which humans and computer systems exchange information with each other.
[1292] An "emotion engine" is a software module that recognizes a person's emotional state in real time based on their facial expressions, tone of voice, typing speed, and other factors.
[1293] The "final list" refers to the final, confirmed list of data after incorporating user modifications.
[1294] A "report" is a document that compiles collected and analyzed information, intended to be presented to relevant parties for specific purposes.
[1295] This invention relates to a system for automatically collecting information on pricing plans and service revisions on official websites and for improving the accuracy of user feedback. This system is configured as follows:
[1296] 1. Crawling Module
[1297] The server sets the official website's starting URL and collects links to all pages within the site. This involves parsing HTML "a" tags and recursively retrieving internal links. For example, it visits "https: / / example.com" and executes a process to collect its internal links.
[1298] 2. Data Acquisition Module
[1299] The server retrieves the markup language data (HTML data) for each page based on the collected links. The retrieved HTML data is obtained from each URL using HTTP requests. For example, an HTTP request is sent to "https: / / example.com / plan1.html" to retrieve the HTML data for that page.
[1300] 3. Natural Language Processing Engine
[1301] The server passes the acquired HTML data to a natural language processing (NLP) engine, which analyzes keywords related to pricing plans and services. For example, based on a list of keywords such as "new pricing plan," "service details," and "price revision," the NLP engine extracts the relevant text.
[1302] 4. List creation module
[1303] The server creates a list of corresponding pages and text based on the extracted text. The list is saved in a structured data format (JSON or CSV). For example, it adds "https: / / example.com / plan1.html" and the text related to "New Pricing Plans" contained on that page to the list.
[1304] 5. User Interface Module
[1305] Users can view and modify the generated list through their browser. For example, if the list contains incorrect information, the user can correct it in their browser.
[1306] 6. Emotion Recognition Module
[1307] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. This emotion engine analyzes the user's facial expressions, tone of voice, typing speed, etc., to evaluate the user's current emotional state. For example, if the user is experiencing high levels of stress, it will detect that.
[1308] 7. Feedback and notification module
[1309] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. For example, if a user expresses strong dissatisfaction, the server uses that information to provide additional information to improve the accuracy of the correction list.
[1310] 8. Report Generation Module
[1311] The server generates a final list and creates a report based on user corrections and sentiment recognition results. The final list is output in CSV or PDF format and distributed to the relevant departments. For example, the final report includes links and content for all relevant pages and is provided in a format that can be easily reviewed by the relevant departments.
[1312] Specific example:
[1313] This system is used by e-commerce site administrators when revising prices for new products. The administrator crawls the entire site to identify pages and text related to price changes. They then review the list and make corrections, and if the sentiment engine detects stress on the administrator, it automatically provides guidelines and support.
[1314] Example of a prompt:
[1315] "Identify the relevant page and text regarding the price revision for the new product. Recognize user sentiment in real time to improve the accuracy of feedback using an emotion engine."
[1316] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1317] Step 1:
[1318] The server sets the official website's starting URL and collects links to all pages within the site. The input is the starting URL, and the output is a list of links to all pages. Specifically, the server parses HTML "a" tags and recursively retrieves internal links. This process allows the server to understand the overall page structure of the site.
[1319] Step 2:
[1320] The server retrieves the markup language data (HTML data) for each page based on the collected links. The input is a list of links, and the output is the HTML data for each page. Specifically, the server uses HTTP requests to retrieve the linked web pages and stores their HTML data locally.
[1321] Step 3:
[1322] The server processes the acquired HTML data using a natural language processing (NLP) engine to analyze keywords related to pricing plans and services. The input is HTML data, and the output is the corresponding text. Specifically, the NLP engine extracts the relevant text based on a predefined list of keywords (e.g., "new pricing plan," "service details," "price revision").
[1323] Step 4:
[1324] The server creates a list of corresponding pages and texts based on the extracted text. The input is the text and the data of the corresponding page, and the output is a list in a structured data format (JSON or CSV). Specifically, the extracted text and its page URL are combined, added to the list, and saved.
[1325] Step 5:
[1326] The user reviews the generated list through a browser and modifies its contents. The input is a structured data list, and the output is the user-modified list. Specifically, the user checks the list contents using a web browser and corrects any inaccuracies.
[1327] Step 6:
[1328] The emotion engine built into the device recognizes the user's emotions in real time as they review and edit lists. Inputs include the user's facial expressions, tone of voice, and typing speed, while output is an evaluation of the user's emotions. Specifically, it monitors the user's emotional state using a camera and microphone and analyzes that state.
[1329] Step 7:
[1330] The server uses the user's emotional state, as recognized by the emotion engine, to improve the accuracy of feedback and provide notifications. The input is the user's emotional evaluation, and the output is suggestions for improving feedback accuracy and notifications. Specifically, if high levels of stress or dissatisfaction are detected, the server provides additional guidance and support to help resolve the issue.
[1331] Step 8:
[1332] The server generates a final list and creates a report based on user-submitted corrections and sentiment recognition results. The input is the user-modified list and sentiment recognition results, and the output is the final report. Specifically, the final list is output in CSV or PDF format and distributed to the relevant departments.
[1333] This series of processing steps allows for the efficient identification and listing of specific strings from all pages of the official website, thereby improving the accuracy of user feedback.
[1334] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1335] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1336] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1337] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1338] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1339] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1340] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1341] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1342] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1343] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1344] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1345] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1346] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1347] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1348] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1349] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1350] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1351] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1352] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1353] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1354] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1355] The following is further disclosed regarding the embodiments described above.
[1356] (Claim 1)
[1357] A means of collecting links to all pages on the official website,
[1358] A means of obtaining the HTML data of each page based on the collected links,
[1359] A method for extracting specific strings from acquired HTML data using a natural language processing engine,
[1360] A method for listing the corresponding pages and strings based on the extracted strings,
[1361] A means of providing an interface that allows the user to modify and verify the generated list,
[1362] A method for generating a final list including user modifications, compiling it into a report, and distributing it,
[1363] A system that includes this.
[1364] (Claim 2)
[1365] The system according to claim 1, characterized in that the user uses an interface through a browser that provides a link to a page containing the extracted string.
[1366] (Claim 3)
[1367] The system according to claim 1, characterized in that a natural language processing engine extracts a string based on a pre-configured keyword list.
[1368] "Example 1"
[1369] (Claim 1)
[1370] A means of collecting links to all pages on the official website,
[1371] A means of obtaining the HTML data of each page based on the collected links,
[1372] A method for extracting specific strings from acquired HTML data using a natural language processing engine,
[1373] A method for listing the corresponding pages and strings based on the extracted strings,
[1374] A means of providing an interface that allows the user to modify and verify the generated list,
[1375] A method for generating a final list including user modifications, compiling it into a report, and distributing it,
[1376] A means of saving data to a local database,
[1377] A means of generating reports in PDF or CSV format,
[1378] A system that includes this.
[1379] (Claim 2)
[1380] The system according to claim 1, characterized in that the user uses an interface through a browser that provides a link to a page containing the extracted string.
[1381] (Claim 3)
[1382] The system according to claim 1, characterized in that a natural language processing engine extracts a string based on a pre-configured keyword list.
[1383] "Application Example 1"
[1384] (Claim 1)
[1385] A means of collecting links to all pages on the official website,
[1386] A means of obtaining the HTML data of each page based on the collected links,
[1387] A method for extracting specific strings from acquired HTML data using a natural language processing engine,
[1388] A method for listing the corresponding pages and strings based on the extracted strings,
[1389] A means of providing an interface that allows the user to modify and verify the generated list,
[1390] A method for generating a final list including user modifications, compiling it into a report, and distributing it,
[1391] A means of automatically parsing and extracting a specific string corresponding to a specific prompt message, which is executed as an application installed on a smartphone.
[1392] A system that includes this.
[1393] (Claim 2)
[1394] The system according to claim 1, characterized in that the user accesses an interface that provides a link to a page containing the extracted string through a smartphone browser.
[1395] (Claim 3)
[1396] The system according to claim 1, characterized in that a natural language processing engine extracts a string based on a pre-configured keyword list.
[1397] "Example 2 of combining an emotion engine"
[1398] (Claim 1)
[1399] A means of collecting links to all pages on the official website,
[1400] A means of obtaining the HTML data of each page based on the collected links,
[1401] A method for extracting specific strings from acquired HTML data using a natural language processing engine,
[1402] A method for listing the corresponding pages and strings based on the extracted strings,
[1403] A means of providing an interface that allows the user to modify and verify the generated list,
[1404] A means to provide an emotion engine that recognizes user emotions in real time,
[1405] A means of improving the accuracy of feedback and providing notifications based on the emotional state recognized by the emotion engine,
[1406] A method for generating a final list including user modifications, compiling it into a report, and distributing it,
[1407] A system that includes this.
[1408] (Claim 2)
[1409] The system according to claim 1, characterized in that the user uses an interface through a browser that provides a link to a page containing the extracted string.
[1410] (Claim 3)
[1411] The system according to claim 1, characterized in that a natural language processing engine extracts a string based on a pre-configured keyword list.
[1412] "Application example 2 when combining with an emotional engine"
[1413] (Claim 1)
[1414] A means of collecting links to all pages of the official website,
[1415] A means of obtaining markup language data for each page based on the collected links,
[1416] A method for extracting specific strings by running the acquired markup language data through a natural language processing engine,
[1417] A method for listing the corresponding pages and strings based on the extracted strings,
[1418] A means of providing an interface that allows the user to modify and verify the generated list,
[1419] A means to improve the accuracy of feedback using an emotion engine that recognizes user emotions,
[1420] A method for generating a final list including user modifications, compiling it into a report, and distributing it,
[1421] A system that includes this.
[1422] (Claim 2)
[1423] The system according to claim 1, characterized in that the user uses a browsing software to access an interface that provides a link to a page containing the extracted string.
[1424] (Claim 3)
[1425] The system according to claim 1, characterized in that a natural language processing engine extracts a string based on a pre-configured keyword list. [Explanation of Symbols]
[1426] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting links to all pages on the official website, A means of obtaining the HTML data of each page based on the collected links, A method for extracting specific strings from acquired HTML data using a natural language processing engine, A method for listing the corresponding pages and strings based on the extracted strings, A means of providing an interface that allows the user to modify and verify the generated list, A method for generating a final list including user modifications, compiling it into a report, and distributing it, A system that includes this.
2. The system according to claim 1, characterized in that the user uses an interface through a browser that provides a link to a page containing the extracted string.
3. The system according to claim 1, characterized in that a natural language processing engine extracts a string based on a pre-configured keyword list.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A