System
The system addresses the challenge of inappropriate and complex internet content for children by converting and filtering content using generative AI and furigana, ensuring safe and educational online experiences.
Patent Information
- Application Number
- JP2024123911
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Existing internet content is often inappropriate and difficult for young children to understand, posing a challenge in providing safe and educational content for kindergarten and elementary school children, with a lack of automated conversion and filtering technologies.
A system that analyzes internet content, uses a generative AI model to convert content according to the user's developmental level, adds furigana, filters inappropriate content, and displays the converted content in an integrated manner, utilizing algorithms and programs for content analysis, transformation, and filtering.
Enables children to access safe and easy-to-understand internet content, promoting learning and protecting them from inappropriate material.
Smart Images

Figure 2026022394000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The majority of content on the Internet is aimed at average adults and is difficult for young people, especially kindergarten and elementary school children, to understand. As a result, when children use the Internet, there is a lack of content at an appropriate level, which hinders their understanding of information and learning, creating an issue. In addition, the risk of inappropriate content appearing on the Internet increases, making it difficult to let children use the Internet safely. [Means for solving the problem]
[0005] The present invention solves the aforementioned problems with a system that includes a means for analyzing acquired internet content, a means for using a generative AI model to convert the content according to the user's developmental level, and a means for displaying the converted content. Specifically, the system includes a means for adding furigana to the acquired content, a means for converting the content's expressions into simpler terms, and a means for filtering inappropriate content. The system also provides a means for analyzing acquired internet content, extracting parts that match specific keywords, converting the extracted parts according to the user's developmental level, and displaying the converted parts in an integrated manner, thereby enabling children to use the internet safely in a way that is easy to understand.
[0006] "Internet content" is any information or data accessible via the Internet, such as web pages, blog posts, videos, images, etc.
[0007] "Means for analyzing" refers to algorithms, databases, and programs used to analyze captured content and understand its structure and content.
[0008] A "generative AI model" is a machine learning model that uses artificial intelligence to automatically generate or convert text and data.
[0009] "User developmental level" refers to the educational and cognitive development stage of the user according to their age and level of understanding.
[0010] "Transformation means" refers to a process or program that changes the original content into a form appropriate for the user's developmental level.
[0011] "Means for adding furigana" refers to algorithms or programs for automatically adding furigana that indicates how to read kanji characters.
[0012] "Means of converting into simple language" refers to algorithms or programs that replace difficult expressions or technical terms with simple expressions that are easy for an average child to understand.
[0013] "Filtering measures" refer to algorithms or programs that filter out inappropriate content and display only safe and appropriate information.
[0014] "Means for extracting parts that match specific keywords" refers to algorithms or programs that automatically search and extract parts of content that match specified keywords.
[0015] "Displaying means" refers to a web browser or application that allows a user to view the converted content. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[0038] Server-side processing
[0039] The server plays an important role in analyzing the acquired Internet content and converting it according to the user's level of development. Specifically, the process is as follows:
[0040] 1. Content Acquisition:
[0041] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and sends an HTTP GET request to download the entire web page.
[0042] 2. Content Analysis:
[0043] To parse the retrieved HTML data, the server uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[0044] 3. Transformation by generative AI models:
[0045] After analyzing the content, the server uses a generative AI model to convert it to suit the user's developmental level. Specifically, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions while adding furigana to kanji. This process utilizes a generative AI model (e.g., GPT-3).
[0046] 4. Filtering inappropriate content:
[0047] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[0048] 5. Return of Converted Content:
[0049] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[0050] Terminal side processing
[0051] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0052] 1. Submit your request:
[0053] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0054] 2. Receiving Content:
[0055] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[0056] 3. Viewing Content:
[0057] The received content is displayed in the browser in a format that is easy for children to understand, such as sentences that make extensive use of hiragana and kanji with furigana.
[0058] User usage scenarios
[0059] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[0060] 1. Initiating access:
[0061] The user opens a browser and enters the URL for "Animal World" to access the site.
[0062] 2. Conversion by the server:
[0063] The device sends a request for content to the server, which retrieves, analyzes, and converts it.
[0064] 3. Viewing Content:
[0065] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[0066] As such, the present invention is a system that provides appropriate Internet usage according to the user's developmental level, and aims to promote children's learning and information acquisition while protecting them from inappropriate content.
[0067] The processing flow will be explained below.
[0068] Step 1:
[0069] A user launches a browser and attempts to access a particular website.
[0070] Step 2:
[0071] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[0072] Step 3:
[0073] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[0074] Step 4:
[0075] Parse the HTML data retrieved by the server using an HTML parser (such as BeautifulSoup) to understand the overall structure of the web page and extract the main content elements (text, images, links, etc.).
[0076] Step 5:
[0077] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level. Specifically, for kindergarteners, the text is converted into hiragana and simple words are used, while for elementary school students, kanji with furigana is used, making it easier to understand.
[0078] Step 6:
[0079] The server filters inappropriate content during the conversion process, using a filtering algorithm to detect inappropriate keywords and phrases and remove them or replace them with harmless alternatives.
[0080] Step 7:
[0081] The server returns the converted content to the device, and sends an HTTP response containing the converted HTML data and other resources.
[0082] Step 8:
[0083] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[0084] Step 9:
[0085] The device then displays the converted content to the user, rendering the web page in a way that is easy for the user to understand, such as displaying sentences that make heavy use of hiragana or kanji with furigana.
[0086] Example 1
[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0088] There is a wide variety of content on the Internet, and understanding and safety of that content poses challenges, especially when children access it. Current technology does not adequately convert content or filter inappropriate content, making it difficult to provide information in a format that is easy for children to understand. Furthermore, the need for manual content filtering and the lack of automated technology to convert content into appropriate expressions according to the user's developmental level result in an inconsistent provision of educational value.
[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0090] In this invention, the server includes means for acquiring accessed internet data, means for analyzing the acquired HTML data and extracting key elements, means for converting the content according to the user's developmental level using a generative AI model, means for filtering the converted content, and means for returning the converted and filtered content to the terminal. This makes it possible to automatically analyze, convert, and filter internet content accessed by children and provide appropriate information according to their developmental level.
[0091] "Accessed Internet Data" is information obtained from the web pages or online resources that a user attempts to access.
[0092] "HTML data" means data written in the HyperText Markup Language used to describe the structure and content of web pages.
[0093] "Key elements" are the core content of a Web page, such as text, images, links, and headings.
[0094] A "generative AI model" is an artificial intelligence algorithm that generates appropriate content based on given input, particularly one based on large-scale natural language processing techniques.
[0095] "User's developmental level" refers to the level of content appropriate for the user's age and learning stage.
[0096] The "means of transforming content" refers to the process of using the generative AI model described above to transform the original content into an appropriate format according to the user's developmental level.
[0097] "Content filtering measures" are the processes that detect inappropriate content or keywords and remove or replace them with harmless forms.
[0098] A "terminal" is an electronic device that a user directly uses as an interface, such as a personal computer, smartphone, or tablet.
[0099] "Means for returning content to the terminal" refers to the process by which the server sends the processed content to the user's terminal.
[0100] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[0101] Server-side processing
[0102] The server retrieves and analyzes data from the internet, uses generative AI models to transform content according to the user's developmental level, and filters out inappropriate content.
[0103] The server retrieves the accessed data on the Internet. Specifically, it receives the URL of the web page the user is trying to access, sends an HTTP GET request to that URL, and retrieves the HTML data. The server uses an HTML parser such as BeautifulSoup to analyze this data, thereby extracting the main elements of the web page (text, images, links, etc.).
[0104] It then uses a generative AI model (e.g., GPT-3) to adapt the parsed content to the user's developmental level, using the following example prompt:
[0105] "Simplify the text on the web page for children. Enter the text below: '{{Web page text}}'"
[0106] "Parse this HTML data and convert the extracted key content elements to child-friendly. HTML data: '{{HTML data}}'"
[0107] The content converted by the generative AI model then goes through a further filtering process: the server detects inappropriate keywords and phrases based on a list and either removes them or replaces them with harmless language.
[0108] After all the processing is complete, the server returns the converted and filtered content to the device, ensuring that the content is presented in a child-friendly and safe manner.
[0109] Terminal side processing
[0110] The terminal is responsible for appropriately displaying the converted content received from the server to the user.
[0111] When a user accesses a particular web page through a browser, the device sends a request to the server, which includes the URL they are trying to access. The server returns the converted content, which the device then displays in the browser.
[0112] For example, if an elementary school student wants to visit a website called "Animal World," the process would go something like this: The user opens a browser and enters the URL for "Animal World." The device sends a request to the server, which retrieves, analyzes, converts, and filters the content. Finally, the device receives the converted content and displays it appropriately to the user.
[0113] User usage scenarios
[0114] Users operate a browser through the system to access specific web pages. For example, on the "Animal World" website, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." This system allows children to obtain appropriate information safely and in an easy-to-understand manner.
[0115] As described above, the present invention aims to provide appropriate Internet usage according to the user's developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[0116] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0117] Server-side processing
[0118] Step 1:
[0119] The server receives the URL of the web page the user is trying to access, which is obtained by an HTTP GET request sent from the device.
[0120] Input: The URL the user types in the browser
[0121] Output: Retrieved URL
[0122] Step 2:
[0123] The server sends an HTTP GET request to the URL and downloads the HTML data of the entire web page. The server issues the request, obtains the response, and saves it.
[0124] Input: Retrieved URL
[0125] Output: Downloaded HTML data
[0126] Step 3:
[0127] To parse the HTML data, the server uses an HTML parser such as BeautifulSoup to parse the data. The server creates a BeautifulSoup object and parses the HTML data to extract the main elements (text, images, links, etc.).
[0128] Input: Downloaded HTML data
[0129] Output: Extracted key elements (text, images, links, etc.)
[0130] Step 4:
[0131] The server sends prompts to a generative AI model (e.g., GPT-3) and converts the analyzed content according to the user's development level. The server constructs an appropriate prompt sentence for the text portion of the content and sends it to the AI model.
[0132] Input: extracted key elements and user's development level
[0133] Output: The converted text content
[0134] Step 5:
[0135] The server scans the converted content and filters out inappropriate content. Using a list of inappropriate keywords, the server detects these words and phrases and either replaces them with appropriate alternatives or removes them.
[0136] Input: The converted text content
[0137] Output: The filtered text content
[0138] Step 6:
[0139] The server returns the content after all conversion and filtering has been completed to the terminal as an HTTP response, so that the converted content can be displayed appropriately to the user.
[0140] Input: Filtered text content
[0141] Output: The transformed content included in the HTTP response
[0142] Terminal side processing
[0143] Step 1:
[0144] When a user accesses a specific web page through a browser, the device sends the request to the server. The browser issues an HTTP GET request, sending the URL that the user wants to access to the server.
[0145] Input: The URL the user types in the browser
[0146] Output: HTTP GET request to the server
[0147] Step 2:
[0148] The terminal receives the converted content returned from the server, receives the HTTP response, and analyzes and obtains the converted HTML data contained in the response body.
[0149] Input: HTTP response from the server
[0150] Output: Converted HTML data
[0151] Step 3:
[0152] The device displays the received content in the browser, and uses the browser's rendering engine to display the converted content in a format that is easy for children to understand.
[0153] Input: Converted HTML data
[0154] Output: Content displayed in the browser
[0155] User usage scenarios
[0156] 1. Initiating Access
[0157] A user opens a browser and enters the URL of the website they want to visit.
[0158] Input: The URL entered by the user
[0159] Output: The browser issues an access request
[0160] 2. Conversion by the server
[0161] The device sends a request for content to the server, which retrieves, analyzes, converts, and filters it.
[0162] Input: The URL the user types into the browser
[0163] Output: The transformed and filtered content
[0164] 3. Display of Content
[0165] The device receives the converted content and displays it appropriately in the browser. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[0166] Input: The converted content sent from the server
[0167] Output: Content displayed in the browser
[0168] The above are the specific processing steps of the system based on the present invention, which enables the conversion of Internet content according to the user's level of development, filtering out inappropriate content, and providing users with safe and easy-to-understand information.
[0169] (Application example 1)
[0170] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0171] When children browse content online, it is difficult to provide them with content that is converted into an appropriate format according to their age and level of understanding. Furthermore, content viewed online may contain inappropriate information, and there is a lack of means to accurately filter and safely provide this information. Furthermore, there is a need to improve learning outcomes by appropriately adjusting the display of content according to the child's developmental level.
[0172] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0173] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for generating a unique prompt sentence and inputting it to the generative AI model, and means for filtering inappropriate content, thereby enabling the content to be converted into a format that is easy for children to understand and displayed safely.
[0174] "Internet content" refers to all information media, such as text, images, video, and audio, that can be obtained through the Internet.
[0175] "Means of analysis" refers to technical methods for structurally understanding acquired Internet content and extracting each element.
[0176] A "generative AI model" is a collection of machine learning algorithms that use artificial intelligence techniques to generate or transform text or images.
[0177] "Means for conversion according to the user's developmental level" refers to technology for converting content into an appropriate format based on the user's age and level of understanding.
[0178] The "means for displaying the converted content" refers to a device or software for appropriately outputting the converted content to the user's terminal.
[0179] "Means for generating unique prompt sentences and inputting them into a generative AI model" refers to a technical method for creating prompt sentences in a form suitable for a generative AI model and providing them to the model.
[0180] "Inappropriate content filtering measures" are technologies that detect inappropriate keywords and phrases from internet content and either remove them or replace them with harmless ones.
[0181] A "means for adding furigana" is a technique or device for adding readings to difficult characters such as kanji.
[0182] "Means of converting expressions into simple terms" are techniques for converting complex sentences and technical terms into simple terms that children can easily understand.
[0183] "Means for extracting parts that match specific keywords" refers to technology for detecting and extracting specific words or phrases within content.
[0184] "Means for integrating and displaying" refers to the technology or device that compiles and presents the analyzed, transformed, and filtered content to the user.
[0185] This invention is a system that converts content on the Internet according to the user's level of development when the user browses the content, and displays the content safely. This system is composed of a server and a terminal.
[0186] Server-side processing
[0187] 1. Content Acquisition
[0188] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data for that URL, which downloads the entire web page data to the server.
[0189] 2. Content Analysis
[0190] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[0191] 3. Conversion using generative AI models
[0192] The server uses a generative AI model (e.g., GPT-3) based on the analyzed content to convert the content to suit the user's developmental level. The conversion is performed by generating appropriate prompts and providing them as input to the AI model. For example, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions, with furigana added to kanji.
[0193] 4. Filtering inappropriate content
[0194] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[0195] 5. Return of Converted Content
[0196] After all conversion and filtering is complete, the server returns the converted content to the terminal.
[0197] Terminal side processing
[0198] 1. Submitting a Request
[0199] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0200] 2. Receiving Content
[0201] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[0202] 3. Display of Content
[0203] The device displays the received content in a way that is easy for children to understand, such as sentences using simple words or text with furigana added to kanji characters.
[0204] Content conversion examples
[0205] For example, if an elementary school student wants to browse a website called "Animal World," using this system, the following flow will be displayed: The sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals," making it easier for elementary school students to understand.
[0206] Hardware and software used
[0207] Hardware: Smartphone (device)
[0208] Software: Python, requests library, BeautifulSoup library, OpenAI GPT-3 API
[0209] Prompt Sentence Examples
[0210] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[0211] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[0212] In this way, the present invention aims to provide users with access to Internet content appropriate to their developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[0213] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0214] Step 1:
[0215] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL into the web browser and clicks the access button. At this time, the device sends an HTTP request to the server, along with information about the user's development level. The input is the URL and development level information entered by the user, and the output is the HTTP request sent to the server.
[0216] Step 2:
[0217] The server sends an HTTP GET request to retrieve the HTML data of the received URL. Specifically, the server downloads the entire web page using the specified URL. At this time, the server receives the HTML content of the web page as a response to the request. The input is the URL, and the output is the retrieved HTML data.
[0218] Step 3:
[0219] The server parses the retrieved HTML data using an HTML parser such as BeautifulSoup. Specifically, it extracts the main content elements such as text, images, and links from the HTML data. The input is the retrieved HTML data, and the output is the extracted content elements.
[0220] Step 4:
[0221] The server uses a generative AI model (e.g., GPT-3) to generate a prompt appropriate for the user's developmental level, and then uses the prompt to convert the content. Specifically, it generates an appropriate prompt based on the analyzed content and inputs it into the generative AI model to convert the content into a format appropriate for the user's developmental level. The input is the prompt and the analyzed content, and the output is the converted content. An example of a prompt is as follows:
[0222] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[0223] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[0224] Step 5:
[0225] The server filters the transformed content to remove or replace inappropriate content, specifically detecting inappropriate keywords or phrases and replacing them with harmless terms. The input is the transformed content, and the output is the filtered content.
[0226] Step 6:
[0227] The server returns the filtered content to the terminal. Specifically, it sends the filtered and converted content data to the terminal as an HTTP response. The input is the filtered content, and the output is the data returned to the terminal.
[0228] Step 7:
[0229] The device displays the converted content received from the server to the user. Specifically, it renders the converted content on the browser and displays it in a format that is easy for children to understand. The input is the content data returned from the server, and the output is the content displayed to the user.
[0230] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0231] A specific system for implementing the present invention is composed of elements of a server, a terminal, a user, and an emotion engine.
[0232] Server-side processing
[0233] The server plays an important role in analyzing the acquired internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are explained below.
[0234] 1. Content Acquisition:
[0235] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[0236] 2. Content Analysis:
[0237] To parse the retrieved HTML data, the server uses an HTML parser to extract the main content elements of the web page, such as text, images, and links.
[0238] 3. Transformation by generative AI models:
[0239] The server analyzes the content and then uses a generative AI model to adapt it to the user's developmental level. Specifically, for kindergarteners, it converts sentences into hiragana and uses simple words, while for elementary school students, it uses kanji with furigana and adapts it to easier-to-understand expressions.
[0240] 4. Emotion Recognition with Emotion Engine:
[0241] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text to recognize the emotion at that time.
[0242] 5. Emotion-based expression adjustment:
[0243] Depending on the emotion recognized, the generative AI model can further adjust the content it outputs. For example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[0244] 6. Filtering inappropriate content:
[0245] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[0246] 7. Return of Converted Content:
[0247] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[0248] Terminal side processing
[0249] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0250] 1. Submit your request:
[0251] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0252] 2. Receiving Content:
[0253] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[0254] 3. Viewing Content:
[0255] The received content is displayed in the browser in a format that is easy for children to understand, such as sentences written in simple language or kanji with furigana.
[0256] User usage scenarios
[0257] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[0258] 1. Initiating access:
[0259] The user opens a browser and enters the URL for "Animal World" to access the site.
[0260] 2. Conversion by the server:
[0261] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[0262] 3. Viewing Content:
[0263] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[0264] As such, the present invention is a system that provides an appropriate Internet usage environment according to the user's developmental level and emotions, and aims to support children's learning and information acquisition, as well as provide a comfortable and secure content experience.
[0265] The processing flow will be explained below.
[0266] Step 1:
[0267] A user launches a browser and attempts to access a particular website (e.g., "Animal World").
[0268] Step 2:
[0269] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[0270] Step 3:
[0271] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[0272] Step 4:
[0273] The server parses the HTML data it retrieves and uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[0274] Step 5:
[0275] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level (kindergartener, elementary school student, etc.) Specifically, it converts sentences into hiragana, uses kanji with furigana, or modifies them into simpler words.
[0276] Step 6:
[0277] The server simultaneously recognizes the user's emotions using an emotion engine, which analyzes facial expressions and tone of voice through a camera and microphone to assess the user's emotional state (e.g., joy, anxiety, excitement, etc.).
[0278] Step 7:
[0279] The server further adjusts the converted content based on the emotions recognized by the emotion engine, for example, changing the text to a more gentle expression and adding an encouraging message if the user is anxious.
[0280] Step 8:
[0281] The server uses filtering algorithms to detect inappropriate keywords and phrases in the content and either remove them or replace them with harmless alternatives.
[0282] Step 9:
[0283] The server sends the converted and filtered content back to the device, which then sends an HTTP response containing the converted HTML data and other resources.
[0284] Step 10:
[0285] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[0286] Step 11:
[0287] The device displays the converted content in the browser. For example, the sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" will be added.
[0288] Through this tailored content, users can access information that is easy to understand and appropriate for their developmental level, and enjoy a comfortable internet experience that is sensitive to their emotions.
[0289] Example 2
[0290] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0291] With conventional internet usage, it has been difficult to display content that takes into account the user's developmental level and emotions, making it difficult to provide appropriate information to children. Furthermore, filtering of inappropriate content is insufficient, creating a need for a safe internet usage environment.
[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0293] In this invention, the server includes a means for analyzing acquired internet content, a means for using a generative AI model to convert the content according to the user's developmental level, and a means for recognizing the user's emotions and adjusting the presentation of the content based on those emotions. This enables the provision of appropriate and safe information according to the user's developmental level and emotions. Furthermore, by including a means for displaying this content and a means for filtering inappropriate content, a safe and comfortable internet environment for children is provided.
[0294] "Internet content" refers to data such as text, images, videos, and links posted on websites and online platforms.
[0295] "Analysis" refers to the process of breaking down acquired content and understanding and extracting its components and meaning.
[0296] "User development level" refers to a concept that indicates the stage of growth of a user's comprehension, knowledge, language ability, etc.
[0297] "Generative AI model" refers to technology that uses artificial intelligence to generate appropriate content or text based on user input.
[0298] "Recognizing the user's emotions" refers to identifying the user's current emotional state from their facial expressions, tone of voice, input characters, etc.
[0299] "Adjusting content presentation" refers to appropriately changing the way content is delivered and the vocabulary used depending on the perceived emotion and developmental level.
[0300] "Inappropriate content filtering" refers to the process of detecting and removing or replacing content that contains violent, adult, or offensive material before it is displayed to users.
[0301] "JSON format data" refers to a data format structured using JavaScript Object Notation that is easy for both humans and machines to read.
[0302] The present invention is composed of a system including a server, a terminal, a user, a generative AI model, and an emotion engine. This system is designed to provide an appropriate Internet environment especially for children, and is capable of displaying content according to the user's developmental level and emotions.
[0303] The server first retrieves content from the Internet. In this process, it uses Python's requests library to send an HTTP GET request to retrieve HTML data from the specified URL. For example, if a user tries to access "https: / / example.com / animal," the server receives this URL and downloads the corresponding HTML data.
[0304] The server then parses the HTML data using the BeautifulSoup library to extract the main content elements of the web page, such as text, images, and links. For example, <title> Animal World< / title> "or" Elephants are very large herbivores. " HTML elements such as are parsed.
[0305] The server then uses a generative AI model to convert the analyzed content according to the user's developmental level. Here, we use OpenAI's GPT-3 as an example. For example, the sentence "Elephants are very large herbivores" is converted to "Elephants are very large grass-eating animals" using the prompt "Please change this sentence to hiragana." The following prompt sentence is used as an example of input to the generative AI model:
[0306] For kindergarteners: "Change this sentence into hiragana. Elephants are very large herbivores."
[0307] For elementary school students: "Please translate this sentence into a form that elementary school students can easily understand, using furigana. Elephants are very large herbivores."
[0308] The server then uses an emotion engine to recognize the user's emotions. Specifically, it uses Microsoft Azure's Face API to analyze the user's facial expressions from image data and identify the emotion at that time. For example, the server can upload an image of the user's facial expression and obtain emotion data such as "anxiety" or "joy." The results are returned in JSON format.
[0309] Based on the results of emotion recognition, the server further adjusts the content output by the generative AI model. For example, if it recognizes that the user is feeling anxious, it adds an encouraging message to the text: "Want to know more about elephants? Good luck!"
[0310] The server then filters out inappropriate content, using Python's re library to perform keyword filtering to detect and remove or modify "violent" or "inappropriate" language.
[0311] Finally, the server returns the converted content to the device in JSON format as an HTTP response, making it displayable on the device.
[0312] When a user accesses a specific web page through a browser, the device sends a request to the server and receives the converted content returned by the server. The received content is displayed in a format that is easy for children to understand using HTML and JavaScript. For example, if the sentence "Elephants are very big animals that eat grass" is displayed, and the user has an "anxious" expression, the message "Want to know more about elephants? Good luck!" is displayed.
[0313] In this way, the present invention can provide appropriate and safe information according to the user's developmental level and emotions, and can provide a safe and comfortable Internet usage environment for children.
[0314] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0315] Step 1:
[0316] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data. The URL entered by the user (e.g., "https: / / example.com / animal") is the input data, and the HTML data returned in the HTTP response is the output. Specifically, the GET request is sent using the Python requests library.
[0317] Step 2:
[0318] The server analyzes the retrieved HTML data. Here, it uses the BeautifulSoup library to extract the main content (text, images, links, etc.) from the HTML data. The HTML data is the input, and the extracted content (e.g., the text "Elephants are very large herbivores") is the output. Specifically, it breaks down the HTML structure and lists each element.
[0319] Step 3:
[0320] The server converts the analyzed text content using a generative AI model. Here, we use OpenAI's GPT-3 model. The input requires the extracted text content (e.g., "Elephants are very large herbivores") and a prompt sentence (e.g., "Please convert this sentence into hiragana"), and the output is the converted text (e.g., "Elephants are very large grass-eating animals"). Specifically, the server calls the API of the generative AI model, sends the prompt, and obtains the conversion result.
[0321] Step 4:
[0322] The server uses an emotion engine to recognize the user's emotions. Here, it uses an emotion recognition algorithm (e.g., Face API) to extract emotions from the user's facial image and text. The input is an image of the user's facial expression and text, and the output is the user's emotional data (e.g., "anxiety," "joy," etc.). Specifically, it calls the emotion engine's API, sends the input data, and retrieves the results.
[0323] Step 5:
[0324] The server adjusts the content output by the generative AI model based on the recognized emotion data. For example, if the input emotion data is "anxiety," an encouraging message (e.g., "Let's do our best!") is added to the output content. Specifically, the server adds a conditional message to the converted text.
[0325] Step 6:
[0326] The server filters inappropriate content by using the Python re library to detect specific keywords or phrases in the text and remove or modify the inappropriate content. The input is the converted and adjusted text data, and the output is the text data with the inappropriate content filtered out. Specifically, it uses regular expressions to find inappropriate keywords and replace them.
[0327] Step 7:
[0328] The server returns the converted content to the device. It sends the converted content in JSON format as an HTTP response, making it displayable on the device. The input is the filtered text data, and the output is JSON data sent to the device. The specific operation is to generate an HTTP response that includes the converted content.
[0329] Step 8:
[0330] The terminal receives the converted content returned from the server. In this case, it receives the HTTP response sent from the server. The input is the HTTP response from the server, and the output is the converted content data. Specifically, it analyzes the HTTP response and extracts the content portion.
[0331] Step 9:
[0332] The device displays the received content in a browser. Specifically, it displays the retrieved text and images using HTML and JavaScript. The input is the extracted content data, and the output is the display content as a user interface. Specifically, it performs DOM operations on the HTML document and inserts the text and images into the specified location.
[0333] (Application example 2)
[0334] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0335] Conventional Internet content distribution systems do not adequately adapt content to suit the user's age and developmental level. They also lack the ability to recognize the user's emotional state in real time and adjust content accordingly. This makes it difficult to provide appropriate content that is easy to understand and safe to use, especially for children. Furthermore, inappropriate content filtering is often inadequate, making it difficult to provide a safe and secure environment. There is a need for a system that solves these issues.
[0336] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0337] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for recognizing the user's emotions, means for adjusting the content based on the recognized emotions, and means for filtering inappropriate content. This makes it possible to provide content appropriate for the user's age and developmental level, as well as to recognize the user's emotional state in real time to display optimal content, providing an environment in which the user can use the content with peace of mind.
[0338] "Internet content" refers to information and multimedia data accessible via the Internet, such as websites, blogs, and video sharing sites.
[0339] "Means of analysis" refers to systems or algorithms that analyze acquired internet content such as text, images, and videos to understand its structure and content.
[0340] "User development level" refers to the stage of cognitive ability according to the user's age, knowledge, and comprehension, and is a standard for providing appropriate content based on this.
[0341] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning techniques to generate new data based on given data.
[0342] A "content transformation means" is a system or algorithm that transforms acquired content into a form appropriate for the user's developmental level.
[0343] A "displaying means" is a system or algorithm that displays the transformed content on an output device (such as a screen or monitor) in a form that is easy for the user to understand.
[0344] "Means for recognizing emotions" refers to systems or algorithms that analyze a user's facial expressions and voice to identify their current emotional state.
[0345] "Adjustment" means a system or algorithm that appropriately modifies the presentation or substance of content based on the perceived sentiment.
[0346] "Inappropriate content" is content that includes information or expressions that are deemed harmful or inappropriate to users.
[0347] "Filtering measures" are systems or algorithms that detect inappropriate content and either remove it or transform it into a harmless form.
[0348] "Furigana" is kana characters added to indicate how to read kanji characters, and is an auxiliary character that makes it easier for users, especially those in the developmental stage, to understand kanji characters.
[0349] A specific system for implementing this invention is composed of elements such as a server, a terminal, a user, and an emotion engine. This system can convert Internet content appropriately according to the user's developmental level and emotional state, and provide it in a form that can be used safely.
[0350] Server-side processing
[0351] The server plays an important role in analyzing the retrieved internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are as follows:
[0352] 1. Content Acquisition:
[0353] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[0354] 2. Content Analysis:
[0355] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[0356] 3. Transformation by generative AI models:
[0357] The server analyzes the content and then uses a generative AI model (e.g., a model using TensorFlow) to adapt the content to the user's developmental level. Specifically, for kindergarteners, the server converts sentences into hiragana and uses simple words. For elementary school students, the server converts the content into easy-to-understand expressions using kanji with furigana.
[0358] 4. Emotion Recognition with Emotion Engine:
[0359] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text, and recognizes the emotion at that time.
[0360] 5. Emotion-based expression adjustment:
[0361] Depending on the recognized emotion, the generative AI model can further adjust the content it outputs: for example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[0362] 6. Filtering inappropriate content:
[0363] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[0364] 7. Return of Converted Content:
[0365] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays the content in a user-friendly format.
[0366] Terminal side processing
[0367] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0368] 1. Submit your request:
[0369] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0370] 2. Receiving Content:
[0371] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[0372] 3. Viewing Content:
[0373] The received content is displayed in the browser in a format that is easy for the user to understand. For example, for kindergarteners, sentences are converted into hiragana, sentences expressed in simple words, and kanji with furigana are used. Also, if the user looks anxious, an encouraging message is added.
[0374] User usage scenarios
[0375] A user accesses a specific web page through the system. For example, if a child wants to browse a website called "Animal World," the following process will occur when using this system:
[0376] Access Start:
[0377] The user opens a browser and enters the URL for "Animal World" to access the site.
[0378] Server-based conversion:
[0379] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[0380] Show content:
[0381] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[0382] Example of generative AI model prompt
[0383] User: 5 years old
[0384] Emotion: Anxiety
[0385] Text: Please convert and display the content of "Dog wagging its tail".
[0386] Prompt result: Generates a kind sentence about a dog wagging its tail and a sentence containing an encouraging message.
[0387] In this way, the system provides appropriate content according to the user's developmental level and emotional state, creating an environment in which the user can use the system with peace of mind.
[0388] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0389] Step 1:
[0390] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL and presses the access button, which generates an HTTP request and sends the request to the server. The input data is the URL, and the output data is the HTTP request.
[0391] Step 2:
[0392] The server receives a request and retrieves HTML data from a specified URL on the Internet. The server sends an HTTP GET request to download the entire specified web page. The input data is the URL and the HTTP request, and the output data is the HTML content. Specifically, the server receives the HTML data as a response.
[0393] Step 3:
[0394] The server uses an HTML parser (e.g., BeautifulSoup) to parse the HTML data it receives. The parsing extracts the main content elements in the web page, such as text, images, and links. The input data is the HTML content, and the output data are the parsed content elements. The server prepares these elements to be stored in a database.
[0395] Step 4:
[0396] The server uses a generative AI model (e.g., a TensorFlow model) to convert the analyzed content elements according to the user's developmental level. For example, for kindergarteners, the server converts sentences into hiragana and uses simple words. The input data are the analyzed content elements and the user's developmental level information, and the output data is the converted content. Specifically, the AI model processes the text data and converts it into an appropriate expression.
[0397] Step 5:
[0398] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze and recognize the user's facial expressions, tone of voice, and emotional expressions in the input text. The input data is the user's facial image and voice data, and the output data is the recognized emotional state. The server prepares to adjust the content based on this data.
[0399] Step 6:
[0400] Based on the recognized emotions, the server further adjusts the content output by the generative AI model. For example, if the user is anxious, it converts the content into gentler language and adds an encouraging message. The input data is the converted content and emotion data, and the output data is the adjusted content. Specifically, the AI model adds text data and makes changes according to the emotion.
[0401] Step 7:
[0402] The server filters inappropriate content through content analysis and conversion. Filtering algorithms are used to detect inappropriate keywords and phrases and remove or replace them with harmless alternatives. The input data is the converted content, and the output data is the filtered, safe content.
[0403] Step 8:
[0404] After all conversion and filtering is complete, the server returns the converted content to the terminal. The input data is the filtered content, and the output data is the HTTP response. This allows the content to be displayed on the terminal in a format that is easy for the user to understand.
[0405] Step 9:
[0406] The device receives the converted content sent back from the server and displays it appropriately to the user. The input data is the converted content, and the output data is the content to be displayed. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." Also, if the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[0407] In this way, the system provides appropriate Internet content according to the user's developmental level and emotional state, creating an environment in which the user can use the content with peace of mind.
[0408] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0409] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0410] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0411] [Second embodiment]
[0412] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0413] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0414] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0415] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0416] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0418] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0419] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0420] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0421] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0422] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0423] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0424] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[0425] Server-side processing
[0426] The server plays an important role in analyzing the acquired Internet content and converting it according to the user's level of development. Specifically, the process is as follows:
[0427] 1. Content Acquisition:
[0428] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and sends an HTTP GET request to download the entire web page.
[0429] 2. Content Analysis:
[0430] To parse the retrieved HTML data, the server uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[0431] 3. Transformation by generative AI models:
[0432] After analyzing the content, the server uses a generative AI model to convert it to suit the user's developmental level. Specifically, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions while adding furigana to kanji. This process utilizes a generative AI model (e.g., GPT-3).
[0433] 4. Filtering inappropriate content:
[0434] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[0435] 5. Return of Converted Content:
[0436] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[0437] Terminal side processing
[0438] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0439] 1. Submit your request:
[0440] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0441] 2. Receiving Content:
[0442] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[0443] 3. Viewing Content:
[0444] The received content is displayed in the browser in a format that is easy for children to understand, such as sentences that make extensive use of hiragana and kanji with furigana.
[0445] User usage scenarios
[0446] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[0447] 1. Initiating access:
[0448] The user opens a browser and enters the URL for "Animal World" to access the site.
[0449] 2. Conversion by the server:
[0450] The device sends a request for content to the server, which retrieves, analyzes, and converts it.
[0451] 3. Viewing Content:
[0452] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[0453] As such, the present invention is a system that provides appropriate Internet usage according to the user's developmental level, and aims to promote children's learning and information acquisition while protecting them from inappropriate content.
[0454] The processing flow will be explained below.
[0455] Step 1:
[0456] A user launches a browser and attempts to access a particular website.
[0457] Step 2:
[0458] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[0459] Step 3:
[0460] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[0461] Step 4:
[0462] Parse the HTML data retrieved by the server using an HTML parser (such as BeautifulSoup) to understand the overall structure of the web page and extract the main content elements (text, images, links, etc.).
[0463] Step 5:
[0464] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level. Specifically, for kindergarteners, the text is converted into hiragana and simple words are used, while for elementary school students, kanji with furigana is used, making it easier to understand.
[0465] Step 6:
[0466] The server filters inappropriate content during the conversion process, using a filtering algorithm to detect inappropriate keywords and phrases and remove them or replace them with harmless alternatives.
[0467] Step 7:
[0468] The server returns the converted content to the device, and sends an HTTP response containing the converted HTML data and other resources.
[0469] Step 8:
[0470] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[0471] Step 9:
[0472] The device then displays the converted content to the user, rendering the web page in a way that is easy for the user to understand, such as displaying sentences that make heavy use of hiragana or kanji with furigana.
[0473] Example 1
[0474] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0475] There is a wide variety of content on the Internet, and understanding and safety of that content poses challenges, especially when children access it. Current technology does not adequately convert content or filter inappropriate content, making it difficult to provide information in a format that is easy for children to understand. Furthermore, the need for manual content filtering and the lack of automated technology to convert content into appropriate expressions according to the user's developmental level result in an inconsistent provision of educational value.
[0476] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0477] In this invention, the server includes means for acquiring accessed internet data, means for analyzing the acquired HTML data and extracting key elements, means for converting the content according to the user's developmental level using a generative AI model, means for filtering the converted content, and means for returning the converted and filtered content to the terminal. This makes it possible to automatically analyze, convert, and filter internet content accessed by children and provide appropriate information according to their developmental level.
[0478] "Accessed Internet Data" is information obtained from the web pages or online resources that a user attempts to access.
[0479] "HTML data" means data written in the HyperText Markup Language used to describe the structure and content of web pages.
[0480] "Key elements" are the core content of a Web page, such as text, images, links, and headings.
[0481] A "generative AI model" is an artificial intelligence algorithm that generates appropriate content based on given input, particularly one based on large-scale natural language processing techniques.
[0482] "User's developmental level" refers to the level of content appropriate for the user's age and learning stage.
[0483] The "means of transforming content" refers to the process of using the generative AI model described above to transform the original content into an appropriate format according to the user's developmental level.
[0484] "Content filtering measures" are the processes that detect inappropriate content or keywords and remove or replace them with harmless forms.
[0485] A "terminal" is an electronic device that a user directly uses as an interface, such as a personal computer, smartphone, or tablet.
[0486] "Means for returning content to the terminal" refers to the process by which the server sends the processed content to the user's terminal.
[0487] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[0488] Server-side processing
[0489] The server retrieves and analyzes data from the internet, uses generative AI models to transform content according to the user's developmental level, and filters out inappropriate content.
[0490] The server retrieves the accessed data on the Internet. Specifically, it receives the URL of the web page the user is trying to access, sends an HTTP GET request to that URL, and retrieves the HTML data. The server uses an HTML parser such as BeautifulSoup to analyze this data, thereby extracting the main elements of the web page (text, images, links, etc.).
[0491] It then uses a generative AI model (e.g., GPT-3) to adapt the parsed content to the user's developmental level, using the following example prompt:
[0492] "Simplify the text on the web page for children. Enter the text below: '{{Web page text}}'"
[0493] "Parse this HTML data and convert the extracted key content elements to child-friendly. HTML data: '{{HTML data}}'"
[0494] The content converted by the generative AI model then goes through a further filtering process: the server detects inappropriate keywords and phrases based on a list and either removes them or replaces them with harmless language.
[0495] After all the processing is complete, the server returns the converted and filtered content to the device, ensuring that the content is presented in a child-friendly and safe manner.
[0496] Terminal side processing
[0497] The terminal is responsible for appropriately displaying the converted content received from the server to the user.
[0498] When a user accesses a particular web page through a browser, the device sends a request to the server, which includes the URL they are trying to access. The server returns the converted content, which the device then displays in the browser.
[0499] For example, if an elementary school student wants to visit a website called "Animal World," the process would go something like this: The user opens a browser and enters the URL for "Animal World." The device sends a request to the server, which retrieves, analyzes, converts, and filters the content. Finally, the device receives the converted content and displays it appropriately to the user.
[0500] User usage scenarios
[0501] Users operate a browser through the system to access specific web pages. For example, on the "Animal World" website, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." This system allows children to obtain appropriate information safely and in an easy-to-understand manner.
[0502] As described above, the present invention aims to provide appropriate Internet usage according to the user's developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[0503] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0504] Server-side processing
[0505] Step 1:
[0506] The server receives the URL of the web page the user is trying to access, which is obtained by an HTTP GET request sent from the device.
[0507] Input: The URL the user types in the browser
[0508] Output: Retrieved URL
[0509] Step 2:
[0510] The server sends an HTTP GET request to the URL and downloads the HTML data of the entire web page. The server issues the request, obtains the response, and saves it.
[0511] Input: Retrieved URL
[0512] Output: Downloaded HTML data
[0513] Step 3:
[0514] To parse the HTML data, the server uses an HTML parser such as BeautifulSoup to parse the data. The server creates a BeautifulSoup object and parses the HTML data to extract the main elements (text, images, links, etc.).
[0515] Input: Downloaded HTML data
[0516] Output: Extracted key elements (text, images, links, etc.)
[0517] Step 4:
[0518] The server sends prompts to a generative AI model (e.g., GPT-3) and converts the analyzed content according to the user's development level. The server constructs an appropriate prompt sentence for the text portion of the content and sends it to the AI model.
[0519] Input: extracted key elements and user's development level
[0520] Output: The converted text content
[0521] Step 5:
[0522] The server scans the converted content and filters out inappropriate content. Using a list of inappropriate keywords, the server detects these words and phrases and either replaces them with appropriate alternatives or removes them.
[0523] Input: The converted text content
[0524] Output: The filtered text content
[0525] Step 6:
[0526] The server returns the content after all conversion and filtering has been completed to the terminal as an HTTP response, so that the converted content can be displayed appropriately to the user.
[0527] Input: Filtered text content
[0528] Output: The transformed content included in the HTTP response
[0529] Terminal side processing
[0530] Step 1:
[0531] When a user accesses a specific web page through a browser, the device sends the request to the server. The browser issues an HTTP GET request, sending the URL that the user wants to access to the server.
[0532] Input: The URL the user types in the browser
[0533] Output: HTTP GET request to the server
[0534] Step 2:
[0535] The terminal receives the converted content returned from the server, receives the HTTP response, and analyzes and obtains the converted HTML data contained in the response body.
[0536] Input: HTTP response from the server
[0537] Output: Converted HTML data
[0538] Step 3:
[0539] The device displays the received content in the browser, and uses the browser's rendering engine to display the converted content in a format that is easy for children to understand.
[0540] Input: Converted HTML data
[0541] Output: Content displayed in the browser
[0542] User usage scenarios
[0543] 1. Initiating Access
[0544] A user opens a browser and enters the URL of the website they want to visit.
[0545] Input: The URL entered by the user
[0546] Output: The browser issues an access request
[0547] 2. Conversion by the server
[0548] The device sends a request for content to the server, which retrieves, analyzes, converts, and filters it.
[0549] Input: The URL the user types into the browser
[0550] Output: The transformed and filtered content
[0551] 3. Display of Content
[0552] The device receives the converted content and displays it appropriately in the browser. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[0553] Input: The converted content sent from the server
[0554] Output: Content displayed in the browser
[0555] The above are the specific processing steps of the system based on the present invention, which enables the conversion of Internet content according to the user's level of development, filtering out inappropriate content, and providing users with safe and easy-to-understand information.
[0556] (Application example 1)
[0557] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0558] When children browse content online, it is difficult to provide them with content that is converted into an appropriate format according to their age and level of understanding. Furthermore, content viewed online may contain inappropriate information, and there is a lack of means to accurately filter and safely provide this information. Furthermore, there is a need to improve learning outcomes by appropriately adjusting the display of content according to the child's developmental level.
[0559] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0560] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for generating a unique prompt sentence and inputting it to the generative AI model, and means for filtering inappropriate content, thereby enabling the content to be converted into a format that is easy for children to understand and displayed safely.
[0561] "Internet content" refers to all information media, such as text, images, video, and audio, that can be obtained through the Internet.
[0562] "Means of analysis" refers to technical methods for structurally understanding acquired Internet content and extracting each element.
[0563] A "generative AI model" is a collection of machine learning algorithms that use artificial intelligence techniques to generate or transform text or images.
[0564] "Means for conversion according to the user's developmental level" refers to technology for converting content into an appropriate format based on the user's age and level of understanding.
[0565] The "means for displaying the converted content" refers to a device or software for appropriately outputting the converted content to the user's terminal.
[0566] "Means for generating unique prompt sentences and inputting them into a generative AI model" refers to a technical method for creating prompt sentences in a form suitable for a generative AI model and providing them to the model.
[0567] "Inappropriate content filtering measures" are technologies that detect inappropriate keywords and phrases from internet content and either remove them or replace them with harmless ones.
[0568] A "means for adding furigana" is a technique or device for adding readings to difficult characters such as kanji.
[0569] "Means of converting expressions into simple terms" are techniques for converting complex sentences and technical terms into simple terms that children can easily understand.
[0570] "Means for extracting parts that match specific keywords" refers to technology for detecting and extracting specific words or phrases within content.
[0571] "Means for integrating and displaying" refers to the technology or device that compiles and presents the analyzed, transformed, and filtered content to the user.
[0572] This invention is a system that converts content on the Internet according to the user's level of development when the user browses the content, and displays the content safely. This system is composed of a server and a terminal.
[0573] Server-side processing
[0574] 1. Content Acquisition
[0575] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data for that URL, which downloads the entire web page data to the server.
[0576] 2. Content Analysis
[0577] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[0578] 3. Conversion using generative AI models
[0579] The server uses a generative AI model (e.g., GPT-3) based on the analyzed content to convert the content to suit the user's developmental level. The conversion is performed by generating appropriate prompts and providing them as input to the AI model. For example, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions, with furigana added to kanji.
[0580] 4. Filtering inappropriate content
[0581] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[0582] 5. Return of Converted Content
[0583] After all conversion and filtering is complete, the server returns the converted content to the terminal.
[0584] Terminal side processing
[0585] 1. Submitting a Request
[0586] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0587] 2. Receiving Content
[0588] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[0589] 3. Display of Content
[0590] The device displays the received content in a way that is easy for children to understand, such as sentences using simple words or text with furigana added to kanji characters.
[0591] Content conversion examples
[0592] For example, if an elementary school student wants to browse a website called "Animal World," using this system, the following flow will be displayed: The sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals," making it easier for elementary school students to understand.
[0593] Hardware and software used
[0594] Hardware: Smartphone (device)
[0595] Software: Python, requests library, BeautifulSoup library, OpenAI GPT-3 API
[0596] Prompt Sentence Examples
[0597] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[0598] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[0599] In this way, the present invention aims to provide users with access to Internet content appropriate to their developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[0600] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0601] Step 1:
[0602] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL into the web browser and clicks the access button. At this time, the device sends an HTTP request to the server, along with information about the user's development level. The input is the URL and development level information entered by the user, and the output is the HTTP request sent to the server.
[0603] Step 2:
[0604] The server sends an HTTP GET request to retrieve the HTML data of the received URL. Specifically, the server downloads the entire web page using the specified URL. At this time, the server receives the HTML content of the web page as a response to the request. The input is the URL, and the output is the retrieved HTML data.
[0605] Step 3:
[0606] The server parses the retrieved HTML data using an HTML parser such as BeautifulSoup. Specifically, it extracts the main content elements such as text, images, and links from the HTML data. The input is the retrieved HTML data, and the output is the extracted content elements.
[0607] Step 4:
[0608] The server uses a generative AI model (e.g., GPT-3) to generate a prompt appropriate for the user's developmental level, and then uses the prompt to convert the content. Specifically, it generates an appropriate prompt based on the analyzed content and inputs it into the generative AI model to convert the content into a format appropriate for the user's developmental level. The input is the prompt and the analyzed content, and the output is the converted content. An example of a prompt is as follows:
[0609] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[0610] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[0611] Step 5:
[0612] The server filters the transformed content to remove or replace inappropriate content, specifically detecting inappropriate keywords or phrases and replacing them with harmless terms. The input is the transformed content, and the output is the filtered content.
[0613] Step 6:
[0614] The server returns the filtered content to the terminal. Specifically, it sends the filtered and converted content data to the terminal as an HTTP response. The input is the filtered content, and the output is the data returned to the terminal.
[0615] Step 7:
[0616] The device displays the converted content received from the server to the user. Specifically, it renders the converted content on the browser and displays it in a format that is easy for children to understand. The input is the content data returned from the server, and the output is the content displayed to the user.
[0617] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0618] A specific system for implementing the present invention is composed of elements of a server, a terminal, a user, and an emotion engine.
[0619] Server-side processing
[0620] The server plays an important role in analyzing the acquired internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are explained below.
[0621] 1. Content Acquisition:
[0622] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[0623] 2. Content Analysis:
[0624] To parse the retrieved HTML data, the server uses an HTML parser to extract the main content elements of the web page, such as text, images, and links.
[0625] 3. Transformation by generative AI models:
[0626] The server analyzes the content and then uses a generative AI model to adapt it to the user's developmental level. Specifically, for kindergarteners, it converts sentences into hiragana and uses simple words, while for elementary school students, it uses kanji with furigana and adapts it to easier-to-understand expressions.
[0627] 4. Emotion Recognition with Emotion Engine:
[0628] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text to recognize the emotion at that time.
[0629] 5. Emotion-based expression adjustment:
[0630] Depending on the emotion recognized, the generative AI model can further adjust the content it outputs. For example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[0631] 6. Filtering inappropriate content:
[0632] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[0633] 7. Return of Converted Content:
[0634] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[0635] Terminal side processing
[0636] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0637] 1. Submit your request:
[0638] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0639] 2. Receiving Content:
[0640] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[0641] 3. Viewing Content:
[0642] The received content is displayed in the browser in a format that is easy for children to understand, such as sentences written in simple language or kanji with furigana.
[0643] User usage scenarios
[0644] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[0645] 1. Initiating access:
[0646] The user opens a browser and enters the URL for "Animal World" to access the site.
[0647] 2. Conversion by the server:
[0648] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[0649] 3. Viewing Content:
[0650] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[0651] As such, the present invention is a system that provides an appropriate Internet usage environment according to the user's developmental level and emotions, and aims to support children's learning and information acquisition, as well as provide a comfortable and secure content experience.
[0652] The processing flow will be explained below.
[0653] Step 1:
[0654] A user launches a browser and attempts to access a particular website (e.g., "Animal World").
[0655] Step 2:
[0656] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[0657] Step 3:
[0658] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[0659] Step 4:
[0660] The server parses the HTML data it retrieves and uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[0661] Step 5:
[0662] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level (kindergartener, elementary school student, etc.) Specifically, it converts sentences into hiragana, uses kanji with furigana, or modifies them into simpler words.
[0663] Step 6:
[0664] The server simultaneously recognizes the user's emotions using an emotion engine, which analyzes facial expressions and tone of voice through a camera and microphone to assess the user's emotional state (e.g., joy, anxiety, excitement, etc.).
[0665] Step 7:
[0666] The server further adjusts the converted content based on the emotions recognized by the emotion engine, for example, changing the text to a more gentle expression and adding an encouraging message if the user is anxious.
[0667] Step 8:
[0668] The server uses filtering algorithms to detect inappropriate keywords and phrases in the content and either remove them or replace them with harmless alternatives.
[0669] Step 9:
[0670] The server sends the converted and filtered content back to the device, which then sends an HTTP response containing the converted HTML data and other resources.
[0671] Step 10:
[0672] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[0673] Step 11:
[0674] The device displays the converted content in the browser. For example, the sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" will be added.
[0675] Through this tailored content, users can access information that is easy to understand and appropriate for their developmental level, and enjoy a comfortable internet experience that is sensitive to their emotions.
[0676] Example 2
[0677] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0678] With conventional internet usage, it has been difficult to display content that takes into account the user's developmental level and emotions, making it difficult to provide appropriate information to children. Furthermore, filtering of inappropriate content is insufficient, creating a need for a safe internet usage environment.
[0679] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0680] In this invention, the server includes a means for analyzing acquired internet content, a means for using a generative AI model to convert the content according to the user's developmental level, and a means for recognizing the user's emotions and adjusting the presentation of the content based on those emotions. This enables the provision of appropriate and safe information according to the user's developmental level and emotions. Furthermore, by including a means for displaying this content and a means for filtering inappropriate content, a safe and comfortable internet environment for children is provided.
[0681] "Internet content" refers to data such as text, images, videos, and links posted on websites and online platforms.
[0682] "Analysis" refers to the process of breaking down acquired content and understanding and extracting its components and meaning.
[0683] "User development level" refers to a concept that indicates the stage of growth of a user's comprehension, knowledge, language ability, etc.
[0684] "Generative AI model" refers to technology that uses artificial intelligence to generate appropriate content or text based on user input.
[0685] "Recognizing the user's emotions" refers to identifying the user's current emotional state from their facial expressions, tone of voice, input characters, etc.
[0686] "Adjusting content presentation" refers to appropriately changing the way content is delivered and the vocabulary used depending on the perceived emotion and developmental level.
[0687] "Inappropriate content filtering" refers to the process of detecting and removing or replacing content that contains violent, adult, or offensive material before it is displayed to users.
[0688] "JSON format data" refers to a data format structured using JavaScript Object Notation that is easy for both humans and machines to read.
[0689] The present invention is composed of a system including a server, a terminal, a user, a generative AI model, and an emotion engine. This system is designed to provide an appropriate Internet environment especially for children, and is capable of displaying content according to the user's developmental level and emotions.
[0690] The server first retrieves content from the Internet. In this process, it uses Python's requests library to send an HTTP GET request to retrieve HTML data from the specified URL. For example, if a user tries to access "https: / / example.com / animal," the server receives this URL and downloads the corresponding HTML data.
[0691] The server then parses the HTML data using the BeautifulSoup library to extract the main content elements of the web page, such as text, images, and links. For example, <title> Animal World< / title> "or" Elephants are very large herbivores. " HTML elements such as are parsed.
[0692] The server then uses a generative AI model to convert the analyzed content according to the user's developmental level. Here, we use OpenAI's GPT-3 as an example. For example, the sentence "Elephants are very large herbivores" is converted to "Elephants are very large grass-eating animals" using the prompt "Please change this sentence to hiragana." The following prompt sentence is used as an example of input to the generative AI model:
[0693] For kindergarteners: "Change this sentence into hiragana. Elephants are very large herbivores."
[0694] For elementary school students: "Please translate this sentence into a form that elementary school students can easily understand, using furigana. Elephants are very large herbivores."
[0695] The server then uses an emotion engine to recognize the user's emotions. Specifically, it uses Microsoft Azure's Face API to analyze the user's facial expressions from image data and identify the emotion at that time. For example, the server can upload an image of the user's facial expression and obtain emotion data such as "anxiety" or "joy." The results are returned in JSON format.
[0696] Based on the results of emotion recognition, the server further adjusts the content output by the generative AI model. For example, if it recognizes that the user is feeling anxious, it adds an encouraging message to the text: "Want to know more about elephants? Good luck!"
[0697] The server then filters out inappropriate content, using Python's re library to perform keyword filtering to detect and remove or modify "violent" or "inappropriate" language.
[0698] Finally, the server returns the converted content to the device in JSON format as an HTTP response, making it displayable on the device.
[0699] When a user accesses a specific web page through a browser, the device sends a request to the server and receives the converted content returned by the server. The received content is displayed in a format that is easy for children to understand using HTML and JavaScript. For example, if the sentence "Elephants are very big animals that eat grass" is displayed, and the user has an "anxious" expression, the message "Want to know more about elephants? Good luck!" is displayed.
[0700] In this way, the present invention can provide appropriate and safe information according to the user's developmental level and emotions, and can provide a safe and comfortable Internet usage environment for children.
[0701] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0702] Step 1:
[0703] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data. The URL entered by the user (e.g., "https: / / example.com / animal") is the input data, and the HTML data returned in the HTTP response is the output. Specifically, the GET request is sent using the Python requests library.
[0704] Step 2:
[0705] The server analyzes the retrieved HTML data. Here, it uses the BeautifulSoup library to extract the main content (text, images, links, etc.) from the HTML data. The HTML data is the input, and the extracted content (e.g., the text "Elephants are very large herbivores") is the output. Specifically, it breaks down the HTML structure and lists each element.
[0706] Step 3:
[0707] The server converts the analyzed text content using a generative AI model. Here, we use OpenAI's GPT-3 model. The input requires the extracted text content (e.g., "Elephants are very large herbivores") and a prompt sentence (e.g., "Please convert this sentence into hiragana"), and the output is the converted text (e.g., "Elephants are very large grass-eating animals"). Specifically, the server calls the API of the generative AI model, sends the prompt, and obtains the conversion result.
[0708] Step 4:
[0709] The server uses an emotion engine to recognize the user's emotions. Here, it uses an emotion recognition algorithm (e.g., Face API) to extract emotions from the user's facial image and text. The input is an image of the user's facial expression and text, and the output is the user's emotional data (e.g., "anxiety," "joy," etc.). Specifically, it calls the emotion engine's API, sends the input data, and retrieves the results.
[0710] Step 5:
[0711] The server adjusts the content output by the generative AI model based on the recognized emotion data. For example, if the input emotion data is "anxiety," an encouraging message (e.g., "Let's do our best!") is added to the output content. Specifically, the server adds a conditional message to the converted text.
[0712] Step 6:
[0713] The server filters inappropriate content by using the Python re library to detect specific keywords or phrases in the text and remove or modify the inappropriate content. The input is the converted and adjusted text data, and the output is the text data with the inappropriate content filtered out. Specifically, it uses regular expressions to find inappropriate keywords and replace them.
[0714] Step 7:
[0715] The server returns the converted content to the device. It sends the converted content in JSON format as an HTTP response, making it displayable on the device. The input is the filtered text data, and the output is JSON data sent to the device. The specific operation is to generate an HTTP response that includes the converted content.
[0716] Step 8:
[0717] The terminal receives the converted content returned from the server. In this case, it receives the HTTP response sent from the server. The input is the HTTP response from the server, and the output is the converted content data. Specifically, it analyzes the HTTP response and extracts the content portion.
[0718] Step 9:
[0719] The device displays the received content in a browser. Specifically, it displays the retrieved text and images using HTML and JavaScript. The input is the extracted content data, and the output is the display content as a user interface. Specifically, it performs DOM operations on the HTML document and inserts the text and images into the specified location.
[0720] (Application example 2)
[0721] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0722] Conventional Internet content distribution systems do not adequately adapt content to suit the user's age and developmental level. They also lack the ability to recognize the user's emotional state in real time and adjust content accordingly. This makes it difficult to provide appropriate content that is easy to understand and safe to use, especially for children. Furthermore, inappropriate content filtering is often inadequate, making it difficult to provide a safe and secure environment. There is a need for a system that solves these issues.
[0723] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0724] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for recognizing the user's emotions, means for adjusting the content based on the recognized emotions, and means for filtering inappropriate content. This makes it possible to provide content appropriate for the user's age and developmental level, as well as to recognize the user's emotional state in real time to display optimal content, providing an environment in which the user can use the content with peace of mind.
[0725] "Internet content" refers to information and multimedia data accessible via the Internet, such as websites, blogs, and video sharing sites.
[0726] "Means of analysis" refers to systems or algorithms that analyze acquired internet content such as text, images, and videos to understand its structure and content.
[0727] "User development level" refers to the stage of cognitive ability according to the user's age, knowledge, and comprehension, and is a standard for providing appropriate content based on this.
[0728] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning techniques to generate new data based on given data.
[0729] A "content transformation means" is a system or algorithm that transforms acquired content into a form appropriate for the user's developmental level.
[0730] A "displaying means" is a system or algorithm that displays the transformed content on an output device (such as a screen or monitor) in a form that is easy for the user to understand.
[0731] "Means for recognizing emotions" refers to systems or algorithms that analyze a user's facial expressions and voice to identify their current emotional state.
[0732] "Adjustment" means a system or algorithm that appropriately modifies the presentation or substance of content based on the perceived sentiment.
[0733] "Inappropriate content" is content that includes information or expressions that are deemed harmful or inappropriate to users.
[0734] "Filtering measures" are systems or algorithms that detect inappropriate content and either remove it or transform it into a harmless form.
[0735] "Furigana" is kana characters added to indicate how to read kanji characters, and is an auxiliary character that makes it easier for users, especially those in the developmental stage, to understand kanji characters.
[0736] A specific system for implementing this invention is composed of elements such as a server, a terminal, a user, and an emotion engine. This system can convert Internet content appropriately according to the user's developmental level and emotional state, and provide it in a form that can be used safely.
[0737] Server-side processing
[0738] The server plays an important role in analyzing the retrieved internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are as follows:
[0739] 1. Content Acquisition:
[0740] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[0741] 2. Content Analysis:
[0742] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[0743] 3. Transformation by generative AI models:
[0744] The server analyzes the content and then uses a generative AI model (e.g., a model using TensorFlow) to adapt the content to the user's developmental level. Specifically, for kindergarteners, the server converts sentences into hiragana and uses simple words. For elementary school students, the server converts the content into easy-to-understand expressions using kanji with furigana.
[0745] 4. Emotion Recognition with Emotion Engine:
[0746] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text, and recognizes the emotion at that time.
[0747] 5. Emotion-based expression adjustment:
[0748] Depending on the recognized emotion, the generative AI model can further adjust the content it outputs: for example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[0749] 6. Filtering inappropriate content:
[0750] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[0751] 7. Return of Converted Content:
[0752] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays the content in a user-friendly format.
[0753] Terminal side processing
[0754] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0755] 1. Submit your request:
[0756] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0757] 2. Receiving Content:
[0758] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[0759] 3. Viewing Content:
[0760] The received content is displayed in the browser in a format that is easy for the user to understand. For example, for kindergarteners, sentences are converted into hiragana, sentences expressed in simple words, and kanji with furigana are used. Also, if the user looks anxious, an encouraging message is added.
[0761] User usage scenarios
[0762] A user accesses a specific web page through the system. For example, if a child wants to browse a website called "Animal World," the following process will occur when using this system:
[0763] Access Start:
[0764] The user opens a browser and enters the URL for "Animal World" to access the site.
[0765] Server-based conversion:
[0766] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[0767] Show content:
[0768] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[0769] Example of generative AI model prompt
[0770] User: 5 years old
[0771] Emotion: Anxiety
[0772] Text: Please convert and display the content of "Dog wagging its tail".
[0773] Prompt result: Generates a kind sentence about a dog wagging its tail and a sentence containing an encouraging message.
[0774] In this way, the system provides appropriate content according to the user's developmental level and emotional state, creating an environment in which the user can use the system with peace of mind.
[0775] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0776] Step 1:
[0777] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL and presses the access button, which generates an HTTP request and sends the request to the server. The input data is the URL, and the output data is the HTTP request.
[0778] Step 2:
[0779] The server receives a request and retrieves HTML data from a specified URL on the Internet. The server sends an HTTP GET request to download the entire specified web page. The input data is the URL and the HTTP request, and the output data is the HTML content. Specifically, the server receives the HTML data as a response.
[0780] Step 3:
[0781] The server uses an HTML parser (e.g., BeautifulSoup) to parse the HTML data it receives. The parsing extracts the main content elements in the web page, such as text, images, and links. The input data is the HTML content, and the output data are the parsed content elements. The server prepares these elements to be stored in a database.
[0782] Step 4:
[0783] The server uses a generative AI model (e.g., a TensorFlow model) to convert the analyzed content elements according to the user's developmental level. For example, for kindergarteners, the server converts sentences into hiragana and uses simple words. The input data are the analyzed content elements and the user's developmental level information, and the output data is the converted content. Specifically, the AI model processes the text data and converts it into an appropriate expression.
[0784] Step 5:
[0785] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze and recognize the user's facial expressions, tone of voice, and emotional expressions in the input text. The input data is the user's facial image and voice data, and the output data is the recognized emotional state. The server prepares to adjust the content based on this data.
[0786] Step 6:
[0787] Based on the recognized emotions, the server further adjusts the content output by the generative AI model. For example, if the user is anxious, it converts the content into gentler language and adds an encouraging message. The input data is the converted content and emotion data, and the output data is the adjusted content. Specifically, the AI model adds text data and makes changes according to the emotion.
[0788] Step 7:
[0789] The server filters inappropriate content through content analysis and conversion. Filtering algorithms are used to detect inappropriate keywords and phrases and remove or replace them with harmless alternatives. The input data is the converted content, and the output data is the filtered, safe content.
[0790] Step 8:
[0791] After all conversion and filtering is complete, the server returns the converted content to the terminal. The input data is the filtered content, and the output data is the HTTP response. This allows the content to be displayed on the terminal in a format that is easy for the user to understand.
[0792] Step 9:
[0793] The device receives the converted content sent back from the server and displays it appropriately to the user. The input data is the converted content, and the output data is the content to be displayed. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." Also, if the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[0794] In this way, the system provides appropriate Internet content according to the user's developmental level and emotional state, creating an environment in which the user can use the content with peace of mind.
[0795] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0796] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0797] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0798] [Third embodiment]
[0799] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0800] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0801] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0802] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0803] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0804] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0805] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0806] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0807] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0808] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0809] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0810] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0811] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[0812] Server-side processing
[0813] The server plays an important role in analyzing the acquired Internet content and converting it according to the user's level of development. Specifically, the process is as follows:
[0814] 1. Content Acquisition:
[0815] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and sends an HTTP GET request to download the entire web page.
[0816] 2. Content Analysis:
[0817] To parse the retrieved HTML data, the server uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[0818] 3. Transformation by generative AI models:
[0819] After analyzing the content, the server uses a generative AI model to convert it to suit the user's developmental level. Specifically, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions while adding furigana to kanji. This process utilizes a generative AI model (e.g., GPT-3).
[0820] 4. Filtering inappropriate content:
[0821] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[0822] 5. Return of Converted Content:
[0823] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[0824] Terminal side processing
[0825] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[0826] 1. Submit your request:
[0827] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0828] 2. Receiving Content:
[0829] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[0830] 3. Viewing Content:
[0831] The received content is displayed in the browser in a format that is easy for children to understand. For example, sentences that make extensive use of hiragana and kanji with furigana are used.
[0832] User usage scenarios
[0833] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[0834] 1. Initiating access:
[0835] The user opens a browser and enters the URL for "Animal World" to access the site.
[0836] 2. Conversion by the server:
[0837] The device sends a request for content to the server, which retrieves, analyzes, and converts it.
[0838] 3. Viewing Content:
[0839] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[0840] As such, the present invention is a system that provides appropriate Internet usage according to the user's developmental level, and aims to promote children's learning and information acquisition while protecting them from inappropriate content.
[0841] The processing flow will be explained below.
[0842] Step 1:
[0843] A user launches a browser and attempts to access a particular website.
[0844] Step 2:
[0845] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[0846] Step 3:
[0847] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[0848] Step 4:
[0849] Parse the HTML data retrieved by the server using an HTML parser (such as BeautifulSoup) to understand the overall structure of the web page and extract the main content elements (text, images, links, etc.).
[0850] Step 5:
[0851] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level. Specifically, for kindergarteners, the text is converted into hiragana and simple words are used, while for elementary school students, kanji with furigana is used, making it easier to understand.
[0852] Step 6:
[0853] The server filters inappropriate content during the conversion process, using a filtering algorithm to detect inappropriate keywords and phrases and remove them or replace them with harmless alternatives.
[0854] Step 7:
[0855] The server returns the converted content to the device, and sends an HTTP response containing the converted HTML data and other resources.
[0856] Step 8:
[0857] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[0858] Step 9:
[0859] The device then displays the converted content to the user, rendering the web page in a way that is easy for the user to understand, such as displaying sentences that make heavy use of hiragana or kanji with furigana.
[0860] Example 1
[0861] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0862] There is a wide variety of content on the Internet, and understanding and safety of that content poses challenges, especially when children access it. Current technology does not adequately convert content or filter inappropriate content, making it difficult to provide information in a format that is easy for children to understand. Furthermore, the need for manual content filtering and the lack of automated technology to convert content into appropriate expressions according to the user's developmental level result in an inconsistent provision of educational value.
[0863] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0864] In this invention, the server includes means for acquiring accessed internet data, means for analyzing the acquired HTML data and extracting key elements, means for converting the content according to the user's developmental level using a generative AI model, means for filtering the converted content, and means for returning the converted and filtered content to the terminal. This makes it possible to automatically analyze, convert, and filter internet content accessed by children and provide appropriate information according to their developmental level.
[0865] "Accessed Internet Data" is information obtained from the web pages or online resources that a user attempts to access.
[0866] "HTML data" means data written in the HyperText Markup Language used to describe the structure and content of web pages.
[0867] "Key elements" are the core content of a Web page, such as text, images, links, and headings.
[0868] A "generative AI model" is an artificial intelligence algorithm that generates appropriate content based on given input, particularly one based on large-scale natural language processing techniques.
[0869] "User's developmental level" refers to the level of content appropriate for the user's age and learning stage.
[0870] The "means of transforming content" refers to the process of using the generative AI model described above to transform the original content into an appropriate format according to the user's developmental level.
[0871] "Content filtering measures" are the processes that detect inappropriate content or keywords and remove or replace them with harmless forms.
[0872] A "terminal" is an electronic device that a user directly uses as an interface, such as a personal computer, smartphone, or tablet.
[0873] "Means for returning content to the terminal" refers to the process by which the server sends the processed content to the user's terminal.
[0874] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[0875] Server-side processing
[0876] The server retrieves and analyzes data from the internet, uses generative AI models to transform content according to the user's developmental level, and filters out inappropriate content.
[0877] The server retrieves the accessed data on the Internet. Specifically, it receives the URL of the web page the user is trying to access, sends an HTTP GET request to that URL, and retrieves the HTML data. The server uses an HTML parser such as BeautifulSoup to analyze this data, thereby extracting the main elements of the web page (text, images, links, etc.).
[0878] It then uses a generative AI model (e.g., GPT-3) to adapt the parsed content to the user's developmental level, using the following example prompt:
[0879] "Simplify the text on the web page for children. Enter the text below: '{{Web page text}}'"
[0880] "Parse this HTML data and convert the extracted key content elements to child-friendly. HTML data: '{{HTML data}}'"
[0881] The content converted by the generative AI model then goes through a further filtering process: the server detects inappropriate keywords and phrases based on a list and either removes them or replaces them with harmless language.
[0882] After all the processing is complete, the server returns the converted and filtered content to the device, ensuring that the content is presented in a child-friendly and safe manner.
[0883] Terminal side processing
[0884] The terminal is responsible for appropriately displaying the converted content received from the server to the user.
[0885] When a user accesses a particular web page through a browser, the device sends a request to the server, which includes the URL they are trying to access. The server returns the converted content, which the device then displays in the browser.
[0886] For example, if an elementary school student wants to visit a website called "Animal World," the process would go something like this: The user opens a browser and enters the URL for "Animal World." The device sends a request to the server, which retrieves, analyzes, converts, and filters the content. Finally, the device receives the converted content and displays it appropriately to the user.
[0887] User usage scenarios
[0888] Users operate a browser through the system to access specific web pages. For example, on the "Animal World" website, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." This system allows children to obtain appropriate information safely and in an easy-to-understand manner.
[0889] As described above, the present invention aims to provide appropriate Internet usage according to the user's developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[0890] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0891] Server-side processing
[0892] Step 1:
[0893] The server receives the URL of the web page the user is trying to access, which is obtained by an HTTP GET request sent from the device.
[0894] Input: The URL the user types in the browser
[0895] Output: Retrieved URL
[0896] Step 2:
[0897] The server sends an HTTP GET request to the URL and downloads the HTML data of the entire web page. The server issues the request, obtains the response, and saves it.
[0898] Input: Retrieved URL
[0899] Output: Downloaded HTML data
[0900] Step 3:
[0901] To parse the HTML data, the server uses an HTML parser such as BeautifulSoup to parse the data. The server creates a BeautifulSoup object and parses the HTML data to extract the main elements (text, images, links, etc.).
[0902] Input: Downloaded HTML data
[0903] Output: Extracted key elements (text, images, links, etc.)
[0904] Step 4:
[0905] The server sends prompts to a generative AI model (e.g., GPT-3) and converts the analyzed content according to the user's development level. The server constructs an appropriate prompt sentence for the text portion of the content and sends it to the AI model.
[0906] Input: extracted key elements and user's development level
[0907] Output: The converted text content
[0908] Step 5:
[0909] The server scans the converted content and filters out inappropriate content. Using a list of inappropriate keywords, the server detects these words and phrases and either replaces them with appropriate alternatives or removes them.
[0910] Input: The converted text content
[0911] Output: The filtered text content
[0912] Step 6:
[0913] The server returns the content after all conversion and filtering has been completed to the terminal as an HTTP response, so that the converted content can be displayed appropriately to the user.
[0914] Input: Filtered text content
[0915] Output: The transformed content included in the HTTP response
[0916] Terminal side processing
[0917] Step 1:
[0918] When a user accesses a specific web page through a browser, the device sends the request to the server. The browser issues an HTTP GET request, sending the URL that the user wants to access to the server.
[0919] Input: The URL the user types in the browser
[0920] Output: HTTP GET request to the server
[0921] Step 2:
[0922] The terminal receives the converted content returned from the server, receives the HTTP response, and analyzes and obtains the converted HTML data contained in the response body.
[0923] Input: HTTP response from the server
[0924] Output: Converted HTML data
[0925] Step 3:
[0926] The device displays the received content in the browser, and uses the browser's rendering engine to display the converted content in a format that is easy for children to understand.
[0927] Input: Converted HTML data
[0928] Output: Content displayed in the browser
[0929] User usage scenarios
[0930] 1. Initiating Access
[0931] A user opens a browser and enters the URL of the website they want to visit.
[0932] Input: The URL entered by the user
[0933] Output: The browser issues an access request
[0934] 2. Conversion by the server
[0935] The device sends a request for content to the server, which retrieves, analyzes, converts, and filters it.
[0936] Input: The URL the user types into the browser
[0937] Output: The transformed and filtered content
[0938] 3. Display of Content
[0939] The device receives the converted content and displays it appropriately in the browser. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[0940] Input: The converted content sent from the server
[0941] Output: Content displayed in the browser
[0942] The above are the specific processing steps of the system based on the present invention, which enables the conversion of Internet content according to the user's level of development, filtering out inappropriate content, and providing users with safe and easy-to-understand information.
[0943] (Application example 1)
[0944] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0945] When children browse content online, it is difficult to provide them with content that is converted into an appropriate format according to their age and level of understanding. Furthermore, content viewed online may contain inappropriate information, and there is a lack of means to accurately filter and safely provide this information. Furthermore, there is a need to improve learning outcomes by appropriately adjusting the display of content according to the child's developmental level.
[0946] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0947] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for generating a unique prompt sentence and inputting it to the generative AI model, and means for filtering inappropriate content, thereby enabling the content to be converted into a format that is easy for children to understand and displayed safely.
[0948] "Internet content" refers to all information media, such as text, images, video, and audio, that can be obtained through the Internet.
[0949] "Means of analysis" refers to technical methods for structurally understanding acquired Internet content and extracting each element.
[0950] A "generative AI model" is a collection of machine learning algorithms that use artificial intelligence techniques to generate or transform text or images.
[0951] "Means for conversion according to the user's developmental level" refers to technology for converting content into an appropriate format based on the user's age and level of understanding.
[0952] The "means for displaying the converted content" refers to a device or software for appropriately outputting the converted content to the user's terminal.
[0953] "Means for generating unique prompt sentences and inputting them into a generative AI model" refers to a technical method for creating prompt sentences in a form suitable for a generative AI model and providing them to the model.
[0954] "Inappropriate content filtering measures" are technologies that detect inappropriate keywords and phrases from internet content and either remove them or replace them with harmless ones.
[0955] A "means for adding furigana" is a technique or device for adding readings to difficult characters such as kanji.
[0956] "Means of converting expressions into simple terms" are techniques for converting complex sentences and technical terms into simple terms that children can easily understand.
[0957] "Means for extracting parts that match specific keywords" refers to technology for detecting and extracting specific words or phrases within content.
[0958] "Means for integrating and displaying" refers to the technology or device that compiles and presents the analyzed, transformed, and filtered content to the user.
[0959] This invention is a system that converts content on the Internet according to the user's level of development when the user browses the content, and displays the content safely. This system is composed of a server and a terminal.
[0960] Server-side processing
[0961] 1. Content Acquisition
[0962] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data for that URL, which downloads the entire web page data to the server.
[0963] 2. Content Analysis
[0964] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[0965] 3. Conversion using generative AI models
[0966] The server uses a generative AI model (e.g., GPT-3) based on the analyzed content to convert the content to suit the user's developmental level. The conversion is performed by generating appropriate prompts and providing them as input to the AI model. For example, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions, with furigana added to kanji.
[0967] 4. Filtering inappropriate content
[0968] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[0969] 5. Return of Converted Content
[0970] After all conversion and filtering is complete, the server returns the converted content to the terminal.
[0971] Terminal side processing
[0972] 1. Submitting a Request
[0973] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[0974] 2. Receiving Content
[0975] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[0976] 3. Display of Content
[0977] The device displays the received content in a way that is easy for children to understand, such as sentences using simple words or text with furigana added to kanji characters.
[0978] Content conversion examples
[0979] For example, if an elementary school student wants to browse a website called "Animal World," using this system, the following flow will be displayed: The sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals," making it easier for elementary school students to understand.
[0980] Hardware and software used
[0981] Hardware: Smartphone (device)
[0982] Software: Python, requests library, BeautifulSoup library, OpenAI GPT-3 API
[0983] Prompt Sentence Examples
[0984] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[0985] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[0986] In this way, the present invention aims to provide users with access to Internet content appropriate to their developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[0987] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0988] Step 1:
[0989] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL into the web browser and clicks the access button. At this time, the device sends an HTTP request to the server, along with information about the user's development level. The input is the URL and development level information entered by the user, and the output is the HTTP request sent to the server.
[0990] Step 2:
[0991] The server sends an HTTP GET request to retrieve the HTML data of the received URL. Specifically, the server downloads the entire web page using the specified URL. At this time, the server receives the HTML content of the web page as a response to the request. The input is the URL, and the output is the retrieved HTML data.
[0992] Step 3:
[0993] The server parses the retrieved HTML data using an HTML parser such as BeautifulSoup. Specifically, it extracts the main content elements such as text, images, and links from the HTML data. The input is the retrieved HTML data, and the output is the extracted content elements.
[0994] Step 4:
[0995] The server uses a generative AI model (e.g., GPT-3) to generate a prompt appropriate for the user's developmental level, and then uses the prompt to convert the content. Specifically, it generates an appropriate prompt based on the analyzed content and inputs it into the generative AI model to convert the content into a format appropriate for the user's developmental level. The input is the prompt and the analyzed content, and the output is the converted content. An example of a prompt is as follows:
[0996] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[0997] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[0998] Step 5:
[0999] The server filters the transformed content to remove or replace inappropriate content, specifically detecting inappropriate keywords or phrases and replacing them with harmless terms. The input is the transformed content, and the output is the filtered content.
[1000] Step 6:
[1001] The server returns the filtered content to the terminal. Specifically, it sends the filtered and converted content data to the terminal as an HTTP response. The input is the filtered content, and the output is the data returned to the terminal.
[1002] Step 7:
[1003] The device displays the converted content received from the server to the user. Specifically, it renders the converted content on the browser and displays it in a format that is easy for children to understand. The input is the content data returned from the server, and the output is the content displayed to the user.
[1004] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1005] A specific system for implementing the present invention is composed of elements of a server, a terminal, a user, and an emotion engine.
[1006] Server-side processing
[1007] The server plays an important role in analyzing the acquired internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are explained below.
[1008] 1. Content Acquisition:
[1009] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[1010] 2. Content Analysis:
[1011] To parse the retrieved HTML data, the server uses an HTML parser to extract the main content elements of the web page, such as text, images, and links.
[1012] 3. Transformation by generative AI models:
[1013] The server analyzes the content and then uses a generative AI model to adapt it to the user's developmental level. Specifically, for kindergarteners, it converts sentences into hiragana and uses simple words. For elementary school students, it uses kanji with furigana and adapts it to easier-to-understand expressions.
[1014] 4. Emotion Recognition with Emotion Engine:
[1015] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text to recognize the emotion at that time.
[1016] 5. Emotion-based expression adjustment:
[1017] Depending on the emotion recognized, the generative AI model can further adjust the content it outputs. For example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[1018] 6. Filtering inappropriate content:
[1019] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[1020] 7. Return of Converted Content:
[1021] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[1022] Terminal side processing
[1023] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[1024] 1. Submit your request:
[1025] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[1026] 2. Receiving Content:
[1027] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[1028] 3. Viewing Content:
[1029] The received content is displayed in the browser in a format that is easy for children to understand, such as sentences written in simple language or kanji with furigana.
[1030] User usage scenarios
[1031] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[1032] 1. Initiating access:
[1033] The user opens a browser and enters the URL for "Animal World" to access the site.
[1034] 2. Conversion by the server:
[1035] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[1036] 3. Viewing Content:
[1037] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[1038] As such, the present invention is a system that provides an appropriate Internet usage environment according to the user's developmental level and emotions, and aims to support children's learning and information acquisition, as well as provide a comfortable and secure content experience.
[1039] The processing flow will be explained below.
[1040] Step 1:
[1041] A user launches a browser and attempts to access a particular website (e.g., "Animal World").
[1042] Step 2:
[1043] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[1044] Step 3:
[1045] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[1046] Step 4:
[1047] The server parses the HTML data it retrieves and uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[1048] Step 5:
[1049] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level (kindergartener, elementary school student, etc.) Specifically, it converts sentences into hiragana, uses kanji with furigana, or modifies them into simpler words.
[1050] Step 6:
[1051] The server simultaneously recognizes the user's emotions using an emotion engine, which analyzes facial expressions and tone of voice through a camera and microphone to assess the user's emotional state (e.g., joy, anxiety, excitement, etc.).
[1052] Step 7:
[1053] The server further adjusts the converted content based on the emotions recognized by the emotion engine, for example, changing the text to a more gentle expression and adding an encouraging message if the user is in an anxious state.
[1054] Step 8:
[1055] The server uses filtering algorithms to detect inappropriate keywords and phrases in the content and either remove them or replace them with harmless alternatives.
[1056] Step 9:
[1057] The server sends the converted and filtered content back to the device, which then sends an HTTP response containing the converted HTML data and other resources.
[1058] Step 10:
[1059] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[1060] Step 11:
[1061] The device displays the converted content in the browser. For example, the sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" will be added.
[1062] Through this tailored content, users can access information that is easy to understand and appropriate for their developmental level, and enjoy a comfortable and emotionally sensitive internet experience.
[1063] Example 2
[1064] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1065] With conventional internet usage, it has been difficult to display content that takes into account the user's developmental level and emotions, making it difficult to provide appropriate information to children. Furthermore, filtering of inappropriate content is insufficient, creating a need for a safe internet usage environment.
[1066] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1067] In this invention, the server includes a means for analyzing acquired internet content, a means for using a generative AI model to convert the content according to the user's developmental level, and a means for recognizing the user's emotions and adjusting the presentation of the content based on those emotions. This enables the provision of appropriate and safe information according to the user's developmental level and emotions. Furthermore, by including a means for displaying this content and a means for filtering inappropriate content, a safe and comfortable internet environment for children is provided.
[1068] "Internet content" refers to data such as text, images, videos, and links posted on websites and online platforms.
[1069] "Analysis" refers to the process of breaking down acquired content and understanding and extracting its components and meaning.
[1070] "User development level" refers to a concept that indicates the stage of growth of a user's comprehension, knowledge, language ability, etc.
[1071] "Generative AI model" refers to technology that uses artificial intelligence to generate appropriate content or text based on user input.
[1072] "Recognizing the user's emotions" refers to identifying the user's current emotional state from their facial expressions, tone of voice, input characters, etc.
[1073] "Adjusting content presentation" refers to appropriately changing the way content is delivered and the vocabulary used depending on the perceived emotion and developmental level.
[1074] "Inappropriate content filtering" refers to the process of detecting and removing or replacing content that contains violent, adult, or offensive material before it is displayed to users.
[1075] "JSON format data" refers to a data format structured using JavaScript Object Notation that is easy for both humans and machines to read.
[1076] The present invention is composed of a system including a server, a terminal, a user, a generative AI model, and an emotion engine. This system is designed to provide an appropriate Internet environment especially for children, and is capable of displaying content according to the user's developmental level and emotions.
[1077] The server first retrieves content from the Internet. In this process, it uses Python's requests library to send an HTTP GET request to retrieve HTML data from the specified URL. For example, if a user tries to access "https: / / example.com / animal," the server receives this URL and downloads the corresponding HTML data.
[1078] The server then parses the HTML data using the BeautifulSoup library to extract the main content elements of the web page, such as text, images, and links. For example, <title> Animal World< / title> "or" Elephants are very large herbivores. " HTML elements such as are parsed.
[1079] The server then uses a generative AI model to convert the analyzed content according to the user's developmental level. Here, we use OpenAI's GPT-3 as an example. For example, the sentence "Elephants are very large herbivores" is converted to "Elephants are very large grass-eating animals" using the prompt "Please change this sentence to hiragana." The following prompt sentence is used as an example of input to the generative AI model:
[1080] For kindergarteners: "Change this sentence into hiragana. Elephants are very large herbivores."
[1081] For elementary school students: "Please translate this sentence into a form that elementary school students can easily understand, using furigana. Elephants are very large herbivores."
[1082] The server then uses an emotion engine to recognize the user's emotions. Specifically, it uses Microsoft Azure's Face API to analyze the user's facial expressions from image data and identify the emotion at that time. For example, the server can upload an image of the user's facial expression and obtain emotion data such as "anxiety" or "joy." The results are returned in JSON format.
[1083] Based on the results of emotion recognition, the server further adjusts the content output by the generative AI model. For example, if it recognizes that the user is feeling anxious, it adds an encouraging message to the text: "Want to know more about elephants? Good luck!"
[1084] The server then filters out inappropriate content, using Python's re library to perform keyword filtering to detect and remove or modify "violent" or "inappropriate" language.
[1085] Finally, the server returns the converted content to the device in JSON format as an HTTP response, making it displayable on the device.
[1086] When a user accesses a specific web page through a browser, the device sends a request to the server and receives the converted content returned by the server. The received content is displayed in a format that is easy for children to understand using HTML and JavaScript. For example, if the sentence "Elephants are very big animals that eat grass" is displayed, and the user has an "anxious" expression, the message "Want to know more about elephants? Good luck!" is displayed.
[1087] In this way, the present invention can provide appropriate and safe information according to the user's developmental level and emotions, and can provide a safe and comfortable Internet usage environment for children.
[1088] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1089] Step 1:
[1090] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data. The URL entered by the user (e.g., "https: / / example.com / animal") is the input data, and the HTML data returned in the HTTP response is the output. Specifically, the GET request is sent using the Python requests library.
[1091] Step 2:
[1092] The server analyzes the retrieved HTML data. Here, it uses the BeautifulSoup library to extract the main content (text, images, links, etc.) from the HTML data. The HTML data is the input, and the extracted content (e.g., the text "Elephants are very large herbivores") is the output. Specifically, it breaks down the HTML structure and lists each element.
[1093] Step 3:
[1094] The server converts the analyzed text content using a generative AI model. Here, we use OpenAI's GPT-3 model. The input requires the extracted text content (e.g., "Elephants are very large herbivores") and a prompt sentence (e.g., "Please convert this sentence into hiragana"), and the output is the converted text (e.g., "Elephants are very large grass-eating animals"). Specifically, the server calls the API of the generative AI model, sends the prompt, and obtains the conversion result.
[1095] Step 4:
[1096] The server uses an emotion engine to recognize the user's emotions. Here, it uses an emotion recognition algorithm (e.g., Face API) to extract emotions from the user's facial image and text. The input is an image of the user's facial expression and text, and the output is the user's emotional data (e.g., "anxiety," "joy," etc.). Specifically, it calls the emotion engine's API, sends the input data, and retrieves the results.
[1097] Step 5:
[1098] The server adjusts the content output by the generative AI model based on the recognized emotion data. For example, if the input emotion data is "anxiety," an encouraging message (e.g., "Let's do our best!") is added to the output content. Specifically, the server adds a conditional message to the converted text.
[1099] Step 6:
[1100] The server filters inappropriate content by using the Python re library to detect specific keywords or phrases in the text and remove or modify the inappropriate content. The input is the converted and adjusted text data, and the output is the text data with the inappropriate content filtered out. Specifically, it uses regular expressions to find inappropriate keywords and replace them.
[1101] Step 7:
[1102] The server returns the converted content to the device. It sends the converted content in JSON format as an HTTP response, making it displayable on the device. The input is the filtered text data, and the output is JSON data sent to the device. The specific operation is to generate an HTTP response that includes the converted content.
[1103] Step 8:
[1104] The terminal receives the converted content returned from the server. In this case, it receives the HTTP response sent from the server. The input is the HTTP response from the server, and the output is the converted content data. Specifically, it analyzes the HTTP response and extracts the content portion.
[1105] Step 9:
[1106] The device displays the received content in a browser. Specifically, it displays the retrieved text and images using HTML and JavaScript. The input is the extracted content data, and the output is the display content as a user interface. Specifically, it performs DOM operations on the HTML document and inserts the text and images into the specified location.
[1107] (Application example 2)
[1108] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1109] Conventional Internet content distribution systems do not adequately adapt content to suit the user's age and developmental level. They also lack the ability to recognize the user's emotional state in real time and adjust content accordingly. This makes it difficult to provide appropriate content that is easy to understand and safe to use, especially for children. Furthermore, inappropriate content filtering is often inadequate, making it difficult to provide a safe and secure environment. There is a need for a system that solves these issues.
[1110] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1111] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for recognizing the user's emotions, means for adjusting the content based on the recognized emotions, and means for filtering inappropriate content. This makes it possible to provide content appropriate for the user's age and developmental level, as well as to recognize the user's emotional state in real time to display optimal content, providing an environment in which the user can use the content with peace of mind.
[1112] "Internet content" refers to information and multimedia data accessible via the Internet, such as websites, blogs, and video sharing sites.
[1113] "Means of analysis" refers to systems or algorithms that analyze acquired internet content such as text, images, and videos to understand its structure and content.
[1114] "User development level" refers to the stage of cognitive ability according to the user's age, knowledge, and comprehension, and is a standard for providing appropriate content based on this.
[1115] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning techniques to generate new data based on given data.
[1116] A "content transformation means" is a system or algorithm that transforms acquired content into a form appropriate for the user's developmental level.
[1117] A "means for displaying" is a system or algorithm that displays the transformed content on an output device (such as a screen or monitor) in a form that is easy for the user to understand.
[1118] "Means for recognizing emotions" refers to systems or algorithms that analyze a user's facial expressions and voice to identify their current emotional state.
[1119] "Adjustment" means a system or algorithm that appropriately modifies the presentation or substance of content based on the perceived sentiment.
[1120] "Inappropriate content" is content that includes information or expressions that are deemed harmful or inappropriate to users.
[1121] "Filtering measures" are systems or algorithms that detect inappropriate content and either remove it or convert it into a harmless form.
[1122] "Furigana" is kana characters added to indicate how to read kanji characters, and is an auxiliary character that makes it easier for users, especially those in the developmental stage, to understand kanji characters.
[1123] A specific system for implementing this invention is composed of elements such as a server, a terminal, a user, and an emotion engine. This system can convert Internet content appropriately according to the user's developmental level and emotional state, and provide it in a form that can be used safely.
[1124] Server-side processing
[1125] The server plays an important role in analyzing the retrieved internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are as follows:
[1126] 1. Content Acquisition:
[1127] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[1128] 2. Content Analysis:
[1129] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[1130] 3. Transformation by generative AI models:
[1131] The server analyzes the content and then uses a generative AI model (e.g., a model using TensorFlow) to adapt the content to the user's developmental level. Specifically, for kindergarteners, the server converts sentences into hiragana and uses simple words. For elementary school students, the server converts the content into easy-to-understand expressions using kanji with furigana.
[1132] 4. Emotion Recognition with Emotion Engine:
[1133] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text, and recognizes the emotion at that time.
[1134] 5. Emotion-based expression adjustment:
[1135] Depending on the recognized emotion, the generative AI model can further adjust the content it outputs. For example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[1136] 6. Filtering inappropriate content:
[1137] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[1138] 7. Return of Converted Content:
[1139] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays the content in a user-friendly format.
[1140] Terminal side processing
[1141] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[1142] 1. Submit your request:
[1143] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[1144] 2. Receiving Content:
[1145] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[1146] 3. Viewing Content:
[1147] The received content is displayed in the browser in a format that is easy for the user to understand. For example, for kindergarteners, sentences are converted into hiragana, sentences expressed in simple words, and kanji with furigana are used. Also, if the user looks anxious, an encouraging message is added.
[1148] User usage scenarios
[1149] A user accesses a specific web page through the system. For example, if a child wants to browse a website called "Animal World," the following process will occur when using this system:
[1150] Access Start:
[1151] The user opens a browser and enters the URL for "Animal World" to access the site.
[1152] Server-based conversion:
[1153] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[1154] Show Content:
[1155] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[1156] Example of generative AI model prompt
[1157] User: 5 years old
[1158] Emotion: Anxiety
[1159] Text: Please convert and display the content of "Dog wagging its tail".
[1160] Prompt result: Generates a kind sentence about a dog wagging its tail and a sentence containing an encouraging message.
[1161] In this way, the system provides appropriate content according to the user's developmental level and emotional state, creating an environment in which the user can use the system with peace of mind.
[1162] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1163] Step 1:
[1164] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL and presses the access button, which generates an HTTP request and sends the request to the server. The input data is the URL, and the output data is the HTTP request.
[1165] Step 2:
[1166] The server receives a request and retrieves HTML data from a specified URL on the Internet. The server sends an HTTP GET request to download the entire specified web page. The input data is the URL and the HTTP request, and the output data is the HTML content. Specifically, the server receives the HTML data as a response.
[1167] Step 3:
[1168] The server uses an HTML parser (e.g., BeautifulSoup) to parse the HTML data it receives. The parsing extracts the main content elements in the web page, such as text, images, and links. The input data is the HTML content, and the output data are the parsed content elements. The server prepares these elements to be stored in a database.
[1169] Step 4:
[1170] The server uses a generative AI model (e.g., a TensorFlow model) to convert the analyzed content elements according to the user's developmental level. For example, for kindergarteners, the server converts sentences into hiragana and uses simple words. The input data are the analyzed content elements and the user's developmental level information, and the output data is the converted content. Specifically, the AI model processes the text data and converts it into an appropriate expression.
[1171] Step 5:
[1172] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze and recognize the user's facial expressions, tone of voice, and emotional expressions in the input text. The input data is the user's facial image and voice data, and the output data is the recognized emotional state. The server prepares to adjust the content based on this data.
[1173] Step 6:
[1174] Based on the recognized emotions, the server further adjusts the content output by the generative AI model. For example, if the user is anxious, it converts the content into gentler language and adds an encouraging message. The input data is the converted content and emotion data, and the output data is the adjusted content. Specifically, the AI model adds text data and makes changes according to the emotion.
[1175] Step 7:
[1176] The server filters inappropriate content through content analysis and conversion. Filtering algorithms are used to detect inappropriate keywords and phrases and remove or replace them with harmless alternatives. The input data is the converted content, and the output data is the filtered, safe content.
[1177] Step 8:
[1178] After all conversion and filtering is complete, the server returns the converted content to the terminal. The input data is the filtered content, and the output data is the HTTP response. This allows the content to be displayed on the terminal in a format that is easy for the user to understand.
[1179] Step 9:
[1180] The device receives the converted content sent back from the server and displays it appropriately to the user. The input data is the converted content, and the output data is the content to be displayed. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." Also, if the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[1181] In this way, the system provides appropriate Internet content according to the user's developmental level and emotional state, creating an environment in which the user can use the content with peace of mind.
[1182] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1183] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1184] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1185] [Fourth embodiment]
[1186] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1187] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1188] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1189] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1190] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1191] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1192] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1193] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1194] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1195] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1196] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1197] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1198] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1199] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[1200] Server-side processing
[1201] The server plays an important role in analyzing the acquired Internet content and converting it according to the user's level of development. Specifically, the process is as follows:
[1202] 1. Content Acquisition:
[1203] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and sends an HTTP GET request to download the entire web page.
[1204] 2. Content Analysis:
[1205] To parse the retrieved HTML data, the server uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[1206] 3. Transformation by generative AI models:
[1207] After analyzing the content, the server uses a generative AI model to convert it to suit the user's developmental level. Specifically, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions while adding furigana to kanji. This process utilizes a generative AI model (e.g., GPT-3).
[1208] 4. Filtering inappropriate content:
[1209] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[1210] 5. Return of Converted Content:
[1211] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[1212] Terminal side processing
[1213] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[1214] 1. Submit your request:
[1215] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[1216] 2. Receiving Content:
[1217] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[1218] 3. Viewing Content:
[1219] The received content is displayed in the browser in a format that is easy for children to understand. For example, sentences that make extensive use of hiragana and kanji with furigana are used.
[1220] User usage scenarios
[1221] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[1222] 1. Initiating access:
[1223] The user opens a browser and enters the URL for "Animal World" to access the site.
[1224] 2. Conversion by the server:
[1225] The device sends a request for content to the server, which retrieves, analyzes, and converts it.
[1226] 3. Viewing Content:
[1227] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[1228] As such, the present invention is a system that provides appropriate Internet usage according to the user's developmental level, and aims to promote children's learning and information acquisition while protecting them from inappropriate content.
[1229] The processing flow will be explained below.
[1230] Step 1:
[1231] A user launches a browser and attempts to access a particular website.
[1232] Step 2:
[1233] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[1234] Step 3:
[1235] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[1236] Step 4:
[1237] Parse the HTML data retrieved by the server using an HTML parser (such as BeautifulSoup) to understand the overall structure of the web page and extract the main content elements (text, images, links, etc.).
[1238] Step 5:
[1239] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level. Specifically, for kindergarteners, the text is converted into hiragana and simple words are used, while for elementary school students, kanji with furigana is used, making it easier to understand.
[1240] Step 6:
[1241] The server filters inappropriate content during the conversion process, using a filtering algorithm to detect inappropriate keywords and phrases and remove them or replace them with harmless alternatives.
[1242] Step 7:
[1243] The server returns the converted content to the device, and sends an HTTP response containing the converted HTML data and other resources.
[1244] Step 8:
[1245] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[1246] Step 9:
[1247] The device then displays the converted content to the user, rendering the web page in a way that is easy for the user to understand, such as displaying sentences that make heavy use of hiragana or kanji with furigana.
[1248] Example 1
[1249] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1250] There is a wide variety of content on the Internet, and understanding and safety of that content poses challenges, especially when children access it. Current technology does not adequately convert content or filter inappropriate content, making it difficult to provide information in a format that is easy for children to understand. Furthermore, the need for manual content filtering and the lack of automated technology to convert content into appropriate expressions according to the user's developmental level result in an inconsistent provision of educational value.
[1251] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1252] In this invention, the server includes means for acquiring accessed internet data, means for analyzing the acquired HTML data and extracting key elements, means for converting the content according to the user's developmental level using a generative AI model, means for filtering the converted content, and means for returning the converted and filtered content to the terminal. This makes it possible to automatically analyze, convert, and filter internet content accessed by children and provide appropriate information according to their developmental level.
[1253] "Accessed Internet Data" is information obtained from the web pages or online resources that a user attempts to access.
[1254] "HTML data" means data written in the HyperText Markup Language used to describe the structure and content of web pages.
[1255] "Key elements" are the core content of a Web page, such as text, images, links, and headings.
[1256] A "generative AI model" is an artificial intelligence algorithm that generates appropriate content based on given input, particularly one based on large-scale natural language processing techniques.
[1257] "User's developmental level" refers to the level of content appropriate for the user's age and learning stage.
[1258] The "means of transforming content" refers to the process of using the generative AI model described above to transform the original content into an appropriate format according to the user's developmental level.
[1259] "Content filtering measures" are the processes that detect inappropriate content or keywords and remove or replace them with harmless forms.
[1260] A "terminal" is an electronic device that a user directly uses as an interface, such as a personal computer, smartphone, or tablet.
[1261] "Means for returning content to the terminal" refers to the process by which the server sends the processed content to the user's terminal.
[1262] A specific system for implementing the present invention comprises elements of a server, a terminal, and a user.
[1263] Server-side processing
[1264] The server retrieves and analyzes data from the internet, uses generative AI models to transform content according to the user's developmental level, and filters out inappropriate content.
[1265] The server retrieves the accessed data on the Internet. Specifically, it receives the URL of the web page the user is trying to access, sends an HTTP GET request to that URL, and retrieves the HTML data. The server uses an HTML parser such as BeautifulSoup to analyze this data, thereby extracting the main elements of the web page (text, images, links, etc.).
[1266] It then uses a generative AI model (e.g., GPT-3) to adapt the parsed content to the user's developmental level, using the following example prompt:
[1267] "Simplify the text on the web page for children. Enter the text below: '{{Web page text}}'"
[1268] "Parse this HTML data and convert the extracted key content elements to child-friendly. HTML data: '{{HTML data}}'"
[1269] The content converted by the generative AI model then goes through a further filtering process: the server detects inappropriate keywords and phrases based on a list and either removes them or replaces them with harmless language.
[1270] After all the processing is complete, the server returns the converted and filtered content to the device, ensuring that the content is presented in a child-friendly and safe manner.
[1271] Terminal side processing
[1272] The terminal is responsible for appropriately displaying the converted content received from the server to the user.
[1273] When a user accesses a particular web page through a browser, the device sends a request to the server, which includes the URL they are trying to access. The server returns the converted content, which the device then displays in the browser.
[1274] For example, if an elementary school student wants to visit a website called "Animal World," the process would go something like this: The user opens a browser and enters the URL for "Animal World." The device sends a request to the server, which retrieves, analyzes, converts, and filters the content. Finally, the device receives the converted content and displays it appropriately to the user.
[1275] User usage scenarios
[1276] Users operate a browser through the system to access specific web pages. For example, on the "Animal World" website, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." This system allows children to obtain appropriate information safely and in an easy-to-understand manner.
[1277] As described above, the present invention aims to provide appropriate Internet usage according to the user's developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[1278] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1279] Server-side processing
[1280] Step 1:
[1281] The server receives the URL of the web page the user is trying to access, which is obtained by an HTTP GET request sent from the device.
[1282] Input: The URL the user types in the browser
[1283] Output: Retrieved URL
[1284] Step 2:
[1285] The server sends an HTTP GET request to the URL and downloads the HTML data of the entire web page. The server issues the request, obtains the response, and saves it.
[1286] Input: Retrieved URL
[1287] Output: Downloaded HTML data
[1288] Step 3:
[1289] To parse the HTML data, the server uses an HTML parser such as BeautifulSoup to parse the data. The server creates a BeautifulSoup object and parses the HTML data to extract the main elements (text, images, links, etc.).
[1290] Input: Downloaded HTML data
[1291] Output: Extracted key elements (text, images, links, etc.)
[1292] Step 4:
[1293] The server sends prompts to a generative AI model (e.g., GPT-3) and converts the analyzed content according to the user's development level. The server constructs an appropriate prompt sentence for the text portion of the content and sends it to the AI model.
[1294] Input: extracted key elements and user's development level
[1295] Output: The converted text content
[1296] Step 5:
[1297] The server scans the converted content and filters out inappropriate content. Using a list of inappropriate keywords, the server detects these words and phrases and either replaces them with appropriate alternatives or removes them.
[1298] Input: The converted text content
[1299] Output: The filtered text content
[1300] Step 6:
[1301] The server returns the content after all conversion and filtering has been completed to the terminal as an HTTP response, so that the converted content can be displayed appropriately to the user.
[1302] Input: Filtered text content
[1303] Output: The transformed content included in the HTTP response
[1304] Terminal side processing
[1305] Step 1:
[1306] When a user accesses a specific web page through a browser, the device sends the request to the server. The browser issues an HTTP GET request, sending the URL that the user wants to access to the server.
[1307] Input: The URL the user types in the browser
[1308] Output: HTTP GET request to the server
[1309] Step 2:
[1310] The terminal receives the converted content returned from the server, receives the HTTP response, and analyzes and obtains the converted HTML data contained in the response body.
[1311] Input: HTTP response from the server
[1312] Output: Converted HTML data
[1313] Step 3:
[1314] The device displays the received content in the browser, and uses the browser's rendering engine to display the converted content in a format that is easy for children to understand.
[1315] Input: Converted HTML data
[1316] Output: Content displayed in the browser
[1317] User usage scenarios
[1318] 1. Initiating Access
[1319] A user opens a browser and enters the URL of the website they want to visit.
[1320] Input: The URL entered by the user
[1321] Output: The browser issues an access request
[1322] 2. Conversion by the server
[1323] The device sends a request for content to the server, which retrieves, analyzes, converts, and filters it.
[1324] Input: The URL the user types into the browser
[1325] Output: The transformed and filtered content
[1326] 3. Display of Content
[1327] The device receives the converted content and displays it appropriately in the browser. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals."
[1328] Input: The converted content sent from the server
[1329] Output: Content displayed in the browser
[1330] The above are the specific processing steps of the system based on the present invention, which enables the conversion of Internet content according to the user's level of development, filtering out inappropriate content, and providing users with safe and easy-to-understand information.
[1331] (Application example 1)
[1332] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1333] When children browse content online, it is difficult to provide them with content that is converted into an appropriate format according to their age and level of understanding. Furthermore, content viewed online may contain inappropriate information, and there is a lack of means to accurately filter and safely provide this information. Furthermore, there is a need to improve learning outcomes by appropriately adjusting the display of content according to the child's developmental level.
[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1335] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for generating a unique prompt sentence and inputting it to the generative AI model, and means for filtering inappropriate content, thereby enabling the content to be converted into a format that is easy for children to understand and displayed safely.
[1336] "Internet content" refers to all information media, such as text, images, video, and audio, that can be obtained through the Internet.
[1337] "Means of analysis" refers to technical methods for structurally understanding acquired Internet content and extracting each element.
[1338] A "generative AI model" is a collection of machine learning algorithms that use artificial intelligence techniques to generate or transform text or images.
[1339] "Means for conversion according to the user's developmental level" refers to technology for converting content into an appropriate format based on the user's age and level of understanding.
[1340] The "means for displaying the converted content" refers to a device or software for appropriately outputting the converted content to the user's terminal.
[1341] "Means for generating unique prompt sentences and inputting them into a generative AI model" refers to a technical method for creating prompt sentences in a form suitable for a generative AI model and providing them to the model.
[1342] "Inappropriate content filtering measures" are technologies that detect inappropriate keywords and phrases from internet content and either remove them or replace them with harmless ones.
[1343] A "means for adding furigana" is a technique or device for adding readings to difficult characters such as kanji.
[1344] "Means of converting expressions into simple terms" are techniques for converting complex sentences and technical terms into simple terms that children can easily understand.
[1345] "Means for extracting parts that match specific keywords" refers to technology for detecting and extracting specific words or phrases within content.
[1346] "Means for integrating and displaying" refers to the technology or device that compiles and presents the analyzed, transformed, and filtered content to the user.
[1347] This invention is a system that converts content on the Internet according to the user's level of development when the user browses the content, and displays the content safely. This system is composed of a server and a terminal.
[1348] Server-side processing
[1349] 1. Content Acquisition
[1350] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data for that URL, which downloads the entire web page data to the server.
[1351] 2. Content Analysis
[1352] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[1353] 3. Conversion using generative AI models
[1354] The server uses a generative AI model (e.g., GPT-3) based on the analyzed content to convert the content to suit the user's developmental level. The conversion is performed by generating appropriate prompts and providing them as input to the AI model. For example, for kindergarteners, the text uses a lot of hiragana and is converted into simpler words. For elementary school students, the text is converted into easy-to-understand expressions, with furigana added to kanji.
[1355] 4. Filtering inappropriate content
[1356] The server filters inappropriate content during content analysis and conversion, removing or replacing inappropriate keywords or phrases with harmless alternatives.
[1357] 5. Return of Converted Content
[1358] After all conversion and filtering is complete, the server returns the converted content to the terminal.
[1359] Terminal side processing
[1360] 1. Submitting a Request
[1361] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[1362] 2. Receiving Content
[1363] The terminal receives the converted content returned from the server, the content being adapted to the user's developmental level.
[1364] 3. Display of Content
[1365] The device displays the received content in a way that is easy for children to understand, such as sentences using simple words or text with furigana added to kanji characters.
[1366] Content conversion examples
[1367] For example, if an elementary school student wants to browse a website called "Animal World," using this system, the following flow will be displayed: The sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals," making it easier for elementary school students to understand.
[1368] Hardware and software used
[1369] Hardware: Smartphone (device)
[1370] Software: Python, requests library, BeautifulSoup library, OpenAI GPT-3 API
[1371] Prompt Sentence Examples
[1372] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[1373] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[1374] In this way, the present invention aims to provide users with access to Internet content appropriate to their developmental level, promote children's learning and information acquisition, and protect them from inappropriate content.
[1375] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1376] Step 1:
[1377] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL into the web browser and clicks the access button. At this time, the device sends an HTTP request to the server, along with information about the user's development level. The input is the URL and development level information entered by the user, and the output is the HTTP request sent to the server.
[1378] Step 2:
[1379] The server sends an HTTP GET request to retrieve the HTML data of the received URL. Specifically, the server downloads the entire web page using the specified URL. At this time, the server receives the HTML content of the web page as a response to the request. The input is the URL, and the output is the retrieved HTML data.
[1380] Step 3:
[1381] The server parses the retrieved HTML data using an HTML parser such as BeautifulSoup. Specifically, it extracts the main content elements such as text, images, and links from the HTML data. The input is the retrieved HTML data, and the output is the extracted content elements.
[1382] Step 4:
[1383] The server uses a generative AI model (e.g., GPT-3) to generate a prompt appropriate for the user's developmental level, and then uses the prompt to convert the content. Specifically, it generates an appropriate prompt based on the analyzed content and inputs it into the generative AI model to convert the content into a format appropriate for the user's developmental level. The input is the prompt and the analyzed content, and the output is the converted content. An example of a prompt is as follows:
[1384] Kindergarten prompt: "Rewrite this sentence in simple language for kindergarteners: {website content}"
[1385] Elementary school prompt: "Please rewrite this sentence for elementary school students, adding furigana to the kanji: {website content}"
[1386] Step 5:
[1387] The server filters the transformed content to remove or replace inappropriate content, specifically detecting inappropriate keywords or phrases and replacing them with harmless terms. The input is the transformed content, and the output is the filtered content.
[1388] Step 6:
[1389] The server returns the filtered content to the terminal. Specifically, it sends the filtered and converted content data to the terminal as an HTTP response. The input is the filtered content, and the output is the data returned to the terminal.
[1390] Step 7:
[1391] The device displays the converted content received from the server to the user. Specifically, it renders the converted content on the browser and displays it in a format that is easy for children to understand. The input is the content data returned from the server, and the output is the content displayed to the user.
[1392] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1393] A specific system for implementing the present invention is composed of elements of a server, a terminal, a user, and an emotion engine.
[1394] Server-side processing
[1395] The server plays an important role in analyzing the acquired internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are explained below.
[1396] 1. Content Acquisition:
[1397] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[1398] 2. Content Analysis:
[1399] To parse the retrieved HTML data, the server uses an HTML parser to extract the main content elements of the web page, such as text, images, and links.
[1400] 3. Transformation by generative AI models:
[1401] The server analyzes the content and then uses a generative AI model to adapt it to the user's developmental level. Specifically, for kindergarteners, it converts sentences into hiragana and uses simple words. For elementary school students, it uses kanji with furigana and adapts it to easier-to-understand expressions.
[1402] 4. Emotion Recognition with Emotion Engine:
[1403] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text to recognize the emotion at that time.
[1404] 5. Emotion-based expression adjustment:
[1405] Depending on the emotion recognized, the generative AI model can further adjust the content it outputs. For example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[1406] 6. Filtering inappropriate content:
[1407] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[1408] 7. Return of Converted Content:
[1409] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays it in a child-friendly format.
[1410] Terminal side processing
[1411] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[1412] 1. Submit your request:
[1413] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[1414] 2. Receiving Content:
[1415] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[1416] 3. Viewing Content:
[1417] The received content is displayed in the browser in a format that is easy for children to understand, such as sentences written in simple language or kanji with furigana.
[1418] User usage scenarios
[1419] The user operates a browser through the system to access a specific web page. As a specific example, if an elementary school student wants to view a website called "Animal World," the following process will occur when using this system.
[1420] 1. Initiating access:
[1421] The user opens a browser and enters the URL for "Animal World" to access the site.
[1422] 2. Conversion by the server:
[1423] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[1424] 3. Viewing Content:
[1425] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[1426] As such, the present invention is a system that provides an appropriate Internet usage environment according to the user's developmental level and emotions, and aims to support children's learning and information acquisition, as well as provide a comfortable and secure content experience.
[1427] The processing flow will be explained below.
[1428] Step 1:
[1429] A user launches a browser and attempts to access a particular website (e.g., "Animal World").
[1430] Step 2:
[1431] The terminal receives a user request and sends a request to the server to acquire the content of the specified URL.
[1432] Step 3:
[1433] The server receives the request and sends an HTTP GET request to the specified URL to retrieve the HTML data of the web page.
[1434] Step 4:
[1435] The server parses the HTML data it retrieves and uses an HTML parser (such as BeautifulSoup) to extract the main content elements of the web page, such as text, images, and links.
[1436] Step 5:
[1437] The server inputs the extracted content into a generative AI model, which converts it according to the user's developmental level (kindergartener, elementary school student, etc.) Specifically, it converts sentences into hiragana, uses kanji with furigana, or modifies them into simpler words.
[1438] Step 6:
[1439] The server simultaneously recognizes the user's emotions using an emotion engine, which analyzes facial expressions and tone of voice through a camera and microphone to assess the user's emotional state (e.g., joy, anxiety, excitement, etc.).
[1440] Step 7:
[1441] The server further adjusts the converted content based on the emotions recognized by the emotion engine, for example, changing the text to a more gentle expression and adding an encouraging message if the user is in an anxious state.
[1442] Step 8:
[1443] The server uses filtering algorithms to detect inappropriate keywords and phrases in the content and either remove them or replace them with harmless alternatives.
[1444] Step 9:
[1445] The server sends the converted and filtered content back to the device, which then sends an HTTP response containing the converted HTML data and other resources.
[1446] Step 10:
[1447] The terminal analyzes the converted content received from the server and prepares it for display in the browser.
[1448] Step 11:
[1449] The device displays the converted content in the browser. For example, the sentence "Elephants are very large herbivores" will be displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" will be added.
[1450] Through this tailored content, users can access information that is easy to understand and appropriate for their developmental level, and enjoy a comfortable and emotionally sensitive internet experience.
[1451] Example 2
[1452] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1453] With conventional internet usage, it has been difficult to display content that takes into account the user's developmental level and emotions, making it difficult to provide appropriate information to children. Furthermore, filtering of inappropriate content is insufficient, creating a need for a safe internet usage environment.
[1454] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1455] In this invention, the server includes a means for analyzing acquired internet content, a means for using a generative AI model to convert the content according to the user's developmental level, and a means for recognizing the user's emotions and adjusting the presentation of the content based on those emotions. This enables the provision of appropriate and safe information according to the user's developmental level and emotions. Furthermore, by including a means for displaying this content and a means for filtering inappropriate content, a safe and comfortable internet environment for children is provided.
[1456] "Internet content" refers to data such as text, images, videos, and links posted on websites and online platforms.
[1457] "Analysis" refers to the process of breaking down acquired content and understanding and extracting its components and meaning.
[1458] "User development level" refers to a concept that indicates the stage of growth of a user's comprehension, knowledge, language ability, etc.
[1459] "Generative AI model" refers to technology that uses artificial intelligence to generate appropriate content or text based on user input.
[1460] "Recognizing the user's emotions" refers to identifying the user's current emotional state from their facial expressions, tone of voice, input characters, etc.
[1461] "Adjusting content presentation" refers to appropriately changing the way content is delivered and the vocabulary used depending on the perceived emotion and developmental level.
[1462] "Inappropriate content filtering" refers to the process of detecting and removing or replacing content that contains violent, adult, or offensive material before it is displayed to users.
[1463] "JSON format data" refers to a data format structured using JavaScript Object Notation that is easy for both humans and machines to read.
[1464] The present invention is composed of a system including a server, a terminal, a user, a generative AI model, and an emotion engine. This system is designed to provide an appropriate Internet environment especially for children, and is capable of displaying content according to the user's developmental level and emotions.
[1465] The server first retrieves content from the Internet. In this process, it uses Python's requests library to send an HTTP GET request to retrieve HTML data from the specified URL. For example, if a user tries to access "https: / / example.com / animal," the server receives this URL and downloads the corresponding HTML data.
[1466] The server then parses the HTML data using the BeautifulSoup library to extract the main content elements of the web page, such as text, images, and links. For example, <title> Animal World< / title> "or" Elephants are very large herbivores. " HTML elements such as are parsed.
[1467] The server then uses a generative AI model to convert the analyzed content according to the user's developmental level. Here, we use OpenAI's GPT-3 as an example. For example, the sentence "Elephants are very large herbivores" is converted to "Elephants are very large grass-eating animals" using the prompt "Please change this sentence to hiragana." The following prompt sentence is used as an example of input to the generative AI model:
[1468] For kindergarteners: "Change this sentence into hiragana. Elephants are very large herbivores."
[1469] For elementary school students: "Please translate this sentence into a form that elementary school students can easily understand, using furigana. Elephants are very large herbivores."
[1470] The server then uses an emotion engine to recognize the user's emotions. Specifically, it uses Microsoft Azure's Face API to analyze the user's facial expressions from image data and identify the emotion at that time. For example, the server can upload an image of the user's facial expression and obtain emotion data such as "anxiety" or "joy." The results are returned in JSON format.
[1471] Based on the results of emotion recognition, the server further adjusts the content output by the generative AI model. For example, if it recognizes that the user is feeling anxious, it adds an encouraging message to the text: "Want to know more about elephants? Good luck!"
[1472] The server then filters out inappropriate content, using Python's re library to perform keyword filtering to detect and remove or modify "violent" or "inappropriate" language.
[1473] Finally, the server returns the converted content to the device in JSON format as an HTTP response, making it displayable on the device.
[1474] When a user accesses a specific web page through a browser, the device sends a request to the server and receives the converted content returned by the server. The received content is displayed in a format that is easy for children to understand using HTML and JavaScript. For example, if the sentence "Elephants are very big animals that eat grass" is displayed, and the user has an "anxious" expression, the message "Want to know more about elephants? Good luck!" is displayed.
[1475] In this way, the present invention can provide appropriate and safe information according to the user's developmental level and emotions, and can provide a safe and comfortable Internet usage environment for children.
[1476] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1477] Step 1:
[1478] The server receives the URL of the web page the user is trying to access and sends an HTTP GET request to retrieve the HTML data. The URL entered by the user (e.g., "https: / / example.com / animal") is the input data, and the HTML data returned in the HTTP response is the output. Specifically, the GET request is sent using the Python requests library.
[1479] Step 2:
[1480] The server analyzes the retrieved HTML data. Here, it uses the BeautifulSoup library to extract the main content (text, images, links, etc.) from the HTML data. The HTML data is the input, and the extracted content (e.g., the text "Elephants are very large herbivores") is the output. Specifically, it breaks down the HTML structure and lists each element.
[1481] Step 3:
[1482] The server converts the analyzed text content using a generative AI model. Here, we use OpenAI's GPT-3 model. The input requires the extracted text content (e.g., "Elephants are very large herbivores") and a prompt sentence (e.g., "Please convert this sentence into hiragana"), and the output is the converted text (e.g., "Elephants are very large grass-eating animals"). Specifically, the server calls the API of the generative AI model, sends the prompt, and obtains the conversion result.
[1483] Step 4:
[1484] The server uses an emotion engine to recognize the user's emotions. Here, it uses an emotion recognition algorithm (e.g., Face API) to extract emotions from the user's facial image and text. The input is an image of the user's facial expression and text, and the output is the user's emotional data (e.g., "anxiety," "joy," etc.). Specifically, it calls the emotion engine's API, sends the input data, and retrieves the results.
[1485] Step 5:
[1486] The server adjusts the content output by the generative AI model based on the recognized emotion data. For example, if the input emotion data is "anxiety," an encouraging message (e.g., "Let's do our best!") is added to the output content. Specifically, the server adds a conditional message to the converted text.
[1487] Step 6:
[1488] The server filters inappropriate content by using the Python re library to detect specific keywords or phrases in the text and remove or modify the inappropriate content. The input is the converted and adjusted text data, and the output is the text data with the inappropriate content filtered out. Specifically, it uses regular expressions to find inappropriate keywords and replace them.
[1489] Step 7:
[1490] The server returns the converted content to the device. It sends the converted content in JSON format as an HTTP response, making it displayable on the device. The input is the filtered text data, and the output is JSON data sent to the device. The specific operation is to generate an HTTP response that includes the converted content.
[1491] Step 8:
[1492] The terminal receives the converted content returned from the server. In this case, it receives the HTTP response sent from the server. The input is the HTTP response from the server, and the output is the converted content data. Specifically, it analyzes the HTTP response and extracts the content portion.
[1493] Step 9:
[1494] The device displays the received content in a browser. Specifically, it displays the retrieved text and images using HTML and JavaScript. The input is the extracted content data, and the output is the display content as a user interface. Specifically, it performs DOM operations on the HTML document and inserts the text and images into the specified location.
[1495] (Application example 2)
[1496] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1497] Conventional Internet content distribution systems do not adequately adapt content to suit the user's age and developmental level. They also lack the ability to recognize the user's emotional state in real time and adjust content accordingly. This makes it difficult to provide appropriate content that is easy to understand and safe to use, especially for children. Furthermore, inappropriate content filtering is often inadequate, making it difficult to provide a safe and secure environment. There is a need for a system that solves these issues.
[1498] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1499] In this invention, the server includes means for analyzing acquired internet content, means for using a generative AI model to convert the content according to the user's developmental level, means for displaying the converted content, means for recognizing the user's emotions, means for adjusting the content based on the recognized emotions, and means for filtering inappropriate content. This makes it possible to provide content appropriate for the user's age and developmental level, as well as to recognize the user's emotional state in real time to display optimal content, providing an environment in which the user can use the content with peace of mind.
[1500] "Internet content" refers to information and multimedia data accessible via the Internet, such as websites, blogs, and video sharing sites.
[1501] "Means of analysis" refers to systems or algorithms that analyze acquired internet content such as text, images, and videos to understand its structure and content.
[1502] "User development level" refers to the stage of cognitive ability according to the user's age, knowledge, and comprehension, and is a standard for providing appropriate content based on this.
[1503] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning techniques to generate new data based on given data.
[1504] A "content transformation means" is a system or algorithm that transforms acquired content into a form appropriate for the user's developmental level.
[1505] A "means for displaying" is a system or algorithm that displays the transformed content on an output device (such as a screen or monitor) in a form that is easy for the user to understand.
[1506] "Means for recognizing emotions" refers to systems or algorithms that analyze a user's facial expressions and voice to identify their current emotional state.
[1507] "Adjustment" means a system or algorithm that appropriately modifies the presentation or substance of content based on the perceived sentiment.
[1508] "Inappropriate content" is content that includes information or expressions that are deemed harmful or inappropriate to users.
[1509] "Filtering measures" are systems or algorithms that detect inappropriate content and either remove it or convert it into a harmless form.
[1510] "Furigana" is kana characters added to indicate how to read kanji characters, and is an auxiliary character that makes it easier for users, especially those in the developmental stage, to understand kanji characters.
[1511] A specific system for implementing this invention is composed of elements such as a server, a terminal, a user, and an emotion engine. This system can convert Internet content appropriately according to the user's developmental level and emotional state, and provide it in a form that can be used safely.
[1512] Server-side processing
[1513] The server plays an important role in analyzing the retrieved internet content and converting it according to the user's developmental level. It also recognizes the user's emotions and adjusts the expressions based on those emotions. The specific steps are as follows:
[1514] 1. Content Acquisition:
[1515] The server receives the URL of the web page the user is trying to access, retrieves the HTML data for that URL, and then sends an HTTP GET request to download the entire web page.
[1516] 2. Content Analysis:
[1517] The server uses an HTML parser (e.g., BeautifulSoup) to parse the retrieved HTML data, extracting the main content elements of the web page, such as text, images, and links.
[1518] 3. Transformation by generative AI models:
[1519] The server analyzes the content and then uses a generative AI model (e.g., a model using TensorFlow) to adapt the content to the user's developmental level. Specifically, for kindergarteners, the server converts sentences into hiragana and uses simple words. For elementary school students, the server converts the content into easy-to-understand expressions using kanji with furigana.
[1520] 4. Emotion Recognition with Emotion Engine:
[1521] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze the user's facial expressions, tone of voice, and emotional expressions in the input text, and recognizes the emotion at that time.
[1522] 5. Emotion-based expression adjustment:
[1523] Depending on the recognized emotion, the generative AI model can further adjust the content it outputs. For example, if the user is anxious, it can convert the content to a more gentle expression and add an encouraging message.
[1524] 6. Filtering inappropriate content:
[1525] The server filters inappropriate content during content analysis and conversion. It uses filtering algorithms to detect inappropriate keywords and phrases and removes them or replaces them with harmless alternatives.
[1526] 7. Return of Converted Content:
[1527] After all the conversion and filtering is complete, the server returns the converted content to the device, which then displays the content in a user-friendly format.
[1528] Terminal side processing
[1529] The terminal system is responsible for properly displaying the converted content received from the server to the user. The specific processing steps are as follows:
[1530] 1. Submit your request:
[1531] When a user accesses a particular web page through a browser, the terminal sends the request to a server.
[1532] 2. Receiving Content:
[1533] The terminal receives the converted content returned from the server, which is converted to suit the user's developmental level and adjusted based on the emotion engine's analysis.
[1534] 3. Viewing Content:
[1535] The received content is displayed in the browser in a format that is easy for the user to understand. For example, for kindergarteners, sentences are converted into hiragana, sentences expressed in simple words, and kanji with furigana are used. Also, if the user looks anxious, an encouraging message is added.
[1536] User usage scenarios
[1537] A user accesses a specific web page through the system. For example, if a child wants to browse a website called "Animal World," the following process will occur when using this system:
[1538] Access Start:
[1539] The user opens a browser and enters the URL for "Animal World" to access the site.
[1540] Server-based conversion:
[1541] The device sends a content request to the server, which then retrieves, analyzes, and converts the content, recognizing the user's emotions and adjusting the conversion content accordingly.
[1542] Show Content:
[1543] The device receives the converted content and displays it appropriately to the user. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." If the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[1544] Example of generative AI model prompt
[1545] User: 5 years old
[1546] Emotion: Anxiety
[1547] Text: Please convert and display the content of "Dog wagging its tail".
[1548] Prompt result: Generates a kind sentence about a dog wagging its tail and a sentence containing an encouraging message.
[1549] In this way, the system provides appropriate content according to the user's developmental level and emotional state, creating an environment in which the user can use the system with peace of mind.
[1550] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1551] Step 1:
[1552] When a user accesses a specific web page through a browser, the device sends the request to the server. Specifically, the user enters a URL and presses the access button, which generates an HTTP request and sends the request to the server. The input data is the URL, and the output data is the HTTP request.
[1553] Step 2:
[1554] The server receives a request and retrieves HTML data from a specified URL on the Internet. The server sends an HTTP GET request to download the entire specified web page. The input data is the URL and the HTTP request, and the output data is the HTML content. Specifically, the server receives the HTML data as a response.
[1555] Step 3:
[1556] The server uses an HTML parser (e.g., BeautifulSoup) to parse the HTML data it receives. The parsing extracts the main content elements in the web page, such as text, images, and links. The input data is the HTML content, and the output data are the parsed content elements. The server prepares these elements to be stored in a database.
[1557] Step 4:
[1558] The server uses a generative AI model (e.g., a TensorFlow model) to convert the analyzed content elements according to the user's developmental level. For example, for kindergarteners, the server converts sentences into hiragana and uses simple words. The input data are the analyzed content elements and the user's developmental level information, and the output data is the converted content. Specifically, the AI model processes the text data and converts it into an appropriate expression.
[1559] Step 5:
[1560] The server uses an emotion engine (e.g., OpenCV or librosa) to analyze and recognize the user's facial expressions, tone of voice, and emotional expressions in the input text. The input data is the user's facial image and voice data, and the output data is the recognized emotional state. The server prepares to adjust the content based on this data.
[1561] Step 6:
[1562] Based on the recognized emotions, the server further adjusts the content output by the generative AI model. For example, if the user is anxious, it converts the content into gentler language and adds an encouraging message. The input data is the converted content and emotion data, and the output data is the adjusted content. Specifically, the AI model adds text data and makes changes according to the emotion.
[1563] Step 7:
[1564] The server filters inappropriate content through content analysis and conversion. Filtering algorithms are used to detect inappropriate keywords and phrases and remove or replace them with harmless alternatives. The input data is the converted content, and the output data is the filtered, safe content.
[1565] Step 8:
[1566] After all conversion and filtering is complete, the server returns the converted content to the terminal. The input data is the filtered content, and the output data is the HTTP response. This allows the content to be displayed on the terminal in a format that is easy for the user to understand.
[1567] Step 9:
[1568] The device receives the converted content sent back from the server and displays it appropriately to the user. The input data is the converted content, and the output data is the content to be displayed. For example, the sentence "Elephants are very large herbivores" is displayed as "Elephants are very large grass-eating animals." Also, if the user looks anxious, an encouraging message such as "Want to know more about elephants? Good luck!" is added.
[1569] In this way, the system provides appropriate Internet content according to the user's developmental level and emotional state, creating an environment in which the user can use the content with peace of mind.
[1570] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1571] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1572] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1573] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1574] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1575] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1576] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1577] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1578] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1579] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1580] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1581] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1582] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1583] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1584] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1585] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1586] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1587] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1588] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1589] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1590] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1591] The following is further disclosed regarding the above embodiment.
[1592] (Claim 1)
[1593] means for analyzing the retrieved internet content;
[1594] a means for using a generative AI model to transform content according to a user's developmental level;
[1595] a means for displaying the converted content;
[1596] A system including:
[1597] (Claim 2)
[1598] means for adding readings to the acquired content;
[1599] A means of translating content expressions into simple language;
[1600] a means of filtering inappropriate content;
[1601] 10. The system of claim 1, comprising:
[1602] (Claim 3)
[1603] A means for analyzing the acquired internet content and extracting portions that match specific keywords;
[1604] means for converting the extracted portion according to the user's development level;
[1605] a means for integrating and displaying the converted portions;
[1606] 10. The system of claim 1, comprising:
[1607] "Example 1"
[1608] (Claim 1)
[1609] A means for obtaining accessed internet data;
[1610] A means for analyzing the acquired HTML data and extracting key elements;
[1611] a means for transforming content according to a user's developmental level using a generative AI model;
[1612] means for filtering the converted content;
[1613] means for returning the converted and filtered content to the terminal;
[1614] A system including:
[1615] (Claim 2)
[1616] A means for adding furigana to the converted content;
[1617] A means of translating content expressions into simple language;
[1618] a means of filtering inappropriate content;
[1619] 10. The system of claim 1, comprising:
[1620] (Claim 3)
[1621] A means for analyzing the acquired internet content and extracting portions that match specific keywords;
[1622] means for converting the extracted portion according to the user's development level;
[1623] a means for integrating and displaying the converted portions;
[1624] 10. The system of claim 1, comprising:
[1625] "Application Example 1"
[1626] (Claim 1)
[1627] means for analyzing the retrieved internet content;
[1628] a means for using a generative AI model to transform content according to a user's developmental level;
[1629] a means for displaying the converted content;
[1630] a means for generating unique prompt sentences to input to the generative AI model;
[1631] a means of filtering inappropriate content;
[1632] A system including:
[1633] (Claim 2)
[1634] means for converting the content into an appropriate format based on the user's age or level of understanding;
[1635] means for adding readings to the acquired content;
[1636] A means of translating content expressions into simple language;
[1637] A means of detecting inappropriate keywords in content and replacing them with appropriate expressions;
[1638] 10. The system of claim 1, comprising:
[1639] (Claim 3)
[1640] A means for analyzing the acquired internet content and extracting portions that match specific keywords;
[1641] means for converting the extracted portion according to the user's development level;
[1642] means for integrating and displaying the filtered extract;
[1643] 10. The system of claim 1, comprising:
[1644] "Example 2: Combining Emotion Engines"
[1645] (Claim 1)
[1646] means for analyzing the retrieved internet content;
[1647] a means for using a generative AI model to transform content according to a user's developmental level;
[1648] means for recognizing a user's emotion and adjusting presentation of content based on the emotion;
[1649] a means for displaying the converted content;
[1650] a means of filtering inappropriate content;
[1651] A system including:
[1652] (Claim 2)
[1653] means for adding readings to the acquired content;
[1654] A means of translating content expressions into simple language;
[1655] a means of filtering inappropriate content;
[1656] 10. The system of claim 1, comprising:
[1657] (Claim 3)
[1658] A means for analyzing the acquired internet content and extracting portions that match specific keywords;
[1659] means for converting the extracted portion according to the user's development level;
[1660] a means for integrating and displaying the converted portions;
[1661] 10. The system of claim 1, comprising:
[1662] "Application example 2 when combining emotion engines"
[1663] (Claim 1)
[1664] means for analyzing the retrieved internet content;
[1665] a means for using a generative AI model to transform content according to a user's developmental level;
[1666] a means for displaying the converted content;
[1667] means for recognizing a user's emotion;
[1668] a means for tailoring content based on the perceived sentiment;
[1669] a means of filtering inappropriate content;
[1670] A system including:
[1671] (Claim 2)
[1672] means for adding readings to the acquired content;
[1673] A means of translating content expressions into simple language;
[1674] a means of filtering inappropriate content;
[1675] A means for recognizing emotions by analyzing the user's facial expressions and voice;
[1676] a means for adjusting content in response to the perceived emotion;
[1677] 10. The system of claim 1, comprising:
[1678] (Claim 3)
[1679] A means for analyzing the acquired internet content and extracting portions that match specific keywords;
[1680] means for transforming the extracted portion according to the user's developmental level and emotions;
[1681] a means for integrating and displaying the converted portions;
[1682] 10. The system of claim 1, comprising: [Explanation of symbols]
[1683] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for analyzing the retrieved internet content; a means for using a generative AI model to transform content according to a user's developmental level; a means for displaying the converted content; A system including:
2. means for adding readings to the acquired content; A means of translating content expressions into simple language; a means of filtering inappropriate content; The system of claim 1 , comprising:
3. A means for analyzing the acquired internet content and extracting portions that match specific keywords; means for converting the extracted portion according to the user's development level; a means for integrating and displaying the converted portions; The system of claim 1 , comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A