Artificial intelligence system, program for artificial intelligence system

The AI system addresses the issue of insufficient response accuracy by allowing users to train and interact with pre-trained AI tools on user-specified data, ensuring accurate and personalized responses.

JP2026083538APending Publication Date: 2026-05-20SATELLITE OFFICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SATELLITE OFFICE CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Conventional information processing systems using large language models generate text-based summaries without learning about user-specified fields, resulting in insufficient response accuracy.

Method used

An artificial intelligence system that allows users to interact with pre-trained AI means by specifying electronic information resources, training the AI on user-specified data, converting it into vector coordinates, and enabling interaction through a user interface for seamless communication.

Benefits of technology

Enables accurate and personalized responses by allowing users to interact with pre-trained AI tools that have learned the specified content, enhancing response accuracy and user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026083538000001_ABST
    Figure 2026083538000001_ABST
Patent Text Reader

Abstract

The system provides an artificial intelligence (AI) system that allows users to start interacting with a pre-trained AI that has already learned from a specified electronic information resource. [Solution] The artificial intelligence system 100 comprises a user terminal 110 and servers 120 and 130. The browser 111 of the user terminal displays a settings screen for the server's pre-trained artificial intelligence means 131. When name information is entered on the settings screen and an electronic information resource is specified, the user terminal uploads the information of the specified electronic information resource along with the name information to the server. The server extracts language character information from the electronic information resource and trains the pre-trained artificial intelligence means. The browser of the user terminal displays the name information NM in a selectable manner on the skill function selection screen 113. When the name information is selected, a talk screen 114 is displayed that allows interaction with the pre-trained artificial intelligence means of the name information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an artificial intelligence system including a user terminal and a server, where the user terminal displays or outputs audio based on the server, and a program thereof.

Background Art

[0002] Conventionally, there is known an information processing apparatus including a language data receiving unit that receives language data in response to a request from a user terminal, a language data processing unit that processes the language data to generate summary basic information serving as a basis for summary generation, a transmission unit for transmitting presentation information including the summary basic information to the user terminal, an editing information receiving unit that receives editing information regarding the presentation information from the user terminal, and a summary generation unit that provides the editing information to a large language model that receives an input of the summary basic information and generates a summary text in which the editing information is reflected (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, the above-described conventional information processing apparatus has a problem that, as long as it directly generates a text-based summary using a large language model from some language text-based summary basic information, it has not learned about any arbitrary field desired by the user and the response accuracy is not sufficient.

[0005] Therefore, the present invention solves the problems of the prior art described above, and that is, the object of the present invention is to provide an artificial intelligence system and program that allows a user to start interacting with a pre-trained artificial intelligence means that has learned a specific electronic information resource specified by the user, from a state in which it has already learned. [Means for solving the problem]

[0006] The invention according to claim 1 is an artificial intelligence system comprising a user terminal and a server, wherein the user terminal displays or outputs audio based on the server, and the user terminal communicates with the server, the browser of the user terminal displays a settings screen for a text generation means which is a pre-trained artificial intelligence means of the server, and when name information is entered on the settings screen and at least one of electronic file information, website URL information, and database information as an electronic information resource is specified, the user terminal uploads the information of the electronic information resource specified along with the name information to the server, the server extracts language character information from the electronic information resource based on the received information of the electronic information resource and trains the language character information in the pre-trained artificial intelligence means, the browser of the user terminal displays the name information of the trained pre-trained artificial intelligence means in a selectable manner on a skill function selection screen, and when the name information of the trained pre-trained artificial intelligence means is selected, a talk screen is displayed which allows interaction with the pre-trained artificial intelligence means of the name information, thereby solving the aforementioned problems.

[0007] The invention according to claim 2 further solves the aforementioned problems by having, in addition to the configuration of the artificial intelligence system described in claim 1, that each time name information is entered and an electronic information resource is specified in the settings screen, the server extracts language character information from the electronic information resource and trains a pre-trained artificial intelligence means, and the browser of the user terminal displays the name information of multiple pre-trained artificial intelligence means that have been trained, in a selectable manner on the skill function selection screen.

[0008] The invention according to claim 3, in addition to the configuration of the artificial intelligence system described in claim 2, includes the following: the server divides the extracted language character information data into a plurality of divided language character file data, converts each of the divided divided language character file data into vector coordinate system divided language data using a vector coordinate transformation means which is a pre-trained artificial intelligence means, stores the converted vector coordinate system divided language data in vector indexes in association with the original divided language character file data, and when a user inputs text or voice in the browser of the user terminal, the user terminal sends the text data or voice data to the server as question information, and the server uses the vector coordinate transformation means to convert the received question information into a vector coordinate system question The system further solves the aforementioned problems by converting the data into vector data, searching for the vector coordinate system segmented language data that is closest to the vector of the vector coordinate system question data in the vector index, selecting the closest vector coordinate system segmented language data, or the segmented language character file data corresponding to the top multiple closest vector coordinate system segmented language data, the server attaching the selected segmented language character file data to the question information, executing text data generation using a pre-trained artificial intelligence means, converting the generated text data obtained into audio data according to the settings and sending it to the user terminal as answer information, and the user terminal displaying the received answer information in a browser or outputting it as audio.

[0009] The invention according to claim 4 further solves the aforementioned problems by having the configuration of the artificial intelligence system described in any one of claims 1 to 3, wherein, in the settings screen, character face image data, character voice data, and speaking style setting information are input or selected, and in the skill function selection screen, name information of a pre-trained artificial intelligence means set for the character is selected, the browser of the user terminal displays the character's face image data on the talk screen based on the setting information for the character, and when the user performs a voice input operation in the browser of the user terminal, the user terminal sends the voice data as question information to the server, the server generates text data as an answer using the pre-trained artificial intelligence means in response to the received question information, converts the generated text data into voice data and sends the voice data to the user terminal, and the browser of the user terminal outputs the received voice data as voice based on the setting information for the character.

[0010] The invention according to claim 5 further solves the aforementioned problems by having a configuration that, in addition to the configuration of the artificial intelligence system described in claim 4, if the character's face image data is the face of a real person and the face of a person registered in the address book, when an email address from the address book is specified in the settings screen, when the user makes a voice input operation as a question in the talk screen, and when the voice data received from the server is output as a voice based on the setting information about the character as an answer, the browser displays a button on the talk screen indicating that the history should be emailed, and when the button to email the history is operated, the server sends text data about the question information and answer information to the email address corresponding to the character.

[0011] The invention according to claim 6 is a program for an artificial intelligence system in which a user terminal displays or outputs audio based on a server, wherein the user terminal communicates with the server, the user terminal's browser displays a settings screen for a text generation means which is a pre-trained artificial intelligence means of the server, and determines whether name information has been entered on the settings screen and whether at least one of electronic file information, website URL information, or database information as an electronic information resource has been specified; an upload step in which the user terminal uploads information of the electronic information resource specified together with the name information to the server; a learning step in which the server extracts language character information from the electronic information resource based on the received information of the electronic information resource and trains the language character information in the pre-trained artificial intelligence means; and a talk screen display step in which the user terminal's browser displays the name information of the trained pre-trained artificial intelligence means in a selectable manner on a skill function selection screen, and when the name information of the trained pre-trained artificial intelligence means is selected, a talk screen display is displayed that allows interaction with the pre-trained artificial intelligence means of the name information, thereby solving the aforementioned problems. [Effects of the Invention]

[0012] The artificial intelligence system of the present invention, comprising a user terminal and a server, not only allows the user terminal to display or output audio based on the server, but also achieves the following unique effects.

[0013] According to the artificial intelligence system of the invention described in claim 1, an electronic information resource is specified by the user along with name information, and a pre-trained artificial intelligence means that has learned that electronic information resource is created with that name information and displayed in the browser at will. As a result, the user can start interacting with the pre-trained artificial intelligence means that has learned the specified electronic information resource from a state in which it has already learned. In other words, the electronic information resources specified by the user become RAG (Retrieval Augmented Generation) data, which are external information sources for the artificial intelligence system. Therefore, the user can create a pre-trained artificial intelligence tool specifically for this purpose, interact with that pre-trained AI tool, and obtain answers based on the specified electronic information resources. Furthermore, since the content of the electronic information resource specified by the user is learned by a pre-trained artificial intelligence system, the user can obtain the desired answer by interacting with the pre-trained artificial intelligence system in a way that is specific to the content of the specified electronic information resource.

[0014] According to the artificial intelligence system of the invention of claim 2, in addition to the effects of the invention of claim 1, the name information of each pre-trained artificial intelligence means that has learned specific electronic information resources is displayed on the settings screen, so that the user can select the name information of a desired pre-trained artificial intelligence means that specializes in a particular information field or content and interact with that pre-trained artificial intelligence means.

[0015] According to the artificial intelligence system of the invention described in claim 3, in addition to the effects of the invention described in claim 2, the RAG-learned content is converted into vectors and stored in a vector index, the question information is converted into vectors and a so-called correlation search is performed, content highly related to the question is selected and an answer is generated based on this, so the user can obtain highly accurate answers to questions.

[0016] According to the artificial intelligence system of the invention of claim 4, in addition to the effects achieved by any one of the inventions of claims 1 to 3, the user can interact with the character as if they were actually talking to the character, because the character's face image data is displayed on the talk screen and when a question is asked by voice input, an answer is obtained by voice output. Furthermore, since the voice data is output as a response based on the setting information for the character's speaking style, it is possible to give the character's speaking style individuality and characteristics.

[0017] According to the artificial intelligence system of the invention of claim 5, in addition to the effects achieved by the invention of claim 4, if the person whose character is real exists, the content of the interaction between the user and the character on the chat screen is sent via email to the email address of the person whose character is real. As a result, the person whose character is real can be informed of the interaction between their character and the user via email. In other words, since users receive email notifications about the content of their avatar's interactions with other users, the character themselves can be aware of the interactions between their avatar and the users.

[0018] According to the program for the artificial intelligence system of the invention of claim 6, similar to the effect achieved by the invention of claim 1, a user can specify an electronic information resource along with name information, and a pre-trained artificial intelligence means that has learned that electronic information resource is created with that name information and displayed in the browser at will. Therefore, the user can start interacting with the pre-trained artificial intelligence means that has learned the specified electronic information resource from a state in which it has already learned. In other words, the electronic information resources specified by the user become RAG (Retrieval Augmented Generation) data, which are external information sources for the artificial intelligence system. Therefore, the user can create a pre-trained artificial intelligence tool specifically for this purpose, interact with that pre-trained AI tool, and obtain answers based on the specified electronic information resources. Furthermore, since the content of the electronic information resource specified by the user is learned by a pre-trained artificial intelligence system, the user can obtain the desired answer by interacting with the pre-trained artificial intelligence system in a way that is specific to the content of the specified electronic information resource. [Brief explanation of the drawing]

[0019] [Figure 1] Figure showing the concept of an artificial intelligence system according to an embodiment of the present invention. [Figure 2] Chart showing an example of the operation of an artificial intelligence system according to an embodiment of the present invention. [Figure 3] Figure showing an example of a setting screen in a browser of a user terminal of an artificial intelligence system according to an embodiment of the present invention. [Figure 4] Explanatory diagram showing how the language text data of an artificial intelligence system according to an embodiment of the present invention is divided and the divided language character file data is converted into language data after being divided in a vector coordinate system. [Figure 5] Explanatory diagram showing how the language data after being divided in a vector coordinate system of an artificial intelligence system according to an embodiment of the present invention is stored in a vector index. [Figure 6] Figure showing an example of a skill function selection screen in a browser of a user terminal of an artificial intelligence system according to an embodiment of the present invention. [Figure 7] Figure showing an example of a talk screen with a selected skill function in a browser of a user terminal of an artificial intelligence system according to an embodiment of the present invention. [Figure 8] Figure showing a state where an answer is displayed on a talk screen with a selected skill function in a browser of a user terminal of an artificial intelligence system according to an embodiment of the present invention. [Figure 9] Explanatory diagram showing how question information, which is a prompt of an artificial intelligence system according to an embodiment of the present invention, is converted into vector coordinate question data. [Figure 10] (A)(B) are explanatory diagrams regarding the correlation search between the vector coordinate question data and the language data after being divided in a vector coordinate system of an artificial intelligence system according to an embodiment of the present invention. [Figure 11] Explanatory diagram showing how the divided language character file data selected for the question information, which is a prompt of an artificial intelligence system according to an embodiment of the present invention, is added and transmitted to the pre-trained artificial intelligence means. [Figure 12] Figure showing an example of an administrator setting screen in a browser of an administrator terminal of an artificial intelligence system according to an embodiment of the present invention. [Figure 13] This figure shows an example of a chat screen in the browser of a smartphone user terminal of an artificial intelligence system, which is an embodiment of the present invention. [Figure 14] This figure shows how a notification email is displayed in the browser of a user terminal of an artificial intelligence system, which is an embodiment of the present invention. [Modes for carrying out the invention]

[0020] The present invention relates to an artificial intelligence system comprising a user terminal and a server, wherein the user terminal displays or outputs audio based on the server, and the user terminal communicates with the server, and the user terminal's browser displays a settings screen for the server's pre-trained artificial intelligence means, which is a text generation means, and on the settings screen, name information is entered and at least one of electronic file information, website URL information, and database information as an electronic information resource is specified, the user terminal uploads the information of the specified electronic information resource along with the name information to the server, and the server extracts language character information from the electronic information resource based on the received information of the electronic information resource and processes the language character information. The system is configured such that a pre-trained artificial intelligence means is trained, and the user's terminal browser displays the name information of the trained pre-trained artificial intelligence means freely on the skill function selection screen. When the name information of the trained pre-trained artificial intelligence means is selected, a talk screen is displayed that allows interaction with the pre-trained artificial intelligence means whose name information is freely accessible. This configuration allows the user to start interacting with a pre-trained artificial intelligence means that has already learned a specific electronic information resource, and the user can obtain a desired response by interacting with the pre-trained artificial intelligence means in a way that is specific to the content of the specified electronic information resource. The specific embodiment of this system is not limited to any particular form. Furthermore, the program of the artificial intelligence system of the present invention includes: an input setting presence determination step in which the user terminal communicates with the server, the user terminal's browser displays a settings screen for the server's pre-trained artificial intelligence means, the settings screen determines whether name information has been entered and whether at least one of electronic file information, website URL information, or database information as an electronic information resource has been specified; an upload step in which the user terminal uploads the information of the specified electronic information resource along with the name information to the server; and a learning step in which the server extracts language character information from the electronic information resource based on the received information of the electronic information resource and causes the pre-trained artificial intelligence means to learn the language character information. The computer executes a step in which the user's terminal browser displays the name information of the pre-trained artificial intelligence means that has been trained, allowing the user to select it freely on the skill function selection screen, and when the name information of the pre-trained artificial intelligence means is selected, it displays a talk screen that allows the user to interact with the pre-trained artificial intelligence means that has learned the specified electronic information resource, enabling the user to start interacting with the pre-trained artificial intelligence means from a state of having already learned it. The specific embodiment of this can be anything, as long as the user can interact with the pre-trained artificial intelligence means in a way that is specific to the content of the specified electronic information resource and obtain the desired answer.

[0021] For example, a user terminal can be any device that sends and receives information, such as a desktop personal computer terminal, a notebook personal computer terminal, a smartphone terminal, or a tablet terminal, as long as it can connect to a server via a communication network including a wide area network (such as the Internet), a local network, or a telephone line. Furthermore, each server can be a single server or multiple servers on the cloud. The pre-trained artificial intelligence means (text generation means) is a text generation means that generates text data and is composed of an interactive model also called a large-scale language model, such as ChatGPT (Generative Pre-trained Transformer) (hereinafter referred to as ChatGPT), and may be configured on one server or on multiple servers on the cloud. Furthermore, the electronic information source can be any electronic data that can be specified, such as electronic file information like document files, website URL information, or database information. Here, a document file refers to an electronic data file that contains text information, and the data file format can be anything from a simple text data format, document data format, cell-display type spreadsheet file format, to pixel-display type image file format such as photos and videos, or any format that allows text to be read using image recognition means or speech recognition means, as long as the text is recognized directly or indirectly by transcribing text from speech. Note that the text also includes program source code. Furthermore, the URL information (Uniform Resource Locator information), which is location information for the specified information that designates the electronic information source, can be any information that can identify an HTML website page or document file obtained by accessing a web server. A chat screen is a screen that displays the exchange of text and file data with one or more parties (including bots and large-scale language models of pre-trained artificial intelligence tools (text generation tools)) in chronological order. [Examples]

[0022] Below, an artificial intelligence system 100, which is an embodiment of the present invention, will be described with reference to Figures 1 to 14. Here, Figure 1 is a diagram showing the concept of an artificial intelligence system 100 which is an embodiment of the present invention; Figure 2 is a chart showing an example of operation of the artificial intelligence system 100 which is an embodiment of the present invention; Figure 3 is a diagram showing an example of the settings screen 112 in the browser 111 of the user terminal 110 of the artificial intelligence system 100 which is an embodiment of the present invention; Figure 4 is an explanatory diagram showing how the language text data of the artificial intelligence system 100 which is an embodiment of the present invention is divided and the divided language character file data DTX is converted into vector coordinate system divided language data VTX; Figure 5 is an explanatory diagram showing how the vector coordinate system divided language data VTX of the artificial intelligence system 100 which is an embodiment of the present invention is stored in the vector index 121; Figure 6 is a diagram showing an example of the skill function selection screen 113 in the browser 111 of the user terminal 110 of the artificial intelligence system 100 which is an embodiment of the present invention; Figure 7 is a diagram showing an example of the talk screen 114 with the selected skill function in the browser 111 of the user terminal 110 of the artificial intelligence system 100 which is an embodiment of the present invention; Figure 8 is a diagram showing the browser of the user terminal 110 of the artificial intelligence system 100 which is an embodiment of the present invention. Figure 111 shows how the answer is displayed on the talk screen 114 with the selected skill function in the browser 111. Figure 9 is an explanatory diagram showing how the prompt question information PT of the artificial intelligence system 100, an embodiment of the present invention, is converted into vector coordinate system question data VPT. Figure 10(A) shows the vector coordinate system question data VPT of the artificial intelligence system 100, an embodiment of the present invention. Figure 10(B) shows the correlation search between the vector coordinate system question data VPT and the vector coordinate system segmented language data VTX of the artificial intelligence system 100, an embodiment of the present invention. Figure 11 is an explanatory diagram showing how the selected segmented language character file data DTX is added to the question information PT, which is a prompt of the artificial intelligence system 100, an embodiment of the present invention, and transmitted to the pre-trained artificial intelligence means. Figure 12 is a diagram showing an example of the administrator settings screen 142 in the browser 141 of the administrator terminal 140 of the artificial intelligence system 100, an embodiment of the present invention. Figure 13 is a diagram showing an example of the talk screen 114 in the browser 111 of the smartphone user terminal 110 of the artificial intelligence system 100, an embodiment of the present invention.Figure 14 shows how a notification email is displayed in the browser 111 of the user terminal 110 of the artificial intelligence system 100, which is an embodiment of the present invention.

[0023] As shown in Figure 1, an embodiment of the present invention, the artificial intelligence system 100, comprises a user terminal 110, a first server 120 and a second server 130 as servers, and, as an example, an administrator terminal 140. The user terminal 110 is configured to display or output audio based on the first server 120 and the second server 130. Of these, the first server 120 has a database, a vector index 121 which also serves as a spatial map, and a bot 122 which is also a display program. The database contains information about the web page of the talk screen 114 of the artificial intelligence system 100 that is displayed by the browser 111 of the user terminal 110. Here, vector index 121 refers to a list, database, or conceptual spatial map in a memory unit used to quickly find a specific object from a large amount of vector coordinate system data.

[0024] Furthermore, the first server 120 has, as an example, a vector coordinate transformation means as a pre-trained artificial intelligence means. Furthermore, the vector coordinate transformation means, also known as the embedding model, is designed to convert input data into data in a vector coordinate system. For example, the vector components are determined based on each element and its intensity in the data being transformed. Furthermore, the second server 130 has ChatGPT, which is an example of a text generation means 131 also known as a Large-Scale Language Model (LLM) as a pre-trained artificial intelligence means. Furthermore, from a technical standpoint, the first server 120 and the second server 130 may have a configuration that is physically common to each other, or they may have a configuration that is physically different.

[0025] Furthermore, the administrator terminal 140 sets permissions for settings such as specifying external information sources referenced by ChatGPT, an example of the text generation means 131, and creating skill functions, in the interaction between the user terminal 110 and ChatGPT, an example of the text generation means 131, including whether to allow only the administrator or also users to do so. The "skill function" refers to a function that uses RAG (Retrieval Augmented Generation) training, which involves specifying external information sources to train a large-scale language model (LLM) as a pre-trained artificial intelligence tool. This creates a large-scale language model (LLM) specialized in the learned content, and the desired answer can be obtained by interacting with this LLM. As an example, let's consider a scenario where the creation of skill functions is permitted not only for administrators but also for users.

[0026] In this embodiment, the user terminal 110 communicates with the first server 120, and the browser 111 of the user terminal 110 displays a settings screen 112 for the text generation means 131, which is a pre-trained artificial intelligence means of the second server 130. Then, on the settings screen 112, the user inputs the name information NM, and at least one of the following is specified as an electronic information resource: electronic file information, website URL information, or database information. Then, the user terminal 110 uploads the information of the specified electronic information resource, along with the name information, to the first server 120.

[0027] Next, the first server 120 extracts language character file data TX as language character information from the electronic information resource based on the information received from the electronic information resource. Then, the first server 120 trains the text generation means 131, which is a pre-trained artificial intelligence means of the second server 130, with language character file data TX as language character information. This is known as RAG learning. Furthermore, the browser 111 of the user terminal 110 displays, as a skill function, the name information NM of the pre-trained artificial intelligence means, the text generation means 131, along with its icon data CN, in a selectable manner on the skill function selection screen 113. Furthermore, when the name information NM or icon data CN of the pre-trained artificial intelligence means, which is a text generation means 131, is selected as a skill function, the system is configured to display a talk screen 114 that allows interaction with the selected name information NM, which is a text generation means 131.

[0028] As a result, the user specifies an electronic information resource along with name information NM, and a pre-trained artificial intelligence means, the text generation means 131, which has learned that electronic information resource, creates a skill function using that name information NM and displays it in the browser 111 at will. As a result, the user can start interacting with the text generation means 131, which is a pre-trained artificial intelligence means that has learned a specific electronic information resource specified by the user, from a state where it has already learned. In other words, the electronic information resources specified by the user become the external information sources for the artificial intelligence system 100, known as RAG (Retrieval Augmented Generation) data. As a result, users can create a pre-trained artificial intelligence means, namely text generation means 131, as a skill function, and interact with this pre-trained artificial intelligence means, namely text generation means 131, to obtain a response based on the specified electronic information resource. Furthermore, the content of the electronic information resource specified by the user is learned by the text generation means 131, which is a pre-trained artificial intelligence means. As a result, the user can obtain the desired response by interacting with the pre-trained artificial intelligence means, the text generation means 131, in a manner specific to the content of the specified electronic information resource.

[0029] Next, we will explain in detail an example of how the artificial intelligence system 100 operates. As shown in Figure 2, in step S1, the user terminal 110 communicates with the first server 120 as a setting screen display step. Then, as shown in Figure 3, the browser 111 of the user terminal 110 displays a settings screen 112 for skill functions using pre-trained artificial intelligence means on the second server 130 via the first server 120.

[0030] For example, a user logs in to the artificial intelligence system 100 by entering their user ID and password on the login screen of the artificial intelligence system 100 in the browser 111 of the user terminal 110. Then, based on the operation to display the skill function settings screen 112, the browser 111 on the user terminal 110 displays the skill function settings screen 112. The settings screen 112 for the skill function includes, as an example, items such as icon, name information for NM, detailed information, language settings, model settings, file upload as electronic information resource, URL as electronic information resource, instructions, rich menu, web browsing, character face image, character voice, character characteristics, history recipient email address, and a "Register Skill" button.

[0031] In the icon section, you can specify the icon data CN. Alternatively, you can choose to have the icon data CN automatically generated by not specifying it. Additionally, the skill name field will be populated with the user's name information (NM). In the "Details" section, users enter a description of the skill's functionality. In the language settings section, you can freely select from various languages. In the model settings section, a large-scale language model, which is a pre-trained artificial intelligence tool, can be freely selected.

[0032] In the file upload section, users can freely select and specify document files and other files by clicking the "Select File" button. For example, you can specify files using drag and drop, or a folder will be displayed, allowing you to freely select and specify document files within that folder. The URL field allows users to freely enter URL information. In the Instructions section, the format of questions and answers for users using the skill function can be freely specified, similar to prompts.

[0033] In the rich menu section, the rich menu button 114c, which is displayed when the chat screen 114 is first displayed, can be freely configured. For example, when the chat screen 114 is initially displayed, pre-written text is entered to make it easier for the user to select questions, etc. When that button is pressed, the text of that button is entered into the input field 114a of the chat screen 114 and sent. In the Web Browsing section, checking "Enable Web Browsing" allows you to freely configure the system to learn the latest information from the internet. Furthermore, in the character face image section, by operating the "Select File" button, you can select a face image data FD, by operating the camera capture button, you can take a picture with the camera and create a face image data FD, or by checking "Use Icon Image" and reusing icon data CN, allowing you to freely specify the character's face image as described later.

[0034] Furthermore, in the character voice section, by operating the "Select File" button, you can select an audio data VD, or by operating the microphone sample input button, you can record with the microphone and create an audio data VD, allowing you to freely specify the character's voice as described later. Furthermore, in the character characteristics section, you can freely specify details about the character, such as gender, age, origin, speaking speed, and the strength of their dialect. Furthermore, in the history recipient section, when a character is set and an interaction takes place with the user on the chat screen 114, the recipient email address for the interaction history can be freely specified.

[0035] In step S2, as an input setting determination step, the browser 111 of the user terminal 110 determines whether the name information NM has been entered on the setting screen 112 and whether at least one of the following has been specified: electronic file information as an electronic information resource, website URL information, or database information. For example, the determination could be made when the "Register Skill" button on the chat screen 114 is clicked. For example, suppose that in settings screen 112, icon data CN is set, "ABC Automobile Public Relations AI" is entered as the skill name, the URL information for the public relations webpage on ABC Automobile's website is entered in the URL field, and the "Register Skill" button is clicked. If it is determined that input or a specified value has been entered, the process proceeds to step S3; otherwise, step S2 is repeated.

[0036] In step S3, as an upload step, the user terminal 110 uploads the information of the specified electronic information resource, along with the name information, to the first server 120. In step S4, as a learning step, the first server 120 extracts language character file data TX as language character information from the electronic information resource based on the information of the received electronic information resource. Then, the first server 120 uses the language character file data TX, which is language character information, to train the pre-trained artificial intelligence means of the second server 130. This creates skill functions using large-scale language models of pre-trained artificial intelligence tools.

[0037] Furthermore, the overarching concept of the technical idea is simply to extract language character information. When this information is extracted, it may be stored in a file format such as a language character file data TX, or it may not be stored in a file format. Furthermore, while the first server 120, as an example of a server, extracted linguistic character information from an electronic information source, the second server 130 may also perform the extraction.

[0038] Furthermore, when extracting language character file data TX as language character information from electronic information resources and training a large-scale language model of a pre-trained artificial intelligence means (RAG), the system may be configured to perform so-called embedding processing, which converts language text data into vector data using a vector coordinate transformation means, which is also a pre-trained artificial intelligence means. For example, as shown in Figure 4, first, as a language character data splitting step, the first server 120, as an example of a server, performs a so-called chunking process in which it splits the extracted language character file data TX into multiple split language character file data DTX. More specifically, as shown in Figure 4, as an example, language character file data TX, which is language character information extracted from a website at a specified URL, is divided into multiple split language character file data DTX, for example, every 500 characters. Furthermore, considering the variation in the length of each sentence, it is possible to divide the text into sections of approximately 400 to 600 characters each, or to divide each sentence into multiple sections. Furthermore, while the first server 120, as an example of a server, performed the chunking process, the second server 130 may also perform the chunking process.

[0039] Then, a pre-trained artificial intelligence means, which is a vector coordinate transformation means, converts the divided language character file data DTX into divided language data VTX in a vector coordinate system, each having multiple parameters. Here, "language data VTX after vector coordinate system partitioning" is, for example, vector data consisting of 1536 dimensions (variables / parameters). In other words, it is vector data identified by 1536 components. The number of components in a vector can be any number. Furthermore, while the first server 120, as an example of a server, performed the embedding process, the second server 130 may also perform the embedding process.

[0040] As a data conversion and storage step, the first server 120, as an example of a server, performs a so-called mapping process, which involves associating the converted vector coordinate system segmented language data VTX with the pre-conversion segmented language character file data DTX and saving them to vector index 121. More specifically, the language data VTX after vector coordinate system division shown in Figure 4 is stored in the vector index 121 of the first server 120, as an example, as shown in Figure 5. In this process, the language data VTX after vector coordinate system division is saved in association with the language character file data DTX before the conversion. Note that while the location of vector index 121 is given as server 120 as an example, it does not have to be server 120. Furthermore, while the first server 120, as an example of a server, performed the mapping process, the second server 130 may also perform the mapping process.

[0041] As a result, when an administrator specifies an electronic information source, the language character file data TX, which is the language character information of the electronic information source, is extracted, a chunking process is performed (which is a splitting process), and then an embedding process, which is a vector transformation process, is performed. The relationship before and after the vector transformation is then linked, and a mapping process, which is a vector index saving process, is performed. As a result, it becomes easy to create so-called RAG (Retrieval Augmented Generation) data, which serves as an external information source for the artificial intelligence system 100.

[0042] In step S5, as shown in Figure 6, as a skill function selection screen display step (talk screen display step), the browser 111 of the user terminal 110 displays pre-trained artificial intelligence means along with name information NM on the skill function selection screen 113, allowing users to select them freely. For example, when the "Skill List" button in the left-hand column is pressed on a screen such as the talk screen 114 of the artificial intelligence system 100, the browser 111 on the user terminal 110 displays the skill function selection screen 113. On the skill function selection screen 113, each created skill function is displayed as a set of name information (NM) and icon data (CN). Then, on the skill function selection screen 113, the skill functions of "ABC Automotive Public Relations AI" that were registered earlier on the settings screen 112 are displayed as a set of name information NM and icon data CN. Furthermore, for the skill function to be displayed, it is sufficient if at least the name information (NM) is shown.

[0043] In step S6, as a skill function selection determination step (talk screen display step), the browser 111 of the user terminal 110 determines whether or not one of the skill functions on the skill function selection screen 113 has been selected. If it is determined that an item has been selected, proceed to step S7; otherwise, repeat step S6. For example, suppose a user selects the "ABC Automotive PR AI" skill function.

[0044] In step S7, as shown in Figure 7, as a step to display the talk screen, the browser 111 of the user terminal 110 displays a talk screen 114 that allows interaction with the skill function of the pre-trained artificial intelligence means for the selected name information NM. For example, the browser 111 on the user terminal 110 displays a chat screen 114 that allows for seamless interaction with the skill functions of "ABC Automotive Public Relations AI". The chat screen 114, for PC use, includes, as an example, an input field 114a, a microphone button 114b, and a rich menu button 114c. The user may use the keyboard to input the instruction information, which is the question information PT, as text into the input field 114a, or they may operate the microphone button 114b to activate the microphone, input the question information PT as voice via the microphone, have the browser 111 convert it to text, and input it as text into the input field 114a. Alternatively, the user may operate the pre-configured rich menu button 114c and input the question information PT into the input field 114a as text.

[0045] As a result, as described above, the user specifies an electronic information resource along with name information NM, and a pre-trained artificial intelligence means, the text generation means 131, which has learned that electronic information resource, creates a skill function using that name information NM and displays it in the browser 111 at will. As a result, the user can start interacting with the text generation means 131, which is a pre-trained artificial intelligence means that has learned a specific electronic information resource specified by the user, from a state where it has already learned. In other words, the electronic information resources specified by the user become the external information sources for the artificial intelligence system 100, known as RAG (Retrieval Augmented Generation) data. As a result, users can create a pre-trained artificial intelligence means, namely text generation means 131, as a skill function, and interact with this pre-trained artificial intelligence means, namely text generation means 131, to obtain a response based on the specified electronic information resource. Furthermore, the content of the electronic information resource specified by the user is learned by the text generation means 131, which is a pre-trained artificial intelligence means. As a result, the user can obtain the desired response by interacting with the pre-trained artificial intelligence means, the text generation means 131, in a manner specific to the content of the specified electronic information resource.

[0046] For example, a user might enter a question PT such as "What are the advantages of electric vehicles compared to gasoline cars?" into input field 114a and then submit the form. Then, the user terminal 110 sends the question information PT to the first server 120. The first server 120 receives the question information PT and uses the skill function of the "ABC Automotive Public Relations AI," which utilizes the pre-trained artificial intelligence means of the second server 130, text generation means 131 (large-scale language model), to generate text data as an answer. Then, the first server 120 sends the generated text data GT obtained as a response to the user terminal 110. Then, as shown in Figure 8, the browser 111 of the user terminal 110 displays the generated text data GT on the chat screen 114.

[0047] Furthermore, in this embodiment, each time name information NM is entered in the settings screen 112 and an electronic information resource is specified, the first server 120 extracts language character file data TX as language character information from that electronic information resource and trains the pre-trained artificial intelligence means, text generation means 131 (large-scale language model). As shown in Figure 6, the browser 111 of the user terminal 110 is configured to display, in a skill function selection screen 113, the icon data CN of multiple pre-trained artificial intelligence means, name information NM, along with the name information CN of the text generation means 131 (large-scale language model), as selectable skill functions.

[0048] As a result, the name information NM of each pre-trained text generation means 131 (large-scale language model), which is a pre-trained artificial intelligence means that has learned specific electronic information resources, is displayed on the skill function selection screen 113. As a result, users can select the name information NM of a pre-trained artificial intelligence means, name information NM, that is a text generation means (large-scale language model) specializing in a particular information field or content, as a skill function, and interact with that pre-trained artificial intelligence means, name information NM, name information NM, name information NM.

[0049] As mentioned above, if the system is configured to perform so-called embedding processing, which converts the language character file data TX, extracted from electronic information resources as language character information, into vector data, then the question information PT will also undergo embedding processing. In other words, when a user inputs text or voice in the browser 111 of the user terminal 110, the user terminal 110 sends the text data or voice data VD to the first server 120 as question information PT, which is an example of instruction information. More specifically, as an example, the question information PT, "What are the advantages of electric vehicles compared to gasoline cars?", is sent to the first server 120.

[0050] Then, as a question information vector transformation step (correlation search step), the first server 120, as an example of a server, transforms the received question information PT into vector coordinate system question data VPT using a vector coordinate transformation means. More specifically, as shown in Figure 9, as an example, the question information PT for "What are the advantages of electric vehicles compared to gasoline cars?" is converted into vector coordinate system question data VPT in the format "2,6,1,1,7,7,6,3,1...". Figure 10(A) shows the conversion of question information PT (text format data) to vector coordinate system question data VPT (Variable Point Data). This representation is two-dimensional as an example, but actual vector coordinates are, for example, 1536-dimensional. In this example, the first server 120 performed the embedding process for the question information PT, but the second server 130 may also perform the embedding process.

[0051] As shown in Figure 10(B), as a correlation search step, the first server 120, as an example of a server, searches for the vector coordinate system segmented language data VTX that has the closest relationship to the vector of the vector coordinate system question data VPT in the vector index 121. Then, the system selects the language data VTX after vector coordinate system segmentation that has the closest relationship, or the language character file data DTX that corresponds to the top-ranking close relationship of the language data VTX after vector coordinate system segmentation. For example, while Server 120 (the first server) performed the correlation search, Server 2 (the second server) 130 could also perform the correlation search.

[0052] As shown in Figure 11, in the reference information attachment transmission step (answer text generation and transmission step), the first server 120, as an example of a server, attaches the selected segmented language character file data DTX to the question information PT and sends it to the skill function of "ABC Automotive Public Relations AI" which uses chat GPT, an example of a pre-trained artificial intelligence means, text generation means 131. Furthermore, since the skill function of "ABC Automotive Public Relations AI" itself is a concept that utilizes the pre-trained artificial intelligence means, the text generation means 131, it is not necessary to consider whether it is installed on the first server 120 or the second server 130. For the user, the skill function of "ABC Automotive Public Relations AI" is operated as a pre-trained artificial intelligence means, namely text generation means 131.

[0053] More specifically, an example of question information PT, "What are the advantages of electric vehicles compared to gasoline cars?", is combined with reference information, an example of selected segmented language character file data DTX, "Electric vehicles are electric...", and sent to the skill function of "ABC Automotive Public Relations AI," which uses chat GPT, an example of a pre-trained artificial intelligence means, text generation means 131. In this case, the bot 122 of the first server 120 adds the selected segmented language character file data DTX to the question information PT and sends it to the skill function of "ABC Automotive Public Relations AI" which uses chat GPT, an example of a pre-trained artificial intelligence means, as a prompt, which is a standard instruction / command information that instructs the AI ​​to refer to the segmented language character file data DTX and answer the question information PT. As an example of a server, the first server 120 sent the data to the skill function of "ABC Automotive Public Relations AI" using ChatGPT, but the second server 130 could also send the data to the skill function of "ABC Automotive Public Relations AI" using ChatGPT.

[0054] As part of the response text generation and transmission step, the first server 120 uses the skill function of "ABC Automotive Public Relations AI," which utilizes Chat GPT, an example of a pre-trained artificial intelligence means, to generate text data. In other words, the skill function of "ABC Automotive Public Relations AI," which uses Chat GPT, an example of a pre-trained artificial intelligence means, text generation means 131, generates generated text data GT as an answer based on the received prompt (an instruction / command to answer the question information PT by referring to the segmented language character file data DTX). Then, as an example of generated text data GT, something like, "Everything is electronically controlled, making it well-suited for autonomous driving. Enter your destination by voice and depart..." is generated.

[0055] Then, the first server 120 sends the generated text data GT obtained from the execution to the user terminal 110. More specifically, an example of generated text data GT is sent to the user terminal 110 with the message: "Everything is electronically controlled, making it well-suited for autonomous driving. Enter your destination by voice to depart..."

[0056] As part of the generated text display / audio output step, the user terminal 110 displays the received generated text data GT in the browser 111 or outputs it as audio. More specifically, the user terminal 110 displays an example of generated text data GT on the talk screen 114 of the browser 111, or outputs it as audio using the speaker, stating, "It is fully electronically controlled, making it well-suited for autonomous driving. Enter your destination by voice to depart..." As a result, the RAG-learned content is converted into a vector and stored at vector index 121, the question information PT is converted into a vector, a so-called correlation search is performed, and content highly related to the question is selected, and an answer is generated based on this. As a result, users can obtain highly accurate answers to their questions.

[0057] In this embodiment, as shown in Figure 12, the browser 141 of the administrator terminal 140 displays the administrator settings screen 142. The administrator settings screen 142 includes a permission setting section 143 for skill function management. For example, there are clauses allowing users to create skill functions, allowing users to share skill functions within the company, allowing users to view the contents of skills, and allowing users to duplicate skill functions. By checking the box that allows access to the content of skill functions, users will be able to freely access the electronic information resource information of registered skill functions. As a result, users can see which skill functions are using what kind of content as electronic information resources.

[0058] Furthermore, in this embodiment, as shown in Figure 3, the character's face image data FD, the character's voice data VD, and the speaking style setting information SD are input or selected on the setting screen 112. Then, as shown in Figure 6, let's assume that a skill function, which is a pre-trained artificial intelligence means set for the character, is selected on the skill function selection screen 113. Then, as shown in Figure 13, the browser 111 of the smartphone user terminal 110 displays the character's face image data FD on the talk screen 114 based on the character's setting information.

[0059] The chat screen 114 of the browser 111 on the smartphone user terminal 110 is equipped with character face image data FD, a microphone button 114b, and a "email history" button 114d. When the "Email History" button 114d is pressed, the microphone becomes active. When a user performs a voice input operation in the browser 111 of the user terminal 110, the voice data VD is not converted to text, for example, and the user terminal 110 sends the voice data VD to the first server 120 as question information PT. So-called streaming transmission is preferred.

[0060] Then, the first server 120 converts the received audio data VD into text data, and uses a pre-trained artificial intelligence skill function to generate text data as an answer to the question information PT. In addition, the first server 120 converts the generated text data GT into audio data VD and sends the audio data VD to the user terminal 110. The browser 111 of the user terminal 110 is configured to output the received audio data VD as audio based on the character's setting information.

[0061] As a result, character face image data FD is displayed on the talk screen 114, and when a question is asked via voice input, an answer is obtained via voice output. As a result, users can interact with the characters as if they were actually talking to them. Furthermore, based on the speech style setting information SD, voice data VD is output as the response. As a result, it becomes possible to give characters individuality and distinctiveness to their way of speaking.

[0062] Furthermore, regarding the face image data FD, the image generation means, which is a pre-trained artificial intelligence means of the first server 120, may be configured to generate a video that appears to be speaking based on the face image data FD of the registered skill function, and when the audio data VD as a response is output as audio on the talk screen 114 of the browser 111 of the user terminal 110, the video that appears to be speaking may be played on the talk screen 114 in synchronization with the audio. In other words, when the audio output is interrupted or stopped, the system may be configured to either stop playing the video that appears to be speaking or to play a video where the mouth is closed, and when the audio output is present, it may be configured to play a video where the mouth is moving as if speaking. In other words, the audio output's pronunciation is linked to the mouth movements in the video, and the video is played back. The first server 120 may use pronunciation recognition to grasp the content of the audio data VD to be sent to the user terminal 110, and link images of mouth openings corresponding to the pronunciation generated by the image generation means to the content grasped by the audio pronunciation recognition, and display them on the user terminal 110's talk screen 114 as a video consisting of a flipbook animation like a GIF.

[0063] In this embodiment, as shown in Figure 3, the character's face image data FD is the face of a real person and is the face of a person registered in the address book, and an email address from the address book is specified in the settings screen 112. Furthermore, as shown in Figure 13, on the chat screen 114 of the browser 111 of the smartphone user terminal 110, if the user performs a voice input operation as a question, and the voice data VD received from the first server 120 is output as a response based on the character's setting information, the browser 111 displays a "Email History" button 114d on the chat screen 114, which indicates that the history will be emailed. When the "Email History" button 114d is pressed, the first server 120 is configured to send text data QA about the question information PT and answer information to the email address corresponding to the character.

[0064] As a result, as shown in Figure 14, if the character is based on a real person, the content of the interaction between the user and the character in the chat screen 114 will be sent via email to the email address of the person associated with that character. As a result, the person behind the character can receive email notifications about the interactions between their character and the user. In other words, you will be notified via email of the content of your avatar's interactions with other users. As a result, the character themselves can understand the interactions between their avatar and the user. Alternatively, you may attach a video of the conversation to the email.

[0065] The artificial intelligence system 100, an embodiment of the present invention obtained in this manner, comprises a user terminal 110 and a first server 120 and a second server 130 as servers. The user terminal 110 communicates with the first server 120, and the browser 111 of the user terminal 110 displays a settings screen 112 for the text generation means 131, which is a pre-trained artificial intelligence means of the second server 130. On the settings screen 112, name information NM is entered, and at least one of electronic file information, website URL information, and database information as an electronic information resource is specified. The user terminal 110 uploads the information of the electronic information resource specified along with the name information to the first server 120. The first server 120 extracts language character file data TX as language character information from the electronic information resource based on the received information of the electronic information resource, and pre-trains the language character file data TX as language character information on the second server 130. The system is configured such that a pre-trained artificial intelligence means, the text generation means 131, is trained, and the browser 111 of the user terminal 110 displays the name information NM of the trained pre-trained artificial intelligence means, the text generation means 131, along with its icon data CN, in a selectable manner on the skill function selection screen 113. When the name information NM of the trained pre-trained artificial intelligence means, the text generation means 131, or its icon data CN, is selected, a talk screen 114 is displayed that allows interaction with the pre-trained artificial intelligence means, the text generation means 131, which has already learned the specified electronic information resource. As a result, the user can start interacting with the pre-trained artificial intelligence means, the text generation means 131, which has already learned the specified electronic information resource, from a pre-trained state. Furthermore, the user can engage in interaction with the pre-trained artificial intelligence means, the text generation means 131, which has learned the specified electronic information resource, and obtain the desired response.

[0066] Furthermore, in the settings screen 112, each time a name information NM is entered and an electronic information resource is specified, the first server 120 extracts language character file data TX as language character information from that electronic information resource and trains the pre-trained artificial intelligence means, text generation means 131, of the second server 130. The browser 111 of the user terminal 110 displays the name information NM of multiple trained pre-trained artificial intelligence means, text generation means 131, freely selectable on the skill function selection screen 113. This configuration allows users to select the name information NM of a desired pre-trained artificial intelligence means, text generation means 131, that is specialized in a particular information field or content, and interact with that pre-trained artificial intelligence means, text generation means 131.

[0067] Furthermore, the first server 120 divides the extracted language character file data TX into multiple divided language character file data DTX, converts each of the divided divided language character file data DTX into vector coordinate system divided language data VTX using a vector coordinate transformation means, which is a pre-trained artificial intelligence means, and stores the converted vector coordinate system divided language data VTX in vector index 121 in association with the original divided language character file data DTX. When a user inputs text or voice in the browser 111 of the user terminal 110, the user terminal 110 sends the text data or voice data VD as question information PT to the first server 120, and the first server 120 converts the received question information PT into vector coordinate system question data VPT using the vector coordinate transformation means, and vector - At index 121, the system searches for the vector coordinate system segmented language data VTX that is closest to the vector of the vector coordinate system question data VPT, selects the closest related vector coordinate system segmented language data VTX, or the segmented language character file data DTX corresponding to the top multiple closest related vector coordinate system segmented language data VTXs, the first server 120 adds the selected segmented language character file data DTX to the question information PT, and uses the pre-trained artificial intelligence means, text generation means 131, of the second server 130 to perform text data generation, converts the generated text data GT obtained from the execution into audio data VD according to the settings and sends it to the user terminal 110 as answer information, and the user terminal 110 displays the received answer information in the browser 111 or outputs it as audio. This configuration allows the user to obtain highly accurate answers to questions.

[0068] Furthermore, when the character's face image data FD, character's voice data VD, and speaking style setting information SD are input or selected on the settings screen 112, and the name information NM of the pre-trained artificial intelligence means, which is the text generation means 131, is selected on the skill function selection screen 113, the browser 111 of the user terminal 110 displays the character's face image data FD on the talk screen 114 based on the setting information for the character, and when the user performs a voice input operation on the browser 111 of the user terminal 110, the user terminal 110 uses the voice data VD as question information PT. The system is configured such that the user sends the information to the first server 120, which uses the pre-trained artificial intelligence means of the second server 130, the text generation means 131, to generate text data as an answer to the received question information PT, converts the generated text data GT into audio data VD, and sends the audio data VD to the user terminal 110, and the browser 111 of the user terminal 110 outputs the received audio data VD as audio based on the character's setting information. This configuration allows the user to interact with the character as if they were actually talking to them, and furthermore, allows the character's way of speaking to have individuality and characteristics.

[0069] Furthermore, if the character's face image data FD is the face of a real person and is registered in the address book, and an email address from the address book is specified in the settings screen 112, then in the talk screen 114, when the user performs a voice input operation as a question, and the voice data VD received from the first server 120 is output as a response based on the character's setting information, the browser 111 displays a button on the talk screen 114 to email the history, and when the button to email the history is operated, the first server 120 sends text data QA about the question information PT and answer information to the email address corresponding to the character. As a result, the person whose character it is can be notified of the interaction between their character and the user via email.

[0070] Furthermore, the program of the artificial intelligence system 100, which is an embodiment of the present invention, includes: an input setting presence determination step S2 in which the user terminal 110 communicates with the first server 120, the browser 111 of the user terminal 110 displays a setting screen 112 for the text generation means 131, which is a pre-trained artificial intelligence means of the second server 130, and determines whether name information NM has been entered on the setting screen 112 and whether at least one of electronic file information as an electronic information resource, website URL information, or database information has been specified; an upload step S3 in which the user terminal 110 uploads the information of the electronic information resource specified along with the name information to the first server 120; and the first server 120 extracts language character file data TX as language character information from the electronic information resource based on the received information of the electronic information resource, and outputs the language character file data TX as language character information to the text generation means of the pre-trained artificial intelligence means of the second server 130. By having the computer execute a learning step S4 in which the text generation means 131 is trained, and a talk screen display step S5 to S7 in which the browser 111 of the user terminal 110 displays the name information NM of the pre-trained artificial intelligence means, which is the text generation means 131, in a selectable manner on the skill function selection screen 113, and when the name information NM of the pre-trained artificial intelligence means, which is the text generation means 131, is selected, a talk screen 114 is displayed in which the user can interact with the pre-trained artificial intelligence means, which is the text generation means 131, in a state where it has already learned the specified electronic information resource, the user can start interacting with the pre-trained artificial intelligence means, which is the text generation means 131, in a state where it has already learned the specified electronic information resource, and furthermore, the user can interact with the pre-trained artificial intelligence means, which is the text generation means 131, in a manner specific to the content of the specified electronic information resource, and obtain the desired answer, the effect is enormous. [Explanation of Symbols]

[0071] 100 ··· Artificial Intelligence System 110 ··· User terminal 111... (User's device) browser 112... Settings screen 113... Skill function selection screen 114 ··· Chat screen 114a... Input field 114b... Microphone button 114c... Rich menu button 114d... "Email History" button 120 ··· Server 1 (Server) 121... Vector index (spatial map) 122 ··· Bot 130 ··· Second Server (Server) 131 ··· Text generation methods (pre-trained artificial intelligence methods, large-scale language model LLM, chat GPT) 140 ··· Administrator terminal 141... (Administrator terminal) browser 142...Administrator settings screen 143 ··· Permission settings section NM... Name information CN ··· Icon data TX ··· (Referenced) Language character file data (language character information) DTX ·· Split language character file data VTX... Language data after vector coordinate system division PT... Question Information (Prompt) VPT · Vector Coordinate System Question Data GT ··· Generated text data FD... (Character) Face image data (face setting information) VD... (Character) Voice data (Voice setting information) SD... Speaking style settings information (settings information) Q&A... Text data about question and answer information

Claims

1. An artificial intelligence system comprising a user terminal and a server, wherein the user terminal displays or outputs audio based on the server, The user terminal communicates with the server, and the user terminal's browser displays a settings screen for the server's pre-trained artificial intelligence means, which is a text generation means. In the aforementioned settings screen, when name information is entered and at least one of the following is specified as an electronic information resource—electronic file information, website URL information, or database information—the user terminal uploads the information of the specified electronic information resource along with the name information to the server. The server extracts language character information from the electronic information resource based on the information received from the electronic information resource, and uses the language character information to train a pre-trained artificial intelligence means. An artificial intelligence system characterized in that the browser of the user terminal displays the name information of the pre-trained artificial intelligence means that has been trained, in a selectable manner on the skill function selection screen, and when the name information of the pre-trained artificial intelligence means is selected, a talk screen is displayed that allows interaction with the pre-trained artificial intelligence means of the name information.

2. In the aforementioned settings screen, each time name information is entered and an electronic information resource is specified, the server extracts linguistic character information from that electronic information resource and uses it to train a pre-trained artificial intelligence means. The artificial intelligence system according to claim 1, characterized in that the browser of the user terminal is configured to display, in a selectable manner, the name information of multiple pre-trained artificial intelligence means on the skill function selection screen.

3. The server divides the extracted language character information data into multiple divided language character file data, converts each of the divided divided language character file data into vector coordinate system divided language data using a vector coordinate transformation means which is a pre-trained artificial intelligence means, and stores the converted vector coordinate system divided language data in vector indexes, relating them to the original divided language character file data. When a user inputs text or voice in the browser of the user terminal, the user terminal sends the text data or voice data to the server as question information. The server converts the received question information into vector coordinate system question data using vector coordinate transformation means, searches the vector index for the vector coordinate system segmented language data that is closest to the vector of the vector coordinate system question data, and selects the segmented language character file data that corresponds to the closest vector coordinate system segmented language data, or the top multiple closest vector coordinate system segmented language data. The server adds the selected segmented language character file data to the question information, uses a pre-trained artificial intelligence means, which is a text generation means, to generate text data, converts the generated text data into audio data according to the settings, and sends it to the user terminal as answer information. The artificial intelligence system according to claim 2, characterized in that the user terminal is configured to display the received response information in a browser or output it as audio.

4. When the character's face image data, character's voice data, and speaking style settings are entered or selected on the settings screen, and the name information of a pre-trained artificial intelligence means set for the character is selected on the skill function selection screen, the browser of the user terminal displays the character's face image data on the talk screen based on the character's settings information. When a user performs a voice input operation in the browser of the user terminal, the user terminal sends the voice data to the server as question information. The server generates text data as an answer to the received question information using a pre-trained artificial intelligence means, converts the generated text data into audio data, and transmits the audio data to the user terminal. The artificial intelligence system according to any one of claims 1 to 3, characterized in that the browser of the user terminal is configured to output the received audio data as audio based on setting information about the character.

5. If the character's facial image data is the face of a real person and is registered in the address book, and an email address from the address book is specified on the settings screen, and the user makes a voice input operation as a question on the chat screen, and the voice data received from the server is output as a response based on the character's setting information, the browser displays a button on the chat screen to email the history. The artificial intelligence system according to claim 4, characterized in that when a button indicating to email the history is pressed, the server sends text data of the question information and answer information to the email address corresponding to the character.

6. A program for an artificial intelligence system that displays or outputs audio based on a server on a user terminal, The user terminal communicates with the server, and the user terminal's browser displays a settings screen for the server's pre-trained artificial intelligence means, which is a text generation means. The input setting presence determination step determines whether name information has been entered on the settings screen and whether at least one of the following has been specified: electronic file information as an electronic information resource, website URL information, or database information. The user terminal uploads information of the specified electronic information resource along with name information to the server in an upload step, The server extracts language character information from the electronic information resource based on the information of the received electronic information resource, and performs a learning step in which it trains a pre-trained artificial intelligence means with the language character information. A program for an artificial intelligence system characterized in that the user terminal's browser displays the name information of the pre-trained artificial intelligence means that has been trained, in a selectable manner on the skill function selection screen, and when the name information of the pre-trained artificial intelligence means is selected, the computer executes a talk screen display step that displays a talk screen in which interaction with the pre-trained artificial intelligence means of the name information is possible.