Intelligent personalized interaction system and construction method thereof
By building an intelligent personalized interaction system that integrates deep learning and natural language processing technology, the shortcomings of the existing system in semantic understanding and personalized services are solved, and accurate understanding and personalized interaction of user needs are achieved, which significantly improves the user experience.
Patent Information
- Application Number
- CN202510227275.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing intelligent interaction system has significant flaws in semantic understanding, personalized services and handling complex problems, and cannot deeply understand user needs, provide accurate and effective solutions and personalized interactions.
Build an intelligent personalized interaction system to achieve accurate understanding and personalized interaction of user problems by integrating deep learning, natural language processing and other technologies. The system includes a problem classification module, a user profile building module, a workflow scheduler and a pre-trained large language model. Through the collaborative work of multiple modules, we generate recommendation questions that meet user interests and continuously optimize using user feedback.
It significantly improves the user experience, can deeply understand user needs, provide accurate personalized interactive services, and meet users' diverse needs in different scenarios.
Smart Images

Figure CN120216630A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing, and particularly focuses on an intelligent system capable of deeply understanding user needs and achieving personalized interaction, as well as the construction and usage methods of this system. By integrating a variety of advanced technologies, it aims to improve the accuracy of intelligent interaction and the user experience, and meet the diverse interaction needs of users in different scenarios. Background Art
[0002] In today's digital age, intelligent interaction systems have been widely penetrated into various fields, such as intelligent customer service, online education, intelligent assistants, etc., and have become important tools for people to obtain information and solve problems. However, there are many significant defects in existing intelligent interaction systems, which seriously limit the improvement of their application effects and user experiences.
[0003] Traditional intelligent interaction systems have serious deficiencies in understanding the semantics and intentions of user questions. Most systems rely on simple keyword matching algorithms and are unable to deeply explore the complex semantic relationships and context information in questions. For example, in the intelligent customer service scenario, when a user asks "The mobile phone I bought last month frequently freezes recently, and it gets extremely hot when charging. How can this be solved?", a conventional intelligent customer service system may only conduct simple searches and responses based on keywords such as "mobile phone", "freezing", "charging and heating", and it is difficult to comprehensively understand key information such as the user's specific usage scenario, purchase time, and mobile phone model, thus being unable to provide accurate and effective solutions. This not only causes users to need to communicate repeatedly to clarify their needs but may also lead to users' dissatisfaction and distrust of the system.
[0004] Existing systems perform poorly in meeting users' personalized needs. They lack in-depth analysis and mining of users' interests, preferences, and historical behaviors, and are unable to provide personalized interaction services according to the unique characteristics of different users. Take an online education platform as an example. The system cannot customize learning plans and recommend suitable learning materials for each student based on their learning progress, knowledge mastery, and learning interests. This makes it difficult for students to obtain targeted learning support, resulting in uneven learning effects and the inability to fully leverage the advantages of online education. In the application of intelligent assistants, the system is also unable to proactively provide personalized function recommendations and services based on users' daily usage habits and needs, leading to a poor user experience and the inability to meet users' growing personalized needs. With the rapid development of big data and artificial intelligence technologies, users have put forward higher requirements for the real-time performance, accuracy, and intelligence level of intelligent interaction systems. Traditional systems are inefficient in processing large-scale and high-dimensional data and are difficult to meet users' demands for instant response. Facing complex and ever-changing user questions, traditional systems have weak generalization capabilities and are prone to errors or inaccurate answers. For example, when dealing with comprehensive questions involving multi-domain knowledge, traditional systems often cannot integrate relevant information to provide comprehensive and accurate answers. In the face of users' vague or metaphorical expressions, the system is even more difficult to understand the true intentions of users, resulting in interaction failures.
[0005] In summary, the existing intelligent interaction systems have serious deficiencies in aspects such as semantic understanding, personalized services, and handling complex problems. There is an urgent need for an innovative intelligent personalized interaction system to fill these gaps, improve the quality and efficiency of intelligent interaction, and meet users' growing diverse needs. Summary of the Invention
[0006] The core objective of the present invention is to provide an intelligent personalized interaction system and its construction method to achieve in-depth understanding, precise analysis, and personalized interaction of user questions, thereby significantly improving the user experience and meeting users' diverse needs in different scenarios.
[0007] To achieve the above objective, the present invention provides a construction method for an intelligent personalized interaction system, including:
[0008] S100: Construct a database;
[0009] S200: Prepare a training set and a test set with question texts, construct a vocabulary of high-frequency words for them, and perform text encoding;
[0010] S300: Construct a workflow scheduler for counting high-frequency question types;
[0011] S400: Establish a user profile construction module, which is used to obtain a user profile based on the historical conversations, user questions, and high-frequency question types in the database;
[0012] S500: Integrate to obtain an intelligent personalized interaction system, which includes a user question input interface, a question classification module, and a database connected in sequence. The database is also connected to a workflow scheduler and a user portrait construction module; both the user portrait and the high-frequency question types are output to a pre-trained large language model to generate recommended questions.
[0013] The table structure of the database includes a user question table and a user portrait information table; the user question table includes: a question primary key field, a user identification field, a question text field, a question time field, and a question type field; the user portrait information table is used to store user portrait information, including a user identification primary key field associated with the user identification field, and an interest tag field and a behavior pattern field; and / or
[0014] The question classification module adopts an architecture that combines a convolutional neural network and a recurrent neural network; and / or
[0015] The user portrait construction module also receives recommended question feedback data from the user, and uses the recommended question feedback data to continuously optimize the user portrait and the question generation strategy of the pre-trained large language model; and / or
[0016] The database, the workflow scheduler are connected to a question generation module, and the question generation module feeds the recommended questions back to the user according to the high-frequency question types and the data in the database.
[0017] The step S200 specifically includes:
[0018] S210: Preprocess the question texts in the training set and the test set, including removing stop words, lemmatization, and stemming;
[0019] S230: Construct a vocabulary based on the high-frequency words in the question texts, and then encode the question texts according to the vocabulary to obtain the encoded training samples and test samples.
[0020] After step S210 and before step S230, there is also step S220: Use a pre-trained language model to perform semantic enhancement processing on the question texts; after the step S230, there is also step S240: Standardize the encoded training samples and test samples; at the same time, use data augmentation techniques to generate new training samples.
[0021] The step S300 specifically includes:
[0022] S310: Set the startup conditions of the workflow scheduler. The startup conditions of the workflow scheduler include starting at a predetermined time interval and / or starting based on specific trigger conditions; the specific trigger conditions include that the number of questions asked by the user within a fixed time period reaches a preset threshold, or the user actively requests personalized recommendations;
[0023] S320: Set the workflow scheduler to use a sliding window algorithm to count high-frequency question types.
[0024] The fixed time period in the specific trigger conditions is adjustable between 10 minutes and 2 hours, and the preset threshold is adjustable between 3 and 20 questions; when using the sliding window algorithm, set a fixed time window and step size.
[0025] The specific steps of step S400 include:
[0026] S410: Establish a user interest analysis module. Based on the user's input questions and historical conversations, using sentiment analysis techniques in natural language processing, obtain keywords such as the user's sentiment tendency towards different questions and the user's interest fields; and use the high-frequency question types from the workflow scheduler as keywords for question frequency;
[0027] S420: Establish a function for generating a user profile. The function for generating a user profile is set to generate a prompt template according to the user's questions, historical conversations, keywords, and combine with the user profile, and call a Transformer analysis model based on the attention mechanism to generate a user profile. The user profile includes the user's interest focus, behavior pattern, and preferred question types.
[0028] The specific steps of step S420 include:
[0029] S421: Select a Transformer analysis model based on the attention mechanism as the basic model for the function of generating a user profile;
[0030] S422: Construct a prompt template for generating a user profile, which is used to incorporate keywords such as the user's interest field, question frequency, and sentiment tendency into the Transformer analysis model based on the attention mechanism in the form of prompt words;
[0031] S423: Define an analysis chain, which is set to pass the user's questions, historical conversations, keywords, and the prompt template for generating a user profile to the Transformer analysis model together to generate a user profile.
[0032] The specific steps of step 500 include: integrating to obtain an intelligent personalized interaction system; conducting system testing and debugging, including unit testing, integration testing, and debugging and optimization. For unit testing, test cases are written for the functions of each module. For integration testing, real user usage scenarios are simulated to verify the overall function of the system and test the high-concurrency performance. For debugging and optimization, debugging tools are used to debug the problem modules, and log information is collected and analyzed for improvement.
[0033] On the other hand, the present invention provides an intelligent personalized interaction system, which is established by adopting the construction method of the intelligent personalized interaction system described above.
[0034] The construction method of the intelligent personalized interaction system of the present invention integrates technologies such as deep learning and natural language processing to achieve accurate understanding of user problems and personalized interaction. The problem classification module uses deep learning algorithms to improve the accuracy of problem processing. The workflow scheduler captures hot demands. The user portrait helps with personalized services. The pre-trained large language model enhances the interaction experience, and it can continuously optimize using user feedback. This system has a wide range of application scenarios, covering fields such as intelligent customer service and online education, can provide efficient services for different users, and promote the intelligent development of various fields. For example, in intelligent customer service, solutions and products can be accurately recommended, and in online education, learning materials can be recommended according to the situation of students. In terms of technical implementation, each step collaborates closely. Database design and index optimization provide a data foundation. The problem classification module accurately understands problems. The workflow scheduler keeps up with demand changes. The user portrait deeply understands users. System integration and optimization give play to the advantages of modules. In the future, the system can combine technologies such as computer vision and speech recognition to achieve multimodal interaction, expand the data source, improve the accuracy of recommendations, and strengthen security and privacy protection. With the progress of technology and the expansion of scenarios, it will play a greater role and bring more convenience. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is an architecture diagram of an intelligent personalized interaction system constructed by the construction method of an intelligent personalized interaction system of the present invention.
[0036] Figure 2 It is a schematic diagram of the sliding window algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following further describes the present invention in conjunction with specific embodiments. It should be understood that the following embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0038] Such as Figure 1The figure shows the architecture diagram of an intelligent personalized interaction system constructed by a method for constructing an intelligent personalized interaction system of the present invention. Among them, the intelligent personalized interaction system includes a user question input interface, a question classification module 10, and a database 20 connected in sequence. The database is also connected to a workflow scheduler 30, a question generation module 40, and a user portrait construction module 50. In addition, the user portrait of the user portrait construction module 50 and the high-frequency question types of the workflow scheduler 30 are both output to a pre-trained large language model 60 to generate recommended questions that match the user's interests. And the pre-trained large language model retrieves relevant historical questions and answers from the database for the user to refer to.
[0039] The user question input interface is used to receive user questions. A question recording and storage module is provided between the user question input interface and the question classification module 10. The question recording and storage module is used to store the user questions in the database. The question classification module 10 is used to obtain the question type according to the user question and store it. The database is used to store user questions and question types. The workflow scheduler counts high-frequency question types when starting. The user portrait construction module is used to obtain the user portrait according to the historical conversations, user questions, and high-frequency question types in the database.
[0040] The question generation module 40 and the pre-trained large language model 60 complement each other. Among them, the question generation module 40 focuses on using the existing high-frequency question types and the information stored in the database in the initial stage of the operation of the intelligent personalized interaction system or when the user's needs are relatively vague to perform preliminary question generation. Specifically, it is used to feed back recommended questions to the user according to the high-frequency question types and the data in the database to obtain new questions and feedback, so as to reduce the burden on the pre-trained large language model. The pre-trained large language model 60 comprehensively considers various factors such as high-frequency question types, the user's interest fields, question preferences, and emotional tendencies to generate targeted and personalized questions. Specifically, it outputs recommended questions that match the user's interests according to the user portrait and high-frequency question types, and retrieves relevant historical questions and answers from the database for reference. The user portrait construction module 50 also receives the recommended question feedback data from the user, and uses the feedback data of the user on the generated questions to continuously optimize the user portrait and the question generation strategy of the pre-trained large language model, and improve the personalized interaction ability of the system.
[0041] The working process of the intelligent personalized interaction system is as follows:
[0042] First, after the user inputs a question to the intelligent personalized interaction system, the question classification module in the intelligent personalized interaction system deeply analyzes the user question based on the fusion architecture of a convolutional neural network and a recurrent neural network, obtains the question type, and stores the question type in the database.
[0043] Through the learning and training of a large number of labeled samples, the problem classification module enables the convolutional neural network to accurately extract local features of the problem text. For example, for technical problems, it can quickly identify key technical terms; the recurrent neural network deeply analyzes the semantics of the text sequence, comprehensively understands the problem logic, and then realizes the accurate classification of the problem and stores the problem type in the database. Among them, the user's problem text data (i.e., the user's question) will be input as the current input data of the recurrent neural network into the hidden layer of the recurrent neural network; the data of the previous time step of the recurrent neural network includes the high-frequency problem type and user portrait information of the previous time step.
[0044] The other modules of the intelligent personalized interaction system other than the problem classification module do not need to use the recurrent neural network and the convolutional neural network.
[0045] Secondly, the workflow scheduler starts according to the set start conditions and counts the high-frequency problem types when starting. The start conditions of the workflow scheduler include starting at a predetermined time interval (6 - 24 hours) and / or starting according to specific trigger conditions. The specific trigger conditions are, for example, that the number of questions asked by the user within a fixed time period (adjustable from 10 minutes to 2 hours) reaches a preset threshold (adjustable from 3 to 20 questions), or the user actively requests personalized recommendations.
[0046] The workflow scheduler internally uses complex and efficient query and statistical logic, adopts the sliding window algorithm to count the high-frequency problem types, and by setting a fixed time window (such as 1 hour) and step size (such as 15 minutes), it updates the problem type statistical results in real time to provide the latest data support for system decision-making.
[0047] In the present invention, the workflow scheduler is implemented through relevant libraries (such as the APScheduler library) and corresponding code logic, which involves the implementation of functions such as setting start conditions and counting high-frequency problem types. The implementation code of these functions and related operations play a role similar to scheduling. In the part of setting the start conditions of the workflow scheduler, by installing the APScheduler library and writing code to start at a predetermined time interval (such as starting the workflow every 12 hours) and / or starting according to specific trigger conditions (such as triggering according to the number of questions of the user within a fixed time period), the start of the workflow is controlled. In the part of counting the high-frequency problem types of the workflow scheduler, through writing code logic for database connection, data acquisition, and statistical calculation using the sliding window algorithm, the query and statistical functions inside the workflow are realized. These operations together constitute the implementation process of the workflow scheduler.
[0048] The user portrait construction module starts when inputting the user's question to interact with the intelligent personalized interaction system, and constructs the user portrait based on the user's questions and historical conversations stored in the database.
[0049] Among them, the user portrait construction module includes a user interest analysis module and a user portrait generation function based on the Transformer analysis model with an attention mechanism as the basic model.
[0050] Based on the user's questions and historical conversations, the user interest analysis module uses sentiment analysis technology in natural language processing to judge the user's sentiment tendency towards different questions, so as to further refine the user interest information. For example, it judges whether the user is actively concerned or negatively treated in a specific field, and at the same time analyzes the user's interest fields, etc., and transfers keyword information such as the user's interest fields and sentiment tendencies to the user portrait generation function. These keywords help the user portrait construction module further analyze the user's interest fields and question preferences, etc., so as to further refine the user interest information and improve the user portrait.
[0051] The user portrait generation function takes the user's questions, historical conversations, and keywords provided by the user interest analysis module as inputs. The user portrait generation function receives the keywords from the user interest analysis module, takes the high-frequency question types from the workflow scheduler as keywords for question frequency, and sets them as the user portrait generation prompt template according to the user's questions, historical conversations, keywords, and combines them with the user portrait generation prompt template. Then it calls the Transformer analysis model with an attention mechanism to generate the user portrait. The user portrait contains the user's interest focus and behavior pattern.
[0052] The user portrait generation function includes a Transformer analysis model with an attention mechanism, a user portrait generation prompt template, and an analysis chain. Among them, the Transformer analysis model can effectively capture the key information and semantic relationships of the text with its powerful sequence modeling ability and attention mechanism; the attention mechanism enables the model to pay more attention to the important parts when processing the user's questions and historical conversations, improving the ability to understand the user's intentions and laying a foundation for generating an accurate user portrait. The user portrait generation prompt template includes keywords such as the user's interest fields, question frequency, and sentiment tendency, thus providing clear guidance for the model to generate the user portrait. The analysis chain is used to transfer the user's questions, historical conversations, keywords as inputs to the user portrait generation prompt template, and the user portrait generation prompt template to the Transformer analysis model together to generate a user portrait containing multi-dimensional information such as the user's interest focus, behavior pattern, and preferred question types.
[0053] The functions of the user interest analysis module include: 1. Deepening the insight into user interests, mining relevant topics in the field that the user is interested in on the basis of classification, so as to enrich the dimensions of the user profile. 2. Optimizing the interaction process and strategies: Based on the output of the user interest analysis module, the system can dynamically adjust the question recommendation and interaction strategies. If it is found that the user's attention to a certain specific field has increased recently, the system will increase the recommendation frequency of questions in this field and optimize the presentation of questions. 3. Enhancing the system's adaptive ability to changes in user interests and maintaining a good interactive relationship with the user.
[0054] The functions of the user profile generation function include: 1. Driving the generation of personalized questions and improving the matching degree between the automatically generated questions and user interests. 2. Helping the system to accurately retrieve answers. If the user is interested in the field of artificial intelligence, the system can quickly provide past high-quality answers in this field, provide more targeted reference materials for the user, help the user to deeply study and explore the content of interest, and improve the efficiency of the user to obtain information.
[0055] That is to say, the data flow direction in the user profile construction module is as follows: The user questions and historical conversations generated by the interaction between the user and the system first flow into the user interest analysis module. After analysis and processing, the obtained sentiment tendency and interest preference information flow to the user profile generation function, and finally the user profile is generated. The generated user profile may be stored in the database or memory for subsequent data interaction with other modules (such as the question classification module, workflow scheduler, pre-trained large language model, etc.), laying a foundation for the system to provide personalized services.
[0056] The question generation module 40 is used to feedback the recommended questions to the user according to the high-frequency question types and the data in the database to obtain new questions and recommended question feedback data. The question generation module 40 is set to provide some common question references for the user based on the statistically high-frequency question types at the initial stage of the operation of the intelligent personalized interaction system or when the user's needs are relatively vague, and obtain question feedback data to help the user further clarify their demand direction. Thus, the question generation module 40 and the pre-trained large language model 60 complement each other. The question generation module 40 focuses on using the existing high-frequency question types in the system and the information stored in the database for preliminary question generation, which is a direct application based on the internal data statistics and storage of the intelligent personalized interaction system. The pre-trained large language model will comprehensively consider various factors such as high-frequency question types, the user's interest fields, question preferences, and sentiment tendencies to generate targeted and personalized questions; the question generation module 40 can reduce the burden on the large language model, enabling the pre-trained large language model 60 to focus more on processing personalized questions that require in-depth analysis of the user profile and complex semantics, thereby optimizing the overall performance and response speed of the system and ensuring the stability of the intelligent personalized interaction system under high concurrency.
[0057] The difference between the question generation module 40 and the pre-trained large language model 60 lies in that:
[0058] The questions generated by the question generation module are relatively basic and general, mainly centered around the high-frequency question types statistically obtained in the system, with relatively weak pertinence.
[0059] However, the questions generated by the pre-trained large language model are highly personalized and targeted. It deeply mines the user profile information, and according to each user's unique interests, preferences and emotional tendencies, combined with the current high-frequency question dynamics, generates questions that can precisely meet the individual needs of users, which is more conducive to improving the user's interaction experience and the efficiency of obtaining information.
[0060] Thus, the intelligent personalized interaction system constructed by the construction method of the present invention integrates the functions of each module. The question classification module stores the question types in the database and passes the question types to the workflow scheduler to statistically obtain the high-frequency question types. The question generation module is used to feed back recommended questions to the user according to the high-frequency question types and the data in the database to obtain new questions and feedback. The user profile construction module receives the user questions and historical conversations provided by the database and the high-frequency question types from the workflow scheduler to obtain the user profile.
[0061] In addition, the workflow scheduler provides the high-frequency question types to the pre-trained large language model, and the user profile construction module feeds back the user profile information (including the user's interest field, question preference, emotional tendency) to the pre-trained large language model. The pre-trained large language model generates recommended questions that fit the user's interests based on the user profile obtained by the user profile construction module and the high-frequency question types obtained by the workflow scheduler, and retrieves relevant historical questions and answers from the database for the user to refer to, thereby realizing the comprehensive consideration of factors such as high-frequency question types, user interest fields, question preferences, and emotional tendencies.
[0062] In addition, during the interaction with the user, the system collects the feedback data of the user on the generated questions, including the recommended question feedback data such as whether the user clicks to view the recommended questions, whether to ask further questions, the content and direction of the questions, the evaluation and feedback on the recommended questions. These recommended question feedback data can also be input into the user profile construction module, so as to continuously optimize the user profile and the question generation strategy of the pre-trained large language model by using these feedback data, such as adjusting the weights of the user profile interest tags, the input parameters and generation algorithms of the pre-trained large language model, the statistical logic and trigger conditions of the workflow scheduler, etc., thereby continuously improving the personalized interaction ability of the system and creating a personalized interaction experience for the user.
[0063] Step S100: Construct the database 20 as the data storage basis;
[0064] Among them, building the data storage foundation mainly refers to selecting a suitable database (such as MySQL) and performing operations at the software level such as table structure design and connection configuration of the database, as well as data initialization and testing.
[0065] Step S100 specifically includes:
[0066] Step S110: Database construction and configuration, selecting a suitable database and performing table structure design and connection configuration.
[0067] Among them, the table structure of the database includes: user question table and user portrait information table;
[0068] The user question table includes: a question primary key field (used to uniquely identify each user question record), a user identification field (of string type to effectively associate user identities), a question text field (of text type to retain the details of user questions completely), a question time field (set as a timestamp type accurate to the second level to precisely track the question time), and a question type field (initially empty for subsequent filling of question types);
[0069] The user portrait information table is used to store user portrait information, including a user identification primary key field associated with the user identification field, as well as an interest tag field and a behavior pattern field.
[0070] Step S110 is specifically broken down into the following sub-steps:
[0071] Step S111: Database selection and installation;
[0072] In the step S111, the MySQL database is selected to build the data storage foundation. The installation operation of the MySQL database is performed on the server to provide platform support for subsequent data storage and management.
[0073] Step S112: Database and table structure creation;
[0074] In step S112, after the installation of the MySQL database, a database named "intelligent_interaction_system" is created to store various types of data involved in the operation of the system. In this database, a user question table "user_questions" is created, and its structure is designed as follows: CREATE TABLE user_questions(question_id INT AUTO_INCREMENT PRIMARY KEY,user_id VARCHAR(255)NOT NULL,question_text TEXT NOT NULL,question_time TIMESTAMP NOT NULL,question_typeVARCHAR(255)DEFAULT NULL); Among them, the question primary key field "question_id" is used as the primary key, with an auto-incrementing integer type, to uniquely identify each user question record; the user identification field "user_id" is of string type, with a length of 255, to store the unique identification of the user, ensuring that each user can be accurately associated; the question text field "question_text" is of text type, to store the details of the questions raised by the user in full; the question time field "question_time" is set to timestamp type, accurate to the second level, to accurately track the specific moment when the user raises the question; the question type field "question_type" is also of string type, with a length of 255, initially empty, and will be used to fill in the question type later.
[0075] At the same time, a user profile information table "user_profiles" is created to store user profile information, and its structure is as follows: CREATE TABLE user_profiles(user_id VARCHAR(255)PRIMARY KEY,interest_tagsVARCHAR(255),behavior_patterns TEXT); Among them, the user identification primary key field "user_id" is used as the primary key and is associated with the user identification field "user_id" in the "user_questions" table to ensure the correspondence between the user profile and the user questions; the interest tag field "interest_tags" is used to store the interest tags of the user; the behavior pattern field "behavior_patterns" is of text type, to describe the behavior patterns of the user in detail.
[0076] In addition, using the index optimization technology of the database, a composite index is established for the user identification field "user_id" and the question time field "question_time" of the user question table "user_questions", and a separate index is established for the question type field "question_type" to improve data storage and query efficiency.
[0077] The field names of the user question table and the user portrait information table are shown in Table 1.
[0078] Table 1: Field Names of User Question Table and User Portrait Information Table
[0079] Step S113: Configure the database connection;
[0080] In the step S113, in the Python project, the pymysql library is used to implement the connection between the input and output of the Python project and the MySQL database.
[0081] The sample code is as follows: import pymysql # Connect to the database conn = pymysql.connect(host='localhost', user='root', password='your_password', database='intelligent_interaction_system', charset='utf8mb4') In the above code, "your_password" needs to be replaced with the actual database password set to ensure a successful connection.
[0082] Step S120 (optional): Data initialization and testing to verify whether the functions of the database are normal.
[0083] Step S120 is specifically divided into the following sub-steps:
[0084] Step S121: Insert test data into the user question table;
[0085] In the step S121, some test data needs to be inserted into the user question table "user_questions" to verify whether the functions of the database are normal.
[0086] The sample code is as follows: cursor = conn.cursor() sql = "INSERT INTO user_questions(user_id,question_text,question_time) VALUES (%s,%s,NOW())" data = ('user1','How to improve Python programming skills?') cursor.execute(sql,data) conn.commit(). The above code first obtains the database cursor, then defines the SQL statement for inserting data and the data values to be inserted, and finally executes the insert operation and commits the transaction to ensure that the data is successfully written to the database.
[0087] Step S122: Test the storage and query functions of the database;
[0088] In step S122, the storage and query functions of the database are tested, and the questions of "user1" are queried to verify whether the data is correctly stored and can be accurately queried.
[0089] The sample code is as follows: sql = "SELECT * FROM user_questions WHERE user_id = %s" cursor.execute(sql,('user1',)) result = cursor.fetchall() for row in result: print(row). This code obtains and prints the question records of "user1" by executing the query statement to verify whether the data is correctly stored and can be accurately queried, ensuring the normal operation of the storage and query functions of the database and providing guarantee for subsequent data operations of the intelligent personalized interaction system.
[0090] Step S200: Problem analysis and processing preparation, including preparing the training set and test set with question texts, constructing a vocabulary of high-frequency words for them, and performing text encoding;
[0091] Step S200 aims to analyze user questions using a deep learning architecture, providing a solid foundation for subsequent model training and problem analysis.
[0092] Step S200 specifically includes:
[0093] Step S210: Data preprocessing, that is, preprocessing the question texts in the training set and test set (including removing stop words, lemmatization, and stemming);
[0094] Thus, high-quality text data can be provided for subsequent operations such as semantic enhancement and vocabulary construction.
[0095] Step S210 specifically includes:
[0096] Step S211: Perform environment preparation and install natural language processing tools;
[0097] Ensure that a properly functioning Python environment has been installed in the system. Execute the command "pip install nltk" in the command line to complete the installation of the natural language processing tool NLTK. Then, run the code in the Python environment: "import nltk; nltk.download('punkt'); nltk.download('stopwords'); nltk.download('wordnet')" to download the relevant corpora and data required by NLTK.
[0098] Step S212: Prepare the training set and the test set, and read the question texts from the training set and the test set;
[0099] Among them, the training set and the test set are generated by a natural language generation model. Among them, first select a natural language generation model based on the Transformer architecture and collect texts such as children's storybooks for pre-training; then generate questions according to themes and requirements, such as generating simple questions like "What does a puppy like to eat" under the theme of "animals", and follow grammar and children's cognitive characteristics during generation; finally, screen and integrate, exclude complex and ambiguous questions, and form the training set and the test set, so as to provide data support for the system.
[0100] Reading the question texts from the training set and the test set specifically includes: Assuming that the data of the training set and the test set are stored in a text file named "questions.txt", each line represents a question, and use the code "with open('questions.txt', 'r') as file: questions = file.readlines()" to complete the data reading.
[0101] Step S213: Use natural language processing tools (such as NLTK) to remove stop words from the question texts to reduce data noise; Stop words such as "de", "shi", "zai", etc., which contribute less to semantic understanding, can reduce data noise.
[0102] First, import stopwords from the NLTK corpus. The code is "from nltk.corpus import stopwords; stop_words = set(stopwords.words('english'))". Then, process the read question text data. The code is "filtered_questions = []; for question in questions: words = question.lower().split(); filtered_words = [word for word in words if word not in stop_words]; filtered_questions.append('.join(filtered_words))". By this operation, stopwords that contribute less to semantic understanding in the text are removed to reduce data noise.
[0103] Step S214: Perform lemmatization and stemming.
[0104] In the step S214, use WordNet for lemmatization. For example, "running" is lemmatized to "run" to unify the word form; use Snowball Stemmer for stemming to extract the stem of the word. For example, "nationality" is extracted as "nation" to simplify the text data and make the model focus more on key information.
[0105] Specifically, import the relevant modules of WordNet and Snowball Stemmer. The code is "from nltk.stem import WordNetLemmatizer, SnowballStemmer; lemmatizer = WordNetLemmatizer(); stemmer = SnowballStemmer('english')". Then, perform lemmatization and stemming operations on the question text after removing stopwords. The code is "preprocessed_questions = []; for question in filtered_questions: words = question.split(); lemmatized_words = [lemmatizer.lemmatize(word) for word in words]; stemmed_words = [stemmer.stem(word) for word in lemmatized_words]; preprocessed_questions.append('.join(stemmed_words))', unifying the vocabulary form to make the model more focused on key information.
[0106] Step S220: Semantic enhancement, that is, using a pre-trained language model to perform semantic enhancement processing on the question text, mining the potential semantic information of the question text, and improving the semantic expression ability of the text.
[0107] For example, for the question "What are the latest products of Apple", the model can recognize that "Apple" may refer to the technology company Apple rather than the fruit, thus enriching the semantic expression of the question and providing more comprehensive information for subsequent analysis. Therefore, by using a pre-trained language model (such as GPT-2, BERT, etc.) to perform semantic enhancement processing on the question text, the pre-trained language model, relying on the knowledge learned from a large-scale corpus, mines the semantic associations and potential semantic information between words, thereby enriching the semantic expression of the question and providing more comprehensive information for subsequent analysis.
[0108] In the step S220, first install the relevant library of the pre-trained language model, and execute the command "pip install transformers" in the command line. Taking the use of the BERT model as an example for semantic enhancement, the code is "from transformers import BertTokenizer, BertModel; import torch; tokenizer = BertTokenizer.from_pretrained('bert-base-uncased'); model = BertModel.from_pretrained('bert-base-uncased'); enhanced_questions = []; for question in preprocessed_questions: inputs = tokenizer(question, return_tensors = 'pt'); outputs = model(**inputs); enhanced_representation = outputs.last_hidden_state.mean(dim=1).detach().numpy(); enhanced_questions.append(enhanced_representation)".
[0109] Step S230: Vocabulary building and text encoding; that is, build a vocabulary based on the high-frequency words in the question text, and then encode the question text according to the vocabulary to obtain the encoded training samples and test samples.
[0110] In the step S230, first use a word segmentation tool to segment the preprocessed text, count the word frequencies, filter out the words with word frequency ≥ 3 and pass the chi-square test (set the p-value threshold to 0.05), and add them to the vocabulary as high-frequency words to build the vocabulary.
[0111] In this embodiment, specifically use the built-in functions of Python or third-party libraries (such as jieba) to segment the preprocessed text. The code is "import jieba; word_counter = {}; for question in preprocessed_questions: words = jieba.cut(question); for word in words: if word in word_counter: word_counter[word] += 1; else: word_counter[word] = 1". The code for filtering out the words with word frequency ≥ 3 is "high_frequency_words = [word for word, count in word_counter.items() if count >= 3]". The chi-square test is used to determine whether a word is closely related to the question type, that is, whether the frequencies of the word in different question types are different (in practical applications, it is necessary to calculate and filter according to the standard method of the chi-square test).
[0112] In step S230, the text encoding adopts the One-Hot encoding method to convert the text into a digital vector, enabling the model to process and learn text information.
[0113] Step S240 (optional): Sample processing; that is, after completing the text encoding, perform standardization processing on the encoded training samples and test samples to avoid affecting the model training effect due to data scale differences, improve the stability and convergence speed of the model; at the same time, use data augmentation techniques to generate new training samples to increase the diversity of training data and enhance the generalization ability of the model.
[0114] The sample standardization process specifically includes normalizing the sample data to a specific interval by calculating the mean and standard deviation of the samples, avoiding the impact of data scale differences on the model training effect, and improving the stability and convergence speed of the model; data augmentation includes transforming the problem texts in the training set by using methods such as randomly replacing, inserting, and deleting words to generate new training samples, increasing the diversity of the training data, enabling the model to learn more different expression forms and semantic information, and enhancing the generalization ability of the model when facing unseen data.
[0115] Step S300: Construct a workflow scheduler for statistically analyzing high-frequency problem types, thereby providing data support for system decision-making.
[0116] Step S300 specifically includes:
[0117] Step S310: Set the startup conditions of the workflow scheduler to ensure that the workflow can be accurately started according to different situations; the startup conditions of the workflow scheduler include starting at a predetermined time interval and / or starting based on specific trigger conditions. The specific trigger conditions include that the number of questions asked by the user within a fixed time period (adjustable from 10 minutes to 2 hours) reaches a preset threshold (adjustable from 3 to 20 questions), or the user actively requests personalized recommendations to adapt to different application scenarios and user groups.
[0118] Step S310 is specifically subdivided into the following sub-steps:
[0119] Step S311: Environment preparation and library installation to facilitate the setting of the startup conditions of the workflow scheduler.
[0120] In the step S311, it is necessary to ensure that a properly functioning Python environment has been installed in the system. Then, execute the command "pip install apscheduler" in the command line to install the APScheduler library, which is used to implement the timing tasks and trigger condition settings of the workflow.
[0121] Step S312: Set the predetermined time interval in the startup conditions. In the said step S312, the following example code is used to start the workflow at a predetermined time interval. The predetermined time interval is 6 - 24 hours. Taking 12 hours as an example of the predetermined time interval, the code is as follows: from apscheduler.schedulers.background import BackgroundScheduler def workflow_function(): # Here is the main logic of the workflow, which will be introduced in detail later print("Workflow started, performing statistics and analysis tasks") scheduler = BackgroundScheduler() # Start the workflow every 12 hours scheduler.add_job(workflow_function, 'interval', hours = 12) scheduler.start()
[0122] Step S313: Set specific trigger conditions in the startup conditions; the specific trigger conditions include that the number of questions asked by the user within a fixed time period (adjustable from 10 minutes to 2 hours) reaches a preset threshold (adjustable from 3 to 20 questions), or the user actively requests personalized recommendations to adapt to different application scenarios and user groups.
[0123] Thus, the fixed time period in the specific trigger conditions can be adjusted between 10 minutes and 2 hours, and the preset threshold can be adjusted between 3 and 20 questions. Among them, according to the characteristics of different application scenarios and user groups, the specific time period and preset threshold are flexibly set. For example, in an active online community, the specific time period can be set shorter and the preset threshold can be increased; on a relatively low-frequency professional platform, the specific time period can be extended and the preset threshold can be decreased to adapt to different application scenarios and user groups. For example, in an active online community, the specific time period can be set to 30 minutes and the preset threshold can be set to 8 questions; on a relatively low-frequency professional platform, the specific time period can be extended to 1 hour and the preset threshold can be decreased to 5 questions.
[0124] In the said step S313, first define variables to record the number of user questions and time. The example code is as follows: user_question_count = {} Then, use the following function to check whether the number of user questions reaches the preset threshold: Then, use the following function to check whether the number of user questions reaches the preset threshold: import time def check_question_threshold(user_id, question): if user_id not in user_question_count: user_question_count[user_id] = {'count': 1,'start_time': time.time()} else: user_question_count[user_id]['count'] += 1 current_time = time.time() if current_time - user_question_count[user_id]['start_time'] > 3600: # within 1 hour if user_question_count[user_id]['count'] >= 10: # reach the threshold of 10 questions workflow_function() user_question_count[user_id] = {'count': 1,'start_time': current_time} At the entrance where the system receives user questions, call the check_question_threshold function, passing in the user ID and the question content.
[0125] Step S320: High-frequency question type statistics; that is, set the workflow scheduler to use a sliding window algorithm to count high-frequency question types.
[0126] Figure 2 It is a schematic diagram of the sliding window algorithm of the workflow scheduler. When setting the workflow scheduler to use a sliding window algorithm to count high-frequency question types, by setting a time window of a fixed size and sliding it forward with a fixed step size, real-time statistics and updates are performed on the question types within the time window. For example, set the time window to 1 hour and the step size to 15 minutes, and update the statistics of various question types within the past 1 hour every 15 minutes, so as to timely capture the dynamic changes of user question types and provide the latest data support for system decision-making.
[0127] Step S320 is specifically divided into the following sub-steps:
[0128] Step S321: Database connection and data acquisition.
[0129] In the step S321, import the database connection library (such as pymysql) and establish a connection to the database. The example code is as follows: import pymysql conn = pymysql.connect(host='localhost', user='root', password='your_password', database='intelligent_interaction_system') Then, write a function to obtain question data from the database. The example code is as follows: def get_questions_from_database(): cursor = conn.cursor() sql = "SELECT question_type FROM user_questions" cursor.execute(sql) result = cursor.fetchall() questions = [row[0] for row in result] cursor.close() return questions
[0130] Step S322: Implement the sliding window algorithm.
[0131] In the step S322, use the following example code to use the sliding window algorithm to count the high-frequency question types: from collections import Counter def workflow_function(): questions = get_questions_from_database() window_size = 100 # The window size is 100 questions step = 10 # The sliding step is 10 questions for i in range(0, len(questions), step): window = questions[i:i + window_size] counter = Counter(window) top_types = counter.most_common(5) # Get the top 5 high-frequency question types print(f"High-frequency question types in the current window: {top_types}").
[0132] Step S400: Establish a user profile construction module. The user profile construction module includes a user interest analysis module and a user profile generation function based on the Transformer analysis model with an attention mechanism as the basic model, which is used to obtain a user profile containing multi-dimensional user information according to the historical conversations and user questions in the database, laying a foundation for the system to provide personalized services;
[0133] As described above, through user interest analysis and the generation mechanism based on the Transformer model,
[0134] Step S400 specifically includes:
[0135] Step S410: Establish a user interest analysis module. Based on the user's input questions and historical conversations, using sentiment analysis technology in natural language processing, obtain keywords such as the user's sentiment tendency towards different questions and the user's interest fields; and use the high-frequency question types from the workflow scheduler as keywords for question frequency;
[0136] For example, when the user asks "I particularly like artificial intelligence. What are its future development directions?", through sentiment analysis, it can be judged that the user has a positive attitude towards artificial intelligence, and then further explore their specific interest points in this field.
[0137] Step S410 is specifically broken down into the following sub-steps:
[0138] Step S411: Environment preparation and library installation: In step S411, it is necessary to ensure that the system has a properly functioning Python environment. Execute the commands "pip install textblob" and "pip install nltk" in the command line to install natural language processing-related libraries (such as TextBlob and NLTK). Then, run the code "import nltk; nltk.download('punkt'); nltk.download('averaged_perceptron_tagger')" in the Python environment to download the relevant corpora and data required by TextBlob and nltk.
[0139] Step S412: Define a function to implement sentiment analysis and extraction of the user's interest fields.
[0140] In the step S412, a function is defined to analyze and obtain the user's sentiment tendency and user interest fields. The example code is as follows: from textblob import TextBlob def analyze_user_interest(user_input, history_dialogue): all_text = "".join([user_input] + history_dialogue) blob = TextBlob(all_text) sentiment = blob.sentiment.polarity if sentiment > 0: interest_sentiment = "positive" elif sentiment < 0: interest_sentiment = "negative" else: interest_sentiment = "neutral" # Further analyze information such as the fields the user is interested in (this is a simple example and can be extended according to actual needs) words = all_text.lower().split() domain_keywords = ["artificial intelligence", "machine learning", "data science", "programming", "history", "science", "technology"] interested_domains = [domain for domain in domain_keywords if any(keyword in words for keyword in domain.split())] return interest_sentiment, interested_domains In the module where the system receives the user's question and obtains the historical dialogue, the analyze_user_interest function is called, passing in the user's current question and the historical dialogue record, obtaining the sentiment tendency and the user interest fields, and storing them in the corresponding data structure for use when generating the user profile later.
[0141] Step S420: Establish a function for generating the user profile. The function for generating the user profile is set to generate a prompt template based on the user's question, historical dialogue, and keywords, and call the Transformer analysis model based on the attention mechanism to generate the user profile. The user profile includes the user's interest focus, behavior pattern, and preferred question types.
[0142] Step S420 is specifically divided into the following sub-steps:
[0143] Step S421: Select the Transformer analysis model based on the attention mechanism as the basic model for generating the user profile, and install the relevant libraries;
[0144] Among them, with its powerful sequence modeling ability and attention mechanism, the Transformer model can effectively capture the key information and semantic relationships in the text. The attention mechanism enables the model to pay more attention to the important parts when processing the user input questions and historical conversations, improving the ability to understand the user's intentions and laying a foundation for generating an accurate user profile;
[0145] In the step S421, execute the command "pip install transformers" in the command line to install the library for loading the Transformer analysis model based on the attention mechanism. Load the pre-trained Transformer model based on the attention mechanism (such as variants of the BERT or GPT series) from the transformers library. The example code is as follows: from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained('bert-base-uncased') model = AutoModelForSequenceClassification.from_pretrained('bert-base-uncased')
[0146] Step S422: Construct a user profile generation prompt template for integrating keywords such as user interest fields, question frequencies, and sentiment tendencies into the Transformer analysis model based on the attention mechanism in the form of prompt words;
[0147] Specifically, this template contains key prompt words for guiding the model to generate the user profile, such as relevant prompts for user interest fields, question frequencies, sentiment tendencies, etc. Design the prompt template carefully to provide clear guidance for the model to generate the user profile.
[0148] For example, the example code of the user profile generation prompt template is as follows: prompt_template = "User interest field: {}\nQuestion frequency: {}\nSentiment tendency: {}\nUser input question: {}\nHistorical conversation: {}"
[0149] Step S423: Define an analysis chain, which is set to pass the user's question, historical dialogue, keywords, and the generated prompt template of the user profile to the Transformer analysis model to generate the user profile.
[0150] In the step S423, the code example is as follows: def generate_user_profile(user_input, history_dialogue, interest_sentiment, interested_domains): input_text = prompt_template.format(",".join(interested_domains), len(history_dialogue), interest_sentiment, user_input, "\n".join(history_dialogue)) inputs = tokenizer(input_text, return_tensors = 'pt') outputs = model(**inputs) # Parse and generate the user profile based on the model output (this is a simple example, and more complex parsing logic may be required in reality) user_profile = {"interest_focus": interested_domains, "behavior_pattern": len(history_dialogue), "preferred_question_type": [] # Can be further determined according to the model output} return user_profile. That is to say, in the link where the system needs to generate the user profile, call the generate_user_profile function, pass in the user's current question, historical dialogue, and keywords, then obtain the generated user profile, and store the user profile in the database or memory for subsequent interaction with other modules.
[0151] Step S500: Integrate to obtain an intelligent personalized interaction system, which includes a user question input interface, a question classification module 10, and a database 20 connected in sequence. The database is also connected to a workflow scheduler 30 and a user profile construction module 50; so that the user profile of the user profile construction module 50 and the high-frequency question types of the workflow scheduler 30 are both output to a pre-trained large language model 60 to generate recommended questions that fit the user's interests. In addition, ensure the stable and efficient operation of the system through comprehensive testing and debugging.
[0152] Among them, the user question input interface is used to receive user questions. The question classification module 10 is used to obtain the question type according to the user question and store it. The database is used to store user questions and question types. The workflow scheduler counts the high-frequency question types when starting. The user portrait construction module is used to obtain the user portrait according to the historical conversations, user questions and high-frequency question types in the database, and retrieve relevant historical questions and answers from the database for reference. The user portrait construction module 50 also receives the recommended question feedback data from the user, and uses the recommended question feedback data to continuously optimize the user portrait and the question generation strategy of the pre-trained large language model, and improve the personalized interaction ability of the system.
[0153] Among them, the question classification module 10 of the present invention is unique in terms of structure and training method.
[0154] The question classification module 10 adopts a fusion neural network architecture, that is, an architecture that combines a convolutional neural network (CNN) and a recurrent neural network (RNN).
[0155] The CNN part extracts local features of the question text through the combination of convolutional layers and pooling layers. For example, for a question like "How to implement the quicksort algorithm in Python?", the CNN can identify key technical terms such as "Python" and "quicksort algorithm" as local features. The RNN part is responsible for processing the semantic information of the text sequence. It takes the question text data as the current input data into the hidden layer, and combines the high-frequency question type and user portrait information of the previous time step. This enables the question classification module to use context information and user historical features to more accurately judge the question type. For example, in a series of questions about programming, it can better understand the relationship and differences between the current question and the previous questions.
[0156] The question classification module 10 is located after the user question input interface and is closely connected to the database. After receiving the user question, it is processed by the internal CNN and RNN, and the classification result is stored in the database. At the same time, it provides data support for the subsequent workflow scheduler and user portrait construction module, forming an organic data processing and transmission chain.
[0157] The training data of the problem classification module 10 comes from a large number of carefully labeled samples, covering various fields and types of problems, ensuring that the model can learn rich and diverse problem feature patterns. Before the data enters the model training, a series of preprocessing operations described above are carried out. First, stop words such as "de", "shi", "zai" and other words that contribute less to semantic understanding are removed to reduce data noise; then word form reduction and stemming are carried out. For example, "running" is reduced to "run", and "nationality" is extracted to "nation" to make the text data more standardized and concise, facilitating the model to focus on key information. Before encoding, semantic enhancement processing is also carried out. Using a pre-trained language model (such as BERT) to mine the potential semantic information of the problem text and improve the semantic expression ability of the text. At the same time, data augmentation techniques are adopted, such as randomly replacing, inserting, deleting words and other methods to transform the problem text in the training set, increasing the diversity of the training data, so that the model can learn more different expression methods and semantic information.
[0158] Based on the preprocessed and enhanced training data, this invention conducts multiple rounds of iterative training on the architecture that fuses the convolutional neural network (CNN) and the recurrent neural network (RNN) adopted by the problem classification module 10 to train the problem classification module 10. During the training process, CNN and RNN work together. The local features extracted by CNN are used as the input supplement for RNN. After RNN combines the context and historical information for semantic analysis, it outputs the problem classification result. The parameters of CNN and RNN are continuously adjusted through the backpropagation algorithm. The cross-entropy loss function is used to measure the difference between the model prediction result and the true label, and the model parameters are continuously updated with the help of an optimizer (such as stochastic gradient descent, Adagrad, etc.) to minimize the loss function and gradually improve the classification accuracy of the model for problems.
[0159] The differences between the problem classification module 10 and the prior art are as follows:
[0160] Many existing problem classification modules adopt a simple keyword matching algorithm. This method only determines the problem type based on specific keywords that appear in the problem and cannot understand the semantics and context information of the problem. For example, for a problem like "The mobile phone I bought last month has been crashing frequently recently, and it gets very hot when charging. How can this be solved?", traditional modules may only extract keywords such as "mobile phone", "crashing", "charging and getting hot", while ignoring important context information such as "bought last month", and it is difficult to provide accurate classification and solutions. However, the module of this invention can deeply analyze the semantic and logical relationships of the text through the fusion architecture of CNN and RNN, and comprehensively consider various factors for accurate classification.
[0161] In addition, some existing problem classification modules adopt a single neural network structure, such as only using CNN or RNN. A single CNN structure may be insufficient in processing semantic information of long text sequences and difficult to capture the overall logic of the text; while a single RNN structure may not be efficient enough in extracting local features. The fusion architecture of the problem classification module of the present invention combines the advantages of both. CNN efficiently extracts local features to provide more accurate input for RNN, and RNN further optimizes the classification results by using context and historical information, thus having stronger capabilities in dealing with complex problems and adapting to different user question styles, being able to classify user questions more accurately, and providing more reliable basic support for the entire intelligent personalized interaction system.
[0162] The large language model will comprehensively consider factors such as high-frequency problem types, the user's interest fields, question preferences, and sentiment tendencies to generate targeted and personalized questions. For example, if the user is a programmer and the current high-frequency problem type is "Python programming optimization", and the user profile shows that the user is more interested in data processing and machine learning, the large language model may generate a question like "How to use Python for efficient data cleaning and feature engineering to optimize machine learning models".
[0163] After generating the questions, the intelligent personalized interaction system retrieves relevant historical questions and answers from the database for the user to refer to. By matching the keywords, types, and user profile information of the recommended questions, the intelligent personalized interaction system can quickly locate relevant historical questions and high-quality answers, providing the user with more comprehensive information and solutions. At the same time, the system presents the generated questions and relevant historical information to the user, interacts with the user, and guides the user to further explore the questions in depth.
[0164] During the interaction with the user, the intelligent personalized interaction system collects feedback data on the recommended questions, including information such as whether the user clicks to view the recommended questions, whether to ask further questions, the content and direction of the questions asked, and the evaluation and feedback on the recommended questions. These feedback data on the recommended questions are important bases for system optimization and can reflect the user's satisfaction and needs with the system's recommended questions.
[0165] Using the user's recommended question feedback data, continuously optimize the user profile and the question generation strategy of the pre-trained large language model. By continuously adjusting the interest tag weights, behavior pattern parameters, and other information in the user profile, make the user profile more accurately reflect the user's true interests and needs. At the same time, adjust the input parameters, training data, and generation algorithms of the large language model according to user feedback to improve the quality and pertinence of question generation. For example, if the user shows strong interest in a recommended question, asks further questions or delves deeper into related topics, the system strengthens the interest tag in this field in the user profile and provides the user with more relevant high-quality content and question recommendations; if the user repeatedly ignores the recommended questions about a certain topic, the system reduces the weight of the interest tag of this topic in the user profile and adjusts the frequency and content of the large language model to generate related questions. In addition, the statistical logic and trigger conditions of the workflow scheduler can be optimized according to user feedback, so that the system can capture the changes in user needs more timely and accurately and provide better personalized interaction services for users. Through continuous optimization, the system can continuously improve its intelligence and adaptability and better meet the diverse needs of users in different scenarios.
[0166] Step S500 specifically includes:
[0167] Step S510: Integrate to obtain an intelligent personalized interaction system and achieve data circulation and collaborative work among modules.
[0168] In step S510, when integrating to obtain an intelligent personalized interaction system, define a unified data transmission format (such as JSON format) and interface specification, implement the corresponding interface functions in each module, and use a message queue system (such as RabbitMQ or Kafka) to build a data circulation pipeline to achieve asynchronous data transmission between modules; therefore, step S510 is specifically divided into the following sub-steps:
[0169] Step S511: Design and implementation of interfaces between modules; that is, define a unified data transmission format (such as JSON format) and interface specification, and implement the corresponding interface functions in each module;
[0170] In the step S511, a unified data transmission format and interface specification are defined. For the data transmitted from the problem classification module to the workflow scheduler and the user profile construction module, the standardized JSON format is adopted. For example, the sample JSON data structure of the problem analysis result is as follows: {"question_id":"12345","category":"Technical - Programming","keywords":["Python","Function call","Error handling"]} When the workflow scheduler provides high - frequency problem type information to the large language model, the JSON format is also used, such as: {"high_frequency_types":["Programming problems","Data analysis problems"],"time_window":"The last 1 hour"} When the user profile construction module feeds back user profile information to the large language model, the format can be: {"user_id":"user001","interest_focus":["Artificial intelligence","Machine learning"],"behavior_pattern":{"average_question_frequency":3,"preferred_time_slot":"8 pm - 10 pm"}} Corresponding interface functions are implemented in each module for sending and receiving this data. For example, in the problem classification module, there is a function responsible for sending the analysis result to a specified queue or buffer, and the workflow scheduler and the user profile construction module have functions to read data from this queue or buffer.
[0171] Step S512: Build a data circulation pipeline; that is, use a message queue system (such as RabbitMQ or Kafka) to build a data circulation pipeline to achieve asynchronous data transmission between modules;
[0172] In the step S512, a message queue system (such as RabbitMQ or Kafka) is used to achieve asynchronous data transmission between modules, ensuring reliable data transfer and system decoupling. Taking RabbitMQ as an example, first install and configure the RabbitMQ server. In the problem classification module, when a new problem analysis result is generated, use the producer client of RabbitMQ to send the data to the corresponding queue. The sample code is as follows: import pika import json connection = pika.BlockingConnection(pika.ConnectionParameters('localhost')) channel = connection.channel() channel.queue_declare(queue='analysis_result_queue') analysis_result = {"question_id":"12345","category":"Technical - Programming","keywords":["Python","Function call","Error handling"]} channel.basic_publish(exchange = "",routing_key = 'analysis_result_queue',body = json.dumps(analysis_result)) connection.close() The workflow scheduler and the user profile construction module act as consumers and read data from the queue. The sample code is as follows: import pika import json def callback(ch,method,properties,body): analysis_result = json.loads(body) # Process the analysis result, update the internal state or perform further calculations channel.queue_declare(queue='analysis_result_queue') channel.basic_consume(queue='analysis_result_queue',on_message_callback = callback,auto_ack = True) channel.start_consuming() Similarly, establish data flow pipelines between the workflow scheduler and the large language model, and between the user profile construction module and the large language model.
[0173] Step S520: System testing and debugging to ensure that the functions of each module of the system are correct and the overall performance is stable.
[0174] System testing and debugging include unit testing, integration testing, and debugging and optimization. Unit tests write test cases for each module's function. Integration testing simulates real user usage scenarios to verify the overall system function and test high-concurrency performance. Debugging and optimization use debugging tools to debug problem modules and collect and analyze log information for improvement.
[0175] Therefore, step S520 is specifically broken down into the following sub-steps:
[0176] Step S521: Unit testing;
[0177] In step S521, unit test cases are written for each module to ensure the functional correctness of each module. For the problem classification module, test datasets including different types of problem texts are used to verify the accuracy of its classification and keyword extraction. For example, using a test framework (such as pytest), the following test function is written: import pytest from problem_analysis_module import analyze_problem @pytest.mark.parametrize("question_text,expected_category,expected_keywords", [("How to implement quicksort algorithm in Python?","Technical - Programming",["Python","quicksort algorithm"]),("What are the latest products of Apple?","Technology - Products",["Apple","products"]),]) def test_analyze_problem(question_text,expected_category,expected_keywords): result = analyze_problem(question_text) assert result["category"] == expected_category assert set(result["keywords"]) == set(expected_keywords) For the workflow scheduler, test the accuracy of its start conditions, time interval settings, and high-frequency problem type statistics. Simulate different time points and problem input situations to verify whether it triggers the workflow as expected and counts high-frequency problem types. For the user profile construction module, test the accuracy of sentiment analysis and user profile generation. Provide different user inputs and historical conversation data to check whether the generated user profile meets expectations.
[0178] Step S522: Integration testing;
[0179] In the step S522, an integration test of the system is carried out, simulating the usage scenarios of real users. Starting from the user inputting questions, track the flow and processing of data among various modules, and verify the overall function of the system. For example, input a series of questions of different types, and check whether the system can correctly analyze the questions, update the workflow statistics, generate appropriate user portraits, and generate relevant answers or recommended questions by the large language model. Test the performance of the system under high concurrency. Use performance testing tools (such as JMeter) to simulate a large number of users asking questions simultaneously, and monitor indicators such as the response time, throughput, and resource utilization rate of the system to ensure that the system can run stably. If performance bottlenecks are found, such as slow database queries or long model inference times, optimize the corresponding modules, such as optimizing database query statements, adjusting model parameters, or adopting distributed computing technologies, etc.
[0180] Step S523: Debugging and optimization;
[0181] In the step S523, during the testing process, use debugging tools (such as the pdb module of Python or the debugging function in the integrated development environment) to debug the modules with problems. For example, if it is found that the user portrait generation is inaccurate, breakpoints can be set in the corresponding code segment, and the input data, model processing process, and output results can be checked step by step to find out the problem and fix it. Continuously collect the log information during the operation of the system, including data transfer logs between modules, model operation logs, and system error logs, etc. Analyze this log information to discover potential problems and optimization points, such as a certain type of error that frequently occurs or data transfer delays, etc., and make targeted improvements and optimizations to continuously improve the stability and performance of the system.
[0182] On the other hand, the present invention provides an intelligent personalized interaction system, which is constructed by adopting the construction method of the intelligent personalized interaction system described above. The obtained intelligent personalized interaction system integrates multiple modules such as a database, a question classification module, a workflow scheduler, a user portrait construction module, and a pre-trained large language model. Through the collaborative work of each module, it realizes precise understanding of user questions, personalized question recommendation, and continuously optimized interaction services.
[0183] The present invention provides an operation method of an intelligent personalized interaction system, which includes:
[0184] Step A1: Construct an intelligent personalized interaction system by adopting the construction method of the intelligent personalized interaction system described above;
[0185] In the system deployment phase, the initialization and configuration of the workflow scheduler, user profile module, and pre-trained large language model need to be completed. The workflow scheduler sets startup rules and trigger condition parameters according to the characteristics of the application scenario; the user profile module constructs an initial profile based on multi-dimensional data mining and analysis; the pre-trained large language model selects an appropriate architecture and fine-tunes for the task, while ensuring that the system hardware environment and network connection meet the performance requirements.
[0186] Step A2: The user inputs a user question and gets a recommended question.
[0187] Thus, the intelligent personalized interaction system uses the workflow scheduler to determine the high-frequency question types, combines the user profile, and drives the pre-trained large language model to generate questions that fit the user's interests. When the user inputs a question, the workflow scheduler will quickly count the current high-frequency question types according to the set rules and time window. At the same time, the user profile module will analyze information such as the user's interest focus and question preference based on the user's historical questions and interaction records. The large language model takes the high-frequency question types and user profile information as inputs, and through complex calculations and inferences, generates questions that highly match the user's interests. For example, if the user is a fitness enthusiast and the current high-frequency question type in the fitness field is "muscle gain diet plan", the system may generate a question like "How can strength trainers develop a high-protein diet plan for muscle gain" by combining the preference for strength training in the user profile.
[0188] Step A3: The intelligent personalized interaction system retrieves relevant historical questions and answers from the database based on the generated recommended questions and the user's historical conversations for the user to refer to; collects feedback data on the recommended questions from the user and uses the feedback data from the user on the generated questions to continuously optimize the user profile and the question generation strategy of the pre-trained large language model, and improve the personalized interaction ability of the system.
[0189] Thus, the intelligent personalized interaction system retrieves relevant content in the database based on the generated questions and historical conversations, organizes the retrieval results and provides them to the user as interaction feedback, and collects the user's responses to the feedback for continuous optimization of the system, so as to continuously optimize the user profile, large language model, workflow scheduler, etc. in the future, and improve the personalized interaction ability and service quality of the system.
[0190] The intelligent personalized interaction system will conduct a comprehensive search in the database to find information such as historical questions, answers, and relevant materials related to the generated recommended questions. After organizing and filtering this information, it will be presented to the user in a clear and understandable manner as interactive feedback. At the same time, closely monitor various responses of the user to the feedback, including the user's expressions, language feedback, operation behaviors, etc. For example, whether the user clicks to view the recommended historical question answers, whether they like or comment on the answers, and whether they raise new questions. Based on these feedbacks, the system can deeply understand whether the user's needs are met and which aspects still need improvement. Using this feedback data, update and improve the user profile, further adjust the parameters and algorithms of the large language model, and optimize the statistical and triggering conditions of the workflow scheduler. For example, if the user feedbacks that a recommended fitness diet plan lacks specific food examples, the system can specifically collect more relevant food example materials, update them to the database, and adjust the large language model to provide more detailed content when generating similar question answers. Through continuous feedback collection and optimization, the system can gradually adapt to the unique needs of each user, provide more accurate and personalized interactive services, and improve user satisfaction and loyalty.
[0191] The construction and usage method of the intelligent personalized interaction system of the present invention integrates technologies such as deep learning and natural language processing to achieve accurate understanding of user questions and personalized interaction. The question classification module uses deep learning algorithms to improve the accuracy of question processing. The workflow scheduler captures hot demands, the user profile helps with personalized services, the pre-trained large language model enhances the interaction experience, and it can be continuously optimized using user feedback. The system has a wide range of application scenarios, covering fields such as intelligent customer service and online education, and can provide efficient services for different users, promoting the intelligent development of various fields. For example, in intelligent customer service, it can accurately recommend solutions and products, and in online education, it can recommend learning materials according to the student's situation. In terms of technical implementation, each step closely cooperates. Database design and index optimization provide a data foundation, the question classification module accurately understands questions, the workflow scheduler keeps up with demand changes, the user profile deeply understands users, and the system integration and optimization play the advantages of modules. In the future, the system can combine technologies such as computer vision and speech recognition to achieve multi-modal interaction, expand the data source to improve the accuracy of recommendations, and strengthen security and privacy protection. With the progress of technology and the expansion of scenarios, it will play a greater role and bring more convenience.
[0192] In addition, the intelligent personalized interaction system of the present invention is applicable to children aged 8 - 13, and its creation plan covers multiple aspects such as methods, systems, and application approaches.
[0193] First, a dedicated database is carefully built, and a table structure that conforms to children's cognitive characteristics is designed. Among them, the question text field retains the details of children's questions in a clear and easy-to-understand text form. The question time field accurately tracks the question time through a time stamp accurate to the second level. The user identification field uses a simple and clear string type to effectively associate the identity of children. The question type field is initially empty and is used to fill in the question type later. At the same time, a table for storing children's user portrait information is created. Through the user identification primary key associated with the question table, as well as the interest tag and behavior pattern fields, the interest preferences and behavior habits of children are comprehensively recorded.
[0194] Next, the question classification module uses advanced deep learning algorithms (such as the fusion architecture of convolutional neural network and recurrent neural network) that are adapted to children's language expressions to deeply analyze and accurately classify children's questions, and properly store the results in the database. In the analysis process, the simplicity, vividness, and possible non-standardization of children's language are fully considered, and the algorithm model is optimized to better understand children's semantics.
[0195] A workflow scheduler with flexible start conditions is introduced. It can be started at a predetermined time interval (such as 4 - 12 hours, adapting to children's daily usage rhythm) or specific trigger conditions (for example, when the number of questions asked by children reaches 3 - 8 within 15 minutes to 1 hour, or when children actively request personalized recommendations). Efficient logic is written inside the workflow. The sliding window algorithm is used to count the high-frequency question types. A fixed-size time window and step size suitable for the frequency of children's questions are set, and the statistical results are updated in real time to capture the dynamic changes of children's question types in a timely manner, providing the latest data support for system decision-making. Combining with a pre-trained large language model specifically for children's language training, parameters are flexibly set according to children's unique interest characteristics and cognitive levels to generate questions that match children's interests.
[0196] The system of the present invention effectively integrates the functions of each module. When applied, a series of processes are triggered by children's questions. The workflow coordinates the working rhythm and data flow of each link. Based on the feedback data of children on the generated questions, the children's user portrait is continuously improved, thereby intelligently adjusting the question recommendation strategy, effectively improving the insight accuracy of children's interests, creating a personalized interaction experience for children, strongly promoting the development of intelligent interaction technology in personalized service fields such as children's education and entertainment, and helping children learn knowledge and expand their thinking in a relaxed and pleasant interaction, meeting their diverse needs in the growth process.
[0197] The above-mentioned is only the preferred embodiment of the present invention and is not used to limit the scope of the present invention. Various changes can be made to the above embodiments of the present invention. That is, all simple, equivalent changes and modifications made according to the claims and the content of the specification of the present invention application fall within the protection scope of the claims of the present invention patent. What the present invention does not describe in detail is all conventional technical content.
Claims
1. A method for constructing an intelligent personalized interactive system, characterized in that: include: Step S100: constructing a database; Step S200: preparing a training set and a test set with question texts, constructing a vocabulary of high-frequency words for them and performing text encoding; Step S300: constructing a workflow scheduler for counting high-frequency problem types; Step S400: Establish a user portrait building module, which is used to obtain a user portrait based on historical conversations, user questions and high-frequency question types in the database; Step S500: An intelligent personalized interactive system is obtained by integration, which includes a user question input interface, a question classification module, and a database connected in sequence, and the database is also connected to a workflow scheduler and a user portrait construction module; the user portrait and high-frequency question types are output to a pre-trained large language model to generate recommended questions.
2. The method for constructing an intelligent personalized interactive system according to claim 1, characterized in that: The table structure of the database includes a user question table and a user portrait information table; The user question table includes: a question primary key field, a user identification field, a question text field, a question time field, and a question type field; the user portrait information table is used to store user portrait information, including a user identification primary key field associated with the user identification field, as well as an interest tag field and a behavior pattern field; and / or The question classification module adopts a convolutional neural network and a recurrent neural network fusion architecture; and / or The user portrait building module also receives recommended question feedback data from users, and uses the recommended question feedback data to continuously optimize the user portrait and the question generation strategy of the pre-trained large language model; and / or The database and the workflow scheduler are connected to a question generation module, and the question generation module feeds back recommended questions to the user based on the high-frequency question types and the data in the database.
3. The method for constructing an intelligent personalized interactive system according to claim 1, characterized in that: The step S200 specifically includes: Step S210: preprocessing the question texts in the training set and the test set, including removing stop words, lemma restoration and stemming; Step S230: construct a vocabulary based on high-frequency words in the question text, and then perform text encoding on the question text according to the vocabulary to obtain encoded training samples and test samples.
4. The method for constructing an intelligent personalized interactive system according to claim 1, characterized in that: After step S210 and before step S230, the method further includes step S220: using a pre-trained language model to perform semantic enhancement processing on the question text; After step S230, the method further includes step S240: performing standardization processing on the encoded training samples and test samples; and at the same time, using data enhancement technology to generate new training samples.
5. The method for constructing an intelligent personalized interactive system according to claim 1, characterized in that: The step S300 specifically includes: Step S310: Setting the start conditions of the workflow scheduler, the start conditions of the workflow scheduler include starting at a predetermined time interval and / or starting according to a specific trigger condition; the specific trigger condition includes the number of questions asked by the user within a fixed time period reaching a preset threshold, or the user actively requests personalized recommendation; Step S320: The workflow scheduler is set to use a sliding window algorithm to count high-frequency problem types.
6. The method for constructing an intelligent personalized interactive system according to claim 5, characterized in that: The fixed time period in the specific trigger condition is adjustable between 10 minutes and 2 hours, and the preset threshold is adjustable between 3 and 20 questions; When using the sliding window algorithm, set a fixed time window and step size.
7. The method for constructing an intelligent personalized interactive system according to claim 1, characterized in that: The step S400 specifically includes: Step S410: Establish a user interest analysis module, which uses sentiment analysis technology in natural language processing to obtain keywords such as user sentiment tendencies towards different questions and user interest areas based on user input questions and historical conversations; and uses high-frequency question types from the workflow scheduler as keywords for question frequency; Step S420: Establish a function for generating a user portrait. The function for generating a user portrait is set to generate a prompt template based on user questions, historical conversations, keywords, and the user portrait, and call the Transformer analysis model based on the attention mechanism to generate a user portrait. The user portrait includes the user's interest focus, behavior pattern, and preferred question type.
8. The method for constructing an intelligent personalized interactive system according to claim 7, characterized in that: The step S420 specifically includes: Step S421: Selecting a Transformer analysis model based on an attention mechanism as a basic model for generating a user portrait function; Step S422: construct a user portrait generation prompt template, which is used to integrate keywords such as user interest areas, question frequency, and emotional tendency into the Transformer analysis model based on the attention mechanism in the form of prompt words; Step S423: define an analysis chain, which is set to pass user questions, historical conversations, keywords, and user portrait generation prompt templates to the Transformer analysis model to generate a user portrait.
9. The method for constructing an intelligent personalized interactive system according to claim 1, characterized in that: The step 500 specifically includes: integrating to obtain an intelligent personalized interactive system; performing system testing and debugging, including unit testing, integration testing, and debugging and optimization. Unit testing writes test cases for the functions of each module. Integration testing simulates real user usage scenarios to verify the overall system functions and test high concurrency performance. Debugging and optimization use debugging tools to debug problem modules, and collect and analyze log information for improvement.
10. An intelligent personalized interactive system, characterized in that: It is established by using the construction method of the intelligent personalized interaction system according to one of claims 1-9.
Citation Information
Cited By
Message template determination method and device, message pushing method and device, equipment and medium
CN120407949A
Message template determination method, apparatus, device, and medium
CN120407949B
Method and system for generating official business management user portrait based on Agent and MCP
CN120873000A