Generative smart campus engine system based on full-text retrieval knowledge enhancement

Through the generative smart campus engine system enhanced by full-text search knowledge, the problem of difficulty in identifying data by users is solved, personalized data screening and multimodal interaction are realized, and the accuracy and convenience of campus information services are improved.

CN120336516APending Publication Date: 2025-07-18ANHUI PATTERN RECOGNITION INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510804132.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When processing user request data, the existing smart campus engine cannot effectively identify a variety of different types of complex and disordered content, resulting in low accuracy of search results and inability to meet personalized needs.

Method used

Through the data acquisition module, database construction module, data screening module, image construction module and data combination module, the full-text search knowledge enhancement is achieved, a knowledge database with notes semantics, analyze user request semantics, build user personality portraits, and conduct multimodal interactions.

Benefits of technology

It realizes accurate analysis of user requested information and personalized data screening, improves data dimensions and interaction effects, and improves the level of campus information service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336516A_ABST
    Figure CN120336516A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data retrieval, and discloses a generation type smart campus engine system based on full-text retrieval knowledge enhancement. Comprising a data acquisition module for acquiring system data of a campus subsystem, a database construction module for constructing a knowledge database, a data screening module for screening out target data from the knowledge database, a portrait construction module for constructing a user personalized portrait, and a data combination module for combining interest data and the target data into engine enhanced data; compared with the prior art, the method has the advantages that complicated and disordered user request data can be simply and accurately represented through the smart campus engine, and a basis can be provided for screening of interest data corresponding to recent behavior modes and behavior preferences of the user side in combination with a mode of constructing the personalized portrait of the user; the combined analysis effect of the real-time request information and the personalized request information of the user side is achieved, and the data dimension of the user request information of the user side is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and more specifically, to a generative intelligent campus engine system based on full-text retrieval knowledge enhancement. Background Art

[0002] With the development of the construction of intelligent campuses, a vast amount of data resources have been accumulated in intelligent campus systems. During the use of intelligent campus systems, it is necessary to perform full-text retrieval processing on the request data of users through natural language technology, and use intelligent models to accurately retrieve the required data from the vast amount of stored data. Therefore, it is necessary to build an intelligent campus engine to meet the retrieval and query needs of users for diverse demand data and information.

[0003] The patent application with the publication number CN115879001A discloses a management method and system for an intelligent campus multimedia comprehensive information service terminal, including an information storage subsystem, an information retrieval subsystem, and an information management subsystem. The information storage subsystem, the information retrieval subsystem, and the information management subsystem are connected in sequence. The information storage subsystem is used to build a database to store text resources, and after extracting the text in the video resources, it is packaged and stored in the database together with the video resources. The information retrieval subsystem retrieves target resources from the database based on resource requests. The information management subsystem is used to perform management operations on the target resources. By extracting the text in the video resources as the retrieval label of the video resources, the video resources can be retrieved by means of text retrieval, solving the problem of low extraction efficiency of target resources. When the existing intelligent campus engine retrieves and queries the request data of users, it collects the real-time request data of users, and after performing correlation analysis on the request data and the stored original data, it filters out the target data associated with the request data. However, when the request data of users contains a variety of different types of complex and disordered content, it is impossible to comprehensively and orderly identify the real content in the request data, resulting in a low degree of matching accuracy between the subsequent retrieved target data and the request data, and it is impossible to perform targeted data retrieval and interaction experience for the personalized requests of different users, resulting in the filtered target data being unable to truly, accurately, and comprehensively meet the personalized needs of users.

[0004] In view of this, the present invention proposes a generative intelligent campus engine system based on full-text retrieval knowledge enhancement to solve the above problems. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A generative intelligent campus engine system based on full-text retrieval knowledge enhancement, which is applied to an intelligent campus platform, includes: The data collection module is used to collect system data of the campus subsystem, determine the update cycle from the system data, and update the system data regularly; The database construction module is used to process the validity of the system data in the data processing mode and construct the system data into a knowledge database with annotation semantics based on the database construction criteria; The data screening module is used to receive user request data from the user end, analyze the request semantics from the user request data through the trained smart campus engine, and match and analyze the request semantics and remark semantics to screen out the target data from the knowledge database; The portrait construction module is used to obtain the user's personality behavior data during the portrait cycle, extract the portrait features from the personality behavior data, and construct a user personality portrait with portrait semantics; The data combination module is used to filter out interest data corresponding to the user's personality portrait from the knowledge database, combine the interest data and target data into engine enhanced data, and control the free switching of interaction modes.

[0006] Furthermore, the campus subsystem includes the academic affairs management subsystem, the library management subsystem, the student information management subsystem, the office subsystem and the environment subsystem; System data includes academic data, borrowing data, student data, office data and environmental data.

[0007] Furthermore, the steps for determining the update cycle are: Record A original data corresponding to the teaching data, borrowing data, student data, office data and environment data as the data to be checked, and query the time when the data to be checked was first recorded one by one through the timestamp to obtain A initial time; Taking A initial moments as the starting point and the current moment as the end point, the data to be checked are compared one by one, and the moment when the data to be checked changes is recorded; In chronological order, the duration between two adjacent moments is recorded as the interval duration, and the minimum value of the interval duration is recorded as the update period.

[0008] Further, the steps of validity processing are: Compare the repeatability of the teaching data, borrowing data, student data, office data and environmental data one by one, record the system data with duplication as duplicate data, and aggregate the same duplicate data to obtain C duplicate sets; Query the recording times of all duplicate data in the C duplicate sets one by one, arrange all the recording times in the C duplicate sets in chronological order, and remove the duplicate data except the last recording time; Perform complete attribute recognition on the remaining educational administration data, borrowing data, student data, office data, and environmental data in C duplicate sets, and record the system data with defective complete attributes as defective data; Based on the calibrated financial complete data, borrowing complete data, student complete data, office complete data, and environmental complete data, automatically fill in the defective data to obtain complete system data.

[0009] Furthermore, the database construction criterion is: one data layer corresponds to one note semantics, and two adjacent data layers are independent of each other; The construction steps of the knowledge database are: Establish a blank database with five vertically distributed data layers, and draw a layer contour in the shape of a circular enclosure on the outside of the five data layers; Count the quantities of the original data corresponding to the educational administration data, borrowing data, student data, office data, and environmental data one by one, and obtain the first quantity value, the second quantity value, the third quantity value, the fourth quantity value, and the fifth quantity value respectively; In the top-down manner, draw data bits with the same quantity as the first quantity value, the second quantity value, the third quantity value, the fourth quantity value, and the fifth quantity value in the five data layers respectively, and import the educational administration data, borrowing data, student data, office data, and environmental data into the data bits of the five data layers one by one to construct an educational administration layer, a borrowing layer, a student layer, an office layer, and an environmental layer; Draw semantic note columns on the layer contours of the educational administration layer, borrowing layer, student layer, office layer, and environmental layer respectively to obtain the first note column, the second note column, the third note column, the fourth note column, and the fifth note column; Sequentially note educational administration, borrowing, students, office, and environment in the first note column, the second note column, the third note column, the fourth note column, and the fifth note column respectively to generate note semantics, and construct a knowledge database with note semantics.

[0010] Furthermore, the steps for training the smart campus engine are: Pre-collect multiple groups of user request data at the user end and the request semantics corresponding to the user request data, mark the user request data as feature vectors, and convert the request semantics into labels corresponding to the feature vectors; One feature vector corresponds to one label, forming a group of training data. Multiple groups of training data form a training set, and the labeled training data is divided into a training set and a test set; Use the feature vectors as the input of the smart campus engine, use the request semantics corresponding to the smart campus engine as the output of the smart campus engine, train the smart campus engine with the training set, test the smart campus engine with the test set, preset an error threshold, and when the mean of the prediction errors of all training data in the test set is less than the preset error threshold, obtain the smart campus engine.

[0011] Further, the steps for screening target data are as follows: Compare each of the D analyzed request semantics with the note semantics in the knowledge database one by one, and mark the data layer where there is an overlap between the note semantics and the request semantics as the target layer; Split the original semantics of the original data in the target layer into a text part and a numerical part through word segmentation technology, and compare the text part with the D request semantics for consistency; Mark the text part that is consistent with the request semantics as the target text, and combine the target text with the corresponding numerical part to form target data, obtaining E target data.

[0012] Further, take the current moment as the end point of the cycle, take the moment when the user request data was first input within the two weeks before the end point of the cycle as the start point of the cycle, and mark the time period between the start point and the end point of the cycle as the portrait cycle.

[0013] Further, the portrait features include interest content, number of views, viewing time, and degree of interest; The steps for constructing a user personality portrait are as follows: Identify the behavioral semantics of the personality behavior data one by one through natural language processing technology, and mark the text part in the behavioral semantics that is consistent with the note semantics as the interest content, obtaining interest contents; Count one by one the number of occurrences of the interest contents in the personality behavior data, obtaining the number of views, and mark one by one through timestamps the occurrence time of the interest contents in the personality behavior data, obtaining the viewing times; Add up one by one the number of views of the interest contents to obtain the total number of times value, and compare one by one the number of views of the interest contents with the total number of times value to obtain the degree of interest; Match the interest contents, the number of views, the viewing times, and the degree of interest one by one and summarize them to obtain portrait units; Simulate a basic portrait with portrait positions, import the portrait units into the portrait positions one by one, and note the interest contents in the corresponding On the image position, prompt image positions to be converted into image semantics; In the order from the largest to the smallest interest degree, arrange the image semantics in descending order to construct a user personality portrait with image semantics.

[0014] Furthermore, the steps to control the free switching of interaction modes are as follows: Construct a data group with a first data packet and a second data packet. In the order of the screening time, import E target data and H interest data into the first data packet and the second data packet one by one to generate engine enhanced data; When text - based interaction is required, directly send the engine enhanced data to the user side for text - mode interaction; When voice - based interaction is required, convert the engine enhanced data into voice output through voice synthesis technology, and send the voice output to the user side for voice - mode interaction; When image - based interaction is required, convert the engine enhanced data into image output through image recognition technology, and send the image output to the user side for image - mode interaction.

[0015] The technical effects and advantages of a generative smart campus engine system based on full - text retrieval knowledge enhancement according to the present invention: (1): Through the operation of processing the effectiveness of system data, it can remove the negative impacts brought by invalid characters and incorrect formats in the system data, ensure the accuracy, integrity and consistency of the system data, and by constructing a knowledge database with the system data, different types of system data can be hierarchically summarized and centrally processed to achieve the integrated effect of system data. By adding note semantics at the data layer, it not only facilitates the accurate search of system data, but also ensures the stability and anti - interference ability of the overall architecture of the knowledge database, and avoids the phenomenon of mutual interference between different types of system data in different data layers; (2): By analyzing the request semantics through the smart campus engine, it can represent the complex and disordered user request data in a concise and accurate manner, providing a true and accurate basis for screening target data from the knowledge database. By combining the construction of a user personality portrait, it can provide a precise basis for screening interest data corresponding to the recent behavior patterns and behavior preferences of the user side, thus achieving the combined analysis effect of the real - time request information and personalized request information of the user side, and improving the data dimension of the user request information of the user side; (3): By combining the target data with the interest data to form the engine-enhanced data and performing multimodal interactive processing operations on it, it is possible to ensure the diversity of the interactive modalities between the engine-enhanced data and the user side, improve the actual interaction effect between the user side and the data, effectively solve many deficiencies of the existing campus information service technology, greatly improve the level of campus informatization services, and bring all-round convenience and optimization effects to the study, work and life of campus teachers and students. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 FIG. 1 is a schematic structural diagram of a generative intelligent campus engine system based on full-text retrieval knowledge enhancement provided by Embodiment 1 of the present invention; Figure 2 FIG. 2 is a schematic diagram of the modules of the intelligent campus platform provided by Embodiment 1 of the present invention; Figure 3 FIG. 3 is a schematic flowchart of a method for a generative intelligent campus engine based on full-text retrieval knowledge enhancement provided by Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1: Please refer to Figure 1 and Figure 2 As shown, the generative intelligent campus engine system based on full-text retrieval knowledge enhancement described in this embodiment is applied to the intelligent campus platform and includes: A data collection module that collects the system data of the campus subsystem, determines the update period from the system data, and takes the update period as the standard to regularly update the system data; The campus subsystem is a subsystem used to record, store and manage various types of data corresponding to different role personnel on campus, so as to effectively manage different roles and different types of data. The system data is the original data stored in the campus subsystem and is used to provide a data basis for subsequent full-text retrieval operations; The campus subsystem includes a teaching affairs management subsystem, a library management subsystem, a student information management subsystem, an office subsystem, and an environment subsystem. Among them, the teaching affairs management subsystem is a subsystem that can store and manage various types of data in teachers' teaching task activities. The library management subsystem is a subsystem that can store and manage various types of book borrowing data in the library. The student information management subsystem is a subsystem that can store and manage various types of data during students' learning process. The office subsystem is a subsystem that can store and manage data of various collective activities in the school. The environment subsystem is a subsystem that can store and manage the occupancy data of the teaching and office sites in the school.

[0019] After clarifying the campus subsystem, it is necessary to collect system data from within the campus subsystem so that the system data can be used as the object for subsequent full-text retrieval knowledge enhancement operations. The system data includes teaching affairs data, borrowing data, student data, office data, and environment data. By comprehensively collecting teaching affairs data, borrowing data, student data, office data, and environment data, various types of data stored in the campus can be comprehensively collected, expanding the data coverage and also increasing the data dimension.

[0020] Specifically, when collecting teaching affairs data, at this time, through natural language processing technology, the original semantics of all the original data in the teaching affairs management subsystem are identified one by one, and after summarizing the original data whose original semantics are associated with the time, place, and personnel of the teaching task, the teaching affairs data is obtained. From the above method of collecting teaching affairs data, it can be seen that when collecting borrowing data, student data, and office data, the original semantics of all the original data in the library management subsystem, the student information management subsystem, and the campus office subsystem are also identified one by one through natural language processing technology, and after summarizing the original data whose original semantics are associated with the time, place, personnel, type, and quantity of book borrowing, student information, and office activities, the borrowing data, student data, and office data are obtained respectively. When collecting environment data, it is necessary to combine the cameras deployed in various teaching activity sites on the campus to obtain the original data of the occupancy time, occupancy quantity, and occupancy location of the teaching activity sites in real time, and after summarizing the original data, the environment data is obtained.

[0021] After collecting all the system data, it is necessary to update the collected system data to ensure that the collected system data can be updated regularly, thereby avoiding the phenomenon that the subsequent retrieval results of the data are inaccurate due to unupdated or untimely updated system data.

[0022] When updating system data, it is necessary to determine the time span between two consecutive updates of the system data, and record the time span as the update period, so as to ensure that the system data can be regularly updated based on one update period; The steps to determine the update cycle are: Record A original data corresponding to the teaching data, borrowing data, student data, office data and environment data as the data to be checked, and query the time when the data to be checked was first recorded one by one through the timestamp to obtain A initial time; Taking A initial moments as the starting point and the current moment as the end point, the data to be checked are compared one by one, and the moment when the data to be checked changes is recorded; In chronological order, the duration between two adjacent moments is recorded as the interval duration, and the minimum value of the interval duration is recorded as the update period.

[0023] After determining the update cycle, it is necessary to regularly collect system data from the campus subsystem based on the duration of an update cycle, so that A system data can be regularly updated and changed to ensure that the system data can maintain a close connection and consistency with the campus subsystem.

[0024] The database construction module processes the validity of the system data in the data processing mode and constructs the system data into a knowledge database with annotation semantics based on the database construction criteria; The data processing mode is a working mode that can perform pre-processing operations on the format, form and content of the collected system data, so that the collected system data can be deleted, filled and re-entered with effectiveness processing such as the system data with defects in format, form and content and that does not meet the requirements under the support of the data processing mode, ensuring that the system data can be in a state of direct use; The steps of validity processing are: Compare the repeatability of the teaching data, borrowing data, student data, office data and environmental data one by one, record the system data with duplication as duplicate data, and aggregate the same duplicate data to obtain C duplicate sets; Query the recording times of all duplicate data in the C duplicate sets one by one, arrange all the recording times in the C duplicate sets in chronological order, and remove the duplicate data except the last recording time; Perform complete attribute recognition on the remaining educational administration data, borrowing data, student data, office data, and environmental data in C duplicate sets, and record the system data with defective complete attributes as defective data; the complete attribute is an attribute representation used to indicate whether there are defective phenomena in the system data. Specifically, the complete attribute includes defective and non-defective. Defective corresponds to the presence of defective phenomena in the system data, and non-defective corresponds to the absence of defective phenomena in the system data; Based on the calibrated financial complete data, borrowing complete data, student complete data, office complete data, and environmental complete data, automatically fill in the defective data to obtain complete system data. When automatically filling in the defective data, the specific form of the original data without defects will be used as the standard to specifically fill in the type, format, and content of the original data corresponding to the defective data, so as to ensure the integrity of the defective data.

[0025] After performing validity processing on the system data, the system data at this time is in a complete and directly usable state in terms of format, form, and content. That is, based on the complete system data, a knowledge database can be constructed that can summarize, store, and manage all types of system data.

[0026] When constructing the knowledge database, in order to ensure that different types of system data can be constructed in an orderly and accurate manner, it is necessary to construct different types of system data in an orderly manner under the restriction of the database construction criteria to ensure the integrity and accuracy of the constructed knowledge database; The database construction criteria are: one data layer corresponds to one note semantics, and adjacent two data layers are independent of each other; the note semantics are words used to concisely and accurately represent the type and meaning of the system data within each data layer, facilitating subsequent query and identification operations of different types of system data in the knowledge database, and the independent architecture between adjacent data layers can ensure the stability and anti-interference ability of the overall architecture of the knowledge database, avoiding the phenomenon of mutual interference between different types of system data in different data layers.

[0027] The construction steps of the knowledge database are: Establish a blank database with five vertically distributed data layers, and draw a layer outline in the shape of a circular closed structure on the outside of the five data layers; the data layer is a structure used to independently and closedly import and store different types of system data and serves as a component of the knowledge database, and the layer outline is a data firewall used to protect the data layer from being closed, avoiding the phenomenon of mutual interference and leakage of system data between adjacent data layers; Count the quantity of the original data corresponding to the educational administration data, borrowing data, student data, office data, and environmental data one by one, and obtain the first quantity value, the second quantity value, the third quantity value, the fourth quantity value, and the fifth quantity value respectively; In the top-down manner, data bits equal in number to the first value, the second value, the third value, the fourth value, and the fifth value are respectively drawn in five data layers, and educational administration data, borrowing data, student data, office data, and environmental data are imported into the data bits of the five data layers one by one to construct an educational administration layer, a borrowing layer, a student layer, an office layer, and an environmental layer; Semantic note columns are respectively drawn on the layer contours of the educational administration layer, the borrowing layer, the student layer, the office layer, and the environmental layer to obtain a first note column, a second note column, a third note column, a fourth note column, and a fifth note column; the semantic note column can provide a position for text notes for the note semantics, thereby ensuring that each data layer has and only has one corresponding note semantics; Educational administration, borrowing, students, office, and environment are respectively noted in the first note column, the second note column, the third note column, the fourth note column, and the fifth note column to generate note semantics, and a knowledge database with note semantics is constructed.

[0028] It should be noted that the layer contour is an electronic fence used to position-limit the position of each data layer and the original data within the data layer, preventing the leakage, loss, and interactive interference of the original data within the data layer, maintaining the independence between data layers. At the same time, the note column is used to provide a position limit for data import for the note semantics of the data layer, ensuring that each data layer can have a unique note semantics.

[0029] A data screening module receives user request data from the user side, analyzes the request semantics from the user request data through a trained smart campus engine, and performs a matching analysis on the request semantics and the note semantics to screen out target data from the knowledge database; The user side is a port used to access the smart campus engine system and input relevant query request information, enabling data information interaction between the user side and the smart campus engine. Specifically, the user side includes, but is not limited to, the electronic mobile terminals carried by students and teachers.

[0030] The user request data is the original user information input by the user side, which can comprehensively represent the query request information of the user side. The request semantics is a concise and accurate representation of the real content in the user request data, that is, keywords can be accurately extracted from the complex, disordered, and messy user request data, and the request semantics can be used as a direct basis for subsequent corresponding data query and screening from the knowledge database.

[0031] The intelligent campus engine is an artificial intelligence model that can analyze key information from the input user request data and output the key and true request semantics in the user request data. To ensure that the intelligent campus engine can accurately and quickly analyze the request semantics from the user request data, it is necessary to collect a large amount of different types of user request data from the user side, as well as the request semantics corresponding to different user request data, and use the user request data and request semantics as the basis for training to perform a large amount of training and optimization processing on the intelligent campus engine; Specifically, the steps for training the intelligent campus engine are as follows: Pre-collect multiple groups of user request data from the user side and the request semantics corresponding to the user request data. Mark a group of user request data as a group of feature vectors to obtain multiple groups of feature vectors, and convert the request semantics into labels corresponding to the feature vectors; One feature vector corresponds to one label, forming a group of training data. Multiple groups of training data form a training set, and the labeled training data is divided into a training set and a test set. 70% of the training data is used as the training set, and 30% of the training data is used as the test set; Use the feature vectors as the input of the intelligent campus engine, and the request semantics corresponding to the intelligent campus engine as the output of the intelligent campus engine. Use the training set to train the intelligent campus engine, and use the test set to test the intelligent campus engine. Preset an error threshold. When the mean of the prediction errors of all training data in the test set is less than the preset error threshold, obtain the intelligent campus engine that analyzes the request semantics based on the user request data.

[0032] Exemplarily, the intelligent campus engine adopts any one of the support vector machine model or the random forest model; the preset error threshold is set in advance according to the actual required accuracy of the intelligent campus engine.

[0033] Based on the training results of the above intelligent campus engine, the following specific embodiments are given. When the user request data on the user side is "What classes do I have tomorrow", the request semantics analyzed by the intelligent campus engine at this time are "tomorrow" and "courses". When the user request data on the user side is "Query which lecture halls are available on Tuesday", the request semantics analyzed by the intelligent campus engine at this time are "Tuesday", "lecture hall", and "available".

[0034] After training the intelligent campus engine, the user request data on the user side can be input into the intelligent campus engine at this time, so as to analyze the request semantics corresponding to the user request data, and based on the analyzed request semantics, perform a matching analysis on the note semantics of the knowledge database, and screen out the required target data from the knowledge database. The target data at this time is recorded as the original data in different data layers in the knowledge database; The steps for screening the target data are as follows: The D request semantics analyzed are compared with the remark semantics in the knowledge database one by one for coincidence matching, and the data layer where the remark semantics and request semantics overlap is recorded as the target layer; The original semantics of the original data in the target layer is split into text part and numeric part through word segmentation technology, and the text part is compared with the D request semantics for consistency; the text part and numeric part are used to separately represent the text information and numeric information in the identified semantics, that is, the semantics can be comprehensively represented in two dimensions, which is convenient for subsequent classification and screening operations of the semantics; The text part that is consistent with the request semantics is recorded as the target text, and the target text and the corresponding digital part are combined into target data to obtain E target data.

[0035] It should be noted that the screened target data are in a disordered and chaotic state, so that multiple target data cannot be sent directly to the user end, and further optimization processing operations need to be performed on the target data.

[0036] The portrait construction module obtains the user's personality behavior data during the portrait cycle, extracts the portrait features from the personality behavior data, and constructs a user personality portrait with portrait semantics; The portrait cycle is used to limit the duration of the user's data screening and retrieval behavior in the smart campus engine, which can provide a limit on the data collection time for the judgment and analysis of the user's recent specific personality behavior performance; Specifically, the current moment is taken as the end point of the cycle, the moment when the user request data is first input within two weeks before the end point of the cycle is taken as the start point of the cycle, and the period from the start point to the end point of the cycle is recorded as the portrait cycle; exemplarily, the portrait cycle can be 3 days, 7 days or 10 days.

[0037] Personality behavior data is used to represent the specific behavior data of the user during the profiling period, so as to provide data support for the description and characterization of the user's recent personality behavior performance and improve the accuracy of subsequent user personality profiling; Specifically, when obtaining personalized behavior data, it is necessary to obtain all user request data input by the user end during the portrait cycle, and then extract the generation time and generation frequency corresponding to the user request data through log analysis technology, and then summarize all user request data, generation time and generation frequency to obtain personalized behavior data.

[0038] After obtaining the personalized behavior data, the data dimensions included in the personalized behavior data are relatively wide and the data content is relatively miscellaneous at this time. In order to accurately identify the personalized behavior data and facilitate the construction of the subsequent user personalized portrait, it is necessary to extract portrait features from the personalized behavior data, so that the portrait features can represent the personalized behaviors of the user side during the portrait cycle, and the constructed user personalized portrait can accurately represent the specific performance of the user side during the portrait cycle.

[0039] The portrait features include interest content, number of accesses, access time, and degree of interest; the interest content is used to represent the content that the user side is interested in during the portrait cycle, the number of accesses is used to represent the number of access operations of the user side to the interest content during the portrait cycle, the access time is used to represent the specific time when the user side accesses the interest content during the portrait cycle, and the degree of interest is used to represent the degree of interest of the user side in each interest content during the portrait cycle.

[0040] The portrait semantics are words used to intuitively and accurately represent the specific content that the user personalized portrait is interested in, and ensure that the user personalized portrait can intuitively represent the specific behavior patterns and behavior preferences of the user side during the portrait cycle.

[0041] The steps to construct a user personalized portrait are as follows: Use natural language processing technology to identify the behavior semantics of the personalized behavior data one by one, and record the text part in the behavior semantics that is consistent with the note semantics as the interest content to obtain interest contents; the behavior semantics are words used to accurately and concisely represent the content in the personalized behavior data, and can provide a comparison basis for the subsequent identification of interest content; Count one by one the number of occurrences of interest contents in the personalized behavior data to obtain the number of accesses; Use timestamps to mark one by one the occurrence time of interest contents in the personalized behavior data to obtain the access time; Accumulate the number of accesses of interest contents one by one to obtain the total number of times value, and compare the number of accesses of interest contents with the total number of times value one by one to obtain the degree of interest; In the formula, is the degree of interest of the th interest content, is the The number of times an interesting content is viewed; Combine interesting contents, number of views, viewing time, and degree of interest one by one, and summarize to obtain image units; Simulate a basic image with image positions, and import image units one by one into image positions, and mark interesting contents on the corresponding image positions respectively, so as to prompt image positions to be converted into image semantics; In the order of decreasing degree of interest, sort image semantics in descending order to construct a user personality portrait with image semantics. By sorting the image semantics in the order of decreasing degree of interest, it is possible to arrange them from high to low according to the degree of interest of the user in different interesting contents, so that the interesting contents with high degree of interest can be more prominent, and thus the effective distinction effect of the user personality portrait on the degree of interest in different interesting contents can be achieved.

[0042] Data combination module, screen out the interest data corresponding to the user personality portrait from the knowledge database, combine the interest data and the target data into engine-enhanced data, and control the free switching of the interaction mode when the engine-enhanced data interacts with the user side; Interest data refers to the original data screened out from the knowledge database corresponding to the image semantics of the user personality portrait, so that the interest data can be associated and matched with the user personality portrait, ensuring that the interest data can truly, comprehensively and accurately summarize the recent behavior patterns and behavior preferences of the user side; When screening interest data, it is necessary to use the image semantics of the user personality portrait as the basis for screening to achieve the accurate screening of interest data; Specifically, compare the text parts of image semantics one by one with the text parts of the original semantics of the original data in the knowledge database, and mark the original data with the text of image semantics as interest data to obtain H interest data.

[0043] After screening out the interest data, it is necessary to summarize and integrate the interest data and the target data at this time, and generate engine-enhanced data that can not only match the user request data but also be adapted to the user personality portrait, so that the engine-enhanced data can be used as the direct object for subsequent interaction with the user side.

[0044] When the engine-enhanced data interacts with the user terminal, in order to ensure the diversity of the interaction between the engine-enhanced data and the user terminal and improve the effect of data interaction, it is necessary to freely adjust the interaction modality of the engine-enhanced data, and finally achieve the multi-modal interaction effect between the engine-enhanced data and the user terminal; The interaction modalities include text modality, voice modality, and image modality; the text modality means that the engine-enhanced data interacts with the user terminal in the form of text, the voice modality means that the engine-enhanced data interacts with the user terminal in the form of voice, and the image modality means that the engine-enhanced data interacts with the user terminal in the form of an image; The steps to control the free switching of interaction modalities are as follows: Construct a data group with a first data packet and a second data packet, and sequentially import E target data and H interest data into the first data packet and the second data packet one by one according to the order of screening time to generate engine-enhanced data; When the engine-enhanced data needs to interact with the user terminal in text form, directly send the engine-enhanced data to the user terminal for text modality interaction; When the engine-enhanced data needs to interact with the user terminal in voice form, convert the engine-enhanced data into voice output through text-to-speech technology, and send the voice output to the user terminal for voice modality interaction; When the engine-enhanced data needs to interact with the user terminal in image form, convert the engine-enhanced data into image output through image recognition technology, and send the image output to the user terminal for image modality interaction.

[0045] It should be noted that text-to-speech technology is a kind of artificial intelligence technology. By converting text information into natural and fluent voice output through algorithms, the effect of converting engine-enhanced data from text to voice can be achieved. Image recognition technology is a kind of artificial intelligence technology, which can convert text information into image output of corresponding image tables, and the effect of converting engine-enhanced data from text to image can be achieved. Both text-to-speech technology and image recognition technology are conventional technologies in this field and are not the invention points of this application, so they will not be elaborated in detail.

[0046] In this embodiment, through the operation of processing the validity of system data, the negative impacts brought by invalid characters and incorrect formats in the system data can be removed, ensuring the accuracy, integrity, and consistency of the system data. By constructing a knowledge database from the system data, different types of system data can be hierarchically summarized and centrally processed to achieve the integrated effect of system data. Combining with the method of adding note semantics at the data layer not only facilitates the accurate search of system data but also ensures the stability and anti-interference ability of the overall architecture of the knowledge database, avoiding the phenomenon of mutual interference between different types of system data in different data layers; By analyzing the request semantics through the intelligent campus engine, the complex and disordered user request data can be concisely and accurately represented, providing a true and accurate basis for screening target data from the knowledge database. Combining with the method of constructing a user personality profile, it can provide a precise basis for screening the interest data corresponding to the recent behavior patterns and behavior preferences of the user side, thus achieving the combined analysis effect of the real-time request information and personalized request information of the user side and improving the data dimension of the user request information on the user side; By combining the target data and the interest data into the engine-enhanced data and performing multimodal interaction processing operations on it, the diversity of the interaction modality between the engine-enhanced data and the user side can be ensured, improving the actual interaction effect between the user side and the data, effectively solving many deficiencies of the existing campus information service technology, greatly enhancing the campus informatization service level, and bringing all-round convenience and optimization effects to the study, work, and life of campus teachers and students.

[0047] Embodiment Two: Please refer to Figure 3 As shown in the figure, the parts not described in detail in this embodiment can be seen in the description of Embodiment One. A generative intelligent campus engine method based on full-text retrieval knowledge enhancement is provided, which is applied to an intelligent campus platform and implemented based on a generative intelligent campus engine system based on full-text retrieval knowledge enhancement, including: S1: Collect the system data of the campus subsystem, determine the update period from the system data, and update the system data regularly; S2: In the data processing mode, perform validity processing on the system data, and based on the database construction criteria, construct the system data into a knowledge database with note semantics; S3: Receive the user request data from the user side, analyze the request semantics from the user request data through the trained intelligent campus engine, perform matching analysis on the request semantics and the note semantics, and screen out the target data from the knowledge database; S4: Obtain the personalized behavior data of the user side within the portrait period, extract the portrait features from the personalized behavior data, and construct a user personality portrait with portrait semantics; S5: Screen out the interest data corresponding to the user's personality portrait from the knowledge database, combine the interest data and the target data into engine enhancement data, and control the free switching of the interaction mode.

[0048] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. A generative intelligent campus engine system based on full-text retrieval knowledge enhancement, which is applied to an intelligent campus platform, and is characterized in that include: The data collection module is used to collect system data of the campus subsystem, determine the update cycle from the system data, and update the system data regularly; The database construction module is used to process the validity of the system data in the data processing mode and construct the system data into a knowledge database with annotation semantics based on the database construction criteria; The data screening module is used to receive user request data from the user end, analyze the request semantics from the user request data through the trained smart campus engine, and match and analyze the request semantics and remark semantics to screen out the target data from the knowledge database; The portrait construction module is used to obtain the user's personality behavior data during the portrait cycle, extract the portrait features from the personality behavior data, and construct a user personality portrait with portrait semantics; The data combination module is used to filter out interest data corresponding to the user's personality portrait from the knowledge database, combine the interest data and target data into engine enhanced data, and control the free switching of interaction modes.

2. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 1, wherein The campus subsystem includes the academic affairs management subsystem, library management subsystem, student information management subsystem, office subsystem and environment subsystem; System data includes academic data, borrowing data, student data, office data and environmental data.

3. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 2, wherein The steps to determine the update cycle are: Record A original data corresponding to the teaching data, borrowing data, student data, office data and environment data as the data to be checked, and query the time when the data to be checked was first recorded one by one through the timestamp to obtain A initial time; Taking A initial moments as the starting point and the current moment as the end point, the data to be checked are compared one by one, and the moment when the data to be checked changes is recorded; In chronological order, the duration between two adjacent moments is recorded as the interval duration, and the minimum value of the interval duration is recorded as the update period.

4. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 3, wherein The steps of validity processing are: Compare the repeatability of the teaching data, borrowing data, student data, office data and environmental data one by one, record the system data with duplication as duplicate data, and aggregate the same duplicate data to obtain C duplicate sets; Query the recording times of all duplicate data in the C duplicate sets one by one, arrange all the recording times in the C duplicate sets in chronological order, and remove the duplicate data except the last recording time; Perform complete attribute recognition on the remaining teaching data, borrowing data, student data, office data, and environment data in the C repeated sets, and record the system data with defective complete attributes as defective data; Based on the calibrated complete financial data, complete borrowing data, complete student data, complete office data and complete environmental data, the defective data is automatically filled in to obtain complete system data.

5. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 4, characterized in that The database construction principle is: one data layer corresponds to one annotation semantics, and two adjacent data layers are independent of each other; The steps to construct the knowledge database are: A blank database with five data layers distributed vertically is established, and layer outlines of annular closed structures are drawn on the outsides of the five data layers; Count the quantities of the original data corresponding to the educational administration data, borrowing data, student data, office data, and environmental data one by one, and obtain the first quantity value, the second quantity value, the third quantity value, the fourth quantity value, and the fifth quantity value respectively; In the top-down manner, draw data bits with the same quantity as the first quantity value, the second quantity value, the third quantity value, the fourth quantity value, and the fifth quantity value in the five data layers respectively, and import the educational administration data, borrowing data, student data, office data, and environmental data into the data bits of the five data layers one by one to construct the educational administration layer, borrowing layer, student layer, office layer, and environmental layer; Draw semantic note columns on the layer outlines of the educational administration layer, borrowing layer, student layer, office layer, and environmental layer respectively to obtain the first note column, the second note column, the third note column, the fourth note column, and the fifth note column; Remark educational administration, borrowing, students, office, and environment in the first note column, the second note column, the third note column, the fourth note column, and the fifth note column in sequence to generate remark semantics, and construct a knowledge database with remark semantics.

6. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 5, wherein The steps to train the smart campus engine are as follows: Pre-collect multiple groups of user request data at the user end and the request semantics corresponding to the user request data, mark the user request data as feature vectors, and convert the request semantics into labels corresponding to the feature vectors; One feature vector corresponds to one label, forming a group of training data. Multiple groups of training data form a training set, and the labeled training data is divided into a training set and a test set; Use the feature vectors as the input of the smart campus engine, use the request semantics corresponding to the smart campus engine as the output of the smart campus engine, train the smart campus engine with the training set, test the smart campus engine with the test set, preset an error threshold, and when the mean of the prediction errors of all training data in the test set is less than the preset error threshold, obtain the smart campus engine.

7. The generative smart campus engine system based on full-text retrieval knowledge enhancement according to claim 6, characterized in that, The steps to screen target data are as follows: Compare each of the analyzed D request semantics with the remark semantics in the knowledge database one by one, and record the data layer where the remark semantics and the request semantics overlap as the target layer; Use the word segmentation technology to split the original semantics of the original data in the target layer into a text part and a numerical part, and compare the text part with the D request semantics for consistency; Record the text part that is consistent with the request semantics as the target text, and combine the target text with the corresponding numerical part to form the target data, obtaining E target data.

8. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 7, wherein Take the current moment as the end point of the period, take the moment when the user request data is first input within the first two weeks before the end point of the period as the start point of the period, and record the time period between the start point and the end point of the period as the portrait period.

9. The generative intelligent campus engine system based on full-text retrieval knowledge enhancement according to claim 8, characterized in that, The portrait features include interest content, number of accesses, access time, and degree of interest; The steps to construct a user personality portrait are as follows: Identify the behavioral semantics of individual behavioral data one by one through natural language processing technology, and record the text part in the behavioral semantics that is consistent with the note semantics as the interest content, obtaining interest contents; Count one by one the number of occurrences of the interest content in the personality behavior data, and obtain the number of access times, and mark one by one through the time stamp the occurrence time of the interest content in the personality behavior data, and obtain the access time; Accumulate the access times of each piece of interesting content one by one to obtain the total number of times, and compare the access times of each piece of interesting content with the total number of times one by one to obtain the degree of interest; Match one by one and summarize interesting contents, access times, access times, interest levels, and portrait units are obtained;​​​​​​​​​​ Simulate a base image with image positions, and import image units one by one into image positions, and respectively mark pieces of interesting content on the corresponding image positions, prompting image positions to be converted into image semantics; Arrange the image semantics in descending order according to the degree of interest, and construct a user personality portrait with image semantics.

10. For a generative smart campus engine system based on full-text retrieval knowledge enhancement according to claim 9, the steps to control the free switching of the interaction mode are as follows: Construct a data group with a first data packet and a second data packet. In the order of the screening time, import the E target data and H interest data into the first data packet and the second data packet one by one to generate engine enhancement data; When text-based interaction is required, the engine-enhanced data is directly sent to the user side for text-modal interaction; When voice-based interaction is required, the engine-enhanced data is converted into voice output through text-to-speech technology, and the voice output is sent to the user side for voice-modal interaction; When image-based interaction is required, the engine-enhanced data is converted into image output through image recognition technology, and the image output is sent to the user side for image-modal interaction.

Citation Information

Patent Citations

  • Smart campus multimedia comprehensive information service terminal management method and system

    CN115879001A

  • Intelligent campus unified portal AI semantic analysis robot system

    CN116932720A

  • Information acquisition method and device based on user portrait, and electronic equipment

    CN117312505A

  • User portrait label rapid matching method based on big data

    CN118012920A

  • Campus man-machine interaction method and system based on artificial intelligence dialogue model

    CN118838990A