Book list generation method and device, electronic equipment, storage medium and product
By identifying authoritative data sources and generating structured datasets, and combining user needs with data update time, the problem of the lack of authority in traditional book list recommendations is solved, and highly reliable and relevant book list generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional book list recommendations are based on content similarity or user behavior analysis, resulting in book lists that lack authority and fail to meet users' needs for highly credible and relevant recommendations.
By identifying authoritative data sources based on user attribute information, obtaining raw recommendation data from these sources, generating a structured dataset, and determining the comprehensive recommendation score for each book based on the type and administrative level of the recommending institutions within the structured dataset, combined with user needs and data update time, a target book list is generated.
This enhances the authority and relevance of the book lists, ensures that the recommended content is highly compatible with the user's identity, and improves the credibility and transparency of the recommendation results.
Smart Images

Figure CN121636829A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence and data fusion technology, and in particular to a book list generation method, apparatus, electronic device, storage medium and product. Background Technology
[0002] With the rapid development of the internet and digital reading, users often face information overload and inconsistent quality of recommended content when confronted with a vast amount of book resources. Traditional book list recommendations are mostly based on content similarity or user behavior analysis, focusing on personalized matching, resulting in book lists lacking authority. Summary of the Invention
[0003] This disclosure provides a book list generation method, apparatus, electronic device, storage medium, and product to address the problem that traditional book list recommendations in related technologies are mostly based on content similarity or user behavior analysis, focusing on personalized matching, resulting in a lack of authority in the generated book lists.
[0004] A first aspect of this disclosure provides a method for generating a book list, the method comprising: Determine authoritative data sources based on user attribute information; We obtain raw recommendation data from authoritative data sources and process it to generate a structured dataset. Based on the type and administrative level of the recommending organization in the structured dataset, the authority score of each book item is determined; Based on user needs, authority ratings, and data update time, a comprehensive recommendation score is determined for each book item, and a target book list is generated based on the comprehensive recommendation score.
[0005] In one embodiment, determining the authoritative data source based on user attribute information includes: Analyze user attribute information to obtain user demand categories; Identify the level of authoritative institutions that corresponds to the user's needs, and determine the data sources belonging to the level of authoritative institutions as authoritative data sources.
[0006] In one embodiment, the original recommendation data is processed to generate a structured dataset, including: The original recommendation data is formatted to obtain intermediate recommendation data. Extract semantic content related to book recommendations from intermediate recommendation data; Generate structured datasets based on semantic content.
[0007] In one embodiment, extracting semantic content associated with book recommendations from intermediate recommendation data includes: In response to the inclusion of non-text data in the intermediate recommendation data, the non-text data is converted into text data. Extract semantic content related to book recommendations from the converted text data.
[0008] In one embodiment, a comprehensive recommendation score for each book is determined based on user needs, authority rating, and data update time, including: Determine the data credibility weight and target weight; The comprehensive recommendation score for each book is determined based on data credibility weight, target weight, user needs, authority score, and data update time.
[0009] In one embodiment, the target weights include user demand weights, authority rating weights, and data update time weights, wherein the data update time weights are determined by the time interval from the current time.
[0010] A second aspect of this disclosure provides a book list generation apparatus, the apparatus comprising: The first determining unit is used to determine the authoritative data source based on user attribute information; The first generation unit is used to obtain raw recommendation data from authoritative data sources and process the raw recommendation data to generate a structured dataset. The second determining unit is used to determine the authority score of each bibliography based on the type and administrative level of the recommending organization in the structured dataset. The second generation unit is used to determine the comprehensive recommendation score for each book based on user needs, authority rating, and data update time, and to generate the target book list based on the comprehensive recommendation score.
[0011] A third aspect of this disclosure provides an electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the first aspect of this disclosure.
[0012] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the first aspect of this disclosure.
[0013] A fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the methods described in the first aspect of this disclosure.
[0014] In summary, this disclosure proposes a book list generation method, which includes: determining authoritative data sources based on user attribute information, which can improve the authority of the book list source; obtaining raw recommendation data from authoritative data sources and processing the raw recommendation data to generate a structured dataset; determining the authority score of each book item based on the type and administrative level of the recommending institution in the structured dataset; determining the comprehensive recommendation score of each book item by considering user needs, authority score, and data update time, and generating a target book list based on the comprehensive recommendation score, which can improve the authority of the generated book list.
[0015] According to the solution provided in this disclosure, the authority of the book list source can be improved by determining the authoritative data source based on user attribute information; the original recommendation data is obtained from the authoritative data source and processed to generate a structured dataset; the authority score of each book item is determined based on the type and administrative level of the recommending institution in the structured dataset; the comprehensive recommendation score of each book item is determined based on user needs, authority score and data update time, and the target book list is generated based on the comprehensive recommendation score, which can improve the authority of the generated book list.
[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0018] Figure 1 A flowchart illustrating the book list generation method provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating the book list generation method provided in this embodiment of the disclosure; Figure 3 This is a schematic diagram of the structure of the book list generation device provided in the embodiments of this disclosure; Figure 4 This is a schematic diagram of the hardware composition structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0019] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0020] E-commerce platforms typically recommend books to users in the form of book lists. These solutions generally filter books from a library based on the reader's attribute tags to generate the book list. For example, if the reader's attribute tag is "elementary school student," this method will select books suitable for elementary school students from the library to generate the book list. However, because reader attribute tags are relatively fixed, the book lists generated by this method are rather limited and lack authority.
[0021] Some companies automatically generate book lists based on trending online events, holidays, or World Book Day activities, resulting in a greater diversity of books on the lists. However, these lists lack authority. Experimental statistics show that the authority and recognition of book lists recommended by existing e-commerce platforms is less than 30%, and the rate of users revisiting them is less than 15%.
[0022] With the rapid development of the internet and digital reading, users often face information overload and inconsistent quality of recommended content when confronted with a vast amount of book resources. Traditional book list recommendations are mostly based on content similarity or user behavior analysis, focusing on personalized matching, resulting in book lists lacking authority.
[0023] To address the shortcomings of related technologies, this disclosure improves the authority of book list sources by determining authoritative data sources based on user attribute information; it obtains raw recommendation data from authoritative data sources and processes the raw recommendation data to generate a structured dataset; it determines the authority score of each book item based on the type and administrative level of the recommending organization in the structured dataset; and it improves the authority of the generated book list by determining the comprehensive recommendation score of each book item based on user needs, authority score, and data update time, and generating the target book list based on the comprehensive recommendation score.
[0024] The present disclosure will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0025] The book list generation method provided in this disclosure is applicable to all intelligent educational scenarios that require filtering highly reliable and relevant recommended content from a vast amount of books, such as personalized reading recommendations in the field of preschool to high school (K12) education, book list recommendations for professional qualification examinations and continuing education, recommendations for reference books in higher education and academic research, and embedded recommendation services in smart terminal devices. The executing entity of the method can be a computer system or service platform, such as an online education platform, a digital library / public library service platform, or a content recommendation system of an e-commerce platform.
[0026] like Figure 1 As shown, Figure 1This is a flowchart illustrating the book list generation method provided in this embodiment of the disclosure. The book list generation method provided in this embodiment of the disclosure includes the following steps: Step 101: Determine the authoritative data source based on user attribute information; In one embodiment, user attribute information refers to data used to describe user characteristics, including but not limited to age, region, education level, occupation, historical reading behavior, search keywords, etc.
[0027] In one embodiment, user attribute information can be obtained by analyzing user profiles or by collecting data through questionnaires.
[0028] In one embodiment, authoritative data sources refer to institutions with official or professional endorsement in a specific knowledge field, whose recommendations are credible. These institutions may include the Ministry of Education, provincial education departments, key primary and secondary schools, industry associations, and government-designated textbook publishers, such as the Ministry of Education's "Reading Guidance Catalog for Primary and Secondary School Students" and the programming learning book list recommended by the China Computer Federation.
[0029] For example, taking user A as an example, the user's attribute information is age 11, location Haidian District, Beijing, and recent search for "cosmic science popularization". Analysis shows that the user's need category is K12 science literacy improvement. Authoritative data sources matching this user are identified, such as the Curriculum and Textbook Development Center of the Ministry of Education, the Beijing Municipal Education Commission, the Haidian District Institute of Educational Sciences, and the Tsinghua University Primary School School-Based Curriculum Group.
[0030] In one embodiment, authoritative data sources are determined based on user attribute information to avoid the blindness of general recommendations, ensure that recommended content is highly adapted to user identity (such as school level, region), and improve relevance and authority.
[0031] Step 102: Obtain raw recommendation data from authoritative data sources and process the raw recommendation data to generate a structured dataset; In one embodiment, the original recommendation data refers to unstructured or semi-structured book list information released by authoritative institutions. Common forms include official website announcements, WeChat public account articles, Portable Document Format (PDF) files, press releases, etc., and the content includes book titles, authors, reasons for recommendation, and target audience.
[0032] In one embodiment, processing the original recommendation data refers to the process of cleaning, format conversion, information extraction, and structured organization of the original recommendation data, with the aim of transforming messy data into readable and analyzable data in a standard format.
[0033] In one embodiment, the structured dataset refers to a standardized set of database tables formed after processing. Unlike traditional book lists that are limited to data related to recommended books, the data table and field planning and design focuses on adding tables and fields such as recommendation organization table, recommendation information table, book recommendation type, and recommendation statistics table on the basis of the detailed recommended book list. Through the planning of these tables and fields, the authority of the recommendation results is improved.
[0034] In one embodiment, the structured dataset typically includes: a bibliographic information table (fields such as book title, author, publisher, publication date, etc.); a recommendation institution table, used to record detailed information about recommending institutions, including but not limited to fields such as recommendation type, administrative level, authority level, institution name, school system, province, city, district / county, school type, institution category, and provincial ranking; a recommendation information table, used to record detailed recommendation information, including but not limited to fields such as recommendation date, recommendation time period, recommended reading type, subject summary, subject, reading target audience, applicable audience, extended category, and book title; a bibliographic type recommendation table, used to record the recommendation type of the books, including but not limited to fields such as book title, subject category, subject subcategory, and content category; and a bibliographic recommendation statistics table, used to record statistical data related to book recommendations, including but not limited to fields such as book name, number of recommendations, number of recommending institutions, number of author recommendations, number of category recommendations, print book matching, and e-book matching.
[0035] In one embodiment, taking the aforementioned user A as an example, the "Reading Guidance Catalog for Primary and Secondary School Students" is obtained from the official website of the Ministry of Education, which is in HyperText Markup Language (HTML) format; an image-text article (image format) is obtained from the WeChat official account of Haidian Academy of Educational Sciences, and Optical Character Recognition (OCR) is used to recognize the text; the PDF version of "Recommended Science and Technology Books for Primary and Secondary Schools in Beijing" is extracted; the extracted text information is cleaned of noise, which may be advertisements or irrelevant paragraphs; after cleaning, the book title format is unified and duplicate entries (book list) are removed; and a structured dataset is established based on the aforementioned data table and fields.
[0036] In one embodiment, raw recommendation data is obtained from authoritative data sources and processed to generate a structured dataset that supports multiple input formats such as web pages, images, and PDFs. This significantly improves the utilization rate of unstructured authoritative data and solves the problem of difficult data collection caused by the diversity of data formats.
[0037] Step 103: Determine the authority score for each bibliography based on the type and administrative level of the recommending organization in the structured dataset; In one embodiment, the type of recommending organization refers to the nature of the organization (such as education administrative departments, schools, industry associations); the administrative level refers to its level in the management system (national, provincial, municipal), and the two together determine its authority and influence.
[0038] In one embodiment, the authority rating is a numerical value that quantifies the authority of a book being recommended. It is usually calculated by weighting the type weight of the recommending institution and the administrative level coefficient. The type weight of the recommending institution is, for example, education authority = 1.0, key school = 0.8, ordinary school = 0.6, and the administrative level is, for example, Ministry of Education (0.9), provincial education department (0.7), municipal education bureau (0.5).
[0039] In one embodiment, taking the type weight of the aforementioned recommending institution, such as education authority = 1.0, key school = 0.8, ordinary school = 0.6, and administrative level such as Ministry of Education (0.9), provincial education department (0.7), municipal education bureau (0.5) as an example, "100,000 Whys" is recommended by the Ministry of Education, and its authority score is 1.0 × 0.9 = 0.9.
[0040] In one embodiment, the authority score of each book is determined by centralizing the type and administrative level of the recommending organization in a structured dataset, which quantitatively assesses the credibility of the recommendation source, transforming authority from a subjective judgment into a calculable indicator and improving the credibility of the recommendation results, i.e., the book list.
[0041] Step 104: Based on user needs, authority rating, and data update time, determine the comprehensive recommendation score for each book item, and generate the target book list based on the comprehensive recommendation score.
[0042] In one embodiment, user needs refer to the user's current knowledge acquisition goals, which can be obtained through user profile classification, such as K12 extracurricular reading, CPA exam preparation, and introductory learning of artificial intelligence.
[0043] In one embodiment, the data update time refers to the publication time of the original recommendation data or the time of the most recent revision, which is used to measure the freshness of the information. The closer the data update time is to the current time, the stronger the timeliness.
[0044] In one embodiment, the comprehensive recommendation score is the final total score used for ranking. It is calculated by weighting three dimensions: user demand matching degree, authority score, and data update time, and reflects the overall recommendation degree of the book.
[0045] In one embodiment, a comprehensive recommendation score for each book is determined based on user needs, authority rating, and data update time. A target book list is then generated based on the comprehensive recommendation score. Each book can be traced back to the original recommending organization, enhancing the system's transparency and resistance to disputes.
[0046] The authority of the book list can be improved by identifying authoritative data sources based on user attribute information; raw recommendation data can be obtained from authoritative data sources and processed to generate a structured dataset; the authority score of each book can be determined based on the type and administrative level of the recommending organization in the structured dataset; and the authority of the generated book list can be improved by determining the comprehensive recommendation score of each book based on user needs, authority score, and data update time, and generating the target book list based on the comprehensive recommendation score.
[0047] In one embodiment, determining the authoritative data source based on user attribute information includes: Analyze user attribute information to obtain user demand categories; In one embodiment, the user demand category refers to mapping the original user attributes to preset knowledge demand types through a classification model or rule engine, such as K12 extracurricular reading, professional qualification exam preparation, higher education general education, and interest-based extended reading.
[0048] In one embodiment, a pre-trained classification model, such as one based on a decision tree or a lightweight neural network, can be used to analyze user attribute information to obtain user demand categories. For example, the input of the classification model is user attribute information, including age 15, location in Pudong New Area, Shanghai, third year of junior high school, recent searches for middle school entrance examination Chinese language questions, classical Chinese translation techniques, middle school entrance examination essay samples, and no clear career direction. The classification model outputs the user demand category as K12 entrance examination preparation - Chinese language subject.
[0049] Identify the level of authoritative institutions that corresponds to the user's needs, and determine the data sources belonging to the level of authoritative institutions as authoritative data sources.
[0050] In one embodiment, the level of authority corresponding to a user need category can be determined through a pre-established mapping relationship between user need categories and the administrative level of authority institutions. Examples include primary school students and municipal or higher-level educational institutions; certified public accountant candidates and national-level industry associations; and artificial intelligence beginners and higher education institutions or science and technology management departments.
[0051] In one embodiment, the mapping relationship between user demand categories and the administrative level of authoritative institutions can be represented in the form of a mapping relationship table or in the form of a mapping relationship array. This application does not limit this.
[0052] In one embodiment, the level of the authoritative institution refers to the level of the recommending entity within the administrative or professional system, typically categorized as: national (e.g., the Ministry of Education, the China Association for Science and Technology), provincial (e.g., provincial education departments, provincial science and technology departments), municipal (e.g., municipal education bureaus), and district / school (e.g., key primary and secondary schools, university departments). The higher the level, the broader the coverage and the stronger the credibility of the recommended content.
[0053] In one embodiment, the original recommendation data is processed to generate a structured dataset, including: The original recommendation data is formatted to obtain intermediate recommendation data. In one embodiment, format unification refers to the process of converting raw data in different formats into a unified intermediate representation, such as converting HTML into plain text, recognizing images as text through OCR, and extracting PDFs into editable text. The purpose is to eliminate differences in data format and facilitate subsequent unified processing.
[0054] In one embodiment, intermediate recommendation data refers to standardized intermediate products obtained after format conversion. These are usually plain text or lightweight markup text, such as JavaScript Object Notation (JSON) or Markdown, which retain the original semantics but remove format noise. They are transitional data before structured processing.
[0055] Extract semantic content related to book recommendations from intermediate recommendation data; Generate structured datasets based on semantic content.
[0056] In one embodiment, extracting semantic content related to book recommendations from intermediate recommendation data refers to identifying and extracting key information fields from the intermediate recommendation data, such as book title, author, publisher, recommending institution, applicable grade level, and reasons for recommendation, while removing irrelevant content (such as advertisements, copyright statements, and navigation bars).
[0057] In one embodiment, the intermediate recommendation data is text. Text content extraction, locating the main text region, and extracting semantic text blocks include: main text region location, semantic tag recognition, and priority extraction. <article> 、 <main> 、 Typical containers, etc. Density analysis: Statistical text density (text / tag ratio), filtering noisy blocks such as navigation bars and advertisements. Text cleaning and normalization can be performed using tag denoising: removing... <script>、<style>等非文本标签及其内容。利用空白处理:合并连续空格、换行符,保留段落分隔符(如标签转\n)。实体解码:转换&、<等HTML实体为可读字符。分段与语义分割,段落划分:根据、标签或换行符生成自然段落。标题层级:提取<h1>至<h6>标签构建内容大纲。
[0058] 在一实施例中,从中间推荐数据中提取与书目推荐关联的语义内容,包括:响应于中间推荐数据包括非文本数据,将非文本数据转换为文本数据;在一实施例中,非文本数据指无法直接通过字符编码读取语义内容的数据类型,主要包括:图像文件、扫描件、嵌入式图表、手写体截图等。在书单生成场景中,常见于微信公众号推文截图、教育局发布的PDF扫描版文件等。将非文本数据转换为文本数据包括利用OCR技术,将图像中的可视文字内容转化为机器可读的字符串序列,转换过程包括图像预处理、字符分割、识别与后处理纠错等步骤。
[0059] 在一实施例中,中间推荐数据为图片,使用OCR引擎,如PaddleOCR或Tesseract逐页处理图像,图像预处理包括去噪、二值化、倾斜校正。
[0060] 在一实施例中,中间推荐数据为表格,识别表格中的每一行文字。
[0061] 在一实施例中,提取与书目推荐相关的语义内容指从OCR输出的文本中,识别并抽取与图书推荐直接相关的关键信息,如书名、作者、出版社、推荐机构、适用学段、推荐理由等,并剔除无关内容,如广告语、页眉页脚、版权声明。
[0062] 从转换后的文本数据中提取与书目推荐关联的语义内容。
[0063] 在一实施例中,中间推荐数据为图片,图片内容提取,HTML / 用户生成内容(User-Generated Content,UGC)原始数据中存在图片格式的内容时,应用内容提取功能。定位图片区域,调用OCR识别能力接口开展图像识别,输出语义化文本数据。应用OCR识别能力时,利用数字(序号)、符号(书名号、括号、标点符号、换行符、空格等)等特殊字符辅助判断并输出书名,同时关联书名、作者等因素与图书数据库比对分析,纠正书名中OCR错误的字。
[0064] 在一实施例中,中间推荐数据为表格,表格内容提取,层级化解析表格结构(表头、数据行、合并单元格),表格定位包括标签选择:识别标签及其嵌套层级。上下文特征:结合id、class属性(如class="data-table")精准定位目标表格。表头与数据行分离,表头识别:优先提取标签内容,若无则默认第一行为表头。数据行遍历:逐行解析内的标签,保留原始行列序号。合并单元格处理包括属性解析:识别rowspan和colspan属性,映射到二维矩阵填补空缺值;动态扩展:生成虚拟单元格填充因合并导致的空缺位置,确保数据结构完整性。输出结构化数据包括表格映射:生成二维数组或键值对集合(表头-数据行关联)。格式兼容:支持导出为逗号分隔值(Comma-Separated Values,CSV)、JSON或数据库表结构。
[0065] 在一实施例中,通过响应于中间推荐数据包括非文本数据,将非文本数据转换为文本数据;从转换后的文本数据中提取与书目推荐关联的语义内容,实现了对多模态输入的智能适配,系统能自动识别数据类型并选择处理路径,无需人工干预,提升自动化水平与鲁棒性。
[0066] 在一实施例中,基于用户需求、权威程度评分和数据更新时间,确定每项书目的综合推荐评分,包括:确定数据可信度权重和目标权重;在一实施例中,数据可信度权重指的是大模型数据补全置信度系数,目标权重指的是指用户需求、权威程度评分、数据更新时间三个维度在综合评分中的权重。
[0067] 基于数据可信度权重、目标权重、用户需求、权威程度评分和数据更新时间,确定每项书目的综合推荐评分。
[0068] 在一实施例中,由于教育部门、学校、培训机构等等权威机构对于推荐的图书每年 / 每2年变化率不超过10%,新增推荐的图书书目基于权威机构对考试大纲的变化、考点的方向变化,有着一定的指引作用,因此需要考虑数据更新时间。
[0069] 在一实施例中,数据更新时间越接近当前时间,时效性越强,对应综合推荐评分越高。
[0070] 在一实施例中,可以通过多源数据融合模块生成三维推荐矩阵确定书目的综合推荐评分,三维推荐矩阵的维度指的是用户需求维度(权重占比α)、权威等级维度(权重占比β)、时效性维度(权重占比γ),其中α+β+γ=1且β≥0.4(β=max(0.4,0.6-0.05×α),权威维度权重不低于40%),λ为大模型数据补全置信度系数。基于强化学习模型动态调整推荐参数,输出包含权威认证标识的个性化书单。
[0071] 综合推荐评分=(α×用户需求)+(β×权威等级)+(γ×时效性)+λ×补全因子。
[0072] 在一实施例中,时效性加权规则可以按照一个学期为一个基础的时效性维度,最近的一个学期之内的数据权重最高1.1γ,按照学期时间依次递减0.01,最小不小于γ。
[0073] 三维推荐矩阵目标函数为;
[0074] 其中,约束条件包括,,,其中,指综合推荐评分,Ui指用户需求匹配度,Ai指权威等级评分,Ti指时效性评分,λ指正则化系数,前述大模型数据补全置信度系数,用于防止过拟合,指补全因子。
[0075] 在一实施例中,目标权重包括用户需求权重、权威程度评分权重和数据更新时间权重,其中,数据更新时间权重由距离当前时间的时间间隔确定。
[0076] 在一实施例中,用户需求权重表示用户需求匹配度在综合评分中的影响比例。反映系统对个性化需求的重视程度,取值范围通常为[0,1],且与其他权重之和为1。权威程度评分权重表示权威性在综合评分中的影响比例。用于强化来自高信用机构,如教育部、省级教研室推荐内容的优先级,体现推荐结果的公信力导向。数据更新时间权重表示信息新鲜度在综合评分中的影响比例。越新的推荐内容,其时间权重越高,确保书单内容紧跟政策与教学实践变化。
[0077] 在一实施例中,用户需求权重、权威程度评分权重和数据更新时间权重可以根据用户需求确定,如考试类强调权威性,可以确定较大的权威程度评分权重,兴趣类强调用户需求的匹配度,即可以确定较大的用户需求权重,提升推荐策略的灵活性与场景适应性。
[0078] 示例性的,用户标签为考研政治,确定候选书目《考研政治红宝书》与该需求的匹配度为0.95;该书被教育部考试中心推荐,对应行政层级为国家级,对应权威程度评分为0.9;推荐发布时间为2025年4月,距当前(2025年10月)6个月,数据更新时间(时效性)评分为0.8。可以采用动态加权策略确定各维度权重,即前述目标权重,可以根据不同用户场景调整权重:对于"升学考试类”需求,权威性优先,确定权威程度评分权重β=0.6,用户需求权重α=0.3,数据更新时间权重γ=0.1,权重之和为1,且权威权重不低于预设阈值0.5。引入评分置信度,若书目的权威评分来源于"单一低层级机构”,则其可信度较低,引入置信度调节因子δ∈[0.8,1.0];其中,教育部为高可信来源,δ=1.0,不作衰减。综合推荐评分=α×用户需求匹配度+β×权威程度评分+γ×时效性评分=0.3×0.95+0.6×0.9+0.1×0.8=0.285+0.54+0.08=0.905,再根据生成推荐结果,所有候选书目按综合评分排序,取前N项生成目标书单,N由用户需求确定,并标注推荐来源与权威等级。
[0079] 在一实施例中,如图2所示,图2为本公开实施例提供的书单生成方法的流程示意图,包括基于用户画像,对需求信息进行分析(需求分析),确定需求分类,并确定需求分类匹配的权威数据源,从权威数据源中获取数据并处理(数据获取整理),构建结构化权威数据集,利用大模型对结构化权威数据集进行处理,生成书单,即前述目标书单。
[0080] 在一实施例中,通过用户画像分析模块对用户属性标签进行多维度分类,生成包含教育阶段、职业领域、知识需求类别的分类矩阵。(1)多模态数据采集子系统:基于联邦学习与迁移学习混合架构的跨平台数据融合机制。采用边缘计算节点部署分布式数据采集代理,通过差分隐私保护技术实现教育机构、职业社交平台、知识付费平台等多源异构数据的安全汇聚(2)智能清洗引擎:基于知识图谱的语义纠错模块构建领域特定的实体关系图谱,采用BERT-BiLSTM-CRF混合模型进行矛盾数据检测与修复,实现清洗策略的动态配置。(3)多维标签分类矩阵,三维动态权重分配模型,包括:教育阶段维度:采用CIPP教育评估模型划分7级梯度(K12 / 本科 / 硕士等)、职业领域维度:融合ISCO-08国际标准与行业知识图谱、知识需求维度:基于Transformer的需求强度预测模块。(4)画像生成模块:基于对抗生成网络(Generative Adversarial Networks,GAN)的个性化画像增强技术,通过生成器构建基础特征向量,判别器引入领域专家验证机制。三维特征雷达图与决策树路径的可交互式呈现输出。(5)画像分析模块:跨维度关联挖掘算法,特征组合:教育-职业匹配度分析模型、知识需求缺口检测矩阵、阶段性能力成长预测曲线等组合特征分析。
[0081] 在一实施例中,构建结构化权威数据库,区别于传统的书单仅限于推荐书目相关数据,本申请中数据表及字段规划设计在推荐书目详表的基础上重点追加设计了推荐机构表、推荐信息表、书目推荐类型、推荐统计表等数据表及字段。通过这些表及字段规划,提升了推荐结果的权威性。(1)推荐书目详表,该表记录推荐书目的详细信息,包括不限于书名、作者、出版社、出版日期等字段。(2)推荐机构表设计,该表记录推荐机构的详细信息,包括不限于推荐类型、行政级别、权威等级、机构名称、学制、省份、地市、区县、办学类型、机构类别、省内排名等字段。(3)推荐信息表,该表记录了详细的推荐信息,包括不限于推荐日期、推荐时间段、推荐阅读类型、科目汇总、科目、阅读对象、适用对象、拓展分类、书名等字段。(3)书目类型推荐表,该表记录了书目的推荐类型,包括不限于书名、科目大类、科目细类、内容分类等字段。(4)书目推荐统计表,该表记录了书目的推荐统计相关数据,包括不限于书目名称、推荐次数、推荐机构数、作者推荐次数、分类推荐次数、纸书匹配、电子书匹配等等字段。
[0082] 在一实施例中,需求分析,基于用户画像分析结果、平台用户标签数据及用户类型分析结果等数据,对用户的潜在书目需求进行分析。匹配相应的权威数据源,基于制定的书目数据表及字段规划设计,收集整理权威数据源数据。比如子女教育与成长,匹配学校、教育局及教育部等推荐书单作为满足用户需求的数据源。
[0083] 在一实施例中,权威数据源,基于需求分析结果,确认用户所需的权威数据源。比如子女教育与成长,则确认对应级别(小学 / 初中 / 高中 / 大学)的全国 / 省级TOP100名校、地区教育局、教育部等单位的推荐数据作为权威数据源。基于权威数据源的数据形态(比如公众号、官网等)开发、使用相关人工智能(Artificial Intelligence,AI)技术工具,建立权威数据获取的规则和周期等,基于制定的书目数据表及字段规划设计表获取相关数据。"机构权威等级字段"(国家级0.9 / 省级0.7 / 市级0.5)与"数据时效性标记"(近3年数据权重+20%)。
[0084] 在一实施例中,数据获取整理指的是权威数据源数据形态各异,导致数据获取整理步骤繁琐。且权威书单数据源中的数据仅具备数据表及字段规划数据中的书目详表和书目推荐信息两部分数据。包括不限于如下方式:(1)公众号,获取并存在HTML格式,python导入数据数据库。对于部分图片数据需采用OCR识别获取相应的字段。处理后的数据匹配数据表及字段规划,进行针对性的数据入库操作。(2)官网,获取并存在HTML格式,python导入数据数据库。对于部分图片数据需采用OCR识别获取相应的字段。处理后的数据匹配数据表及字段规划,进行针对性的数据入库操作。(3)用户UGC数据,制定用户UGC导入的数据规则,处理后的数据匹配数据表及字段规划,进行针对性的数据入库操作。
[0085] 在一实施例中,大模型数据处理指基于预训练大语言模型构建数据补全引擎,引擎配置有:(1)权威数据优先级规则(教育部门数据>省级名校数据>地市教育机构数据)。
[0086] (2)异构数据处理规则(HTML文本解析规则、OCR图像识别规则、UGC数据过滤规则)。
[0087] 处理流程为:原始数据格式解析(HTML / OCR / UGC)、字段对齐、权威等级标注,在处理流程中可能还包括文本清洗、图像坐标映射、语义合规过滤等操作。对原始数据进行文本清晰,筛选标题、文档名称、文档内容等包括"书单”、"推荐书单”、"书目”"推荐书目”等字段的数据。
[0088] 可以包括以下步骤:(2.1)原始数据预处理,如HTML数据处理:解析HTTP头部或HTML元标签声明的编码,未声明时采用自适应编码检测。基于HTML5标准规范构造结构化DOM树,修复标签未闭合、错误嵌套等语法问题。UGC文件(PDF / word / excel等格式)数据处理:PDF文件使用PDF格式转换文件,将PDF文件格式转换为word文件格式。word和excel格式文件保持原格式。(2.2)内容提取包括表格内容提取、文本内容提取、图片内容提取。(3)定义权威评分函数:A=Σ(w_i×l_i),其中w_i为机构权重,l_i为行政层级系数。
[0089] 在一实施例中,书单生成指通过多源数据融合模块生成三维推荐矩阵,其维度包括:用户需求维度(权重占比α)、权威等级维度(权重占比β)、时效性维度(权重占比γ),其中α+β+γ=1且β≥0.4(β=max(0.4,0.6-0.05×α),强制权威维度权重不低于40%),λ为大模型数据补全置信度系数。基于强化学习模型动态调整推荐参数,输出包含权威认证标识的个性化书单。最终评分=(α×用户需求)+(β×权威等级)+(γ×时效性)+λ×补全因子,其中,由于教育部门、学校、培训机构等等权威机构对于推荐的图书每年 / 每2年变化率不超过10%,新增推荐的图书书目基于权威机构对考试大纲的变化、考点的方向变化,有着一定的指引作用。
[0090] 时效性加权规则:按照一个学期为一个基础的时效性维度,最近的一个学期之内的数据权重最高1.1γ,按照学期时间依次递减0.01,最小不小于γ。
[0091] 在一实施例中,基础教育场景(K12),某小学五年级学生用户的书单生成过程,包含:用户标签:年龄11岁、所在地区北京市海淀区、近期搜索关键词"科普读物"。权威数据源:海淀区教委推荐书目(权重0.6)、清华附小校本阅读清单(权重0.3)。大模型补充:近三年科普类获奖图书(权重0.1)。最终输出书单包含教育部认证标识及来源说明。
[0092] 在一实施例中,职业资格考试场景,注册会计师考试备考书单生成过程,重点说明:如何融合财政部指定教材(权威系数0.7)、省级会计协会推荐教辅(权威系数0.2)、历年考生高分通过书单(权威系数0.1)的智能加权算法。
[0093] 综上,本公开提供的方案:首先,通过基于用户属性信息确定权威数据来源,可以提高书单来源的权威性;从权威数据来源获取原始推荐数据,并对原始推荐数据进行处理,生成结构化数据集;基于结构化数据集中推荐机构的类型和行政层级,确定每项书目的权威程度评分;通过根据用户需求、权威程度评分和数据更新时间,确定每项书目的综合推荐评分,并根据综合推荐评分生成目标书单,可以提高生成的书单的权威性。
[0094] 其次,通过响应于中间推荐数据包括非文本数据,将非文本数据转换为文本数据;从转换后的文本数据中提取与书目推荐关联的语义内容,实现了对多模态输入的智能适配,系统能自动识别数据类型并选择处理路径,无需人工干预,提升自动化水平与鲁棒性。
[0095] 为了实现本公开实施例提供的书单生成方法,本公开实施例还提供一种书单生成装置,如图3所示。图3为本公开实施例提供的书单生成装置的结构示意图,书单生成装置300,包括:第一确定单元301,用于基于用户属性信息确定权威数据来源;第一生成单元302,用于从权威数据来源获取原始推荐数据,并对原始推荐数据进行处理,生成结构化数据集;第二确定单元303,用于基于结构化数据集中推荐机构的类型和行政层级,确定每项书目的权威程度评分;第二生成单元304,用于基于用户需求、权威程度评分和数据更新时间,确定每项书目的综合推荐评分,并根据综合推荐评分生成目标书单。
[0096] 在一实施例中,第一确定单元301,具体用于:对用户属性信息进行分析,得到用户需求类别;确定与用户需求类别对应的权威机构层级,及将属于权威机构层级的数据来源确定为权威数据来源。
[0097] 在一实施例中,第一生成单元302,具体用于:对原始推荐数据进行格式统一化处理,得到中间推荐数据;从中间推荐数据中提取与书目推荐关联的语义内容;基于语义内容,生成结构化数据集。
[0098] 在一实施例中,第一生成单元302,具体用于:响应于中间推荐数据包括非文本数据,将非文本数据转换为文本数据;从转换后的文本数据中提取与书目推荐关联的语义内容。
[0099] 在一实施例中,第二生成单元304,包括:确定数据可信度权重和目标权重;基于数据可信度权重、目标权重、用户需求、权威程度评分和数据更新时间,确定每项书目的综合推荐评分。
[0100] 在一实施例中,目标权重包括用户需求权重、权威程度评分权重和数据更新时间权重,其中,数据更新时间权重由距离当前时间的时间间隔确定。
[0101] 需要说明的是:上述实施例提供的书单生成装置在进行书单生成时,仅以上述各程序模块的划分进行举例说明,实际应用中,可以根据需要而将上述处理分配由不同的程序模块完成,即将书单生成装置的内部结构划分成不同的程序模块,以完成以上描述的全部或者部分处理。另外,上述实施例提供的书单生成装置与本公开实施例提供书单生成方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
[0102] 图4为本公开实施例提供的电子设备的硬件组成结构示意图,如图4所示,电子设备400包括至少一个处理器402;以及与至少一个处理器402通信连接的存储器401;其中,存储器401存储有可被至少一个处理器402执行的指令,指令被至少一个处理器402执行,以实现本公开实施例的书单生成方法的步骤。
[0103] 可选地,该电子设备具体可为本申请实施例的书单生成装置,并且该电子设备可以实现本申请实施例的各个方法中由书单生成装置实现的相应流程,为了简洁,在此不再赘述。
[0104] 可理解,电子设备中还包括通信接口403。电子设备中的各个组件通过总线系统404耦合在一起。可理解,总线系统404用于实现这些组件之间的连接通信。总线系统404除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图4中将各种总线都标为总线系统404。
[0105] 可以理解,存储器401可以是易失性存储器或非易失性存储器,也可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(ROM,Read Only Memory)、可编程只读存储器(PROM,Programmable Read-Only Memory)、可擦除可编程只读存储器(EPROM,Erasable Programmable Read-Only Memory)、电可擦除可编程只读存储器(EEPROM,Electrically Erasable Programmable Read-Only Memory)、磁性随机存取存储器(FRAM,ferromagnetic random access memory)、快闪存储器(Flash Memory)、磁表面存储器、光盘、或只读光盘(CD-ROM,Compact Disc Read-Only Memory);磁表面存储器可以是磁盘存储器或磁带存储器。易失性存储器可以是随机存取存储器(RAM,Random AccessMemory),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(SRAM,Static Random Access Memory)、同步静态随机存取存储器(SSRAM,Synchronous Static Random Access Memory)、动态随机存取存储器(DRAM,Dynamic Random Access Memory)、同步动态随机存取存储器(SDRAM,SynchronousDynamic Random Access Memory)、双倍数据速率同步动态随机存取存储器(DDRSDRAM,Double Data Rate Synchronous Dynamic Random Access Memory)、增强型同步动态随机存取存储器(ESDRAM,Enhanced Synchronous Dynamic Random Access Memory)、同步连接动态随机存取存储器(SLDRAM,SyncLink Dynamic Random Access Memory)、直接内存总线随机存取存储器(DRRAM,Direct Rambus Random Access Memory)。本发明实施例描述的存储器401旨在包括但不限于这些和任意其它适合类型的存储器。
[0106] 上述本公开实施例揭示的方法可以应用于处理器402中,或者由处理器402实现。处理器402可能是一种集成电路芯片,具有信号的处理能力。在实现过程中,上述方法的各步骤可以通过处理器402中的硬件的集成逻辑电路或者软件形式的指令完成。上述的处理器402可以是通用处理器、DSP,或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。处理器402可以实现或者执行本发明实施例中的公开的各方法、步骤及逻辑框图。通用处理器可以是微处理器或者任何常规的处理器等。结合本发明实施例所公开的方法的步骤,可以直接体现为硬件译码处理器执行完成,或者用译码处理器中的硬件及软件模块组合执行完成。软件模块可以位于存储介质中,该存储介质位于存储器401,处理器402读取存储器401中的信息,结合其硬件完成前述方法的步骤。
[0107] 在示例性实施例中,电子设备可以被一个或多个应用专用集成电路(ASIC,Application Specific Integrated Circuit)、DSP、可编程逻辑器件(PLD,ProgrammableLogic Device)、复杂可编程逻辑器件(CPLD,Complex Programmable Logic Device)、FPGA、通用处理器、控制器、MCU、微处理器(Microprocessor)、或其他电子元件实现,用于执行前述方法。
[0108] 本公开实施例还提供了一种存储有计算机指令的非瞬时计算机可读存储介质,计算机指令用于使计算机执行时实现本发明实施例的书单生成方法的步骤。
[0109] 可选的,该计算机可读存储介质可应用于本申请实施例中的书单生成装置,并且该计算机指令使得计算机执行本申请实施例的各个方法中由书单生成装置实现的相应流程,为了简洁,在此不再赘述。
[0110] 本公开实施例还提供了一种计算机程序产品,包括计算机程序,计算机程序在被处理器执行时实现本发明实施例提供的书单生成方法的步骤。
[0111] 在本申请所提供的几个实施例中,应该理解到,所揭露的设备和方法,可以通过其它的方式实现。以上所描述的设备实施例仅仅是示意性的,例如,单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,如:多个单元或组件可以结合,或可以集成到另一个系统,或一些特征可以忽略,或不执行。另外,所显示或讨论的各组成部分相互之间的耦合、或直接耦合、或通信连接可以是通过一些接口,设备或单元的间接耦合或通信连接,可以是电性的、机械的或其它形式的。
[0112] 上述作为分离部件说明的单元可以是、或也可以不是物理上分开的,作为单元显示的部件可以是、或也可以不是物理单元,即可以位于一个地方,也可以分布到多个网络单元上;可以根据实际的需要选择其中的部分或全部单元来实现本实施例方案的目的。
[0113] 另外,在本发明各实施例中的各功能单元可以全部集成在一个处理单元中,也可以是各单元分别单独作为一个单元,也可以两个或两个以上单元集成在一个单元中;上述集成的单元既可以采用硬件的形式实现,也可以采用硬件加软件功能单元的形式实现。
[0114] 本领域普通技术人员可以理解:实现上述方法实施例的全部或部分步骤可以通过程序指令相关的硬件来完成,前述的程序可以存储于一计算机可读取存储介质中,该程序在执行时,执行包括上述方法实施例的步骤;而前述的存储介质包括:移动存储设备、ROM、RAM、磁碟或者光盘等各种可以存储程序代码的介质。
[0115] 或者,本发明上述集成的单元如果以软件功能模块的形式实现并作为独立的产品销售或使用时,也可以存储在一个计算机可读取存储介质中。基于这样的理解,本发明实施例的技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机、服务器、或者网络设备等)执行本发明各个实施例方法的全部或部分。而前述的存储介质包括:移动存储设备、ROM、RAM、磁碟或者光盘等各种可以存储程序代码的介质。
[0116] 以上仅为本发明的具体实施方式,但本发明的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本发明揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本发明的保护范围之内。因此,本发明的保护范围应以权利要求的保护范围为准。< / script> < / main> < / article>
Claims
1. A book list generation method characterized by comprising: The method comprises: determining an authoritative data source based on user attribute information; obtaining original recommendation data from the authoritative data source and processing the original recommendation data to generate a structured data set; determining an authoritative degree score of each book based on the type and administrative level of a recommendation institution in the structured data set; determining a comprehensive recommendation score of each book based on user demand, the authoritative degree score and data update time, and generating a target book list according to the comprehensive recommendation score.
2. The method of claim 1, wherein, The method of determining an authoritative data source based on user attribute information comprises: analyzing the user attribute information to obtain a user demand category; determining an authoritative institution level corresponding to the user demand category, and determining a data source belonging to the authoritative institution level as the authoritative data source.
3. The method of claim 1, wherein, The method of processing the original recommendation data to generate a structured data set comprises: performing format unification processing on the original recommendation data to obtain intermediate recommendation data; extracting semantic content associated with book recommendation from the intermediate recommendation data; generating the structured data set based on the semantic content.
4. The method of claim 3, wherein, The method of extracting semantic content associated with book recommendation from the intermediate recommendation data comprises: in response to the intermediate recommendation data including non-text data, converting the non-text data into text data; extracting semantic content associated with book recommendation from the converted text data.
5. The method of claim 1, wherein, The method of determining a comprehensive recommendation score of each book based on user demand, the authoritative degree score and data update time comprises: determining a data credibility weight and a target weight; determining a comprehensive recommendation score of each book based on the data credibility weight, the target weight, user demand, the authoritative degree score and data update time.
6. The method of claim 5, wherein, The target weight comprises a user demand weight, an authoritative degree score weight and a data update time weight, wherein the data update time weight is determined by a time interval from the current time.
7. A book list generating apparatus characterized by comprising: The method comprises: a first determination unit configured to determine an authoritative data source based on user attribute information; a first generation unit configured to obtain original recommendation data from the authoritative data source and process the original recommendation data to generate a structured data set; a second determination unit configured to determine an authoritative degree score of each book based on the type and administrative level of a recommendation institution in the structured data set; a second generation unit configured to determine a comprehensive recommendation score of each book based on user demand, the authoritative degree score and data update time, and generate a target book list according to the comprehensive recommendation score.
8. An electronic device, comprising: The method comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1 to 6.
10. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 6.