Electronic archive personalized recommendation method and system based on user behavior

By constructing document jump sequences and semantic path fragments based on user access order, the problem of content fragmentation when the topic spans a large range is solved in existing technologies. This achieves the coherence and structural consistency of document paths, and improves the effect of personalized recommendations for electronic archives.

CN122432419APending Publication Date: 2026-07-21TONGLUE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGLUE TECH CO LTD
Filing Date
2026-06-08
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies are prone to content fragmentation and broken recommendation chains when processing user behavior sequences with large thematic spans, failing to effectively support the complete presentation of the hierarchical structure in the document system and affecting the efficiency and depth of users' continuous exploration of relevant archival resources.

Method used

By constructing a document jump sequence based on user access order, semantic path fragments are extracted. Combining content feature transformation and structural continuity, path segments with consistent performance are selected, topic positions are marked, duplicate content is excluded, the arrangement logic of recommended content is optimized, and the semantic connection between documents and the overall consistency of path arrangement are enhanced.

Benefits of technology

It has achieved the construction of document chains with coherent content and stable structure, which has improved the overall expressiveness of topic derivation, structural continuity and information organization in the document path, and enhanced the logical coherence of recommended content and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432419A_ABST
    Figure CN122432419A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of information retrieval, in particular to an electronic archive personalized recommendation method and system based on user behavior, which comprises obtaining access records and constructing semantic jump segments, aggregating structure continuous content to form convergence chain groups, marking levels and extracting jump sections to generate cross-layer paths, supplementing unvisited documents to form extension lists, and sorting to retain theme direction to generate personalized recommendation content. The present application constructs document jump sequences under access sequence and extracts semantic path segments, combines structure continuity to aggregate and exclude content clutter paragraphs, marks theme positions based on directory levels and extracts jump paragraphs, filters un-presented documents at the tail of the path to retain theme consistent parts, constructs document chains with stable structure, strengthens semantic connection and path consistency, optimizes the arrangement logic of recommended content, and enhances the expressiveness of theme derivation, structure continuation and information organization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval technology, and in particular to a method and system for personalized recommendation of electronic archives based on user behavior. Background Technology

[0002] The field of information retrieval technology encompasses the research and application of expressing information needs, acquiring relevant data, and presenting information. Its core components include key aspects such as query understanding, document modeling, retrieval algorithm design, and matching strategy optimization. Information retrieval not only relies on keyword matching but also integrates user behavior modeling, semantic understanding, and machine learning, gradually evolving towards personalization and intelligence. In environments with frequent user-system interactions, information retrieval systems need to dynamically respond to changes in user interests, analyze historical behavior, and achieve more efficient information matching and recommendation capabilities to meet the diverse information needs of different users.

[0003] Personalized electronic archive recommendation methods based on user behavior involve analyzing user search history, browsing behavior, and archive utilization patterns, combined with collaborative filtering and content mining algorithms, to construct a model reflecting users' potential interests and preferences, and then proactively pushing electronic archive resources accordingly. This mainly encompasses processes such as user behavior sequence extraction, interest feature vector construction, matching calculations of similar users or similar content, and dynamic updating of the recommended archive set. The entire recommendation process relies on historical behavioral data for reasoning and correlation analysis, typically employing methods such as behavioral log parsing, content attribute annotation, and similarity ranking to select and present personalized resources.

[0004] Existing technologies are mostly based on behavior frequency and similarity calculations, ignoring the semantic jump features and structural coherence between documents during user visits. When dealing with behavior sequences with large topic spans, content fragmentation problems are prone to occur, resulting in a lack of logical continuity between recommended content and the topics that users are currently interested in. In scenarios with long browsing paths or frequent topic shifts, it is easy to cause the recommendation chain to break, which cannot effectively support the complete presentation of the hierarchical structure in the document system and affects the efficiency and depth of users' continuous exploration of relevant archival resources. Summary of the Invention

[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a method and system for personalized recommendation of electronic records based on user behavior. The technical solution is as follows: A personalized recommendation method for electronic records based on user behavior includes the following steps: S1: Obtain the user's access records in the archive platform, organize the behavior sequence according to the user's identity and rearrange the archive behavior of each user, record the jump combination of adjacent archives, extract the content change direction between jumps, construct the access behavior chain, and obtain a set of semantic jump path fragments; S2: Call each path segment in the semantic jump path fragment set, extract the content expression mode of the document within the path, perform structural judgment on the content performance of the document in sequence, filter out the path segments with consistent continuous performance, remove the parts with inconsistent topic directions, and obtain the convergent semantic behavior chain group. S3: Call the document content in the convergent semantic behavior chain group, retrieve the topic identifier of the document in the directory structure, mark the topic belonging position of each document in the path, identify paragraphs with hierarchical changes, extract the parts with clear topic spans separately, and obtain the cross-level belonging topic path set; S4: Call the document number in the cross-layer subject path set, compare it with the file number accessed by the user, exclude duplicate number content, retain the document record of the first appearance, and organize it into an updated document chain according to its subject path identifier to obtain a candidate subject extension document list.

[0006] As a further aspect of the present invention, the file access sequence clues include user access identity sequence, document jump logic relationship, jump content feature transformation fragments, the convergent semantic behavior chain group includes adjacent content description structure, semantic concentrated expression segment, and structural continuity fragment set, the cross-layer affiliation topic path set includes topic level jump marker, continuous behavior topic span segment, and document topic level mapping, and the candidate topic extended document list includes newly added unread document number, topic path completion information, and extended chain structure document set.

[0007] As a further aspect of the present invention, the step of obtaining S1 is as follows: S101: Obtain all access records of the user in the archive platform, and aggregate the user identity field and access timestamp field in the records. Arrange the access behavior of the same user in chronological order to build a continuous document access behavior chain and generate a user time-series access sequence set. S102: Based on the user time-series access sequence set, extract the adjacent document access pairs for each user, call the document content feature vector and access time interval data, and perform joint filtering based on the feature vector difference and time interval threshold to obtain a time-constrained jump combination set; S103: Call the content feature vector sequence in the time-constrained jump combination set, extract the feature change information in the document jump path, and perform sequence splicing and semantic consistency judgment on the continuous feature change path to establish a semantic jump path fragment set.

[0008] As a further aspect of the present invention, the process of jointly filtering based on feature vector differences and time interval thresholds limits the temporal continuity of document jump behavior by setting a time interval threshold range. The determination of the difference in feature vectors is based on the distance calculation results between the feature vectors of document content, and combined with the preset feature similarity judgment criteria to determine whether the screening conditions are met. The adjacent document access pairs included in the time-constrained jump combination set must simultaneously meet two conditions: the time interval does not exceed the time interval threshold and the difference in content feature vectors does not exceed the feature similarity judgment criteria. During the establishment of the semantic jump path fragment set, only document jump path combinations that satisfy the coherent trend of content feature changes are retained.

[0009] As a further aspect of the present invention, the step of obtaining S2 is as follows: S201: Call each access path in the semantic jump path fragment set, extract the content description data frame corresponding to each document in the path, perform structural comparison of the content description between adjacent documents, determine the continuous relationship of the description structure within the path range, perform annotation processing based on the consistency of structural order position, and generate a structural continuous tag matrix. S202: Based on the structural continuity marker matrix, retrieve the document fragment sequence that exhibits a structural continuity state in the path, and perform position merging processing on the content description vector corresponding to the sequence. Then, aggregate and encode multiple structural continuity fragments merged within the same path according to the adjacent structural dimensions to obtain a set of structural aggregated fragments. S203: For the structured aggregated fragment set, call the content description span parameter of each fragment, and perform joint screening on the span parameter and content similarity index to remove combined fragments with abrupt span changes and semantic discrepancies, retain path segments that are structurally continuous and content-convergent, and generate convergent semantic behavior chain groups.

[0010] As a further aspect of the present invention, the step of obtaining S3 is as follows: S301: Call all document content in the convergent semantic behavior chain group, retrieve the corresponding topic level position of each document in the archive directory, bind the topic level with the document index, establish the topic level index matrix of the path document, and generate the path topic level annotation set. S302: Based on the path topic level annotation set, call the document level identifier in each path group, sequentially scan the level index values ​​of adjacent documents, detect positions where there are discontinuous changes, and extract the document paragraph index with jump characteristics to obtain a set of level jump paragraphs. S303: Based on the set of hierarchical jump paragraphs, extract the topic information of the corresponding documents and their position numbers in the path, and classify them in combination with the complete structure of the path, mark the document segments that show topic span changes in continuous behavior paths, and establish a set of cross-level affiliation topic paths.

[0011] As a further aspect of the present invention, the step of obtaining S4 is as follows: S401: Call all document numbers in the cross-layer subject path set, obtain the list of files that the user has completed accessing, perform number matching operation on the two sets of numbers, mark the document records with overlapping numbers, and exclude them from the path set to generate a set of unaccessed document numbers. S402: Based on the set of unaccessed document numbers, retrieve the document content associated with the corresponding number, extract the topic path location registered in the original archive directory, construct a one-to-one mapping table between documents and topic path locations, and establish a set of unaccessed document topic mappings; S403: Based on the path structure order in the unvisited document topic mapping set, sort all unvisited document numbers hierarchically by path, and aggregate document records under each topic position in sequence to construct a structurally continuous document sequence path and generate a candidate topic extension document list.

[0012] As a further aspect of the present invention, the method further includes: S5: Call the document set in the candidate topic extended document list, arrange the documents according to their path order, identify the unvisited part at the end of the path, retain the document paragraphs that are in the same order and have the same direction as the output sequence, and obtain the personalized recommendation content of electronic archives. The personalized recommendations for electronic archives include paragraphs arranged in structural order, segments displayed with thematic consistency, and document completion at the end of the path.

[0013] As a further aspect of the present invention, the step of obtaining S5 is as follows: S501: Call the document set in the candidate topic extended document list, rearrange the document content according to the order index parameter of each document in the chain path, construct a linear content sequence structure, perform position alignment processing on it, and generate a chain content sequence set; S502: Based on the chained content sequence set, extract the sequence fragments at the end of the path, compare the position index of the fragments in the content sequence, filter the document paragraphs that have not yet appeared in the complete path, and classify them in combination with the sequential number to obtain a set of newly added paragraphs at the end of the path. S503: Based on the document content characteristics in the newly added paragraph set at the end of the path, retrieve its topic tags, compare them with the topic direction vectors in the preceding path, filter out segments with inconsistent directions, and retain consistent content in the original order to establish personalized recommendation content for electronic archives.

[0014] A personalized recommendation system for electronic records based on user behavior, the system comprising: The behavior jump relationship module obtains user access records and extracts user identifiers, access times and document numbers. It constructs an operation sequence based on time sorting, calls the summaries, types and category tags of adjacent documents, judges the changes in jump content, marks the jump combination type, reorganizes it into a chain structure, and generates a set of semantic jump path fragments. The path convergence aggregation module calls the document paths in the set of semantic jump path fragments, extracts the content description of each document, determines whether the adjacent description structures are continuous, retains the paths with concentrated structures, removes paragraphs with chaotic jump spans, aggregates fragments with consistent structures, and generates convergent semantic behavior chain groups. The hierarchical span extraction module calls the document number in the convergent semantic behavior chain group, retrieves the topic level position in its directory, compares the adjacent level numbers according to the path order, determines whether there is a hierarchical jump behavior, extracts and marks the segments with large spans, and generates a cross-level belonging topic path set. The extended document construction module calls the document number in the cross-layered subject path set, compares it with the user access file list, removes duplicate numbers, organizes unaccessed documents and matches their path level positions, reorganizes the chain structure according to the path order, and generates a candidate subject extended document list. The personalized recommendation output module calls the candidate topic extended document list, extracts document content and path order information, filters documents that have not yet appeared at the end of the path, retains paragraph content with the same topic direction, rearranges the structural path for display, and generates personalized electronic archive recommendation content.

[0015] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this invention, a document jump sequence based on access order is constructed, semantic path fragments are extracted by combining content feature transformation, and behavior segments with chaotic content spans are eliminated by using structural continuity aggregation. Furthermore, the topic positions are marked based on directory hierarchy information and jump paragraphs are extracted. Unseen documents are filtered at the end of the path and the parts with consistent themes are retained. This completes the construction of a document chain with coherent content and stable structure, enhances the semantic connection between documents and the overall consistency of path arrangement, optimizes the arrangement logic of recommended content through semantic trend recognition and cross-layer topic aggregation, and improves the comprehensive expressiveness of topic derivation, structural continuation and information organization in the document path. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart illustrating the process of obtaining the semantic jump path fragment set in this invention. Figure 3 This is a flowchart illustrating the process of obtaining convergent semantic behavior chain groups in this invention. Figure 4This is a flowchart illustrating the process of obtaining the cross-layer affiliation topic path set in this invention. Figure 5 Flowchart for obtaining the list of extended documents for candidate topics of this invention; Figure 6 This is a flowchart illustrating the process of obtaining personalized recommendation content for electronic archives in this invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] Please see Figure 1 This invention provides a technical solution: a personalized recommendation method for electronic records based on user behavior, comprising the following steps: S1: Obtain all user access records in the archive platform, organize the behavior sequence according to user identity, rearrange the archive operation order of each user according to time order, record the jump combination between adjacent documents, establish a sequence chain based on the changes in the content characteristics of document jumps, generate archive access order clues, and obtain a set of semantic jump path fragments; S2: Call each access path in the semantic jump path fragment set, extract the content description of each document in the path, judge whether the behavior shows a convergent tendency according to the adjacent structure of the content description, aggregate the fragments that show structural continuity in the same path, remove the parts with messy content span, and obtain the convergent semantic behavior chain group. S3: Call all document content in the convergent semantic behavior chain group, retrieve the topic level position of the document in the archive directory, mark the document level position in the same path item by item, extract paragraphs with hierarchical jump characteristics, mark the document segments with prominent topic span in continuous behavior, and obtain the cross-level belonging topic path set. S4: Call the document numbers involved in the cross-layer belonging topic path set, compare them with the list of archives that the user has completed accessing, exclude records with overlapping numbers, centrally organize the documents that have not appeared in the behavior record, and re-match their corresponding topic path positions to build an updated document chain and obtain a candidate topic extension document list. S5: Call the document set in the candidate topic extended document list, sort the content according to the order of the documents in the chain path, identify the document paragraphs that have not yet appeared at the end of the path, and retain the parts that are consistent with the topic direction of the previous section from the path structure, and form a continuous display result in order to obtain personalized electronic archive recommendation content.

[0023] The archive access sequence clues include user access identity sequence, document jump logic relationship, and jump content feature change fragments; the convergent semantic behavior chain group includes adjacent content description structure, semantically concentrated expression segment, and structural continuity fragment set; the cross-level belonging topic path set includes topic level jump marker, continuous behavior topic span segment, and document topic level mapping; the candidate topic extended document list includes newly added unread document number, topic path completion information, and extended chain structure document set; the personalized electronic archive recommendation content includes structurally ordered paragraphs, topic consistency display fragments, and path tail completion documents.

[0024] Please see Figure 2 The steps to obtain S1 are as follows: S101: Obtain all access records of the user in the archive platform, and aggregate the user identity field and access timestamp field in the records. Arrange the access behavior of the same user in chronological order to build a continuous document access behavior chain and generate a user time-series access sequence set. To obtain all user access records in the archive platform, the raw data of the fields "User Identifier," "Unique Document Number," and "Access Timestamp" must first be extracted from the log database. By traversing the access records in the database, access events corresponding to different users are grouped. The data table is then partitioned based on the "User Identifier," constructing an access event set for each user. For the timestamp field of each user's access event set, the timestamps are sorted in ascending order, and the order in which each user accessed documents is arranged chronologically to generate a document access sequence. For example, if user A accessed records with document IDs D01, D03, and D05 in sequence, with access timestamps of 12:01, 12:03, and 12:10 respectively, then the organized chronological access sequence for user A would be [D01, D03, ...]. [D05] In this process, it is also necessary to determine whether there are duplicate or abnormal values ​​in the user identity field. For example, whether there is a situation where the same user accesses different documents multiple times in the same time period. For such conflicting behaviors, rules can be set to retain the first access record and discard the rest. In this way, the time series access behavior chain corresponding to each user is finally constructed to obtain the user time series access sequence set.

[0025] S102: Based on the user time-series access sequence set, extract the adjacent document access pairs of each user, call the content feature vector of the document and the access time interval data, and perform joint filtering based on the feature vector difference and the time interval threshold to obtain the time-constrained jump combination set; Based on the user's temporal access sequence set, adjacent document access pairs are extracted for each user. For each accessed document pair, a document content feature vector is extracted. For example, the vector can be obtained based on TF-IDF encoding or embedded semantic representation. The content feature vector of each document is denoted as... The access time interval is represented by the difference in timestamps between two adjacent document accesses. The two content feature vectors in the document pair and Euclidean or cosine distance between As a basis for content differences, a standard for judging feature similarity is also set. With time interval threshold For each document access pair, execute the conditional judgment. and If both conditions are met, the document access pair is considered a valid redirect, for example, by setting... , If a user's access time is 450 seconds between two consecutive visits, and the vector difference between the two documents is... If a document meets the criteria, it is retained; otherwise, it is discarded. This process involves setting the threshold. The value of can be obtained by analyzing the mean of the differences between all document feature vectors. and standard deviation ,make For example, if the average difference among all documents is 0.45 and the standard deviation is 0.15, then... Set to 0.30, time interval threshold The average access frequency of user behavior can be used as a reference to determine this. For example, if the average interval between document accesses by platform users is 500 seconds, then... It can be set to 600 seconds, and finally all jump pairs that satisfy the two conditions are retained to construct a set of time-constrained jump combinations.

[0026] S103: Call the content feature vector sequence in the time-constrained jump combination set, extract the feature change information in the document jump path, and perform sequence splicing and semantic consistency judgment on continuous feature change paths to establish a semantic jump path fragment set. The system retrieves the content feature vector sequence from the time-constrained jump combination set, extracts the features of each vector in the jump path sequentially, and records the direction and magnitude of feature changes involved in each jump. For example, if the vector change from document A to B is represented as... The change from B to C is as follows: The path change sequence is then... Determine if there is a consistent trend of change along the path, for example... and If the direction is the same in most dimensions and the change is not drastic, the sign and magnitude difference of the feature change value in each dimension are compared for judgment. If the direction is consistent in more than 80% of the dimensions and the magnitude difference is within 0.1, the two jumps are considered to have semantic consistency. Then, a splicing operation is performed to connect multiple consecutive jump paths into a complete path sequence. For example, if the path from document A→B→C meets the consistency condition, it is spliced ​​into path A→B→C. If there is a sudden change in direction or an abnormal change in magnitude in a certain jump segment, such as a change in a dimension from positive to negative or a change in magnitude greater than 0.3, the path splicing is interrupted and the path ends. For all spliced ​​path sequence sets, a duplicate screening is performed to retain jump paths with consistent feature change trends and conforming to continuous change patterns. Finally, a set of semantic jump path fragments is established.

[0027] Please see Figure 3 The steps to obtain S2 are as follows: S201: Call each access path in the semantic jump path fragment set, extract the content description data frame corresponding to each document in the path, perform structural comparison of the content description between adjacent documents, determine the continuous relationship of the description structure within the path range, perform annotation processing based on the consistency of structural order position, and generate a structural continuous tag matrix. The system calls each access path in the semantic jump path fragment set, sequentially extracting the document numbers contained in each access path. Based on the document number, it retrieves the corresponding content description data frame. Each content description data frame uses fixed fields to represent semantic blocks such as the document's title, paragraph tags, keyword tags, and summary paragraphs. Each field has a clear start and end number. After extracting the data frames of all documents in the path, a field comparison operation is performed on any two adjacent documents in the access path. If the corresponding fields have field names, tag identifiers, and content structure positions that are completely identical, they are determined to be structurally consistent. After performing the comparison operation on all fields, the proportion of consistent fields is counted. If the consistency proportion exceeds 75%, it is recorded as a structurally continuous state. In a real-world scenario, if the first document description field sequence in the path is [title="A1", paragraph tag="L1", paragraph content="B1"], and the next document is [title="A2", paragraph tag="L1", ...], ... If the segment content = "B2", then since the segment labels are consistent and the structural positions are the same, it can be determined that the structure is continuous. In the comparison operation, it is necessary to compare whether the field names are completely consistent, whether the field position indexes match, and whether the semantics of the field content have the same type of label. After each group of paths completes the structural comparison, the mark value "1" is set for the structurally consistent field positions according to the comparison results, and the structurally inconsistent fields are marked as "0". The mark results are mapped to the document positions in the access path to form a two-dimensional Boolean mark matrix. The rows of this matrix represent the path sequence number, the columns represent the document position index, and the values ​​are whether the structure is continuous. Finally, the structural continuity mark matrix is ​​output.

[0028] S202: Based on the structural continuity marker matrix, retrieve the document fragment sequence that exhibits a structural continuity state in the path, and perform position merging processing on the content description vector corresponding to the sequence. Then, aggregate and encode multiple structural continuity fragments merged within the same path according to the adjacent structural dimensions to obtain a set of structural aggregated fragments. Based on the continuous regions marked with a value of "1" in the structural continuous label matrix, each access path is scanned, and documents continuously marked with "1" are grouped into a structural continuous segment. For example, if the corresponding position in a certain row of the path structural label matrix is ​​[1,1,0,1,1,1], then two structural continuous segments can be identified: [Document 1, Document 2] and [Document 4, Document 5, Document 6]. For the content description data frames involved in each segment, the corresponding description vector sequence is extracted, for example, represented in TF-IDF vector format, with a fixed vector dimension of d. After concatenating the document vectors in each structural continuous segment in order, a position merging operation is performed based on the vector index position. The position merging operation refers to merging multiple document vectors in the sequence along the same dimension. The values ​​at each position are weighted averaged. For example, if the vector values ​​of three documents in the i-th dimension are 0.3, 0.5, and 0.4 respectively, the averaged value is 0.4. The merged vector is used as the total vector representation of the structural fragment. Adjacency aggregation is performed on the merged vectors of all structurally continuous fragments in the same path. Fragments that are continuous in position and have consistent field structure encoding are concatenated, and the concatenated vector is retained as the encoding representation of the new fragment. During the aggregation process, if the last document position of two fragments differs from the first document position of another fragment by no more than 1, they are considered to be continuous in position. Field structure encoding refers to the identical numbering sequence of fields of the same type in the content description field after concatenation. For example, if two fragments are both combinations of [title, section, paragraph], they are considered to have the same structure. Finally, the set of all merged and aggregated fragments that satisfy both positional continuity and consistent structure is output.

[0029] S203: For the set of structurally aggregated fragments, call the content description span parameter of each fragment, and jointly filter the span parameter and content similarity index to remove combined fragments with abrupt span changes and semantic discrepancies, retain path segments that are structurally continuous and content-convergent, and generate convergent semantic behavior chain groups. For each fragment in the structured aggregate fragment set, extract its corresponding content description span parameter. The span parameter refers to the index spacing between the first and last documents within the fragment in the access path, for example, the path length is... The starting document index of the fragment is Ending The span is Simultaneously, the semantic description vector corresponding to each segment is extracted, and the semantic similarity index between the semantic description vector and other segment vectors is calculated. Similarity can be measured using cosine similarity. The criteria for determining the similarity index value are as follows: If the span of any segment exceeds the set span threshold Or, its similarity index is less than the average similarity of other segments in the path. If a segment does not meet the criteria, it should be removed. For example, if a segment has a span of... And the average similarity is Since neither of the two conditions is met, it is removed from the set of structurally aggregated fragments, subject to the span threshold. The settings can be referenced from the average span of all aggregated segments. with standard deviation The setting method is as follows: ,like , ,but Rounding to By combining the judgment logic to perform a joint filtering operation, only the set of fragments with a reasonable span and semantic similarity higher than the set threshold is retained, resulting in path segments with continuous structure and similar content, and outputting a chain group of similar semantic behaviors.

[0030] Please see Figure 4 The steps to obtain S3 are as follows: S301: Call all document content in the convergent semantic behavior chain group, retrieve the corresponding topic level position of each document in the archive directory, bind the topic level with the document index, establish the topic level index matrix of the path document, and generate the path topic level annotation set; The algorithm retrieves the content of all documents in the convergent semantic behavior chain group, obtains the topic path position of each document in the archive directory, and performs a topic level retrieval operation on each document to determine its topic level number. If the archive directory adopts a three-level hierarchical structure, each document can correspond to a topic level number such as (1,2,4), where the first level represents the first-level topic index, the second level represents the sub-topic, and the third level represents the specific document category position. After extraction, the index position of the document in the behavior path is combined and bound with its topic level number. For example, if document A is the 3rd document in the path and its corresponding topic level position is (2,1,5), it is recorded as (3,2,1,5). After performing this operation on all documents, all documents in each path and their topic level binding values ​​are arranged in the path order to construct a multi-dimensional array, where the rows represent the path sequence number, the columns represent the document index, and the cell value is the corresponding topic level number, forming a topic level index matrix of path documents. This matrix structure is used to describe the mapping relationship between the structural position of a document in the path and its semantic attribution. Finally, the path topic level annotation set corresponding to each path is output.

[0031] S302: Based on the path topic level annotation set, call the document level identifier in each path, sequentially scan the level index values ​​of adjacent documents, detect the positions where there are discontinuous changes, and extract the document paragraph index with jump characteristics to obtain a set of level jump paragraphs. Based on the path topic hierarchy annotation set, the hierarchical identifier sequence of documents is extracted from each path. The numerical numbers representing topic categories in the hierarchical identifiers are scanned in document order. The difference in hierarchical numbers between adjacent documents is defined as the jump criterion. When performing jump detection, the absolute difference between the hierarchical number of the current document and the number of the previous document is calculated. If the difference exceeds a set threshold ΔL, it is determined that a hierarchical jump has occurred. The threshold ΔL can be set to 1, indicating that adjacent documents are allowed to cross at most one topic. For example, if the hierarchical number of the previous document is (2,3,1) and the current document is (2,5,1), the difference is 2, which exceeds the threshold, so the jump event is recorded. After performing the above jump detection on each path, the index numbers of all document paragraphs that have jumped are extracted and recorded. For example, if document indexes [3→4] and [6→7] in path A jump, the jump paragraph index is output as [4,7]. The jump paragraph set only retains the document indexes that have jumped, as an important structural identifier of hierarchical changes, and finally generates a hierarchical jump paragraph set.

[0032] S303: Based on the set of hierarchical jump paragraphs, extract the topic information of the corresponding documents and their position numbers in the path, and classify them in combination with the complete structure of the path. Mark the document segments that show the topic span change in the continuous behavior path and establish a cross-level topic path set. Based on the hierarchical jump paragraph set, extract the document content corresponding to each jump paragraph from the convergent semantic behavior chain group, obtain its topic information fields, such as the topic name and topic tag number, and record the document's position number in the original path. For example, if document B is the 5th position in path B and its topic number is (3,2,1), then it is recorded as (path B, index 5, (Topic 3.2.1) This recording operation is performed sequentially for all skipped paragraphs, and the skipped paragraphs are clustered and organized according to the path identifier to form a path-level skipped document set. During the organization process, the difference between the topic number of the skipped paragraph and the topic number of the preceding and following documents in the path is calculated. If the topic number changes in the main category dimension (i.e., the first level number), the paragraph is marked as a cross-topic skip. If it only changes in the sub-category dimension (the second or third level number), it is recorded as a detail skip. After classification, the path structure is merged. Paragraphs with multiple main category skips in the same path are aggregated to form a topic span interval, and their starting index and ending index number are recorded. For example, if document 3 to document 6 in path C are both main category skipped paragraphs, then the interval [3,6] is formed. Finally, the set of the belonging topics and their position numbers of all cross-level paragraphs is established by path, and the cross-level belonging topic path set is output.

[0033] Please see Figure 5 The steps to obtain S4 are as follows: S401: Call all document numbers in the cross-layer subject path set, obtain the list of files that the user has completed accessing, perform number matching operation on the two sets of numbers, mark the document records with overlapping numbers, exclude them from the path set, and generate a set of unaccessed document numbers. To retrieve all document IDs from the cross-level subject path set, first extract all document IDs from this path set into set A. Then, retrieve the list of archives accessed by the current user and extract the document IDs from it into set B. Perform an ID matching operation between set A and set B, that is, for each document ID in set A... The system checks if a document number exists in set B. If it does, it marks it as a duplicate. The matching operation is achieved by comparing the number strings one by one to see if they are completely identical. For example, if set A contains document numbers [101, 102, 105, 110] and set B contains document numbers [100, 102, 103], then number 102 is a duplicate. The system then removes number 102 from set A, generating a new set C = [101, 105, 110]. This set is the set of document numbers that the user has not visited. The removal process sets a flag for all matching numbers, removes numbers with the flag "visited", and constructs a new set of numbers, ultimately forming the set of unvisited document numbers.

[0034] S402: Based on the set of unaccessed document numbers, retrieve the document content associated with the corresponding number, extract the subject path location registered in the original archive directory, construct a one-to-one mapping table between documents and subject path locations, and establish a set of unaccessed document subject mappings; Based on the set of unvisited document IDs, content retrieval is performed on each ID pair. The document records corresponding to the IDs are extracted from the archive database. For each record, the subject path location information in the original archive directory needs to be obtained. Subject paths are usually represented by a three-level classification. For example, if the directory path corresponding to ID 105 is "Policy Documents / Industry Standards / Technical Standards", the extraction result is the path node index (1,3,2). A one-to-one relationship is established between each ID and its corresponding path index, recorded as a mapping pair (105→1.3.2). After processing all IDs in this way, a mapping set is obtained. For each mapping pair, the validity of the path and the document matching need to be verified, that is, to confirm that the ID does exist under the path. If there are abnormal directory structures or conflicting ID records, the mapping is invalidated. The verified mapping pairs are arranged in sequence to form a complete mapping table. The structure is a pair record structure of document ID and subject path location information. In this table structure, the ID is the key and the path location is the value. Finally, the unvisited document subject mapping set is generated.

[0035] S403: Based on the path structure order in the unvisited document topic mapping set, sort all unvisited document numbers hierarchically by path, aggregate document records under each topic position in sequence, construct a document sequence path with continuous structure, and generate a list of candidate topic extension documents. Based on the path structure order of the unvisited document topic mapping set, all mapping records are extracted. The topic path position field in each record is sorted, and the sorting rule follows the hierarchical order, prioritizing the main category. That is, it is sorted firstly by the first-level category number in ascending order, and then by the second-level and third-level categories in turn. For example, if there are path numbers (1,2,3), (1,2,1), and (1,3,1), the sorting result should be (1,2,1), (1,2,3), and (1,3,1). After sorting, the document numbers under each topic path are aggregated in order. In the aggregation operation, all document numbers belonging to the same path structure are grouped together. For example, document numbers 105, 106, and 109 under path (1,2,1) are grouped into a subsequence. A new document sequence group is generated at the path change point. This type of aggregation is performed on all path groups, and finally several document sequence paths are formed. The document number order in each path is strictly consistent with its corresponding topic level, thus generating a structurally continuous document sequence path and outputting a list of candidate topic extension documents.

[0036] Please see Figure 6 The steps to obtain S5 are as follows: S501: Call the document collection in the candidate topic extended document list, rearrange the document content according to the order index parameter of each document in the chain path, construct a linear content sequence structure, perform position alignment processing on it, and generate a chain content sequence set; The process involves retrieving the document set from the candidate topic extended document list. First, the sequential index number of each document in the set is extracted from its chained path. This number represents the logical position of the document in the path sequence. For example, document number D105 has a sequential index of 8 in the path, which is recorded as (105, 8). After extracting the indexes for all documents, the documents are sorted in ascending order by index number to form an ordered document sequence. Next, content aggregation is performed on the sorted document set, that is, the content fields of each document are extracted one by one, such as title, abstract, and body paragraphs, and these fields are concatenated in sequence. The content is structured as a linear sequence, where each concatenated segment retains the original document number identifier, which is fixedly appended to the beginning of the content segment. Then, position alignment processing is performed on the concatenated sequence structure. The processing method is to mark the starting character position of each segment in the overall sequence. For example, if the content of the first document is 500 characters, then the starting position of the content of the second document is 501, and so on, to build a global position information table. The position information table uniquely determines the relative position of each document in the entire sequence, and finally forms a chained content sequence set containing sequential index, document number, and start and end positions.

[0037] S502: Based on the chained content sequence set, extract the sequence fragments at the end of the path, compare the position index of the fragments in the content sequence, filter the document paragraphs that have not yet appeared in the complete path, and classify them in combination with the sequential number to obtain the set of newly added paragraphs at the end of the path. Based on a chained content sequence set, sequence segments at the end of the path are extracted. The end range can be set as a percentage window based on the total number of documents. For example, if the total number of documents in the path is 20, and the end range is set to 20%, then the last 4 documents are extracted to form a set of end content segments. The position index of each document in the set in the sequence structure is extracted and compared with the position of all documents in the chained content sequence set. If the document number has already appeared in the previous segment, it is marked as a duplicate number. If it has not appeared, it is considered a new segment. The new segment number and its position number in the path are combined into a record. For example, if document number 110 first appears at position 18 and does not appear in the first 17 positions, it is recorded as (110, 18). After screening all the end documents, all segments marked as new are sorted in ascending order according to their path position numbers and classified into a set of new end segments. Each record in the set must contain the document number, path index, and original content reference for subsequent comparison processing.

[0038] S503: Based on the document content characteristics in the newly added paragraph set at the end of the path, retrieve its topic tags, compare them with the topic direction vectors in the preceding path, filter out segments with inconsistent directions, and retain consistent content in the original order to create personalized electronic archive recommendation content. Based on the document content features of the newly added paragraphs at the end of the path, the topic tags corresponding to each document are extracted. These tags can be obtained from the pre-defined topic fields in the document content. For example, if a document's topic is marked as "intelligent manufacturing," then the topic tag for that document is "intelligent manufacturing." Simultaneously, topic direction vectors are extracted from the preceding documents in the chained path. These direction vectors can be formed by aggregating and averaging the word vectors of multiple document topic terms. For example, if the topics of the five preceding documents are "industrial internet," "digital factory," "industrial automation," "intelligent manufacturing," and "control system," then their direction vectors are the average of the word vectors of the five topic tags. Subsequently, a similarity comparison is performed between the topic vectors of the newly added paragraphs at the end and the direction vectors of the preceding path. If the similarity is less than a set threshold... If the direction is inconsistent, it is determined that the direction is inconsistent; otherwise, it is determined that the direction is consistent. Segments with inconsistent directions are filtered out, that is, paragraph records with insufficient similarity are deleted. The paragraph numbers and their order positions that are consistent with the path topic direction are retained. Finally, the retained paragraphs are sorted and spliced ​​in the original order to create personalized recommendation content for electronic archives.

[0039] A personalized recommendation system for electronic records based on user behavior, comprising: The behavior jump relationship module obtains user access records and extracts user identifiers, access times and document numbers. It constructs an operation sequence based on time sorting, calls the summaries, types and category tags of adjacent documents, judges the changes in jump content, marks the jump combination type, reorganizes it into a chain structure, and generates a set of semantic jump path fragments. The path convergence aggregation module calls the document paths in the semantic jump path fragment set, extracts the content description of each document, determines whether the adjacent description structures are continuous, retains the paths in the structure set, removes paragraphs with chaotic jump spans, aggregates fragments with consistent structures, and generates convergent semantic behavior chain groups. The hierarchical span extraction module calls the document number in the convergent semantic behavior chain group, retrieves the topic level position in its directory, compares the adjacent level numbers according to the path order, determines whether there is a hierarchical jump behavior, extracts and marks the segments with large spans, and generates a cross-level belonging topic path set. The extended document building module calls the document number in the cross-level subject path set, compares it with the user access archive list, removes duplicate numbers, organizes unaccessed documents and matches their path level position, reorganizes the chain structure according to the path order, and generates a candidate subject extended document list. The personalized recommendation output module calls up a list of candidate topic extension documents, extracts document content and path order information, filters documents that have not yet appeared at the end of the path, retains paragraph content with the same topic direction, rearranges the structural path for display, and generates personalized electronic archive recommendation content.

[0040] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A personalized recommendation method for electronic records based on user behavior, characterized in that, Includes the following steps: S1: Obtain the user's access records in the archive platform, organize the behavior sequence according to the user's identity and rearrange the archive behavior of each user, record the jump combination of adjacent archives, extract the content change direction between jumps, construct the access behavior chain, and obtain a set of semantic jump path fragments; S2: Call each path segment in the semantic jump path fragment set, extract the content expression mode of the document within the path, perform structural judgment on the content performance of the document in sequence, filter out the path segments with consistent continuous performance, remove the parts with inconsistent topic directions, and obtain the convergent semantic behavior chain group. S3: Call the document content in the convergent semantic behavior chain group, retrieve the topic identifier of the document in the directory structure, mark the topic belonging position of each document in the path, identify paragraphs with hierarchical changes, extract the parts with clear topic spans separately, and obtain the cross-level belonging topic path set; S4: Call the document number in the cross-layer belonging topic path set, compare it with the file number accessed by the user, exclude duplicate number content, retain the document record of the first appearance, and organize it into an updated document chain according to its topic path identifier to obtain a candidate topic extension document list.

2. The personalized recommendation method for electronic records based on user behavior according to claim 1, characterized in that: The file access sequence clues include user access identity sequence, document jump logic relationship, and jump content feature transformation fragments. The convergent semantic behavior chain group includes adjacent content description structure, semantic concentrated expression segment, and structural continuity fragment set. The cross-level affiliation topic path set includes topic level jump marker, continuous behavior topic span segment, and document topic level mapping. The candidate topic extended document list includes newly added unread document number, topic path completion information, and extended chain structure document set.

3. The personalized recommendation method for electronic records based on user behavior according to claim 1, characterized in that, The steps for obtaining S1 are as follows: S101: Obtain all access records of the user in the archive platform, and aggregate the user identity field and access timestamp field in the records. Arrange the access behavior of the same user in chronological order to build a continuous document access behavior chain and generate a user time-series access sequence set. S102: Based on the user time-series access sequence set, extract the adjacent document access pairs for each user, call the document content feature vector and access time interval data, and perform joint filtering based on the feature vector difference and time interval threshold to obtain a time-constrained jump combination set; S103: Call the content feature vector sequence in the time-constrained jump combination set, extract the feature change information in the document jump path, and perform sequence splicing and semantic consistency judgment on the continuous feature change path to establish a semantic jump path fragment set.

4. The personalized recommendation method for electronic records based on user behavior according to claim 3, characterized in that: The process of jointly filtering based on feature vector differences and time interval thresholds limits the temporal continuity of document jump behavior by setting a time interval threshold range. The determination of the difference in feature vectors is based on the distance calculation results between the feature vectors of document content, and combined with the preset feature similarity judgment criteria to determine whether the screening conditions are met. The adjacent document access pairs included in the time-constrained jump combination set must simultaneously meet two conditions: the time interval does not exceed the time interval threshold and the difference in content feature vectors does not exceed the feature similarity judgment criteria. In the process of establishing the set of semantic jump path fragments, only document jump path combinations that satisfy the coherent trend of content feature changes are retained.

5. The personalized recommendation method for electronic records based on user behavior according to claim 1, characterized in that, The steps for obtaining S2 are as follows: S201: Call each access path in the semantic jump path fragment set, extract the content description data frame corresponding to each document in the path, perform structural comparison of the content description between adjacent documents, determine the continuous relationship of the description structure within the path range, perform annotation processing based on the consistency of structural order position, and generate a structural continuous tag matrix. S202: Based on the structural continuity marker matrix, retrieve the document fragment sequence that exhibits a structural continuity state in the path, and perform position merging processing on the content description vector corresponding to the sequence. Then, aggregate and encode multiple structural continuity fragments merged within the same path according to the adjacent structural dimensions to obtain a set of structural aggregated fragments. S203: For the structured aggregated fragment set, call the content description span parameter of each fragment, and perform joint screening on the span parameter and content similarity index to remove combined fragments with abrupt span changes and semantic discrepancies, retain path segments that are structurally continuous and content-convergent, and generate convergent semantic behavior chain groups.

6. The personalized recommendation method for electronic records based on user behavior according to claim 1, characterized in that, The steps for obtaining S3 are as follows: S301: Call all document content in the convergent semantic behavior chain group, retrieve the corresponding topic level position of each document in the archive directory, bind the topic level with the document index, establish the topic level index matrix of the path document, and generate the path topic level annotation set. S302: Based on the path topic level annotation set, call the document level identifier in each path group, sequentially scan the level index values ​​of adjacent documents, detect positions where there are discontinuous changes, and extract the document paragraph index with jump characteristics to obtain a set of level jump paragraphs. S303: Based on the set of hierarchical jump paragraphs, extract the topic information of the corresponding documents and their position numbers in the path, and classify them in combination with the complete structure of the path, mark the document segments that show topic span changes in continuous behavior paths, and establish a set of cross-level affiliation topic paths.

7. The personalized recommendation method for electronic records based on user behavior according to claim 1, characterized in that, The steps for obtaining S4 are as follows: S401: Call all document numbers in the cross-layer subject path set, obtain the list of files that the user has completed accessing, perform number matching operation on the two sets of numbers, mark the document records with overlapping numbers, and exclude them from the path set to generate a set of unaccessed document numbers. S402: Based on the set of unaccessed document numbers, retrieve the document content associated with the corresponding number, extract the topic path location registered in the original archive directory, construct a one-to-one mapping table between documents and topic path locations, and establish a set of unaccessed document topic mappings; S403: Based on the path structure order in the unvisited document topic mapping set, sort all unvisited document numbers hierarchically by path, and aggregate document records under each topic position in sequence to construct a structurally continuous document sequence path and generate a candidate topic extension document list.

8. The personalized recommendation method for electronic records based on user behavior according to claim 1, characterized in that, The method further includes: S5: Call the document set in the candidate topic extended document list, arrange the documents according to their path order, identify the unvisited part at the end of the path, retain the document paragraphs that are in the same order and have the same direction as the output sequence, and obtain the personalized recommendation content of electronic archives. The personalized recommendations for electronic archives include paragraphs arranged in structural order, segments displayed with thematic consistency, and document completion at the end of the path.

9. The personalized recommendation method for electronic records based on user behavior according to claim 8, characterized in that, The steps for obtaining S5 are as follows: S501: Call the document set in the candidate topic extended document list, rearrange the document content according to the order index parameter of each document in the chain path, construct a linear content sequence structure, perform position alignment processing on it, and generate a chain content sequence set; S502: Based on the chained content sequence set, extract the sequence fragments at the end of the path, compare the position index of the fragments in the content sequence, filter the document paragraphs that have not yet appeared in the complete path, and classify them in combination with the sequential number to obtain a set of newly added paragraphs at the end of the path. S503: Based on the document content characteristics in the newly added paragraph set at the end of the path, retrieve its topic tags, compare them with the topic direction vectors in the preceding path, filter out segments with inconsistent directions, and retain consistent content in the original order to establish personalized recommendation content for electronic archives.

10. A personalized recommendation system for electronic records based on user behavior, characterized in that: The system is used in the personalized recommendation method for electronic records based on user behavior as described in any one of claims 1-9, the system comprising: The behavior jump relationship module obtains user access records and extracts user identifiers, access times and document numbers. It constructs an operation sequence based on time sorting, calls the summaries, types and category tags of adjacent documents, judges the changes in jump content, marks the jump combination type, reorganizes it into a chain structure, and generates a set of semantic jump path fragments. The path convergence aggregation module calls the document paths in the set of semantic jump path fragments, extracts the content description of each document, determines whether the adjacent description structures are continuous, retains the paths with concentrated structures, removes paragraphs with chaotic jump spans, aggregates fragments with consistent structures, and generates convergent semantic behavior chain groups. The hierarchical span extraction module calls the document number in the convergent semantic behavior chain group, retrieves the topic level position in its directory, compares the adjacent level numbers according to the path order, determines whether there is a hierarchical jump behavior, extracts and marks the segments with large spans, and generates a cross-level belonging topic path set. The extended document construction module calls the document number in the cross-layered subject path set, compares it with the user access file list, removes duplicate numbers, organizes unaccessed documents and matches their path level positions, reorganizes the chain structure according to the path order, and generates a candidate subject extended document list. The personalized recommendation output module calls the candidate topic extended document list, extracts document content and path order information, filters documents that have not yet appeared at the end of the path, retains paragraph content with the same topic direction, rearranges the structural path for display, and generates personalized electronic archive recommendation content.