Online content recommendation method and system based on semantic discovery

Through data cleaning and natural language processing, a fusion recommendation list is generated through data cleaning and natural language processing, which solves the problem that user intentions are difficult to understand in short text retrieval, and efficient and accurate literature retrieval is achieved, improving user experience and retrieval accuracy.

CN120492737AActive Publication Date: 2025-08-15BEIJING YINGKE QIANXIN TECH CO LTD

Patent Information

Application Number
CN202510668296.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-15
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing semantic discovery technologies are difficult to accurately understand user intentions in short text retrieval, which increases user learning costs, and traditional search tools are difficult to meet the needs of accurate and efficient literature retrieval.

Method used

Through data cleaning and natural language processing, a vector semantic model and BM25 algorithm are combined to generate a fusion recommendation list, and a vector semantic model is used to capture semantic correlation. The BM25 algorithm captures text similarity, dynamically adjusts weights to generate comprehensive scores, lowers the threshold for user use, and improves retrieval accuracy.

Benefits of technology

A more accurate and diverse recommendation list was generated, which improved the comprehensive performance and user experience of the search system, reduced user learning costs, improved the accuracy of short text retrieval, and retained the advantages of traditional retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492737A_ABST
    Figure CN120492737A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an online content recommendation method and system based on semantic discovery, and the method comprises the steps: obtaining literature data, and carrying out the data cleaning; performing natural language processing on the literature data after data cleaning, and performing vector conversion on the literature data after natural language processing; training and updating a vector semantic model in real time by using the literature data subjected to vector conversion, and generating a vector semantic recommendation list by using the trained vector semantic model; carrying out keyword frequency calculation on the basis of the literature data after data cleaning, obtaining a text similarity result based on a result of keyword frequency calculation in combination with a BM25 algorithm, and generating a text similarity recommendation list by using the text similarity result; and fusing the vector semantic recommendation list and the text similarity recommendation list to generate a fusion recommendation list. According to the method, the relevance between the literatures can be comprehensively captured from the two dimensions of semantics and text similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an online content recommendation method and system based on semantic discovery. Background Art

[0002] With the rapid development of information technology and the explosive growth of information resources, traditional keyword search technology has been unable to meet users' needs for accurate and efficient document retrieval. However, semantic discovery technology, as an important innovation in the field of information retrieval, has been widely used in various document retrieval scenarios such as academic research, legal case retrieval, patent analysis, and medical literature review due to its outstanding performance in improving retrieval accuracy and relevance.

[0003] However, for users accustomed to using traditional search tools or literature databases, their operation methods and search logic are very familiar. However, semantic discovery technology often requires users to master new search logic and interaction methods, which increases the learning cost. Short text search is a common scenario in literature search, but because short texts often lack sufficient contextual information, existing semantic discovery technology has difficulty accurately understanding their semantics. Words in short texts may have multiple meanings, and semantic discovery technology may not be able to accurately determine the user's true intention. As a result, existing short text search returns results that are less relevant to user needs.

[0004] Therefore, how to provide an online content recommendation method and system based on semantic discovery is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0005] In view of this, the present invention proposes an online content recommendation method and system based on semantic discovery, aiming to solve the problem of lowering the user's usage threshold and improving its retrieval accuracy.

[0006] In one aspect, the present invention proposes an online content recommendation method based on semantic discovery, comprising: Obtain literature data and perform data cleaning; Performing natural language processing on the document data after data cleaning, and performing vector conversion on the document data after natural language processing; Using the document data after vector conversion to train the vector semantic model and update it in real time; Based on the user search data, generate a vector semantic recommendation list using the trained vector semantic model; Perform keyword frequency calculation based on the document data after data cleaning and the user search data, obtain text similarity results based on the results of the keyword frequency calculation in combination with the BM25 algorithm, and generate a text similarity recommendation list using the text similarity results; The vector semantic recommendation list and the text similarity recommendation list are fused to generate a fused recommendation list.

[0007] Furthermore, the natural language processing includes: performing word segmentation, part-of-speech tagging and named entity recognition on the document data after data cleaning.

[0008] Furthermore, the vector semantic model is trained using the document data after vector conversion, and the vector semantic recommendation list is generated using the trained vector semantic model, including: Splicing the independent semantic units after the entity meaning is marked into a single character stream; Generate subword units based on the Byte Pair Encoding algorithm and obtain the cross-language BPE vocabulary and subword segmenter; Constructing an embedding matrix based on the cross-language BPE vocabulary and the subword segmenter; Based on the Transformer architecture, the Transformer encoder is jointly trained, the model parameters are trained through cross-lingual MLM and contrastive alignment loss, and the progressive masking and language balancing strategy is applied to align the semantic space; The trained Transformer encoder is hierarchically expanded, and a hierarchical training strategy is implemented based on a multi-head attention mechanism and a feedforward neural network to obtain the vector semantic model.

[0009] Furthermore, when the vector semantic model is trained using the document data after vector conversion and the vector semantic recommendation list is generated using the trained vector semantic model, the method further includes: Inputting the document data after vector conversion into the vector semantic model, identifying the input language according to subword distribution, and performing mixed language encoding on paragraphs containing multi-language mixtures; Retrieve relevant entities to inject context and trace the cross-lingual evidence path through the multi-head attention mechanism; Freeze the original parameters, only train the newly added language adapter, and monitor the multi-dimensional indicators of the semantic space to trigger the update of the vector semantic model.

[0010] Furthermore, when performing keyword frequency calculation based on the literature data after data cleaning, it includes: Obtaining the frequency of words appearing in the text content of the document data after natural language processing; Obtaining an inverse document frequency based on the text content of the document data after natural language processing; The result of the keyword frequency calculation is obtained by multiplying the word frequency by the inverse document frequency.

[0011] Furthermore, when obtaining text similarity results based on the result of the keyword frequency calculation in combination with the BM25 algorithm, it includes: ; in, is the inverse document frequency, For words The adjusted term frequencies in document d.

[0012] Furthermore, when the vector semantic recommendation list and the text similarity recommendation list are integrated to generate a fused recommendation list based on dynamic weights, the following steps are included: For each candidate document, extract its semantic score from the vector semantic recommendation list; For each candidate document, extract its text similarity score from the text similarity recommendation list; Determine the weight of the query based on its length and obtain its comprehensive score: Comprehensive score = *Semantic score + (1- )*Text similarity score; in Dynamically adjust with query length; Arrange in descending order according to the comprehensive scores and generate the fusion recommendation list.

[0013] Furthermore, freezing the original parameters, training only the newly added language adapter and monitoring the multi-dimensional indicators of the semantic space to trigger the update of the vector semantic model includes: Obtaining the retrieval accuracy of the vector semantic model; Get the semantic alignment error rate of the vector semantic model: ; in, and is the vector representation of the same concept in different languages in the parallel corpus, and cos is the cosine similarity; Acquire the multi-dimensional index of the semantic space based on the retrieval accuracy and the alignment error rate; Monitor multi-dimensional indicators of the semantic space and trigger a circuit breaker mechanism based on the multi-dimensional indicators.

[0014] Furthermore, monitoring the multi-dimensional indicators of the semantic space and triggering the circuit breaker mechanism based on the multi-dimensional indicators includes: When the multi-dimensional index is higher than a threshold, triggering the vector semantic model to be updated; When the multi-dimensional indicator is lower than the threshold, the fuse rollback instruction is triggered and the incremental training instruction is started.

[0015] Compared with existing technologies, the present invention has the following advantages: through data cleaning and natural language processing, it ensures data accuracy and processability, improving the reliability of subsequent analysis; combining the vector semantic model and the BM25 algorithm, it can comprehensively capture the relevance between documents from the two dimensions of semantics and text similarity, thereby generating a more accurate and diverse recommendation list; the fused recommendation list not only meets the user's demand for semantic relevance, but also provides supplementary recommendations based on text similarity, improving the overall performance and user experience of the recommendation system, and enhancing the retrieval accuracy of short texts. Furthermore, while retaining the user's traditional keyword search habits, it uses natural language processing technology to perform contextual semantic enhancement on short texts, mapping short queries into a semantic space through the vector semantic model, addressing the sparsity problem of short text data and improving retrieval accuracy; at the same time, the text similarity calculation based on the BM25 algorithm retains sensitivity to precise term matching and document structural features, ensuring that the advantages of classic search are not lost. The hybrid recommendation list generated by the fusion of the two not only presents accurate content recommendations in real time on the web page, but also reduces the user's query expression precision requirements through background semantic understanding, allowing users to obtain multi-dimensional enhanced results without adjusting the traditional "keyword + phrase" input mode.

[0016] On the other hand, the present application also provides an online content recommendation system based on semantic discovery, which is used to apply the above-mentioned online content recommendation method based on semantic discovery, including: The acquisition module is configured to acquire literature data and perform natural language processing; A monitoring module is configured to monitor the state of the semantic space in real time and generate a state change event stream; A vector semantic module is configured to retrieve literature data and generate a vector semantic recommendation list; A text similarity module is configured to retrieve document data and generate a text similarity list; A storage module is configured to store the collected literature data and intermediate results; The fusion module is configured to fuse the vector semantic recommendation list and the text similarity recommendation list and generate a fused recommendation list.

[0017] It is understandable that the above-mentioned online content recommendation method and system based on semantic discovery have the same beneficial effects, and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 A flowchart of an online content recommendation method based on semantic discovery provided by an embodiment of the present invention; Figure 2 This is a functional block diagram of an online real-time accurate content recommendation system based on intelligent semantic discovery provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0020] See Figure 1 As shown, in some embodiments of the present application, this embodiment provides an online content recommendation method based on semantic discovery, including: S100: Obtain literature data and perform data cleaning; S200: performing natural language processing on the document data after data cleaning, and performing vector conversion on the document data after natural language processing; S300: Using the document data after vector conversion to train the vector semantic model and update it in real time, and using the trained vector semantic model to generate a vector semantic recommendation list; S400: Calculate keyword frequency based on the document data after data cleaning, obtain text similarity results based on the results of the keyword frequency calculation and the BM25 algorithm, and use the text similarity results to generate a text similarity recommendation list; S500: Fusing the vector semantic recommendation list and the text similarity recommendation list to generate a fused recommendation list.

[0021] Specifically, we deploy the vector semantic model and BM25 full-text retrieval model server through the browser. The front-end receives search requests, performs search and fusion operations, and returns the search results to the front-end for display. The browser retrieves document data, including but not limited to document titles, abstracts, and sources. The document data is cleaned to remove noise and invalid information, such as formatting errors and garbled characters. Natural language processing is then performed on the cleaned document data, pre-annotating key data and phrases in the documents. The annotated document data is then converted into a format suitable for model input, preparing for subsequent retrieval and analysis. The document data is then fed into a vector semantic model. Using the trained vector semantic model, the user's search data is retrieved and a vector semantic recommendation list is generated. Term Frequency-Inverse Document Frequency (TF-IDF) is also calculated on the document data. TF-IDF is a commonly used weighting technique in information retrieval and text mining. It is used to assess the importance of a term within a document set or corpus. The importance of a term increases proportionally with its occurrence in a document, but decreases inversely with its frequency in the corpus. If a word is relatively rare but appears multiple times in the article, it likely reflects the characteristics of the article and is exactly the keyword data we need. We then generate a text similarity recommendation list based on the keyword data. When the user enters a search keyword or sentence, two searches are initiated simultaneously. After the search is completed, the results obtained from the two searches are fused. The semantic similarity and text relevance are comprehensively considered through weighted averaging (using dynamic weights: short queries are given a higher BM25 weight, and long content queries use a higher semantic search weight). Based on the results of the fused search, we screen out several documents that are most relevant to the user input, sort these documents, recommend them, and display them to the user.

[0022] As can be seen, data cleaning and natural language processing ensure data accuracy and processability, improving the reliability of subsequent analysis. Data cleaning is the first step in data processing, improving data quality by removing noise and invalid information, such as formatting errors and garbled characters. Natural language processing further processes the cleaned document data, pre-annotating key data and phrases within the documents and converting the annotated document data into a format suitable for model input. This prepares the data for subsequent retrieval and analysis, improving data processability and providing a high-quality data foundation for subsequent vector conversion and model training. Secondly, combining the vector semantic model with the BM25 algorithm comprehensively captures the relevance between documents from two dimensions: semantics and text similarity, thereby generating more accurate and diverse recommendation lists. By converting document data into vector form, the vector semantic model captures semantic relationships between documents and generates semantically-based recommendation lists. The BM25 algorithm is a classic text similarity calculation method. It combines the results of keyword frequency calculation with the text similarity results obtained by the BM25 algorithm to generate a recommendation list based on text similarity. This dual-dimensional recommendation approach can improve the overall performance of the recommendation system and the user experience. The fused recommendation list not only meets users' needs for semantic relevance but also provides supplementary recommendations based on text similarity, improving the overall performance and user experience of the recommendation system. When a user enters a search keyword or sentence, the system simultaneously initiates two search methods. After the search is complete, the results of the two search methods are merged. By comprehensively considering semantic similarity and text relevance through a weighted average method, the system selects the most relevant documents to the user input based on the results of the fused search. These documents are ranked, recommended, and presented to the user, improving the retrieval accuracy of short texts. Without changing the user's original search habits, the system lowers the user's usage threshold through the form of real-time, accurate online content recommendations on web pages, and further improves retrieval accuracy through the fused search strategy. Retrieving short texts has always been a difficult problem in information retrieval. Due to the short length of texts, traditional keyword matching methods often have difficulty capturing the semantic information of the text. By combining the vector semantic model and the BM25 algorithm, it is possible to comprehensively capture the relevance between short texts from both semantic and text similarity dimensions, thereby generating a more accurate and diverse recommendation list. At the same time, the solution deploys a vector semantic model and a BM25 full-text retrieval model server through the browser, receives retrieval requests from the front end, performs retrieval and fusion operations, and returns the retrieval results to the front end for display. This form of online real-time and accurate content recommendation not only lowers the user's usage threshold, but also improves the real-time and accuracy of retrieval.

[0023] In some embodiments of the present application, performing natural language processing on the document data after data cleaning and performing vector conversion on the document data after natural language processing includes: Based on the domain dictionary, the document data after data cleaning is divided into independent semantic units; Label each independent semantic unit with its grammatical category; Identify and annotate the entity meaning of the independent semantic units after the grammatical category is marked; Convert the independent semantic units after annotating the entity meaning into a format suitable for model input.

[0024] It is understandable that when performing natural language processing on document data, the document data is subjected to natural language processing operations such as word segmentation (dividing continuous natural language text into the smallest units with semantics. For example, using Jieba word segmentation), part-of-speech tagging (part-of-speech tagging: marking the grammatical category of each word after word segmentation, such as noun, verb, adjective, etc.) and named entity recognition (named entity recognition: identifying entities with specific meanings in the text, such as names of people, places, organizations, time, and currency).

[0025] In some embodiments of the present application, the vector semantic model is trained and updated in real time using the document data after vector conversion. When the trained vector semantic model is used to generate a vector semantic recommendation list, the following steps are included: Splice the independent semantic units after the annotated entity meaning into a single character stream; Generate subword units based on the Byte Pair Encoding algorithm and obtain the cross-language BPE vocabulary and subword segmenter; Build an embedding matrix based on the cross-lingual BPE vocabulary and subword segmenter; Based on the Transformer architecture, the Transformer encoder is jointly trained, the model parameters are trained through cross-lingual MLM and contrastive alignment loss, and the progressive masking and language balancing strategy is applied to align the semantic space; The trained Transformer encoder is hierarchically expanded, and a hierarchical training strategy is implemented based on the multi-head attention mechanism and feedforward neural network to obtain a vector semantic model.

[0026] Specifically, the RoBERTa model, based on the Transformer architecture, performs well in natural language processing tasks, capturing long-range semantic dependencies and using word embedding technology to represent each word as a low-dimensional vector. (Word embedding is a technique for mapping words (or phrases) in natural language to a low-dimensional continuous vector space. Its core idea is to use machine learning models to learn the semantic and grammatical features of words from a large amount of text, so that semantically similar words are close in distance in the vector space, while semantically unrelated words are farther apart.) The model encodes the input through multiple Transformer layers, each of which includes a multi-headed attention mechanism and a feed-forward neural network.

[0027] For multilingual literature data, a cross-language pre-trained model is used or multilingual corpora are introduced during the training process. RoBERTa can directly process text in multiple languages and map it to a unified semantic space. RoBERTa's multilingual capabilities are achieved through data blending, a shared vocabulary, and cross-lingual training objectives. Its core goal is to map different languages to a unified vector space. This multilingual mapping to the same semantic space is achieved by sharing a common vocabulary, constructing a dictionary by sampling text from all languages using Byte Pair Encoding (BPE), and introducing a Translation Language Model during the training phase to align multilingual relationships.

[0028] In some embodiments of the present application, the vector semantic model is trained and updated in real time using the document data after vector conversion. When the trained vector semantic model is used to generate a vector semantic recommendation list, the following steps are also included: The document data after vector conversion is input into the vector semantic model, the input language is identified based on the subword distribution, and mixed language encoding is performed on paragraphs containing multilingual mixtures; Retrieve relevant entities to inject context and trace cross-lingual evidence paths through a multi-head attention mechanism; Freeze the original parameters, train only the newly added language adapter, and monitor the multi-dimensional indicators of the semantic space to trigger the update of the vector semantic model.

[0029] Specifically, hybrid language encoding is achieved through a hierarchical attention mechanism: local attention is used at the word level to capture intra-language dependencies, global attention is used to aggregate cross-language semantics at the paragraph level, and relative position encoding is used to maintain the sequential relationship of alternating language segments. When new data needs to be trained, the original parameters are frozen and only the newly added language adapter is trained. The multidimensional indicators of the semantic space of the newly added language adapter are monitored, and the vector semantic model is updated based on the multidimensional indicators.

[0030] It is understandable that based on the subword distribution analysis of the cross-language BPE vocabulary, the system can detect the dominant language components of the input text in real time (such as identifying 60% Chinese, 30% English, and 10% professional symbols in a mixed Chinese-English paragraph). Compared with traditional language detection methods based on n-gram or dictionary matching, it has improved the accuracy and is particularly good at processing terminology-intensive academic texts. Subword granularity modeling can solve the out-of-vocabulary (OOV) problem. For example, the new Chinese word "quantum supremacy" is split into "##quantum" + "##hegemony", and shares the subword "##quant" with the English "quantum supremacy", realizing cross-language semantic equivalence mapping, and improving the retrieval recall rate of low-resource language documents. The RotaE knowledge graph embedding and gated fusion strategy are used to dynamically inject external knowledge (such as "CRISPR→gene editing→Nobel Prize 2020") into text representation, and visualize cross-language evidence paths (such as Chinese "neural network" and English "neural The system uses a high-intensity attention link ("high-intensity attention link" in the "network") to provide explainable recommendation basis, allowing users to intuitively understand the recommendation logic. Through the triple innovation of multilingual hybrid encoding, knowledge-enhanced reasoning, and sustainable learning mechanisms, it has built an efficient, stable, and explainable cross-language document processing system. Its core value lies in: it breaks through the limitations of traditional single-language models and realizes "unbounded input and precise output" multilingual free interaction; it maintains system stability during continuous evolution through parameter freezing and adapter fine-tuning; and it also uses hierarchical attention and entity injection technology.

[0031] In some embodiments of the present application, when calculating the keyword frequency based on the literature data after data cleaning, the following steps are included: Obtain the frequency of words appearing in the text content of the document data after natural language processing; Obtaining the inverse document frequency based on the text content of the document data after natural language processing; Multiply the word frequency by the inverse document frequency to get the result of keyword frequency calculation.

[0032] Specifically, get the word frequency: , inverse document frequency: , and then obtain the final keyword frequency calculation result by multiplying the word frequency and the inverse document frequency.

[0033] In some embodiments of the present application, when obtaining text similarity results based on the results of keyword frequency calculation combined with the BM25 algorithm, the following steps are included: ; in, is the inverse document frequency, For words The adjusted term frequencies in document d.

[0034] It's understandable that by combining keyword frequency with the text similarity results obtained using the BM25 algorithm, a text similarity recommendation list is generated based on these results. By counting the frequency of word occurrences in a single document, the document's core subject terms can be identified. Furthermore, by calculating the rarity of words across the entire corpus, common terms (such as "research" and "methods") can be filtered out and domain-specific terms can be highlighted, improving differentiation and reducing interference from common terms. The deep integration of TF-IDF and BM25 algorithms creates an efficient and adaptable text similarity calculation framework. This framework not only extracts high-quality keywords through TF-IDF weighting, but also addresses document length and word frequency discrepancies through BM25 optimization. Furthermore, dynamic updates and parameter optimization achieve domain adaptation.

[0035] In some embodiments of the present application, when the vector semantic recommendation list and the text similarity recommendation list are fused and the fused recommendation list is generated based on the dynamic weight, the following steps are included: For each candidate document, extract its semantic score from the vector semantic recommendation list; For each candidate document, extract its text similarity score from the text similarity recommendation list; Determine the dynamic weight based on the query length and obtain its comprehensive score: ; in Dynamically adjust with query length; Sort by comprehensive score in descending order and generate a fusion recommendation list.

[0036] Specifically, by extracting the semantic score of each candidate document from the vector semantic recommendation list, the semantic relationship between documents can be captured, generating a semantic-based recommendation list. By extracting the text similarity score of each candidate document from the text similarity recommendation list, the text similarity between documents can be captured, generating a recommendation list based on text similarity. By determining a dynamic weight based on the query length and obtaining its comprehensive score, it is possible to comprehensively consider semantic similarity and text relevance to generate a more accurate and diverse recommendation list. The dynamic weighting approach can automatically adjust the weights of the semantic score and text similarity score based on the length of the content entered by the user. For short queries, a higher text similarity score weight is assigned because keyword matching is more important for short queries; for long content queries, a higher semantic score weight is used because the semantic relationships of long content are more complex. This dynamic weighting approach can improve the accuracy and diversity of the recommendation list and meet the diverse needs of users.

[0037] By sorting by comprehensive score in descending order and generating a fusion recommendation list, the efficiency and accuracy of document retrieval and recommendation can be improved. Sorting by comprehensive score in descending order can filter out the most relevant documents to the user input based on the comprehensive score, sort these documents, recommend them, and display them to the user.

[0038] In some embodiments of the present application, freezing the original parameters, training only the newly added language adapter, and monitoring the multi-dimensional indicators of the semantic space to trigger the vector semantic model update include: Get the vector semantic model retrieval accuracy: ; Get the semantic alignment error rate of the vector semantic model: ; in, and is the vector representation of the same concept in different languages in the parallel corpus, and cos is the cosine similarity; Obtain multi-dimensional indicators of semantic space based on retrieval accuracy and alignment error rate; Monitor multi-dimensional indicators of the semantic space and trigger the circuit breaker mechanism based on these multi-dimensional indicators.

[0039] In some embodiments of the present application, monitoring multi-dimensional indicators of the semantic space and triggering a circuit breaker mechanism based on the multi-dimensional indicators includes: When the multi-dimensional index is higher than the threshold, the vector semantic model is updated; When the multi-dimensional indicators are lower than the threshold, the circuit breaker rollback instruction is triggered and the incremental training instruction is started.

[0040] Specifically, multi-dimensional indicator calculation: , where CLSR is the cross-lingual retrieval recall rate and SAE is the semantic alignment error. When training a new language adapter, the semantic space metrics are monitored. When the metrics are greater than or equal to 0.7, the vector semantic model is updated. When the metrics are less than 0.7, a circuit breaker rollback instruction is triggered and incremental training mode is started. Incremental training is a machine learning method that aims to adapt model parameters locally to new data or new tasks without retraining the entire model. In the circuit breaker mechanism, incremental training is used to repair model performance degradation or expand new language capabilities while maintaining the stability of the original knowledge. Its principle is to lock most of the parameters of the main model (such as the pre-trained Transformer layer) and only open a small number of new modules (such as adapters) for training. Targeted adjustments are made to specific problems (such as new language adaptation and domain drift) to avoid global parameter perturbations. By freezing the main parameters, the model's performance on the original task is maintained and catastrophic forgetting is prevented. Lightweight bottleneck structures (such as 768-dimensional → 256-dimensional → 768-dimensional datasets), only the adapter parameters are trained. New data (such as new language corpora) is then mixed with a small amount of old data (to prevent forgetting). High-error samples (abnormal data during circuit breaker triggering) are weighted and sampled to obtain the main task loss (cross-lingual MLM) and the contrastive alignment loss. The gradient magnitudes of these two losses are normalized before summing to avoid optimization bias caused by dimensionality differences. Specifically, for example, samples are drawn from the mixed data pool according to weights. Each batch contains: 64 new language parallel data, 32 new language monolingual data, 16 old data, and 8 high-error samples. A loss weight of 2 is applied to high-error samples. The MLM loss is calculated separately for the new and old data (only the original language portion is calculated for the old data). The contrastive alignment loss is calculated for the parallel data, and the monolingual data is back-translated to generate pseudo-parallel pairs for computation. Gradients are only received by the adapter and the top-level LayerNorm parameters. After gradient clipping, an AdamW update is performed, and the parameter changes are recorded (e.g., if the adapter weight change is < 0.001). The CLSR and SAE indicators of the validation set are calculated every 1000 steps. If they decrease for three consecutive times, early stopping is triggered (saving the current best checkpoint). When the multi-dimensional index of the semantic space of the newly added language adapter is greater than or equal to 0.7, the vector semantic model is updated.

[0041] In summary, the present invention has the following beneficial effects: through data cleaning and natural language processing, the accuracy and processability of data are ensured, and the reliability of subsequent analysis is improved; the combination of vector semantic model and BM25 algorithm can comprehensively capture the relevance between documents from the two dimensions of semantics and text similarity, thereby generating a more accurate and diverse recommendation list; the fused recommendation list not only meets the user's demand for semantic relevance, but also provides supplementary recommendations based on text similarity, improving the overall performance and user experience of the recommendation system, and improving the retrieval accuracy of short texts. While retaining the user's traditional keyword search habits, natural language processing technology is used to perform contextual semantic enhancement on short texts, and short queries are mapped into a semantic space through the vector semantic model, solving the problem of short text data sparsity and improving retrieval accuracy; at the same time, the text similarity calculation based on the BM25 algorithm retains sensitivity to precise term matching and document structure features, ensuring that the advantages of classic search are not lost. The hybrid recommendation list generated by the fusion of the two can not only present accurate content recommendations in real time on the web page, but also reduce the user's query expression precision requirements through background semantic understanding, allowing users to obtain multi-dimensional enhanced results without adjusting the traditional "keyword + phrase" input mode.

[0042] In another preferred embodiment based on the above embodiment, refer to Figure 2 As shown, this embodiment provides a dynamic system configuration and state change notification system for applying the above-mentioned online content recommendation method based on semantic discovery, including: The acquisition module is configured to acquire literature data and perform natural language processing; A monitoring module is configured to monitor the state of the semantic space in real time and generate a state change event stream; A vector semantic module is configured to retrieve literature data and generate a vector semantic recommendation list; A text similarity module is configured to retrieve document data and generate a text similarity list; A storage module is configured to store the collected literature data and intermediate results; The fusion module is configured to fuse the vector semantic recommendation list and the text similarity recommendation list and generate a fused recommendation list.

[0043] Specifically, the acquisition module collects document data and performs natural language processing, and then transmits the document data to the vector semantic module and the text similarity module, and then transmits the results to the fusion module for fusion, and finally obtains a fusion recommendation list. During the processing, the document data and intermediate results and other data are stored in the storage module for subsequent analysis and training.

[0044] As you can understand, the browser plug-in automatically reads the content of the currently viewed page (via the browser's official plug-in SDK or similar extraction function) and extracts key elements, such as the document title, abstract, and source, without requiring manual input from the user. The browser plug-in interface displays a list of search results, including automatically extracted information such as the document title, abstract, and source. Relevance sorting options are provided, allowing users to choose to sort by semantic similarity, year in ascending and descending order, and so on. When a user clicks on a recommendation in the search results list, a description of the recommended document and related information, including title, summary, author, year, and DOI, is displayed. Language filtering allows users to select only Chinese or English documents based on their preferences or needs. On the backend, a server deploying a vector semantic model and a BM25 full-text search model (e.g., using a microservices architecture and a RESTful API) receives search requests from the frontend, performs search and fusion operations, and returns the search results to the frontend for display. The storage module uses a database to store the collected document data and pre-processed intermediate results, and regularly updates the document data in the database to ensure the timeliness of search results.

[0045] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0046] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0047] These computer program instructions may also be stored in a computer-readable storage device that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable storage device produce an article of manufacture comprising an instruction device that implements the process. Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. An online content recommendation method based on semantic discovery, characterized in that: include: Obtain literature data and perform data cleaning; Performing natural language processing on the document data after data cleaning, and performing vector conversion on the document data after natural language processing; Using the document data after vector conversion to train the vector semantic model and update it in real time; Based on the user search data, generate a vector semantic recommendation list using the trained vector semantic model; Perform keyword frequency calculation based on the document data after data cleaning and the user search data, obtain text similarity results based on the results of the keyword frequency calculation in combination with the BM25 algorithm, and generate a text similarity recommendation list using the text similarity results; The vector semantic recommendation list and the text similarity recommendation list are fused to generate a fused recommendation list.

2. The online content recommendation method based on semantic discovery according to claim 1, characterized in that The natural language processing includes: performing word segmentation, part-of-speech tagging and named entity recognition on the document data after data cleaning.

3. The online content recommendation method based on semantic discovery according to claim 2, characterized in that The vector semantic model is trained using the document data after vector conversion, and the vector semantic recommendation list is generated using the trained vector semantic model, including: Splicing the independent semantic units after the entity meaning is marked into a single character stream; Generate subword units based on the Byte Pair Encoding algorithm and obtain the cross-language BPE vocabulary and subword segmenter; Constructing an embedding matrix based on the cross-language BPE vocabulary and the subword segmenter; Based on the Transformer architecture, the Transformer encoder is jointly trained, the model parameters are trained through cross-lingual MLM and contrastive alignment loss, and the progressive masking and language balancing strategy is applied to align the semantic space; The trained Transformer encoder is hierarchically expanded, and a hierarchical training strategy is implemented based on a multi-head attention mechanism and a feedforward neural network to obtain the vector semantic model.

4. The online content recommendation method based on semantic discovery according to claim 3, characterized in that The vector semantic model is trained using the document data after vector conversion, and when the vector semantic model is used to generate a vector semantic recommendation list, the method further includes: Inputting the document data after vector conversion into the vector semantic model, identifying the input language according to subword distribution, and performing mixed language encoding on paragraphs containing multi-language mixtures; Retrieve relevant entities to inject context and trace the cross-lingual evidence path through the multi-head attention mechanism; Freeze the original parameters, only train the newly added language adapter, and monitor the multi-dimensional indicators of the semantic space to trigger the update of the vector semantic model.

5. The online content recommendation method based on semantic discovery according to claim 4, characterized in that When calculating the keyword frequency based on the literature data after data cleaning, it includes: Obtaining the frequency of words appearing in the text content of the document data after natural language processing; Obtaining an inverse document frequency based on the text content of the document data after natural language processing; The result of the keyword frequency calculation is obtained by multiplying the word frequency by the inverse document frequency.

6. The online content recommendation method based on semantic discovery according to claim 5, characterized in that When the text similarity result is obtained based on the result of the keyword frequency calculation combined with the BM25 algorithm, it includes: ; in, is the inverse document frequency, For words The adjusted term frequencies in document d.

7. The online content recommendation method based on semantic discovery according to claim 6, characterized in that: When the vector semantic recommendation list and the text similarity recommendation list are integrated to generate a fused recommendation list based on a dynamic weight, the method includes: For each candidate document, extract its semantic score from the vector semantic recommendation list; For each candidate document, extract its text similarity score from the text similarity recommendation list; Determine the weight of the query based on its length and obtain its comprehensive score: Comprehensive score = *Semantic score + (1- )*Text similarity score; in Dynamically adjust with query length; Arrange in descending order according to the comprehensive scores and generate the fusion recommendation list.

8. The online content recommendation method based on semantic discovery according to claim 7, characterized in that: Freezing the original parameters, training only the newly added language adapter and monitoring the multi-dimensional indicators of the semantic space to trigger the update of the vector semantic model includes: Obtaining the retrieval accuracy of the vector semantic model; Get the semantic alignment error rate of the vector semantic model: ; in, and is the vector representation of the same concept in different languages in the parallel corpus, and cos is the cosine similarity; Acquire the multi-dimensional index of the semantic space based on the retrieval accuracy and the alignment error rate; Monitor multi-dimensional indicators of the semantic space and trigger a circuit breaker mechanism based on the multi-dimensional indicators.

9. The online content recommendation method based on semantic discovery according to claim 8, characterized in that Monitoring the multi-dimensional indicators of the semantic space and triggering the circuit breaker mechanism based on the multi-dimensional indicators includes: When the multi-dimensional index is higher than a threshold, triggering the vector semantic model to be updated; When the multi-dimensional indicator is lower than the threshold, the fuse rollback instruction is triggered and the incremental training instruction is started.

10. An online content recommendation system based on semantic discovery, applied to the online content recommendation method based on semantic discovery according to any one of claims 1 to 9, characterized in that: include: The acquisition module is configured to acquire literature data and perform natural language processing; A monitoring module is configured to monitor the state of the semantic space in real time and generate a state change event stream; A vector semantic module is configured to retrieve literature data and generate a vector semantic recommendation list; A text similarity module is configured to retrieve document data and generate a text similarity list; A storage module is configured to store the collected literature data and intermediate results; The fusion module is configured to fuse the vector semantic recommendation list and the text similarity recommendation list and generate a fused recommendation list.

Citation Information

Patent Citations

  • A bilingual word embedding-based cross-language text similarity assessment technique

    CN109213995A

  • Scientific and technical literature quotation recommendation method based on deep learning

    CN113239181A

  • Cross-language text semantic model generation method and device and electronic equipment

    CN114417879A

  • Purchase data structured processing method and system

    CN119903062A

  • Hybrid enhanced indexing method and system based on vector retrieval and BM25 algorithm

    CN119961376A

Cited By

  • Music recommendation method and device and XR equipment

    CN120744169A

  • Music recommendation method and device, and XR device

    CN120744169B

  • Data mining system and method for mail data

    CN121070985A

  • Service recommendation method and device, storage medium and electronic equipment

    CN121880662A

  • A service recommendation method and device, a storage medium and an electronic device

    CN121880662B