Dialect region feature recognition method and system based on multi-modal user interest vector calculation and application

By constructing a user dialect community matrix, the interest of users in multiple dialect regions is quantified, solving the problem of insufficient dialect culture preference identification in existing technologies, and achieving more accurate user recommendations and advertising.

CN121884777APending Publication Date: 2026-04-17BESTTONE HOLDING
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511903661.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify users' dialect and cultural preferences, as well as the dynamic drift of dialect preferences due to changes in life scenarios, resulting in low accuracy of regional recommendations and an inability to meet users' refined perception of interests.

Method used

By constructing a seed user interest feature matrix and a target user behavior matrix, and combining multimodal data analysis, a user dialect community matrix is ​​generated to quantify users' interests in multiple dialect regions. The dialect content feature matrix is ​​then used for recommendations.

Benefits of technology

It improved the accuracy of user dialect identification, adapted to users' multi-dialect scenarios, enhanced the accuracy of recommendations and user satisfaction, and expanded the platform's business model, such as dialect marketing and advertising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884777A_ABST
    Figure CN121884777A_ABST
Patent Text Reader

Abstract

The invention relates to the field of emerging technologies, in particular to a dialect region feature recognition method and system based on multi-mode user interest vector calculation and application. Comprising the following steps: constructing a seed user interest feature matrix: selecting seed users in each dialect region according to a preset rule, extracting interactive behavior features of the seed users on platform contents, and forming the seed user interest feature matrix; acquiring target user behavior data: acquiring an interaction behavior log of a target user and platform dialect related content, and constructing a user behavior feature matrix based on the target user behavior data; and generating a user dialect community matrix: performing similarity calculation and clustering analysis on the user behavior feature matrix and the seed user interest feature matrix to obtain a dialect community matrix of the target user. According to the method, the dialect features related to the content are quantized as independent dimensions and fused into the feature vectors, the recommendation list containing various dialect contents and recommendation degrees can be generated for the user through recognition and matching of the dialect features of the content and the behavior features of the user, and the recommendation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to emerging technology fields, and in particular to a dialect region feature recognition method, system, and application based on multimodal user interest vector calculation. Background Technology

[0002] Current mainstream short video / social media platforms use user behavior feature analysis and similar content recommendation to recommend content that different user groups may be interested in, thereby increasing content click-through rates and user engagement time, and meeting the platform's need for precise operation. In terms of technical implementation, platform recommendation functions generally construct feature vectors based on content tags and user behavior tags. They then employ collaborative filtering algorithms and deep learning techniques to calculate the distribution of users' potential preferences for certain content based on the similarity of these feature vectors, thus enabling the recommendation of relevant content.

[0003] In practical applications, to build a content service platform covering the entire country with tens of millions of users, its recommendation module can introduce a region-related dimension in addition to user interests and behavioral characteristics. By automatically identifying the regional cultural characteristics of the content, it can improve users' refined perception of their social circles. Existing technologies can perform relatively coarse-grained regional recommendations through IP address or mobile phone number segment identification, but they have low accuracy, are prone to errors, and cannot identify users' cultural preferences (e.g., people working and living in other places want to see content from their hometown) or solve the problem of user interest drift.

[0004] Among the published existing technologies, CN111382310A and CN108920503A calculate similarity based on general interest tags and further introduce the trust factor of social networks to optimize the recommendation algorithm, so as to reduce the impact of data sparsity and improve the efficiency of the algorithm. This type of method mainly considers user characteristics, content characteristics and the influence of social networks, but does not solve the regional recommendation problem. CN116089726A proposes a multi-dialect and multimodal resource recommendation method for Chinese and Tibetan languages. It realizes dialect content matching recommendation based on machine translation and similar group diffusion technology. This method focuses on automatic identification and labeling of content, but does not quantify dialect as an independent dimension, nor does it consider the dynamic drift of dialect preferences caused by changes in users' life scenarios. It cannot make full use of the regional cultural characteristics of the content to improve users' refined perception of identity circles.

[0005] In summary, for content platforms with a large user base and wide coverage, it is necessary to construct a user interest vector representation model that integrates users' multimodal behavioral data and linguistic and cultural characteristics to cluster and identify user groups with the same or similar dialect regional characteristics, thereby solving the regional recommendation problem. Summary of the Invention

[0006] The purpose of this invention is to address the problems existing in the background technology by proposing a dialect region feature recognition method based on multimodal user interest vector calculation. This method can automatically identify and quantify users' interests in content from multiple dialect regions, solving problems such as the inability of existing recommendation technologies to effectively identify users' dialect cultural preferences and the dynamic drift of dialect preferences caused by changes in living scenarios.

[0007] The technical solution of the present invention, in its first aspect, provides a dialect region feature recognition method based on multimodal user interest vector calculation, comprising the following specific steps: S1. Construct a seed user interest feature matrix: Select seed users from various dialect regions according to preset rules, extract their interaction behavior features with platform content, and form a seed user interest feature matrix. S2. Obtain target user behavior data: Collect interaction behavior logs of target users with platform dialect-related content, and construct a user behavior feature matrix based on the target user behavior data; S3. Generate a user dialect community matrix: Perform similarity calculation and cluster analysis on the user behavior feature matrix and the seed user interest feature matrix to obtain the dialect community matrix of the target user; the matrix contains at least one dialect and its corresponding preference weight, which is used to characterize the dynamic dialect interest preference of the target user.

[0008] Preferably, in step S1, the content is labeled offline using a combination of automatic and manual methods to determine the dialect attribute; and the interest feature vector of the seed users is calculated.

[0009] Preferably, the preset rules for selecting seed users in step S1 include: The stability of its geographical attributes is confirmed based on the IP address and / or the H code of the mobile phone number segment; The frequency of their use of dialect in platform comments is determined based on the number of comments; Set a threshold for platform usage activity.

[0010] Preferably, in step S2, the dialect-related content includes videos, posts, audio, or comments; Interactive behaviors include watching, liking, sharing, commenting, creating, and searching.

[0011] Preferably, in step S2, the factors in the user behavior features are mapped to construct a user behavior vector; The user behavior vectors are processed by concatenation and weighted averaging to generate a multimodal, comprehensive user behavior feature matrix.

[0012] Preferably, in the process of constructing user behavior vectors, user behavior is used for modeling, the user behavior sequence is mapped into a low-dimensional dense vector through an embedding model, and the behavior vectors of different modalities are spliced ​​and weighted and fused.

[0013] Preferably, the clustering analysis in step S3 includes the K-Means algorithm and normalization processing.

[0014] A second aspect of the present invention provides a content recommendation method based on dialect region features calculated from multimodal user interest vectors. The method uses the aforementioned dialect region feature recognition method to identify dialect content, and includes the following specific steps: Constructing a dialect content feature matrix: For platform content, identify and quantify the various dialect features it contains to form a dialect content feature matrix; Identify the dialect region characteristics of the target users to obtain the dialect community matrix of the target users; Fusion recommendation calculation: Based on the dialect weights in the dialect community matrix of the target user and the correlation strength between the content and each dialect in the dialect content feature matrix, calculate the recommendation probability of each content for the target user; Generate and output a recommendation list: Based on the recommendation probability, generate a recommendation list of dialect content for the target user; Data storage: Save the recommendation results and update the behavioral attributes of ordinary users.

[0015] Preferably, a natural language processing model is first used to analyze dialect words in the text, and / or a speech recognition model is used to analyze dialect speech features in the audio; and a preset dialect feature weight parameter group is used to perform weighted calculation on the identified features to obtain quantified dialect feature values ​​to construct the matrix. Subsequently, matrix operations were performed on the weight vectors in the dialect community matrix and the dialect content feature matrix, and combined with the operation metrics of the newness and popularity of the content, to obtain the comprehensive recommendation probability.

[0016] A third aspect of the present invention provides a dialect region feature recognition system based on multimodal user interest vector calculation, which identifies dialect content according to the above-described recognition method, including: The data acquisition module is used to collect logs of interaction behaviors between platform users and dialect-related content; The feature extraction and vectorization module is used to construct and store the dialect content feature matrix and the seed user interest feature matrix. Based on the logs collected by the data acquisition module, construct a target user behavior feature matrix; The vector calculation and clustering module is used to perform similarity calculation and cluster analysis between the target user behavior feature matrix and the seed user interest feature matrix to generate the dialect community matrix of the target user. The application recommendation module calculates and outputs recommendation results for the target user based on the dialect community matrix and the dialect content feature matrix.

[0017] Compared with the prior art, the present invention has the following beneficial technical effects: 1. Technically, compared with existing content recommendation and user profiling technologies, this invention quantifies and integrates content-related dialect features as an independent dimension into the feature vector. By recognizing and matching the dialect features of the content with the user's behavioral features, a recommendation list containing multiple dialects and recommendation levels can be generated for the user, which can improve the accuracy of the recommendation. 2. For platform users, this invention generates a dialect community matrix containing multiple dialect features for a single user, and can automatically determine which dialect has a higher weight for that user, improving the accuracy of user dialect identity classification, and is more adaptable to users' multi-dialect scenarios, such as changes in dialect preferences after working and living in other places. It fully considers the dynamic interest drift of users, improves the accuracy of user dialect identity classification, and can improve user satisfaction. 3. For platform operators, this invention expands the business model for content platforms and opens up a new dimension of precise advertising: "dialect marketing". Attached Figure Description

[0018] Figure 1 This is a system structure diagram of the dialect region feature recognition system in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the dialect region feature recognition method in an embodiment of the present invention. Detailed Implementation Example 1

[0019] This embodiment proposes a dialect region feature recognition method based on multimodal user interest vector calculation, which includes the following specific steps: S1. Construct a seed user interest feature matrix: Select seed users from various dialect regions according to preset rules, extract their interaction behavior features with platform content, and form a seed user interest feature matrix. S2. Obtain target user behavior data: Collect interaction behavior logs of target users with platform dialect-related content, and construct a user behavior feature matrix based on the target user behavior data; S3. Generate a user dialect community matrix: Perform similarity calculation and cluster analysis on the user behavior feature matrix and the seed user interest feature matrix to obtain the dialect community matrix of the target user; the matrix contains at least one dialect and its corresponding preference weight, which is used to characterize the dynamic dialect interest preference of the target user.

[0020] To facilitate understanding of this solution, a specific example is used below to provide a detailed introduction: Content is labeled offline, and the interest feature vectors of seed users are calculated. The content and behavioral preferences of dialect users are output as a prerequisite and basis for the cold start of the recommendation system. Figure 2 As shown on the left.

[0021] Step 1: Using a combination of automatic and manual methods, label dialect-related content and calculate the dialect feature values ​​of the video content.

[0022] Before the video content is uploaded, the speech recognition model is used to analyze the dialect speech features in the video, and the video subtitles are analyzed using the NLP model. At the same time, the dialect attributes are manually identified by the platform operators, and the manual identification is given the highest weight. After the video content is uploaded, the platform automatically collects dialect words from the comments text of the content.

[0023] definition: (1) Dialect feature weight parameter set, representing the weight of various dialect features in content tagging:

[0024] Among them, w1 represents the weight of video content containing dialect, which is a strong identifier and has the highest weight; w2 represents the weight of video subtitles containing dialect; w3 represents the weight of comments containing dialect, and so on. The specific weight parameter group is determined by the system backend and can be adjusted in a timely manner according to the actual operation results; (2) The dialect feature vector matrix corresponding to the content, representing the dialect features related to a certain video content:

[0025] Where dIDn is a specific dialect ID; fIDn represents the numerical value of a certain feature of that dialect, such as: Is it a video in a dialect? (0 / 1) Number of times dialect appears in the video subtitles; The number of times dialects appeared in the comments section, etc.

[0026] One piece of content can correspond to multiple dialect features; therefore, each row of matrix F3 represents the quantitative feature of a certain dialect presented by a certain video content. The fID factor corresponds one-to-one with the parameters in wParamSet.

[0027] Each row of matrix F3, representing a dialect feature, is used to construct a vector fIDSet. This vector is then multiplied by wParamSet, yielding a dialect feature value fValue for each calculation. Finally, a dialect feature value matrix F4 is constructed.

[0028]

[0029] F4 corresponds to the ID (cID) of a video content, indicating that the content has quantitative values ​​of several dialect features, but the strength of each dialect is different.

[0030] Step 2: Seed user selection and feature extraction.

[0031] Based on operational needs, a certain number of seed users from different dialect regions are selected, and their interest characteristics are extracted as the basis for identifying dialect user attributes in the recommendation module. Seed users are limited to dialect characteristics of one region. To accurately extract user characteristics, a certain number of seed users needs to be maintained, generally no less than 100. Selection rules include: (1) Regional characteristics: Confirmed by IP address and H code, and has not changed for a long time; dialect appeared multiple times in the comments; (2) Activity level; The platform usage time in the past three years has reached a certain standard; (3) Social attributes: strong social attributes, such as a large number of comments and shares.

[0032] First, we analyzed the top 100 most-used video content by seed users in each dialect region. For each video content, we extracted the most popular user behaviors (watching videos, commenting, sharing, etc.) and recorded them as follows: {cID,bID1,bID2,bID3} Considering computational complexity, each video content (cID) corresponds to no more than 3 feature behaviors (bID).

[0033] Thus, the interest feature matrix of seed users in the dialect region is obtained as follows:

[0034] Wherein, dID is the dialect ID, with one line for each seed user in each dialect region; cID is the top content ID, which can be empty if the behavior does not involve specific content; and bID is the feature behavior ID.

[0035] Step 3: Following steps 1 and 2, the interest feature vectors of seed users in each dialect region have been constructed. Combined with the dialect feature vector F4 of the video content, they can jointly represent the correspondence between dialect regions and content and behavior. Save F4 and F5 as the basis for subsequent user recommendations and update them regularly.

[0036] In this embodiment, for video content, dialect features contained in the content are identified and quantified using a combination of automatic and manual methods. A dialect feature matrix is ​​used to represent multiple dialect features of varying strengths possessed by the content. Based on the content's dialect feature matrix, combined with the user's dialect community matrix, the recommendation probability of each dialect content for a user can be calculated, improving the accuracy of dialect recommendations. 2. Seed user interest feature extraction model: The interest feature vectors of seed users in each dialect region are selected and constructed. Combined with the dialect feature vectors of the video content, they can jointly represent the correspondence between dialect regions and content and behavior, serving as the basis for user recommendations. Example 2

[0037] This embodiment provides a content recommendation method based on dialect region features calculated from multimodal user interest vectors. It uses the dialect region feature recognition method from Embodiment 1 to identify dialect content, and includes the following specific steps: Constructing a dialect content feature matrix: For platform content, identify and quantify the various dialect features it contains to form a dialect content feature matrix; Identify the dialect region characteristics of the target users to obtain the dialect community matrix of the target users; Fusion recommendation calculation: Based on the dialect weights in the dialect community matrix of the target user and the correlation strength between the content and each dialect in the dialect content feature matrix, calculate the recommendation probability of each content for the target user; Generate and output a recommendation list: Based on the recommendation probability, generate a recommendation list of dialect content for the target user; Data storage: Save the recommendation results and update the behavioral attributes of ordinary users.

[0038] To facilitate understanding of this solution, a specific example is used below to provide a detailed introduction: like Figure 2 As shown, based on dialect feature values ​​and seed user interest features, a content recommendation method based on multimodal user interest vector calculation can be obtained and applied to the recommendation module of a content platform. The process steps are as follows: Figure 2 As shown on the right.

[0039] Step 1: Collect user behavior characteristics, mainly extracted from relevant content and social network dimensions.

[0040] Collect user interaction behaviors (viewing, liking, sharing, commenting, creating, searching) with dialect-related content (videos, posts, audio, comments) from platform logs.

[0041] Step 2: Map the factors in the user behavior features to construct the user behavior vector.

[0042] By modeling user behavior, user behavior sequences are mapped into low-dimensional dense vectors through an embedding model (such as Word2Vec).

[0043] Step 3: Process the user behavior vectors by means of concatenation, weighted averaging, etc., to generate a multimodal, comprehensive user behavior feature matrix.

[0044] The user behavior vector obtained in step two is processed by merging similar behaviors and retaining the video content ID (cID) and its corresponding feature behavior bID. Each cID corresponds to no more than 3 features, resulting in a two-dimensional matrix containing user behavior features, represented as follows:

[0045] Step 4: Based on the user behavior feature matrix F6, calculate the cosine similarity of each row with the seed user interest feature matrix F5 and sort them to obtain the similarity between the user and each dialect seed user. After processing such as cluster analysis (e.g., K-Means algorithm) and data normalization, the final dialect community matrix F2 corresponding to a specific user is obtained, dividing the user into different dialect communities:

[0046] F2 contains quantified values ​​for several dialect regions, representing the user's preference for those regions. Each row corresponds to a specific dialect, sorted by weight. Generally, no more than three dialects are recommended to each user.

[0047] In practical applications, if the cID has strong dialect characteristics, the related calculations can be further optimized to reduce the amount of computation.

[0048] Step 5: At the application layer, the dialect community matrix F2 of a specific user and the dialect feature matrix F4 of a specific content are used as the basis for content recommendation. At the same time, combined with operational demand metrics (such as the freshness of content, etc.), the recommendation probability of each content is further calculated, and finally a list of dialect content recommended to users is obtained. It can include multiple dialects to adapt to the actual preferences of users.

[0049] Step 6: Save the recommendation results and update the behavioral attributes of ordinary users to form a closed loop. Example 3

[0050] This embodiment provides a dialect region feature recognition system based on multimodal user interest vector calculation. It identifies dialect content according to the recognition method in Embodiment 1, including: The data acquisition module is used to collect logs of interaction behaviors between platform users and dialect-related content; The feature extraction and vectorization module is used to construct and store the dialect content feature matrix and the seed user interest feature matrix. Based on the logs collected by the data acquisition module, construct a target user behavior feature matrix; The vector calculation and clustering module is used to perform similarity calculation and cluster analysis between the target user behavior feature matrix and the seed user interest feature matrix to generate the dialect community matrix of the target user. The application recommendation module calculates and outputs recommendation results for the target user based on the dialect community matrix and the dialect content feature matrix.

[0051] like Figure 1 As shown, the data acquisition module is responsible for collecting user interaction behaviors (including watching, liking, sharing, commenting, creating, searching, etc.) corresponding to dialect-related content (including: videos / audio, posts, comments, etc.) from the platform log system.

[0052] From a technical implementation perspective, log collection can be achieved through various methods, such as real-time collection via log server clusters, real-time printing via business interface backend services, or transmission via external platform business interfaces. For recommendation services, if the platform operator does not have strong real-time requirements, logs can also be imported in batches asynchronously for processing. IP address ranges and H-code ranges are obtained through query interfaces provided by operator platforms or third-party external platforms.

[0053] After the raw logs are collected, they undergo parsing, cleaning, and default value filling operations, while discarding irrelevant entries, before proceeding to the next step of storage and aggregation processing.

[0054] The feature extraction and vectorization module is responsible for extracting dialect content features from the collected log data, generating user behavior features based on user behavior modeling, and fusing the two to obtain a multimodal, comprehensive user dialect interest vector. Specifically, this includes: (1) Content feature extraction The NLP model was used to analyze dialect words in video subtitles and comment text; the speech recognition model was used to analyze dialect speech features in audio to obtain dialect feature vectors. (2) User behavior modeling The user's behavior sequence (such as a list of dialect video IDs watched) is mapped into a low-dimensional, dense user behavior vector through an embedding model (such as Word2Vec). (3) Multimodal vector fusion By concatenating and weighting the dialect feature vector with the user behavior vector, a comprehensive user dialect interest vector matrix is ​​obtained. This matrix can characterize the interest features of users in a dialect region towards various dialect contents, similar to:

[0055] F1 corresponds to the interest feature matrix of dialect region seed users described in the Implementation Examples section.

[0056] In the vector calculation and clustering module, the obtained user dialect interest vector matrix is ​​compared with the pre-generated dialect seed user interest vectors according to dialect ID, and the distance between the vectors is measured by the cosine similarity algorithm.

[0057] Using this distance, cluster analysis and normalization are performed on user vectors to classify users into multiple dialect groups, ultimately forming a dialect group matrix corresponding to a specific user:

[0058] Each row of matrix F2 corresponds to a specific dialect, and they are sorted by weight.

[0059] The application recommendation module, based on a dialect community matrix, automatically implements user services and precision operation-related application functions, including: Precise content recommendation: Recommending popular or newly published content from the user's dialect community; Social connection discovery recommends other users in the same dialect group to promote social connections; Regional advertising: Providing advertisers with the ability to conduct precise marketing based on dialect regions.

[0060] The application layer is also responsible for updating the user behavior log data related to this application to the data collection layer, thus achieving a closed business loop.

[0061] Unlike existing content recommendation and user profiling technologies, this invention quantifies dialect features as an independent dimension and integrates them into the feature vector. By automatically identifying users' dialect identities and even their cultural and linguistic circles, it improves the accuracy of user dialect identity classification. Based on the dynamic fusion and weighting of multimodal dialect features, it generates a dialect community matrix containing multiple dialect features for a single user and can automatically determine which dialect has a higher weight for that user, making it more suitable for the user's multi-dialect scenarios. It fully considers the user's dynamic interest drift and can more effectively connect users to social networks and deliver regional advertising.

[0062] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. A dialect area feature recognition method based on a multi-modal user interest vector calculation, characterized in that, The specific steps include the following: S1. Construct a seed user interest feature matrix: Select seed users from various dialect regions according to preset rules, extract their interaction behavior features with platform content, and form a seed user interest feature matrix. S2. Obtain target user behavior data: Collect interaction logs of target users with platform dialect-related content, and construct a user behavior feature matrix based on the target user behavior data; S3. Generate user dialect community matrix: Perform similarity calculation and cluster analysis on the user behavior feature matrix and the seed user interest feature matrix to obtain the target user's dialect community matrix; this matrix contains at least one dialect and its corresponding preference weight, which is used to represent the target user's dynamic dialect interest preference. 2.The dialect region feature recognition method based on the multi-modal user interest vector calculation of claim 1, characterized in that, In step S1, the content is labeled offline using a combination of automatic and manual methods to determine the dialect attribute; and the interest feature vector of the seed users is calculated. 3.The dialect region feature recognition method based on the multi-modal user interest vector calculation of claim 1, characterized in that, The preset rules for selecting seed users in step S1 include: The stability of its geographical attributes is confirmed based on the IP address and / or the H code of the mobile phone number segment; The frequency of their use of dialect in platform comments is determined based on the number of comments; Set a threshold for platform usage activity. 4.The dialect region feature recognition method based on the multi-modal user interest vector calculation of claim 1, characterized in that, Step S2: Dialect-related content includes videos, posts, audio, or comments; Interactive behaviors include watching, liking, sharing, commenting, creating, and searching.

5. The dialectal regional feature recognition method based on the multi-modal user interest vector calculation according to claim 1, characterized in that, In step S2, the factors in the user behavior features are mapped to construct a user behavior vector; The user behavior vectors are processed by concatenation and weighted averaging to generate a multimodal, comprehensive user behavior feature matrix.

6. The dialectal regional feature recognition method based on the multi-modal user interest vector calculation according to claim 5, characterized in that, In the process of constructing user behavior vectors, user behavior is used for modeling. The user behavior sequence is mapped into a low-dimensional dense vector through an embedding model, and the behavior vectors of different modalities are spliced ​​and weighted and fused.

7. The dialectal regional feature recognition method based on the multi-modal user interest vector calculation according to claim 1, characterized in that, The cluster analysis in step S3 includes the K-Means algorithm and normalization processing.

8. A dialect area feature-based content recommendation method based on a multi-modal user interest vector calculation, using the dialect area feature identification method of any one of claims 1-7 to identify dialect content, characterized in that, The specific steps include the following: Constructing a dialect content feature matrix: For platform content, identify and quantify the various dialect features it contains to form a dialect content feature matrix; Identify the dialect region characteristics of the target users to obtain the dialect community matrix of the target users; Fusion recommendation calculation: Based on the dialect weights in the dialect community matrix of the target user and the correlation strength between the content and each dialect in the dialect content feature matrix, calculate the recommendation probability of each content for the target user; Generate and output a recommendation list: Based on the recommendation probability, generate a recommendation list of dialect content for the target user; Data storage: Save the recommendation results and update the behavioral attributes of ordinary users. 9.The dialectal region-specific content recommendation method based on the multi-modal user interest vector computation of claim 8, characterized in that, First, use a natural language processing model to analyze dialect words in the text, and / or use a speech recognition model to analyze dialect speech features in the audio; The identified features are weighted using a preset dialect feature weight parameter set to obtain quantified dialect feature values ​​to construct the matrix. Subsequently, matrix operations were performed on the weight vectors in the dialect community matrix and the dialect content feature matrix, and combined with the operation metrics of the newness and popularity of the content, to obtain the comprehensive recommendation probability.

10. A dialect region feature recognition system based on multi-modal user interest vector calculation, which recognizes dialect content according to the recognition method of any one of claims 1-7, characterized in that, include: The data acquisition module is used to collect logs of interaction behaviors between platform users and dialect-related content; The feature extraction and vectorization module is used to construct and store the dialect content feature matrix and the seed user interest feature matrix. Based on the logs collected by the data acquisition module, construct a target user behavior feature matrix; The vector calculation and clustering module is used to perform similarity calculation and cluster analysis between the target user behavior feature matrix and the seed user interest feature matrix to generate the dialect community matrix of the target user. The application recommendation module calculates and outputs recommendation results for the target user based on the dialect community matrix and the dialect content feature matrix.

Citation Information

Patent Citations

  • Micro-video personalized recommendation algorithm based on social network credibility

    CN108920503A

  • Short video recommendation method based on label semantic similarity

    CN111382310A

  • Chinese and Tibetan multi-dialect and multi-mode resource recommendation method and device

    CN116089726A