A Multimodal Data-Driven Modeling Optimization Method in the Field of Social Work
Through the multimodal data-driven modeling optimization method in the social work field, the problems of multimodal data processing complexity and topic modeling instability are solved, and more efficient and accurate social work data analysis is achieved.
Patent Information
- Application Number
- CN202510424096.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art faces the problems of complexity, unstable theme modeling effects, dimensionality reduction and clustering instability, and inefficient analysis when processing multimodal social work data.
The multimodal data-driven modeling optimization method in the social work field includes building multimodal data sets, performing text embedding, using popularity bias regularization processors and dynamic document embedding optimizers, clustering through the HDBSCAN algorithm, and optimizing topic modeling through the c-TF-IDF and probability reassignment matrix methods.
It improves the fusion ability of multimodal data, enhances the accuracy and interpretability of topic modeling, ensures the stability and consistency of analysis results, and improves the data analysis ability of social work research.
Smart Images

Figure CN119939482B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a method for optimizing the modeling of the social work field driven by multi-modal data. Background Art
[0002] In recent years, social work research and strategy analysis have relied on multi-channel data collection, including social questionnaires, interview records, social work case reports, strategy documents, social media discussions, hotline voice data, community research videos, etc. These data provide key support for administrative departments, research institutions, and social organizations to deeply understand social issues and optimize the social service system. However, with the expansion of the data scale and the diversification of data sources, the existing analysis methods face the following challenges when processing this information: 1. The complexity of multi-modal data: Traditional social work research mainly relies on structured data and unstructured text, but modern social work data has gradually expanded to multi-modal data, including text data, voice data, image / video data. These different modal data contain rich social information, but their heterogeneity makes it difficult for traditional analysis methods to uniformly process, limiting the overall data utilization ability. 2. The limitations of topic modeling in social work data: Traditional topic modeling methods are mainly based on the bag-of-words model, which is difficult to process short texts, open-ended questionnaire answers, and social media discussions in social work data, resulting in unstable topic modeling effects; Social work data often contains a large number of high-frequency general words, such as "survey" and "problem", etc. These words may dominate the topic extraction process, covering up truly important social issues and affecting the specificity and interpretability of topic modeling. 3. The instability of dimensionality reduction and clustering: Topic modeling usually relies on dimensionality reduction techniques and clustering methods for text classification, but in social work data, different types of texts have large differences. Fixed-parameter dimensionality reduction methods may lead to information loss and unstable short text clustering effects. In addition, due to the small number of samples for some social issues, they may be judged as outliers during the clustering process and thus ignored, resulting in the analysis results lacking attention to the problems of marginalized groups. 4. The inefficiency and lack of adaptability of existing methods: Existing social work data analysis methods usually rely on manual collation and qualitative analysis, which are difficult to handle large-scale and multi-source data. The analysis process is time-consuming and subjective, affecting the efficiency of data-driven decision-making; Moreover, traditional topic modeling methods have poor adaptability to social work data, lacking mechanisms for enhancing social academic terms, retaining marginalized group issues, and automatically optimizing dimensionality reduction, resulting in poor interpretability and applicability of topic modeling results.
[0003] Therefore, social work data analysis faces challenges such as difficulties in multi-modal data fusion, limited topic modeling methods, instability in dimensionality reduction and clustering, and inefficiency of existing methods. There is an urgent need for more efficient, accurate, and intelligent technical means to enhance the data analysis ability of social work research and provide strong support for social strategy formulation, public service optimization, and social problem research. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide an optimized method for modeling in the field of social work driven by multi-modal data, which solves the deficiencies existing in the prior art.
[0005] The purpose of the present invention is achieved through the following technical solutions: An optimized method for modeling in the field of social work driven by multi-modal data, the optimization method includes:
[0006] S1. Construct a social work multi-modal data containing text, audio, and image / video data, and after data preprocessing, perform text embedding to convert the text into an embedding vector;
[0007] S2. Use a popularity bias regularization processor to process the data input in the field of social work through a dual optimization strategy of high-frequency common word attenuation weighting and social academic term enhancement;
[0008] S3. Dynamically select the dimensionality reduction dimension through a dynamic document embedding optimizer to optimize the effect of topic modeling, and use the UMAP algorithm to reduce the dimensionality of the vectors in the BERT embedding space;
[0009] S4. Use the HDBSCAN algorithm to cluster the documents into similar embedding groups to extract topics from them, form a hierarchical clustering structure by calculating the density relationship between data points, and divide the final clusters through a density threshold;
[0010] S5. Use c-TF-IDF to measure the importance of vocabulary by calculating the word frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus, and reassign outliers by calculating the semantic similarity through the probability reassignment matrix method to ensure the maximum utilization of data.
[0011] The high-frequency common word attenuation weighting includes the following:
[0012] Perform exponential attenuation on the top M high-frequency words in the corpus in terms of frequency, so that the influence of these high-frequency words gradually weakens during the topic modeling process;
[0013] The social academic term enhancement includes the following:
[0014] Construct a term dictionary related to social work, social strategy, and social governance, and set a multiple magnification of the weights of these terms during subsequent calculations.
[0015] The effects of optimizing topic modeling by dynamically selecting the dimensionality reduction dimension through a dynamic document embedding optimizer include:
[0016] Dimensional range setting: The dynamic document embedding optimizer sets the range of UMAP output dimensions to cover the semantic representation requirements of typical text embeddings, while avoiding information loss caused by inappropriate dimensions, enabling the model to retain key information in the text data during dimensionality reduction, thereby improving the quality and accuracy of clustering;
[0017] Silhouette coefficient evaluation: The dynamic document embedding optimizer evaluates each candidate dimension selected from the set dimensional range and calculates the silhouette coefficient of the corresponding clustering result;
[0018] Optimal dimension selection: By comparing the silhouette coefficients of different candidate dimensions, the dimension that maximizes is selected.
[0019] The specific content of S4 includes the following:
[0020] A1. Calculate the density reachability distance between two data points and ; ;
[0021] A2. Through the density reachability distance, construct a weighted complete graph and construct a minimum spanning tree for the complete graph to form a hierarchical eye clustering tree;
[0022] A3. By analyzing the persistence of each cluster in the hierarchical clustering tree, set the stability of each cluster as the sum of the weights of all edges in the cluster;
[0023] A4. Mark the outliers that cannot meet the density requirements as noise points;
[0024] A5. Divide the text into different clusters and a cluster containing noise points.
[0025] In S5, the importance of each word in the topic cluster is measured by calculating the term frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus through the formula , where , , is the category-weighted term frequency, is the inverse document frequency, is the combination of the category-weighted term frequency and the inverse document frequency, used to measure the importance of the word in the topic , is the word in the topic The number of occurrences in, is the subject The total word frequency of all words in is the total number of documents, is the number of documents containing the word .
[0026] In the above S5, the maximum marginal relevance search method is also combined to avoid selecting duplicate words and ensure the diversity of topic words. The calculation formula of the maximum marginal relevance search includes , where S is the set of topic keywords that have been selected currently, represents the topic center vector of the current topic, is the word and the topic between the similarities, is a hyperparameter used to control the balance between the relevance and diversity of keywords.
[0027] The method of calculating semantic similarity through the probability reassignment matrix to reassign outliers includes:
[0028] Construct a document - topic probability matrix: The probability reassignment matrix calculates the probability distribution of each document belonging to each topic. The probability distribution of each document represents the probability that the document belongs to each topic;
[0029] Calculate the probability of each document for each topic: Calculate the similarity between each document and each topic to obtain the probability that the document belongs to each topic;
[0030] Update the topic assignment of outliers: Determine whether a document should be reassigned to a certain topic according to the probability distribution of each document for each topic. If the probability of a document on a certain topic is high and exceeds the set threshold, the document is reassigned to the closest topic cluster.
[0031] The present invention has the following advantages: A multi - modal data - driven modeling optimization method for the social work field optimizes the multi - modal compatibility of the BERTopic model, supports extracting and fusing topic information from different data sources, and enhances the analysis depth of social issues; provides a more intelligent text optimization mechanism to ensure that topic modeling can highlight the important terms in the social work field, thereby improving the accuracy and interpretability of social problem analysis; optimizes the adaptive adjustment mechanism of dimensionality reduction parameters to ensure that different types of social work data can retain complete semantic information, improves the stability of topic clustering, and makes the analysis results more consistent and generalizable; optimizes the topic clustering strategy of social issues to ensure that the issues of these groups will not be automatically discarded by the model, but are effectively classified through the probability reassignment mechanism, improving the fairness and comprehensiveness of social issue research. Description of the Drawings
[0032] Figure 1 This is a schematic diagram of the process of the present invention. Specific embodiments
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are only some of the embodiments of the present application, rather than all of them. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the protection scope of the present application claimed, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application. The following further describes the present invention with reference to the accompanying drawings.
[0034] As Figure 1 shown, the present invention specifically relates to a method for optimizing the modeling in the field of social work driven by multimodal data based on BERTopic to improve the accuracy, interpretability, and robustness of data analysis, and provide intelligent technical support for social strategy research, social service optimization, and public governance decision-making. It specifically includes the following content:
[0035] Step 1: Multimodal data construction;
[0036] The purpose of this step is to construct a multimodal dataset of social work that includes text, audio, and image / video data, and perform standardization processing on it so as to uniformly input it into the topic modeling process, improving the availability of the data and the modeling effect.
[0037] (1) Text data: Text data is the main data source for social work research, including:
[0038] Social questionnaires: Open-ended question-and-answer data, reflecting the public's views on social issues.
[0039] Social worker reports: Text information recorded by social workers during case management, social surveys, and interventions.
[0040] Policy documents: Social policies, welfare measures, and legal and regulatory texts issued by administrative departments, non-profit organizations (NGOs), etc.
[0041] Social media discussions: Discussions related to Weibo, Xiaohongshu, Reddit, Twitter, etc., containing information such as the public's views on social issues and strategy feedback.
[0042] (2)Audio data: The voice data mainly comes from the following categories:
[0043] Interview records: Interview records conducted by social researchers, administrative departments or non-profit organizations on interviewees (such as social vulnerable groups, strategy beneficiaries).
[0044] Help hotline calls: The telephone recordings of mental health hotlines, social assistance hotlines, etc., which can reflect the real needs of the public in social issues.
[0045] Voice interpretations of strategies: Voice interpretations of social strategies by strategy formulators or experts, such as press conferences, radio interviews, etc.
[0046] (3)Image / video data: The image and video data in the field of social work are mainly used to record social surveys and strategy implementation situations, including:
[0047] Community research images, strategy promotion materials and screenshots of news reports.
[0048] Step 2: Data preprocessing;
[0049] (1)Text data preprocessing;
[0050] Stop word removal: Use the stop word dictionaries of natural language processing tools such as NLTK and Jieba to remove high-frequency words that are semantically irrelevant and filter the stop words from the tweet text.
[0051] Duplicate removal: Remove texts with too high similarity through text hash values (MinHash) and semantic similarity (BERT cosine similarity) to avoid data redundancy.
[0052] Processing of emojis and links: Use regular expressions (Regex) to detect and remove this noise content. For emojis, detect the Unicode emoji range and replace it with an empty character; for URL links, match the strings starting with "http(s): / / " or "www" and remove the links.
[0053] (2)Voice data processing: Since voice data cannot be directly input into the topic modeling process, it needs to be converted into text first, and at the same time, clean the noise that may be generated during the conversion process.
[0054] Speech to text: Use Whisper (OpenAI) or Wav2Vec2.0 (Facebook AI) for automatic speech recognition (ASR) to convert interview recordings and help hotline audio into text. Text cleaning: Use TextBlob and SymSpell to correct the spelling of the text generated by ASR to ensure the accuracy of the converted text; remove voice stop words. Remove misrecognized words and background noise in the voice signal to improve the quality of ASR.
[0055] (3)Image / Video Data Processing: The text information in images and video data needs to be extracted for topic modeling.
[0056] OCR Text Extraction: Use Tesseract OCR or Easy OCR to extract text from pictures. Cleaning OCR Errors: OCR may misidentify characters, and the Levenshtein edit distance can be used for spelling correction.
[0057] Video Subtitle Processing: ASR Speech-to-Text: For videos with speech, use Whisper / Wav2Vec2.0 to extract the audio and convert it to text; Timeline Alignment: Align the text extracted by ASR to the timeline to ensure that the text content of different segments is consistent with the video content; OCR Extract Text from Video Frames: If the video frame contains text, use OCR to extract the text information and perform deduplication and cleaning.
[0058] (4)Unified Format after Multimodal Data Preprocessing;
[0059] Finally, after the preprocessing of the above text, audio, and image / video, all data will be converted into a standard text format. After the above multimodal data collection and preprocessing, all data is standardized into text format and can be uniformly input into the BERTopic topic modeling process to ensure that the model can be compatible with different data sources and improve the mining ability of social issues.
[0060] Step Three: Text Embedding;
[0061] Use the paraphrase-multilingual-MiniLM-L12-v2 embedding model to convert the text into embedding vectors. An embedding vector is a vector represented by mapping complex, high-dimensional, or sparse data (such as text, images, categorical features, etc.) to a low-dimensional, dense vector space through an embedding function. This operation maps the text to a low-dimensional vector space and avoids the curse of dimensionality problem that may occur in high-dimensional vector models.
[0062] Specifically, assume the text contains n terms, where the BERT vector corresponding to each term is . The vector of the text can be calculated by averaging the BERT vectors of all terms, that is: , where is the topic The i-th word in is the word vector representation, is the vector representation of the text, is the BERT embedding vector of the entry.
[0063] Step 4: Perform Popularity Deviation Regularizer (PDR) processing;
[0064] In text data such as social work questionnaires, social worker reports, strategy documents, and social media data, there are often a large number of high-frequency general words. These words have a high frequency of occurrence in the corpus, but they do not have sufficient topic characteristics. If these words dominate the topic modeling process, it may lead to the extracted topics being too broad and lacking fine-grained discrimination ability, thus affecting the accuracy and interpretability of the topic model.
[0065] To solve this problem, the present invention proposes a Popularity Deviation Regularization (PDR) mechanism to process the data input in the field of social work through a dual optimization strategy.
[0066] Among them, high-frequency general word attenuation weighting: reducing the interference of irrelevant general words to topic modeling and increasing the model's attention to important social issues.
[0067] Social academic term enhancement: enhancing the weight of professional terms in the fields of social strategies and social work during the topic modeling process to ensure that key topics are not overlooked.
[0068] (1) High-frequency general word attenuation weighting;
[0069] High-frequency general words in social work data usually belong to background vocabulary. Their frequent occurrence may cause topic modeling to tend towards non-professional content and ignore truly important social issues. For example, in social questionnaires, words such as "survey", "respondent", and "strategy" may appear in almost all texts, but these words themselves do not carry actual social issue information.
[0070] To solve this problem, PDR reduces its impact on the model by implementing attenuation weighting for high-frequency words. The specific method is: exponentially attenuating the high-frequency words ranked in the top 10% of the frequency in the corpus, so that the influence of these high-frequency words gradually weakens during the topic modeling process. Through attenuation processing, the interference of these words can be effectively reduced, preventing them from dominating the model.
[0071] (2) Social academic term enhancement;
[0072] In social work data, in addition to high-frequency general words, some important domain terms (such as "childneglect", "food insecurity", "domestic violence", "housing insecurity") may appear less frequently in the overall data set, but play an important role in indicating specific topics. If these terms are not specially processed, they may be ignored in the topic modeling process, resulting in the loss or ambiguity of key social issues.
[0073] Construct a dictionary of terms related to social work, social strategy, and social governance
[0074] In the subsequent c-TF-IDF calculation stage, the weights of these terms are amplified by 1.5 times.
[0075] Popularity bias regularization (PDR) optimizes the text input of social work data and improves the accuracy and interpretability of topic modeling through two mechanisms: high-frequency common word attenuation weighting and sociological term enhancement. This optimization mechanism ensures that BERTopic can more accurately identify core issues such as social strategies, social equity, and child protection, providing more powerful data support for social work research and strategy analysis.
[0076] Step 5: Perform dynamic document embedding optimizer (DDEO) processing;
[0077] The UMAP process (Uniform Manifold Approximation and Projection) is a nonlinear dimensionality reduction algorithm used to map high-dimensional data to a low-dimensional space while maintaining the local and global structure of the data as much as possible.
[0078] In order to solve the problem of large fluctuations in semantic information loss during the UMAP dimensionality reduction process, this paper proposes a dynamic document embedding optimizer (DDEO) to optimize the effect of topic modeling by dynamically selecting the dimension of UMAP dimensionality reduction. Specifically, DDEO aims to adaptively select the best dimension of UMAP dimensionality reduction, thereby maximizing the retention of the semantic information of the text and improving the final topic differentiation and explanatory power. The implementation process of DDEO is as follows:
[0079] (1) Dimension range setting;
[0080] First, to ensure the adaptability and optimization of UMAP dimensionality reduction, DDEO presets the range of UMAP output dimensions. , it is generally selected to be between. This dimensional range can effectively cover the semantic representation requirements of typical text embeddings, while avoiding information loss or overfitting problems caused by too high or too low dimensions. Selecting an appropriate dimensional range can enable the model to retain as much key information as possible in the text data during the dimensionality reduction process, thereby improving the quality and accuracy of clustering.
[0081] (2) Silhouette coefficient evaluation;
[0082] During the selection process, DDEO evaluates each candidate dimension and calculates the silhouette coefficient of the corresponding clustering result. This metric can reflect the clustering quality of the data at different dimensions. The formula for calculating the silhouette coefficient is:
[0083] ,
[0084] where N is the total number of all samples in the dataset, is the i-th sample, is the average distance from the sample to other points in the same cluster, is the minimum average distance to the nearest other cluster.
[0085] (3) Optimal dimension selection;
[0086] DDEO selects the dimension that maximizes the silhouette coefficient by comparing the silhouette coefficients of different candidate dimensions:
[0087] ,
[0088] where, is the d value that maximizes S(d), that is, find the d value when maximizing the objective function S(d).
[0089] This selection ensures the semantic integrity of the text data after dimensionality reduction and improves the stability of clustering and the consistency of themes. The selection of the optimal dimension can not only reduce information loss but also enhance the interpretability of themes and the generalization ability of the model.
[0090] Step Six: Text dimensionality reduction;
[0091] Use the UMAP algorithm to reduce the dimensionality of the vectors in the BERT embedding space. This algorithm can effectively retain the local and global features of high-dimensional data.
[0092] ,
[0093] where, is the text embedding vector without dimensionality reduction, is the text embedding vector after dimensionality reduction.
[0094] Step Seven: Text clustering;
[0095] Use the HDBSCAN algorithm (density-based hierarchical clustering algorithm) to cluster documents into similar embedding groups for extracting themes. By calculating the density relationship between data points, a hierarchical clustering structure is formed, and the final clusters are divided by a density threshold. The clustering label for each text is obtained through this density calculation, thus reflecting the similarity and potential theme structure among texts. The specific implementation is as follows:
[0096] (1) Calculate the mutual reachability distance between two points as:
[0097] ,
[0098] where, represents and the Euclidean distance between represents the distance to the farthest point among its nearest k neighbors, represents the distance to the farthest point among its nearest k neighbors, and k is a parameter of the minimum sample number (MinSamples) (i.e., the MinSamples parameter set by HDBSCAN).
[0099] (2) Through the mutual reachability distance, construct a weighted complete graph G-(V,E), where the point set V is the documents and the edge weight E is the mutual reachability distance. Then construct a minimum spanning tree (MST) for this graph to form a hierarchical dendrogram. The weight of the minimum spanning tree is as follows:
[0100] ,
[0101] where, is the mutual reachability distance between two points.
[0102] (3) By analyzing the persistence of each cluster in the hierarchical clustering tree, define the stability of each cluster is the sum of the weights of all edges in the cluster.
[0103] ,
[0104] where, and represent the distance thresholds for the appearance and disappearance of the cluster respectively.
[0105] (4) HDBSCAN will mark the outliers that do not meet the density requirements (i.e., the clusters with insufficient core points) as noise points. The processing of outliers is based on the local density relationship of the core distance:
[0106] ,
[0107] Among them, A higher value means that this point is more likely to be noise, is the data point to the distance of the farthest point among its nearest k neighbors (the number of neighbors set by MinSamples), is the average of the core distances of all points in the cluster to which this point belongs.
[0108] (5) HDBSCAN finally divides the text into different clusters , and a cluster containing noise points , and the clustering label assignment is:
[0109] ,
[0110] Among them, represents the kth cluster, is the different texts in the dataset.
[0111] Step Eight, Topic Representation;
[0112] The c-TF-IDF method is adopted to evaluate the importance of words in each topic. Specifically, c-TF-IDF measures the importance of words by calculating the term frequency (TF) of each word in the topic cluster and the inverse document frequency (IDF) of this word in the entire corpus. The formula is as follows:
[0113] ,
[0114] ,
[0115] ,
[0116] Among them, is the class-based term frequency, is the inverse document frequency, is the combination of class-based term frequency (c-TF) and inverse document frequency (IDF), used to measure the word in the topic importance, is the word in the topic appearance times in, is the topic total term frequency of all words in, is the total number of documents, is the number of documents containing the word
[0117] In addition, to extract the most representative topic keywords, the present invention combines the maximum marginal relevance search (MMR) method to avoid selecting duplicate words and ensure the diversity of topic words.
[0118] ,
[0119] where S is the set of topic keywords that have been selected currently, represents the topic center vector of the current topic, is the word and the topic between the similarity, is a hyperparameter that controls the balance between the relevance of keywords (similarity to the topic) and the diversity (similarity to other keywords), is to find the one that is most similar to the current candidate word , that is, the maximum value of its similarity, is a keyword in the set S of selected keywords.
[0120] Step Nine: Perform probability redistribution matrix (PRM) redistribution;
[0121] In the traditional BERTopic model, the HDBSCAN clustering algorithm will automatically identify some data points as outliers and exclude them from the topics. However, in applications such as social work questionnaire surveys, policy document analysis, and social media public feedback information monitoring, these outliers may represent potential important social issues. If simply discarded, it may lead to the loss of key information such as minority group issues and special social phenomena. To solve this problem, the present invention proposes the probability redistribution matrix (PRM) method, which redistributes outliers through semantic similarity calculation to ensure the maximum utilization of data and improve the coverage and robustness of topic modeling.
[0122] Furthermore, PRM redistributes outliers through the following steps:
[0123] (1) Construct a document-topic probability matrix;
[0124] First, PRM calculates the probability distribution of each document belonging to each topic. This process is based on the HDBSCAN clustering results, where the probability distribution of each document represents the probability that the document belongs to each topic. For each document, the sum of the probabilities of all topics is 1, that is .
[0125] Among them, M is the total number of topics, representing the distribution of each document among all topics. For documents marked as outliers, PRM will calculate their similarity to each topic and probabilistically reassign them to the closest topic.
[0126] (2) Calculate the probability of each document for each topic;
[0127] The PRM method obtains the probability that a document belongs to each topic by calculating the similarity between each document and each topic. Specifically, PRM uses a distance metric function to measure the similarity between a document and a topic. Suppose the similarity between a document and a topic is represented by PRM uses the following formula to calculate the probability that a document belongs to a topic:
[0128] ,
[0129] where is the distance between the current document and a specific topic , is the distance between the current document and all topics , which is used to measure the matching degree between the document and each topic. The smaller the distance, the more relevant it is, and finally it is used to be converted into a probability through Softmax.
[0130] (3) Update the topic assignment of outliers;
[0131] Once the topic probability distribution of each document is calculated, PRM will determine whether a certain document should be reassigned to a certain topic according to the probability distribution of each document for each topic, especially those documents marked as outliers by HDBSCAN. If the probability of an outlier document on a certain topic is high and exceeds the set threshold, the document will be reassigned to the closest topic cluster. By this method, PRM can not only retain the information of outliers, but also improve the model's ability to identify potential important topics.
[0132] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and improvements, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. And the changes and alterations made by those skilled in the art without departing from the spirit and scope of the present invention shall all fall within the protection scope of the appended claims of the present invention.
Claims
1. A multimodal data-driven social work field modeling optimization method, characterized by: The optimization method comprises: S1. Construct a social work multimodal data set containing text, audio, and image / video data, and perform text embedding after data preprocessing to convert the text into an embedding vector; S2, using the popularity bias regularization processor to process the data input in the social work field through the dual optimization strategy of high-frequency common word attenuation weighting and sociological term enhancement; S3, dynamically select the dimension reduction dimension through the dynamic document embedding optimizer to optimize the effect of topic modeling, and use the UMAP algorithm to reduce the dimension of the vector in the BERT embedding space; S4. Use the HDBSCAN algorithm to cluster documents into similar embedding groups in order to extract topics from them, form a hierarchical clustering structure by calculating the density relationship between data points, and divide the final clusters by density thresholds; S5. The importance of vocabulary is measured by calculating the word frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus through c-TF-IDF, and the semantic similarity is calculated by the probability redistribution matrix method to redistribute outliers to ensure maximum utilization of data; The method of calculating semantic similarity by a probability redistribution matrix method to redistribute outliers includes: Constructing the document-topic probability matrix: The probability redistribution matrix calculates the probability distribution of each document belonging to each topic. The probability distribution of each document represents the probability that the document belongs to each topic. For each document, the sum of the probabilities of all topics is 1, that is, , where M is the total number of topics, indicating the distribution of each document among all topics; Calculate the probability of each document for each topic: For documents marked as outliers, the probability redistribution matrix calculates the similarity between each document and each topic and probabilistically redistributes it to the closest topic. The probability redistribution matrix uses a distance metric function to measure the similarity between documents and topics. Let the similarity between documents and topics be expressed as To express it, PRM uses the following formula to calculate the probability that a document belongs to a topic: , and then get the probability that the document belongs to each topic, where Is the current document With a specific topic The distance The current document and all topics The distance is used to measure the matching degree between the document and each topic. The smaller the distance, the more relevant it is. It is finally used to convert into probability through Softmax. Update the topic assignment of outliers: Once the topic probability distribution of each document is calculated, the probability redistribution matrix determines whether the document should be reallocated to a topic based on the probability distribution of each document to each topic, especially those documents marked as outliers by HDBSCAN. If the probability of a document on a topic is high and exceeds the set threshold, the document is reallocated to the closest topic cluster.
2. A multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The high-frequency common word attenuation weighting includes the following: Exponential decay is performed on the top M high-frequency words in the corpus, so that the influence of these high-frequency words gradually weakens during the topic modeling process; The Sociology Terminology Enhancement includes the following: Construct a dictionary of terms related to social work, social strategy and social governance, and amplify the weights of these terms by set multiples in the subsequent calculation process.
3. The multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The effects of optimizing topic modeling by dynamically selecting the dimension reduction dimension through the dynamic document embedding optimizer include: Dimension range setting: The dynamic document embedding optimizer sets the range of UMAP output dimensions to cover the semantic representation requirements of typical text embeddings while avoiding information loss caused by inappropriate dimensions, so that the model retains key information in the text data during dimensionality reduction, thereby improving the quality and accuracy of clustering; Silhouette coefficient evaluation: The dynamic document embedding optimizer evaluates each candidate dimension selected from the set dimension range and calculates the silhouette coefficient of the corresponding clustering result; Optimal dimension selection: By comparing the silhouette coefficients of different candidate dimensions, the dimension with the maximum value is selected.
4. The multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The S4 specifically includes the following contents: A1. Calculate two data points and Density accessibility distance ,in, express and The Euclidean distance between express The distance to the farthest point among its k nearest neighbors, express The distance to the farthest point among its k nearest neighbors; A2. Construct a weighted complete graph through density reachability distance, and construct a minimum spanning tree for the complete graph to form a hierarchical eye tree; A3. By analyzing the persistence of each cluster in the hierarchical clustering tree, the stability of each cluster is set to be the sum of the weights of all edges in the cluster; A4. Mark outliers that fail to meet density requirements as noise points; A5. Divide the text into different clusters and a cluster containing noise points.
5. The multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The formula in S5 is To calculate the frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus to measure the importance of the word, where , , is the category-weighted term frequency, is the inverse document frequency, It is a combination of category-weighted word frequency and inverse document frequency to measure word On topic The importance of Yes word On topic The number of occurrences in Is the theme The total frequency of all words in , is the total number of documents, Contains words The number of documents.
6. A multimodal data-driven social work field modeling optimization method according to claim 5, characterized in that: The S5 also combines the maximum marginal relevance search method to avoid selecting repeated words and ensure the diversity of subject words. The calculation formula of the maximum marginal relevance search includes: , where S is the currently selected topic keyword set, The topic center vector representing the current topic, It's a word and Theme The similarity between is a hyperparameter that controls the balance between keyword relevance and diversity. Indicates finding words that match the current candidate The value with the largest similarity is It is a keyword in the selected keyword set S.
Citation Information
Patent Citations
Theme modeling and sentiment analysis method and system based on deep learning
CN118394944A
Graph network infringement behavior classification method based on text expansion and label information fusion
CN119046461A