Social work field modeling optimization method driven by multi-modal data
Through the multimodal data-driven modeling optimization method in the social work field, the problems of multimodal data processing complexity and instability in the topic modeling in the existing technology are solved, and more efficient and accurate social work data analysis is achieved.
Patent Information
- Application Number
- CN202510424096.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art faces the problems of complexity, instability in thematic modeling, dimensionality reduction and clustering instability, and inefficient analysis when processing multimodal social work data.
The multimodal data-driven modeling and optimization method in the social work field is adopted, and the topic modeling and clustering effects are optimized by constructing multimodal data sets, performing text embedding, using popularity deviation regularization processors, dynamic document embedding optimizers, HDBSCAN algorithms and probability reassignment matrix methods.
It improves the ability to integrate and utilize multimodal data, enhances the accuracy and interpretability of topic modeling, ensures the stability and consistency of analysis results, and improves the efficiency and adaptability of social work data analysis.
Smart Images

Figure CN119939482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a multimodal data-driven modeling optimization method in the field of social work. Background Art
[0002] In recent years, social work research and strategy analysis have relied on multi-channel data collection, including social survey questionnaires, interview records, social work case reports, strategy documents, social media discussions, hotline voice data, community survey images, etc. These data provide key support for administrative departments, research institutions and social organizations to deeply understand social issues and optimize the social service system. However, with the expansion of data scale and the diversification of data sources, existing analysis methods face the following challenges in processing this information: 1. The complexity of multimodal data: Traditional social work research mainly relies on structured data and unstructured text, but modern social work data has gradually expanded to multimodal data, including text data, voice data, image / video data. These different modal data contain rich social information, but their heterogeneity makes it difficult for traditional analysis methods to process them in a unified manner, limiting the ability to utilize the data in aggregate. 2. Limitations of topic modeling on social work data: Traditional topic modeling methods are mainly based on bag-of-words models, which are difficult to handle short texts, open-ended questionnaire answers, and social media discussions in social work data, resulting in unstable topic modeling results; social work data often contain a large number of high-frequency common words, such as "survey" and "problem", etc. These words may dominate the topic extraction process, cover up truly important social issues, and affect the specificity and interpretability of topic modeling. 3. Instability of dimensionality reduction and clustering: Topic modeling usually relies on dimensionality reduction technology and clustering methods for text classification, but in social work data, different types of text have large differences. Fixed parameter dimensionality reduction methods may lead to information loss and unstable clustering effects of short texts. In addition, some social issues may be judged as outliers in the clustering process due to the small number of samples, and thus be ignored, resulting in the lack of attention to marginal group issues in the analysis results. 4. Insufficient efficiency and adaptability of existing methods: Existing social work data analysis methods usually rely on manual collation and qualitative analysis, which are difficult to cope with large-scale, multi-source data. The analysis process is time-consuming and highly subjective, affecting the efficiency of data-driven decision-making; and traditional topic modeling methods have poor adaptability to social work data, lacking mechanisms for enhancing sociological terminology, retaining marginal group issues, and automatic dimensionality reduction optimization, resulting in poor interpretability and applicability of topic modeling results.
[0003] Therefore, social work data analysis faces challenges such as difficulty in multimodal data fusion, limitations of topic modeling methods, instability of dimensionality reduction and clustering, and low efficiency of existing methods. There is an urgent need for more efficient, accurate, and intelligent technical means to enhance the data analysis capabilities of social work research and provide strong support for social strategy formulation, public service optimization, and social problem research. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a multimodal data-driven social work field modeling optimization method to solve the deficiencies of the prior art.
[0005] The purpose of the present invention is achieved through the following technical solution: a multimodal data-driven social work field modeling optimization method, the optimization method comprising: S1. Build a social work multimodal data set containing text, audio, and image / video data, and perform text embedding after data preprocessing to convert text into embedding vectors; S2, using the popularity bias regularization processor to process the data input in the social work field through the dual optimization strategy of high-frequency common word attenuation weighting and sociological term enhancement; S3, dynamically select the dimension reduction dimension through the dynamic document embedding optimizer to optimize the effect of topic modeling, and use the UMAP algorithm to reduce the dimension of the vector in the BERT embedding space; S4. Use the HDBSCAN algorithm to cluster documents into similar embedding groups in order to extract topics from them, form a hierarchical clustering structure by calculating the density relationship between data points, and divide the final clusters by density thresholds; S5. c-TF-IDF is used to measure the importance of vocabulary by calculating the word frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus, and the semantic similarity is calculated by the probability redistribution matrix method to redistribute outliers to ensure maximum utilization of data.
[0006] The high-frequency common word attenuation weighting includes the following: Exponential decay is performed on the top M high-frequency words in the corpus, so that the influence of these high-frequency words gradually weakens during the topic modeling process; The Sociology Terminology Enhancement includes the following: Construct a dictionary of terms related to social work, social strategy and social governance, and amplify the weights of these terms by set multiples in the subsequent calculation process.
[0007] The effects of optimizing topic modeling by dynamically selecting the dimension reduction dimension through the dynamic document embedding optimizer include: Dimension range setting: The dynamic document embedding optimizer sets the range of UMAP output dimensions to cover the semantic representation requirements of typical text embeddings while avoiding information loss caused by inappropriate dimensions, so that the model retains key information in the text data during dimensionality reduction, thereby improving the quality and accuracy of clustering; Silhouette coefficient evaluation: The dynamic document embedding optimizer evaluates each candidate dimension selected from the set dimension range and calculates the silhouette coefficient of the corresponding clustering result; Optimal dimension selection: By comparing the silhouette coefficients of different candidate dimensions, the dimension with the maximum value is selected.
[0008] The S4 specifically includes the following contents: A1. Calculate two data points and Density accessibility distance ; A2. Construct a weighted complete graph through density reachability distance, and construct a minimum spanning tree for the complete graph to form a hierarchical eye tree; A3. By analyzing the persistence of each cluster in the hierarchical clustering tree, the stability of each cluster is set to be the sum of the weights of all edges in the cluster; A4. Mark outliers that fail to meet density requirements as noise points; A5. Divide the text into different clusters and a cluster containing noise points.
[0009] The formula in S5 is To calculate the frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus to measure the importance of the word, where , , is the category-weighted term frequency, is the inverse document frequency, It is a combination of category-weighted word frequency and inverse document frequency to measure word In Theme The importance of Yes word In Theme The number of occurrences in Is the theme The total frequency of all words in , is the total number of documents, Contains words The number of documents.
[0010] The S5 also combines the maximum marginal relevance search method to avoid selecting repeated words and ensure the diversity of subject words. The calculation formula of the maximum marginal relevance search includes: , where S is the currently selected topic keyword set, The topic center vector representing the current topic, It's a word and Theme The similarity between is a hyperparameter that controls the balance between keyword relevance and diversity.
[0011] The method of calculating semantic similarity by a probability redistribution matrix method to redistribute outliers includes: Constructing the document-topic probability matrix: The probability redistribution matrix calculates the probability distribution of each document belonging to each topic. The probability distribution of each document represents the probability that the document belongs to each topic. Calculate the probability of each document for each topic: Calculate the similarity between each document and each topic to get the probability that the document belongs to each topic; Update the topic assignment of outliers: Determine whether the document should be reallocated to a topic based on the probability distribution of each document to each topic. If the probability of a document being on a topic is high and exceeds the set threshold, the document is reallocated to the closest topic cluster.
[0012] The present invention has the following advantages: a multimodal data-driven social work field modeling optimization method, which optimizes the multimodal compatibility of the BERTopic model, supports the extraction and fusion of topic information from different data sources, and enhances the depth of analysis of social issues; provides a more intelligent text optimization mechanism to ensure that topic modeling can highlight important terms in the field of social work, thereby improving the accuracy and interpretability of social problem analysis; optimizes the adaptive adjustment mechanism of dimensionality reduction parameters to ensure that different types of social work data can retain complete semantic information, improve the stability of topic clustering, and make the analysis results more consistent and generalizable; optimizes the topic clustering strategy of social issues to ensure that the topics of these groups will not be automatically discarded by the model, but will be effectively classified through the probability redistribution mechanism, thereby improving the fairness and comprehensiveness of social issue research. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided below in conjunction with the drawings is not intended to limit the scope of protection of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application. The present invention is further described below in conjunction with the drawings.
[0015] like Figure 1 As shown, the present invention specifically relates to a social work field modeling optimization method driven by multimodal data based on BERTopic, so as to improve the accuracy, interpretability and robustness of data analysis, and provide intelligent technical support for social strategy research, social service optimization, and public governance decision-making. It specifically includes the following contents: Step 1: Multimodal data construction; This step aims to construct a multimodal social work dataset containing text, audio, and image / video data, and standardize it so that it can be uniformly input into the topic modeling process to improve data availability and modeling effects.
[0016] (1) Text data: Text data is the main data source for social work research, including: Social survey questionnaires: open-ended question and answer data that reflect the public’s views on social issues.
[0017] Social work report: text information recorded by social workers during case management, social investigation and intervention.
[0018] Strategy documents: social strategies, welfare measures, and legal and regulatory texts issued by administrative departments, non-profit organizations (NGOs), etc.
[0019] Social media discussions: such as Weibo, Xiaohongshu, Reddit, and Twitter, including information such as the public’s views on social issues and strategic feedback.
[0020] (2) Audio data: Voice data mainly comes from the following categories: Interview records: records of interviews conducted by social researchers, administrative departments or non-profit organizations with respondents (such as vulnerable groups in society, beneficiaries of strategies).
[0021] Helpline calls: Recordings of phone calls to helpline hotlines, such as mental health hotlines and social assistance hotlines, can reflect the public’s real needs in social issues.
[0022] Strategy interpretation audio: audio interpretation of social strategies by strategists or experts, such as press conferences, radio interviews, etc.
[0023] (3) Image / video data: Image and video data in the field of social work are mainly used to record social surveys and strategy implementation, including: Community research images, strategy communications materials, and news coverage screenshots.
[0024] Step 2: Data preprocessing; (1) Text data preprocessing; Remove stop words: Use the stop word dictionary of natural language processing tools such as NLTK and Jieba to remove semantically irrelevant high-frequency words and filter the stop words in the tweet text.
[0025] Deduplication: Use text hash value (MinHash) and semantic similarity (BERT cosine similarity) to remove texts with too high similarity to avoid data redundancy.
[0026] Processing emoticons and links: Use regular expressions (Regex) to detect and remove these noise contents. For emoticons, detect the Unicode emoticon range and replace it with empty characters. For URL links, match strings starting with "http(s): / / " or "www" and remove the links.
[0027] (2) Speech data processing: Since speech data cannot be directly input into the topic modeling process, it must first be converted into text and the noise that may be generated during the conversion process must be cleaned up.
[0028] Speech-to-text: Use Whisper (OpenAI) or Wav2Vec2.0 (Facebook AI) for automatic speech recognition (ASR) to convert interview recordings and helpline audio to text. Text cleaning: Use TextBlob and SymSpell to correct the spelling of the text generated by ASR to ensure the accuracy of the converted text; remove speech stop words. Remove misrecognized words and background noise in the speech signal to improve ASR quality.
[0029] (3) Image / video data processing: Text information in image and video data needs to be extracted for topic modeling.
[0030] OCR text extraction: Use Tesseract OCR or Easy OCR to extract text from images. Clean up OCR errors: OCR may misrecognize characters, and Levenshtein edit distance can be used for spelling correction.
[0031] Video subtitle processing: ASR speech to text: For videos with speech, use Whisper / Wav2Vec2.0 to extract audio and convert it into text; Timeline alignment: Align the text extracted by ASR to ensure that the text content of different clips is consistent with the video content; OCR extracts text from video frames: If the video contains text, use OCR to extract the text information, deduplicate it, and clean it up.
[0032] (4) Unified format after multimodal data preprocessing; Finally, after the above-mentioned preprocessing of text, audio, and images / video, all data will be converted into a standard text format. After the above-mentioned multimodal data collection and preprocessing, all data are standardized into text format and can be uniformly input into the BERTopic topic modeling process to ensure that the model is compatible with different data sources and improve the ability to mine social issues.
[0033] Step 3: Text embedding; The paraphrase-multilingual-MiniLM-L12-v2 embedding model is used to convert text into an embedding vector. An embedding vector is a vector that maps complex, high-dimensional or sparse data (such as text, images, classification features, etc.) to a low-dimensional, dense vector space through an embedding function. This operation maps text to a low-dimensional vector space while avoiding the dimensionality disaster problem that may occur in high-dimensional vector models.
[0034] Specifically, assuming the text Contains n terms, where the BERT vector corresponding to each term is .text Vector The representation can be calculated by averaging the BERT vectors of all terms, namely: ,in, Is the theme The i-th word in Yes word The vector representation of is the vector representation of the text, is the BERT embedding vector of the term.
[0035] Step 4: Perform popularity deviation regularizer (PDR) processing; Text data such as social work questionnaires, social work reports, strategy documents, and social media data often contain a large number of high-frequency common words. These words appear frequently in the corpus, but they do not have sufficient topic characteristics. If these words dominate the topic modeling process, the extracted topics may be too broad and lack fine-grained differentiation capabilities, thus affecting the accuracy and interpretability of the topic model.
[0036] To solve this problem, this paper proposes a popularity deviation regularization (PDR) mechanism to process data input in the field of social work through a dual optimization strategy.
[0037] Among them, high-frequency common words are attenuated and weighted: reducing the interference of irrelevant common words on topic modeling and increasing the model's attention to important social issues.
[0038] Enhanced sociological terminology: Increase the weight of professional terms in fields such as social strategy and social work in the topic modeling process to ensure that key topics are not overlooked.
[0039] (1) Attenuation weighting of high-frequency common words; High-frequency common words in social work data are usually background words. Their frequent appearance may cause topic modeling to be biased towards non-professional content and ignore truly important social issues. For example, in a social survey questionnaire, the words "survey", "respondent", and "strategy" may appear in almost all texts, but these words themselves do not contain actual social issue information.
[0040] To solve this problem, PDR reduces the influence of high-frequency words on the model by implementing attenuation weighting on high-frequency words. The specific method is to perform exponential attenuation on the top 10% of the high-frequency words in the corpus, so that the influence of these high-frequency words gradually weakens during the topic modeling process. Through attenuation processing, the interference of these words can be effectively reduced to prevent them from dominating the model.
[0041] (2) Enhancement of sociological terminology; In social work data, in addition to high-frequency general words, some important domain terms (such as "childneglect", "food insecurity", "domestic violence", "housing insecurity") may appear less frequently in the overall data set, but play an important role in indicating specific topics. If these terms are not specially processed, they may be ignored in the topic modeling process, resulting in the loss or ambiguity of key social issues.
[0042] Construct a dictionary of terms related to social work, social strategy, and social governance In the subsequent c-TF-IDF calculation stage, the weights of these terms are amplified by 1.5 times.
[0043] Popularity bias regularization (PDR) optimizes the text input of social work data and improves the accuracy and interpretability of topic modeling through two mechanisms: high-frequency common word attenuation weighting and sociological term enhancement. This optimization mechanism ensures that BERTopic can more accurately identify core issues such as social strategies, social equity, and child protection, providing more powerful data support for social work research and strategy analysis.
[0044] Step 5: Perform dynamic document embedding optimizer (DDEO) processing; The UMAP process (Uniform Manifold Approximation and Projection) is a nonlinear dimensionality reduction algorithm used to map high-dimensional data to a low-dimensional space while maintaining the local and global structure of the data as much as possible.
[0045] In order to solve the problem of large fluctuations in semantic information loss during the UMAP dimensionality reduction process, this paper proposes a dynamic document embedding optimizer (DDEO) to optimize the effect of topic modeling by dynamically selecting the dimension of UMAP dimensionality reduction. Specifically, DDEO aims to adaptively select the best dimension of UMAP dimensionality reduction, thereby maximizing the retention of the semantic information of the text and improving the final topic differentiation and explanatory power. The implementation process of DDEO is as follows: (1) Dimension range setting; First, to ensure the adaptability and optimization of UMAP dimensionality reduction, DDEO presets the range of UMAP output dimensions. , generally between . This dimensional range can effectively cover the semantic representation requirements of typical text embeddings, while avoiding information loss or overfitting problems caused by too high or too low dimensions. Choosing a suitable dimensional range can enable the model to retain as much key information in the text data as possible during the dimensionality reduction process, thereby improving the quality and accuracy of clustering.
[0046] (2) Silhouette coefficient evaluation; During the selection process, DDEO evaluates each candidate dimension and calculates the silhouette coefficient of the corresponding clustering result. This indicator can reflect the clustering quality of data in different dimensions. The calculation formula of the silhouette coefficient is: , Where N is the total number of samples in the dataset. is the i-th sample, is the average distance from the sample to other points in the same cluster, is the minimum average distance to the nearest other cluster.
[0047] (3) Optimal dimension selection; DDEO compares the silhouette coefficients of different candidate dimensions and selects the dimension that maximizes the value: , in, It is the value of d that maximizes S(d), that is, find the value of d that maximizes the objective function S(d).
[0048] This choice ensures the semantic integrity of the text data after dimensionality reduction and improves the stability of clustering and the consistency of topics. The selection of the optimal dimension can not only reduce information loss, but also enhance the interpretability of the topic and the generalization ability of the model.
[0049] Step 6: Text dimensionality reduction; The UMAP algorithm is used to reduce the dimensionality of the vectors in the BERT embedding space. This algorithm can effectively retain the local and global features of high-dimensional data.
[0050] , in, is the text embedding vector without dimensionality reduction, is the text embedding vector after dimensionality reduction.
[0051] Step 7: Text clustering; The HDBSCAN algorithm (density-based hierarchical clustering algorithm) is used to cluster documents into similar embedding groups in order to extract topics from them. By calculating the density relationship between data points, a hierarchical clustering structure is formed, and the final clusters are divided by density thresholds. The cluster label of each text is obtained through this density calculation, which can reflect the similarity and potential topic structure between texts. The specific implementation is as follows: (1) Calculate the mutual reachability distance between two points as: , in, express and The Euclidean distance between express The distance to the farthest point among its k nearest neighbors, express The distance to the farthest point among its nearest k neighbors, where k is the parameter for the minimum number of samples (MinSamples) (i.e., the MinSamples parameter set by HDBSCAN).
[0052] (2) Through the density reachability distance, a weighted complete graph G-(V, E) is constructed, where the point set V is the document and the edge weight E is the density reachability distance. Then the minimum spanning tree (MST) is constructed on the graph to form a hierarchical eye tree (Dendrogram). The minimum spanning tree weight is as follows: , in, is the density reachability distance between two points.
[0053] (3) Define the stability of each cluster by analyzing the persistence of each cluster in the hierarchical clustering tree is the sum of the weights of all edges in the cluster.
[0054] , in, and Represents the distance threshold for cluster appearance and disappearance respectively.
[0055] (4) HDBSCAN will mark outliers that fail to meet density requirements (i.e., clusters with insufficient core points) as noise points. The processing of outliers is based on the local density relationship of the core distance: , in, A higher value of means that the point is more likely to be noise. It is a data point The distance to the farthest point among its nearest k neighbors (the number of neighbors set by MinSamples), It is the mean of the core distances of all points in the cluster to which the point belongs.
[0056] (5) HDBSCAN finally divides the text into different clusters , and a cluster containing noise points , the cluster label assignment is: , in, represents the kth cluster (Cluster), are the different texts in the dataset.
[0057] Step 8: Theme representation; The c-TF-IDF method is used to evaluate the importance of words in each topic. Specifically, c-TF-IDF measures the importance of words by calculating the term frequency (TF) of each word in the topic cluster and the inverse document frequency (IDF) of the word in the entire corpus. The formula is as follows: , , , in, is the class-based term frequency, is the Inverse Document Frequency, It is a combination of category-weighted term frequency (c-TF) and inverse document frequency (IDF) to measure the word In Theme The importance of Yes word In Theme The number of occurrences in Is the theme The total frequency of all words in , is the total number of documents, Contains words The number of documents.
[0058] In addition, in order to extract the most representative topic keywords, the present invention combines the maximum marginal relevance search (MMR) method to avoid selecting repeated words and ensure the diversity of topic words.
[0059] , Among them, S is the currently selected theme keyword set, The topic center vector representing the current topic, It's a word and Theme The similarity between is a hyperparameter that controls the balance between the relevance (similarity to the topic) and diversity (similarity to other keywords) of a keyword. is to find the word that matches the current candidate The most similar one, that is, the one with the maximum similarity, It is a keyword in the selected keyword set S.
[0060] Step 9: Perform probability redistribution matrix (PRM) redistribution; In the traditional BERTopic model, the HDBSCAN clustering algorithm will automatically identify some data points as outliers and exclude them from the topic. However, in applications such as social work questionnaires, policy document analysis, and social media public feedback information monitoring, these outliers may represent potentially important social issues. If simply discarded, it may lead to the loss of key information such as minority issues and special social phenomena. In order to solve this problem, the present invention proposes a probability redistribution matrix (PRM) method, which redistributes outliers through semantic similarity calculation to ensure maximum utilization of data and improve the coverage and robustness of topic modeling.
[0061] Furthermore, PRM redistributes outliers through the following steps: (1) Construct a document-topic probability matrix; First, PRM calculates the probability distribution of each document belonging to each topic. This process is based on the HDBSCAN clustering results, where the probability distribution of each document represents the probability that the document belongs to each topic. For each document, the sum of the probabilities of all topics is 1, that is, .
[0062] Among them, M is the total number of topics, indicating the distribution of each document in all topics. For documents marked as outliers, PRM will calculate the similarity between them and each topic and probabilistically reassign them to the closest topic.
[0063] (2) Calculate the probability of each document for each topic; The PRM method calculates the similarity between each document and each topic to obtain the probability that the document belongs to each topic. Specifically, PRM uses a distance metric function to measure the similarity between documents and topics. Assume that the similarity between documents and topics is calculated by To express it, PRM uses the following formula to calculate the probability that a document belongs to a topic: , in, Is the current document With a specific topic The distance The current document and all topics The distance is used to measure the matching degree between the document and each topic. The smaller the distance, the more relevant it is. It is finally used to convert into probability through Softmax.
[0064] (3) Update the topic assignment of outliers; Once the topic probability distribution of each document is calculated, PRM will decide whether a document should be reassigned to a topic based on the probability distribution of each document to each topic, especially those documents marked as outliers by HDBSCAN. If the probability of an outlier document on a topic is high and exceeds the set threshold, the document will be reassigned to the closest topic cluster. In this way, PRM can not only retain the information of outliers, but also improve the model's ability to identify potentially important topics.
[0065] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used for various other combinations, modifications and improvements, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art do not deviate from the spirit and scope of the present invention, and should be within the scope of protection of the claims attached to the present invention.
Claims
1. A multimodal data-driven social work field modeling optimization method, characterized by: The optimization method comprises: S1. Build a social work multimodal data set containing text, audio, and image / video data, and perform text embedding after data preprocessing to convert text into embedding vectors; S2, using the popularity bias regularization processor to process the data input in the social work field through the dual optimization strategy of high-frequency common word attenuation weighting and sociological term enhancement; S3, dynamically select the dimension reduction dimension through the dynamic document embedding optimizer to optimize the effect of topic modeling, and use the UMAP algorithm to reduce the dimension of the vector in the BERT embedding space; S4. Use the HDBSCAN algorithm to cluster documents into similar embedding groups in order to extract topics from them, form a hierarchical clustering structure by calculating the density relationship between data points, and divide the final clusters by density thresholds; S5. c-TF-IDF is used to measure the importance of vocabulary by calculating the word frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus, and the semantic similarity is calculated by the probability redistribution matrix method to redistribute outliers to ensure maximum utilization of data.
2. A multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The high-frequency common word attenuation weighting includes the following: Exponential decay is performed on the top M high-frequency words in the corpus, so that the influence of these high-frequency words gradually weakens during the topic modeling process; The Sociology Terminology Enhancement includes the following: Construct a dictionary of terms related to social work, social strategy and social governance, and amplify the weights of these terms by set multiples in the subsequent calculation process.
3. The multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The effects of optimizing topic modeling by dynamically selecting the dimension reduction dimension through the dynamic document embedding optimizer include: Dimension range setting: The dynamic document embedding optimizer sets the range of UMAP output dimensions to cover the semantic representation requirements of typical text embeddings while avoiding information loss caused by inappropriate dimensions, so that the model retains key information in the text data during dimensionality reduction, thereby improving the quality and accuracy of clustering; Silhouette coefficient evaluation: The dynamic document embedding optimizer evaluates each candidate dimension selected from the set dimension range and calculates the silhouette coefficient of the corresponding clustering result; Optimal dimension selection: By comparing the silhouette coefficients of different candidate dimensions, the dimension with the maximum value is selected.
4. The multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The S4 specifically includes the following contents: A1. Calculate two data points and Density accessibility distance ,in, express and The Euclidean distance between express The distance to the farthest point among its k nearest neighbors, express The distance to the farthest point among its k nearest neighbors; A2. Construct a weighted complete graph through density reachability distance, and construct a minimum spanning tree for the complete graph to form a hierarchical eye tree; A3. By analyzing the persistence of each cluster in the hierarchical clustering tree, the stability of each cluster is set to be the sum of the weights of all edges in the cluster; A4. Mark outliers that fail to meet density requirements as noise points; A5. Divide the text into different clusters and a cluster containing noise points.
5. The multimodal data-driven social work field modeling optimization method according to claim 1, characterized in that: The formula in S5 is To calculate the frequency of each word in the topic cluster and the inverse document frequency of the word in the entire corpus to measure the importance of the word, where , , is the category-weighted term frequency, is the inverse document frequency, It is a combination of category-weighted word frequency and inverse document frequency to measure word On topic The importance of Yes word On topic The number of occurrences in Is the theme The total frequency of all words in , is the total number of documents, Contains words The number of documents.
6. A multimodal data-driven social work field modeling optimization method according to claim 5, characterized in that: The S5 also combines the maximum marginal relevance search method to avoid selecting repeated words and ensure the diversity of subject words. The calculation formula of the maximum marginal relevance search includes: , where S is the currently selected topic keyword set, The topic center vector representing the current topic, It's a word and Theme The similarity between is a hyperparameter that controls the balance between keyword relevance and diversity. Indicates finding words that match the current candidate The value with the largest similarity is It is a keyword in the selected keyword set S.
7. The multimodal data-driven social work field modeling optimization method according to claim 1 is characterized by: The method of calculating semantic similarity by a probability redistribution matrix method to redistribute outliers includes: Constructing the document-topic probability matrix: The probability redistribution matrix calculates the probability distribution of each document belonging to each topic. The probability distribution of each document represents the probability that the document belongs to each topic. Calculate the probability of each document for each topic: Calculate the similarity between each document and each topic to get the probability that the document belongs to each topic; Update the topic assignment of outliers: Determine whether the document should be reallocated to a topic based on the probability distribution of each document to each topic. If the probability of a document being on a topic is high and exceeds the set threshold, the document is reallocated to the closest topic cluster.
Citation Information
Patent Citations
Automatic text summarization method based on fusion semantic clustering
CN108197111A
Data cleaning method and system, computer equipment and storage medium
CN112364935A
Unsupervised keyword extraction method based on gated topic model
CN117390157A
Small sample point cloud semantic segmentation method based on difference enhancement and related equipment
CN118015262A
Theme modeling and sentiment analysis method and system based on deep learning
CN118394944A
Cited By
Document blood relationship analysis method based on BERT model
CN120523943A