Big data-based humanity and community data acquisition system and method thereof

By designing a humanities and social science data acquisition system based on big data, the problems of dispersed data sources, low processing efficiency and limited analysis capabilities in the existing technology are solved, efficient data acquisition and accurate analysis are achieved, and a powerful user interaction interface is provided.

CN120067195AInactive Publication Date: 2025-05-30SANYA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510147348.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing methods of obtaining humanities and social science data have problems such as dispersed data sources, low data processing efficiency and limited analysis capabilities, which are difficult to meet the rapidly developing humanities and social science research needs.

Method used

A humanities and social science data acquisition system based on big data is designed, including humanities and social science data extraction module, data post-processing module, data storage module, analysis module and user port module. Cross-platform resources are obtained through API interface, network crawler and expert information entry units, data cleaning, deduplication and classification are carried out, and semantic understanding and knowledge reconstruction are realized.

Benefits of technology

Real-time data capture and incremental synchronization is realized, the data surface is expanded, the data accuracy and processing efficiency is ensured, and powerful analysis capabilities and visual interactive interface are provided to facilitate user operations and data exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067195A_ABST
    Figure CN120067195A_ABST
Patent Text Reader

Abstract

The invention discloses a human and social data acquisition system based on big data. The system comprises a human and social data extraction module, a data post-processing module, a data storage module, an analysis module and a user port module. The humanistic and social data extraction module is used for acquiring real-time capture and incremental synchronization of cross-platform resources, and the data post-processing module is used for performing data cleaning, duplicate removal and classification on data contents of the humanistic and social data extraction module; the data storage module is used for storing data information processed by the data post-processing module, the analysis module acquires the stored data information to perform semantic understanding and reconstruction and complete data information extraction, and the user port module comprises user registration, login, verification and data retrieval and acquisition. Compared with the prior art, the human and social data acquisition system and method based on the big data have the advantages that application is convenient, data acquisition is convenient, and data are comprehensive and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and in particular to a humanities and social sciences information acquisition system based on big data and a method thereof. Background Art

[0002] With the rapid development of humanities and social sciences research, the number of academic resources has exploded, and traditional data acquisition and processing methods can no longer meet research needs.

[0003] Existing methods of obtaining humanities and social science data mostly rely on traditional means such as literature retrieval and questionnaire surveys, which have defects such as scattered data sources, low data processing efficiency, and limited analytical capabilities. With the continuous development of big data technology, how to efficiently and accurately obtain the data needed for humanities and social science research from massive data has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a humanities and social science information acquisition system and method based on big data, which is easy to use, convenient to acquire data, and has comprehensive and accurate data.

[0005] In order to solve the above technical problems, the technical solution provided by the present invention is: a humanities and social sciences data acquisition system based on big data, including a humanities and social sciences data extraction module, a data post-processing module, a data storage module, an analysis module and a user port module;

[0006] The humanities and social sciences data extraction module includes real-time capture and incremental synchronization of cross-platform resources, and the data post-processing module includes data cleaning, deduplication and classification of the data content of the humanities and social sciences data extraction module;

[0007] The data storage module is used to store the data information processed by the data post-processing module. The analysis module obtains the stored data information for semantic understanding and reconstruction to complete data information extraction. The user port module includes user registration, login, verification and data retrieval and acquisition.

[0008] Preferably, the humanities and social sciences data extraction module includes an API interface unit, a web crawler unit, and an expert information entry unit.

[0009] Preferably, the API interface unit includes access to libraries, academic journal databases, government public information websites, and social media platforms to collect literature, policies, news, and public hot spots.

[0010] Preferably, the web crawler unit includes configuring a dynamic IP proxy, sending an HTTP request after acquiring a crawl target, parsing web page content, extracting information and storing data.

[0011] Preferably, the expert information input unit includes obtaining information of experts and scholars, and sorting, classifying, and inputting the materials of experts and scholars.

[0012] Preferably, the user port module includes a visual interaction unit, and the visual interaction unit performs data retrieval input and visual display of retrieval results.

[0013] Preferably, the analysis module includes a semantic understanding unit, a knowledge reconstruction unit, and a knowledge reconstruction unit.

[0014] Preferably, the post-processing module for materials includes a data cleaning unit, a data deduplication unit, and a data classification unit.

[0015] On the other hand, the present invention discloses a method for using a system for obtaining humanities and social science materials based on big data, including the following steps:

[0016] S1: Raw data collection;

[0017] S2: Raw data processing;

[0018] S3: Data analysis of the processed data;

[0019] S4: Outputting the analyzed data information to the user port module;

[0020] S5: Using the user port module for retrieval and application of data information.

[0021] Preferably, the raw data collection in S1 includes using a humanities and social science material extraction module to obtain cross-platform resources;

[0022] The raw data processing in S2 includes using the post-processing module for materials to perform data cleaning, deduplication, and classification;

[0023] The data analysis of the processed data in S3 includes using the analysis module to extract key data information.

[0024] The advantages of the present invention compared with the prior art are as follows: In the present invention, the basic information is obtained through the API interface unit, the network crawler unit, and the expert information input unit, thereby greatly expanding the data surface. Regular data acquisition can ensure real-time capture and incremental synchronization of cross-platform resources, ensuring data iteration. In the present invention, an analysis module including semantic understanding and reconstruction is also constructed to complete data information extraction, ensuring the accuracy of the processed data, and the visual interface can realize multi-data exploration applications, which is convenient for operation and use. Brief Description of the Drawings

[0025] Figure 1 is a structural schematic diagram of a system for obtaining humanities and social science materials based on big data. Detailed Description of the Invention

[0026] The present invention will be further described in detail below with reference to the accompanying drawings.

[0027] A system for obtaining humanities and social science materials based on big data, characterized in that it includes a humanities and social science material extraction module, a post-processing module for materials, a data storage module, an analysis module, and a user port module; the humanities and social science material extraction module includes real-time capture and incremental synchronization for obtaining cross-platform resources, and the post-processing module for materials includes data cleaning, deduplication, and classification of the data content of the humanities and social science material extraction module; the data storage module is used to store the data information processed by the post-processing module for materials, the analysis module obtains the stored data information for semantic understanding and reconstruction to complete data information extraction, and the user port module includes user registration, login, verification, and data retrieval and acquisition.

[0028] In one embodiment:

[0029] The humanities and social science material extraction module includes an API interface unit, a web crawler unit, and an expert information entry unit. Among them: the API interface unit is an important bridge for the humanities and social science material extraction module to connect with external data sources, including accessing libraries, academic journal databases, government public information websites, and social media platforms to collect literature, policies, news, and public hot topic materials, automatically collecting rich literature materials, policy documents, news reports, and public hot topic materials, including but not limited to history, economics, sociology, cultural studies, etc.;

[0030] The web crawler unit includes configuring a dynamic IP proxy, sending an HTTP request after obtaining the crawling target, parsing the web page content to extract information and store data, ensuring the efficiency and stability of data collection;

[0031] The expert information entry unit includes obtaining information of experts and scholars, sorting, classifying, and entering the materials of experts and scholars, and collecting and entering first-hand materials from experts and scholars through manual means, such as interview records, conference papers, and unpublished works, so as to continuously enrich and improve its humanities and social science database.

[0032] In one embodiment:

[0033] The analysis module includes a semantic understanding unit, a knowledge reconstruction unit, and a knowledge reconstruction unit, and the post-processing module for materials includes a data cleaning unit, a data deduplication unit, and a data classification unit;

[0034] The data cleaning unit includes removing noise data such as HTML tags and advertisement information, detecting and correcting spelling and grammar errors in the text. Meanwhile, a text error correction model is used to perform semantic correction on the OCR recognition results. Through approximate text detection, semantically similar documents are identified and merged, and cross-language duplicate documents are also identified. During this process, a version control mechanism based on timestamps is used to retain the latest version of the data.

[0035] During classification, automatic classification of documents is achieved based on a deep learning model. Meanwhile, an active learning strategy is adopted to optimize the performance of the classification model through manual annotation.

[0036] When conducting analysis and utilization, a pre-trained language model is used to achieve deep semantic representation of the text. Dependency parsing technology can be adopted to extract entity relation triples in the text, identify and extract key academic elements such as research methods and conclusions. Based on the ontology-based academic knowledge representation framework, discrete knowledge units are mapped to a continuous vector space to refine core academic viewpoints and form knowledge cards for application. When extracting information, rule templates can be specified for key information extraction, including authors, institutions, time, etc.

[0037] A method for using a humanistic and social science data acquisition system based on big data includes the following steps:

[0038] S1: Raw data collection;

[0039] S2: Raw data processing;

[0040] S3: Data analysis of the processed data;

[0041] S4: Outputting the analyzed data information to the user port module;

[0042] S5: Using the user port module for retrieval and application of data information.

[0043] When the present invention is specifically implemented, the raw data collection in S1 includes using a humanistic and social science data extraction module to obtain cross-platform resources;

[0044] The raw data processing in S2 includes using a post-data processing module for data cleaning, deduplication, and classification;

[0045] The data analysis of the processed data in S3 includes using an analysis module to extract key data information. The system obtains humanistic and social science literature data, relevant discussion content, etc. through API interfaces, etc., performs deduplication and noise filtering on the collected data, retains high-quality literature, uses models to extract research methods and conclusions in the literature, and simultaneously generates a domain knowledge graph based on the extracted entities and relationships, thereby displaying the knowledge graph and trend prediction results through an interactive interface to support users' exploration and analysis.

[0046] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0047] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the description in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

[0048] The scope of the present invention claimed is defined by the appended claims and their equivalents. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention.

[0049] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0050] In the present invention, unless otherwise clearly specified and defined, the first feature being “above” or “below” the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being “above”, “over” and “on” the second feature includes that the first feature is directly above and obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being “below”, “under” and “beneath” the second feature includes that the first feature is directly below and obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature

[0051] The above has described the present invention and its embodiments. Such description is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, design similar structural modes and embodiments without creative efforts, they should all belong to the scope of protection of the present invention.

Claims

1. A humanities and social sciences data acquisition system based on big data, characterized by: It includes humanities and social science data extraction module, data post-processing module, data storage module, analysis module and user port module; The humanities and social sciences data extraction module includes real-time capture and incremental synchronization of cross-platform resources, and the data post-processing module includes data cleaning, deduplication and classification of the data content of the humanities and social sciences data extraction module; The data storage module is used to store the data information processed by the data post-processing module. The analysis module obtains the stored data information for semantic understanding and reconstruction to complete data information extraction. The user port module includes user registration, login, verification and data retrieval and acquisition.

2. According to the humanities and social sciences data acquisition system based on big data as described in claim 1, it is characterized by: The humanities and social sciences data extraction module includes an API interface unit, a web crawler unit, and an expert information input unit.

3. According to the humanities and social sciences data acquisition system based on big data as described in claim 2, it is characterized by: The API interface unit includes access to libraries, academic journal databases, government public information websites, and social media platforms to collect documents, policies, news, and public hot spots.

4. According to the humanities and social sciences data acquisition system based on big data as described in claim 2, it is characterized by: The web crawler unit includes configuring a dynamic IP proxy, sending an HTTP request after obtaining a crawl target, parsing web page content, extracting information and storing data.

5. According to the big data-based humanities and social sciences data acquisition system of claim 2, it is characterized by: The expert information input unit includes obtaining information of experts and scholars, and arranging, classifying and inputting the information of experts and scholars.

6. A humanities and social sciences data acquisition system based on big data according to claim 1 or 2, characterized in that: The user port module includes a visualization interaction unit, which performs data retrieval input and visualization display of retrieval results.

7. The humanities and social sciences data acquisition system based on big data according to claim 6, characterized in that: The analysis module includes a semantic understanding unit, a knowledge reconstruction unit and a knowledge reconstruction unit.

8. The humanities and social sciences data acquisition system based on big data according to claim 6, characterized in that: The data post-processing module includes a data cleaning unit, a data deduplication unit and a data classification unit.

9. A method for using a humanities and social sciences data acquisition system based on big data, applied to the humanities and social sciences data acquisition system according to any one of claims 1 to 8, characterized in that: The steps include: S1: Raw data collection; S2: Raw data processing; S3: Data analysis of processed data; S4: Output the analyzed data information to the user port module; S5: Use the user port module to perform data information retrieval application.

10. The method for using a humanities and social sciences data acquisition system based on big data according to claim 9, characterized in that: The raw data collection in S1 includes obtaining cross-platform resources using a humanities and social sciences data extraction module; Raw data processing in S2 includes data cleaning, deduplication, and classification using the data post-processing module; Data analysis of processed data in S3 involves extracting key data information using analysis modules.