Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

225 results about "Web crawler" patented technology

A Web crawler, sometimes called a spider or spiderbot and often shortened to crawler, is an Internet bot that systematically browses the World Wide Web, typically for the purpose of Web indexing (web spidering).

Electronic product personalized marketing method based on big data accurate portrait

The invention discloses an electronic product personalized marketing method based on a big data accurate portrait, and relates to the technical field of big data processing, and the method comprises the following steps: multi-source data collection and cross-domain fusion processing: obtaining a multi-source data set related to an electronic product from each platform of the Internet by using a web crawler and a streaming processing engine, the multi-source data set is subjected to cleaning classification and alignment fusion processing to generate multi-modal data, and then the multi-modal data is updated in real time to dynamically reflect data changes, so that a foundation is laid for subsequent data analysis; according to the method, multi-source data is collected by using a web crawler and a streaming processing engine, quantitative analysis is performed on the data by using a Hofscode culture dimension model and a Maslow demand level theory, a user portrait is constructed by using a distributed data flow engine Apache Flink to dynamically update a label, and a GAN (GAN Adversarial Generative Network) anomaly detection and evaluation analysis method is applied, so that the user portrait is obtained. And comprehensive and accurate data acquisition and processing are realized.
Owner:SUZHOU HEHEYI TECH CO LTD

Web automatic penetration testing method and device, electronic equipment and storage medium

The invention provides a Web automatic penetration test method and device, electronic equipment and a storage medium, and belongs to the technical field of website vulnerability tests.The method comprises the steps that a web crawler collects target Web information to form a basic set; detecting vulnerability acquisition information; matching the adaptive script, and generating an intelligent load to verify vulnerability availability; constructing a knowledge graph to determine vulnerability association, and evaluating success rate and risk; planning an optimal attack route by using reinforcement learning, and determining attack depth and range, namely attack path risk features; and evaluating the risk by the dynamic scoring model, and generating a visual report containing detailed information, results and repair and improvement suggestions. According to the automatic penetration testing method, manual intervention is reduced in the whole process, dependence on professional experience is reduced, efficiency is improved, and the problems that traditional testing is tedious and low in efficiency are solved.
Owner:SICHUAN BANWOHUI TECHNOLOGY CO LTD +1

Large-model-driven agricultural knowledge graph analysis method and system

The invention relates to the technical field of agriculture, in particular to a large-model-driven agricultural knowledge graph analysis method and system, and aims to realize standardized processing of multi-source heterogeneous data through a three-stage preprocessing process and combine with a field adaptive large-model training technology so as to realize the large-model-driven agricultural knowledge graph analysis method and the large-model-driven agricultural knowledge graph analysis method and the large-model-driven agricultural knowledge graph analysis system. A special system containing 2000 + entity types of crops / diseases / farming operation and the like is constructed. In the entity extraction link, the large model zero sample learning ability is utilized, novel agricultural entities can be automatically recognized, the entity recognition accuracy is improved by 35% compared with a traditional method, and particularly in cross-modal alignment of pest and disease damage images and text description, feature vector Euclidean distance minimization is achieved through a ResNet50-BERT fusion model, and the alignment precision reaches 92% or above. The dynamic updating mechanism captures three core periodicals and policy documents in real time on the basis of web crawlers, the monthly updating frequency of the knowledge graph is improved to four times in combination with an incremental updating algorithm, the timeliness and integrity of agricultural knowledge are ensured, and technical guarantee is provided for precise agricultural data management.
Owner:ZHENGZHOU DIGITAL INTELLIGENCE TECH RES INST CO LTD

Project cloud platform data management method and system based on process market

InactiveCN120430757AWeb data indexingSemantic analysisBusiness enterpriseEnterprise process
The invention relates to the technical field of information management, and discloses a project cloud platform data management method based on a process market, and the method comprises the steps: obtaining process demand information inputted by an enterprise user, and obtaining an enterprise name of the enterprise user; crawling enterprise process information through a web crawler tool, and performing multi-level classification to construct a process market platform; obtaining the business scope and the enterprise scale of the enterprise user according to the enterprise name, and determining the enterprise type of the enterprise user according to the business scope; performing keyword extraction on the process demand information input by the enterprise user to determine a target keyword of the process demand information; and matching a demand flow list in a flow market platform according to the enterprise type, the enterprise scale and the target keyword, and pushing the demand flow list to the enterprise user. According to the method, an intelligent process market platform is constructed, keyword extraction and analysis are carried out on demands of enterprise users, automatic aggregation and accurate matching of process resources are realized, and the process adaptation accuracy is remarkably improved.
Owner:SUZHOU HUIQIDA TECH TRANSFER CO LTD

Multi-source data federated governance method and system for vocational education

The invention relates to a data processing method and system, in particular to a vocational education multi-source data federation treatment method and system, and belongs to the technical field of vocational education informatization and big data. Then constructing a knowledge graph in a triple form and mapping the knowledge graph to a multi-dimensional tensor space; secondly, multi-source data compatibility is analyzed and evaluated through topological features, and feature fusion maintaining topological invariants is executed; finally, knowledge reasoning and service are provided based on a quantum probability field model, the complex relation expression ability is improved through tensor representation, the heterogeneous data fusion quality is guaranteed through topology durability, the uncertainty in knowledge service is processed by introducing the quantum probability theory, and the reliability of the knowledge reasoning and service is improved. The problems of multi-source heterogeneous, complex relation, insufficient service accuracy and the like of vocational education data are effectively solved, and a comprehensive data federation governance solution is provided for the vocational education field.
Owner:CHONGQING HANHAI RUIZHI BIG DATA TECH CO LTD +1

Ransomware automatic acquisition and analysis method and system based on large language model

The invention discloses an automatic ransomware collection and analysis method and system based on a large language model in the technical field of network security, and the method comprises the steps: constructing a target data source, collecting a first sample of a ransomware, and capturing the first sample which is actually triggered in an enterprise-level network environment and context behavior data of the first sample from the first sample; an automatic crawler system is established by adopting an open source web crawler framework to establish real-time connection with a malicious code platform, the popularity of a first sample on the malicious code platform is analyzed, and an optimization task is set. According to the method, the natural language processing capacity and the automation technology of a large language model are combined, efficient capture, accurate analysis and quick response of the samples are achieved, an anti-virtualization detection mechanism of the blackmail virus is effectively coped with by implementing an anti-virtualization detection confrontation strategy, and an intelligent protection chain is constructed for network security.
Owner:JIANGSU TAIHU HUIYUN DATA SYST CO LTD

Construction method of time sequence knowledge graph in water conservancy field based on T-ATLOP model

The invention provides a method and a system for constructing a time sequence knowledge graph in the water conservancy field based on a T-ATLOP model, which are applied to the water conservancy field, mainly comprise two key levels, namely a data layer and a knowledge extraction layer, and realize semi-automatic time sequence knowledge graph construction. The method comprises the following steps: carrying out batch document processing on water conservancy census data by utilizing a large language model through a prompt project, crawling water conservancy consultation of water conservancy halls of all provinces and cities by a web crawler, constructing a time sequence knowledge graph data set HE-DOCU through two-step method labeling of large language model assistance and heuristic strategy labeling, and carrying out a knowledge extraction layer part on the basis of an ATLOP model. A new model T-ATLOP is provided for knowledge extraction. And finally, storing the knowledge tetrad of the knowledge extraction layer by using a graph database to form a water conservancy field time sequence knowledge graph. According to the method, the more perfect water conservancy field time sequence knowledge graph is constructed by fully utilizing the water conservancy information, so that the accuracy and the intelligent level of querying related knowledge by hydrological employees can be improved.
Owner:HOHAI UNIV

Production safety accident key cause link identification method based on complex network

The invention relates to the technical field of safety production management and intelligent analysis, in particular to a production safety accident key cause link identification method based on a complex network, which comprises the following steps of: acquiring structured production safety accident report data from a safety production supervision platform by utilizing a web crawler technology; a TF-IDF algorithm is combined with an industrial standard term library to optimize Jieba word segmentation, and a chi-square statistical method is adopted to calculate the correlation between each risk keyword and an accident category. By fusing text mining, statistical analysis and complex network theories and combining field dictionary optimization word segmentation, multi-accident model adaptation and triple centrality index analysis, the problems that a traditional method is single in analysis dimension, depends on subjective experience and is insufficient in technology integration are solved; efficient utilization of unstructured data, scientific layering of risk factors and accurate identification of key cause links are realized, and objectivity and effectiveness of safety management of a complex production system are effectively improved.
Owner:CIVIL AVIATION UNIV OF CHINA

Multi-source data federation governance method and system for vocational education

The present application relates to a data processing method and system, in particular to a vocational education multi-source data federal governance method and system, belonging to the field of vocational education informatization and big data technology, which first acquires and preprocesses vocational education multi-source data through network crawler and OCR technology; then constructs a knowledge graph in the form of triplets and maps it to a multi-dimensional tensor space; then evaluates the compatibility of multi-source data through topological feature analysis, and performs feature fusion that maintains topological invariants; finally, based on a quantum probability field model, knowledge reasoning and services are provided, the present application uses tensor representation to improve the expression ability of complex relationships, ensures the quality of heterogeneous data fusion through topological persistence, and introduces quantum probability theory to handle uncertainty in knowledge services, effectively solving the problems of multi-source heterogeneous vocational education data, complex relationships and insufficient service accuracy, and providing a comprehensive data federal governance solution for the vocational education field.
Owner:CHONGQING HANHAI RUIZHI BIG DATA TECH CO LTD +1

Fast evaluation method of site seismic liquefaction hazard based on artificial intelligence algorithm

Disclosed in the present disclosure is a fast evaluation method of site seismic liquefaction hazard based on artificial intelligence algorithm: establishing a historical seismic and site information database, the database including a demand input module, a web crawler module, a data processing module, and a database module connected in sequence; a neural network model performs prediction to acquire a post-earthquake site dominant frequency; and, on the basis of the post-earthquake site dominant frequency, acquiring a site earthquake damage degree and seismic performance parameters. The present disclosure solves the problem of fast evaluating post-earthquake site earthquake damage and site seismic performance parameters, and can rapidly evaluate the site liquefaction or softening earthquake damage degree and site seismic performance parameters in given earthquake conditions.
Owner:ZHEJIANG UNIV

Multi-dimensional hotspot extraction method based on data analysis

The invention discloses a multi-dimensional hotspot extraction method based on data analysis, relates to the technical field of data processing, and solves the problem that features and hotspot feature sets of text information data, image video data and voice information data are difficult to extract respectively after data is automatically captured in real time and preprocessed by using a web crawler technology. The hot events are difficult to identify by using a clustering algorithm, and analysis and evaluation of time dimension, space dimension and emotion dimension are performed on the hot events; and the development trend of the hot event is difficult to predict. According to the method, text information, image video and voice information data are automatically captured through a web crawler, multi-modal fusion is carried out after preprocessing and hot spot feature extraction, a hot spot event is identified by utilizing a clustering algorithm, and multi-dimensional analysis, evaluation and prediction are carried out on the hot spot event.
Owner:School of Political Science, National Defense University of the Chinese People's Liberation Army

Knowledge graph generation method and system fusing multi-source science and technology data visualization results

The invention relates to a knowledge graph generation method and system fusing multi-source science and technology data visualization results. The invention provides a systematic solution for solving the problems that in the prior art, multi-source heterogeneous data fusion is difficult, preprocessing intellectualization is insufficient, implicit relation mining is insufficient, and the visualization effect is poor. The method comprises the steps that a data import module obtains data through local and web crawlers, and a keyword is expanded through a GPT model; the preprocessing module is used for cleaning, synonym merging and field standardization; the information processing module generates one-dimensional graph data and a two-dimensional triple, wherein an implicit relationship is predicted by a BERT model and a graph neural network; the graph structure planning module adopts force-oriented layout and dimensionality reduction optimization visualization; and the co-imported network graph is divided into communities through a difference filter and a Louvain algorithm. According to the method, multi-source data is effectively integrated, noise is automatically processed, semantic association is mined, the high-readability chart is generated, and the construction efficiency and the visualization effect of the knowledge graph are remarkably improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST +1

Dynamic cybersecurity scoring using traffic fingerprinting and risk score improvement

A system for dynamic cybersecurity scoring using traffic fingerprinting and score improvement, that uses a web crawler that sends message prompts to external hosts and receives responses from external hosts, a time-series data store that produces time-series data from the message responses, and a directed computational graph module that analyzes the time-series data to produce a weighted score representing the overall cybersecurity state of an organization.
Owner:QOMPLX INC

Self-adaptive multi-dimensional evaluation high-quality scientific and technological information screening method and system

The invention discloses a self-adaptive multi-dimensional evaluation high-quality scientific and technological information screening method and system, and relates to the technical field of information screening, and the method comprises the following steps: collecting data from a database by using an API interface and a web crawler technology, and establishing a multi-source data set; establishing a preprocessing data set; performing feature extraction on the preprocessed data set, and establishing a multi-dimensional feature set after standardization processing; carrying out dimension evaluation under a six-dimensional evaluation channel on the multi-dimensional feature set, and establishing a dimension score; and after the dynamic weight mapped by the dimension score is adaptively configured, weighted calculation is executed, and a scientific and technological information screening result is generated. The technical problems that in the prior art, due to the fact that the scientific and technological information screening dimension is single, the evaluation weight is fixed and different scenes are difficult to adapt, the screening result is insufficient in accuracy and comprehensiveness are solved, and the purposes of achieving self-adaptive multi-dimensional evaluation and high-quality screening of the scientific and technological information and improving the screening efficiency are achieved. And the accuracy and comprehensiveness of scientific and technological information screening are improved.
Owner:DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Asset evaluation intelligent system based on AI technology

The invention relates to the technical field of asset assessment. The asset evaluation intelligent system based on the AI technology is provided, and the system comprises the steps that web crawler capture processing is conducted on a data source of a public network channel, and a first type of heterogeneous data is obtained; aPI interface docking processing is carried out on the data source of the enterprise internal system, and second-class heterogeneous data is obtained; performing logic splicing processing on the first type of heterogeneous data and the second type of heterogeneous data to generate multi-source heterogeneous asset evaluation data; performing model matching processing on the asset value characteristic matrix based on the asset type, and selecting an estimation model architecture; performing interpretability analysis processing on the dynamic evaluation parameters to generate a key factor contribution degree and an interpretable evaluation conclusion; and performing document processing on the contribution degree of the key factor and the interpretable evaluation conclusion to generate an interactive evaluation document so as to improve the multi-source data integration capability, enhance the risk dynamic evaluation adaptability and strengthen the real-time relevance and interpretability of the evaluation document.
Owner:CHINA SECURITIES REAL ESTATE APPRAISAL & COST GROUP CO LTD

Processing for spam detection of untrusted domains

Embodiments of the technology described programmatically decrease the number of spam Uniform Resource locators (URLs) that are accessed from untrusted domains when the subdomain prefix is above a threshold probability of having been randomly generated. In this regard, prior to adding a discovered set of URLs to a crawl queue of a web crawler, the URLs are filtered into URLs from trusted domains and untrusted domains determined by a statistical model. The trusted domain URLs are added to the crawl queue, and the remaining URLs are sandboxed to filter out spam URLs. The subdomain prefixes of the sandboxed URLs are applied to a neural network to determine the probability that the subdomain prefixes are randomly generated. When a subdomain prefix is above a threshold probability of having been randomly generated, the subdomain is determined to be a spam subdomain and can be blocked.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Talent data and post matching method based on weight adjustment and related equipment

The invention discloses a talent data and post matching method based on weight adjustment and related equipment, which can effectively improve the matching accuracy, help recruiters to quickly position adaptive candidates, reduce the manual screening cost and improve the recruitment efficiency, thereby meeting the requirements of quick recruitment of enterprises and efficient job hunting of talents. The method comprises the following steps: capturing multi-source data by utilizing a web crawler, and carrying out natural language processing on the multi-source data to obtain structured data; post time sequence characteristics of the structured post data in the structured data are extracted; when it is determined that the post time sequence features meet post requirements, performing multi-dimensional analysis on the post time sequence features to correct deviations and fuse weights, and then generating target post time sequence features; extracting talent features of the structured talent data in the structured data; and converting the target post time sequence features and the talent features into target vectors, and inputting the target vectors into a deep learning model to obtain a prediction matching degree of talent data and posts.
Owner:CETC BIGDATA RES INST CO LTD +1

Large language model-based vulnerability remediation action descriptions

A vulnerability documentation system detects vulnerabilities having outdated or undocumented formatted descriptions for corresponding remediation actions. A web crawler crawls the Internet for configuration data for software / firmware affected by the detected vulnerabilities and descriptive content for the remediation actions. The vulnerability documentation system prompts and LLM with a prompt for each detected vulnerability comprising instructions to generate a formatted description for remediation actions using the crawled configuration data / descriptive content. The vulnerability documentation system then populates natural language descriptions of remediation actions from the formatted descriptions and pushes the natural language descriptions to affected devices.
Owner:PALO ALTO NETWORKS INC

Vietnamese grammar error correction corpus construction method based on error type perception of large model

The invention relates to a Vietnamese grammar error correction corpus construction method based on error type perception of a large model, and belongs to the field of natural language processing. According to the method, firstly, a voice recognition model is used for simulating Vietnamese grammar errors in a real scene, a preliminary error correction data set is generated, then through deep analysis of distribution rules and grammar structure features of typical errors in the data set, a chain thinking cue (CoT) mechanism fusing error type features is designed in a targeted mode, and the error type features of the Vietnamese grammar errors are extracted. Guiding a large language model (LLM) to generate synthetic statements containing predetermined grammar errors in batches; thirdly, in order to enhance corpus quality, synchronously implementing a web crawler to collect a native Vietnamese text, and constructing a pure monolingual corpus through multi-layer filtering and cleaning; and finally, the generated synthetic data needs to be strictly verified and processed to ensure that the error type is consistent with a preset target, and a pre-training model normal form and a large model normal form are strengthened in a two-stage fine tuning manner, so that the generalization ability of a grammar error correction model is effectively improved, and the problem that Vietnamese grammar error correction corpus is deficient is solved.
Owner:KUNMING UNIV OF SCI & TECH

College student psychological health service system and method based on AI Agent

The invention relates to the field of intelligent psychological health management service, in particular to an AI Agent-based college student psychological health service system and method, and the system comprises a psychological health measurement module, a social personality portrait module and a psychological health intervention module. According to the method, a measurer agent measures the current psychological health condition of a college student from a transverse angle through a standardized psychological health scale; the portrayman agent analyzes the historical psychological health conditions and MBTI personality portraits of the college student users from the longitudinal angle by using the social media data of the college student users obtained by the web crawler; based on psychological health scale data and social media data, the system adaptively recommends a consultant agent for college students, the consultant agent is externally connected with a psychological consultation expert guide and real consultation case data, professional and effective psychological health intervention services are provided for the college students, and the system is suitable for college student emotion accompanying, college auxiliary decision making and other scenes.
Owner:RENMIN UNIVERSITY OF CHINA

Database construction and data standardization method based on MOF proton conductor material

The invention discloses a database construction and data standardization method based on an MOF proton conductor material, and belongs to the crossing field of material science and database technology. Aiming at the problems of data dispersion and non-uniform formats of the current MOF proton conductor material, data is collected through multiple channels, and a web crawler is innovatively adopted to improve the efficiency. In the aspect of database construction, a distributed database is selected and combined with a block chain technology. In the aspect of data standardization, multi-dimensional data formats are standardized, complex representation data are processed through artificial intelligence, quality control is enhanced, data reasonability is predicted through machine learning, and real-time updating and correction are achieved through big data. According to the method, dispersed data are integrated, consistency and comparability are ensured, data quality and intelligent level are improved, and powerful support is provided for related research and application.
Owner:XI'AN POLYTECHNIC UNIVERSITY

Systems and methods for dynamically updating data for course generation

A system for dynamically updating data for course generation, the system including a web crawler operating on a server, wherein the web crawler is configured to identify one or more predetermined HTML elements on a plurality of web pages, identify isolated data as a function of the one or more predetermined HTML elements, compare the isolated data to a modification baseline and generate regulatory data as a function of the isolated data and the comparison, and a memory the memory containing instructions configuring at least a processor to receive the regulatory data, classify the regulatory data to one or more guideline categorizations, identify a plurality of course modules, modify the one or more course modules as a function of the regulatory data and transmit a notification associated with the one or more modified course modules to an end user.
Owner:BH OPERATIONS LLC

Method and system for detecting unauthorized URL (Uniform Resource Locator) of website based on automatic operation of common user

The invention provides a method and system for detecting an unauthorized URL of a website based on automatic operation of a common user, belongs to the technical field of network security, and can at least partially solve the problems of low detection efficiency, poor accuracy and excessive dependence on manual operation in the prior art. Automatically logging in a website by using the registration information, processing a verification link, and obtaining session information; obtaining a JS file through a web crawler in a normal user login state; analyzing the JS file, extracting static, dynamic and relative URLs in combination with lexical analysis, grammatical analysis and sandbox simulation execution, and unifying the static, dynamic and relative URLs as absolute URLs; de-duplicating and normalizing the URL; accessing the URL through a common user session by using a network test tool, obtaining response data, comparing the response data with a high-authority user reference response, and judging an unauthorized URL; the method has the beneficial effects that automatic and comprehensive detection of the unauthorized URL is realized, and the detection efficiency and accuracy are remarkably improved.
Owner:HUANENG POWER INT INC +1

Landslide distribution probability prediction method and system based on remote sensing data

The invention belongs to the technical field of landslide identification and prediction, and particularly relates to a landslide distribution probability prediction method and system based on remote sensing data, and the prediction method comprises the following steps: step 1 to step 6, selecting a research area, capturing remote sensing data, describing a curved surface form of the research area, and constructing a digital elevation model; after screening, introducing elevation data into MATLAB software, and drawing a three-dimensional topographic map in a three-dimensional coordinate system; constructing a landslide probability model, simulating variables in the landslide, deducing seismic data, and deducing slope data; constructing a WGEN model, and deducing a precipitation density function of the research area based on GAMMA distribution to realize precipitation simulation; the prediction system comprises a web crawler module, a local database, a processor and a touch screen. According to the method, the elevation and rainfall of the plateau area are associated, multiple models are constructed, the landslide distribution probability accuracy is high, and risk early warning can be carried out within 3 hours after an earthquake.
Owner:WUHAN GUOAN INTELLIGENT EQUIP CO LTD

Systems and methods for interactive scheduling

Disclosed herein are embodiments of systems, methods, and products comprises an analytic server, which automatically manages appointment scheduling. The analytic server receives a customer request to schedule an appointment. The analytic server determines the required data from both customer and service provider for making the appointment. The analytic server retrieves customer data comprising requested service attributes, user preferences, users attributes from internal database and external data source. The analytic server retrieves service providers' data comprising provider service attributes, providers' attributes from internal database and external data sources. The analytic server accesses external data source by web crawling various websites. The analytic server executes an artificial intelligence model to predict user preferences and needs. The analytic server determines potential service providers best matching the customer's input or predicted preferences. The analytic server generates an appointment for each matching service provider and transmits an electronic message comprising the appointments to customer device.
Owner:UNITED SERVICES AUTOMOBILE ASSOCIATION (USAA)

LLM-based engineering project scheme compliance examination method

The invention provides an LLM-based engineering project scheme compliance examination method. The method comprises the following steps: Step 1, data preparation; implementing multi-dimensional data acquisition through a distributed web crawler system; step 2, processing various types of data of images, texts and tables by using a large language model (LLM) technology; step 3, performing Prompt fine tuning optimization on the data, so that the large language model can better understand and generate related contents of compliance review of the engineering project scheme; step 4, performing automatic compliance check based on rule and mode matching; by defining a series of rules and modes, automatic compliance review of project documents is realized; and Step 5, establishing a risk generation assessment and early warning mechanism based on the large model. Through automatic compliance inspection, the workload of manual auditing is reduced, and document specifications are ensured; and an intelligent risk assessment and early warning mechanism identifies and notifies potential risks in real time, so that compliance risks are effectively reduced, and the safety and the overall level of project management are improved.
Owner:THREE GORGES GROUP IND DEVELOPMENT (BEIJING) CO LTD +1

Mobile application privacy policy compliance detection method based on pre-training model

The invention discloses a mobile application privacy policy compliance detection method based on a pre-training model, and aims to realize efficient and intelligent privacy policy compliance automatic detection. The method comprises the following steps: firstly, constructing a hierarchical privacy policy compliance detection index system according to domestic related laws and regulations and standards; secondly, collecting an original text of the privacy policy through a web crawler technology, and constructing an unlabeled corpus after cleaning and structured processing; thirdly, constructing a multi-label classification data set based on a mode of combining a large language model and manual review; mapping the text and the label to a unified semantic vector space by adopting a text-label joint embedding strategy, and inputting the text and the label into a multi-granularity classification model; according to the model, on the basis of an ERNIE pre-training model, context feature enhancement and deep semantic interaction are realized through a bidirectional long-short-term memory network, a self-attention mechanism and a text-label cross attention mechanism, so that the multi-label classification performance is remarkably improved; finally, according to a label prediction result output by the model and a preset index system, compliance judgment is automatically completed, and a structured detection report is generated. According to the invention, the automation degree and efficiency of privacy policy compliance detection are effectively improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Underground space entity identification method and system based on open source spatio-temporal data

The invention relates to the technical field of information perception and intelligent processing, in particular to an underground space entity recognition method and system based on open source spatio-temporal data, and the method comprises the steps: collecting multi-modal spatio-temporal data containing underground space information from a multi-source open source platform through a web crawler and an API interface; performing standardized preprocessing, including denoising, word segmentation, part-of-speech tagging and syntactic analysis, on an original text, and improving semantic recognition capability in combination with a BiLSTM-CRF model and a BERT model; identifying an underground space related entity through fusion of a BERT + CRF model and a domain rule; and complementing missing space-time attribute information based on a space-time clustering algorithm and a graph neural network, and constructing an entity-space-time structure. According to the method, through fusion of multi-source open source data, deep semantic analysis and space-time intelligent modeling, high-precision dynamic identification and full-dimensional information completion of underground space entities are realized, and the intelligent level and decision-making efficiency of underground facility management are remarkably improved.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Product Label Extraction Method Based on Internet Big Data and AI Large Language Model

The present invention relates to the technical field of product label extraction, and specifically, to a product label extraction method based on Internet big data and AI large language models. It includes the following steps: S1. Use web crawler technology to capture the text data of products on the Internet; S2. Adopt the TF-IDF algorithm to determine the important words in the text data, and combine the Skip-Gram model to capture the semantic associations between words. In the process of capturing the semantic associations between words, introduce the weight reflecting the user browsing frequency and the user behavior feature vector to optimize the capturing process; S3. Based on the extracted important words and the semantic association information between words, use a large-scale pre-trained language model to generate product labels; S4. Combine the sequence annotation model BERT and the conditional random field CRF to locate and classify the product labels, and output the finally extracted product labels. The technology of the present invention can effectively locate and classify product labels by combining the BERT model and the conditional random field (CRF) layer.
Owner:BEIJING TAOMI TECHNOLOGY CO LTD

Hotel quotation dimension multi-supply chain bidding rotation system

The invention discloses a hotel quotation dimension multi-supply chain bidding rotation system, and relates to the technical field of hotel supply chain management. Comprising a data acquisition and integration module for collecting related data of supply chain services of departments in a hotel, docking with a management system in the hotel through an interface, realizing automatic acquisition and transmission of the data, and obtaining external market information through a web crawler technology and an industry data platform interface; and the multi-dimensional evaluation calculation module comprises a price evaluation sub-module, a quality evaluation sub-module, a delivery date evaluation sub-module and an after-sales service evaluation sub-module. According to the method, key topics and key words in policies, regulations and news are extracted, core key points are quickly grasped, and hotels are helped to insight into potential influences of policy and regulation changes on business fields of the hotels in advance; the emotional tendency of policies, laws and hot spot information can be quantified based on emotional analysis, hotels can visually understand market emotions, supply chain strategies can be adjusted in time, and market adaptability is enhanced.
Owner:HANGZHOU CHUXUAN INFORMATION TECHNOLOGY CO LTD