Panoramic data visualization method and system based on multi-source heterogeneous data
Through graph theory algorithm and machine learning, the correlation diagram of enterprise-level data is constructed, which solves the problems of in-depth mining and interactive display of multi-source heterogeneous data, improves data readability and user experience, and optimizes the performance of large-scale data processing.
Patent Information
- Application Number
- CN202510842443.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing multi-source heterogeneous data, the existing visualization platform lacks in-depth mining and visualization of complex relationships between data, which is difficult to meet users' needs for in-depth exploration and personalized insights on data, and there are bottlenecks in performance and efficiency during large-scale data processing.
Graph theory algorithm and machine learning are used to mine the enterprise-level data sets, build an association relationship diagram, and generate a dynamic interactive visual interface. Through data cleaning, clustering and neural network model training, node and edge weights are dynamically adjusted to support user interaction operations.
In-depth correlation mining of multi-source heterogeneous data is realized, the readability and user interactivity of data are improved, richer information display is provided, and the performance of large-scale data processing is optimized.
Smart Images

Figure CN120353791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of enterprise panoramic data visualization, and in particular to a panoramic data visualization method and system based on multi-source heterogeneous data. Background Art
[0002] In an enterprise panoramic data visualization platform, it is usually necessary to integrate and display data from different data sources (such as databases, file systems, external APIs, etc.). These data often have the characteristics of multi-source heterogeneity, including structured data (such as table data in a relational database), semi-structured data (such as data in JSON or XML format), and unstructured data (such as text files, images, etc.). When dealing with multi-source heterogeneous data, existing visualization platforms mainly focus on simply summarizing and presenting the data, lacking in-depth mining and visualization of the complex relationships between data. In terms of user interactivity, most only provide basic operations such as zooming, panning, and filtering, making it difficult to meet users' needs for in-depth exploration and personalized insights of data. In addition, current visualization methods often use static visualization templates and cannot adjust the visualization presentation in real time according to the dynamic changes of the data. There are also certain bottlenecks in performance and efficiency when dealing with large-scale data.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a comprehensive review, nor is it intended to identify key / important elements or delineate the scope of protection of these embodiments. Instead, it serves as a preamble to the subsequent detailed description.
[0005] Embodiments of this application provide a panoramic data visualization method and system based on multi-source heterogeneous data to effectively construct a panoramic view of multi-source heterogeneous data.
[0006] In some embodiments, the panoramic data visualization method based on multi-source heterogeneous data includes: extracting enterprise data and industry data to form an enterprise-level data set; cleaning the enterprise-level data set; performing relationship mining on the enterprise-level data set according to graph theory algorithms and machine learning to construct an association relationship graph; and generating a dynamic interactive visualization interface based on the association relationship graph for displaying enterprise data information to users.
[0007] Optionally, the data cleaning of the enterprise-level data set includes: inputting the enterprise-level data set into a language processor to obtain an output vector of each language processor, where there are multiple language processors; calculating a legitimacy coefficient of each language processor according to the output vector; screening out special elements in the output vector among the language processors with a legitimacy coefficient higher than a preset value, where the special elements include elements with an occurrence frequency lower than a preset value; and deleting the special elements from the enterprise-level data set.
[0008] Optionally, the constructing of the association relationship graph includes: distinguishing unstructured data in the enterprise-level data set and structuring the unstructured data; grouping the data using a clustering algorithm to determine multiple word sets; training a neural network model using the word sets, and constructing an association relationship graph based on the neural network model.
[0009] Optionally, grouping the data using a clustering algorithm to determine multiple word sets includes: determining sentence similarity using the following formula:
[0010] where C_main is the similarity between subjects, C 动 represents the similarity between verbs, and C 副 represents the similarity between adverbs. a1, a2, and a3 are weights respectively, and the relationship is a1 > a2 > a3; among the sentences with a sentence similarity greater than a preset threshold, determine word similarity using the following formula:
[0011] where C 词 represents the similarity between word A and word B, and Threshold represents the similarity threshold; construct one or more word sets based on the word similarity.
[0012] Optionally, the neural network training process includes: in a first time period, training the model using a first learning rate; in a second time period, training the model using a second learning rate; where the end time of the first time period is less than or equal to the start time of the second time period.
[0013] Optionally, it further includes: dynamically adjusting the node weights and / or edge weights in the association relationship graph according to data information, where the data confidence includes one or more of data importance, user query frequency, and user marking times.
[0014] Optionally, it further includes: rendering the visualization interface according to the current screen display range when the data access volume is greater than a threshold.
[0015] In some embodiments, a panoramic data visualization system based on multi-source heterogeneous data is further provided, including: a set construction module for extracting enterprise data and industry data to form an enterprise-level data set; a cleaning module for cleaning the enterprise-level data set; a relationship mining module for mining relationships in the enterprise-level data set according to graph theory algorithms and machine learning to construct an association relationship graph; and a display module for generating a dynamic interactive visualization interface based on the association relationship graph to display enterprise data information to users.
[0016] In some embodiments, an electronic device is further provided, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned panoramic data visualization method based on multi-source heterogeneous data.
[0017] In some embodiments, a computer-readable storage medium is further provided. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above-mentioned panoramic data visualization method based on multi-source heterogeneous data.
[0018] The panoramic data visualization method and system based on multi-source heterogeneous data provided by the embodiments of the present application can achieve the following technical effects: Quickly and deeply mine the association relationships between heterogeneous data according to graph theory algorithms and machine learning, and dynamically display the association relationships to users, effectively improving the readability of the data and providing users with richer information display.
[0019] The above general description and the following description are only exemplary and explanatory, and are not used to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and among them: Figure 1 is a flowchart of a panoramic data visualization method based on multi-source heterogeneous data according to an embodiment of the present application; Figure 2 is a structural schematic diagram of a panoramic data visualization system based on multi-source heterogeneous data according to an embodiment of the present application; Figure 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The attached drawings are for reference and illustration only and are not used to limit the embodiments of the present application. In the following technical description, for the sake of explanation, numerous details are provided to give a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0022] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0023] Unless otherwise specified, the term "plurality" means two or more.
[0024] In the embodiments of the present application, the character " / " indicates that the objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0025] The term "and / or" is an associative relationship describing an object and indicates that three relationships can exist. For example, A and / or B means: A or B, or, A and B these three relationships.
[0026] Combined Figure 1 As shown, it is a flowchart of a panoramic data visualization method based on multi-source heterogeneous data provided by the embodiments of the present application. The method includes the following steps: S101: Extract enterprise data and industry data to form an enterprise-level data set; Specifically, the above steps are a data collection and preprocessing process. Data is extracted from the enterprise's relational database (such as MySQL) using an adapter, and at the same time, web crawler technology is used to obtain the enterprise's industry data from external websites. For different formats of data, such as sales data extracted from a CSV file and report data extracted from a PDF file, different conversion rules are used to convert them into a unified data structure.
[0027] The data sources collected in the embodiments of the present application include but are not limited to the enterprise's internal database, file system, data provided by third-party data services, etc. The adapter pattern is used to extract data from different data sources and convert it into a unified data format according to the different types and structures of the data. For example, structured data is converted into a general data table format, and key information is extracted from semi-structured and unstructured data and converted into a feature vector representation.
[0028] In a power enterprise or power system, multi-source heterogeneous data usually includes the following types of key data: device monitoring data, operating parameters such as transformer temperature, line current and voltage collected in real time by sensors, device status detection data such as infrared imaging and partial discharge, environmental and meteorological data, lightning location system, meteorological monitoring data such as wind speed, temperature and humidity, geological disaster monitoring data (such as earthquake and landslide warnings), business management data, structured database information such as power grid topology and asset ledgers, semi-structured text records such as inspection reports and maintenance work orders, power market transaction data, user power consumption behavior data, satellite remote sensing, image / video data from drone inspections, cross-system data, heterogeneous data streams from different platforms such as SCADA systems and EMS systems, and also includes device inspection reports (unstructured data such as PDF / images) from third-party suppliers. After these data are integrated through dynamic visualization technology, a panoramic view covering "devices - environment - business" can be formed to assist in fault prediction and dispatching decisions.
[0029] S102: Clean the enterprise-level data set; This includes removing noise data, handling missing values and outliers, and performing standardization or normalization on the data to ensure that data from different sources are processed under the same dimension.
[0030] S103: Conduct relationship mining on the structured enterprise-level data set based on graph theory algorithms and machine learning to construct an association graph; S104: Generate a dynamic interactive visualization interface based on the association graph for presenting enterprise data information to users.
[0031] The panoramic data visualization method based on multi-source heterogeneous data proposed in the embodiments of this application can quickly and deeply mine the association relationships between heterogeneous data according to graph theory algorithms and machine learning, and dynamically display the association relationships to users, effectively improving the readability of the data and providing users with richer information display.
[0032] In a feasible implementation manner, first, the multi-source heterogeneous data in the enterprise-level data combination is input into a language processor to obtain an output vector for each language processor, where there are multiple language processors; according to the output vector, calculate the legitimacy coefficient of the language processor; among the language processors with a legitimacy coefficient higher than the preset value, screen out the special elements in the output vector, where the special elements include elements with an appearance frequency lower than the preset value; delete the special elements from the multi-source heterogeneous data.
[0033] Specifically, the above language processors include one or more of DeepSeek, GPT-4, ERNIE Bot, Baichuan, and ChatGLM. Multiple language processors can be set. For example, DeepSeek, ERNIE Bot, and GPT-4 can be set as the language processors in the embodiments of the present application. Multisource heterogeneous data is input into these language processors to obtain output vectors of the language processors. Each output vector includes multiple elements. Since each type of language processor has different capabilities in processing various structural data, first, the legitimacy of the language processor is evaluated, that is, whether the language processor can effectively perform heterogeneous data processing. The specific method for determining legitimacy can adopt the consistency of the output vectors. When the output vectors of individual language processors are significantly less consistent with the output vectors of other language processors, then that language processor is excluded.
[0034] Further, for the language processors whose legitimacy coefficients meet the preset requirements, special elements in the output vectors, that is, elements with significantly lower occurrence frequencies, which may be noise data, abnormal data, etc., are screened out and deleted. The final heterogeneous data obtained is the cleaned heterogeneous data set.
[0035] Through the above data cleaning process, known language models can be used for data preprocessing, improving data processing efficiency and accuracy, and further screening multiple language models, further ensuring data accuracy.
[0036] In some feasible embodiments, constructing the association relationship graph includes: distinguishing unstructured data in the enterprise-level data set and structuring the unstructured data; using a clustering algorithm to group the data to determine multiple word sets; training a neural network model using the word sets and constructing an association relationship graph based on the neural network model.
[0037] In some feasible embodiments of the present invention, using a clustering algorithm to group the data to determine multiple word sets includes: Using the following formula to determine sentence similarity:
[0038] Among them, C_sub is the similarity between subjects, C 动 represents the similarity between verbs, and C 副 represents the similarity between adverbs. a1, a2, and a3 are weights respectively, and the relationship is a1 > a2 > a3; Among the sentences where the sentence similarity is greater than the preset threshold, use the following formula to determine word similarity:
[0039] Among them, C 词Indicates the similarity between word A and word B, and Threshold indicates the similarity threshold; Based on the word similarity, construct one or more word sets.
[0040] For the employee data and project data of an enterprise, group the employees by department and project through a clustering algorithm, use named entity recognition technology to extract the project names and person in charge from text reports, and establish the connections between them.
[0041] In the embodiments of the present application, the relationship between data usually manifests as the correlation between data. The specific steps include: S1: Distinguish structured data and unstructured data, and structure the unstructured data. Specifically, use natural language processing technology to extract entities and the relationships between entities, such as extracting entities such as company departments, personnel, and business processes from text files, and establish the connections between them through entity recognition and relationship extraction.
[0042] S2: Use a clustering algorithm to group the data, identify different word sets, and at the same time quantitatively evaluate the data relationships within and between the word sets to provide a basis for subsequent visual layout.
[0043] Specifically, in the embodiments of the present application, for the knowledge fusion in the process of graph construction, it is a process of aggregating words with higher semantic similarity. For example, A = {a1, a2, a3,..., am}, where a1 to am represent m words with the same meaning, and A is the fused word set. Therefore, how to determine semantic similarity in the data is an important link in constructing the management relationship graph.
[0044] The embodiments of the present application adopt a two-stage method to determine the word set.
[0045] In the first stage, determine the sentence similarity. Use the following formula to calculate the sentence similarity:
[0046] Among them, C_main is the similarity between the subjects, C 动 represents the similarity between the verbs, C 副 represents the similarity between the adverbs. a1, a2, and a3 are weights respectively, and the relationship is a1 > a2 > a3. This can ensure the accuracy of the sentence similarity degree.
[0047] In the second stage, on the basis of determining the sentence similarity, when the sentence similarity is greater than the first threshold, determine the similar words in the sentence, that is, determine the word similarity.
[0048] Specifically, use the following formula to determine the similarity between word A and word B:
[0049] Among them, C 词 represents the similarity between word A and word B, and Threshold represents the similarity threshold.
[0050] Through the two-stage similarity determination, it is possible to avoid words in different contexts from being classified into the same word set, improving the determination effect of the word set.
[0051] In some feasible embodiments, the neural network training process includes: in the first time period, training the model with the first learning rate; in the second time period, training the model with the second learning rate; wherein, the end time of the first time period is less than or equal to the start time of the second time period.
[0052] Specifically, an improved CNN neural network is used for data clustering. In traditional algorithms, the learning rate is usually fixed, and over time, it may lead to missing the optimal solution during the learning process. To address the above problem, the system provided in the embodiments of the present application uses a dynamic learning rate that changes over time. The relational expression of the learning rate a is as follows:
[0053] where t is the time period, t set is the time threshold, and t max is the maximum time period. It can be seen from the above expression that the learning rate a in the embodiments of the present application changes in segments over time, and the exponential iterative step size decay process is adopted for the learning rate, which can effectively avoid the system from missing the optimal solution, while accelerating the calculation time and effectively suppressing the system oscillation divergence.
[0054] Optionally, the above method further includes: dynamically adjusting the node weights and / or edge weights in the association relationship graph according to data information, where the data confidence includes one or more of data importance, user query frequency, and user marking times.
[0055] Specifically, the force-directed layout algorithm can be used to generate an initial visual layout according to the data relationship and the distribution of clusters. At the same time, according to the data information, the weights of the nodes and edges are dynamically adjusted. For example, for the data nodes that the user focuses on, increase their node size and the thickness of the edges to improve their display priority. The data information includes data importance, user query times, user marking times, etc.
[0056] Furthermore, time series analysis is introduced. For data with a time dimension, the visual layout is dynamically adjusted according to the passage of time, so that the visual elements can change according to the evolution of time. The user can perform interactive operations through the time slider or time axis to view the data distribution and relationship at different time points.
[0057] Through the above dynamic layout and interactive functions, users can explore data in a more flexible way, meeting the personalized needs of users and the analysis needs in different scenarios.
[0058] In some feasible embodiments, the above method further includes: when the data access volume is greater than a threshold, rendering the visualization interface according to the display range of the current screen.
[0059] In actual application, a Web-based visualization interface can be developed to support users to achieve complex interactive functions through mouse and keyboard operations. Users can click on nodes to view detailed information, rearrange nodes through drag-and-drop operations, and can fold and expand data clusters to control the information density displayed.
[0060] Furthermore, a data filtering and screening function is implemented. Users can filter data according to specific conditions (such as time range, data type, data attributes, etc.), and the visualization interface will update the layout and display content in real time according to the filtering conditions.
[0061] To increase the user-friendliness of the system, users can be allowed to customize the styles and display attributes of visualization elements, including colors, shapes, transparencies, etc., to meet the personalized needs of different users.
[0062] The above interactive visualization interface allows users to view the detailed information of employees and projects under a department by clicking on the department node, view the project progress in different months using the timeline. Users can adjust the layout by dragging nodes, and can also screen out employees within a specific performance range.
[0063] In some feasible embodiments, data sampling and progressive rendering techniques can be further adopted. For large-scale data, only the data in the currently visible area of the user is rendered. When the user scrolls or pans the interface, the rendering range is updated according to the user's operation, improving the rendering performance of the visualization.
[0064] At the same time, a data update mechanism is established. When the data in the data source changes, the visualization layout and data display are updated in real time. Through an incremental update algorithm, only the changed data part is updated, reducing the latency and performance overhead of data update.
[0065] In terms of performance optimization, when users view a large amount of data, only the data in the visible area of the user is rendered. When the user scrolls the screen, the rendering area is dynamically updated. When new project data is added to the database, the visualization display is updated in real time.
[0066] Through the above performance optimization techniques, it is ensured that when dealing with large-scale data, a good user experience can still be maintained, ensuring the fluency and real-time nature of the visualization platform.
[0067] Based on the above method embodiments, an embodiment of the present invention further provides a panoramic data visualization system 200 based on multi-source heterogeneous data. Referring to Figure 2 as shown, the system includes: A set construction module 201, configured to extract enterprise data and industry data to form an enterprise-level data set; A cleaning module 202, configured to perform data cleaning on the enterprise-level data set; A relationship mining module 203, configured to perform relationship mining on the enterprise-level data set according to graph theory algorithms and machine learning, and construct an association relationship graph; A display module 204, configured to generate a dynamic interactive visualization interface based on the association relationship graph for displaying enterprise data information to users.
[0068] The above-mentioned panoramic data visualization system based on multi-source heterogeneous data proposed in the embodiments of the present application quickly and deeply mines the association relationships between heterogeneous data according to graph theory algorithms and machine learning, and dynamically displays the association relationships to users, effectively improving the readability of the data and providing users with richer information display.
[0069] An embodiment of the present invention further provides an electronic device. As Figure 3 shown, it is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 301 and a memory 302. The memory 302 stores computer-executable instructions that can be executed by the processor 301, and the processor 301 executes the computer-executable instructions to implement the above-mentioned panoramic data visualization method based on multi-source heterogeneous data.
[0070] In Figure 3 the illustrated embodiment, the electronic device further includes a bus 303 and a communication interface 304. Among them, the processor 301, the communication interface 304, and the memory 302 are connected through the bus 303.
[0071] Among them, the memory 302 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 304 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 303 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 403 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a bidirectional arrow is used in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0072] The processor 301 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 301. The above-mentioned processor 301 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor 301 reads the information in the memory and combines its hardware to complete the steps of the panoramic data visualization method based on multi-source heterogeneous data in the foregoing embodiments.
[0073] An embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above panoramic data visualization method based on multi-source heterogeneous data. For specific implementation, reference may be made to the foregoing method embodiments, which will not be elaborated herein.
[0074] The above description and drawings fully illustrate the embodiments of the present application, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or substituted for parts and features of other embodiments. Moreover, the terms used in this application are only for describing embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, or apparatus comprising the element. Herein, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.
[0075] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. The skilled person can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present application. The skilled person can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0076] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. In addition, the functional units in the embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A panoramic data visualization method based on multi-source heterogeneous data, characterized in that, It includes: Extract enterprise data and industry data to form an enterprise-level data set; Clean the enterprise-level data set; Conduct relationship mining on the enterprise-level data set based on graph theory algorithms and machine learning to construct an association relationship graph; Generate a dynamic interactive visualization interface based on the association relationship graph for displaying enterprise data information to users.
2. The method according to claim 1, characterized in that, The cleaning of the enterprise-level data set includes: Input the enterprise-level data set into a language processor to obtain the output vectors of each language processor, where there are multiple language processors; Calculate the legitimacy coefficient of each language processor according to the output vector; In the language processors with a legitimacy coefficient higher than the preset value, screen out the special elements in the output vector, where the special elements include elements with an appearance frequency lower than the preset value; Delete the special elements from the enterprise-level data set.
3. The method according to claim 1, wherein The construction of the association relationship graph includes: Distinguish the unstructured data in the enterprise-level data set and structure the unstructured data; Use a clustering algorithm to group the data to determine multiple word sets; Train a neural network model using the word sets and construct an association relationship graph based on the neural network model.
4. The method according to claim 3, wherein Using a clustering algorithm to group the data to determine multiple word sets includes: Use the following formula to determine the sentence similarity: Among them, C_main is the similarity between subjects, C 动 represents the similarity between verbs, C 副 represents the similarity between adverbs, a1, a2, and a3 are weights respectively, and the relationship is a1 > a2 > a3; In the sentences with a sentence similarity greater than the preset threshold, use the following formula to determine the word similarity: Among them, C 词 represents the similarity between word A and word B, and Threshold represents the similarity threshold; Construct one or more word sets based on the word similarity.
5. The method according to claim 3, characterized in that, The neural network training process includes: In the first time period, train the model using the first learning rate; In the second time period, train the model using the second learning rate; where the end time of the first time period is less than or equal to the start time of the second time period.
6. The method according to any one of claims 1 to 5, characterized in that, It also includes: Dynamically adjust the node weights and / or edge weights in the association relationship graph according to the data information, where the data confidence includes one or more of data importance, user query frequency, and user marking times.
7. The method according to any one of claims 1 to 5, characterized in that It also includes: When the data access volume is greater than the threshold, render the visualization interface according to the current screen display range.
8. A panoramic data visualization system based on multi-source heterogeneous data, characterized in that, It includes: A set construction module for extracting enterprise data and industry data to form an enterprise-level data set; A cleaning module for cleaning the enterprise-level data set; A relationship mining module for conducting relationship mining on the enterprise-level data set based on graph theory algorithms and machine learning to construct an association relationship graph; A display module for generating a dynamic interactive visualization interface based on the association relationship graph for displaying enterprise data information to users.
9. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores computer executable instructions that can be executed by the processor. The processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer executable instructions. When the computer executable instructions are called and executed by the processor, the computer executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for extracting sentences with similar meanings and standard grammar from academic documents
CN105677634A
Text infringement detection method and device, electronic equipment and storage medium
CN114564936A
Sentence vector generation method and device, statement similarity determination method and device and electronic equipment
CN114625841A
Enterprise information management method and system
CN118277638A
Cited By
Dynamic interactive enterprise panoramic visualization system based on multi-source heterogeneous data
CN120892495A