A Visualization Method for Data Analysis Pipeline of Space Science Experiments
By generating data, models and visualization nodes on the user interface, interactive data analysis pipeline operation is implemented, and the user interface does not support interactive parameter configuration and visualization of analysis results is solved, improving the convenience and intuitiveness of data analysis.
Patent Information
- Application Number
- CN202411444408.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-16
AI Technical Summary
In the prior art, the user interface does not support interactive parameter configuration, and users cannot directly view the analysis results in visualization in the current interface, and the scheduling efficiency of algorithm analysis jobs is also low, resulting in poor user data analysis convenience.
Provide a visual spatial science experimental data analysis pipeline method, which generates data nodes, model nodes and visualization nodes on the same user interface, supports interactive node selection and layout, node relationship connection, data and algorithm configuration, analysis progress monitoring and analysis results visualization.
It realizes that users can interactively complete data selection, algorithm configuration and result visualization in the same interface, simplifies the operation process, improves the intuitiveness and convenience of data analysis, and supports multi-process parallel orchestration and automated management, reducing the learning cost and usage threshold of the system.
Smart Images

Figure CN119398477B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data analysis, and particularly to a visual data analysis pipeline method for space science experiments. Background Art
[0002] The online analysis and application research of multi-task data in space science plays a decisive role in improving the research and utilization level of massive and multi-source space science experiment data, promoting the output of results and knowledge discovery, and better exerting the application benefits of space science experiment data.
[0003] Currently, with the use of computer technology, the visualization of electronic information on computer devices can be achieved, and then the man-machine interactive data analysis research can be realized.
[0004] However, in the related art, the user interface does not support interactive parameter configuration for each node, the user cannot directly view the analysis results visually in the current user interface, the user needs to manually manage the algorithm modules in the server, and the scheduling efficiency of algorithm analysis jobs also needs to be improved, which makes the convenience of user data analysis poor. Summary of the Invention
[0005] The embodiments of this application provide a visual data analysis pipeline method for space science experiments, which can solve the technical problem of poor convenience in space science experiment data analysis.
[0006] To achieve the above object, the embodiments of this application adopt the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a method for a visual data analysis pipeline of space science experiments. The method for the visual data analysis pipeline of space science experiments includes: generating a data node, a model node, and a visualization node in response to an input data selection operation, an algorithm model selection operation, and a visualization component selection operation; the data node, the model node, and the visualization node are located on the same user interface; performing parameter configuration on the data node, the model node, and the visualization node, and performing a pipeline connection operation to generate a flowchart; wherein, the output of the data node is connected to the input of the model node, or the output of the data node is connected to the input of the visualization node, and the output of the model node is connected to the input of the visualization node; based on the flowchart, receiving an input job submission operation; parsing the flowchart to obtain a parsing result; the parsing result is used to represent the dependency relationship between the data node, the model node, and the visualization node and the sorting result of the data stream; based on the parsing result, obtaining a job scheduling result; based on the job scheduling result, retrieving an algorithm image deployment package corresponding to the algorithm model selection operation and file data corresponding to the data selection operation; the file data includes video data and image data; performing algorithm analysis on the file data according to the algorithm image deployment package to obtain an analysis result; visually displaying the analysis result; wherein, the analysis result includes images, videos, strings, two-dimensional charts, and / or three-dimensional charts.
[0008] Based on the above description of the method for the visual data analysis pipeline of space science experiments provided by the embodiment of the present application, it can be seen that the method for the visual data analysis pipeline of space science experiments includes interactively completing node selection and layout, node relationship connection, data and algorithm configuration, analysis progress monitoring, and analysis result visualization in the same user interface, so that the entire data algorithm analysis process forms a visual and seamless whole, improving the intuitiveness of user data analysis.
[0009] Moreover, the embodiment of the present application realizes multi-process pipeline analysis jobs and their automated management, supports parallel arrangement of multiple algorithm analysis processes. After the arrangement is completed, the system automatically establishes the dependency relationship of the algorithm analysis jobs, calls the analysis algorithm to run the jobs, and collects the analysis progress, analysis results, and operation logs, greatly simplifying the operation complexity of users and reducing the learning cost and usage threshold of the system.
[0010] Moreover, the embodiment of the present application realizes displaying the original data and the analysis result on the same interface. When the analysis job is completed, the user can synchronously view the visualization views of the original data and the analysis result in the current user interface, intuitively comparing and displaying the differences and characteristics between the analysis result and the original data, improving the convenience of user data analysis.
[0011] Moreover, through the automated deployment of the algorithm model and the automated management of analysis jobs in the embodiments of the present application, any programming language and framework are supported, enabling algorithm developers to focus on the development of the algorithm itself without concerning about the underlying software and hardware resources of the system and the operation scheduling issues of the analysis algorithm, which helps to improve the scalability of the system and the utilization rate of the underlying software and hardware resources.
[0012] In a feasible implementation manner of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: based on web page embedding, setting the data node, the model node, and the visualization node on the same user interface.
[0013] In a feasible implementation manner of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: the pipeline connection operation includes a multi-process pipeline node connection operation; the multi-process pipeline node connection operation includes a node connection method based on a directed acyclic graph structure to connect the isolated data node, model node, and visualization node into a complete algorithm analysis pipeline.
[0014] In a feasible implementation manner of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: based on the flowchart, verifying whether there are isolated nodes; if there are no isolated nodes, then based on the flowchart, verifying whether there are missing input edges; if there are no missing input edges, then based on the flowchart, verifying whether the data types match; if the data types match, then based on the flowchart, verifying whether the required parameters are filled; if all the required parameters are filled, then based on the flowchart, verifying whether there is a loop; if there is no loop, then based on the flowchart, receiving the input job submission operation.
[0015] In a feasible implementation manner of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: based on the flowchart, obtaining multiple jobs; for the jobs with a first relationship, sending them to the analysis algorithm on the computing instance one by one in a serial form.
[0016] In a feasible implementation manner of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: based on the flowchart, obtaining multiple jobs; for the jobs with a second relationship, sending them to the analysis algorithm on the computing instance in a parallel form.
[0017] In a feasible implementation manner of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: the algorithm image deployment package is used to cooperate in generating an execution container; the execution container includes a processor, memory, graphics processor resources, and a mounted file directory.
[0018] In a feasible implementation of the first aspect, the method for visualizing the data analysis pipeline of space science experiments further includes: presetting an image template; according to an algorithm configuration file, reading the image template, copying the code and data of the algorithm model into the image, automatically installing the dependent third-party libraries, and constructing the algorithm image deployment package.
[0019] In a second aspect, an embodiment of the present application provides a system for visualizing the data analysis pipeline of space science experiments. The system for visualizing the data analysis pipeline of space science experiments includes: at least one processor; a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided in the first aspect.
[0020] In a third aspect, an embodiment of the present application provides a computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method provided in the first aspect. Description of the Drawings
[0021] Figure 1 It is a schematic structural diagram of a system for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0022] Figure 2 It is a schematic structural diagram of a system for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0023] Figure 3 It is a schematic flowchart of a method for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0024] Figure 4 It is a schematic diagram of a data selection node in a method for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0025] Figure 5 It is a schematic diagram of an algorithm model node in a method for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0026] Figure 6 It is a schematic diagram of a data visualization node in a method for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0027] Figure 7 It is a schematic diagram of the log of algorithm model automatic deployment in a method for visualizing the data analysis pipeline of space science experiments provided by an embodiment of the present application;
[0028] Figure 8Schematic diagram of a computing instance in a visual spatial science experiment data analysis pipeline method provided by an embodiment of the present application;
[0029] Figure 9 Flow chart of a visual spatial science experiment data analysis pipeline method provided by an embodiment of the present application;
[0030] Figure 10 Visual overall interface of the flow chart of a visual spatial science experiment data analysis pipeline method provided by an embodiment of the present application. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention. Among them, in the description of the embodiments of the present invention, unless otherwise specified, "a plurality of" means two or more than two. "At least one (piece)" or its similar expression refers to any combination of these items, including any combination of a single item (piece) or plural items (pieces). For example, at least one (piece) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.
[0032] In addition, in order to clearly describe the technical solutions in the embodiments of the present invention, in the embodiments of the present invention, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different. At the same time, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0033] The principles and features of the present application are described below, and the examples given are only used to explain the present application and are not used to limit the scope of the present application.
[0034] Visualization refers to the technology of presenting data, information, and knowledge in an intuitive and understandable way through visualization tools such as charts, graphs, and maps. It helps people better understand the process of event development and discover the laws contained in the data. A user interface based on visualization enables users to find the required information or functions more quickly, reducing the time required to complete tasks. Completing data selection, data analysis configuration, and result visualization in the same user interface is a necessary approach to achieving efficient and convenient data analysis and also the development trend of current data visualization analysis software.
[0035] Automated deployment is a process of automatically deploying application programs from the development environment to the production or other environments through software tools. This process reduces the need for manual operations, thereby reducing the risk of human errors and increasing the deployment efficiency and frequency. Automated deployment usually includes steps such as code building, testing, configuration management, and application release, ensuring that software updates can be deployed to the target system quickly and reliably.
[0036] Dynamic scheduling is a method of adjusting and optimizing workload distribution according to resource availability and job requirements at runtime. This method is particularly suitable for environments where there are dependencies between jobs and resources are limited. It can adjust the scheduling strategy based on the real-time system state and resource usage to cope with unforeseen changes. Since dynamic scheduling makes decisions based on real-time information, it can adjust task priorities and execution orders during operation. By dynamically reallocating jobs to the most suitable resources, the overall performance of the system can be maximized. If a job fails or lacks necessary resources, dynamic scheduling can reschedule the affected jobs. Therefore, dynamic scheduling is often used to manage the execution of computationally intensive tasks, especially in fields such as cloud computing and data analysis.
[0037] The embodiments of this application provide a method for a visualization-based data analysis pipeline for space science experiments, which is applicable to engineering application scenarios of data types such as manned space science videos and images. By designing a flowchart user interface, automated deployment of algorithm models, and dynamic scheduling of analysis jobs, multi-process, interactive, intuitive, and quantitative algorithm analysis of space science experiment data is achieved.
[0038] The embodiments of this application provide a system for a visualization-based data analysis pipeline for space science experiments, which can execute the method for a visualization-based data analysis pipeline for space science experiments provided by the embodiments of this application. Figure 1 It is a schematic structural diagram of a system for a visualization-based data analysis pipeline for space science experiments provided by the embodiments of this application.
[0039] As Figure 1As shown, the visualization spatial science experiment data analysis pipeline system includes at least one processor 011 and a memory 012 communicatively connected to the at least one processor; wherein, the memory 012 stores instructions executable by the at least one processor 011, and when the instructions are executed by the at least one processor 011, the at least one processor 011 is enabled to execute the visualization spatial science experiment data analysis pipeline method provided in the embodiments of the present application.
[0040] In some embodiments, the visualization spatial science experiment data analysis pipeline system may be an electronic device provided with a display screen.
[0041] The display screen can be used to display a user interface to implement visualization and human-computer interaction functions. For example, display images, videos, etc. The display screen includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include 1 or N display screens, where N is an integer greater than 1.
[0042] Exemplarily, the electronic device in the embodiments of the present application may include, but is not limited to: a folding screen device (such as: a folding screen mobile phone, a folding screen computer, etc.), a tablet computer, a handheld computer, a netbook, a personal computer (PC), a personal digital assistant (PDA), a portable multimedia player (PMP), an augmented reality (AR) / virtual reality (VR) device, a television, a vehicle-mounted computer, etc. Alternatively, the electronic device may also be an electronic device with a display screen of other types or structures. The embodiments of the present application do not limit the specific device form of the electronic device.
[0043] Such as Figure 2As shown, in some embodiments, the visualized spatial science experiment data analysis pipeline system further includes a flowchart user interface module, a distributed database and a file system, an algorithm model automatic deployment module, an analysis job dynamic scheduling module, and a computing instance middleware module.
[0044] In one implementation, the flowchart user interface module includes a data selection component, an algorithm model component, and a data visualization component.
[0045] In one implementation, the flowchart user interface module further includes a multi-process pipeline node connection component, each node parameter configuration component, and a multi-process pipeline node layout component.
[0046] Figure 3 It is a schematic flowchart of a visualized spatial science experiment data analysis pipeline method provided by an embodiment of the present application. As Figure 3 shown, in some embodiments, the visualized spatial science experiment data analysis pipeline method includes the following steps:
[0047] S1, in response to an input data selection operation, algorithm model selection operation, and visualization component selection operation, generate a data node, a model node, and a visualization node.
[0048] The data node, the model node, and the visualization node are located on the same user interface.
[0049] The data selection operation refers to selecting video data or image data as the data item to be analyzed by the algorithm. In some embodiments, the video data or image data can be pre-stored in the distributed database and the file system. When performing the data selection operation, it can be achieved by selecting the file name of the video data or image data.
[0050] As Figure 4 shown, in one implementation, the user can select "Original Video Data.mp4" by clicking on the selection box "Open Data Panel" on the user interface to implement the user's data selection operation. It can be understood that the data selection operation can also be implemented by non-contact operations such as the user's eye movement, language, or gesture. The specific form of the data selection operation is not limited in this application.
[0051] In one implementation, in response to the input data selection operation, generate a data node as Figure 4 shown. This data node component has only one output. This output is used to connect to the input of the model node or the visualization node.
[0052] In some embodiments, drag the data node component to the canvas, and select a piece of data through the data menu on the data node. After saving, this data node represents this piece of data. The data node has only an output interface.
[0053] The algorithm model selection operation refers to selecting a model. In some embodiments, data related to the algorithm can be pre-stored in the algorithm model automatic deployment module. When performing the model selection operation, it can be achieved by selecting the name of the model. As Figure 5 shown, in one implementation, the user realizes the algorithm model selection operation by clicking on the selection box "General Semantic Segmentation Model" on the user interface. It can be understood that the algorithm model selection operation can also be achieved through non-contact operations such as the user's eye movement, language, or gesture. The specific form of the algorithm model selection operation is not limited in this application.
[0054] In one implementation, in response to the input algorithm model selection operation, a model node as Figure 5 shown is generated. This model node component has one or more inputs, and can have one or more outputs. Among them, one output can be connected to multiple downstream inputs, and one input can only be connected to one upstream output, that is, the input can only have one source, and the output can be copied multiple times for sharing by multiple downstream nodes. In one implementation, a model is selected from the algorithm model list and dragged onto the canvas to form an algorithm model node. The model node has input and output interfaces, as well as a parameter form.
[0055] In some embodiments, the data node, model node, and visualization node are uniformly displayed on the display screen. The visual display can be achieved through a data visualization component. In one implementation, the data visualization component is implemented in the form of a web (WEB) page embedding, supporting direct data visualization displays such as image comparison visualization and video visualization in the current user interface, and also supporting opening the data visualization view in the visualization node in full-screen mode in a separate window.
[0056] S2. Configure the parameters of the data node, model node, and visualization node, and perform pipeline connection operations to generate a flowchart.
[0057] The parameter configuration operation includes setting algorithm model parameters. There are various data types supported by the parameter configuration. For example, the data type can be a number. Exemplarily, an integer or a floating-point number. The visual display is a number input box. Select the number input box and enter a number in the number input box. Also, for example, the data type can be a string. The visual display is a character input box. Select the character input box and enter a string in the character input box. Again, for example, the data type can be a boolean. The visual display is a checkbox. Also, for example, the data type can be a single selection. The visual display is a dropdown box. As Figure 5As shown, in one implementation, the user configures the parameter of the intersection over union threshold by clicking on the input box of "iou" on the user interface and then entering the value "0.6". It can be understood that the parameter configuration operation can also be implemented through non-contact operations such as the user's eye movement, language, or gesture. The specific form of the parameter configuration operation is not limited in this application.
[0058] In one implementation, each node parameter configuration module is used to configure parameters for specific node components.
[0059] Pipeline connection operation, such as Figure 10 As shown, in some embodiments, drag the connection line between each node (data node, model node, and visualization node). When connecting, the output interface of the node is connected to the input interface of the next node. In this way, all the node connection lines form a directed acyclic graph. For example, the output of the data node is connected to the input of the model node, or the output of the data node is connected to the input of the visualization node.
[0060] In some embodiments, the pipeline connection operation can be implemented through a flowchart component. The user generates a flowchart by connecting each data node, model node, or visualization node on the display screen. After connecting to any eligible node, the node component has one or more inputs. The flowchart view can be scaled and moved as a whole with the flowchart.
[0061] In some embodiments, the pipeline connection operation can be multi-process pipeline node connection. In one implementation, the multi-process pipeline node connection connects isolated nodes into a complete algorithm analysis pipeline in the form of connection lines, and supports the parallel layout and relationship connection of nodes based on the node connection method of the directed acyclic graph structure.
[0062] It can be understood that each process pipeline in the flowchart can be used as a job (also called an algorithm analysis job).
[0063] In some embodiments, the flowchart and form are saved in the database in JSON format, and the information includes node information, node position, and edge information. The node information includes the data corresponding to the node, the algorithm model, the form content, and the visualization component ID.
[0064] The node position includes the horizontal coordinate and vertical coordinate of the node on the canvas. The edge information includes the source node and target node of each edge, and the source interface and target interface of each edge. In one implementation, the multi-process pipeline layout uses an automatic layout algorithm to determine the horizontal coordinate and vertical coordinate of each node in the user interface, so that they are neatly arranged without overlap.
[0065] S3. Based on the flowchart, receive the input job submission operation.
[0066] The operation of submitting a job, including the one input by the user by clicking the "Submit" button on the display screen after the flowchart is generated. It can be understood that the data selection operation can also be implemented through non-contact operations such as the user's eye movement, language, or gesture. For the specific form of the data selection operation, this application does not make any limitations.
[0067] In some embodiments, the operation of submitting a job can be input through a job submission interface. The job submission interface is developed and implemented according to the interface specification algorithm.
[0068] In some embodiments, in addition to the job submission interface, a job query interface and a job cancellation interface can also be set up.
[0069] In some embodiments, when performing step S3, the method for visualizing the spatial science experiment data analysis pipeline further includes:
[0070] S31, based on the flowchart, verify the correctness of its parameters and topological relationships.
[0071] In one implementation, when performing step S31, the method for visualizing the spatial science experiment data analysis pipeline further includes:
[0072] S311, based on the flowchart, verify whether there are isolated nodes.
[0073] S312, if there are no isolated nodes, then based on the flowchart, verify whether there are missing input edges.
[0074] S313, if there are no missing input edges, then based on the flowchart, verify whether the data types match.
[0075] S314, if the data types match, then based on the flowchart, verify whether the required parameters are filled in.
[0076] S315, if all the required parameters are filled in, then based on the flowchart, verify whether there is a loop.
[0077] A loop, also known as a cycle, is a concept in graph theory. If task x must be completed before task y: x → y, and y → z, while z → x. Then a cycle will be formed: x → y → z → x. There should be no loops in a valid flowchart, that is, a directed acyclic graph.
[0078] S316, if there is no cycle, then based on the flowchart, receive the input job submission operation.
[0079] S4, parse the flowchart to obtain the parsing result.
[0080] The parsing result is used to represent the dependency relationship between the data nodes, model nodes, and visualization nodes and the sorting result of the data flow.
[0081] Parse the flowchart, including analyzing the algorithm analysis flowchart arranged by the user, extracting input data, algorithm models, configuration parameters, and topological relationships therefrom, and sorting the nodes in the flowchart according to their dependencies and data flows.
[0082] In some embodiments, hand over the flowchart information to the analysis job dynamic scheduling module for flowchart parsing.
[0083] S5. Based on the parsing result, obtain the job scheduling result.
[0084] According to the parsing result, including the dependencies of the nodes, issue the algorithm analysis jobs to the corresponding computing instances in a serial or parallel manner and run them. At the same time, monitor the job execution progress and perform job scheduling. In this way, the job scheduling result is obtained. Figure 8 Give an example of a computing instance.
[0085] In some embodiments, perform a topological sort on all algorithm model nodes in the flowchart according to the node information and edge information to determine the sequence and dependencies of each analysis job.
[0086] In one implementation, for jobs with a first relationship (i.e., having a dependency relationship), issue them to the analysis algorithms on the computing instances one by one in series. After the analysis algorithm completes the job, the system will update the process progress and execute the subsequent jobs until all the analysis jobs corresponding to the process nodes are completed.
[0087] In another implementation, for jobs with a second relationship (i.e., having no dependency relationship), issue them to the analysis algorithms on the computing instances in parallel. After the analysis algorithm completes the job, the system will update the process progress and execute the subsequent jobs until all the analysis jobs corresponding to the process nodes are completed.
[0088] In some embodiments, when performing step S5, the visual spatial science experiment data analysis pipeline method further includes:
[0089] S51. According to the job scheduling result, display the progress of each job.
[0090] In some embodiments, the progress monitoring can display the progress of the analysis jobs in the algorithm model nodes in real time. Exemplarily, represent the progress of the job in the form of the percentage of the completed amount to the total amount.
[0091] In some embodiments, upload the progress of each job to the distributed database and file system.
[0092] S6. Based on the job scheduling result, retrieve the algorithm image deployment package corresponding to the algorithm model selection operation and the file data corresponding to the data selection operation.
[0093] The file data includes video data and image data. In some embodiments, the data to be analyzed is downloaded from a distributed database and a file system.
[0094] Among them, the video data is a data file stored in a video format, such as mp4, mpeg, avi, etc. The image data supports data files stored in an image compression package format such as png, jpg, bmp, etc., such as rar, zip, etc.
[0095] The algorithm image deployment package is used to cooperate with the generation of execution containers. The algorithm image deployment package can be pre-stored in a visual space science experiment data analysis pipeline system. In some embodiments, the algorithm image deployment package is pre-stored in an algorithm model automated deployment module.
[0096] In some embodiments, an interface and parameter definition file is written according to the system interface specification and packaged together with the algorithm program into an algorithm program to be deployed. In this way, the algorithm image deployment package is obtained.
[0097] Among them, the execution container includes a processor (CPU), memory, a graphics processing unit (GPU) resource, and a mounted file directory, and has functions of generation, start, stop, restart, and deletion. It can also schedule the execution of algorithm containers and cancel analysis tasks. In some embodiments, the container engine (Docker, an operating system-level virtualization container engine) commands of the execution container are executed by the computing instance middleware to implement operations such as starting, stopping, restarting, deleting, and reading logs as Figure 7 shown.
[0098] The construction of the algorithm image deployment package will be described in detail below.
[0099] In some embodiments, before performing step S1, the visual space science experiment data analysis pipeline method further includes:
[0100] S611, preset an image template.
[0101] The algorithm model automated deployment module stores an image template library. The image template library includes multiple image templates. In some embodiments, based on the Docker container engine, according to different hardware and software environments, more than a dozen basic container image templates combined by different hardware and software environments can be preset. For example: Python language program image template that requires GPU, Python language program image template that does not require GPU, Java language program image template that requires GPU, Java language program image template that does not require GPU, etc. In one implementation, the selected basic image is copied to the algorithm program directory, and the algorithm program is built into a Docker image according to the process described in the basic image, and the installation script is run. Then, the running environment is automatically installed according to the environment configuration file in the algorithm program. After the image is built, the image is exported as a compressed file for deployment.
[0102] The image template provides the operating system, system-level libraries, and corresponding installation scripts.
[0103] S612. Based on the image template, construct an algorithm image deployment package.
[0104] The image template is used for the construction of the algorithm image deployment package. In some embodiments, the image template is read from the image template library according to the algorithm configuration file, the code and data of the algorithm model are copied into the image template, and the third-party programs and components (such as: MMCV for computer vision, NumPy for numerical calculation) on which the algorithm model depends are automatically installed. Finally, an algorithm image is generated and packaged and stored at a specified location on the server. It can be understood that the requirements of the third-party programs on which the algorithm model program depends can be described through the environment file. When building the image, reading this environment file can be used to install through the network.
[0105] To ensure the accuracy of constructing the algorithm image deployment package, in some embodiments, before presetting the image template, test whether the interface information and data type of the algorithm meet the preset standards. Only after the test is successful can steps S611 and S612 be executed. As Figure 9 shown, in one implementation, by reading the interfaces and parameter definitions in the algorithm program, its correctness is verified.
[0106] In some embodiments, the algorithm program is uploaded to the system, and information such as name and description is input, and the required image template and the computing instance to be deployed are selected. After the algorithm is uploaded, the visual spatial science experiment data analysis pipeline system will automatically execute the test, packaging, and deployment processes. In one implementation, the exported image is copied to the computing instance and imported into Docker to form a container. Input and output data directories are mounted for the container, and then the container is started to complete the deployment.
[0107] S7. According to the algorithm mirror deployment package, perform algorithm analysis on the file data to obtain the analysis result.
[0108] For the analysis result, store all the data generated by the algorithm, which can be directly visualized and viewed in this system, and data download is also supported.
[0109] S8. Visualize the analysis result.
[0110] In some embodiments, the analysis result is uploaded to the distributed database and the file system.
[0111] In some embodiments, the visualization can be implemented through the data visualization component. In one implementation, the data visualization component is implemented in the form of a web (WEB) page embedding, supporting data visualization displays such as image comparison visualization and video visualization directly in the current user interface. The data visualization view can be scaled and moved as a whole with the flow chart, and also supports opening the data visualization view in a separate window in full screen mode.
[0112] In some embodiments, select a visualization component from the visualization component list and drag it to the canvas to form a visualization node. The visualization node has only input interfaces. As Figure 9 shown, integrate the developed visualization components into the system. The integrated visualization components will appear in the visualization component list.
[0113] To meet the needs of users, as Figure 9 shown, in one implementation, declare the data types and quantities required for the input of the visualization component.
[0114] In some embodiments, develop an interactive data visualization component, including image visualization function, video playback function, string (such as JSON) viewing function, two-dimensional chart visualization function, and three-dimensional chart visualization function. The image visualization function can be used to view a series of images, and if there is an algorithm analysis result, it can be displayed synchronously with the original data. The video playback function can be used to play videos. The JSON viewing function displays the attributes and values of the JSON file in a tree structure. The two-dimensional chart visualization function displays data in the form of scatter points, broken lines, parallel coordinate axes, etc. on the two-dimensional coordinate system. The three-dimensional chart visualization function displays data in the form of scatter points, broken lines, curved surfaces, etc. on the three-dimensional coordinate system.
[0115] The embodiments of the present application implement a unified interactive visualization user interface. By designing a flowchart-style user interface form, node views containing different information such as editable parameters or viewing status are designed for data selection, algorithm configuration, and visualization of analysis results respectively, and support interactive node layout, relationship connection, and parameter setting to integrate all information onto the same page. The design determines the data source and flow direction based on the connection relationship between nodes and determines the running order of the analysis jobs.
[0116] The embodiments of the present application implement a multi-process data algorithm analysis pipeline through the design of a node connection method with a directed acyclic graph structure and an analysis job dynamic scheduling technology. The design uses a directed acyclic graph structure for the parallel layout and relationship connection of nodes. Each node connection branch forms an algorithm analysis pipeline to implement a multi-process algorithm analysis pipeline, constituting an algorithm analysis flowchart integrating data selection, algorithm configuration, and visualization of analysis results. After the algorithm analysis flowchart is set up, the system will sort the nodes in the flowchart according to their dependency relationships and data flow directions and create algorithm analysis jobs. The analysis job dynamic scheduling technology will issue the algorithm analysis jobs to the corresponding computing instances in a serial or parallel manner according to the dependency relationships of the nodes and run them. At the same time, monitor the job execution progress and perform job scheduling until all the analysis jobs corresponding to the nodes are completed.
[0117] The embodiments of the present application implement the visualization of analysis results on the same interface by designing a WEB page embedded data visualization view and directly embedding this view into the current user interface. This view can accept multiple data sources through node connections and synchronously display the original data and analysis results in the embedded visualization view. The visualization view can be scaled and moved as a whole with the flowchart and also supports opening the data visualization view in full-screen mode in a separate window.
[0118] The embodiments of the present application implement the automatic deployment of algorithm models by designing an algorithm model automatic deployment module and defining the algorithm model interface declaration standard. After the algorithm development is completed, the system first tests whether the algorithm model runs normally and verifies whether the algorithm model interface declaration file meets the standard. If the algorithm model passes the test, the system automatically installs the dependent components for it and packages it as a Docker image. Subsequently, the packaged image is automatically sent to the computing instance, and the computing instance is responsible for loading the algorithm model image and providing analysis services externally.
[0119] Based on the same application concept, an embodiment of the present application further provides a visual spatial science experiment data analysis pipeline system. The method corresponding to the visual spatial science experiment data analysis pipeline system may be the visual spatial science experiment data analysis pipeline method in the foregoing embodiments, and the principle of solving problems is similar to that of this method. The visual spatial science experiment data analysis pipeline system provided by the embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of multiple embodiments of the present application described above.
[0120] Another embodiment of the present application further provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application described above.
[0121] Specifically, this embodiment can adopt any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0122] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0123] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including - but not limited to - wireless, wire, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0124] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0125] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0126] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0127] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0128] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0130] The above-mentioned integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units stored in a storage medium include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0131] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.
[0132] In addition, it is obvious that the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. A plurality of elements or devices recited in the apparatus claims may also be implemented by one element or device through software or hardware. Terms such as first, second, etc. are used to denote names and do not denote any particular order.
Claims
1. A visual pipeline method for analyzing space science experiment data, characterized in that: include: In response to input data selection operations, algorithm model selection operations, and visualization component selection operations, data nodes, model nodes, and visualization nodes are generated; The data node, the model node and the visualization node are located on the same user interface; Parameter configuration is performed on the data node, the model node and the visualization node, and a pipeline connection operation is performed to generate a flow chart; wherein the output of the data node is connected to the input of the model node, or the output of the data node is connected to the input of the visualization node, and the output of the model node is connected to the input of the visualization node; the pipeline connection operation includes multi-process connection; Based on the flow chart, receiving an input job submission operation; Parsing the flowchart to obtain parsing results; the parsing results are used to characterize the dependency relationship between the data nodes, the model nodes and the visualization nodes and the sorting result of the data flow; Based on the analysis result, a job scheduling result is obtained; Based on the job scheduling result, the algorithm image deployment package corresponding to the algorithm model selection operation and the file data corresponding to the data selection operation are retrieved; the file data includes video data and image data; According to the algorithm image deployment package, performing algorithm analysis on the file data to obtain an analysis result; Visually displaying the analysis results; wherein the analysis results include images, videos, character strings, two-dimensional charts and / or three-dimensional charts; The parsing of the flowchart to obtain the parsing result includes: analyzing the algorithm analysis flowchart arranged by the user, extracting input data, algorithm model, configuration parameters and topological relationship therefrom, and sorting the nodes in the flowchart according to their dependency relationship and data flow direction to obtain the parsing result; The job scheduling result obtained based on the analysis result includes: according to the analysis result, including the dependency of the nodes, the algorithm analysis job is sent to the corresponding computing instance in a serial or parallel manner and runs, and at the same time, the job execution progress is monitored, the job is scheduled, and the job scheduling result is obtained.
2. The visualized space science experiment data analysis pipeline method according to claim 1 is characterized in that: The visualized space science experiment data analysis pipeline method also includes: based on web page embedding, setting the data node, the model node and the visualization node on the same user interface.
3. The visualized space science experiment data analysis pipeline method according to claim 1 or 2, characterized in that: The visualized space science experiment data analysis pipeline method also includes: The pipeline connection operation includes a multi-process pipeline node connection operation; the multi-process pipeline node connection operation includes a node connection method based on a directed acyclic graph structure, connecting the isolated data nodes, the model nodes and the visualization nodes into a complete algorithm analysis pipeline.
4. The visualized space science experiment data analysis pipeline method according to claim 1 or 2, characterized in that: The visualized space science experiment data analysis pipeline method also includes: Based on the flow chart, verify whether there is an isolated node; If the isolated node does not exist, verifying whether there is a missing input edge based on the flow chart; If the missing input edge does not exist, verify whether the data types match based on the flow chart; If the data types match, verify whether the required parameters are filled in based on the flow chart; If the required parameters are all filled in, verify whether there is a loop based on the flow chart; If the loop does not exist, then based on the flow chart, the job submission operation is received as input.
5. The visualized space science experiment data analysis pipeline method according to claim 1 or 2, characterized in that: The visualized space science experiment data analysis pipeline method also includes: Based on the flow chart, a plurality of jobs are obtained; The jobs having the first relationship are sent to the analysis algorithms on the computing instances one by one in a serial manner.
6. The visualized space science experiment data analysis pipeline method according to claim 1 or 2, characterized in that: The visualized space science experiment data analysis pipeline method also includes: Based on the flowchart, a plurality of jobs are obtained; For jobs with the second relationship, the analysis algorithm is sent to the computing instance in parallel.
7. The visualized space science experiment data analysis pipeline method according to claim 1 or 2, characterized in that: The visualized space science experiment data analysis pipeline method also includes: The algorithm image deployment package is used to cooperate in generating an execution container; the execution container includes a processor, memory, graphics processor resources and a mount file directory.
8. The visualized space science experiment data analysis pipeline method according to claim 7, characterized in that: The visualized space science experiment data analysis pipeline method also includes: Preset image templates; According to the algorithm configuration file, the image template is read, the code and data of the algorithm model are copied to the image, the dependent third-party library is automatically installed, and the algorithm image deployment package is built.
9. A visual space science experiment data analysis pipeline system, characterized in that: include: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.
10. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Visual data analysis process modeling method and system
CN114816374A
Industry-university-research integrated platform supporting visual dragging to carry out Internet of Vehicles data fusion analysis
CN116009844A